Path: csiph.com!usenet.pasdenom.info!dedibox.gegeweb.org!gegeweb.eu!nntpfeed.proxad.net!proxad.net!feeder1-2.proxad.net!usenet-fr.net!nerim.net!novso.com!newsfeed.xs4all.nl!newsfeed5.news.xs4all.nl!xs4all!post.news.xs4all.nl!not-for-mail Return-Path: X-Original-To: python-list@python.org Delivered-To: python-list@mail.python.org X-Spam-Status: OK 0.004 X-Spam-Evidence: '*H*': 0.99; '*S*': 0.00; 'bytes.': 0.07; 'escape': 0.07; 'users,': 0.07; 'python': 0.09; 'any.': 0.09; 'byte,': 0.09; 'creator': 0.09; 'lawrence': 0.09; 'subject:string': 0.09; 'to:addr:comp.lang.python': 0.09; 'cc:addr:python-list': 0.10; 'stored': 0.10; 'language': 0.14; '"pay"': 0.16; '3.3,': 0.16; '4-byte': 0.16; 'benefit.': 0.16; 'design:': 0.16; 'grounds': 0.16; 'hint:': 0.16; 'saying.': 0.16; 'subject:unicode': 0.16; 'string': 0.17; 'wrote:': 0.17; 'bytes': 0.17; 'detect': 0.17; 'mathematical': 0.17; 'tend': 0.17; 'unicode': 0.17; 'saying': 0.18; 'memory': 0.18; 'discussion': 0.20; 'preferred': 0.20; 'thanks.': 0.21; 'affected': 0.22; 'required.': 0.22; 'cc:2**0': 0.23; 'cc:no real name:2**0': 0.24; 'least': 0.25; 'cc:addr:python.org': 0.25; 'header:In-Reply-To:1': 0.25; 'header :User-Agent:1': 0.26; '---': 0.26; '(see': 0.27; 'plain': 0.27; 'correct': 0.28; 'rest': 0.28; 'overhead': 0.29; 'represented': 0.29; 'sleep': 0.29; 'character': 0.29; 'points': 0.29; "i'm": 0.29; 'related': 0.30; "skip:' 10": 0.30; 'code': 0.31; 'point': 0.31; 'could': 0.32; 'getting': 0.33; 'hopefully': 0.33; 'me?': 0.33; 'received:google.com': 0.34; 'wrong': 0.34; 'posting': 0.35; 'table': 0.35; 'received:209.85': 0.35; 'there': 0.35; 'but': 0.36; 'characters': 0.36; 'test': 0.36; 'correctly': 0.37; 'does': 0.37; 'being': 0.37; 'received:209': 0.37; 'far': 0.37; 'subject:: ': 0.38; 'mark': 0.38; 'performance': 0.39; 'google': 0.39; 'where': 0.40; 'think': 0.40; 'from:no real name:2**0': 0.60; 'subject:, ': 0.61; 'time,': 0.62; 'world': 0.63; 'subject:...': 0.63; 'email addr:gmail.com': 0.63; 'more': 0.63; 'making': 0.64; 'compliant': 0.65; 'saturday': 0.65; 'subject.': 0.65; 'soon': 0.70; 'benefit': 0.70; 'saving': 0.72; 'day': 0.73; 'frank': 0.75; '100%': 0.76; 'topic,': 0.78; '"das': 0.84; 'grosse': 0.84; 'overhead,': 0.84; 'points,': 0.84; 'samedi': 0.84; 'subject:, ...': 0.84; 'water.': 0.84 Newsgroups: comp.lang.python Date: Sat, 25 Aug 2012 08:47:52 -0700 (PDT) In-Reply-To: Complaints-To: groups-abuse@google.com Injection-Info: glegroupsg2000goo.googlegroups.com; posting-host=83.79.66.203; posting-account=ung4FAoAAAC46zhHJ0Nsnuox7M5gDvs_ References: <1874857c-68ef-4c1b-b15a-46ef47df9445@googlegroups.com> <1cb3f062-eb45-4b0c-977b-76afb099923c@googlegroups.com> User-Agent: G2/1.0 X-Google-Web-Client: true X-Google-IP: 83.79.66.203 MIME-Version: 1.0 Subject: Re: Flexible string representation, unicode, typography, ... From: wxjmfauth@gmail.com To: comp.lang.python@googlegroups.com Content-Type: text/plain; charset=ISO-8859-1 Content-Transfer-Encoding: quoted-printable Cc: python-list@python.org X-BeenThere: python-list@python.org X-Mailman-Version: 2.1.12 Precedence: list List-Id: General discussion list for the Python programming language List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Message-ID: Lines: 114 NNTP-Posting-Host: 2001:888:2000:d::a6 X-Trace: 1345909675 news.xs4all.nl 6938 [2001:888:2000:d::a6]:35161 X-Complaints-To: abuse@xs4all.nl Xref: csiph.com comp.lang.python:27876 Le samedi 25 ao=FBt 2012 11:46:34 UTC+2, Frank Millman a =E9crit=A0: > On 25/08/2012 10:58, Mark Lawrence wrote: >=20 > > On 25/08/2012 08:27, wxjmfauth@gmail.com wrote: >=20 > >> >=20 > >> Unicode design: a flat table of code points, where all code >=20 > >> points are "equals". >=20 > >> As soon as one attempts to escape from this rule, one has to >=20 > >> "pay" for it. >=20 > >> The creator of this machinery (flexible string representation) >=20 > >> can not even benefit from it in his native language (I think >=20 > >> I'm correctly informed). >=20 > >> >=20 > >> Hint: Google -> "Das grosse Eszett" >=20 > >> >=20 > >> jmf >=20 > >> >=20 > > >=20 > > It's Saturday morning, I'm stone cold sober, had a good sleep and I'm >=20 > > still baffled as to the point if any. Could someone please enlightem m= e? >=20 > > >=20 >=20 >=20 > Here's what I think he is saying. I am posting this to test the water. I= =20 >=20 > am also confused, and if I have got it wrong hopefully someone will=20 >=20 > correct me. >=20 >=20 >=20 > In python 3.3, unicode strings are now stored as follows - >=20 > if all characters can be represented by 1 byte, the entire string is= =20 >=20 > composed of 1-byte characters >=20 > else if all characters can be represented by 1 or 2 bytea, the entire= =20 >=20 > string is composed of 2-byte characters >=20 > else the entire string is composed of 4-byte characters >=20 >=20 >=20 > There is an overhead in making this choice, to detect the lowest number= =20 >=20 > of bytes required. >=20 >=20 >=20 > jmfauth believes that this only benefits 'english-speaking' users, as=20 >=20 > the rest of the world will tend to have strings where at least one=20 >=20 > character requires 2 or 4 bytes. So they incur the overhead, without=20 >=20 > getting any benefit. >=20 >=20 >=20 > Therefore, I think he is saying that he would have preferred that python= =20 >=20 > standardise on 4-byte characters, on the grounds that the saving in=20 >=20 > memory does not justify the performance overhead. >=20 >=20 >=20 > Frank Millman Very well explained. Thanks. More precisely, affected are not only the 'english-speaking' users, but all the users who are using not latin-1 characters. (See the title of this topic, ... typography). Being at the same time, latin-1 and unicode compliant is a plain absurdity in the mathematical sense. --- For those you do not know, the go language has introduced the rune type. As far as I know, nobody is complaining, I have not even seen a discussion related to this subject. 100% Unicode compliant from the day 0. Congratulations. jmf