Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.lang.python > #197866 > unrolled thread

open: 'ascii', 'backslashreplace' not behaving as expected - why?

Started byVeek M <veekjunk@foobar.com>
First post2026-08-08 00:38 +0000
Last post2026-08-10 12:08 +1200
Articles 6 — 3 participants

Back to article view | Back to comp.lang.python


Contents

  open: 'ascii', 'backslashreplace' not behaving as expected - why? Veek M <veekjunk@foobar.com> - 2026-08-08 00:38 +0000
    Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-08 01:19 +0000
      Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Veek M <veekjunk@foobar.com> - 2026-08-08 04:22 +0000
        Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Veek M <veekjunk@foobar.com> - 2026-08-08 04:25 +0000
          Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Veek M <veekjunk@foobar.com> - 2026-08-08 04:33 +0000
            Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Greg Ewing <greg.ewing@canterbury.ac.nz> - 2026-08-10 12:08 +1200

#197866 — open: 'ascii', 'backslashreplace' not behaving as expected - why?

FromVeek M <veekjunk@foobar.com>
Date2026-08-08 00:38 +0000
Subjectopen: 'ascii', 'backslashreplace' not behaving as expected - why?
Message-ID<1155tpv$13on7$1@dont-email.me>
So in kate, i Ctrl-Shift-U and type ffff to enter a unicode codepoint  of 
0xffff. Then at the REPL prompt i do 
fh = open('/tmp/x', 'rt', -1, 'ascii', 'backslashreplace', None)

and i get 
fh.readline()
'\\xef\\xbf\\xbf\n'

Since I wrote two bytes 0xff and 0xff into Kate - why am i getting 0xef 
0xbf and 0xbf ?

[toc] | [next] | [standalone]


#197867

FromLawrence D’Oliveiro <ldo@nz.invalid>
Date2026-08-08 01:19 +0000
Message-ID<1156074$14bmn$1@dont-email.me>
In reply to#197866
On Sat, 8 Aug 2026 00:38:24 -0000 (UTC), Veek M wrote:

> So in kate, i Ctrl-Shift-U and type ffff to enter a unicode
> codepoint of 0xffff. Then at the REPL prompt i do
> fh = open('/tmp/x', 'rt', -1, 'ascii', 'backslashreplace', None)
>
> and i get
> fh.readline()
> '\\xef\\xbf\\xbf\n'
>
> Since I wrote two bytes 0xff and 0xff into Kate - why am i getting
> 0xef 0xbf and 0xbf ?

    Python 3.14.6 (main, Jun 10 2026, 18:54:31) [GCC 15.2.0] on linux
    Type "help", "copyright", "credits" or "license" for more information.
    >>> b'\xef\xbf\xbf\n'.decode()
    '\uffff\n'

[toc] | [prev] | [next] | [standalone]


#197868

FromVeek M <veekjunk@foobar.com>
Date2026-08-08 04:22 +0000
Message-ID<1156av3$15v7o$1@dont-email.me>
In reply to#197867
On Sat, 8 Aug 2026 01:19:32 -0000 (UTC), Lawrence D’Oliveiro wrote:

> b'\xef\xbf\xbf\n'.decode()

Could you explain how it works and what exactly is going on?

fh.readline() returns a unicode string with the funny chars (bytes 0xff 
0xff) encoded as \\xef \\xbf \\xbf - why is it \\? why not just use a 
single u'\xef\xbf\xbf' - why is he escaping the '\'.

Also - how exactly is he getting ef bf bf and not ff ff?

[toc] | [prev] | [next] | [standalone]


#197869

FromVeek M <veekjunk@foobar.com>
Date2026-08-08 04:25 +0000
Message-ID<1156b47$15v7o$2@dont-email.me>
In reply to#197868
On Sat, 8 Aug 2026 04:22:59 -0000 (UTC), Veek M wrote:

> On Sat, 8 Aug 2026 01:19:32 -0000 (UTC), Lawrence D’Oliveiro wrote:
> 
>> b'\xef\xbf\xbf\n'.decode()
> 
> Could you explain how it works and what exactly is going on?
> 
> fh.readline() returns a unicode string with the funny chars (bytes 0xff
> 0xff) encoded as \\xef \\xbf \\xbf - why is it \\? why not just use a
> single u'\xef\xbf\xbf' - why is he escaping the '\'.
> 
> Also - how exactly is he getting ef bf bf and not ff ff?

oh is 0xff 0xff when encoded to disk in utf-8 
(sys.getsystemdefaultencoding) 0xef 0xbf 0xbf?

[toc] | [prev] | [next] | [standalone]


#197870

FromVeek M <veekjunk@foobar.com>
Date2026-08-08 04:33 +0000
Message-ID<1156bib$15v7o$3@dont-email.me>
In reply to#197869
On Sat, 8 Aug 2026 04:25:43 -0000 (UTC), Veek M wrote:

> On Sat, 8 Aug 2026 04:22:59 -0000 (UTC), Veek M wrote:
> 
>> On Sat, 8 Aug 2026 01:19:32 -0000 (UTC), Lawrence D’Oliveiro wrote:
>> 
>>> b'\xef\xbf\xbf\n'.decode()
>> 
>> Could you explain how it works and what exactly is going on?
>> 
>> fh.readline() returns a unicode string with the funny chars (bytes 0xff
>> 0xff) encoded as \\xef \\xbf \\xbf - why is it \\? why not just use a
>> single u'\xef\xbf\xbf' - why is he escaping the '\'.
>> 
>> Also - how exactly is he getting ef bf bf and not ff ff?
> 
> oh is 0xff 0xff when encoded to disk in utf-8
> (sys.getsystemdefaultencoding) 0xef 0xbf 0xbf?

yes,
root@laptopveek:/tmp# od -x /tmp/x
0000000 bfef 0abf
0000004

it's the raw utf-8 encoded as bytes but since it is a unicode string why 
doesn't he save it as u'\xef\xbf\xbf' why does he escape the '\' and make 
it '\\x'

[toc] | [prev] | [next] | [standalone]


#197876

FromGreg Ewing <greg.ewing@canterbury.ac.nz>
Date2026-08-10 12:08 +1200
Message-ID<ndsj32Fq77oU1@mid.individual.net>
In reply to#197870
On 8/08/26 4:33 pm, Veek M wrote:
> it's the raw utf-8 encoded as bytes but since it is a unicode string why
> doesn't he save it as u'\xef\xbf\xbf' why does he escape the '\' and make
> it '\\x'

Because you decoded it as ascii with backslashreplace. It's replacing 
each byte that's outside the ascii range with four characters: a 
backslash, an 'x', and two hex digits. The backslashes are being doubled 
when you print the string and its repr() gets computed.

Since the file is actually utf-8 and not ascii, that's the appropriate 
way to decode it:

   fh = open('/tmp/x', 'rt', encoding = 'utf-8')

Then your ffff should come through as a single character in the string 
and print as '\uffff'.

-- 
Greg



[toc] | [prev] | [standalone]


Back to top | Article view | comp.lang.python


csiph-web