Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.lang.python > #197866 > unrolled thread
| Started by | Veek M <veekjunk@foobar.com> |
|---|---|
| First post | 2026-08-08 00:38 +0000 |
| Last post | 2026-08-10 12:08 +1200 |
| Articles | 6 — 3 participants |
Back to article view | Back to comp.lang.python
open: 'ascii', 'backslashreplace' not behaving as expected - why? Veek M <veekjunk@foobar.com> - 2026-08-08 00:38 +0000
Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Lawrence D’Oliveiro <ldo@nz.invalid> - 2026-08-08 01:19 +0000
Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Veek M <veekjunk@foobar.com> - 2026-08-08 04:22 +0000
Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Veek M <veekjunk@foobar.com> - 2026-08-08 04:25 +0000
Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Veek M <veekjunk@foobar.com> - 2026-08-08 04:33 +0000
Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Greg Ewing <greg.ewing@canterbury.ac.nz> - 2026-08-10 12:08 +1200
| From | Veek M <veekjunk@foobar.com> |
|---|---|
| Date | 2026-08-08 00:38 +0000 |
| Subject | open: 'ascii', 'backslashreplace' not behaving as expected - why? |
| Message-ID | <1155tpv$13on7$1@dont-email.me> |
So in kate, i Ctrl-Shift-U and type ffff to enter a unicode codepoint of
0xffff. Then at the REPL prompt i do
fh = open('/tmp/x', 'rt', -1, 'ascii', 'backslashreplace', None)
and i get
fh.readline()
'\\xef\\xbf\\xbf\n'
Since I wrote two bytes 0xff and 0xff into Kate - why am i getting 0xef
0xbf and 0xbf ?
[toc] | [next] | [standalone]
| From | Lawrence D’Oliveiro <ldo@nz.invalid> |
|---|---|
| Date | 2026-08-08 01:19 +0000 |
| Message-ID | <1156074$14bmn$1@dont-email.me> |
| In reply to | #197866 |
On Sat, 8 Aug 2026 00:38:24 -0000 (UTC), Veek M wrote:
> So in kate, i Ctrl-Shift-U and type ffff to enter a unicode
> codepoint of 0xffff. Then at the REPL prompt i do
> fh = open('/tmp/x', 'rt', -1, 'ascii', 'backslashreplace', None)
>
> and i get
> fh.readline()
> '\\xef\\xbf\\xbf\n'
>
> Since I wrote two bytes 0xff and 0xff into Kate - why am i getting
> 0xef 0xbf and 0xbf ?
Python 3.14.6 (main, Jun 10 2026, 18:54:31) [GCC 15.2.0] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> b'\xef\xbf\xbf\n'.decode()
'\uffff\n'
[toc] | [prev] | [next] | [standalone]
| From | Veek M <veekjunk@foobar.com> |
|---|---|
| Date | 2026-08-08 04:22 +0000 |
| Message-ID | <1156av3$15v7o$1@dont-email.me> |
| In reply to | #197867 |
On Sat, 8 Aug 2026 01:19:32 -0000 (UTC), Lawrence D’Oliveiro wrote: > b'\xef\xbf\xbf\n'.decode() Could you explain how it works and what exactly is going on? fh.readline() returns a unicode string with the funny chars (bytes 0xff 0xff) encoded as \\xef \\xbf \\xbf - why is it \\? why not just use a single u'\xef\xbf\xbf' - why is he escaping the '\'. Also - how exactly is he getting ef bf bf and not ff ff?
[toc] | [prev] | [next] | [standalone]
| From | Veek M <veekjunk@foobar.com> |
|---|---|
| Date | 2026-08-08 04:25 +0000 |
| Message-ID | <1156b47$15v7o$2@dont-email.me> |
| In reply to | #197868 |
On Sat, 8 Aug 2026 04:22:59 -0000 (UTC), Veek M wrote: > On Sat, 8 Aug 2026 01:19:32 -0000 (UTC), Lawrence D’Oliveiro wrote: > >> b'\xef\xbf\xbf\n'.decode() > > Could you explain how it works and what exactly is going on? > > fh.readline() returns a unicode string with the funny chars (bytes 0xff > 0xff) encoded as \\xef \\xbf \\xbf - why is it \\? why not just use a > single u'\xef\xbf\xbf' - why is he escaping the '\'. > > Also - how exactly is he getting ef bf bf and not ff ff? oh is 0xff 0xff when encoded to disk in utf-8 (sys.getsystemdefaultencoding) 0xef 0xbf 0xbf?
[toc] | [prev] | [next] | [standalone]
| From | Veek M <veekjunk@foobar.com> |
|---|---|
| Date | 2026-08-08 04:33 +0000 |
| Message-ID | <1156bib$15v7o$3@dont-email.me> |
| In reply to | #197869 |
On Sat, 8 Aug 2026 04:25:43 -0000 (UTC), Veek M wrote: > On Sat, 8 Aug 2026 04:22:59 -0000 (UTC), Veek M wrote: > >> On Sat, 8 Aug 2026 01:19:32 -0000 (UTC), Lawrence D’Oliveiro wrote: >> >>> b'\xef\xbf\xbf\n'.decode() >> >> Could you explain how it works and what exactly is going on? >> >> fh.readline() returns a unicode string with the funny chars (bytes 0xff >> 0xff) encoded as \\xef \\xbf \\xbf - why is it \\? why not just use a >> single u'\xef\xbf\xbf' - why is he escaping the '\'. >> >> Also - how exactly is he getting ef bf bf and not ff ff? > > oh is 0xff 0xff when encoded to disk in utf-8 > (sys.getsystemdefaultencoding) 0xef 0xbf 0xbf? yes, root@laptopveek:/tmp# od -x /tmp/x 0000000 bfef 0abf 0000004 it's the raw utf-8 encoded as bytes but since it is a unicode string why doesn't he save it as u'\xef\xbf\xbf' why does he escape the '\' and make it '\\x'
[toc] | [prev] | [next] | [standalone]
| From | Greg Ewing <greg.ewing@canterbury.ac.nz> |
|---|---|
| Date | 2026-08-10 12:08 +1200 |
| Message-ID | <ndsj32Fq77oU1@mid.individual.net> |
| In reply to | #197870 |
On 8/08/26 4:33 pm, Veek M wrote:
> it's the raw utf-8 encoded as bytes but since it is a unicode string why
> doesn't he save it as u'\xef\xbf\xbf' why does he escape the '\' and make
> it '\\x'
Because you decoded it as ascii with backslashreplace. It's replacing
each byte that's outside the ascii range with four characters: a
backslash, an 'x', and two hex digits. The backslashes are being doubled
when you print the string and its repr() gets computed.
Since the file is actually utf-8 and not ascii, that's the appropriate
way to decode it:
fh = open('/tmp/x', 'rt', encoding = 'utf-8')
Then your ffff should come through as a single character in the string
and print as '\uffff'.
--
Greg
[toc] | [prev] | [standalone]
Back to top | Article view | comp.lang.python
csiph-web