Path: csiph.com!fu-berlin.de!uni-berlin.de!individual.net!not-for-mail From: Greg Ewing Newsgroups: comp.lang.python Subject: Re: open: 'ascii', 'backslashreplace' not behaving as expected - why? Date: Mon, 10 Aug 2026 12:08:00 +1200 Lines: 24 Message-ID: References: <1155tpv$13on7$1@dont-email.me> <1156074$14bmn$1@dont-email.me> <1156av3$15v7o$1@dont-email.me> <1156b47$15v7o$2@dont-email.me> <1156bib$15v7o$3@dont-email.me> Mime-Version: 1.0 Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit X-Trace: individual.net Aa6v8VrhAjC8w0qqaKpW4g+yMbB5YhTQoawmitlOgl3eY5084E Cancel-Lock: sha1:/YLGh3/camLmBtAMcBDRQj1iIPA= sha256:wM9Mc83loJkB1HzrgU2jV25RVt7Wa5XsMQnde84frKE= User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.13; rv:91.0) Gecko/20100101 Thunderbird/91.3.2 Content-Language: en-US In-Reply-To: <1156bib$15v7o$3@dont-email.me> Xref: csiph.com comp.lang.python:197876 On 8/08/26 4:33 pm, Veek M wrote: > it's the raw utf-8 encoded as bytes but since it is a unicode string why > doesn't he save it as u'\xef\xbf\xbf' why does he escape the '\' and make > it '\\x' Because you decoded it as ascii with backslashreplace. It's replacing each byte that's outside the ascii range with four characters: a backslash, an 'x', and two hex digits. The backslashes are being doubled when you print the string and its repr() gets computed. Since the file is actually utf-8 and not ascii, that's the appropriate way to decode it: fh = open('/tmp/x', 'rt', encoding = 'utf-8') Then your ffff should come through as a single character in the string and print as '\uffff'. -- Greg