Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > muc.lists.netbsd.tech.userlevel > #11863 > unrolled thread

Re: CVS commit: src/external/bsd/libarchive/dist/libarchive

Started byTaylor R Campbell <riastradh@NetBSD.org>
First post2026-08-24 12:42 +0000
Last post2026-08-24 14:38 +0000
Articles 14 — 7 participants

Back to article view | Back to muc.lists.netbsd.tech.userlevel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Taylor R Campbell <riastradh@NetBSD.org> - 2026-08-24 12:42 +0000
    Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Thomas Klausner <wiz@netbsd.org> - 2026-08-24 14:55 +0200
      Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Martin Husemann <martin@duskware.de> - 2026-08-24 15:15 +0200
        Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Greg Troxel <gdt@lexort.com> - 2026-08-24 09:23 -0400
          Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Martin Husemann <martin@duskware.de> - 2026-08-24 15:57 +0200
            Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Mouse <mouse@Rodents-Montreal.ORG> - 2026-08-24 11:35 -0400
              Re: CVS commit: src/external/bsd/libarchive/dist/libarchive tlaronde@kergis.com - 2026-08-24 17:53 +0200
        Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Crystal Kolipe <kolipe.c@exoticsilicon.com> - 2026-08-24 13:46 +0000
      Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Mouse <mouse@Rodents-Montreal.ORG> - 2026-08-24 09:22 -0400
      Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Taylor R Campbell <riastradh@NetBSD.org> - 2026-08-24 14:01 +0000
        Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Thomas Klausner <wiz@netbsd.org> - 2026-08-24 16:23 +0200
          Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Taylor R Campbell <riastradh@NetBSD.org> - 2026-08-24 14:52 +0000
            Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Thomas Klausner <wiz@netbsd.org> - 2026-08-24 20:23 +0200
        Re: CVS commit: src/external/bsd/libarchive/dist/libarchive Taylor R Campbell <riastradh@NetBSD.org> - 2026-08-24 14:38 +0000

#11863 — Re: CVS commit: src/external/bsd/libarchive/dist/libarchive

FromTaylor R Campbell <riastradh@NetBSD.org>
Date2026-08-24 12:42 +0000
SubjectRe: CVS commit: src/external/bsd/libarchive/dist/libarchive
Message-ID<20260824124235.D8CE284E54@mail.netbsd.org>
> Date: Mon, 24 Aug 2026 11:36:32 +0200
> From: Thomas Klausner <wiz@netbsd.org>
> 
> On Fri, Aug 21, 2026 at 10:12:25PM +0100, Taylor R Campbell wrote:
> > libarchive: Patch iconv() use to handle POSIX semantics.
> 
> A userland from August 19 was fine, but one from today is not, so I
> suspect this commit to cause an extraction error for
> pkgsrc/devel/py-meson_python.
> 
> (no LANG, LC_* set):
> 
> # tar xvzf meson_python-0.20.0.tar.gz
> tar: Pathname can't be converted from UTF-8 to current locale
> tar: Error exit delayed from previous errors
> 
> The relevant file name is
> 
> meson_python-0.20.0/tests/packages/encoding/\343\203\206\343\202\271\343\203\210.py
> 
> or
> 
> meson_python-0.20.0/tests/packages/encoding/ใƒ†ใ‚นใƒˆ.py
> 
> in a UTF-8 locale.
> 
> Is this an intended consequence of the change?

How was the archive extracted before?  Did it turn into a file called
`???.py' (i.e., with literal question marks)?

I suspect that the code was only `working' by accident before, i.e.,
failing to report failure and blithely barging ahead with output that
the libarchive authors thought of as corrupted, having had replacement
characters substituted where the input can't be represented.

On a Debian system (where bsdtar is presumably using GNU iconv), if I
try this, it fails in exactly the same way with LANG=C and works with
LANG=C.UTF-8.  archivers/bsdtar is built without any iconv at all, so
it rejects the archive altogether no matter what your locale is!  And
as far as I know, there's no way to ask bsdtar/libarchive to just
treat each path as a sequence of bytes without interpretation.

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [next] | [standalone]


#11864

FromThomas Klausner <wiz@netbsd.org>
Date2026-08-24 14:55 +0200
Message-ID<aow96bwmZq_NnLOR@exadelic.gatalith.at>
In reply to#11863
On Mon, Aug 24, 2026 at 12:42:33PM +0100, Taylor R Campbell wrote:
> How was the archive extracted before?  Did it turn into a file called
> `???.py' (i.e., with literal question marks)?

Yes:
meson_python-0.20.0/tests/packages/encoding/???.py

> I suspect that the code was only `working' by accident before, i.e.,
> failing to report failure and blithely barging ahead with output that
> the libarchive authors thought of as corrupted, having had replacement
> characters substituted where the input can't be represented.
> 
> On a Debian system (where bsdtar is presumably using GNU iconv), if I
> try this, it fails in exactly the same way with LANG=C and works with
> LANG=C.UTF-8.  archivers/bsdtar is built without any iconv at all, so
> it rejects the archive altogether no matter what your locale is!  And
> as far as I know, there's no way to ask bsdtar/libarchive to just
> treat each path as a sequence of bytes without interpretation.

I guess I'll switch the affected packages to GNU tar then, which
extracts the exact binary sequence and reports no error:

meson_python-0.20.0/tests/packages/encoding/<E3><83><86><E3><82><B9><E3><83><88>.py

# ls -al
total 40
drwxr-xr-x   2 root  wheel   144 Aug 24 12:51 .
drwxr-xr-x  41 root  wheel  1872 Aug 24 12:51 ..
-rw-rw-r--   1 root  wheel   210 Jun 10 19:15 meson.build
-rw-rw-r--   1 root  wheel   162 Jun 10 19:15 pyproject.toml
-rw-rw-r--   1 root  wheel    92 Jun 10 19:15 ?????????.py
# ls | hexdump -C
00000000  6d 65 73 6f 6e 2e 62 75  69 6c 64 0a 70 79 70 72  |meson.build.pypr|
00000010  6f 6a 65 63 74 2e 74 6f  6d 6c 0a e3 83 86 e3 82  |oject.toml......|
00000020  b9 e3 83 88 2e 70 79 0a                           |.....py.|
00000028
#

Personally I'm not quite sure if I want a tar tool to convert between
encodings. Is the base encoding included in the tar file? If not, how
does the unpacking tar know how it should interpret the byte sequence?
It could be EUC-JP for all we know.
 Thomas

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11865

FromMartin Husemann <martin@duskware.de>
Date2026-08-24 15:15 +0200
Message-ID<aoxD33xvDytFYOY7@big-apple.aprisoft.de>
In reply to#11864
On Mon, Aug 24, 2026 at 02:55:06PM +0200, Thomas Klausner wrote:
> Personally I'm not quite sure if I want a tar tool to convert between
> encodings. Is the base encoding included in the tar file?

Officially (AFAICT) ustar format only supports ASCII, everything else
is a grey area. Best option is to use pax format, where the filename
is guaranteed to be UTF8 encoded.

Martin

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11867

FromGreg Troxel <gdt@lexort.com>
Date2026-08-24 09:23 -0400
Message-ID<rmizeyb1wk3.fsf@s1.lexort.com>
In reply to#11865
Martin Husemann <martin@duskware.de> writes:

> On Mon, Aug 24, 2026 at 02:55:06PM +0200, Thomas Klausner wrote:
>> Personally I'm not quite sure if I want a tar tool to convert between
>> encodings. Is the base encoding included in the tar file?
>
> Officially (AFAICT) ustar format only supports ASCII, everything else
> is a grey area.

I can see that, tar being from long ago.  Within UB, though, it seems
really obvious to just take the bytes in the filename in the archive and
use that with syscalls and not do any character set stuff, unless
explicitly requested.

What kind of conversions do you think are happening with libarchive, and
why do you think those conversions are a good idea?  (Genuine question;
I am fuzzy about beyond-ascii in filenames in general)



> Best option is to use pax format, where the filename is guaranteed to
> be UTF8 encoded.


That's not a question on the table - this is about dealing with tar
files others create.

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11868

FromMartin Husemann <martin@duskware.de>
Date2026-08-24 15:57 +0200
Message-ID<aoxNrCSZfMTxoVBg@big-apple.aprisoft.de>
In reply to#11867
On Mon, Aug 24, 2026 at 09:23:56AM -0400, Greg Troxel wrote:
> What kind of conversions do you think are happening with libarchive, and
> why do you think those conversions are a good idea?  (Genuine question;
> I am fuzzy about beyond-ascii in filenames in general)

The way (I think) bsdtar and libarchive are structured, the filename
conversion (depending on locale) happens in a totaly different part
from the storage/decoding.

Many other archive formats (e.g. zip, rar and pax) define UTF8 as internal
file name encoding, so assuming that for the undefined parts of ancient
things like ustar is a (maybe unpopular but) reasonable choice.

While the unix-ish interpretation that mouse potinted at is great
when using the archive in the same context it was created, it
forces other extractors to guess the encoding, which is pretty bad.

I am one of the few remaining users (I guess) of non-ASCII and non-UTF8
locales. And one day I will go (of course not verbatim; have a script
go) through all files stored on my NASes and rename them from
ISO-8859-1 to UTF8.

Untill then I will have to live with e.g. "unrar" failiing to extract
archives with filenames containing funny printing marks or air quotes
(or I guess emojies). Workaround: temporarily set LC_CTYPE to UTF8
and manually rename afterwards.


Martin

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11877

FromMouse <mouse@Rodents-Montreal.ORG>
Date2026-08-24 11:35 -0400
Message-ID<202608241535.LAA22520@Stone.Rodents-Montreal.ORG>
In reply to#11868
> While the unix-ish interpretation that mouse potinted at is great
> when using the archive in the same context it was created, it forces
> other extractors to guess the encoding, which is pretty bad.

Unix pathnames have historically been octet strings, uninterpreted
except for slash and NUL.  I would certainly *hope* that they still are
at the syscall level; anything else is IMO a "renders it unusable" bug.
Treating filenames as character strings anywhere but user presentation
leads directly to a large morass of trouble.

"[O]ther extractors" shouldn't be doing anything with encodings; they
should be taking the octet string given and using it as the pathname to
the relevant syscall(s).  (Unless specifically told to remap, but I'm
not sure I'd consider even that much sane.)

If there are extractors running on systems where filenames are
character strings instead of octet strings, they will need impedance
matching, but that's always been true.

> I am one of the few remaining users (I guess) of non-ASCII and
> non-UTF8 locales.

I don't know about "few", but I'm another.  I use 8859-1 routinely,
8859-14 a bit less routinely but still relatively commonly.  In a few
cases I even have both in the same directory, I think.  Maybe a few
others; not sure.  My reaction to variable-size characters is very much
"delenda est".

I'm not going to stop doing so, either.  Anything that can't handle the
resulting octet strings will, on my systems, get fixed, replaced, or at
worst used only for use cases where its brokenness is tolerable.

/~\ The ASCII				  Mouse
\ / Ribbon Campaign
 X  Against HTML		mouse@rodents-montreal.org
/ \ Email!	     7D C8 61 52 5D E7 2D 39  4E F1 31 3E E8 B3 27 4B

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11879

Fromtlaronde@kergis.com
Date2026-08-24 17:53 +0200
Message-ID<aoxo_xqeGi7bmPyZ@kergis.com>
In reply to#11877
On Mon, Aug 24, 2026 at 11:35:02AM -0400, Mouse wrote:
> > While the unix-ish interpretation that mouse potinted at is great
> > when using the archive in the same context it was created, it forces
> > other extractors to guess the encoding, which is pretty bad.
> 
> Unix pathnames have historically been octet strings, uninterpreted
> except for slash and NUL.  I would certainly *hope* that they still are
> at the syscall level; anything else is IMO a "renders it unusable" bug.
> Treating filenames as character strings anywhere but user presentation
> leads directly to a large morass of trouble.
> 

Plus, IMHO, that one legitimately expects an archiver to be "lossless"
i.e. it extracts exactly what it has instructed to archive. If there is
any modification (except if explicitely requested so definitively not
by default) or if there is any loss of information, the utility should
not be called an archiver.
-- 
        Thierry Laronde <tlaronde +AT+ kergis +dot+ com>
                     http://www.kergis.com/
                    http://kertex.kergis.com/
Key fingerprint = 0FF7 E906 FBAF FE95 FD89  250D 52B1 AE95 6006 F40C

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11871

FromCrystal Kolipe <kolipe.c@exoticsilicon.com>
Date2026-08-24 13:46 +0000
Message-ID<aoxLMa8wDmxow2LS@exoticsilicon.com>
In reply to#11865
On Mon, Aug 24, 2026 at 03:15:11PM +0200, Martin Husemann wrote:
> On Mon, Aug 24, 2026 at 02:55:06PM +0200, Thomas Klausner wrote:
> > Personally I'm not quite sure if I want a tar tool to convert between
> > encodings. Is the base encoding included in the tar file?
> 
> Officially (AFAICT) ustar format only supports ASCII, everything else
> is a grey area. Best option is to use pax format, where the filename
> is guaranteed to be UTF8 encoded.

That's the default - pax extended headers are assumed to be utf-8 encoded
_unless_ hdrcharset is set to binary, in which case the data is just an octet
stream.

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11866

FromMouse <mouse@Rodents-Montreal.ORG>
Date2026-08-24 09:22 -0400
Message-ID<202608241322.JAA24740@Stone.Rodents-Montreal.ORG>
In reply to#11864
>> On a Debian system (where bsdtar is presumably using GNU iconv), if
>> I try this, it fails in exactly the same way with LANG=C and works
>> with LANG=C.UTF-8.

!

>> archivers/bsdtar is built without any iconv at
>> all, so it rejects the archive altogether no matter what your locale
>> is!

!!  So, it either (a) refuses to tar up certain pathnames or (b)
refuses to extract the resulting archive or (c) generates an archive
containing names different from the original names?  I'd call any of
those catastrophically broken.  Extremely non-Unixy.

>> And as far as I know, there's no way to ask bsdtar/libarchive to
>> just treat each path as a sequence of bytes without interpretation.

I find it astonishing that _any_ tar implementation would treat
pathnames as anything but opaque octet strings.  If I ran into this I'd
be filing "renders it unusable" bug reports for the implementation in
question (well, if I could find a usable bug-reporting address for it).

> Personally I'm not quite sure if I want a tar tool to convert between
> encodings.

I'm quite sure I don't.  If I ever personally run into a tar that
treats pathnames as anything but opaque (except for slash and NUL)
octet strings, I will consider it outright broken.

Makes me glad I have my own tar implementation.

Actually, it reminds me of the sqlite people.  I asked them about using
non-UTF-8 filenames for databases (the doc says that using non-UTF-8
filenames provokes undefined behaviour) and the response - and this is
a direct quote - was "What's the problem?  Are you wanting to name a
database with some octet sequence that is not valid UTF-8?  Why would
you want to do that?".  Absolutely jaw-dropping.

/~\ The ASCII				  Mouse
\ / Ribbon Campaign
 X  Against HTML		mouse@rodents-montreal.org
/ \ Email!	     7D C8 61 52 5D E7 2D 39  4E F1 31 3E E8 B3 27 4B

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11869

FromTaylor R Campbell <riastradh@NetBSD.org>
Date2026-08-24 14:01 +0000
Message-ID<20260824140118.D9E6384E5A@mail.netbsd.org>
In reply to#11864
> Date: Mon, 24 Aug 2026 14:55:06 +0200
> From: Thomas Klausner <wiz@netbsd.org>
> 
> Personally I'm not quite sure if I want a tar tool to convert between
> encodings. Is the base encoding included in the tar file? If not, how
> does the unpacking tar know how it should interpret the byte sequence?
> It could be EUC-JP for all we know.

The archive in question is in POSIX pax interchange format, which has
_two paths_ for the file in question (written here in vis(3)
notation):

- (pax extended header `path' attribute, always UTF-8)

  meson_python-0.19.0/tests/packages/encoding/\xe3\x83\x86\xe3\x82\xb9\xe3\x83\x88.py

  Reference:
  https://pubs.opengroup.org/onlinepubs/9799919799/utilities/pax.html#tag_20_94_13_03

- (ustar record path field, exactly 100 octets, unspecified encoding)

  meson_python-0.19.0/tests/packages/encoding/???.py

  Reference:
  https://pubs.opengroup.org/onlinepubs/9799919799/utilities/pax.html#tag_20_94_13_06

POSIX prescribes of pax(1) that, for the `path' extended header
attribute:

	The pax utility shall translate the pathname of the file from
	the encoding in the header to the character set appropriate
	for the local file system.

(Our pax(1) does not understand extended headers, which is a bug we
should have fixed a long time ago but most effort has gone into
bsdtar/libarchive instead.)

This isn't authoritative for tar(1), of course.

What happened before the change I made, though, was worse than either:

(a) taking the pax extended header `path' attribute verbatim (which is
    what I would guess, but have not verified, that GNU tar does),
_or_
(b) taking the ustar path verbatim (which is what our pax(1) does).

Instead, the unpatched bsdtar would:

1. Take the pax extended header `path' attribute.

2. iconv -f UTF-8 -t 646 (i.e., convert from UTF-8 to US-ASCII, which
   is the character encoding in the C locale)

3. Use the result _despite_ iconv's report that it had to substitute
   replacement characters, in this case `?' (which happens to coincide
   with path in the ustar record, but that's just a coincidence, so to
   speak).

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11873

FromThomas Klausner <wiz@netbsd.org>
Date2026-08-24 16:23 +0200
Message-ID<aoxTQE0zy28TcZo1@exadelic.gatalith.at>
In reply to#11869
So what's the best way forward?

pkgsrc sets
ALL_ENV+=  LANG=C
by default, and this now breaks unpacking on NetBSD-current for at least

devel/py-meson_python
textproc/py-sphinx
lang/go124
devel/qt6-qttools

I see

a) do something about tar(1) in -current
b) do something about LANG=C in pkgsrc/mk
c) extract them using gtar, which doesn't care

 Thomas

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11876

FromTaylor R Campbell <riastradh@NetBSD.org>
Date2026-08-24 14:52 +0000
Message-ID<20260824145212.84E0F84D81@mail.netbsd.org>
In reply to#11873
> Date: Mon, 24 Aug 2026 16:23:38 +0200
> From: Thomas Klausner <wiz@netbsd.org>
> 
> So what's the best way forward?
> 
> pkgsrc sets
> ALL_ENV+=  LANG=C
> by default, and this now breaks unpacking on NetBSD-current for at least
> 
> devel/py-meson_python
> textproc/py-sphinx
> lang/go124
> devel/qt6-qttools
> 
> I see
> 
> a) do something about tar(1) in -current

This is likely to be a nonstarter unless there is an explicit option
instructing tar(1) to take the pax extended header `path' verbatim or
something.  We shouldn't diverge from upstream libarchive bsdtar
semantics unless we have an extremely compelling reason to do so.

(The patch I applied was to _reduce_ divergence from what upstream's
intent appeared to be, and I submitted it upstream for review in order
to make the semantics agree on more platfoms as verified by the tests:
<https://github.com/libarchive/libarchive/issues/3413>,
<https://github.com/libarchive/libarchive/pull/3416>.)

> b) do something about LANG=C in pkgsrc/mk

I suggest we set LC_CTYPE to a UTF-8 locale in EXTRACT_ENV for these
packages, if not also in ALL_ENV and/or likewise in all packages.

Previously discussed on tech-pkg@ in 2022, leading us to set
EXTRACT_ENV+= LC_CTYPE=en_US.UTF-8 in all packages on NetBSD<=8:

https://mail-index.NetBSD.org/tech-pkg/2022/05/05/msg026252.html
https://mail-index.netbsd.org/pkgsrc-changes/2022/05/21/msg254803.html

For NetBSD>=10, we can use C.UTF-8 instead of en_US.UTF-8 (though if
it's only for LC_CTYPE, using en_US.UTF-8 won't hurt, distressingly
angloyankocentric though that may be).

> c) extract them using gtar, which doesn't care

This is likely not reliable as noted at:
https://mail-index.NetBSD.org/tech-userlevel/2026/08/24/msg015063.html

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [next] | [standalone]


#11880

FromThomas Klausner <wiz@netbsd.org>
Date2026-08-24 20:23 +0200
Message-ID<aoyL-f-r75Zt7Q2v@exadelic.gatalith.at>
In reply to#11876

[Multipart message โ€” attachments visible in raw view] — view raw

On Mon, Aug 24, 2026 at 02:52:11PM +0100, Taylor R Campbell wrote:
> > b) do something about LANG=C in pkgsrc/mk
> 
> I suggest we set LC_CTYPE to a UTF-8 locale in EXTRACT_ENV for these
> packages, if not also in ALL_ENV and/or likewise in all packages.
> 
> Previously discussed on tech-pkg@ in 2022, leading us to set
> EXTRACT_ENV+= LC_CTYPE=en_US.UTF-8 in all packages on NetBSD<=8:
> 
> https://mail-index.NetBSD.org/tech-pkg/2022/05/05/msg026252.html
> https://mail-index.netbsd.org/pkgsrc-changes/2022/05/21/msg254803.html
> 
> For NetBSD>=10, we can use C.UTF-8 instead of en_US.UTF-8 (though if
> it's only for LC_CTYPE, using en_US.UTF-8 won't hurt, distressingly
> angloyankocentric though that may be).

I'm testing the attached diff - so far it looks good, I'll commit it
once my bulk build is done.
 Thomas

[toc] | [prev] | [next] | [standalone]


#11874

FromTaylor R Campbell <riastradh@NetBSD.org>
Date2026-08-24 14:38 +0000
Message-ID<20260824143820.28F0E84E59@mail.netbsd.org>
In reply to#11869
> Date: Mon, 24 Aug 2026 14:01:17 +0000
> From: Taylor R Campbell <riastradh@NetBSD.org>
> 
> > Date: Mon, 24 Aug 2026 14:55:06 +0200
> > From: Thomas Klausner <wiz@netbsd.org>
> > 
> > Personally I'm not quite sure if I want a tar tool to convert between
> > encodings. Is the base encoding included in the tar file? If not, how
> > does the unpacking tar know how it should interpret the byte sequence?
> > It could be EUC-JP for all we know.
> 
> The archive in question is in POSIX pax interchange format, which has
> _two paths_ for the file in question (written here in vis(3)
> notation):
> 
> - (pax extended header `path' attribute, always UTF-8)

Correction: not `always UTF-8' -- rather, UTF-8 by default, unless
overridden by hdrcharset.  But there is no hdrcharset in this input,
so we know the input is supposed to be UTF-8, not EUC-JP.

Reference:
https://pubs.opengroup.org/onlinepubs/9799919799/utilities/pax.html#tag_20_94_13_03

> (a) taking the pax extended header `path' attribute verbatim (which is
>     what I would guess, but have not verified, that GNU tar does),

I reviewed the GNU tar code, and I believe that:

1. this is, in fact, what GNU tar does, but

2. the GNU tar authors may consider it a bug and intend to fix or at
   least change it in some way later.

So I'm not sure the GNU tar behaviour is any more reliable than the
earlier bsdtar behaviour.


This is the subroutine in GNU tar that decodes the pax extended header
`path' attribute (among other things), from src/xheader.c in v1.35
(commit e545d446dfe6564265cdf4186641ee76f4acc7fa):

   1024 static void
   1025 decode_string (char **string, char const *arg)
   1026 {
   1027   if (*string)
   1028     {
   1029       free (*string);
   1030       *string = NULL;
   1031     }
   1032   if (!utf8_convert (false, arg, string))
   1033     {
   1034       /* FIXME: report error and act accordingly to --pax invalid=UTF-8 */
   1035       assign_string (string, arg);
   1036     }
   1037 }

https://cgit.git.savannah.gnu.org/cgit/tar.git/tree/src/xheader.c?h=v1.35&id=e545d446dfe6564265cdf4186641ee76f4acc7fa

(assign_string under the hood boils down to strdup.)

--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-admin@muc.de

[toc] | [prev] | [standalone]


Back to top | Article view | muc.lists.netbsd.tech.userlevel


csiph-web