Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #85537 > unrolled thread

Bug#1098226: linux-image-6.12.13-amd64: newfstatat syscall ENOENT failure in multithreaded program

Started byVincent Lefevre <vincent@vinc17.net>
First post2025-02-18 14:00 +0100
Last post2025-02-18 17:10 +0100
Articles 4 — 1 participant

Back to article view | Back to linux.debian.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Bug#1098226: linux-image-6.12.13-amd64: newfstatat syscall ENOENT failure in multithreaded program Vincent Lefevre <vincent@vinc17.net> - 2025-02-18 14:00 +0100
    Bug#1098226: linux-image-6.12.13-amd64: newfstatat syscall ENOENT failure in multithreaded program Vincent Lefevre <vincent@vinc17.net> - 2025-02-18 14:10 +0100
      Bug#1098226: linux-image-6.12.13-amd64: newfstatat syscall ENOENT failure in multithreaded program Vincent Lefevre <vincent@vinc17.net> - 2025-02-18 14:40 +0100
        Bug#1098226: perl: thread creation while a directory handle is open does a fchdir, affecting other threads Vincent Lefevre <vincent@vinc17.net> - 2025-02-18 17:10 +0100

#85537 — Bug#1098226: linux-image-6.12.13-amd64: newfstatat syscall ENOENT failure in multithreaded program

FromVincent Lefevre <vincent@vinc17.net>
Date2025-02-18 14:00 +0100
SubjectBug#1098226: linux-image-6.12.13-amd64: newfstatat syscall ENOENT failure in multithreaded program
Message-ID<KhvmF-8v9-19@gated-at.bofh.it>
FYI, I can also reproduce this issue on my Samsung Galaxy S23 Ultra
phone under Termux/Android. Similarly, running the Perl script with
"strace -f" also yields a lot of stat failures.

-- 
Vincent Lefèvre <vincent@vinc17.net> - Web: <https://www.vinc17.net/>
100% accessible validated (X)HTML - Blog: <https://www.vinc17.net/blog/>
Work: CR INRIA - computer arithmetic / Pascaline project (LIP, ENS-Lyon)

[toc] | [next] | [standalone]


#85538

FromVincent Lefevre <vincent@vinc17.net>
Date2025-02-18 14:10 +0100
Message-ID<Khvwl-8On-1@gated-at.bofh.it>
In reply to#85537
On 2025-02-18 13:49:18 +0100, Vincent Lefevre wrote:
> FYI, I can also reproduce this issue on my Samsung Galaxy S23 Ultra
> phone under Termux/Android. Similarly, running the Perl script with
> "strace -f" also yields a lot of stat failures.

I can also reproduce it under macOS (cfarm104.cfarm.net), so I'm
wondering whether the bug could actually be in Perl. But in any case,
I don't understand how this couldn't be a bug in the Linux kernel,
according to the strace output.

-- 
Vincent Lefèvre <vincent@vinc17.net> - Web: <https://www.vinc17.net/>
100% accessible validated (X)HTML - Blog: <https://www.vinc17.net/blog/>
Work: CR INRIA - computer arithmetic / Pascaline project (LIP, ENS-Lyon)

[toc] | [prev] | [next] | [standalone]


#85539

FromVincent Lefevre <vincent@vinc17.net>
Date2025-02-18 14:40 +0100
Message-ID<KhvZn-8YF-1@gated-at.bofh.it>
In reply to#85538
On 2025-02-18 14:07:36 +0100, Vincent Lefevre wrote:
> On 2025-02-18 13:49:18 +0100, Vincent Lefevre wrote:
> > FYI, I can also reproduce this issue on my Samsung Galaxy S23 Ultra
> > phone under Termux/Android. Similarly, running the Perl script with
> > "strace -f" also yields a lot of stat failures.
> 
> I can also reproduce it under macOS (cfarm104.cfarm.net), so I'm
> wondering whether the bug could actually be in Perl. But in any case,
> I don't understand how this couldn't be a bug in the Linux kernel,
> according to the strace output.

Hmm... There's a fchdir in the strace output. If the current directory
is global to the process, this could be an issue. I now really suspect
a bug in perl.

If I change my script to do

opendir DIR, $dir or die "$0: opendir failed ($!)\n";
my @files = readdir DIR;
foreach my $file (@files)
  {
    $nthreads < $maxthreads or join_threads;
    $nthreads++ < $maxthreads or die "$0: internal error\n";
    threads->create(\&stat_test, $file);
  }
closedir DIR or die "$0: closedir failed ($!)\n";

then the failures still occur.

But if I then move the closedir as follows

opendir DIR, $dir or die "$0: opendir failed ($!)\n";
my @files = readdir DIR;
closedir DIR or die "$0: closedir failed ($!)\n";
foreach my $file (@files)
  {
    $nthreads < $maxthreads or join_threads;
    $nthreads++ < $maxthreads or die "$0: internal error\n";
    threads->create(\&stat_test, $file);
  }

the failures no longer occur.

-- 
Vincent Lefèvre <vincent@vinc17.net> - Web: <https://www.vinc17.net/>
100% accessible validated (X)HTML - Blog: <https://www.vinc17.net/blog/>
Work: CR INRIA - computer arithmetic / Pascaline project (LIP, ENS-Lyon)

[toc] | [prev] | [next] | [standalone]


#85555 — Bug#1098226: perl: thread creation while a directory handle is open does a fchdir, affecting other threads

FromVincent Lefevre <vincent@vinc17.net>
Date2025-02-18 17:10 +0100
SubjectBug#1098226: perl: thread creation while a directory handle is open does a fchdir, affecting other threads
Message-ID<Khykx-aB2-1@gated-at.bofh.it>
In reply to#85539
Control: reassign -1 perl 5.40.1-2
Control: retitle -1 perl: thread creation while a directory handle is open does a fchdir, affecting other threads (race condition)
Control: tags -1 security upstream
Control: severity -1 grave
Control: forwarded -1 https://github.com/Perl/perl5/issues/23010

This is a bug visible in the perl code, so I've just reported the bug
upstream.

(Not sure about the severity, but this can yield incorrect file
operations in the involved directory, which may be very problematic
if this directory is untrusted.)

On 2025-02-18 14:26:54 +0100, Vincent Lefevre wrote:
> Hmm... There's a fchdir in the strace output. If the current directory
> is global to the process, this could be an issue. I now really suspect
> a bug in perl.

Yes, thread creation does a chdir when a directory handle is open.
As the current working directory is global to the process, this
can affect other threads, if they do file operations with relative
pathnames. Even though the current working directory is set back
to the old value, this is a race condition, which can affect real
scripts (this is how I identified this bug).

-- 
Vincent Lefèvre <vincent@vinc17.net> - Web: <https://www.vinc17.net/>
100% accessible validated (X)HTML - Blog: <https://www.vinc17.net/blog/>
Work: CR INRIA - computer arithmetic / Pascaline project (LIP, ENS-Lyon)

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web