Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #89372 > unrolled thread

Bug#1116065: linux: kernel oops with rsync on MSI X99A (regression since 6.12)

Started bySalvatore Bonaccorso <carnil@debian.org>
First post2025-09-23 21:10 +0200
Last post2025-09-26 07:00 +0200
Articles 3 — 1 participant

Back to article view | Back to linux.debian.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Bug#1116065: linux: kernel oops with rsync on MSI X99A (regression since 6.12) Salvatore Bonaccorso <carnil@debian.org> - 2025-09-23 21:10 +0200
    Bug#1116065: linux: kernel oops with rsync on MSI X99A (regression since 6.12) Salvatore Bonaccorso <carnil@debian.org> - 2025-09-24 09:50 +0200
      Bug#1116065: linux: kernel oops with rsync on MSI X99A (regression since 6.12) Salvatore Bonaccorso <carnil@debian.org> - 2025-09-26 07:00 +0200

#89372 — Bug#1116065: linux: kernel oops with rsync on MSI X99A (regression since 6.12)

FromSalvatore Bonaccorso <carnil@debian.org>
Date2025-09-23 21:10 +0200
SubjectBug#1116065: linux: kernel oops with rsync on MSI X99A (regression since 6.12)
Message-ID<LygyJ-2iU-5@gated-at.bofh.it>
Control: tags -1 + moreinfo

Hi,

On Tue, Sep 23, 2025 at 05:59:30PM +0000, Antwerpen, G. (Gert) van wrote:
> Package: linux-image-amd64
> Version: 6.12.0-1 (also reproducible with 6.16.3-1~bpo13+1)
> Severity: important
> Tags: upstream, regression
> 
> Summary:
> Kernel oops / crash when running rsync on MSI X99A SLI PLUS (MS-7885).
> System was stable for years with Debian stable kernel 6.1.x.
> The problem only appears after upgrading to Trixie (kernel 6.12 and newer).
> 
> System information:
> - Machine: MSI X99A SLI PLUS (MS-7885)
> - BIOS: 1.D0 (07/15/2016)
> - CPU: Intel Xeon (Haswell-E, socket 2011-3)
> - RAM: [please fill in, e.g. 64 GB DDR4 ECC/non-ECC]
> - Debian version: Trixie (testing)
> - Kernel versions tested:
>   - 6.1.x (Debian bookworm stable) → works fine
>   - 6.12.0-1 (Debian Trixie) → oops/crash
>   - 6.16.3-1~bpo13+1 (Debian Trixie backports) → oops/crash
> 
> Kernel log excerpt:
> BUG: unable to handle page fault for address: ffffffaabfe73e00
> #PF: supervisor instruction fetch in kernel mode
> #PF: error_code(0x0010) - not-present page
> PGD 1de231067 P4D 1de231067 PUD 0
> Oops: 0010 [#2] SMP PTI
> CPU: 9 UID: 40001 PID: 230866 Comm: rsync Tainted: G      D             6.16.3+deb13-amd64 #1 PREEMPT(lazy)  Debian 6.16.3-1~bpo13+1
> Hardware name: MSI MS-7885/X99A SLI PLUS(MS-7885), BIOS 1.D0 07/15/2016
> RIP: 0010:0xffffffaabfe73e00
> Code: Unable to access opcode bytes at 0xffffffaabfe73dd6.
> Call Trace:
>  filemap_readahead.isra.0+0x75/0xb0
>  filemap_get_pages+0x3ed/0x770
>  sock_write_iter+0x18e/0x1a0
>  ...
>  note: rsync[230866] exited with irqs disabled
> 
> (Full logs can be provided if required.)
> 
> Steps to reproduce:
> 1. Run rsync on large data sets (local disk to remote).
> 2. After some time, system crashes with kernel oops (see logs above).
> 3. Always reproducible on kernel >= 6.12, never seen on 6.1.
> 
> Expected result:
> No kernel oops — rsync should run reliably.
> 
> Actual result:
> Kernel crashes with page fault in kernel mode, requiring system restart.
> 
> Additional notes:
> - Hardware tested with memtest86+ (no errors).
> - No overclocking.
> - Issue seems to be a regression introduced in Linux 6.12.
> - Possibly related to filesystem or networking modules, but exact trigger unknown.

Can you please provide full kernel logs of the problem happening. If
you do not get access to the machine after oops'ing the you might
attach a netconsole to get the relevant logs.

Additionally to the logs ideally you provide all the meta information
collected by running reportbug's bugscripts for the kernel reports.

As for the regreesion itself and identify the breaking commit: Can you
bisect the upstream changes. Ideallally you first can range bit closer
the upstream versions where it is regressing. You can use for that the
snapshot.debian.org service to fetch older linux-image versions. Once
you have  close enough range, then bisect the upstream changes (would
you need help and have instructions to do that?).

> -- This message may contain information that is not intended for you. If you are not the addressee or if this message was sent to you by mistake, you are requested to inform the sender and delete the message. TNO accepts no liability for the content of this e-mail, for the manner in which you use it and for damage of any kind resulting from the risks inherent to the electronic transmission of messages.

YOu might want to drop this when filling a public bugreport ;-)

Regards,
Salvatore

[toc] | [next] | [standalone]


#89390

FromSalvatore Bonaccorso <carnil@debian.org>
Date2025-09-24 09:50 +0200
Message-ID<Lysqd-aN1-7@gated-at.bofh.it>
In reply to#89372
Hi Gert,

On Tue, Sep 23, 2025 at 07:11:54PM +0000, Antwerpen, G. (Gert) van wrote:
> Thanks to your fast reaction. I am not sure it's possible to pinpoint the exact releases where this occurred first time.
> 
> The problem occurs on a system that is in full operation, where I can't do all kinds of reboots to test different kernels.
> 
> I only see it on the Trixie kernel and the Trixie backports kernel now, the Bookworm kernel didn't have it.

Okay that makes things more difficult.

> Here is the complete relevant part of the kernel log:
> 
> 
> 
> BUG: unable to handle page fault for address: ffffffaabfe73e00
> 
> #PF: supervisor instruction fetch in kernel mode
> 
> #PF: error_code(0x0010) - not-present page
> 
> PGD 1de231067 P4D 1de231067 PUD 0
> 
> Oops: Oops: 0010 [#2] SMP PTI

Unfortunately that is not even the first Ooops. So there should be
more earlier already in the log. Are you able to provide the full
kernel log from once the problem is happening? 

If the system is not accessible anymore, then consider attaching a
netconsole so we get logs from as early as possible from boot and then
logged over the network. 

The full log would be better, if not, then at least we should get the
first oops and see from there.


> 
> CPU: 9 UID: 40001 PID: 230866 Comm: rsync Tainted: G      D             6.16.3+deb13-amd64 #1 PREEMPT(lazy)  Debian 6.16.3-1~bpo13+1
> 
> Tainted: [D]=DIE
> 
> Hardware name: MSI MS-7885/X99A SLI PLUS(MS-7885), BIOS 1.D0 07/15/2016

This sesm quite an old system, nwer BIOS version seems available still
slightly more recent taht the 07/15/2016 version, so you might
consider updating it.

Regards,
Salvatore

[toc] | [prev] | [next] | [standalone]


#89413

FromSalvatore Bonaccorso <carnil@debian.org>
Date2025-09-26 07:00 +0200
Message-ID<Lz8IN-Dtc-1@gated-at.bofh.it>
In reply to#89390
Control: retitle linux: kernel oops with rsync on MSI X99A with ntfs3

Hi

here is a very similar report in upstream's bugzilla:
https://bugzilla.kernel.org/show_bug.cgi?format=multiple&id=215460
(though it mentions to be fixed already, or occuring more rarely)

I will drop "regression since 6.12" from the subject, since it is
related to using ntfs3 driver in kernel, which was not available
before, so it's not directly a kernel regression we can bisect from
before situation.

Regards,
Salvatore

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web