Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #270702 > unrolled thread

6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off

Started byJustin Piszcz <jpiszcz@lucidpixels.com>
First post2024-07-01 14:40 +0200
Last post2024-07-03 17:30 +0200
Articles 7 — 3 participants

Back to article view | Back to linux.debian.user


Contents

  6.1.0: NVME drive goes offline randomly even with:  nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Justin Piszcz <jpiszcz@lucidpixels.com> - 2024-07-01 14:40 +0200
    Re: 6.1.0: NVME drive goes offline randomly even with:  nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Michael Kjörling <c9bc136c6063@ewoof.net> - 2024-07-01 14:40 +0200
      Re: 6.1.0: NVME drive goes offline randomly even with:  nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Justin Piszcz <jpiszcz@lucidpixels.com> - 2024-07-01 17:50 +0200
        Re: 6.1.0: NVME drive goes offline randomly even with:  nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Dmitrii Odintcov <cyprussocialite@gmail.com> - 2024-07-01 18:00 +0200
        Re: 6.1.0: NVME drive goes offline randomly even with:  nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Michael Kjörling <c9bc136c6063@ewoof.net> - 2024-07-01 19:10 +0200
          Re: 6.1.0: NVME drive goes offline randomly even with:  nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Dmitrii Odintcov <cyprussocialite@gmail.com> - 2024-07-02 11:20 +0200
            Re: 6.1.0: NVME drive goes offline randomly even with:  nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Justin Piszcz <jpiszcz@lucidpixels.com> - 2024-07-03 17:30 +0200

#270702 — 6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off

FromJustin Piszcz <jpiszcz@lucidpixels.com>
Date2024-07-01 14:40 +0200
Subject6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off
Message-ID<IVou5-6VSL-5@gated-at.bofh.it>
Hello,

Note: I've also followed up on the LKML with this inquiry but as
Debian stable uses an older kernel (6.1.0), I was wondering if anyone
on this list has run into this problem?

Kernel: 6.1.0-17-amd64
Distribution: Debian stable
Arch: x86_64

I have 2 NVME drives as part of a BTRFS RAID-1, initially when this
happened the first time I added the following to the kernel cmdline at
boot:
nvme_core.default_ps_max_latency_us=0 pcie_aspm=off

This greatly reduced the frequency of this issue (last uptime was ~70
days).  However, it has occurred twice since then, this time I had
netconsole up to capture the crash.

The full kernel netconsole before during and after the crash:
https://installkernel.tripod.com/20240701-6.1.0-crash.txt

The model & firmware version of both drives are identical:
Model Number:                       Samsung SSD 990 PRO with Heatsink 4TB
Firmware Version:                   4B2QJXD7

Motherboard being used:
Manufacturer: ASUSTeK COMPUTER INC.
Product Name: Pro WS W680-ACE IPMI

Is there a workaround or potential fix for this issue?

The issue starts when this occurs:
[6078737.345641] nvme nvme2: I/O 154 (I/O Cmd) QID 6 timeout, aborting
[6078737.348143] nvme nvme2: I/O 155 (I/O Cmd) QID 6 timeout, aborting

Then later, a kernel panic:
[6078894.702941] BTRFS error (device nvme0n1p2): error writing primary
super block to device 2
[6078894.707920] BTRFS warning (device nvme0n1p2): csum hole found for
disk bytenr range [3659038877598419968, 3659038877598424064)
[6078894.708310] BTRFS critical (device nvme0n1p2): unable to find
chunk map for logical 3659038877598419968 length 4096
[6078894.708652] BUG: kernel NULL pointer dereference, address: 000000000000005a
[6078894.708879] #PF: supervisor read access in kernel mode
[6078894.709107] #PF: error_code(0x0000) - not-present page
[6078894.709292] PGD 0 P4D 0
[6078894.709509] Oops: 0000 [#1] PREEMPT SMP NOPTI
[6078894.709692] CPU: 12 PID: 3349611 Comm: kworker/u64:18 Not tainted
6.1.0-17-amd64 #1  Debian 6.1.69-1
[6078894.709856] Hardware name: ASUSTeK COMPUTER INC. System Product
Name/Pro WS W680-ACE IPMI, BIOS 3401 03/19/2024
[6078894.710022] Workqueue: btrfs-endio btrfs_end_bio_work [btrfs]
[6078894.710267] RIP: 0010:btrfs_get_io_geometry+0x13/0xf0 [btrfs]
[6078894.710483] Code: f4 ff ff ff e9 67 ff ff ff 66 66 2e 0f 1f 84 00
00 00 00 00 0f 1f 00 0f 1f 44 00 00 41 56 49 89 c9 48 89 cf 41 55 41
54 55 53 <4c> 8b 76 70 89 d3 31 d2 4c 8b 5e 18 41 8b 4e 10 45 8b 6e 14
4d 29
[6078894.710692] RSP: 0018:ffffa9cfc6657c08 EFLAGS: 00010286
[6078894.710711] BTRFS error (device nvme0n1p2): error writing primary
super block to device 2
[6078894.710876] RAX: ffffffffffffffea RBX: ffffffffffffffea RCX:
32c7847906990c00
[6078894.710882] RDX: 0000000000000000 RSI: ffffffffffffffea RDI:
32c7847906990c00
[6078894.710882] RBP: ffffa9cfc6657d28 R08: ffffa9cfc6657cc8 R09:
32c7847906990c00
[6078894.710882] R10: 0000000000000003 R11: ffff9a3efff6dc28 R12:
ffff9a2018195000
[6078894.710883] R13: 0000000000000001 R14: 0000000000001000 R15:
ffffa9cfc6657d50
[6078894.710884] FS:  0000000000000000(0000) GS:ffff9a3e7fb00000(0000)
knlGS:0000000000000000
[6078894.710884] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[6078894.710885] CR2: 000000000000005a CR3: 0000000bfc210000 CR4:
0000000000750ee0
[6078894.710885] PKRU: 55555554
[6078894.710885] Call Trace:
[6078894.710887]  <TASK>
[6078894.710891]  ? page_fault_oops+0xd2/0x2b0
[6078894.710889]  ? __die_body.cold+0x1a/0x1f
[6078894.710893]  ? exc_page_fault+0x70/0x170
[6078894.715724]  ? asm_exc_page_fault+0x22/0x30
[6078894.716084]  ? btrfs_get_io_geometry+0x13/0xf0 [btrfs]
[6078894.716470] BTRFS error (device nvme0n1p2): error writing primary
super block to device 2
[6078894.716462]  ? btrfs_get_chunk_map.cold+0x15/0x42 [btrfs]
[6078894.717384]  __btrfs_map_block+0xc4/0xe40 [btrfs]
[6078894.717771]  ? kmem_cache_free+0x15/0x310
[6078894.718147]  btrfs_submit_bio+0xa2/0x240 [btrfs]
[6078894.718571]  btrfs_repair_one_sector+0x29f/0x3a0 [btrfs]
[6078894.718972]  ? btrfs_submit_data_write_bio+0x110/0x110 [btrfs]
[6078894.719364]  end_compressed_bio_read+0x118/0x2f0 [btrfs]
[6078894.719753]  process_one_work+0x1c4/0x380
[6078894.720135]  worker_thread+0x4d/0x380
[6078894.720510]  ? rescuer_thread+0x3a0/0x3a0
[6078894.720864]  kthread+0xd7/0x100
[6078894.721224]  ? kthread_complete_and_exit+0x20/0x20
[6078894.721579]  ret_from_fork+0x1f/0x30
[6078894.721984]  </TASK>

Justin

[toc] | [next] | [standalone]


#270703

FromMichael Kjörling <c9bc136c6063@ewoof.net>
Date2024-07-01 14:40 +0200
Message-ID<IVou6-6VSL-7@gated-at.bofh.it>
In reply to#270702
On 1 Jul 2024 08:23 -0400, from jpiszcz@lucidpixels.com (Justin Piszcz):
> Kernel: 6.1.0-17-amd64
> Distribution: Debian stable
> Arch: x86_64

Your system is about half a year out of date. For Bookworm, 6.1.0-17
(6.1.69) is from early January; 6.1.0-18 (6.1.76) is from about a week
into February; and current is now 6.1.0-22 (6.1.94). (Upstream 6.1 is
at 6.1.96 since a few days ago.)

I suggest upgrading first, and seeing if the problem persists.

-- 
Michael Kjörling                     🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”

[toc] | [prev] | [next] | [standalone]


#270708

FromJustin Piszcz <jpiszcz@lucidpixels.com>
Date2024-07-01 17:50 +0200
Message-ID<IVrrY-6XC7-11@gated-at.bofh.it>
In reply to#270703
Hello,

Thanks, I've upgraded to the latest kernel version and will see if the
issue recurs.

$ uname -a
Linux int 6.1.0-22-amd64 #1 SMP PREEMPT_DYNAMIC Debian 6.1.94-1
(2024-06-21) x86_64 GNU/Linux

Regards,
Justin


On Mon, Jul 1, 2024 at 8:39 AM Michael Kjörling <c9bc136c6063@ewoof.net> wrote:
>
> On 1 Jul 2024 08:23 -0400, from jpiszcz@lucidpixels.com (Justin Piszcz):
> > Kernel: 6.1.0-17-amd64
> > Distribution: Debian stable
> > Arch: x86_64
>
> Your system is about half a year out of date. For Bookworm, 6.1.0-17
> (6.1.69) is from early January; 6.1.0-18 (6.1.76) is from about a week
> into February; and current is now 6.1.0-22 (6.1.94). (Upstream 6.1 is
> at 6.1.96 since a few days ago.)
>
> I suggest upgrading first, and seeing if the problem persists.
>
> --
> Michael Kjörling                     🔗 https://michael.kjorling.se
> “Remember when, on the Internet, nobody cared that you were a dog?”
>

[toc] | [prev] | [next] | [standalone]


#270709

FromDmitrii Odintcov <cyprussocialite@gmail.com>
Date2024-07-01 18:00 +0200
Message-ID<IVrBE-6XFj-7@gated-at.bofh.it>
In reply to#270708
Hey all,


Recently started experiencing the exact same issue.

Context:

  6.6.13+bpo-amd64 #1 SMP PREEMPT_DYNAMIC Debian 6.6.13-1~bpo12+1
(2024-02-15) x86_64 GNU/Linux

  Samsung SSD 970 Evo Plus 2TB

Happened a few times during heightened use (eg Steam installing a game
on the drive). Now I've added the recommended kernel options and will
see what happens.

[toc] | [prev] | [next] | [standalone]


#270712

FromMichael Kjörling <c9bc136c6063@ewoof.net>
Date2024-07-01 19:10 +0200
Message-ID<IVsHn-6YwV-7@gated-at.bofh.it>
In reply to#270708
On 1 Jul 2024 12:52 -0400, from jpiszcz@lucidpixels.com (Justin Piszcz):
> With the latest stable kernel (6.1.0-22), it crashed (below) shortly
> after boot (1-2hr), with the prior version (6.1.0-17) it had been
> stable other than the NVME dropping out.  Will try/test with a newer
> bpo kernel or similar..

Thank you for testing with an up-to-date kernel.

If you are able to, consider testing with a vanilla upstream kernel
compiled using the Debian kernel configuration settings for that same
branch. If that crashes too, then it's not something that Debian has
introduced but is rather either an upstream issue or something about
your hardware.

-- 
Michael Kjörling                     🔗 https://michael.kjorling.se
“Remember when, on the Internet, nobody cared that you were a dog?”

[toc] | [prev] | [next] | [standalone]


#270728

FromDmitrii Odintcov <cyprussocialite@gmail.com>
Date2024-07-02 11:20 +0200
Message-ID<IVHQ5-77U9-1@gated-at.bofh.it>
In reply to#270712
Hi,


Just had it happen to me again - on boot this time - with the
recommended `nvme_core.default_ps_max_latency_us=0 pcie_aspm=off`.

Worth noting that it's only happening with one of two SSDs I have
installed - the other being Samsung SSD 980 PRO 2TB - and that they've
been working fine for several years until this started happening
sometime last month.

[toc] | [prev] | [next] | [standalone]


#270772

FromJustin Piszcz <jpiszcz@lucidpixels.com>
Date2024-07-03 17:30 +0200
Message-ID<IWa5I-7pWx-5@gated-at.bofh.it>
In reply to#270728
On Tue, Jul 2, 2024 at 5:17 AM Dmitrii Odintcov
<cyprussocialite@gmail.com> wrote:
>
> Hi,
>
>
> Just had it happen to me again - on boot this time - with the
> recommended `nvme_core.default_ps_max_latency_us=0 pcie_aspm=off`.
>
> Worth noting that it's only happening with one of two SSDs I have
> installed - the other being Samsung SSD 980 PRO 2TB - and that they've
> been working fine for several years until this started happening
> sometime last month.

One thing I noticed is when I boot without any special kernel
parameters I was seeing the following, I've read that this error can
be due to a buggy BIOS:

$ lspci | grep -i 1b.4
00:1b.4 PCI bridge: Intel Corporation Alder Lake-S PCH PCI Express
Root Port (rev 11)

Jul  2 14:09:00 x kernel: [ 1213.547609] pcieport 0000:00:1b.4: AER:
Multiple Corrected error message received from 0000:00:1b.4
Jul  2 14:09:00 x kernel: [ 1213.547843] pcieport 0000:00:1b.4: PCIe
Bus Error: severity=Corrected, type=Data Link Layer, (Receiver ID)
Jul  2 14:09:00 x kernel: [ 1213.548011] pcieport 0000:00:1b.4:
device [8086:7ac4] error status/mask=00000040/00002000
Jul  2 14:09:00 x kernel: [ 1213.548182] pcieport 0000:00:1b.4:    [ 6] BadTLP
Jul  2 14:09:43 x kernel: [ 1256.906830] pcieport 0000:00:1b.4: AER:
Multiple Corrected error message received from 0000:00:1b.4
Jul  2 14:09:43 x kernel: [ 1256.907169] pcieport 0000:00:1b.4: PCIe
Bus Error: severity=Corrected, type=Data Link Layer, (Receiver ID)
Jul  2 14:09:43 x kernel: [ 1256.907395] pcieport 0000:00:1b.4:
device [8086:7ac4] error status/mask=00000040/00002000
Jul  2 14:09:43 x kernel: [ 1256.907648] pcieport 0000:00:1b.4:    [ 6] BadTLP

I am now testing with pci=nommconf, so far it has been 20 hours and no
crash yet nor Corrected errors as I showed above, but I need to give
it more time.  I am curious if there is any improvement on your side
if only testing with pci=nommconf ?

Justin

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web