Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.user > #270702 > unrolled thread
| Started by | Justin Piszcz <jpiszcz@lucidpixels.com> |
|---|---|
| First post | 2024-07-01 14:40 +0200 |
| Last post | 2024-07-03 17:30 +0200 |
| Articles | 7 — 3 participants |
Back to article view | Back to linux.debian.user
6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Justin Piszcz <jpiszcz@lucidpixels.com> - 2024-07-01 14:40 +0200
Re: 6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Michael Kjörling <c9bc136c6063@ewoof.net> - 2024-07-01 14:40 +0200
Re: 6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Justin Piszcz <jpiszcz@lucidpixels.com> - 2024-07-01 17:50 +0200
Re: 6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Dmitrii Odintcov <cyprussocialite@gmail.com> - 2024-07-01 18:00 +0200
Re: 6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Michael Kjörling <c9bc136c6063@ewoof.net> - 2024-07-01 19:10 +0200
Re: 6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Dmitrii Odintcov <cyprussocialite@gmail.com> - 2024-07-02 11:20 +0200
Re: 6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off Justin Piszcz <jpiszcz@lucidpixels.com> - 2024-07-03 17:30 +0200
| From | Justin Piszcz <jpiszcz@lucidpixels.com> |
|---|---|
| Date | 2024-07-01 14:40 +0200 |
| Subject | 6.1.0: NVME drive goes offline randomly even with: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off |
| Message-ID | <IVou5-6VSL-5@gated-at.bofh.it> |
Hello, Note: I've also followed up on the LKML with this inquiry but as Debian stable uses an older kernel (6.1.0), I was wondering if anyone on this list has run into this problem? Kernel: 6.1.0-17-amd64 Distribution: Debian stable Arch: x86_64 I have 2 NVME drives as part of a BTRFS RAID-1, initially when this happened the first time I added the following to the kernel cmdline at boot: nvme_core.default_ps_max_latency_us=0 pcie_aspm=off This greatly reduced the frequency of this issue (last uptime was ~70 days). However, it has occurred twice since then, this time I had netconsole up to capture the crash. The full kernel netconsole before during and after the crash: https://installkernel.tripod.com/20240701-6.1.0-crash.txt The model & firmware version of both drives are identical: Model Number: Samsung SSD 990 PRO with Heatsink 4TB Firmware Version: 4B2QJXD7 Motherboard being used: Manufacturer: ASUSTeK COMPUTER INC. Product Name: Pro WS W680-ACE IPMI Is there a workaround or potential fix for this issue? The issue starts when this occurs: [6078737.345641] nvme nvme2: I/O 154 (I/O Cmd) QID 6 timeout, aborting [6078737.348143] nvme nvme2: I/O 155 (I/O Cmd) QID 6 timeout, aborting Then later, a kernel panic: [6078894.702941] BTRFS error (device nvme0n1p2): error writing primary super block to device 2 [6078894.707920] BTRFS warning (device nvme0n1p2): csum hole found for disk bytenr range [3659038877598419968, 3659038877598424064) [6078894.708310] BTRFS critical (device nvme0n1p2): unable to find chunk map for logical 3659038877598419968 length 4096 [6078894.708652] BUG: kernel NULL pointer dereference, address: 000000000000005a [6078894.708879] #PF: supervisor read access in kernel mode [6078894.709107] #PF: error_code(0x0000) - not-present page [6078894.709292] PGD 0 P4D 0 [6078894.709509] Oops: 0000 [#1] PREEMPT SMP NOPTI [6078894.709692] CPU: 12 PID: 3349611 Comm: kworker/u64:18 Not tainted 6.1.0-17-amd64 #1 Debian 6.1.69-1 [6078894.709856] Hardware name: ASUSTeK COMPUTER INC. System Product Name/Pro WS W680-ACE IPMI, BIOS 3401 03/19/2024 [6078894.710022] Workqueue: btrfs-endio btrfs_end_bio_work [btrfs] [6078894.710267] RIP: 0010:btrfs_get_io_geometry+0x13/0xf0 [btrfs] [6078894.710483] Code: f4 ff ff ff e9 67 ff ff ff 66 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 00 0f 1f 44 00 00 41 56 49 89 c9 48 89 cf 41 55 41 54 55 53 <4c> 8b 76 70 89 d3 31 d2 4c 8b 5e 18 41 8b 4e 10 45 8b 6e 14 4d 29 [6078894.710692] RSP: 0018:ffffa9cfc6657c08 EFLAGS: 00010286 [6078894.710711] BTRFS error (device nvme0n1p2): error writing primary super block to device 2 [6078894.710876] RAX: ffffffffffffffea RBX: ffffffffffffffea RCX: 32c7847906990c00 [6078894.710882] RDX: 0000000000000000 RSI: ffffffffffffffea RDI: 32c7847906990c00 [6078894.710882] RBP: ffffa9cfc6657d28 R08: ffffa9cfc6657cc8 R09: 32c7847906990c00 [6078894.710882] R10: 0000000000000003 R11: ffff9a3efff6dc28 R12: ffff9a2018195000 [6078894.710883] R13: 0000000000000001 R14: 0000000000001000 R15: ffffa9cfc6657d50 [6078894.710884] FS: 0000000000000000(0000) GS:ffff9a3e7fb00000(0000) knlGS:0000000000000000 [6078894.710884] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [6078894.710885] CR2: 000000000000005a CR3: 0000000bfc210000 CR4: 0000000000750ee0 [6078894.710885] PKRU: 55555554 [6078894.710885] Call Trace: [6078894.710887] <TASK> [6078894.710891] ? page_fault_oops+0xd2/0x2b0 [6078894.710889] ? __die_body.cold+0x1a/0x1f [6078894.710893] ? exc_page_fault+0x70/0x170 [6078894.715724] ? asm_exc_page_fault+0x22/0x30 [6078894.716084] ? btrfs_get_io_geometry+0x13/0xf0 [btrfs] [6078894.716470] BTRFS error (device nvme0n1p2): error writing primary super block to device 2 [6078894.716462] ? btrfs_get_chunk_map.cold+0x15/0x42 [btrfs] [6078894.717384] __btrfs_map_block+0xc4/0xe40 [btrfs] [6078894.717771] ? kmem_cache_free+0x15/0x310 [6078894.718147] btrfs_submit_bio+0xa2/0x240 [btrfs] [6078894.718571] btrfs_repair_one_sector+0x29f/0x3a0 [btrfs] [6078894.718972] ? btrfs_submit_data_write_bio+0x110/0x110 [btrfs] [6078894.719364] end_compressed_bio_read+0x118/0x2f0 [btrfs] [6078894.719753] process_one_work+0x1c4/0x380 [6078894.720135] worker_thread+0x4d/0x380 [6078894.720510] ? rescuer_thread+0x3a0/0x3a0 [6078894.720864] kthread+0xd7/0x100 [6078894.721224] ? kthread_complete_and_exit+0x20/0x20 [6078894.721579] ret_from_fork+0x1f/0x30 [6078894.721984] </TASK> Justin
[toc] | [next] | [standalone]
| From | Michael Kjörling <c9bc136c6063@ewoof.net> |
|---|---|
| Date | 2024-07-01 14:40 +0200 |
| Message-ID | <IVou6-6VSL-7@gated-at.bofh.it> |
| In reply to | #270702 |
On 1 Jul 2024 08:23 -0400, from jpiszcz@lucidpixels.com (Justin Piszcz): > Kernel: 6.1.0-17-amd64 > Distribution: Debian stable > Arch: x86_64 Your system is about half a year out of date. For Bookworm, 6.1.0-17 (6.1.69) is from early January; 6.1.0-18 (6.1.76) is from about a week into February; and current is now 6.1.0-22 (6.1.94). (Upstream 6.1 is at 6.1.96 since a few days ago.) I suggest upgrading first, and seeing if the problem persists. -- Michael Kjörling 🔗 https://michael.kjorling.se “Remember when, on the Internet, nobody cared that you were a dog?”
[toc] | [prev] | [next] | [standalone]
| From | Justin Piszcz <jpiszcz@lucidpixels.com> |
|---|---|
| Date | 2024-07-01 17:50 +0200 |
| Message-ID | <IVrrY-6XC7-11@gated-at.bofh.it> |
| In reply to | #270703 |
Hello, Thanks, I've upgraded to the latest kernel version and will see if the issue recurs. $ uname -a Linux int 6.1.0-22-amd64 #1 SMP PREEMPT_DYNAMIC Debian 6.1.94-1 (2024-06-21) x86_64 GNU/Linux Regards, Justin On Mon, Jul 1, 2024 at 8:39 AM Michael Kjörling <c9bc136c6063@ewoof.net> wrote: > > On 1 Jul 2024 08:23 -0400, from jpiszcz@lucidpixels.com (Justin Piszcz): > > Kernel: 6.1.0-17-amd64 > > Distribution: Debian stable > > Arch: x86_64 > > Your system is about half a year out of date. For Bookworm, 6.1.0-17 > (6.1.69) is from early January; 6.1.0-18 (6.1.76) is from about a week > into February; and current is now 6.1.0-22 (6.1.94). (Upstream 6.1 is > at 6.1.96 since a few days ago.) > > I suggest upgrading first, and seeing if the problem persists. > > -- > Michael Kjörling 🔗 https://michael.kjorling.se > “Remember when, on the Internet, nobody cared that you were a dog?” >
[toc] | [prev] | [next] | [standalone]
| From | Dmitrii Odintcov <cyprussocialite@gmail.com> |
|---|---|
| Date | 2024-07-01 18:00 +0200 |
| Message-ID | <IVrBE-6XFj-7@gated-at.bofh.it> |
| In reply to | #270708 |
Hey all, Recently started experiencing the exact same issue. Context: 6.6.13+bpo-amd64 #1 SMP PREEMPT_DYNAMIC Debian 6.6.13-1~bpo12+1 (2024-02-15) x86_64 GNU/Linux Samsung SSD 970 Evo Plus 2TB Happened a few times during heightened use (eg Steam installing a game on the drive). Now I've added the recommended kernel options and will see what happens.
[toc] | [prev] | [next] | [standalone]
| From | Michael Kjörling <c9bc136c6063@ewoof.net> |
|---|---|
| Date | 2024-07-01 19:10 +0200 |
| Message-ID | <IVsHn-6YwV-7@gated-at.bofh.it> |
| In reply to | #270708 |
On 1 Jul 2024 12:52 -0400, from jpiszcz@lucidpixels.com (Justin Piszcz): > With the latest stable kernel (6.1.0-22), it crashed (below) shortly > after boot (1-2hr), with the prior version (6.1.0-17) it had been > stable other than the NVME dropping out. Will try/test with a newer > bpo kernel or similar.. Thank you for testing with an up-to-date kernel. If you are able to, consider testing with a vanilla upstream kernel compiled using the Debian kernel configuration settings for that same branch. If that crashes too, then it's not something that Debian has introduced but is rather either an upstream issue or something about your hardware. -- Michael Kjörling 🔗 https://michael.kjorling.se “Remember when, on the Internet, nobody cared that you were a dog?”
[toc] | [prev] | [next] | [standalone]
| From | Dmitrii Odintcov <cyprussocialite@gmail.com> |
|---|---|
| Date | 2024-07-02 11:20 +0200 |
| Message-ID | <IVHQ5-77U9-1@gated-at.bofh.it> |
| In reply to | #270712 |
Hi, Just had it happen to me again - on boot this time - with the recommended `nvme_core.default_ps_max_latency_us=0 pcie_aspm=off`. Worth noting that it's only happening with one of two SSDs I have installed - the other being Samsung SSD 980 PRO 2TB - and that they've been working fine for several years until this started happening sometime last month.
[toc] | [prev] | [next] | [standalone]
| From | Justin Piszcz <jpiszcz@lucidpixels.com> |
|---|---|
| Date | 2024-07-03 17:30 +0200 |
| Message-ID | <IWa5I-7pWx-5@gated-at.bofh.it> |
| In reply to | #270728 |
On Tue, Jul 2, 2024 at 5:17 AM Dmitrii Odintcov <cyprussocialite@gmail.com> wrote: > > Hi, > > > Just had it happen to me again - on boot this time - with the > recommended `nvme_core.default_ps_max_latency_us=0 pcie_aspm=off`. > > Worth noting that it's only happening with one of two SSDs I have > installed - the other being Samsung SSD 980 PRO 2TB - and that they've > been working fine for several years until this started happening > sometime last month. One thing I noticed is when I boot without any special kernel parameters I was seeing the following, I've read that this error can be due to a buggy BIOS: $ lspci | grep -i 1b.4 00:1b.4 PCI bridge: Intel Corporation Alder Lake-S PCH PCI Express Root Port (rev 11) Jul 2 14:09:00 x kernel: [ 1213.547609] pcieport 0000:00:1b.4: AER: Multiple Corrected error message received from 0000:00:1b.4 Jul 2 14:09:00 x kernel: [ 1213.547843] pcieport 0000:00:1b.4: PCIe Bus Error: severity=Corrected, type=Data Link Layer, (Receiver ID) Jul 2 14:09:00 x kernel: [ 1213.548011] pcieport 0000:00:1b.4: device [8086:7ac4] error status/mask=00000040/00002000 Jul 2 14:09:00 x kernel: [ 1213.548182] pcieport 0000:00:1b.4: [ 6] BadTLP Jul 2 14:09:43 x kernel: [ 1256.906830] pcieport 0000:00:1b.4: AER: Multiple Corrected error message received from 0000:00:1b.4 Jul 2 14:09:43 x kernel: [ 1256.907169] pcieport 0000:00:1b.4: PCIe Bus Error: severity=Corrected, type=Data Link Layer, (Receiver ID) Jul 2 14:09:43 x kernel: [ 1256.907395] pcieport 0000:00:1b.4: device [8086:7ac4] error status/mask=00000040/00002000 Jul 2 14:09:43 x kernel: [ 1256.907648] pcieport 0000:00:1b.4: [ 6] BadTLP I am now testing with pci=nommconf, so far it has been 20 hours and no crash yet nor Corrected errors as I showed above, but I need to give it more time. I am curious if there is any improvement on your side if only testing with pci=nommconf ? Justin
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.user
csiph-web