Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.bugs.dist > #1260791 > unrolled thread

Bug#1114806: Regression: Suspend to RAM fails on kernel 6.16.3 (AMDGPU hang)

Started byMarcel Jira <marcel.jira@gmail.com>
First post2025-09-10 08:10 +0200
Last post2025-09-13 20:50 +0200
Articles 4 — 2 participants

Back to article view | Back to linux.debian.bugs.dist


Contents

  Bug#1114806: Regression: Suspend to RAM fails on kernel 6.16.3 (AMDGPU hang) Marcel Jira <marcel.jira@gmail.com> - 2025-09-10 08:10 +0200
    Bug#1114806: Regression: Suspend to RAM fails on kernel 6.16.3 (AMDGPU hang) Salvatore Bonaccorso <carnil@debian.org> - 2025-09-10 08:20 +0200
    Bug#1114806: Bisect result Salvatore Bonaccorso <carnil@debian.org> - 2025-09-13 20:30 +0200
      Bug#1114806: Bisect result Salvatore Bonaccorso <carnil@debian.org> - 2025-09-13 20:50 +0200

#1260791 — Bug#1114806: Regression: Suspend to RAM fails on kernel 6.16.3 (AMDGPU hang)

FromMarcel Jira <marcel.jira@gmail.com>
Date2025-09-10 08:10 +0200
SubjectBug#1114806: Regression: Suspend to RAM fails on kernel 6.16.3 (AMDGPU hang)
Message-ID<LtmbL-edrA-1@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Package: inux-image-6.16.3+deb14-amd64
Version: linux-image-6.16.3+deb14-amd64
Severity: important
Tags: upstream
X-Debbugs-Cc: marcel.jira@gmail.com

Dear Maintainer,

since the upgrade to Linux kernel 6.16.3, suspend to RAM no longer works on my
system.
With kernel 6.12.38 (previous version in testing) suspend still works as
expected, so this is a regression introduced between 6.12 and 6.16.

---

Observed behavior (user perspective):
- When suspending (on timeout or via power button), the screen turns off as
expected.
- Mouse and keyboard LEDs flash a few times, then remain dark (I did not
observer those flashes on previous kernels, but I am unsure).
- Screens go dark as expected.
- The system power button LED stays permanently on (it should normally turn
off).
- The system never resumes from suspend.
- Only a hard power-off (long press on power button) allows restarting the
machine.

**Note:** Any unsaved work in RAM is lost due to the required hard reset.

Expected behavior:
- System should suspend and resume properly, like under kernel 6.12.38.

---

Log excerpt (from `journalctl -b -1`):

Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: MODE1 reset
Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: GPU mode1
reset
Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: GPU psp mode1
reset
Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: psp reg
(0x16061) wait timed out, mask: 8000ffff, read: ffffffff exp: 80000000


This strongly suggests the issue is in the amdgpu driver.

---

Reproducibility:
- Always, 100% reproducible.

Hardware:
CPU: AMD Ryzen 5 3600
GPU: Radeon RX 5500 XT

Additional information:
- systemd was upgraded to 258~rc3-1 at the same time, but suspend works fine
when booting into kernel 6.12.38 with the new systemd.
- Logs from the failing suspend attempt (`journalctl -b -1`) are attached in
full.

Thank you for your work on the Debian kernel!

Best regards,

Marcel Jira


-- System Information:
Debian Release: forky/sid
  APT prefers testing
  APT policy: (990, 'testing')
Architecture: amd64 (x86_64)
Foreign Architectures: i386

Kernel: Linux 6.16.3+deb14-amd64 (SMP w/12 CPU threads; PREEMPT)
Locale: LANG=de_AT.UTF-8, LC_CTYPE=de_AT.UTF-8 (charmap=UTF-8), LANGUAGE=de_AT:de
Shell: /bin/sh linked to /usr/bin/dash
Init: systemd (via /run/systemd/system)
LSM: AppArmor: enabled

[toc] | [next] | [standalone]


#1260793

FromSalvatore Bonaccorso <carnil@debian.org>
Date2025-09-10 08:20 +0200
Message-ID<Ltmlr-edve-5@gated-at.bofh.it>
In reply to#1260791
Control: reassign -1 src:linux 6.16.3-1
Control: tags -1 + moreinfo upstream

Hi Marcel,

On Wed, Sep 10, 2025 at 08:06:10AM +0200, Marcel Jira wrote:
> Package: inux-image-6.16.3+deb14-amd64
> Version: linux-image-6.16.3+deb14-amd64
> Severity: important
> Tags: upstream
> X-Debbugs-Cc: marcel.jira@gmail.com
> 
> Dear Maintainer,
> 
> since the upgrade to Linux kernel 6.16.3, suspend to RAM no longer works on my
> system.
> With kernel 6.12.38 (previous version in testing) suspend still works as
> expected, so this is a regression introduced between 6.12 and 6.16.
> 
> ---
> 
> Observed behavior (user perspective):
> - When suspending (on timeout or via power button), the screen turns off as
> expected.
> - Mouse and keyboard LEDs flash a few times, then remain dark (I did not
> observer those flashes on previous kernels, but I am unsure).
> - Screens go dark as expected.
> - The system power button LED stays permanently on (it should normally turn
> off).
> - The system never resumes from suspend.
> - Only a hard power-off (long press on power button) allows restarting the
> machine.
> 
> **Note:** Any unsaved work in RAM is lost due to the required hard reset.
> 
> Expected behavior:
> - System should suspend and resume properly, like under kernel 6.12.38.
> 
> ---
> 
> Log excerpt (from `journalctl -b -1`):
> 
> Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: MODE1 reset
> Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: GPU mode1
> reset
> Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: GPU psp mode1
> reset
> Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: psp reg
> (0x16061) wait timed out, mask: 8000ffff, read: ffffffff exp: 80000000
> 
> 
> This strongly suggests the issue is in the amdgpu driver.
> 
> ---
> 
> Reproducibility:
> - Always, 100% reproducible.
> 
> Hardware:
> CPU: AMD Ryzen 5 3600
> GPU: Radeon RX 5500 XT
> 
> Additional information:
> - systemd was upgraded to 258~rc3-1 at the same time, but suspend works fine
> when booting into kernel 6.12.38 with the new systemd.
> - Logs from the failing suspend attempt (`journalctl -b -1`) are attached in
> full.
> 
> Thank you for your work on the Debian kernel!

Thanks for your report. Ideally please attach the full kernel log to
the bug. Along though I have two questions:

Can you please test the current version in unstable (6.16.5-1), is
the problem solved there?

If not: Can you bisect the upstream changes to isolate the commit
which breaks the behaviour? (Do you need rough instructions on how to
do it?)

Regards,
Salvatore

[toc] | [prev] | [next] | [standalone]


#1261262 — Bug#1114806: Bisect result

FromSalvatore Bonaccorso <carnil@debian.org>
Date2025-09-13 20:30 +0200
SubjectBug#1114806: Bisect result
Message-ID<LuDax-f415-5@gated-at.bofh.it>
In reply to#1260791
Hi Niklas,

On Fri, Sep 12, 2025 at 08:02:02PM +0200, Niklas Cathor wrote:
> Hi Salvatore,
> 
> I encountered the same issue, and was able to bisect. I'm pasting the result
> below.
> Thank you for looking into this. Let me know if I should report it upstream
> instead.
> 
> cheers,
> Niklas
> 
> 
> 165a69a87d6bde85cac2c051fa6da611ca4524f6 is the first bad commit
> commit 165a69a87d6bde85cac2c051fa6da611ca4524f6 (HEAD)
> Author: Lijo Lazar <lijo.lazar@amd.com>
> Date:   Mon Jun 2 12:55:14 2025 +0530
> 
>     drm/amdgpu: Add more checks to PSP mailbox
> 
>     [ Upstream commit 8345a71fc54b28e4d13a759c45ce2664d8540d28 ]
> 
>     Instead of checking the response flag, use status mask also to check
>     against any unexpected failures like a device drop. Also, log error if
>     waiting on a psp response fails/times out.
> 
>     Signed-off-by: Lijo Lazar <lijo.lazar@amd.com>
>     Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
>     Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
>     Signed-off-by: Sasha Levin <sashal@kernel.org>
> 
>  drivers/gpu/drm/amd/amdgpu/amdgpu_psp.c  |  4 ++++
>  drivers/gpu/drm/amd/amdgpu/amdgpu_psp.h  | 11 +++++++++++
>  drivers/gpu/drm/amd/amdgpu/psp_v10_0.c   |  4 ++--
>  drivers/gpu/drm/amd/amdgpu/psp_v11_0.c   | 31
> +++++++++++++++++++------------
>  drivers/gpu/drm/amd/amdgpu/psp_v11_0_8.c | 25 +++++++++++++++----------
>  drivers/gpu/drm/amd/amdgpu/psp_v12_0.c   | 18 +++++++++++-------
>  drivers/gpu/drm/amd/amdgpu/psp_v13_0.c   | 25 +++++++++++++++----------
>  drivers/gpu/drm/amd/amdgpu/psp_v13_0_4.c | 25 +++++++++++++++----------
>  drivers/gpu/drm/amd/amdgpu/psp_v14_0.c   | 25 +++++++++++++++----------
>  9 files changed, 107 insertions(+), 61 deletions(-)

Ok that is great you found the offending commit. Can you try if
applying 440cec4ca1c2 ("drm/amdgpu: Wait for bootloader after PSPv11
reset") fixes the issue?

Regards,
Salvatore

[toc] | [prev] | [next] | [standalone]


#1261265 — Bug#1114806: Bisect result

FromSalvatore Bonaccorso <carnil@debian.org>
Date2025-09-13 20:50 +0200
SubjectBug#1114806: Bisect result
Message-ID<LuDtT-f489-1@gated-at.bofh.it>
In reply to#1261262

[Multipart message — attachments visible in raw view] — view raw

Hi Niklas,

On Sat, Sep 13, 2025 at 08:23:01PM +0200, Salvatore Bonaccorso wrote:
> Hi Niklas,
> 
> On Fri, Sep 12, 2025 at 08:02:02PM +0200, Niklas Cathor wrote:
> > Hi Salvatore,
> > 
> > I encountered the same issue, and was able to bisect. I'm pasting the result
> > below.
> > Thank you for looking into this. Let me know if I should report it upstream
> > instead.
> > 
> > cheers,
> > Niklas
> > 
> > 
> > 165a69a87d6bde85cac2c051fa6da611ca4524f6 is the first bad commit
> > commit 165a69a87d6bde85cac2c051fa6da611ca4524f6 (HEAD)
> > Author: Lijo Lazar <lijo.lazar@amd.com>
> > Date:   Mon Jun 2 12:55:14 2025 +0530
> > 
> >     drm/amdgpu: Add more checks to PSP mailbox
> > 
> >     [ Upstream commit 8345a71fc54b28e4d13a759c45ce2664d8540d28 ]
> > 
> >     Instead of checking the response flag, use status mask also to check
> >     against any unexpected failures like a device drop. Also, log error if
> >     waiting on a psp response fails/times out.
> > 
> >     Signed-off-by: Lijo Lazar <lijo.lazar@amd.com>
> >     Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
> >     Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
> >     Signed-off-by: Sasha Levin <sashal@kernel.org>
> > 
> >  drivers/gpu/drm/amd/amdgpu/amdgpu_psp.c  |  4 ++++
> >  drivers/gpu/drm/amd/amdgpu/amdgpu_psp.h  | 11 +++++++++++
> >  drivers/gpu/drm/amd/amdgpu/psp_v10_0.c   |  4 ++--
> >  drivers/gpu/drm/amd/amdgpu/psp_v11_0.c   | 31
> > +++++++++++++++++++------------
> >  drivers/gpu/drm/amd/amdgpu/psp_v11_0_8.c | 25 +++++++++++++++----------
> >  drivers/gpu/drm/amd/amdgpu/psp_v12_0.c   | 18 +++++++++++-------
> >  drivers/gpu/drm/amd/amdgpu/psp_v13_0.c   | 25 +++++++++++++++----------
> >  drivers/gpu/drm/amd/amdgpu/psp_v13_0_4.c | 25 +++++++++++++++----------
> >  drivers/gpu/drm/amd/amdgpu/psp_v14_0.c   | 25 +++++++++++++++----------
> >  9 files changed, 107 insertions(+), 61 deletions(-)
> 
> Ok that is great you found the offending commit. Can you try if
> applying 440cec4ca1c2 ("drm/amdgpu: Wait for bootloader after PSPv11
> reset") fixes the issue?

One thing: the commit won't apply cleanly pre 9888f73679b7
("drm/amdgpu: Add a noverbose flag to psp_wait_for") changes. So
either test mainline at the commit and the previous comit to confirm
the fix, and if possible then still with a backported variant.

An attempt of it is attached here which should apply on top of
6.16.7-1.

Regards,
Salvatore

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.bugs.dist


csiph-web