Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.bugs.dist > #1260791 > unrolled thread
| Started by | Marcel Jira <marcel.jira@gmail.com> |
|---|---|
| First post | 2025-09-10 08:10 +0200 |
| Last post | 2025-09-13 20:50 +0200 |
| Articles | 4 — 2 participants |
Back to article view | Back to linux.debian.bugs.dist
Bug#1114806: Regression: Suspend to RAM fails on kernel 6.16.3 (AMDGPU hang) Marcel Jira <marcel.jira@gmail.com> - 2025-09-10 08:10 +0200
Bug#1114806: Regression: Suspend to RAM fails on kernel 6.16.3 (AMDGPU hang) Salvatore Bonaccorso <carnil@debian.org> - 2025-09-10 08:20 +0200
Bug#1114806: Bisect result Salvatore Bonaccorso <carnil@debian.org> - 2025-09-13 20:30 +0200
Bug#1114806: Bisect result Salvatore Bonaccorso <carnil@debian.org> - 2025-09-13 20:50 +0200
| From | Marcel Jira <marcel.jira@gmail.com> |
|---|---|
| Date | 2025-09-10 08:10 +0200 |
| Subject | Bug#1114806: Regression: Suspend to RAM fails on kernel 6.16.3 (AMDGPU hang) |
| Message-ID | <LtmbL-edrA-1@gated-at.bofh.it> |
[Multipart message — attachments visible in raw view] — view raw
Package: inux-image-6.16.3+deb14-amd64 Version: linux-image-6.16.3+deb14-amd64 Severity: important Tags: upstream X-Debbugs-Cc: marcel.jira@gmail.com Dear Maintainer, since the upgrade to Linux kernel 6.16.3, suspend to RAM no longer works on my system. With kernel 6.12.38 (previous version in testing) suspend still works as expected, so this is a regression introduced between 6.12 and 6.16. --- Observed behavior (user perspective): - When suspending (on timeout or via power button), the screen turns off as expected. - Mouse and keyboard LEDs flash a few times, then remain dark (I did not observer those flashes on previous kernels, but I am unsure). - Screens go dark as expected. - The system power button LED stays permanently on (it should normally turn off). - The system never resumes from suspend. - Only a hard power-off (long press on power button) allows restarting the machine. **Note:** Any unsaved work in RAM is lost due to the required hard reset. Expected behavior: - System should suspend and resume properly, like under kernel 6.12.38. --- Log excerpt (from `journalctl -b -1`): Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: MODE1 reset Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: GPU mode1 reset Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: GPU psp mode1 reset Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: psp reg (0x16061) wait timed out, mask: 8000ffff, read: ffffffff exp: 80000000 This strongly suggests the issue is in the amdgpu driver. --- Reproducibility: - Always, 100% reproducible. Hardware: CPU: AMD Ryzen 5 3600 GPU: Radeon RX 5500 XT Additional information: - systemd was upgraded to 258~rc3-1 at the same time, but suspend works fine when booting into kernel 6.12.38 with the new systemd. - Logs from the failing suspend attempt (`journalctl -b -1`) are attached in full. Thank you for your work on the Debian kernel! Best regards, Marcel Jira -- System Information: Debian Release: forky/sid APT prefers testing APT policy: (990, 'testing') Architecture: amd64 (x86_64) Foreign Architectures: i386 Kernel: Linux 6.16.3+deb14-amd64 (SMP w/12 CPU threads; PREEMPT) Locale: LANG=de_AT.UTF-8, LC_CTYPE=de_AT.UTF-8 (charmap=UTF-8), LANGUAGE=de_AT:de Shell: /bin/sh linked to /usr/bin/dash Init: systemd (via /run/systemd/system) LSM: AppArmor: enabled
[toc] | [next] | [standalone]
| From | Salvatore Bonaccorso <carnil@debian.org> |
|---|---|
| Date | 2025-09-10 08:20 +0200 |
| Message-ID | <Ltmlr-edve-5@gated-at.bofh.it> |
| In reply to | #1260791 |
Control: reassign -1 src:linux 6.16.3-1 Control: tags -1 + moreinfo upstream Hi Marcel, On Wed, Sep 10, 2025 at 08:06:10AM +0200, Marcel Jira wrote: > Package: inux-image-6.16.3+deb14-amd64 > Version: linux-image-6.16.3+deb14-amd64 > Severity: important > Tags: upstream > X-Debbugs-Cc: marcel.jira@gmail.com > > Dear Maintainer, > > since the upgrade to Linux kernel 6.16.3, suspend to RAM no longer works on my > system. > With kernel 6.12.38 (previous version in testing) suspend still works as > expected, so this is a regression introduced between 6.12 and 6.16. > > --- > > Observed behavior (user perspective): > - When suspending (on timeout or via power button), the screen turns off as > expected. > - Mouse and keyboard LEDs flash a few times, then remain dark (I did not > observer those flashes on previous kernels, but I am unsure). > - Screens go dark as expected. > - The system power button LED stays permanently on (it should normally turn > off). > - The system never resumes from suspend. > - Only a hard power-off (long press on power button) allows restarting the > machine. > > **Note:** Any unsaved work in RAM is lost due to the required hard reset. > > Expected behavior: > - System should suspend and resume properly, like under kernel 6.12.38. > > --- > > Log excerpt (from `journalctl -b -1`): > > Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: MODE1 reset > Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: GPU mode1 > reset > Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: GPU psp mode1 > reset > Sep 10 07:33:02 computer-id kernel: amdgpu 0000:0b:00.0: amdgpu: psp reg > (0x16061) wait timed out, mask: 8000ffff, read: ffffffff exp: 80000000 > > > This strongly suggests the issue is in the amdgpu driver. > > --- > > Reproducibility: > - Always, 100% reproducible. > > Hardware: > CPU: AMD Ryzen 5 3600 > GPU: Radeon RX 5500 XT > > Additional information: > - systemd was upgraded to 258~rc3-1 at the same time, but suspend works fine > when booting into kernel 6.12.38 with the new systemd. > - Logs from the failing suspend attempt (`journalctl -b -1`) are attached in > full. > > Thank you for your work on the Debian kernel! Thanks for your report. Ideally please attach the full kernel log to the bug. Along though I have two questions: Can you please test the current version in unstable (6.16.5-1), is the problem solved there? If not: Can you bisect the upstream changes to isolate the commit which breaks the behaviour? (Do you need rough instructions on how to do it?) Regards, Salvatore
[toc] | [prev] | [next] | [standalone]
| From | Salvatore Bonaccorso <carnil@debian.org> |
|---|---|
| Date | 2025-09-13 20:30 +0200 |
| Subject | Bug#1114806: Bisect result |
| Message-ID | <LuDax-f415-5@gated-at.bofh.it> |
| In reply to | #1260791 |
Hi Niklas,
On Fri, Sep 12, 2025 at 08:02:02PM +0200, Niklas Cathor wrote:
> Hi Salvatore,
>
> I encountered the same issue, and was able to bisect. I'm pasting the result
> below.
> Thank you for looking into this. Let me know if I should report it upstream
> instead.
>
> cheers,
> Niklas
>
>
> 165a69a87d6bde85cac2c051fa6da611ca4524f6 is the first bad commit
> commit 165a69a87d6bde85cac2c051fa6da611ca4524f6 (HEAD)
> Author: Lijo Lazar <lijo.lazar@amd.com>
> Date: Mon Jun 2 12:55:14 2025 +0530
>
> drm/amdgpu: Add more checks to PSP mailbox
>
> [ Upstream commit 8345a71fc54b28e4d13a759c45ce2664d8540d28 ]
>
> Instead of checking the response flag, use status mask also to check
> against any unexpected failures like a device drop. Also, log error if
> waiting on a psp response fails/times out.
>
> Signed-off-by: Lijo Lazar <lijo.lazar@amd.com>
> Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
> Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
> Signed-off-by: Sasha Levin <sashal@kernel.org>
>
> drivers/gpu/drm/amd/amdgpu/amdgpu_psp.c | 4 ++++
> drivers/gpu/drm/amd/amdgpu/amdgpu_psp.h | 11 +++++++++++
> drivers/gpu/drm/amd/amdgpu/psp_v10_0.c | 4 ++--
> drivers/gpu/drm/amd/amdgpu/psp_v11_0.c | 31
> +++++++++++++++++++------------
> drivers/gpu/drm/amd/amdgpu/psp_v11_0_8.c | 25 +++++++++++++++----------
> drivers/gpu/drm/amd/amdgpu/psp_v12_0.c | 18 +++++++++++-------
> drivers/gpu/drm/amd/amdgpu/psp_v13_0.c | 25 +++++++++++++++----------
> drivers/gpu/drm/amd/amdgpu/psp_v13_0_4.c | 25 +++++++++++++++----------
> drivers/gpu/drm/amd/amdgpu/psp_v14_0.c | 25 +++++++++++++++----------
> 9 files changed, 107 insertions(+), 61 deletions(-)
Ok that is great you found the offending commit. Can you try if
applying 440cec4ca1c2 ("drm/amdgpu: Wait for bootloader after PSPv11
reset") fixes the issue?
Regards,
Salvatore
[toc] | [prev] | [next] | [standalone]
| From | Salvatore Bonaccorso <carnil@debian.org> |
|---|---|
| Date | 2025-09-13 20:50 +0200 |
| Subject | Bug#1114806: Bisect result |
| Message-ID | <LuDtT-f489-1@gated-at.bofh.it> |
| In reply to | #1261262 |
[Multipart message — attachments visible in raw view] — view raw
Hi Niklas,
On Sat, Sep 13, 2025 at 08:23:01PM +0200, Salvatore Bonaccorso wrote:
> Hi Niklas,
>
> On Fri, Sep 12, 2025 at 08:02:02PM +0200, Niklas Cathor wrote:
> > Hi Salvatore,
> >
> > I encountered the same issue, and was able to bisect. I'm pasting the result
> > below.
> > Thank you for looking into this. Let me know if I should report it upstream
> > instead.
> >
> > cheers,
> > Niklas
> >
> >
> > 165a69a87d6bde85cac2c051fa6da611ca4524f6 is the first bad commit
> > commit 165a69a87d6bde85cac2c051fa6da611ca4524f6 (HEAD)
> > Author: Lijo Lazar <lijo.lazar@amd.com>
> > Date: Mon Jun 2 12:55:14 2025 +0530
> >
> > drm/amdgpu: Add more checks to PSP mailbox
> >
> > [ Upstream commit 8345a71fc54b28e4d13a759c45ce2664d8540d28 ]
> >
> > Instead of checking the response flag, use status mask also to check
> > against any unexpected failures like a device drop. Also, log error if
> > waiting on a psp response fails/times out.
> >
> > Signed-off-by: Lijo Lazar <lijo.lazar@amd.com>
> > Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
> > Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
> > Signed-off-by: Sasha Levin <sashal@kernel.org>
> >
> > drivers/gpu/drm/amd/amdgpu/amdgpu_psp.c | 4 ++++
> > drivers/gpu/drm/amd/amdgpu/amdgpu_psp.h | 11 +++++++++++
> > drivers/gpu/drm/amd/amdgpu/psp_v10_0.c | 4 ++--
> > drivers/gpu/drm/amd/amdgpu/psp_v11_0.c | 31
> > +++++++++++++++++++------------
> > drivers/gpu/drm/amd/amdgpu/psp_v11_0_8.c | 25 +++++++++++++++----------
> > drivers/gpu/drm/amd/amdgpu/psp_v12_0.c | 18 +++++++++++-------
> > drivers/gpu/drm/amd/amdgpu/psp_v13_0.c | 25 +++++++++++++++----------
> > drivers/gpu/drm/amd/amdgpu/psp_v13_0_4.c | 25 +++++++++++++++----------
> > drivers/gpu/drm/amd/amdgpu/psp_v14_0.c | 25 +++++++++++++++----------
> > 9 files changed, 107 insertions(+), 61 deletions(-)
>
> Ok that is great you found the offending commit. Can you try if
> applying 440cec4ca1c2 ("drm/amdgpu: Wait for bootloader after PSPv11
> reset") fixes the issue?
One thing: the commit won't apply cleanly pre 9888f73679b7
("drm/amdgpu: Add a noverbose flag to psp_wait_for") changes. So
either test mainline at the commit and the previous comit to confirm
the fix, and if possible then still with a backported variant.
An attempt of it is attached here which should apply on top of
6.16.7-1.
Regards,
Salvatore
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.bugs.dist
csiph-web