Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #93953 > unrolled thread

Bug#1146343: amdgpu: RustDesk triggers VCE ring timeout and GPU reset on Polaris RX 480 since 6.12.105

Started bySalvatore Bonaccorso <carnil@debian.org>
First post2026-09-01 07:00 +0200
Last post2026-09-13 09:00 +0200
Articles 4 — 3 participants

Back to article view | Back to linux.debian.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Bug#1146343: amdgpu: RustDesk triggers VCE ring timeout and GPU reset on Polaris RX 480 since 6.12.105 Salvatore Bonaccorso <carnil@debian.org> - 2026-09-01 07:00 +0200
    Processed: Re: Bug#1146343: amdgpu: RustDesk triggers VCE ring  timeout and GPU reset on Polaris RX 480 since 6.12.105 "Debian Bug Tracking System" <owner@bugs.debian.org> - 2026-09-01 07:00 +0200
    Bug#1146343: amdgpu: RustDesk triggers VCE ring timeout and GPU reset on Polaris RX 480 since 6.12.105 Salvatore Bonaccorso <carnil@debian.org> - 2026-09-01 23:10 +0200
      Bug#1146343: amdgpu: RustDesk triggers VCE ring timeout and GPU reset on Polaris RX 480 since 6.12.105 Uwe Kleine-König <ukleinek@debian.org> - 2026-09-13 09:00 +0200

#93953 — Bug#1146343: amdgpu: RustDesk triggers VCE ring timeout and GPU reset on Polaris RX 480 since 6.12.105

FromSalvatore Bonaccorso <carnil@debian.org>
Date2026-09-01 07:00 +0200
SubjectBug#1146343: amdgpu: RustDesk triggers VCE ring timeout and GPU reset on Polaris RX 480 since 6.12.105
Message-ID<NyoLf-euXB-1@gated-at.bofh.it>
Control: tags -1 + moreinfo

Hi,

On Mon, Aug 31, 2026 at 10:41:47AM -0500, Manuel R. Buffa wrote:
> Package: src:linux
> Version: 6.12.107-1
> Severity: important
> X-Debbugs-Cc: manuel@manbpro.com
> 
> Dear Maintainer,
> 
> I am reporting a reproducible amdgpu/VCE regression on Debian 13 (trixie)
> with an AMD Radeon RX 480 / Polaris10 GPU. Active RustDesk remote-control
> sessions are stable on Debian kernel 6.12.101-1 but trigger a VCE ring timeout
> and GPU reset on 6.12.105-1 and 6.12.107-1.
[...]

As you can reproduce your issue reliably, please bisect the changes
between 6.12.101 and 6.12.105 to identify the breaking commit. That
given we can have a look if it is a known regression already or needs
to be reported upstream yet.

Let me know if you need instructions for the bisecting steps!

Regards,
Salvatore

[toc] | [next] | [standalone]


#93955 — Processed: Re: Bug#1146343: amdgpu: RustDesk triggers VCE ring timeout and GPU reset on Polaris RX 480 since 6.12.105

From"Debian Bug Tracking System" <owner@bugs.debian.org>
Date2026-09-01 07:00 +0200
SubjectProcessed: Re: Bug#1146343: amdgpu: RustDesk triggers VCE ring timeout and GPU reset on Polaris RX 480 since 6.12.105
Message-ID<NyoLf-euXB-11@gated-at.bofh.it>
In reply to#93953
Processing control commands:

> tags -1 + moreinfo
Bug #1146343 [src:linux] amdgpu: RustDesk triggers VCE ring timeout and GPU reset on Polaris RX 480 since 6.12.105
Added tag(s) moreinfo.

-- 
1146343: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1146343
Debian Bug Tracking System
Contact owner@bugs.debian.org with problems

[toc] | [prev] | [next] | [standalone]


#93964

FromSalvatore Bonaccorso <carnil@debian.org>
Date2026-09-01 23:10 +0200
Message-ID<NyDTX-eD2s-1@gated-at.bofh.it>
In reply to#93953
Hi Manuel, all

On Tue, Sep 01, 2026 at 06:52:55AM +0200, Salvatore Bonaccorso wrote:
> Control: tags -1 + moreinfo
> 
> Hi,
> 
> On Mon, Aug 31, 2026 at 10:41:47AM -0500, Manuel R. Buffa wrote:
> > Package: src:linux
> > Version: 6.12.107-1
> > Severity: important
> > X-Debbugs-Cc: manuel@manbpro.com
> > 
> > Dear Maintainer,
> > 
> > I am reporting a reproducible amdgpu/VCE regression on Debian 13 (trixie)
> > with an AMD Radeon RX 480 / Polaris10 GPU. Active RustDesk remote-control
> > sessions are stable on Debian kernel 6.12.101-1 but trigger a VCE ring timeout
> > and GPU reset on 6.12.105-1 and 6.12.107-1.
> [...]
> 
> As you can reproduce your issue reliably, please bisect the changes
> between 6.12.101 and 6.12.105 to identify the breaking commit. That
> given we can have a look if it is a known regression already or needs
> to be reported upstream yet.
> 
> Let me know if you need instructions for the bisecting steps!

Manuel asked off-bug about the instructions, so here we go. To bisect
between 6.12.101 and 6.12.105 proceed as follows (once you have the
my_config, compilation could be as well on a more powerful machine).

    git clone --single-branch -b linux-6.12.y https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux-stable.git
    cd linux-stable
    git checkout v6.12.101
    cp /boot/config-$(uname -r) .config
    yes '' | make localmodconfig
    make savedefconfig
    mv defconfig arch/x86/configs/my_defconfig

    # test 6.12.101 to ensure this is "good"
    make my_defconfig
    make -j $(nproc) bindeb-pkg
    ... install the resulting .deb package and confirm problem does not exist

    # test 6.12.105 to ensure this is "bad"
    git checkout v6.12.105
    make my_defconfig
    make -j $(nproc) bindeb-pkg
    ... install the resulting .deb package and confirm problem exists

With that confirmed, the bisection can start:

    git bisect start
    git bisect good v6.12.101
    git bisect bad v6.12.105

In each bisection step git checks out a state between the oldest
known-bad and the newest known-good commit. In each step test using:

    make my_defconfig
    make -j $(nproc) bindeb-pkg
    ... install, verify if problem exists

and if the problem is hit run:

    git bisect bad

and if the problem doesn't trigger run:

    git bisect good

. Please pay attention to always select the just built kernel for
booting, it won't always be the default kernel picked up by grub.

Iterate until git announces to have identified the first bad commit.

Then provide the output of

    git bisect log

In the course of the bisection you might have to uninstall previous
kernels again to not exhaust the disk space in /boot. Also in the end
uninstall all self-built kernels again.

Regards,
Salvatore

[toc] | [prev] | [next] | [standalone]


#94079

FromUwe Kleine-König <ukleinek@debian.org>
Date2026-09-13 09:00 +0200
Message-ID<NCMlX-gTib-1@gated-at.bofh.it>
In reply to#93964

[Multipart message — attachments visible in raw view] — view raw

Hello Manuel,

On Sat, Sep 05, 2026 at 06:44:58PM -0500, Manuel Buffa wrote:
> I completed the requested bisection between stable v6.12.101 and v6.12.105
> on the affected machine.
> 
> The kernel configuration was derived on the affected work-box while it was
> running Debian 6.12.101 using the requested localmodconfig / savedefconfig
> procedure. I froze that my_defconfig and used it for every build.
> Compilation and packaging occurred on a separate amd64 builder; every
> GOOD/BAD classification came only from the work-box with its RX480/Polaris10
> GPU under real RustDesk H.264 VAAPI use.
> 
> The endpoints reproduced as follows:
> 
>   v6.12.101  GOOD
>   v6.12.105  BAD
> 
> The complete git bisect log is attached. Git mechanically identified:
> 
>   52566c150cadb1a17df840dd25862a61dcf84bed is the first bad commit
>   drm/amdgpu: check ASPM on the dGPU host link
> 
> Its immediate parent, fa9624ad4d6d4b25e448dbc550ddc69e6f849640, tested GOOD
> during the bisection. I subsequently qualified that exact parent in two real
> RustDesk H.264 VAAPI sessions totalling about 43 minutes, including a
> deliberate disconnect/reconnect. All 288 periodic samples and the complete
> kernel journal were free of the characteristic timeout/reset signatures. It
> is now the pinned operational kernel on this work-box, and two consecutive
> normal boots plus a further 10-minute H.264 VAAPI canary were clean.

Thanks for going through that!
 
> There is an important causal qualification. I also built a clean v6.12.105
> tree with only 52566c150cad... reverted. That kernel remained BAD: after
> about 79 seconds of real RustDesk H.264 VAAPI activity, the work-box
> reproduced the same sequence:
> 
>   ring vce0 timeout
>   GPU reset begin!
>   VRAM is lost due to GPU reset!

:-(. So 52566c150cad wasn't the real (or single) reason for the
regression you see. If I understand correctly the symptoms on
52566c150cad were also slightly different.

Can you please test 5f46322e0b84af29e10eb951ff45bd6ea40640de,
869eb255e499112f9ffb42e7fa96ff2fca0df54b and
81b5af1fb0f14cace6c3b3130a05e5602a597820 with 52566c150cad reverted?

(That is

	git checkout 5f46322e0b84af29e10eb951ff45bd6ea40640de,
	git revert 52566c150cad
	... build and test according to Salvatore's instructions

and the same for the other two commits.)

I think with these results we can report upstream and let the amdgpu
experts work out the rest (or ask smarter questions).

Best regards
Uwe

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web