Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #85312 > unrolled thread

Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs

Started bySalvatore Bonaccorso <carnil@debian.org>
First post2025-01-29 16:00 +0100
Last post2025-03-11 13:00 +0100
Articles 12 — 4 participants

Back to article view | Back to linux.debian.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Salvatore Bonaccorso <carnil@debian.org> - 2025-01-29 16:00 +0100
    Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Cordell Bloor <cgmb@slerp.xyz> - 2025-02-02 00:40 +0100
      Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Dieter Faulbaum <dieter@faulbaum.in-berlin.de> - 2025-02-02 11:30 +0100
        Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Dieter Faulbaum <dieter@faulbaum.in-berlin.de> - 2025-02-03 11:00 +0100
        Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Uwe Kleine-König <ukleinek@debian.org> - 2025-02-11 22:30 +0100
          Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Dieter Faulbaum <dieter@faulbaum.in-berlin.de> - 2025-02-12 11:40 +0100
            Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Dieter Faulbaum <dieter@faulbaum.in-berlin.de> - 2025-02-12 12:10 +0100
              Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Uwe Kleine-König <ukleinek@debian.org> - 2025-02-12 16:30 +0100
                Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Dieter Faulbaum <dieter@faulbaum.in-berlin.de> - 2025-02-12 18:10 +0100
                  Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Uwe Kleine-König <ukleinek@debian.org> - 2025-02-12 18:40 +0100
                    Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Dieter Faulbaum <dieter@faulbaum.in-berlin.de> - 2025-02-12 19:40 +0100
              Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs Dieter Faulbaum <dieter@faulbaum.in-berlin.de> - 2025-03-11 13:00 +0100

#85312 — Bug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs

FromSalvatore Bonaccorso <carnil@debian.org>
Date2025-01-29 16:00 +0100
SubjectBug#1093124: libhsa-runtime64-1: HSA exception: Queue create failed at hsaKmtCreateQueue with multiple programs
Message-ID<KahHP-cRx9-3@gated-at.bofh.it>
Control: tags -1 + moreinfo

Hi,

On Wed, Jan 29, 2025 at 05:16:23AM -0700, Cordell Bloor wrote:
> Control: reassign -1 src:linux 6.12.3-1
> Control: notfound linux/6.11.10-1
> Control: affects libhsa-runtime64-1
> 
> The HSA failures in OpenCL on Polaris 20 hardware (e.g. the Radeon RX 570)
> appear to be due to regression from Linux 6.11 to Linux 6.12. It works fine
> on 6.11.10-1.
> 
> I hope I've done the BTS control right.

It was mentioned to see some errors in dmesg output, can you attach
the full log please to the bug?

Would you be able to bisect the changes upstream between 6.11 and 6.12
to identify the breaking commit?

Can you actually test as well the newest version from 6.12.y in
unstable? Does this expose the problem as well?

Regards,
Salvatore

[toc] | [next] | [standalone]


#85349

FromCordell Bloor <cgmb@slerp.xyz>
Date2025-02-02 00:40 +0100
Message-ID<KbvfH-dFdg-1@gated-at.bofh.it>
In reply to#85312
Hi Salvatore,

Thanks for looking at this.

On 2025-01-29 07:54, Salvatore Bonaccorso wrote:
> It was mentioned to see some errors in dmesg output, can you attach
> the full log please to the bug?
>
> Would you be able to bisect the changes upstream between 6.11 and 6.12
> to identify the breaking commit?

Perhaps Dieter can give it a shot. I should be focused on getting my 
packages ready for the Trixie freeze.

> Can you actually test as well the newest version from 6.12.y in
> unstable? Does this expose the problem as well?

Dieter mentioned observing the same problem with 6.12.10. Though, I only 
just now noticed that it didn't make it onto the bug tracker:

On 2025-01-28 02:35, Dieter Faulbaum wrote:
> a little update: in the meantime a newer kernel (6.12.10) landed in 
> sid, but that doesn't help.
Sincerely,
Cory Bloor

[toc] | [prev] | [next] | [standalone]


#85352

FromDieter Faulbaum <dieter@faulbaum.in-berlin.de>
Date2025-02-02 11:30 +0100
Message-ID<KbFoJ-dMQU-1@gated-at.bofh.it>
In reply to#85349
Hello Cory and Salvatore,

Cordell Bloor <cgmb@slerp.xyz> writes:

>> Would you be able to bisect the changes upstream between 6.11 
>> and 6.12
>> to identify the breaking commit?
I think, that it works with kernels in testing/sid before more 
than 4 weeks ago.
(3 weeks ago I issued this at darktable first).

If I understand the list 
https://deb.debian.org/debian/pool/main/l/linux-signed-amd64/ 
right,
the last linux-image-6.11.* is from 2024-12-22 
(linux-image-6.11.10+bpo-amd64_6.11.10-1~bpo12+1_amd64.deb).
I'm relatively sure that with this version it works yet.

> Perhaps Dieter can give it a shot. I should be focused on 
> getting my packages ready for the Trixie freeze.
Sorry I'm not really sure, what I can do to find the breaking 
commit.
I can download the version 6.11.10 and can try it out (if you 
think this would be useful).

At the moment I "haven't" any kernels 6.11.* (and no sources 
"normally").
But with newest kernel 6.12.11 (from testing) the problem still 
exists.


Dieter

[toc] | [prev] | [next] | [standalone]


#85363

FromDieter Faulbaum <dieter@faulbaum.in-berlin.de>
Date2025-02-03 11:00 +0100
Message-ID<Kc1pf-e2tK-3@gated-at.bofh.it>
In reply to#85352
Hello Cory and Salvatore,

I (installed) and checked the kernel linux-image-6.11.10 and yes 
it works with this kernel.
But this has Cory done before (I think).

Dieter

[toc] | [prev] | [next] | [standalone]


#85483

FromUwe Kleine-König <ukleinek@debian.org>
Date2025-02-11 22:30 +0100
Message-ID<Kf5Zn-g4Cc-23@gated-at.bofh.it>
In reply to#85352

[Multipart message — attachments visible in raw view] — view raw

Hello Dieter,

On Sun, Feb 02, 2025 at 11:22:49AM +0100, Dieter Faulbaum wrote:
> Sorry I'm not really sure, what I can do to find the breaking commit.

Bisecting would work as follows:

 - Cloning the upstream source:

 	git clone https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
	cd linux
	git checkout v6.11

 - Reusing the Debian configuration and disable all modules not in use:

 	cp /boot/config-$(uname -r) .config
	make localmodconfig
	cp .config arch/x86/configs/my_defconfig

 - Building a kernel package:

 	make -j10 my_defconfig bindeb-pkg

The last step creates a kernel package. Install it and test if the issue
occurs. For 6.11 I would expect this to work (but in general it's good
to make sure one's assumptions are good, so please don't skip that
test).

Then do

	git checkout v6.12
	make -j10 my_defconfig bindeb-pkg

which creates another kernel package. Test this in the same way.
If this shows the problems you're set for a bisection:

 - Start the bisection:

 	git bisect start v6.12 v6.11

 - Build and test the version that git suggested:

 	make -j10 my_defconfig bindeb-pkg
	... install + boot

 - Depending on the outcome of your test do either

 	git bisect good

   (if the tested kernel package doesn't show the problem) or

   	git bisect bad

   and repeat starting at "Build and test the version that git
   suggested" until git identfied the first bad commit.

I skipped some steps (like maybe you need to install build dependencies
and how you install the resulting package). Also you probably need to
deinstall the test packages again to not fill your /boot partition (if
you have one). If you still have questions, don't hesitate to ask.

Best regards
Uwe

[toc] | [prev] | [next] | [standalone]


#85489

FromDieter Faulbaum <dieter@faulbaum.in-berlin.de>
Date2025-02-12 11:40 +0100
Message-ID<KfijT-gdRw-15@gated-at.bofh.it>
In reply to#85483
Hello Uwe,

thanks for the intro (it's a long time ago that I compiled my last 
kernel.-)

Uwe Kleine-König <ukleinek@debian.org> writes:
...
>  - Building a kernel package:
>
>  	make -j10 my_defconfig bindeb-pkg
>
> The last step creates a kernel package. Install it and test if 
> the issue
> occurs. For 6.11 I would expect this to work (but in general 
> it's good
> to make sure one's assumptions are good, so please don't skip 
> that
> test).

I failed at this step.-(
The last lines are:
  LD [M]  net/vmw_vsock/vsock.ko
  LD [M]  net/vmw_vsock/vmw_vsock_virtio_transport_common.ko
make[4]: *** [debian/rules:74: build-arch] Error 2
dpkg-buildpackage: error: make -f debian/rules binary subprocess 
returned exit status 2
make[3]: *** [scripts/Makefile.package:121: bindeb-pkg] Error 2
make[2]: *** [Makefile:1547: bindeb-pkg] Error 2
make[1]: *** [/home/dieter/tmp/linux/Makefile:347: 
__build_one_by_one] Error 2
make: *** [Makefile:224: __sub-make] Error 2

Do you have an idea what I made wrong (or are there any missing 
build dependencies)?

The steps before looked good (for me), only some questions in the 
'make localmodconfig',
which I replied with a RET.

With regards
Dieter

[toc] | [prev] | [next] | [standalone]


#85490

FromDieter Faulbaum <dieter@faulbaum.in-berlin.de>
Date2025-02-12 12:10 +0100
Message-ID<KfiMV-gehJ-11@gated-at.bofh.it>
In reply to#85489
Hello Uwe,

sorry for the "too fast" question,
I found my fault (a missing dependency) myself:
There were these lines (a little bit "hidden" between the many 
"CC-lines"):

BTF: .tmp_vmlinux1: pahole (pahole) is not available
Failed to generate BTF for vmlinux
Try to disable CONFIG_DEBUG_INFO_BTF
make[6]: *** [scripts/Makefile.vmlinux:34: vmlinux] Error 1
make[5]: *** [Makefile:1157: vmlinux] Error 2

Salvatore mentioned this too (thank you).

After installing pahole, all went fine (but my CPU needs a long 
time for the build process,-)
Need some more time for the next steps.


With regards
Dieter

[toc] | [prev] | [next] | [standalone]


#85491

FromUwe Kleine-König <ukleinek@debian.org>
Date2025-02-12 16:30 +0100
Message-ID<KfmQx-ggPL-5@gated-at.bofh.it>
In reply to#85490
Hello Dieter,

On 2/12/25 12:05, Dieter Faulbaum wrote:
> sorry for the "too fast" question,
> I found my fault (a missing dependency) myself:
> There were these lines (a little bit "hidden" between the many "CC-lines"):
> 
> BTF: .tmp_vmlinux1: pahole (pahole) is not available
> Failed to generate BTF for vmlinux
> Try to disable CONFIG_DEBUG_INFO_BTF
> make[6]: *** [scripts/Makefile.vmlinux:34: vmlinux] Error 1
> make[5]: *** [Makefile:1157: vmlinux] Error 2
> 
> Salvatore mentioned this too (thank you).
> 
> After installing pahole, all went fine (but my CPU needs a long time for 
> the build process,-)

I think you can also disable some config item and then pahole isn't 
needed. Anyhow, you seem to cope, that's fine.

> Need some more time for the next steps.

That's expected, I'm not holding my breath :-) Thanks for going through 
that even though it's time-consuming. Just report back when (and if) 
you're through.

Best regards
Uwe

[toc] | [prev] | [next] | [standalone]


#85492

FromDieter Faulbaum <dieter@faulbaum.in-berlin.de>
Date2025-02-12 18:10 +0100
Message-ID<Kfopj-ghWz-1@gated-at.bofh.it>
In reply to#85491
Hello

Uwe Kleine-König <ukleinek@debian.org> writes:
...
> I think you can also disable some config item and then pahole 
> isn't needed. Anyhow, you seem to cope, that's fine.
>
>> Need some more time for the next steps.
>
> That's expected, I'm not holding my breath :-) Thanks for going 
> through that even though it's time-consuming. Just report back 
> when (and if) you're through.

I ran into problems at "stage" 10 (after creating and testing 9 
kernels) like these:

make -j16 my_defconfig bindeb-pkg
#
# No change to .config
#
  GEN     debian
dpkg-buildpackage --build=binary --no-pre-clean --unsigned-changes 
-R'make -f debian/rules' -j1 -a$(cat debian/arch)
dpkg-buildpackage: info: source package linux-upstream
dpkg-buildpackage: info: source version 
6.10.0-rc6-02741-g722e96c99f1d-11
dpkg-buildpackage: info: source distribution trixie
dpkg-buildpackage: info: source changed by Dieter Faulbaum 
<dieter@faulbaum.in-berlin.de>
dpkg-buildpackage: info: host architecture amd64
 dpkg-source --before-build .
 make -f debian/rules binary
#
# No change to .config
#
mkdir -p /home/dieter/tmp/linux/tools/objtool && make 
O=/home/dieter/tmp/linux subdir=tools/objtool --no-print-directory 
-C objtool 
mkdir -p /home/dieter/tmp/linux/tools/bpf/resolve_btfids && make 
O=/home/dieter/tmp/linux subdir=tools/bpf/resolve_btfids 
--no-print-directory -C bpf/resolve_btfids 
  INSTALL libsubcmd_headers
  INSTALL libsubcmd_headers
  CALL    scripts/checksyscalls.sh
  UPD     init/utsversion-tmp.h
  CC      init/version.o
  AR      init/built-in.a
  CC [M]  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_queue.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_device_queue_manager.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_device_queue_manager_cik.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_device_queue_manager_vi.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_device_queue_manager_v9.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_device_queue_manager_v10.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_device_queue_manager_v11.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_device_queue_manager_v12.o
  CC [M]  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_interrupt.o
  CC [M]  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_events.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/cik_event_interrupt.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_int_process_v9.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_int_process_v10.o
  CC [M] 
  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_int_process_v11.o
  CC [M]  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_smi_events.o
  CC [M]  drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_crat.o
drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_queue.c: In function 
‘kfd_queue_buffer_svm_get’:
drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_queue.c:108:26: error: 
implicit declaration of function ‘svm_range_from_addr’ 
[-Wimplicit-function-declaration]
  108 |                 prange = svm_range_from_addr(&p->svms, 
  addr, NULL);
      |                          ^~~~~~~~~~~~~~~~~~~
drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_queue.c:108:24: error: 
assignment to ‘struct svm_range *’ from ‘int’ makes pointer from 
integer without a cast [-Wint-conversion]
  108 |                 prange = svm_range_from_addr(&p->svms, 
  addr, NULL);
      |                        ^
drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_queue.c:112:28: error: 
invalid use of undefined type ‘struct svm_range’
  112 |                 if (!prange->mapped_to_gpu)
      |                            ^~
In file included from ./include/linux/kernel.h:23,
                 from ./include/linux/cpumask.h:11,
                 from ./arch/x86/include/asm/paravirt.h:21,
                 from ./arch/x86/include/asm/irqflags.h:60,
                 from ./include/linux/irqflags.h:18,
                 from ./include/linux/spinlock.h:59,
                 from ./include/linux/mmzone.h:8,
                 from ./include/linux/gfp.h:7,
                 from ./include/linux/slab.h:16,
                 from 
                 drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_queue.c:25:
drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_queue.c:118:45: error: 
invalid use of undefined type ‘struct svm_range’
  118 |                 if (!test_bit(gpuidx, 
  prange->bitmap_access) &&
      |                                             ^~
./include/linux/bitops.h:45:44: note: in definition of macro 
‘bitop’
   45 |           __builtin_constant_p((uintptr_t)(addr) != 
   (uintptr_t)NULL) && \
      |                                            ^~~~
drivers/gpu/drm/amd/amdgpu/../amdkfd/kfd_queue.c:118:22: note: in 
expansion of macro ‘test_bit’
  118 |                 if (!test_bit(gpuidx, 
  prange->bitmap_access) &&
      |                      ^~~~~~~~

There are more errors similar to these (is the full output 
needed?)
Any idea what I can do (better)?

With regards
Dieter

[toc] | [prev] | [next] | [standalone]


#85493

FromUwe Kleine-König <ukleinek@debian.org>
Date2025-02-12 18:40 +0100
Message-ID<KfoSl-gi8u-9@gated-at.bofh.it>
In reply to#85492

[Multipart message — attachments visible in raw view] — view raw

Hello Dieter,

On 2/12/25 18:05, Dieter Faulbaum wrote:
> There are more errors similar to these (is the full output needed?)
> Any idea what I can do (better)?

Try if

	git bisect skip

gets you out of that. It makes git select a different revision to test 
for you. If needed repeat it.

`git bisect skip` the thing to do in general if you cannot find out for 
a suggested revision if the problem is in there or not. Something like 
that shouldn't happen, but it does at times. That makes the bisection a 
bit ineffective, but typically it converges anyhow. Alternatively you 
can try to be smart about the selection of the next revision to test and 
just do git checkout of an alternative revision.

Best regards
Uwe

[toc] | [prev] | [next] | [standalone]


#85495

FromDieter Faulbaum <dieter@faulbaum.in-berlin.de>
Date2025-02-12 19:40 +0100
Message-ID<KfpOp-giK9-9@gated-at.bofh.it>
In reply to#85493
Hello Uwe,

Thank you for your patience and help!

Uwe Kleine-König <ukleinek@debian.org> writes:

> 	git bisect skip
>
> gets you out of that. It makes git select a different revision 
> to test for you. If needed repeat it.
>
> `git bisect skip` the thing to do in general if you cannot find 
> out for a suggested revision if the problem is in there or not. 
> Something like that shouldn't happen, but it does at times. That 
> makes the bisection a bit ineffective, but typically it 
> converges anyhow. Alternatively you can try to be smart about 
> the selection of the next revision to test and just do git 
> checkout of an alternative revision.

(after some skips) this should be the "bad" one:

68e599db7a549f010a329515f3508d8a8c3467a4 is the first bad commit
commit 68e599db7a549f010a329515f3508d8a8c3467a4 (HEAD)
Author: Philip Yang <Philip.Yang@amd.com>
Date:   Thu Jun 20 12:21:57 2024 -0400

    drm/amdkfd: Validate user queue buffers
    
    Find user queue rptr, ring buf, eop buffer and cwsr area BOs, 
    and
    check BOs are mapped on the GPU with correct size and take the 
    BO
    reference.
    
    Signed-off-by: Philip Yang <Philip.Yang@amd.com>
    Reviewed-by: Felix Kuehling <felix.kuehling@amd.com>
    Acked-by: Christian König <christian.koenig@amd.com>
    Signed-off-by: Alex Deucher <alexander.deucher@amd.com>

 drivers/gpu/drm/amd/amdkfd/kfd_priv.h  |  4 ++++
 drivers/gpu/drm/amd/amdkfd/kfd_queue.c | 38 
 ++++++++++++++++++++++++++++++++++++--
 2 files changed, 40 insertions(+), 2 deletions(-)


With regards
Dieter

[toc] | [prev] | [next] | [standalone]


#86416

FromDieter Faulbaum <dieter@faulbaum.in-berlin.de>
Date2025-03-11 13:00 +0100
Message-ID<Kp6r7-5f0r-1@gated-at.bofh.it>
In reply to#85490
Hello to all who are involved,

there is a newer kernel (6.10.17) in trixie but it seems for me, 
that the patch
https://lore.kernel.org/all/20250130000412.29812-1-Philip.Yang@amd.com/T/
from Philip Yang hasn't found its way into this kernel.
The problem still exists.

With regards
Dieter

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web