Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #83286 > unrolled thread

Bug#1076555: linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc

Started byBastian Blank <waldi@debian.org>
First post2024-07-29 13:50 +0200
Last post2024-09-07 18:20 +0200
Articles 6 — 4 participants

Back to article view | Back to linux.debian.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Bug#1076555: linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc Bastian Blank <waldi@debian.org> - 2024-07-29 13:50 +0200
    Bug#1076555: linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc Ben Hutchings <ben@decadent.org.uk> - 2024-09-02 01:10 +0200
      Processed: Re: Bug#1076555: linux-image-6.9.9-amd64: boot crash  RIP: 0010:kmem_cache_alloc "Debian Bug Tracking System" <owner@bugs.debian.org> - 2024-09-02 01:10 +0200
      Bug#1076555: linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc Patrice Duroux <patrice.duroux@gmail.com> - 2024-09-02 13:40 +0200
        Bug#1076555: linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc Ben Hutchings <ben@decadent.org.uk> - 2024-09-05 02:20 +0200
          Bug#1076555: linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc Patrice Duroux <patrice.duroux@gmail.com> - 2024-09-07 18:20 +0200

#83286 — Bug#1076555: linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc

FromBastian Blank <waldi@debian.org>
Date2024-07-29 13:50 +0200
SubjectBug#1076555: linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc
Message-ID<J5x33-1f10-5@gated-at.bofh.it>
Control: tags -1 moreinfo

On Thu, Jul 18, 2024 at 07:56:32PM +0200, Patrice Duroux wrote:
> I just faced this boot problem on my sid system for the first time
> since I updated the Linux kernel a few days ago.

Could you please provide an unfiltered kernel log?  The one you attached
does not even contain the full crash message.  Nor is it the first crash
(visible by the D in the taint string).

Regards,
Bastian

-- 
Beam me up, Scotty!

[toc] | [next] | [standalone]


#83804

FromBen Hutchings <ben@decadent.org.uk>
Date2024-09-02 01:10 +0200
Message-ID<Ji1RL-9jHB-5@gated-at.bofh.it>
In reply to#83286

[Multipart message — attachments visible in raw view] — view raw

Control: tag -1 moreinfo

On Mon, 29 Jul 2024 18:07:37 +0200 Patrice Duroux
<patrice.duroux@gmail.com> wrote:
> Hi Bastian,
> Sorry I probably did something wrong with gnome-logs. Let's go then
> with journalctl
> and here are attached the complete log files, or at least I hope so!
[...]

OK, so we have:

> WARNING: CPU: 11 PID: 541 at mm/page_alloc.c:4556 __alloc_pages+0x2a3/0x340

This means something tried to allocate a chunk of kernel memory that is
larger than the kernel page allocator supports.

[...]
>  acpi_ds_build_internal_buffer_obj+0xa5/0x170
>  acpi_ds_eval_data_object_operands+0x13a/0x140
>  acpi_ds_exec_end_op+0x440/0x500
>  acpi_ps_parse_loop+0xfb/0x6b0
>  acpi_ps_parse_aml+0x80/0x3d0
>  acpi_ps_execute_method+0x13f/0x270
>  acpi_ns_evaluate+0x128/0x2d0
>  acpi_evaluate_object+0x14d/0x2f0
>  __query_block+0x10a/0x1e0 [wmi]
>  wmi_query_block+0x88/0xd0 [wmi]
>  init_bios_attributes.part.0+0x55/0x2f0 [dell_wmi_sysman]
>  sysman_init+0x158/0xff0 [dell_wmi_sysman]
>  ? __pfx_sysman_init+0x10/0x10 [dell_wmi_sysman]
>  do_one_initcall+0x58/0x320
>  do_init_module+0x60/0x240
>  init_module_from_file+0x89/0xe0
>  idempotent_init_module+0x120/0x2b0
>  __x64_sys_finit_module+0x5e/0xb0

That happened during initialisation of the dell_wmi_sysman module.

[...]
> general protection fault, probably for non-canonical address 0x800771b66d9d7272: 0000 [#1] PREEMPT SMP NOPTI
> CPU: 11 PID: 521 Comm: (udev-worker) Tainted: G        W          6.9.10-amd64 #1  Debian 6.9.10-1
> Hardware name: Dell Inc. Precision 7540/0T2FXT, BIOS 1.32.0 04/01/2024
> RIP: 0010:kmem_cache_alloc_node+0xed/0x360
[...]
> general protection fault, probably for non-canonical address 0x800771b66d9d7272: 0000 [#2] PREEMPT SMP NOPTI
> CPU: 11 PID: 696 Comm: fsck.ext4 Tainted: G      D W          6.9.10-amd64 #1  Debian 6.9.10-1
> Hardware name: Dell Inc. Precision 7540/0T2FXT, BIOS 1.32.0 04/01/2024
> RIP: 0010:kmem_cache_alloc+0xd7/0x340
[...]

Then later on the kernel heap allocation crashes, suggesting memory
corruption.  Maybe related to the first failure, maybe not.

- There is a newer kernel version available in unstable now (6.10.6 or
maybe 6.10.7 by the time you read this).  Does that fix the issue?

- Do these same error messages appear on every boot?

- If you prevent the dell_wmi_sysman module loading, by adding
"blacklist=dell_wmi_sysman" to the kernel command line, do all of the
error messages stop appearing?

Ben.

-- 
Ben Hutchings
Computers are not intelligent.	They only think they are.

[toc] | [prev] | [next] | [standalone]


#83805 — Processed: Re: Bug#1076555: linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc

From"Debian Bug Tracking System" <owner@bugs.debian.org>
Date2024-09-02 01:10 +0200
SubjectProcessed: Re: Bug#1076555: linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc
Message-ID<Ji1RL-9jHB-7@gated-at.bofh.it>
In reply to#83804
Processing control commands:

> tag -1 moreinfo
Bug #1076555 [src:linux] linux-image-6.9.9-amd64: boot crash RIP: 0010:kmem_cache_alloc
Added tag(s) moreinfo.

-- 
1076555: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1076555
Debian Bug Tracking System
Contact owner@bugs.debian.org with problems

[toc] | [prev] | [next] | [standalone]


#83807

FromPatrice Duroux <patrice.duroux@gmail.com>
Date2024-09-02 13:40 +0200
Message-ID<JidzA-9r80-3@gated-at.bofh.it>
In reply to#83804
Hi Ben,

Le lun. 2 sept. 2024 à 00:59, Ben Hutchings <ben@decadent.org.uk> a écrit :
>
> Control: tag -1 moreinfo
>
> On Mon, 29 Jul 2024 18:07:37 +0200 Patrice Duroux
> <patrice.duroux@gmail.com> wrote:
> > Hi Bastian,
> > Sorry I probably did something wrong with gnome-logs. Let's go then
> > with journalctl
> > and here are attached the complete log files, or at least I hope so!
> [...]
>
> OK, so we have:
>
> > WARNING: CPU: 11 PID: 541 at mm/page_alloc.c:4556 __alloc_pages+0x2a3/0x340
>
> This means something tried to allocate a chunk of kernel memory that is
> larger than the kernel page allocator supports.
>
> [...]
> >  acpi_ds_build_internal_buffer_obj+0xa5/0x170
> >  acpi_ds_eval_data_object_operands+0x13a/0x140
> >  acpi_ds_exec_end_op+0x440/0x500
> >  acpi_ps_parse_loop+0xfb/0x6b0
> >  acpi_ps_parse_aml+0x80/0x3d0
> >  acpi_ps_execute_method+0x13f/0x270
> >  acpi_ns_evaluate+0x128/0x2d0
> >  acpi_evaluate_object+0x14d/0x2f0
> >  __query_block+0x10a/0x1e0 [wmi]
> >  wmi_query_block+0x88/0xd0 [wmi]
> >  init_bios_attributes.part.0+0x55/0x2f0 [dell_wmi_sysman]
> >  sysman_init+0x158/0xff0 [dell_wmi_sysman]
> >  ? __pfx_sysman_init+0x10/0x10 [dell_wmi_sysman]
> >  do_one_initcall+0x58/0x320
> >  do_init_module+0x60/0x240
> >  init_module_from_file+0x89/0xe0
> >  idempotent_init_module+0x120/0x2b0
> >  __x64_sys_finit_module+0x5e/0xb0
>
> That happened during initialisation of the dell_wmi_sysman module.
>
> [...]
> > general protection fault, probably for non-canonical address 0x800771b66d9d7272: 0000 [#1] PREEMPT SMP NOPTI
> > CPU: 11 PID: 521 Comm: (udev-worker) Tainted: G        W          6.9.10-amd64 #1  Debian 6.9.10-1
> > Hardware name: Dell Inc. Precision 7540/0T2FXT, BIOS 1.32.0 04/01/2024
> > RIP: 0010:kmem_cache_alloc_node+0xed/0x360
> [...]
> > general protection fault, probably for non-canonical address 0x800771b66d9d7272: 0000 [#2] PREEMPT SMP NOPTI
> > CPU: 11 PID: 696 Comm: fsck.ext4 Tainted: G      D W          6.9.10-amd64 #1  Debian 6.9.10-1
> > Hardware name: Dell Inc. Precision 7540/0T2FXT, BIOS 1.32.0 04/01/2024
> > RIP: 0010:kmem_cache_alloc+0xd7/0x340
> [...]
>
> Then later on the kernel heap allocation crashes, suggesting memory
> corruption.  Maybe related to the first failure, maybe not.
>
> - There is a newer kernel version available in unstable now (6.10.6 or
> maybe 6.10.7 by the time you read this).  Does that fix the issue?

Since then, I have upgraded and am now using the 6.11.x series from experimental
without experiencing this issue anymore.

> - Do these same error messages appear on every boot?

Often, but not always. I was also not sure if it could be related to #1076561
that I had seen pass by at the same time.

> - If you prevent the dell_wmi_sysman module loading, by adding
> "blacklist=dell_wmi_sysman" to the kernel command line, do all of the
> error messages stop appearing?

Would you like me to reinstall one of those kernel versions so as to
reproduce this issue
and try adding the option?

Many thanks,
Patrice

[toc] | [prev] | [next] | [standalone]


#83839

FromBen Hutchings <ben@decadent.org.uk>
Date2024-09-05 02:20 +0200
Message-ID<Jj8o9-a0WR-9@gated-at.bofh.it>
In reply to#83807

[Multipart message — attachments visible in raw view] — view raw

On Mon, 2024-09-02 at 13:33 +0200, Patrice Duroux wrote:
> Hi Ben,
> 
> Le lun. 2 sept. 2024 à 00:59, Ben Hutchings <ben@decadent.org.uk> a écrit :
[...]
> > - There is a newer kernel version available in unstable now (6.10.6 or
> > maybe 6.10.7 by the time you read this).  Does that fix the issue?
> 
> Since then, I have upgraded and am now using the 6.11.x series from experimental
> without experiencing this issue anymore.

That's good news!  But 6.11 will not reach unstable for a few weeks.

> > - Do these same error messages appear on every boot?
> 
> Often, but not always. I was also not sure if it could be related to #1076561
> that I had seen pass by at the same time.

I don't see any connection with that bug.

> > - If you prevent the dell_wmi_sysman module loading, by adding
> > "blacklist=dell_wmi_sysman" to the kernel command line, do all of the
> > error messages stop appearing?
> 
> Would you like me to reinstall one of those kernel versions so as to
> reproduce this issue
> and try adding the option?

Yes please.

Ben.

-- 
Ben Hutchings
Theory and practice are closer in theory than in practice - John Levine

[toc] | [prev] | [next] | [standalone]


#83868

FromPatrice Duroux <patrice.duroux@gmail.com>
Date2024-09-07 18:20 +0200
Message-ID<Jk6kh-aFHK-1@gated-at.bofh.it>
In reply to#83839
Hi Ben,

All my attempts both after installing 6.9.9 on my current Sid system,
or using a snapshot of Sid at that time when linux-image-6.9.9 was pushed
(fastidious by the way) have failed. I am not able to get the issue again.

But since then it occurs to me that there was also a firmware update (fwupd)
for my Dell laptop system (1.32.0 to 1.34.0)

$ LC_ALL="C.UTF-8" fwupdmgr get-history
Dell Inc. Precision 7540
│
└─System Firmware:
  │   Device ID:          591f43f1e1ff1092daa9c95dd2b19488a82ae9e4
  │   Previous version:   1.32.0
  │   Update State:       Success
  │   Last modified:      2024-08-24 11:04
  │   GUID:               74992bad-d49d-4f82-bb77-bbe865bea34e
  │   Device Flags:       • Internal device
  │                       • Updatable
  │                       • System requires external power source
  │                       • Supported on remote server
  │                       • Needs a reboot after installation
  │                       • Device is usable for the duration of the update
  │
  └─Precision 7X40 System Update:
        New version:      1.34.0
        Remote ID:        lvfs
        Release ID:       95810
        Summary:          Firmware for the Dell Precision 7X40
        License:          Proprietary
        Size:             22.5 MB
        Created:          2024-07-11
        Urgency:          Critical
        Vendor:           Dell
        Release Flags:    • Trusted metadata
        Description:
        This stable release fixes the following issues:

        • This release contains security updates as disclosed in the
Dell Security Advisory.
        Issue:            DSA-2024-243
        Checksum:
6a5fa4d418e93826767d42405c2a7ad8236647fa4fe85d570d3ae1ccc5dafafb

Could this actually be a fix?

Regards,
Patrice

Le jeu. 5 sept. 2024 à 02:14, Ben Hutchings <ben@decadent.org.uk> a écrit :
>
> On Mon, 2024-09-02 at 13:33 +0200, Patrice Duroux wrote:
> > Hi Ben,
> >
> > Le lun. 2 sept. 2024 à 00:59, Ben Hutchings <ben@decadent.org.uk> a écrit :
> [...]
> > > - There is a newer kernel version available in unstable now (6.10.6 or
> > > maybe 6.10.7 by the time you read this).  Does that fix the issue?
> >
> > Since then, I have upgraded and am now using the 6.11.x series from experimental
> > without experiencing this issue anymore.
>
> That's good news!  But 6.11 will not reach unstable for a few weeks.
>
> > > - Do these same error messages appear on every boot?
> >
> > Often, but not always. I was also not sure if it could be related to #1076561
> > that I had seen pass by at the same time.
>
> I don't see any connection with that bug.
>
> > > - If you prevent the dell_wmi_sysman module loading, by adding
> > > "blacklist=dell_wmi_sysman" to the kernel command line, do all of the
> > > error messages stop appearing?
> >
> > Would you like me to reinstall one of those kernel versions so as to
> > reproduce this issue
> > and try adding the option?
>
> Yes please.
>
> Ben.
>
> --
> Ben Hutchings
> Theory and practice are closer in theory than in practice - John Levine
>

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web