Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #82984 > unrolled thread

Bug#1076309: [s390x] lots of "User process fault: interruption code XXXX"

Started byPaul Gevers <elbrus@debian.org>
First post2024-07-14 09:40 +0200
Last post2024-09-30 21:10 +0200
Articles 5 — 3 participants

Back to article view | Back to linux.debian.kernel


Contents

  Bug#1076309: [s390x] lots of "User process fault: interruption code XXXX" Paul Gevers <elbrus@debian.org> - 2024-07-14 09:40 +0200
    Bug#1076309: [s390x] lots of "User process fault: interruption code XXXX" Bastian Blank <waldi@debian.org> - 2024-07-24 11:00 +0200
    Bug#1076309: [s390x] lots of "User process fault: interruption code XXXX" Bastian Blank <waldi@debian.org> - 2024-07-24 11:10 +0200
      Bug#1076309: [s390x] lots of "User process fault: interruption code XXXX" Paul Gevers <elbrus@debian.org> - 2024-07-24 21:10 +0200
    Processed: Re: [s390x] lots of "User process fault: interruption  code XXXX" "Debian Bug Tracking System" <owner@bugs.debian.org> - 2024-09-30 21:10 +0200

#82984 — Bug#1076309: [s390x] lots of "User process fault: interruption code XXXX"

FromPaul Gevers <elbrus@debian.org>
Date2024-07-14 09:40 +0200
SubjectBug#1076309: [s390x] lots of "User process fault: interruption code XXXX"
Message-ID<J01ZT-1ZVg-1@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Package: src:linux
Version: 6.7.12-1~bpo12+1
X-Debbugs-CC: debian-ci@lists.debian.org, debian-s390@lists.debian.org
Severity: normal
X-Debbugs-Cc: elbrus@debian.org

Hi linux maintainers,

Today I restarted the s390x host of ci.d.n because I lost access. I 
inspected the journal and noticed there were a lot kernel messages 
(several tens to hundreds per day) like "User process fault: 
interruption code 003b ilc:3 in my_kmcdump[2aa07980000+f000]". I 
attached the first block I found in the journal after the reboot.

Please let me know if you need more information.

Paul
PS: I checked with carnil before filing this report, he thought it was 
worth reporting.

-- Package-specific info:
** Version:
Linux version 6.7.12+bpo-s390x (debian-kernel@lists.debian.org) 
(s390x-linux-gnu-gcc-12 (Debian 12.2.0-14) 12.2.0, GNU ld (GNU Binutils 
for Debian) 2.40) #1 SMP Debian 6.7.12-1~bpo12+1 (2024-05-06)

** Command line:
root=/dev/mapper/sysvg-root BOOT_IMAGE=0

** Not tainted

** Kernel log:
Unable to read kernel log; any relevant messages should be attached

** Model information
processor 0: version = FF,  identification = 010000,  machine = 8561
processor 1: version = FF,  identification = 020000,  machine = 8561
processor 2: version = FF,  identification = 030000,  machine = 8561
processor 3: version = FF,  identification = 040000,  machine = 8561
processor 4: version = FF,  identification = 050000,  machine = 8561
processor 5: version = FF,  identification = 060000,  machine = 8561
processor 6: version = FF,  identification = 070000,  machine = 8561
processor 7: version = FF,  identification = 080000,  machine = 8561
processor 8: version = FF,  identification = 090000,  machine = 8561
processor 9: version = FF,  identification = 0A0000,  machine = 8561

** Loaded modules:
nfsd
auth_rpcgss
nfs_acl
lockd
grace
tls
sunrpc
tcp_diag
inet_diag
veth
nft_masq
nft_chain_nat
nf_nat
nf_conntrack
nf_defrag_ipv6
nf_defrag_ipv4
nf_tables
libcrc32c
nfnetlink
binfmt_misc
qeth_l2
bridge
s390_trng
prng
stp
ctr
llc
sg
xts
aes_s390
des_s390
libdes
sha512_s390
qeth
sha256_s390
sha1_s390
sha_common
vmur
ccwgroup
scsi_dh_alua
dm_service_time
dm_multipath
loop
configfs
ip_tables
x_tables
autofs4
ext4
crc16
mbcache
jbd2
crc32c_generic
sd_mod
t10_pi
crc64_rocksoft
crc64
crc_t10dif
crct10dif_generic
crct10dif_common
dm_mod
zfcp
scsi_transport_fc
dasd_eckd_mod
scsi_mod
chacha_s390
libchacha
dasd_fba_mod
dasd_mod
scsi_common

** PCI devices:

** USB devices:
not available


-- System Information:
Debian Release: 12.6
   APT prefers stable-updates
   APT policy: (990, 'stable-updates'), (990, 'stable-security'), (990, 
'stable'), (500, 'unstable'), (1, 'experimental')
Architecture: s390x

Kernel: Linux 6.7.12+bpo-s390x (SMP w/10 CPU threads)
Locale: LANG=C, LC_CTYPE=C.UTF-8 (charmap=UTF-8), LANGUAGE not set
Shell: /bin/sh linked to /usr/bin/dash
Init: systemd (via /run/systemd/system)
LSM: AppArmor: enabled

Versions of packages linux-image-6.7.12+bpo-s390x depends on:
ii  initramfs-tools [linux-initramfs-tool]  0.142
ii  kmod                                    30+20221128-1
ii  linux-base                              4.9

Versions of packages linux-image-6.7.12+bpo-s390x recommends:
ii  apparmor             3.0.8-3
ii  firmware-linux-free  20200122-1

Versions of packages linux-image-6.7.12+bpo-s390x suggests:
pn  debian-kernel-handbook  <none>
pn  linux-doc-6.7           <none>
ii  s390-tools              2.16.0-2

Versions of packages linux-image-6.7.12+bpo-s390x is related to:
pn  firmware-amd-graphics     <none>
pn  firmware-atheros          <none>
pn  firmware-bnx2             <none>
pn  firmware-bnx2x            <none>
pn  firmware-brcm80211        <none>
pn  firmware-cavium           <none>
pn  firmware-intel-sound      <none>
pn  firmware-intelwimax       <none>
pn  firmware-ipw2x00          <none>
pn  firmware-ivtv             <none>
pn  firmware-iwlwifi          <none>
pn  firmware-libertas         <none>
pn  firmware-linux-nonfree    <none>
pn  firmware-misc-nonfree     <none>
pn  firmware-myricom          <none>
pn  firmware-netxen           <none>
pn  firmware-qlogic           <none>
pn  firmware-realtek          <none>
pn  firmware-samsung          <none>
pn  firmware-siano            <none>
pn  firmware-ti-connectivity  <none>
pn  xen-hypervisor            <none>

-- no debconf information

[toc] | [next] | [standalone]


#83166

FromBastian Blank <waldi@debian.org>
Date2024-07-24 11:00 +0200
Message-ID<J3G0N-4kg8-1@gated-at.bofh.it>
In reply to#82984
Hi

On Sun, Jul 14, 2024 at 09:22:32AM +0200, Paul Gevers wrote:
> Today I restarted the s390x host of ci.d.n because I lost access. I
> inspected the journal and noticed there were a lot kernel messages (several
> tens to hundreds per day) like "User process fault: interruption code 003b
> ilc:3 in my_kmcdump[2aa07980000+f000]". I attached the first block I found
> in the journal after the reboot.

I see "003b" and "0007" as interruption codes.  All those are page
faults (and are reported to the user space process via SIGSEGV).  Aka
the process tries to access data that is not available in the page
table.

Why s390 decides to always dump them to the kernel log if the
signal is unhandled is currently over me.  There is similar code on
other architectures, and also enabled by default, but I frankly have
never seen it.

Right now I have no idea what this could tell.

Bastian

-- 
Four thousand throats may be cut in one night by a running man.
		-- Klingon Soldier, "Day of the Dove", stardate unknown

[toc] | [prev] | [next] | [standalone]


#83167

FromBastian Blank <waldi@debian.org>
Date2024-07-24 11:10 +0200
Message-ID<J3Gat-4kyJ-3@gated-at.bofh.it>
In reply to#82984
On Thu, Jul 18, 2024 at 08:06:46AM +0200, Paul Gevers wrote:
> However, that doesn't seem to work on our s390x host as it seems to freeze
> instead. Is this something known? Something I'm doing wrong (E.g. these
> options behaving differently on s390x)? Is this a s390x kernel bug?

This now points to a kernel bug.  Which requires a new kernel first for
further debugging.

What does "freeze" mean"?  Also no sysrq?

See https://www.kernel.org/doc/html/v5.3/s390/debugging390.html#sysrq,
but you need to enable that before.

Bastian

-- 
Not one hundred percent efficient, of course ... but nothing ever is.
		-- Kirk, "Metamorphosis", stardate 3219.8

[toc] | [prev] | [next] | [standalone]


#83183

FromPaul Gevers <elbrus@debian.org>
Date2024-07-24 21:10 +0200
Message-ID<J3Px8-3K8-13@gated-at.bofh.it>
In reply to#83167

[Multipart message — attachments visible in raw view] — view raw

Hi waldi,

On 24-07-2024 10:57 a.m., Bastian Blank wrote:
> On Thu, Jul 18, 2024 at 08:06:46AM +0200, Paul Gevers wrote:
>> However, that doesn't seem to work on our s390x host as it seems to freeze
>> instead. Is this something known? Something I'm doing wrong (E.g. these
>> options behaving differently on s390x)? Is this a s390x kernel bug?
> 
> This now points to a kernel bug.  Which requires a new kernel first for
> further debugging.
> 
> What does "freeze" mean"?  Also no sysrq?

With freeze I mean I have no connection anymore. And when I reboot, 
there's nothing in the journal since I lost connection. I don't know yet 
what sysrq means, so I'll look into that.

> See https://www.kernel.org/doc/html/v5.3/s390/debugging390.html#sysrq,
> but you need to enable that before.

Maybe tomorrow.

Paul

[toc] | [prev] | [next] | [standalone]


#84146 — Processed: Re: [s390x] lots of "User process fault: interruption code XXXX"

From"Debian Bug Tracking System" <owner@bugs.debian.org>
Date2024-09-30 21:10 +0200
SubjectProcessed: Re: [s390x] lots of "User process fault: interruption code XXXX"
Message-ID<JstWp-g07U-9@gated-at.bofh.it>
In reply to#82984
Processing control commands:

> found -1 6.10.6-1~bpo12+1
Bug #1076309 [src:linux] [s390x] lots of "User process fault: interruption code XXXX"
Marked as found in versions linux/6.10.6-1~bpo12+1.
> clone -1 -2
Bug #1076309 [src:linux] [s390x] lots of "User process fault: interruption code XXXX"
Bug 1076309 cloned as bug 1083062
> retitle -2 s390x fails with vm.panic_on_oom=1 & kernel.panic=10
Bug #1083062 [src:linux] [s390x] lots of "User process fault: interruption code XXXX"
Changed Bug title to 's390x fails with vm.panic_on_oom=1 & kernel.panic=10' from '[s390x] lots of "User process fault: interruption code XXXX"'.

-- 
1076309: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1076309
1083062: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1083062
Debian Bug Tracking System
Contact owner@bugs.debian.org with problems

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web