Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #82499 > unrolled thread

Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate

Started byRichard Kojedzinszky <richard+debian+bugreport@kojedz.in>
First post2024-05-20 11:40 +0200
Last post2024-06-28 00:20 +0200
Articles 9 — 5 participants

Back to article view | Back to linux.debian.kernel


Contents

  Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> - 2024-05-20 11:40 +0200
    Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate Salvatore Bonaccorso <carnil@debian.org> - 2024-05-20 21:10 +0200
      Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate Diederik de Haas <didi.debian@cknow.org> - 2024-05-20 23:40 +0200
      Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> - 2024-05-21 10:10 +0200
    Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> - 2024-05-23 14:20 +0200
    Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> - 2024-05-24 07:40 +0200
      Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> - 2024-05-25 18:10 +0200
    Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard@kojedz.in> - 2024-05-27 06:10 +0200
    Bug#1071501: marked as done (linux-image-6.1.0-21-arm64: Linux  NFS client hangs in nfs4_lookup_revalidate) "Debian Bug Tracking System" <owner@bugs.debian.org> - 2024-06-28 00:20 +0200

#82499 — Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate

FromRichard Kojedzinszky <richard+debian+bugreport@kojedz.in>
Date2024-05-20 11:40 +0200
SubjectBug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate
Message-ID<IG7ER-eBme-1@gated-at.bofh.it>
Package: src:linux
Version: 6.1.90-1
Severity: normal
X-Debbugs-Cc: richard+debian+bugreport@kojedz.in

Dear Maintainer,

I am running kubernetes on debian, and pods are mounting multiple nfs
shares. I am running dovecot processes in PODs, which receive mails from
the internet, and also serves as imap server for clients. I am
monitoring my mail system by sending mails periodically (15 seconds) and
also downloading them via imap. I found a few times that some dovecot process
stuck in D state, a reboot was always needed to recover from that state.

Unfortunately, I was not able to trigger the bug really fast, I dont
really know what operations does dovecot issue and in what order to trigger
this behavior. So until I get closer, I've set up a similar, but smaller
environment with just a single dovecot process, and it also does the
same work, delivering only test mails locally, and serving them via imap
to the monitoring client, storing everything on NFS. Fortunately, this also
triggers the bug, after a few hours one of the dovecot processes is stuck
in D state. Kernel also shows blocked state:

May 19 12:16:49 k8s-node07 kernel: INFO: task lmtp:665683 blocked for more than 120 seconds.
May 19 12:16:49 k8s-node07 kernel:       Not tainted 6.1.0-21-arm64 #1 Debian 6.1.90-1
May 19 12:16:49 k8s-node07 kernel: "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
May 19 12:16:49 k8s-node07 kernel: task:lmtp            state:D stack:0     pid:665683 ppid:2881   flags:0x00000000
May 19 12:16:49 k8s-node07 kernel: Call trace:
May 19 12:16:49 k8s-node07 kernel:  __switch_to+0xf0/0x170
May 19 12:16:49 k8s-node07 kernel:  __schedule+0x340/0x940
May 19 12:16:49 k8s-node07 kernel:  schedule+0x58/0xf0
May 19 12:16:49 k8s-node07 kernel:  __nfs_lookup_revalidate+0x118/0x160 [nfs]
May 19 12:16:49 k8s-node07 kernel:  nfs4_lookup_revalidate+0x20/0x30 [nfs]
May 19 12:16:49 k8s-node07 kernel:  lookup_fast+0x138/0x150
May 19 12:16:49 k8s-node07 kernel:  walk_component+0x30/0x1a0
May 19 12:16:49 k8s-node07 kernel:  path_lookupat+0x80/0x1a4
May 19 12:16:49 k8s-node07 kernel:  filename_lookup+0xb4/0x1b0
May 19 12:16:49 k8s-node07 kernel:  vfs_statx+0x94/0x19c
May 19 12:16:49 k8s-node07 kernel:  vfs_fstatat+0x68/0x90
May 19 12:16:49 k8s-node07 kernel:  __do_sys_newfstatat+0x58/0xa0
May 19 12:16:49 k8s-node07 kernel:  __arm64_sys_newfstatat+0x28/0x34
May 19 12:16:49 k8s-node07 kernel:  invoke_syscall+0x78/0x100
May 19 12:16:49 k8s-node07 kernel:  el0_svc_common.constprop.0+0x4c/0xf4
May 19 12:16:49 k8s-node07 kernel:  do_el0_svc+0x34/0xd0
May 19 12:16:49 k8s-node07 kernel:  el0_svc+0x34/0xd4
May 19 12:16:49 k8s-node07 kernel:  el0t_64_sync_handler+0xf4/0x120
May 19 12:16:49 k8s-node07 kernel:  el0t_64_sync+0x18c/0x190

Or, for another process:

May 20 04:50:01 k8s-node07 kernel: INFO: task imap:8337 blocked for more than 120 seconds.
May 20 04:50:01 k8s-node07 kernel:       Not tainted 6.1.0-21-arm64 #1 Debian 6.1.90-1
May 20 04:50:01 k8s-node07 kernel: "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
May 20 04:50:01 k8s-node07 kernel: task:imap            state:D stack:0     pid:8337  ppid:3164   flags:0x00000000
May 20 04:50:01 k8s-node07 kernel: Call trace:
May 20 04:50:01 k8s-node07 kernel:  __switch_to+0xf0/0x170
May 20 04:50:01 k8s-node07 kernel:  __schedule+0x340/0x940
May 20 04:50:01 k8s-node07 kernel:  schedule+0x58/0xf0
May 20 04:50:01 k8s-node07 kernel:  __nfs_lookup_revalidate+0x118/0x160 [nfs]
May 20 04:50:01 k8s-node07 kernel:  nfs4_lookup_revalidate+0x20/0x30 [nfs]
May 20 04:50:01 k8s-node07 kernel:  lookup_fast+0x138/0x150
May 20 04:50:01 k8s-node07 kernel:  walk_component+0x30/0x1a0
May 20 04:50:01 k8s-node07 kernel:  path_lookupat+0x80/0x1a4
May 20 04:50:01 k8s-node07 kernel:  filename_lookup+0xb4/0x1b0
May 20 04:50:01 k8s-node07 kernel:  vfs_statx+0x94/0x19c
May 20 04:50:01 k8s-node07 kernel:  vfs_fstatat+0x68/0x90
May 20 04:50:01 k8s-node07 kernel:  __do_sys_newfstatat+0x58/0xa0
May 20 04:50:01 k8s-node07 kernel:  __arm64_sys_newfstatat+0x28/0x34
May 20 04:50:01 k8s-node07 kernel:  invoke_syscall+0x78/0x100
May 20 04:50:01 k8s-node07 kernel:  el0_svc_common.constprop.0+0x4c/0xf4
May 20 04:50:01 k8s-node07 kernel:  do_el0_svc+0x34/0xd0
May 20 04:50:01 k8s-node07 kernel:  el0_svc+0x34/0xd4
May 20 04:50:01 k8s-node07 kernel:  el0t_64_sync_handler+0xf4/0x120
May 20 04:50:01 k8s-node07 kernel:  el0t_64_sync+0x18c/0x190


Of course the NFS server is running, and other NFS mounts are still
working from the node. Also, this started to happen with Debian's
kernel. Before that, I was compiling my own upstream kernel, version
5.15. With that, I've never experienced such a lockup.

Unfortunately, I dont know, how to go further, how shall I collect more
relevant debugging information.

I expect thet dovecot is just an application, which should not cause any
kernel-side lockups. In my test lab, this specific NFS mount is just
mounted on one machine, so it really suggests me a linux nfs-client side
issue, not related to caching coherency between multiple clients.

-- Package-specific info:
** Version:
Linux version 6.1.0-21-arm64 (debian-kernel@lists.debian.org) (gcc-12 (Debian 12.2.0-14) 12.2.0, GNU ld (GNU Binutils for Debian) 2.40) #1 SMP Debian 6.1.90-1 (2024-05-03)

** Command line:
net.ifnames=0 console=ttyS2,1500000 console=tty1 root=UUID=b4ff4167-1fe9-4fd6-9b9c-c3c68d98108b rw rootwait panic=10

** Not tainted

** Kernel log:
May 20 04:52:02 k8s-node07 kernel: INFO: task imap:8337 blocked for more than 241 seconds.
May 20 04:52:02 k8s-node07 kernel:       Not tainted 6.1.0-21-arm64 #1 Debian 6.1.90-1
May 20 04:52:02 k8s-node07 kernel: "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message.
May 20 04:52:02 k8s-node07 kernel: task:imap            state:D stack:0     pid:8337  ppid:3164   flags:0x00000000
May 20 04:52:02 k8s-node07 kernel: Call trace:
May 20 04:52:02 k8s-node07 kernel:  __switch_to+0xf0/0x170
May 20 04:52:02 k8s-node07 kernel:  __schedule+0x340/0x940
May 20 04:52:02 k8s-node07 kernel:  schedule+0x58/0xf0
May 20 04:52:02 k8s-node07 kernel:  __nfs_lookup_revalidate+0x118/0x160 [nfs]
May 20 04:52:02 k8s-node07 kernel:  nfs4_lookup_revalidate+0x20/0x30 [nfs]
May 20 04:52:02 k8s-node07 kernel:  lookup_fast+0x138/0x150
May 20 04:52:02 k8s-node07 kernel:  walk_component+0x30/0x1a0
May 20 04:52:02 k8s-node07 kernel:  path_lookupat+0x80/0x1a4
May 20 04:52:02 k8s-node07 kernel:  filename_lookup+0xb4/0x1b0
May 20 04:52:02 k8s-node07 kernel:  vfs_statx+0x94/0x19c
May 20 04:52:02 k8s-node07 kernel:  vfs_fstatat+0x68/0x90
May 20 04:52:02 k8s-node07 kernel:  __do_sys_newfstatat+0x58/0xa0
May 20 04:52:02 k8s-node07 kernel:  __arm64_sys_newfstatat+0x28/0x34
May 20 04:52:02 k8s-node07 kernel:  invoke_syscall+0x78/0x100
May 20 04:52:02 k8s-node07 kernel:  el0_svc_common.constprop.0+0x4c/0xf4
May 20 04:52:02 k8s-node07 kernel:  do_el0_svc+0x34/0xd0
May 20 04:52:02 k8s-node07 kernel:  el0_svc+0x34/0xd4
May 20 04:52:02 k8s-node07 kernel:  el0t_64_sync_handler+0xf4/0x120
May 20 04:52:02 k8s-node07 kernel:  el0t_64_sync+0x18c/0x190

** Model information

** Loaded modules:
sd_mod
t10_pi
crc64_rocksoft_generic
crc64_rocksoft
crc_t10dif
crct10dif_generic
crc64
sg
iscsi_tcp
libiscsi_tcp
libiscsi
scsi_transport_iscsi
scsi_mod
scsi_common
nf_conntrack_netlink
rpcsec_gss_krb5
auth_rpcgss
nfsv4
dns_resolver
nfs
lockd
grace
fscache
netfs
nft_log
nft_limit
xt_limit
xt_NFLOG
nfnetlink_log
xt_physdev
xt_TCPMSS
xt_tcpudp
xt_mark
xt_multiport
xt_addrtype
dummy
ipt_REJECT
nf_reject_ipv4
ip_set_hash_ipport
nft_chain_nat
xt_nat
xt_MASQUERADE
xt_ipvs
nf_nat
xt_set
ip_set_hash_ip
ip_set_hash_net
ip_set
veth
xt_conntrack
xt_comment
nft_compat
nf_tables
nfnetlink
overlay
sunrpc
binfmt_misc
evdev
aes_ce_blk
snd_soc_rk817
aes_ce_cipher
polyval_ce
snd_soc_core
polyval_generic
snd_pcm_dmaengine
ext4
ghash_ce
gf128mul
sha2_ce
leds_gpio
snd_pcm
sha256_arm64
sha1_ce
rockchip_thermal
crc16
mbcache
snd_timer
jbd2
snd
dw_wdt
soundcore
rk817_charger
rk805_pwrkey
cpufreq_dt
br_netfilter
bridge
stp
llc
ip_vs_sh
ip_vs_wrr
ip_vs_rr
ip_vs
nf_conntrack
nf_defrag_ipv6
nf_defrag_ipv4
drm
loop
fuse
efi_pstore
dm_mod
dax
configfs
ip_tables
x_tables
autofs4
xfs
libcrc32c
crc32c_generic
realtek
rk808_regulator
fan53555
dwmac_rk
stmmac_platform
stmmac
pcs_xpcs
spi_rockchip
phylink
dw_mmc_rockchip
dw_mmc_pltfm
of_mdio
dw_mmc
fixed
crct10dif_ce
crct10dif_common
fixed_phy
fwnode_mdio
pl330
i2c_rk3x
io_domain
libphy

** PCI devices:
not available

** USB devices:
not available


-- System Information:
Debian Release: 12.5
  APT prefers stable-updates
  APT policy: (500, 'stable-updates'), (500, 'stable-security'), (500, 'stable')
Architecture: arm64 (aarch64)

Kernel: Linux 6.1.0-21-arm64 (SMP w/4 CPU threads)
Locale: LANG=C, LC_CTYPE=C.UTF-8 (charmap=UTF-8), LANGUAGE not set
Shell: /bin/sh linked to /usr/bin/dash
Init: unable to detect

Versions of packages linux-image-6.1.0-21-arm64 depends on:
ii  initramfs-tools [linux-initramfs-tool]  0.142
ii  kmod                                    30+20221128-1
ii  linux-base                              4.9

Versions of packages linux-image-6.1.0-21-arm64 recommends:
ii  apparmor             3.0.8-3
ii  firmware-linux-free  20200122-1

Versions of packages linux-image-6.1.0-21-arm64 suggests:
pn  debian-kernel-handbook  <none>
pn  linux-doc-6.1           <none>

Versions of packages linux-image-6.1.0-21-arm64 is related to:
pn  firmware-amd-graphics     <none>
pn  firmware-atheros          <none>
pn  firmware-bnx2             <none>
pn  firmware-bnx2x            <none>
pn  firmware-brcm80211        <none>
pn  firmware-cavium           <none>
pn  firmware-intel-sound      <none>
pn  firmware-intelwimax       <none>
pn  firmware-ipw2x00          <none>
pn  firmware-ivtv             <none>
pn  firmware-iwlwifi          <none>
pn  firmware-libertas         <none>
pn  firmware-linux-nonfree    <none>
pn  firmware-misc-nonfree     <none>
pn  firmware-myricom          <none>
pn  firmware-netxen           <none>
pn  firmware-qlogic           <none>
pn  firmware-realtek          <none>
pn  firmware-samsung          <none>
pn  firmware-siano            <none>
pn  firmware-ti-connectivity  <none>
pn  xen-hypervisor            <none>

-- no debconf information

[toc] | [next] | [standalone]


#82512

FromSalvatore Bonaccorso <carnil@debian.org>
Date2024-05-20 21:10 +0200
Message-ID<IGgyu-eGN0-1@gated-at.bofh.it>
In reply to#82499
Hi Richard,

On Mon, May 20, 2024 at 09:27:24AM +0000, Richard Kojedzinszky wrote:
> Package: src:linux
> Version: 6.1.90-1
> Severity: normal
> X-Debbugs-Cc: richard+debian+bugreport@kojedz.in
> 
> Dear Maintainer,
> 
> I am running kubernetes on debian, and pods are mounting multiple nfs
> shares. I am running dovecot processes in PODs, which receive mails from
> the internet, and also serves as imap server for clients. I am
> monitoring my mail system by sending mails periodically (15 seconds) and
> also downloading them via imap. I found a few times that some dovecot process
> stuck in D state, a reboot was always needed to recover from that state.
> 
> Unfortunately, I was not able to trigger the bug really fast, I dont
> really know what operations does dovecot issue and in what order to trigger
> this behavior. So until I get closer, I've set up a similar, but smaller
> environment with just a single dovecot process, and it also does the
> same work, delivering only test mails locally, and serving them via imap
> to the monitoring client, storing everything on NFS. Fortunately, this also
> triggers the bug, after a few hours one of the dovecot processes is stuck
> in D state. Kernel also shows blocked state:

As you seem in the lucky position to be able to trigger the issue in a
more localized setup, might you:

- try as well more recent kernels from upper suites (6.8.9-1 in
  unstable would be ideal to check if the issue is there as well).
- I did read you cannot trigger with 5.15. If you build 6.1.90 from
  upstream without Debian patches I assume you can trigger the issue
  likewise? If so could you bisect the changes introducing the issue?
  This is a cumbersome process in particular if you need few hours to
  trigger it  So maybe the following point could be done first:
- Can you report the issue to the linux-nfs list, keeping us in the
  loop?

Regards,
Salvatore

[toc] | [prev] | [next] | [standalone]


#82514

FromDiederik de Haas <didi.debian@cknow.org>
Date2024-05-20 23:40 +0200
Message-ID<IGiTD-eI4n-3@gated-at.bofh.it>
In reply to#82512

[Multipart message — attachments visible in raw view] — view raw

On Monday, 20 May 2024 21:07:49 CEST Salvatore Bonaccorso wrote:
> - I did read you cannot trigger with 5.15. If you build 6.1.90 from
>   upstream without Debian patches I assume you can trigger the issue
>   likewise? If so could you bisect the changes introducing the issue?

If the test with the upstream 6.1.90 version also has this problem, there's 
another (series of) test(s) worth doing, which could shorten the bisect 
operation significantly.

I got the impression that you have only tried it with version 6.1.90.
Have you tried it with earlier versions in the 6.1 series to see if the issue 
is present there? 

Via https://snapshot.debian.org/package/linux-signed-arm64/ you can find 
earlier versions from the 6.1 series already compiled and packaged.
To take version 6.1.52 as (random) example:
- click on the ``6.1.52+1`` link
- In the ``Binary packages`` list, click on the linux-image-6.1.0-X-arm64 
link, where 'X' is 12 in this case
- Click the ``linux-image-6.1.0-12-arm64_6.1.52-1_arm64.deb`` link to download 
the deb file which you can then install as root or with sudo by executing
``apt install ./linux-image-<version-info>_arm64.deb``

If the problem does NOT occur with 6.1.52-1, then you try a higher version. 
Continue that process until you've found the latest version that works and the 
earliest version where it stopped working.

If the problem also occurs with 6.1.52-1, then you try an (even) older 
version.

This is to test whether it was a regression *within* the 6.1 series and if so, 
to get the narrowest range without having to compile yourself.

HTH

[toc] | [prev] | [next] | [standalone]


#82517

FromRichard Kojedzinszky <richard+debian+bugreport@kojedz.in>
Date2024-05-21 10:10 +0200
Message-ID<IGsJj-eOf0-1@gated-at.bofh.it>
In reply to#82512
Dear Salvatore,

I've already started bisecting. It will take some time. Usually the bug 
appears after a few hours, unfortunately I am not able to trigger it 
faster. So, if the bug appears, I can step forward easily, but if not, 
its hard to decide if it is still present and simply just have not 
occured, or if the current version is a good one. I'll try to do my 
best.

I will also contact linux-nfs mailing list.

As I remember, it started nearly a year ago, when I switched to Debian's 
kernel. I dont know exactly what version was at that time. Howewer, I've 
checked Debian's patches, and I did not find anything related to NFS.

Regards,
Richard


2024-05-20 21:07 időpontban Salvatore Bonaccorso ezt írta:
> Hi Richard,
> 
> On Mon, May 20, 2024 at 09:27:24AM +0000, Richard Kojedzinszky wrote:
>> Package: src:linux
>> Version: 6.1.90-1
>> Severity: normal
>> X-Debbugs-Cc: richard+debian+bugreport@kojedz.in
>> 
>> Dear Maintainer,
>> 
>> I am running kubernetes on debian, and pods are mounting multiple nfs
>> shares. I am running dovecot processes in PODs, which receive mails 
>> from
>> the internet, and also serves as imap server for clients. I am
>> monitoring my mail system by sending mails periodically (15 seconds) 
>> and
>> also downloading them via imap. I found a few times that some dovecot 
>> process
>> stuck in D state, a reboot was always needed to recover from that 
>> state.
>> 
>> Unfortunately, I was not able to trigger the bug really fast, I dont
>> really know what operations does dovecot issue and in what order to 
>> trigger
>> this behavior. So until I get closer, I've set up a similar, but 
>> smaller
>> environment with just a single dovecot process, and it also does the
>> same work, delivering only test mails locally, and serving them via 
>> imap
>> to the monitoring client, storing everything on NFS. Fortunately, this 
>> also
>> triggers the bug, after a few hours one of the dovecot processes is 
>> stuck
>> in D state. Kernel also shows blocked state:
> 
> As you seem in the lucky position to be able to trigger the issue in a
> more localized setup, might you:
> 
> - try as well more recent kernels from upper suites (6.8.9-1 in
>   unstable would be ideal to check if the issue is there as well).
> - I did read you cannot trigger with 5.15. If you build 6.1.90 from
>   upstream without Debian patches I assume you can trigger the issue
>   likewise? If so could you bisect the changes introducing the issue?
>   This is a cumbersome process in particular if you need few hours to
>   trigger it  So maybe the following point could be done first:
> - Can you report the issue to the linux-nfs list, keeping us in the
>   loop?
> 
> Regards,
> Salvatore

[toc] | [prev] | [next] | [standalone]


#82552 — Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate

FromRichard Kojedzinszky <richard+debian+bugreport@kojedz.in>
Date2024-05-23 14:20 +0200
SubjectBug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate
Message-ID<IHfAl-fhqp-1@gated-at.bofh.it>
In reply to#82499
Dear devs,

Now bisecting turned out that 3c59366c207e4c6c6569524af606baf017a55c61 
is the bad commit for me. Strangely it only affects my dovecot process 
accessing data over NFS.

Can you please confirm that this may be a bad commit?

My earlier attached programs may be used to demonstrate/trigger the 
issue. It even could be stripped down to minimal operations to trigger 
the bug.

Thanks in advance,
Richard


2024-05-23 09:10 időpontban Richard Kojedzinszky ezt írta:
> Dear NFS developers,
> 
> I am running multiple PODs on a Kubernetes node, they all mount 
> different NFS shares from the same nfs server. I started to notice 
> hangups in my dovecot process after I switched to Debian's kernel from 
> upstream 5.15. You can find Debian bugreport at 
> https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1071501.
> 
> So, effectively I am running dovecot in Kubernetes, and dovecot's data 
> directory is accessed over NFS. Eventually one dovecot process stucks 
> in nfs4_lookup_revalidate(). From that point, that process cannot be 
> killed, howewer, other processes can access NFS as normal. Also, 
> another dovecot process running on the very same node accessing the 
> same NFS share works too.
> 
> Now, I am still in the process of bisecting, howewer, I cannot reliably 
> trigger the bug. Originally it took a few days after I've noticed a 
> hanging process. Now I am trying to mimic file operations what dovecot 
> does in a faster way. Now it seems that it triggers the bug in a few 
> hours, howewer, during bisects, I can still make mistakes.
> 
> I've scheduled many of my applications which use NFS shares to the same 
> node, to have more NFS load on that node.
> 
> I am attaching my simple app which triggers the bug in a few hours, at 
> least in my lab. I have two dedicated NFS shares for this test case, 
> and I am running 3 instances of the applications for both shares. Also, 
> I am running other production applications on the same node which also 
> use NFS, howewer, I dont experience lockups with them. They are 
> librenms, prometheus, and a docker private registry. This way I dont 
> know if running the attached app only is enough to trigger the bug.
> 
> Once I have a suspectible commit based on my bisecting process, I will 
> report it here.
> 
> My NFS server is a TrueNAS, based on FreeBSD 13.3.
> 
> Thanks in advance,
> Richard

[toc] | [prev] | [next] | [standalone]


#82562 — Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate

FromRichard Kojedzinszky <richard+debian+bugreport@kojedz.in>
Date2024-05-24 07:40 +0200
SubjectBug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate
Message-ID<IHvON-frao-1@gated-at.bofh.it>
In reply to#82499
Dear Neil,

I've applied your patch, and since then there are no lockups. Before 
that my application reported a lockup in a minute or two, now it has 
been running for half an hour, and still running.

Thanks,
Richard

2024-05-24 01:31 időpontban NeilBrown ezt írta:
> On Fri, 24 May 2024, Richard Kojedzinszky wrote:
>> Dear devs,
>> 
>> I am attaching a stripped down version of the little program which
>> triggers the bug very quickly, in a few minutes in my test lab. It
>> turned out that a single NFS mountpoint is enough. Just start the
>> program giving it the NFS mount as first argument. It will chdir 
>> there,
>> and do file operations, which will trigger a lockup in a few minutes.
> 
> I couldn't get the go code to run.  But then it is a long time since I
> played with go and I didn't try very hard.
> If you could provide simple instructions and a list of package
> dependencies that I need to install (on Debian), I can give it a try.
> 
> Or you could try this patch.  It might help, but I don't have high
> hopes.  It adds some memory barriers and fixes a bug which would cause 
> a
> problem if memory allocation failed (but memory allocation never 
> fails).
> 
> NeilBrown
> 
> diff --git a/fs/nfs/dir.c b/fs/nfs/dir.c
> index ac505671efbd..5bcc0d14d519 100644
> --- a/fs/nfs/dir.c
> +++ b/fs/nfs/dir.c
> @@ -1804,7 +1804,7 @@ __nfs_lookup_revalidate(struct dentry *dentry, 
> unsigned int flags,
>  	} else {
>  		/* Wait for unlink to complete */
>  		wait_var_event(&dentry->d_fsdata,
> -			       dentry->d_fsdata != NFS_FSDATA_BLOCKED);
> +			       smp_load_acquire(&dentry->d_fsdata) != NFS_FSDATA_BLOCKED);
>  		parent = dget_parent(dentry);
>  		ret = reval(d_inode(parent), dentry, flags);
>  		dput(parent);
> @@ -2508,7 +2508,7 @@ int nfs_unlink(struct inode *dir, struct dentry 
> *dentry)
>  	spin_unlock(&dentry->d_lock);
>  	error = nfs_safe_remove(dentry);
>  	nfs_dentry_remove_handle_error(dir, dentry, error);
> -	dentry->d_fsdata = NULL;
> +	smp_store_release(&dentry->d_fsdata, NULL);
>  	wake_up_var(&dentry->d_fsdata);
>  out:
>  	trace_nfs_unlink_exit(dir, dentry, error);
> @@ -2616,7 +2616,7 @@ nfs_unblock_rename(struct rpc_task *task, struct 
> nfs_renamedata *data)
>  {
>  	struct dentry *new_dentry = data->new_dentry;
> 
> -	new_dentry->d_fsdata = NULL;
> +	smp_store_release(&new_dentry->d_fsdata, NULL);
>  	wake_up_var(&new_dentry->d_fsdata);
>  }
> 
> @@ -2717,6 +2717,10 @@ int nfs_rename(struct mnt_idmap *idmap, struct 
> inode *old_dir,
>  	task = nfs_async_rename(old_dir, new_dir, old_dentry, new_dentry,
>  				must_unblock ? nfs_unblock_rename : NULL);
>  	if (IS_ERR(task)) {
> +		if (must_unlock) {
> +			smp_store_release(&new_dentry->d_fsdata, NULL);
> +			wake_up_var(&new_dentry->d_fsdata);
> +		}
>  		error = PTR_ERR(task);
>  		goto out;
>  	}

[toc] | [prev] | [next] | [standalone]


#82568 — Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate

FromRichard Kojedzinszky <richard+debian+bugreport@kojedz.in>
Date2024-05-25 18:10 +0200
SubjectBug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate
Message-ID<II281-fKXU-1@gated-at.bofh.it>
In reply to#82562
Dear Neil,

According to my quick tests, your patch seems to fix this bug. Could you 
also manage to try my attached code, could you also reproduce the bug?

Thanks,
Richard

2024-05-24 07:29 időpontban Richard Kojedzinszky ezt írta:
> Dear Neil,
> 
> I've applied your patch, and since then there are no lockups. Before 
> that my application reported a lockup in a minute or two, now it has 
> been running for half an hour, and still running.
> 
> Thanks,
> Richard
> 
> 2024-05-24 01:31 időpontban NeilBrown ezt írta:
>> On Fri, 24 May 2024, Richard Kojedzinszky wrote:
>>> Dear devs,
>>> 
>>> I am attaching a stripped down version of the little program which
>>> triggers the bug very quickly, in a few minutes in my test lab. It
>>> turned out that a single NFS mountpoint is enough. Just start the
>>> program giving it the NFS mount as first argument. It will chdir 
>>> there,
>>> and do file operations, which will trigger a lockup in a few minutes.
>> 
>> I couldn't get the go code to run.  But then it is a long time since I
>> played with go and I didn't try very hard.
>> If you could provide simple instructions and a list of package
>> dependencies that I need to install (on Debian), I can give it a try.
>> 
>> Or you could try this patch.  It might help, but I don't have high
>> hopes.  It adds some memory barriers and fixes a bug which would cause 
>> a
>> problem if memory allocation failed (but memory allocation never 
>> fails).
>> 
>> NeilBrown
>> 
>> diff --git a/fs/nfs/dir.c b/fs/nfs/dir.c
>> index ac505671efbd..5bcc0d14d519 100644
>> --- a/fs/nfs/dir.c
>> +++ b/fs/nfs/dir.c
>> @@ -1804,7 +1804,7 @@ __nfs_lookup_revalidate(struct dentry *dentry, 
>> unsigned int flags,
>>  	} else {
>>  		/* Wait for unlink to complete */
>>  		wait_var_event(&dentry->d_fsdata,
>> -			       dentry->d_fsdata != NFS_FSDATA_BLOCKED);
>> +			       smp_load_acquire(&dentry->d_fsdata) != NFS_FSDATA_BLOCKED);
>>  		parent = dget_parent(dentry);
>>  		ret = reval(d_inode(parent), dentry, flags);
>>  		dput(parent);
>> @@ -2508,7 +2508,7 @@ int nfs_unlink(struct inode *dir, struct dentry 
>> *dentry)
>>  	spin_unlock(&dentry->d_lock);
>>  	error = nfs_safe_remove(dentry);
>>  	nfs_dentry_remove_handle_error(dir, dentry, error);
>> -	dentry->d_fsdata = NULL;
>> +	smp_store_release(&dentry->d_fsdata, NULL);
>>  	wake_up_var(&dentry->d_fsdata);
>>  out:
>>  	trace_nfs_unlink_exit(dir, dentry, error);
>> @@ -2616,7 +2616,7 @@ nfs_unblock_rename(struct rpc_task *task, struct 
>> nfs_renamedata *data)
>>  {
>>  	struct dentry *new_dentry = data->new_dentry;
>> 
>> -	new_dentry->d_fsdata = NULL;
>> +	smp_store_release(&new_dentry->d_fsdata, NULL);
>>  	wake_up_var(&new_dentry->d_fsdata);
>>  }
>> 
>> @@ -2717,6 +2717,10 @@ int nfs_rename(struct mnt_idmap *idmap, struct 
>> inode *old_dir,
>>  	task = nfs_async_rename(old_dir, new_dir, old_dentry, new_dentry,
>>  				must_unblock ? nfs_unblock_rename : NULL);
>>  	if (IS_ERR(task)) {
>> +		if (must_unlock) {
>> +			smp_store_release(&new_dentry->d_fsdata, NULL);
>> +			wake_up_var(&new_dentry->d_fsdata);
>> +		}
>>  		error = PTR_ERR(task);
>>  		goto out;
>>  	}

[toc] | [prev] | [next] | [standalone]


#82588 — Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate

FromRichard Kojedzinszky <richard@kojedz.in>
Date2024-05-27 06:10 +0200
SubjectBug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate
Message-ID<IIzQl-g70J-1@gated-at.bofh.it>
In reply to#82499

[Multipart message — attachments visible in raw view] — view raw

Dear Neil,

I was running it on arm64, may that be the reason?

Regards,
Richard

On May 27, 2024 4:02:32 AM GMT+02:00, NeilBrown <neilb@suse.de> wrote:
>On Sun, 26 May 2024, Richard Kojedzinszky wrote:
>> Dear Neil,
>> 
>> According to my quick tests, your patch seems to fix this bug. Could you 
>> also manage to try my attached code, could you also reproduce the bug?
>
>Thanks for testing.
>
>I can run your test code but it isn't triggering the bug (90 minutes so
>far).  Possibly a different compiler used for the kernel, possibly
>hardware differences (I'm running under qemu).  Bugs related to barriers
>(which this one seems to be) need just the right circumstances to
>trigger so they can be hard to reproduce on a different system.
>
>I've made some cosmetic improvements to the patch and will post it to
>the NFS maintainers.
>
>Thanks again,
>NeilBrown
>
>
>> 
>> Thanks,
>> Richard
>> 
>> 2024-05-24 07:29 időpontban Richard Kojedzinszky ezt írta:
>> > Dear Neil,
>> > 
>> > I've applied your patch, and since then there are no lockups. Before 
>> > that my application reported a lockup in a minute or two, now it has 
>> > been running for half an hour, and still running.
>> > 
>> > Thanks,
>> > Richard
>> > 
>> > 2024-05-24 01:31 időpontban NeilBrown ezt írta:
>> >> On Fri, 24 May 2024, Richard Kojedzinszky wrote:
>> >>> Dear devs,
>> >>> 
>> >>> I am attaching a stripped down version of the little program which
>> >>> triggers the bug very quickly, in a few minutes in my test lab. It
>> >>> turned out that a single NFS mountpoint is enough. Just start the
>> >>> program giving it the NFS mount as first argument. It will chdir 
>> >>> there,
>> >>> and do file operations, which will trigger a lockup in a few minutes.
>> >> 
>> >> I couldn't get the go code to run.  But then it is a long time since I
>> >> played with go and I didn't try very hard.
>> >> If you could provide simple instructions and a list of package
>> >> dependencies that I need to install (on Debian), I can give it a try.
>> >> 
>> >> Or you could try this patch.  It might help, but I don't have high
>> >> hopes.  It adds some memory barriers and fixes a bug which would cause 
>> >> a
>> >> problem if memory allocation failed (but memory allocation never 
>> >> fails).
>> >> 
>> >> NeilBrown
>> >> 
>> >> diff --git a/fs/nfs/dir.c b/fs/nfs/dir.c
>> >> index ac505671efbd..5bcc0d14d519 100644
>> >> --- a/fs/nfs/dir.c
>> >> +++ b/fs/nfs/dir.c
>> >> @@ -1804,7 +1804,7 @@ __nfs_lookup_revalidate(struct dentry *dentry, 
>> >> unsigned int flags,
>> >>  	} else {
>> >>  		/* Wait for unlink to complete */
>> >>  		wait_var_event(&dentry->d_fsdata,
>> >> -			       dentry->d_fsdata != NFS_FSDATA_BLOCKED);
>> >> +			       smp_load_acquire(&dentry->d_fsdata) != NFS_FSDATA_BLOCKED);
>> >>  		parent = dget_parent(dentry);
>> >>  		ret = reval(d_inode(parent), dentry, flags);
>> >>  		dput(parent);
>> >> @@ -2508,7 +2508,7 @@ int nfs_unlink(struct inode *dir, struct dentry 
>> >> *dentry)
>> >>  	spin_unlock(&dentry->d_lock);
>> >>  	error = nfs_safe_remove(dentry);
>> >>  	nfs_dentry_remove_handle_error(dir, dentry, error);
>> >> -	dentry->d_fsdata = NULL;
>> >> +	smp_store_release(&dentry->d_fsdata, NULL);
>> >>  	wake_up_var(&dentry->d_fsdata);
>> >>  out:
>> >>  	trace_nfs_unlink_exit(dir, dentry, error);
>> >> @@ -2616,7 +2616,7 @@ nfs_unblock_rename(struct rpc_task *task, struct 
>> >> nfs_renamedata *data)
>> >>  {
>> >>  	struct dentry *new_dentry = data->new_dentry;
>> >> 
>> >> -	new_dentry->d_fsdata = NULL;
>> >> +	smp_store_release(&new_dentry->d_fsdata, NULL);
>> >>  	wake_up_var(&new_dentry->d_fsdata);
>> >>  }
>> >> 
>> >> @@ -2717,6 +2717,10 @@ int nfs_rename(struct mnt_idmap *idmap, struct 
>> >> inode *old_dir,
>> >>  	task = nfs_async_rename(old_dir, new_dir, old_dentry, new_dentry,
>> >>  				must_unblock ? nfs_unblock_rename : NULL);
>> >>  	if (IS_ERR(task)) {
>> >> +		if (must_unlock) {
>> >> +			smp_store_release(&new_dentry->d_fsdata, NULL);
>> >> +			wake_up_var(&new_dentry->d_fsdata);
>> >> +		}
>> >>  		error = PTR_ERR(task);
>> >>  		goto out;
>> >>  	}
>> 
>

[toc] | [prev] | [next] | [standalone]


#82859 — Bug#1071501: marked as done (linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate)

From"Debian Bug Tracking System" <owner@bugs.debian.org>
Date2024-06-28 00:20 +0200
SubjectBug#1071501: marked as done (linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate)
Message-ID<IU5Dc-64H3-15@gated-at.bofh.it>
In reply to#82499

[Multipart message — attachments visible in raw view] — view raw

Your message dated Thu, 27 Jun 2024 22:10:11 +0000
with message-id <E1sMxJj-000sZ1-JY@fasolo.debian.org>
and subject line Bug#1071501: fixed in linux 6.9.7-1
has caused the Debian Bug report #1071501,
regarding linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate
to be marked as done.

This means that you claim that the problem has been dealt with.
If this is not the case it is now your responsibility to reopen the
Bug report if necessary, and/or fix the problem forthwith.

(NB: If you are a system administrator and have no idea what this
message is talking about, this may indicate a serious mail system
misconfiguration somewhere. Please contact owner@bugs.debian.org
immediately.)


-- 
1071501: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1071501
Debian Bug Tracking System
Contact owner@bugs.debian.org with problems

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web