Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.kernel > #82499 > unrolled thread
| Started by | Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> |
|---|---|
| First post | 2024-05-20 11:40 +0200 |
| Last post | 2024-06-28 00:20 +0200 |
| Articles | 9 — 5 participants |
Back to article view | Back to linux.debian.kernel
Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> - 2024-05-20 11:40 +0200
Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate Salvatore Bonaccorso <carnil@debian.org> - 2024-05-20 21:10 +0200
Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate Diederik de Haas <didi.debian@cknow.org> - 2024-05-20 23:40 +0200
Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> - 2024-05-21 10:10 +0200
Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> - 2024-05-23 14:20 +0200
Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> - 2024-05-24 07:40 +0200
Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> - 2024-05-25 18:10 +0200
Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate Richard Kojedzinszky <richard@kojedz.in> - 2024-05-27 06:10 +0200
Bug#1071501: marked as done (linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate) "Debian Bug Tracking System" <owner@bugs.debian.org> - 2024-06-28 00:20 +0200
| From | Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> |
|---|---|
| Date | 2024-05-20 11:40 +0200 |
| Subject | Bug#1071501: linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate |
| Message-ID | <IG7ER-eBme-1@gated-at.bofh.it> |
Package: src:linux Version: 6.1.90-1 Severity: normal X-Debbugs-Cc: richard+debian+bugreport@kojedz.in Dear Maintainer, I am running kubernetes on debian, and pods are mounting multiple nfs shares. I am running dovecot processes in PODs, which receive mails from the internet, and also serves as imap server for clients. I am monitoring my mail system by sending mails periodically (15 seconds) and also downloading them via imap. I found a few times that some dovecot process stuck in D state, a reboot was always needed to recover from that state. Unfortunately, I was not able to trigger the bug really fast, I dont really know what operations does dovecot issue and in what order to trigger this behavior. So until I get closer, I've set up a similar, but smaller environment with just a single dovecot process, and it also does the same work, delivering only test mails locally, and serving them via imap to the monitoring client, storing everything on NFS. Fortunately, this also triggers the bug, after a few hours one of the dovecot processes is stuck in D state. Kernel also shows blocked state: May 19 12:16:49 k8s-node07 kernel: INFO: task lmtp:665683 blocked for more than 120 seconds. May 19 12:16:49 k8s-node07 kernel: Not tainted 6.1.0-21-arm64 #1 Debian 6.1.90-1 May 19 12:16:49 k8s-node07 kernel: "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. May 19 12:16:49 k8s-node07 kernel: task:lmtp state:D stack:0 pid:665683 ppid:2881 flags:0x00000000 May 19 12:16:49 k8s-node07 kernel: Call trace: May 19 12:16:49 k8s-node07 kernel: __switch_to+0xf0/0x170 May 19 12:16:49 k8s-node07 kernel: __schedule+0x340/0x940 May 19 12:16:49 k8s-node07 kernel: schedule+0x58/0xf0 May 19 12:16:49 k8s-node07 kernel: __nfs_lookup_revalidate+0x118/0x160 [nfs] May 19 12:16:49 k8s-node07 kernel: nfs4_lookup_revalidate+0x20/0x30 [nfs] May 19 12:16:49 k8s-node07 kernel: lookup_fast+0x138/0x150 May 19 12:16:49 k8s-node07 kernel: walk_component+0x30/0x1a0 May 19 12:16:49 k8s-node07 kernel: path_lookupat+0x80/0x1a4 May 19 12:16:49 k8s-node07 kernel: filename_lookup+0xb4/0x1b0 May 19 12:16:49 k8s-node07 kernel: vfs_statx+0x94/0x19c May 19 12:16:49 k8s-node07 kernel: vfs_fstatat+0x68/0x90 May 19 12:16:49 k8s-node07 kernel: __do_sys_newfstatat+0x58/0xa0 May 19 12:16:49 k8s-node07 kernel: __arm64_sys_newfstatat+0x28/0x34 May 19 12:16:49 k8s-node07 kernel: invoke_syscall+0x78/0x100 May 19 12:16:49 k8s-node07 kernel: el0_svc_common.constprop.0+0x4c/0xf4 May 19 12:16:49 k8s-node07 kernel: do_el0_svc+0x34/0xd0 May 19 12:16:49 k8s-node07 kernel: el0_svc+0x34/0xd4 May 19 12:16:49 k8s-node07 kernel: el0t_64_sync_handler+0xf4/0x120 May 19 12:16:49 k8s-node07 kernel: el0t_64_sync+0x18c/0x190 Or, for another process: May 20 04:50:01 k8s-node07 kernel: INFO: task imap:8337 blocked for more than 120 seconds. May 20 04:50:01 k8s-node07 kernel: Not tainted 6.1.0-21-arm64 #1 Debian 6.1.90-1 May 20 04:50:01 k8s-node07 kernel: "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. May 20 04:50:01 k8s-node07 kernel: task:imap state:D stack:0 pid:8337 ppid:3164 flags:0x00000000 May 20 04:50:01 k8s-node07 kernel: Call trace: May 20 04:50:01 k8s-node07 kernel: __switch_to+0xf0/0x170 May 20 04:50:01 k8s-node07 kernel: __schedule+0x340/0x940 May 20 04:50:01 k8s-node07 kernel: schedule+0x58/0xf0 May 20 04:50:01 k8s-node07 kernel: __nfs_lookup_revalidate+0x118/0x160 [nfs] May 20 04:50:01 k8s-node07 kernel: nfs4_lookup_revalidate+0x20/0x30 [nfs] May 20 04:50:01 k8s-node07 kernel: lookup_fast+0x138/0x150 May 20 04:50:01 k8s-node07 kernel: walk_component+0x30/0x1a0 May 20 04:50:01 k8s-node07 kernel: path_lookupat+0x80/0x1a4 May 20 04:50:01 k8s-node07 kernel: filename_lookup+0xb4/0x1b0 May 20 04:50:01 k8s-node07 kernel: vfs_statx+0x94/0x19c May 20 04:50:01 k8s-node07 kernel: vfs_fstatat+0x68/0x90 May 20 04:50:01 k8s-node07 kernel: __do_sys_newfstatat+0x58/0xa0 May 20 04:50:01 k8s-node07 kernel: __arm64_sys_newfstatat+0x28/0x34 May 20 04:50:01 k8s-node07 kernel: invoke_syscall+0x78/0x100 May 20 04:50:01 k8s-node07 kernel: el0_svc_common.constprop.0+0x4c/0xf4 May 20 04:50:01 k8s-node07 kernel: do_el0_svc+0x34/0xd0 May 20 04:50:01 k8s-node07 kernel: el0_svc+0x34/0xd4 May 20 04:50:01 k8s-node07 kernel: el0t_64_sync_handler+0xf4/0x120 May 20 04:50:01 k8s-node07 kernel: el0t_64_sync+0x18c/0x190 Of course the NFS server is running, and other NFS mounts are still working from the node. Also, this started to happen with Debian's kernel. Before that, I was compiling my own upstream kernel, version 5.15. With that, I've never experienced such a lockup. Unfortunately, I dont know, how to go further, how shall I collect more relevant debugging information. I expect thet dovecot is just an application, which should not cause any kernel-side lockups. In my test lab, this specific NFS mount is just mounted on one machine, so it really suggests me a linux nfs-client side issue, not related to caching coherency between multiple clients. -- Package-specific info: ** Version: Linux version 6.1.0-21-arm64 (debian-kernel@lists.debian.org) (gcc-12 (Debian 12.2.0-14) 12.2.0, GNU ld (GNU Binutils for Debian) 2.40) #1 SMP Debian 6.1.90-1 (2024-05-03) ** Command line: net.ifnames=0 console=ttyS2,1500000 console=tty1 root=UUID=b4ff4167-1fe9-4fd6-9b9c-c3c68d98108b rw rootwait panic=10 ** Not tainted ** Kernel log: May 20 04:52:02 k8s-node07 kernel: INFO: task imap:8337 blocked for more than 241 seconds. May 20 04:52:02 k8s-node07 kernel: Not tainted 6.1.0-21-arm64 #1 Debian 6.1.90-1 May 20 04:52:02 k8s-node07 kernel: "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. May 20 04:52:02 k8s-node07 kernel: task:imap state:D stack:0 pid:8337 ppid:3164 flags:0x00000000 May 20 04:52:02 k8s-node07 kernel: Call trace: May 20 04:52:02 k8s-node07 kernel: __switch_to+0xf0/0x170 May 20 04:52:02 k8s-node07 kernel: __schedule+0x340/0x940 May 20 04:52:02 k8s-node07 kernel: schedule+0x58/0xf0 May 20 04:52:02 k8s-node07 kernel: __nfs_lookup_revalidate+0x118/0x160 [nfs] May 20 04:52:02 k8s-node07 kernel: nfs4_lookup_revalidate+0x20/0x30 [nfs] May 20 04:52:02 k8s-node07 kernel: lookup_fast+0x138/0x150 May 20 04:52:02 k8s-node07 kernel: walk_component+0x30/0x1a0 May 20 04:52:02 k8s-node07 kernel: path_lookupat+0x80/0x1a4 May 20 04:52:02 k8s-node07 kernel: filename_lookup+0xb4/0x1b0 May 20 04:52:02 k8s-node07 kernel: vfs_statx+0x94/0x19c May 20 04:52:02 k8s-node07 kernel: vfs_fstatat+0x68/0x90 May 20 04:52:02 k8s-node07 kernel: __do_sys_newfstatat+0x58/0xa0 May 20 04:52:02 k8s-node07 kernel: __arm64_sys_newfstatat+0x28/0x34 May 20 04:52:02 k8s-node07 kernel: invoke_syscall+0x78/0x100 May 20 04:52:02 k8s-node07 kernel: el0_svc_common.constprop.0+0x4c/0xf4 May 20 04:52:02 k8s-node07 kernel: do_el0_svc+0x34/0xd0 May 20 04:52:02 k8s-node07 kernel: el0_svc+0x34/0xd4 May 20 04:52:02 k8s-node07 kernel: el0t_64_sync_handler+0xf4/0x120 May 20 04:52:02 k8s-node07 kernel: el0t_64_sync+0x18c/0x190 ** Model information ** Loaded modules: sd_mod t10_pi crc64_rocksoft_generic crc64_rocksoft crc_t10dif crct10dif_generic crc64 sg iscsi_tcp libiscsi_tcp libiscsi scsi_transport_iscsi scsi_mod scsi_common nf_conntrack_netlink rpcsec_gss_krb5 auth_rpcgss nfsv4 dns_resolver nfs lockd grace fscache netfs nft_log nft_limit xt_limit xt_NFLOG nfnetlink_log xt_physdev xt_TCPMSS xt_tcpudp xt_mark xt_multiport xt_addrtype dummy ipt_REJECT nf_reject_ipv4 ip_set_hash_ipport nft_chain_nat xt_nat xt_MASQUERADE xt_ipvs nf_nat xt_set ip_set_hash_ip ip_set_hash_net ip_set veth xt_conntrack xt_comment nft_compat nf_tables nfnetlink overlay sunrpc binfmt_misc evdev aes_ce_blk snd_soc_rk817 aes_ce_cipher polyval_ce snd_soc_core polyval_generic snd_pcm_dmaengine ext4 ghash_ce gf128mul sha2_ce leds_gpio snd_pcm sha256_arm64 sha1_ce rockchip_thermal crc16 mbcache snd_timer jbd2 snd dw_wdt soundcore rk817_charger rk805_pwrkey cpufreq_dt br_netfilter bridge stp llc ip_vs_sh ip_vs_wrr ip_vs_rr ip_vs nf_conntrack nf_defrag_ipv6 nf_defrag_ipv4 drm loop fuse efi_pstore dm_mod dax configfs ip_tables x_tables autofs4 xfs libcrc32c crc32c_generic realtek rk808_regulator fan53555 dwmac_rk stmmac_platform stmmac pcs_xpcs spi_rockchip phylink dw_mmc_rockchip dw_mmc_pltfm of_mdio dw_mmc fixed crct10dif_ce crct10dif_common fixed_phy fwnode_mdio pl330 i2c_rk3x io_domain libphy ** PCI devices: not available ** USB devices: not available -- System Information: Debian Release: 12.5 APT prefers stable-updates APT policy: (500, 'stable-updates'), (500, 'stable-security'), (500, 'stable') Architecture: arm64 (aarch64) Kernel: Linux 6.1.0-21-arm64 (SMP w/4 CPU threads) Locale: LANG=C, LC_CTYPE=C.UTF-8 (charmap=UTF-8), LANGUAGE not set Shell: /bin/sh linked to /usr/bin/dash Init: unable to detect Versions of packages linux-image-6.1.0-21-arm64 depends on: ii initramfs-tools [linux-initramfs-tool] 0.142 ii kmod 30+20221128-1 ii linux-base 4.9 Versions of packages linux-image-6.1.0-21-arm64 recommends: ii apparmor 3.0.8-3 ii firmware-linux-free 20200122-1 Versions of packages linux-image-6.1.0-21-arm64 suggests: pn debian-kernel-handbook <none> pn linux-doc-6.1 <none> Versions of packages linux-image-6.1.0-21-arm64 is related to: pn firmware-amd-graphics <none> pn firmware-atheros <none> pn firmware-bnx2 <none> pn firmware-bnx2x <none> pn firmware-brcm80211 <none> pn firmware-cavium <none> pn firmware-intel-sound <none> pn firmware-intelwimax <none> pn firmware-ipw2x00 <none> pn firmware-ivtv <none> pn firmware-iwlwifi <none> pn firmware-libertas <none> pn firmware-linux-nonfree <none> pn firmware-misc-nonfree <none> pn firmware-myricom <none> pn firmware-netxen <none> pn firmware-qlogic <none> pn firmware-realtek <none> pn firmware-samsung <none> pn firmware-siano <none> pn firmware-ti-connectivity <none> pn xen-hypervisor <none> -- no debconf information
[toc] | [next] | [standalone]
| From | Salvatore Bonaccorso <carnil@debian.org> |
|---|---|
| Date | 2024-05-20 21:10 +0200 |
| Message-ID | <IGgyu-eGN0-1@gated-at.bofh.it> |
| In reply to | #82499 |
Hi Richard, On Mon, May 20, 2024 at 09:27:24AM +0000, Richard Kojedzinszky wrote: > Package: src:linux > Version: 6.1.90-1 > Severity: normal > X-Debbugs-Cc: richard+debian+bugreport@kojedz.in > > Dear Maintainer, > > I am running kubernetes on debian, and pods are mounting multiple nfs > shares. I am running dovecot processes in PODs, which receive mails from > the internet, and also serves as imap server for clients. I am > monitoring my mail system by sending mails periodically (15 seconds) and > also downloading them via imap. I found a few times that some dovecot process > stuck in D state, a reboot was always needed to recover from that state. > > Unfortunately, I was not able to trigger the bug really fast, I dont > really know what operations does dovecot issue and in what order to trigger > this behavior. So until I get closer, I've set up a similar, but smaller > environment with just a single dovecot process, and it also does the > same work, delivering only test mails locally, and serving them via imap > to the monitoring client, storing everything on NFS. Fortunately, this also > triggers the bug, after a few hours one of the dovecot processes is stuck > in D state. Kernel also shows blocked state: As you seem in the lucky position to be able to trigger the issue in a more localized setup, might you: - try as well more recent kernels from upper suites (6.8.9-1 in unstable would be ideal to check if the issue is there as well). - I did read you cannot trigger with 5.15. If you build 6.1.90 from upstream without Debian patches I assume you can trigger the issue likewise? If so could you bisect the changes introducing the issue? This is a cumbersome process in particular if you need few hours to trigger it So maybe the following point could be done first: - Can you report the issue to the linux-nfs list, keeping us in the loop? Regards, Salvatore
[toc] | [prev] | [next] | [standalone]
| From | Diederik de Haas <didi.debian@cknow.org> |
|---|---|
| Date | 2024-05-20 23:40 +0200 |
| Message-ID | <IGiTD-eI4n-3@gated-at.bofh.it> |
| In reply to | #82512 |
[Multipart message — attachments visible in raw view] — view raw
On Monday, 20 May 2024 21:07:49 CEST Salvatore Bonaccorso wrote: > - I did read you cannot trigger with 5.15. If you build 6.1.90 from > upstream without Debian patches I assume you can trigger the issue > likewise? If so could you bisect the changes introducing the issue? If the test with the upstream 6.1.90 version also has this problem, there's another (series of) test(s) worth doing, which could shorten the bisect operation significantly. I got the impression that you have only tried it with version 6.1.90. Have you tried it with earlier versions in the 6.1 series to see if the issue is present there? Via https://snapshot.debian.org/package/linux-signed-arm64/ you can find earlier versions from the 6.1 series already compiled and packaged. To take version 6.1.52 as (random) example: - click on the ``6.1.52+1`` link - In the ``Binary packages`` list, click on the linux-image-6.1.0-X-arm64 link, where 'X' is 12 in this case - Click the ``linux-image-6.1.0-12-arm64_6.1.52-1_arm64.deb`` link to download the deb file which you can then install as root or with sudo by executing ``apt install ./linux-image-<version-info>_arm64.deb`` If the problem does NOT occur with 6.1.52-1, then you try a higher version. Continue that process until you've found the latest version that works and the earliest version where it stopped working. If the problem also occurs with 6.1.52-1, then you try an (even) older version. This is to test whether it was a regression *within* the 6.1 series and if so, to get the narrowest range without having to compile yourself. HTH
[toc] | [prev] | [next] | [standalone]
| From | Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> |
|---|---|
| Date | 2024-05-21 10:10 +0200 |
| Message-ID | <IGsJj-eOf0-1@gated-at.bofh.it> |
| In reply to | #82512 |
Dear Salvatore, I've already started bisecting. It will take some time. Usually the bug appears after a few hours, unfortunately I am not able to trigger it faster. So, if the bug appears, I can step forward easily, but if not, its hard to decide if it is still present and simply just have not occured, or if the current version is a good one. I'll try to do my best. I will also contact linux-nfs mailing list. As I remember, it started nearly a year ago, when I switched to Debian's kernel. I dont know exactly what version was at that time. Howewer, I've checked Debian's patches, and I did not find anything related to NFS. Regards, Richard 2024-05-20 21:07 időpontban Salvatore Bonaccorso ezt írta: > Hi Richard, > > On Mon, May 20, 2024 at 09:27:24AM +0000, Richard Kojedzinszky wrote: >> Package: src:linux >> Version: 6.1.90-1 >> Severity: normal >> X-Debbugs-Cc: richard+debian+bugreport@kojedz.in >> >> Dear Maintainer, >> >> I am running kubernetes on debian, and pods are mounting multiple nfs >> shares. I am running dovecot processes in PODs, which receive mails >> from >> the internet, and also serves as imap server for clients. I am >> monitoring my mail system by sending mails periodically (15 seconds) >> and >> also downloading them via imap. I found a few times that some dovecot >> process >> stuck in D state, a reboot was always needed to recover from that >> state. >> >> Unfortunately, I was not able to trigger the bug really fast, I dont >> really know what operations does dovecot issue and in what order to >> trigger >> this behavior. So until I get closer, I've set up a similar, but >> smaller >> environment with just a single dovecot process, and it also does the >> same work, delivering only test mails locally, and serving them via >> imap >> to the monitoring client, storing everything on NFS. Fortunately, this >> also >> triggers the bug, after a few hours one of the dovecot processes is >> stuck >> in D state. Kernel also shows blocked state: > > As you seem in the lucky position to be able to trigger the issue in a > more localized setup, might you: > > - try as well more recent kernels from upper suites (6.8.9-1 in > unstable would be ideal to check if the issue is there as well). > - I did read you cannot trigger with 5.15. If you build 6.1.90 from > upstream without Debian patches I assume you can trigger the issue > likewise? If so could you bisect the changes introducing the issue? > This is a cumbersome process in particular if you need few hours to > trigger it So maybe the following point could be done first: > - Can you report the issue to the linux-nfs list, keeping us in the > loop? > > Regards, > Salvatore
[toc] | [prev] | [next] | [standalone]
| From | Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> |
|---|---|
| Date | 2024-05-23 14:20 +0200 |
| Subject | Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate |
| Message-ID | <IHfAl-fhqp-1@gated-at.bofh.it> |
| In reply to | #82499 |
Dear devs, Now bisecting turned out that 3c59366c207e4c6c6569524af606baf017a55c61 is the bad commit for me. Strangely it only affects my dovecot process accessing data over NFS. Can you please confirm that this may be a bad commit? My earlier attached programs may be used to demonstrate/trigger the issue. It even could be stripped down to minimal operations to trigger the bug. Thanks in advance, Richard 2024-05-23 09:10 időpontban Richard Kojedzinszky ezt írta: > Dear NFS developers, > > I am running multiple PODs on a Kubernetes node, they all mount > different NFS shares from the same nfs server. I started to notice > hangups in my dovecot process after I switched to Debian's kernel from > upstream 5.15. You can find Debian bugreport at > https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1071501. > > So, effectively I am running dovecot in Kubernetes, and dovecot's data > directory is accessed over NFS. Eventually one dovecot process stucks > in nfs4_lookup_revalidate(). From that point, that process cannot be > killed, howewer, other processes can access NFS as normal. Also, > another dovecot process running on the very same node accessing the > same NFS share works too. > > Now, I am still in the process of bisecting, howewer, I cannot reliably > trigger the bug. Originally it took a few days after I've noticed a > hanging process. Now I am trying to mimic file operations what dovecot > does in a faster way. Now it seems that it triggers the bug in a few > hours, howewer, during bisects, I can still make mistakes. > > I've scheduled many of my applications which use NFS shares to the same > node, to have more NFS load on that node. > > I am attaching my simple app which triggers the bug in a few hours, at > least in my lab. I have two dedicated NFS shares for this test case, > and I am running 3 instances of the applications for both shares. Also, > I am running other production applications on the same node which also > use NFS, howewer, I dont experience lockups with them. They are > librenms, prometheus, and a docker private registry. This way I dont > know if running the attached app only is enough to trigger the bug. > > Once I have a suspectible commit based on my bisecting process, I will > report it here. > > My NFS server is a TrueNAS, based on FreeBSD 13.3. > > Thanks in advance, > Richard
[toc] | [prev] | [next] | [standalone]
| From | Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> |
|---|---|
| Date | 2024-05-24 07:40 +0200 |
| Subject | Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate |
| Message-ID | <IHvON-frao-1@gated-at.bofh.it> |
| In reply to | #82499 |
Dear Neil,
I've applied your patch, and since then there are no lockups. Before
that my application reported a lockup in a minute or two, now it has
been running for half an hour, and still running.
Thanks,
Richard
2024-05-24 01:31 időpontban NeilBrown ezt írta:
> On Fri, 24 May 2024, Richard Kojedzinszky wrote:
>> Dear devs,
>>
>> I am attaching a stripped down version of the little program which
>> triggers the bug very quickly, in a few minutes in my test lab. It
>> turned out that a single NFS mountpoint is enough. Just start the
>> program giving it the NFS mount as first argument. It will chdir
>> there,
>> and do file operations, which will trigger a lockup in a few minutes.
>
> I couldn't get the go code to run. But then it is a long time since I
> played with go and I didn't try very hard.
> If you could provide simple instructions and a list of package
> dependencies that I need to install (on Debian), I can give it a try.
>
> Or you could try this patch. It might help, but I don't have high
> hopes. It adds some memory barriers and fixes a bug which would cause
> a
> problem if memory allocation failed (but memory allocation never
> fails).
>
> NeilBrown
>
> diff --git a/fs/nfs/dir.c b/fs/nfs/dir.c
> index ac505671efbd..5bcc0d14d519 100644
> --- a/fs/nfs/dir.c
> +++ b/fs/nfs/dir.c
> @@ -1804,7 +1804,7 @@ __nfs_lookup_revalidate(struct dentry *dentry,
> unsigned int flags,
> } else {
> /* Wait for unlink to complete */
> wait_var_event(&dentry->d_fsdata,
> - dentry->d_fsdata != NFS_FSDATA_BLOCKED);
> + smp_load_acquire(&dentry->d_fsdata) != NFS_FSDATA_BLOCKED);
> parent = dget_parent(dentry);
> ret = reval(d_inode(parent), dentry, flags);
> dput(parent);
> @@ -2508,7 +2508,7 @@ int nfs_unlink(struct inode *dir, struct dentry
> *dentry)
> spin_unlock(&dentry->d_lock);
> error = nfs_safe_remove(dentry);
> nfs_dentry_remove_handle_error(dir, dentry, error);
> - dentry->d_fsdata = NULL;
> + smp_store_release(&dentry->d_fsdata, NULL);
> wake_up_var(&dentry->d_fsdata);
> out:
> trace_nfs_unlink_exit(dir, dentry, error);
> @@ -2616,7 +2616,7 @@ nfs_unblock_rename(struct rpc_task *task, struct
> nfs_renamedata *data)
> {
> struct dentry *new_dentry = data->new_dentry;
>
> - new_dentry->d_fsdata = NULL;
> + smp_store_release(&new_dentry->d_fsdata, NULL);
> wake_up_var(&new_dentry->d_fsdata);
> }
>
> @@ -2717,6 +2717,10 @@ int nfs_rename(struct mnt_idmap *idmap, struct
> inode *old_dir,
> task = nfs_async_rename(old_dir, new_dir, old_dentry, new_dentry,
> must_unblock ? nfs_unblock_rename : NULL);
> if (IS_ERR(task)) {
> + if (must_unlock) {
> + smp_store_release(&new_dentry->d_fsdata, NULL);
> + wake_up_var(&new_dentry->d_fsdata);
> + }
> error = PTR_ERR(task);
> goto out;
> }
[toc] | [prev] | [next] | [standalone]
| From | Richard Kojedzinszky <richard+debian+bugreport@kojedz.in> |
|---|---|
| Date | 2024-05-25 18:10 +0200 |
| Subject | Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate |
| Message-ID | <II281-fKXU-1@gated-at.bofh.it> |
| In reply to | #82562 |
Dear Neil,
According to my quick tests, your patch seems to fix this bug. Could you
also manage to try my attached code, could you also reproduce the bug?
Thanks,
Richard
2024-05-24 07:29 időpontban Richard Kojedzinszky ezt írta:
> Dear Neil,
>
> I've applied your patch, and since then there are no lockups. Before
> that my application reported a lockup in a minute or two, now it has
> been running for half an hour, and still running.
>
> Thanks,
> Richard
>
> 2024-05-24 01:31 időpontban NeilBrown ezt írta:
>> On Fri, 24 May 2024, Richard Kojedzinszky wrote:
>>> Dear devs,
>>>
>>> I am attaching a stripped down version of the little program which
>>> triggers the bug very quickly, in a few minutes in my test lab. It
>>> turned out that a single NFS mountpoint is enough. Just start the
>>> program giving it the NFS mount as first argument. It will chdir
>>> there,
>>> and do file operations, which will trigger a lockup in a few minutes.
>>
>> I couldn't get the go code to run. But then it is a long time since I
>> played with go and I didn't try very hard.
>> If you could provide simple instructions and a list of package
>> dependencies that I need to install (on Debian), I can give it a try.
>>
>> Or you could try this patch. It might help, but I don't have high
>> hopes. It adds some memory barriers and fixes a bug which would cause
>> a
>> problem if memory allocation failed (but memory allocation never
>> fails).
>>
>> NeilBrown
>>
>> diff --git a/fs/nfs/dir.c b/fs/nfs/dir.c
>> index ac505671efbd..5bcc0d14d519 100644
>> --- a/fs/nfs/dir.c
>> +++ b/fs/nfs/dir.c
>> @@ -1804,7 +1804,7 @@ __nfs_lookup_revalidate(struct dentry *dentry,
>> unsigned int flags,
>> } else {
>> /* Wait for unlink to complete */
>> wait_var_event(&dentry->d_fsdata,
>> - dentry->d_fsdata != NFS_FSDATA_BLOCKED);
>> + smp_load_acquire(&dentry->d_fsdata) != NFS_FSDATA_BLOCKED);
>> parent = dget_parent(dentry);
>> ret = reval(d_inode(parent), dentry, flags);
>> dput(parent);
>> @@ -2508,7 +2508,7 @@ int nfs_unlink(struct inode *dir, struct dentry
>> *dentry)
>> spin_unlock(&dentry->d_lock);
>> error = nfs_safe_remove(dentry);
>> nfs_dentry_remove_handle_error(dir, dentry, error);
>> - dentry->d_fsdata = NULL;
>> + smp_store_release(&dentry->d_fsdata, NULL);
>> wake_up_var(&dentry->d_fsdata);
>> out:
>> trace_nfs_unlink_exit(dir, dentry, error);
>> @@ -2616,7 +2616,7 @@ nfs_unblock_rename(struct rpc_task *task, struct
>> nfs_renamedata *data)
>> {
>> struct dentry *new_dentry = data->new_dentry;
>>
>> - new_dentry->d_fsdata = NULL;
>> + smp_store_release(&new_dentry->d_fsdata, NULL);
>> wake_up_var(&new_dentry->d_fsdata);
>> }
>>
>> @@ -2717,6 +2717,10 @@ int nfs_rename(struct mnt_idmap *idmap, struct
>> inode *old_dir,
>> task = nfs_async_rename(old_dir, new_dir, old_dentry, new_dentry,
>> must_unblock ? nfs_unblock_rename : NULL);
>> if (IS_ERR(task)) {
>> + if (must_unlock) {
>> + smp_store_release(&new_dentry->d_fsdata, NULL);
>> + wake_up_var(&new_dentry->d_fsdata);
>> + }
>> error = PTR_ERR(task);
>> goto out;
>> }
[toc] | [prev] | [next] | [standalone]
| From | Richard Kojedzinszky <richard@kojedz.in> |
|---|---|
| Date | 2024-05-27 06:10 +0200 |
| Subject | Bug#1071501: Linux NFS client hangs in nfs4_lookup_revalidate |
| Message-ID | <IIzQl-g70J-1@gated-at.bofh.it> |
| In reply to | #82499 |
[Multipart message — attachments visible in raw view] — view raw
Dear Neil,
I was running it on arm64, may that be the reason?
Regards,
Richard
On May 27, 2024 4:02:32 AM GMT+02:00, NeilBrown <neilb@suse.de> wrote:
>On Sun, 26 May 2024, Richard Kojedzinszky wrote:
>> Dear Neil,
>>
>> According to my quick tests, your patch seems to fix this bug. Could you
>> also manage to try my attached code, could you also reproduce the bug?
>
>Thanks for testing.
>
>I can run your test code but it isn't triggering the bug (90 minutes so
>far). Possibly a different compiler used for the kernel, possibly
>hardware differences (I'm running under qemu). Bugs related to barriers
>(which this one seems to be) need just the right circumstances to
>trigger so they can be hard to reproduce on a different system.
>
>I've made some cosmetic improvements to the patch and will post it to
>the NFS maintainers.
>
>Thanks again,
>NeilBrown
>
>
>>
>> Thanks,
>> Richard
>>
>> 2024-05-24 07:29 időpontban Richard Kojedzinszky ezt írta:
>> > Dear Neil,
>> >
>> > I've applied your patch, and since then there are no lockups. Before
>> > that my application reported a lockup in a minute or two, now it has
>> > been running for half an hour, and still running.
>> >
>> > Thanks,
>> > Richard
>> >
>> > 2024-05-24 01:31 időpontban NeilBrown ezt írta:
>> >> On Fri, 24 May 2024, Richard Kojedzinszky wrote:
>> >>> Dear devs,
>> >>>
>> >>> I am attaching a stripped down version of the little program which
>> >>> triggers the bug very quickly, in a few minutes in my test lab. It
>> >>> turned out that a single NFS mountpoint is enough. Just start the
>> >>> program giving it the NFS mount as first argument. It will chdir
>> >>> there,
>> >>> and do file operations, which will trigger a lockup in a few minutes.
>> >>
>> >> I couldn't get the go code to run. But then it is a long time since I
>> >> played with go and I didn't try very hard.
>> >> If you could provide simple instructions and a list of package
>> >> dependencies that I need to install (on Debian), I can give it a try.
>> >>
>> >> Or you could try this patch. It might help, but I don't have high
>> >> hopes. It adds some memory barriers and fixes a bug which would cause
>> >> a
>> >> problem if memory allocation failed (but memory allocation never
>> >> fails).
>> >>
>> >> NeilBrown
>> >>
>> >> diff --git a/fs/nfs/dir.c b/fs/nfs/dir.c
>> >> index ac505671efbd..5bcc0d14d519 100644
>> >> --- a/fs/nfs/dir.c
>> >> +++ b/fs/nfs/dir.c
>> >> @@ -1804,7 +1804,7 @@ __nfs_lookup_revalidate(struct dentry *dentry,
>> >> unsigned int flags,
>> >> } else {
>> >> /* Wait for unlink to complete */
>> >> wait_var_event(&dentry->d_fsdata,
>> >> - dentry->d_fsdata != NFS_FSDATA_BLOCKED);
>> >> + smp_load_acquire(&dentry->d_fsdata) != NFS_FSDATA_BLOCKED);
>> >> parent = dget_parent(dentry);
>> >> ret = reval(d_inode(parent), dentry, flags);
>> >> dput(parent);
>> >> @@ -2508,7 +2508,7 @@ int nfs_unlink(struct inode *dir, struct dentry
>> >> *dentry)
>> >> spin_unlock(&dentry->d_lock);
>> >> error = nfs_safe_remove(dentry);
>> >> nfs_dentry_remove_handle_error(dir, dentry, error);
>> >> - dentry->d_fsdata = NULL;
>> >> + smp_store_release(&dentry->d_fsdata, NULL);
>> >> wake_up_var(&dentry->d_fsdata);
>> >> out:
>> >> trace_nfs_unlink_exit(dir, dentry, error);
>> >> @@ -2616,7 +2616,7 @@ nfs_unblock_rename(struct rpc_task *task, struct
>> >> nfs_renamedata *data)
>> >> {
>> >> struct dentry *new_dentry = data->new_dentry;
>> >>
>> >> - new_dentry->d_fsdata = NULL;
>> >> + smp_store_release(&new_dentry->d_fsdata, NULL);
>> >> wake_up_var(&new_dentry->d_fsdata);
>> >> }
>> >>
>> >> @@ -2717,6 +2717,10 @@ int nfs_rename(struct mnt_idmap *idmap, struct
>> >> inode *old_dir,
>> >> task = nfs_async_rename(old_dir, new_dir, old_dentry, new_dentry,
>> >> must_unblock ? nfs_unblock_rename : NULL);
>> >> if (IS_ERR(task)) {
>> >> + if (must_unlock) {
>> >> + smp_store_release(&new_dentry->d_fsdata, NULL);
>> >> + wake_up_var(&new_dentry->d_fsdata);
>> >> + }
>> >> error = PTR_ERR(task);
>> >> goto out;
>> >> }
>>
>
[toc] | [prev] | [next] | [standalone]
| From | "Debian Bug Tracking System" <owner@bugs.debian.org> |
|---|---|
| Date | 2024-06-28 00:20 +0200 |
| Subject | Bug#1071501: marked as done (linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate) |
| Message-ID | <IU5Dc-64H3-15@gated-at.bofh.it> |
| In reply to | #82499 |
[Multipart message — attachments visible in raw view] — view raw
Your message dated Thu, 27 Jun 2024 22:10:11 +0000 with message-id <E1sMxJj-000sZ1-JY@fasolo.debian.org> and subject line Bug#1071501: fixed in linux 6.9.7-1 has caused the Debian Bug report #1071501, regarding linux-image-6.1.0-21-arm64: Linux NFS client hangs in nfs4_lookup_revalidate to be marked as done. This means that you claim that the problem has been dealt with. If this is not the case it is now your responsibility to reopen the Bug report if necessary, and/or fix the problem forthwith. (NB: If you are a system administrator and have no idea what this message is talking about, this may indicate a serious mail system misconfiguration somewhere. Please contact owner@bugs.debian.org immediately.) -- 1071501: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1071501 Debian Bug Tracking System Contact owner@bugs.debian.org with problems
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.kernel
csiph-web