Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.kernel > #84819 > unrolled thread
| Started by | Bill Brelsford <wb@k2di.net> |
|---|---|
| First post | 2024-12-15 02:30 +0100 |
| Last post | 2025-01-17 04:30 +0100 |
| Articles | 14 — 3 participants |
Back to article view | Back to linux.debian.kernel
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2024-12-15 02:30 +0100
Processed: Re: Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault "Debian Bug Tracking System" <owner@bugs.debian.org> - 2024-12-15 14:20 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Salvatore Bonaccorso <carnil@debian.org> - 2024-12-15 14:20 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2024-12-16 04:30 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Salvatore Bonaccorso <carnil@debian.org> - 2024-12-16 21:40 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2024-12-17 03:40 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Salvatore Bonaccorso <carnil@debian.org> - 2024-12-17 23:20 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2024-12-19 06:50 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2024-12-19 16:30 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2025-01-15 03:00 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Salvatore Bonaccorso <carnil@debian.org> - 2025-01-15 07:20 +0100
Processed: Re: Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault "Debian Bug Tracking System" <owner@bugs.debian.org> - 2025-01-16 09:40 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Salvatore Bonaccorso <carnil@debian.org> - 2025-01-16 09:40 +0100
Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2025-01-17 04:30 +0100
| From | Bill Brelsford <wb@k2di.net> |
|---|---|
| Date | 2024-12-15 02:30 +0100 |
| Subject | Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault |
| Message-ID | <JTLCh-h5XL-1@gated-at.bofh.it> |
Package: nfs-kernel-server
Version: 1:2.8.2-1
Severity: important
Dear Maintainer,
Since upgrading from 1:2.8.1-2, nfs-kernel-server fails to start.
From the bootlog:
Sat Dec 14 15:35:50 2024: Starting NFS kernel daemon: nfsd/etc/init.d/nfs-kernel-server: line 58: 3169 Segmentation fault start-stop-daemon --start --oknodo --quiet --nicelevel $RPCNFSDPRIORITY --exec $PREFIX/sbin/rpc.nfsd
Sat Dec 14 15:35:50 2024: failed!
Tracing the init.d script shows that the segfault appears to be in
rpc.nfsd:
+ start-stop-daemon --start --oknodo --quiet --nicelevel 0 --exec /usr/sbin/rpc.nfsd
./nfs-kernel-server: line 61: 3694 Segmentation fault start-stop-daemon --start --oknodo --quiet --nicelevel $RPCNFSDPRIORITY --exec $PREFIX/sbin/rpc.nfs
Downgrading to 1:2.8.1-1 (from trixie; also nfs-common and
libnfsidmap1) resolves the problem.
Thanks.. Bill
-- Package-specific info:
-- rpcinfo --
program vers proto port service
100000 4 tcp 111 portmapper
100000 3 tcp 111 portmapper
100000 2 tcp 111 portmapper
100000 4 udp 111 portmapper
100000 3 udp 111 portmapper
100000 2 udp 111 portmapper
100024 1 udp 34570 status
100024 1 tcp 46351 status
100021 1 udp 51436 nlockmgr
100021 3 udp 51436 nlockmgr
100021 4 udp 51436 nlockmgr
100021 1 tcp 33553 nlockmgr
100021 3 tcp 33553 nlockmgr
100021 4 tcp 33553 nlockmgr
-- /etc/default/nfs-kernel-server --
RPCNFSDPRIORITY=0
NEED_SVCGSSD=""
-- /etc/nfs.conf --
[general]
pipefs-directory=/run/rpc_pipefs
[nfsrahead]
[exports]
[exportfs]
[gssd]
[lockd]
[exportd]
[mountd]
manage-gids=y
[nfsdcld]
[nfsdcltrack]
[nfsd]
[statd]
[sm-notify]
[svcgssd]
-- /etc/nfs.conf.d/*.conf --
-- System Information:
Debian Release: trixie/sid
APT prefers unstable
APT policy: (500, 'unstable')
Architecture: amd64 (x86_64)
Kernel: Linux 6.12.3-amd64 (SMP w/8 CPU threads; PREEMPT)
Locale: LANG=en_US.UTF-8, LC_CTYPE=en_US.UTF-8 (charmap=UTF-8), LANGUAGE not set
Shell: /bin/sh linked to /usr/bin/dash
Init: sysvinit (via /sbin/init)
LSM: AppArmor: enabled
Versions of packages nfs-kernel-server depends on:
ii keyutils 1.6.3-4
ii libblkid1 2.40.2-12
ii libc6 2.40-4
ii libcap2 1:2.66-5+b1
ii libevent-core-2.1-7t64 2.1.12-stable-10+b1
ii libnl-3-200 3.7.0-0.3+b1
ii libnl-genl-3-200 3.7.0-0.3+b1
ii libreadline8t64 8.2-6
ii libsqlite3-0 3.46.1-1
ii libtirpc3t64 1.3.4+ds-1.3+b1
ii libuuid1 2.40.2-12
ii libwrap0 7.6.q-34
ii libxml2 2.12.7+dfsg+really2.9.14-0.2+b1
ii netbase 6.4
ii nfs-common 1:2.8.2-1
ii ucf 3.0045
Versions of packages nfs-kernel-server recommends:
ii python3 3.12.7-1
pn python3-yaml <none>
Versions of packages nfs-kernel-server suggests:
ii procps 2:4.0.4-6
-- no debconf information
[toc] | [next] | [standalone]
| From | "Debian Bug Tracking System" <owner@bugs.debian.org> |
|---|---|
| Date | 2024-12-15 14:20 +0100 |
| Subject | Processed: Re: Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault |
| Message-ID | <JTWHn-hcYL-1@gated-at.bofh.it> |
| In reply to | #84819 |
Processing control commands: > tags -1 + unreproducible moreinfo Bug #1089976 [nfs-kernel-server] nfs-kernel-server: Fails to start - segmentation fault Added tag(s) moreinfo and unreproducible. -- 1089976: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1089976 Debian Bug Tracking System Contact owner@bugs.debian.org with problems
[toc] | [prev] | [next] | [standalone]
| From | Salvatore Bonaccorso <carnil@debian.org> |
|---|---|
| Date | 2024-12-15 14:20 +0100 |
| Message-ID | <JTWHn-hcYL-3@gated-at.bofh.it> |
| In reply to | #84819 |
Control: tags -1 + unreproducible moreinfo Hi Bill, On Sat, Dec 14, 2024 at 05:00:34PM -0800, Bill Brelsford wrote: > Package: nfs-kernel-server > Version: 1:2.8.2-1 > Severity: important > > Dear Maintainer, > > Since upgrading from 1:2.8.1-2, nfs-kernel-server fails to start. > >From the bootlog: > > Sat Dec 14 15:35:50 2024: Starting NFS kernel daemon: nfsd/etc/init.d/nfs-kernel-server: line 58: 3169 Segmentation fault start-stop-daemon --start --oknodo --quiet --nicelevel $RPCNFSDPRIORITY --exec $PREFIX/sbin/rpc.nfsd > Sat Dec 14 15:35:50 2024: failed! > > Tracing the init.d script shows that the segfault appears to be in > rpc.nfsd: > > + start-stop-daemon --start --oknodo --quiet --nicelevel 0 --exec /usr/sbin/rpc.nfsd > ./nfs-kernel-server: line 61: 3694 Segmentation fault start-stop-daemon --start --oknodo --quiet --nicelevel $RPCNFSDPRIORITY --exec $PREFIX/sbin/rpc.nfs > > Downgrading to 1:2.8.1-1 (from trixie; also nfs-common and > libnfsidmap1) resolves the problem. I'm not able to reproduce it here. Can you please Install the dbgsym packages as well and get more information by making sure the service is stopped and start it by hand under debugger. This might give some more clue for upstream. Regards, Salvatore
[toc] | [prev] | [next] | [standalone]
| From | Bill Brelsford <wb@k2di.net> |
|---|---|
| Date | 2024-12-16 04:30 +0100 |
| Message-ID | <JU9XX-hppY-1@gated-at.bofh.it> |
| In reply to | #84824 |
Hi Salvatore, On Sun Dec 15 2024 at 02:15 PM +0100, Salvatore Bonaccorso wrote: > I'm not able to reproduce it here. Can you please Install the dbgsym > packages as well and get more information by making sure the service > is stopped and start it by hand under debugger. > > This might give some more clue for upstream. I apparently need some help with gdb to get useful output. After running "find-dbgsym-packages /usr/sbin/rpc.nfsd", I installed libc6-dbg and nfs-kernel-server-dbgsym. Then I stopped /etc/init.d/nfs-kernel-server just before the call to rpc.nfsd and ran gdb: # gdb /usr/sbin/rpc.nfsd GNU gdb (Debian 15.2-1) 15.2 ... Reading symbols from /usr/sbin/rpc.nfsd... Reading symbols from /usr/lib/debug/.build-id/4c/400a0c5314bb3884d7adbde7889a7bbc3a0eaa.debug... (gdb) run Starting program: /usr/sbin/rpc.nfsd [Thread debugging using libthread_db enabled] Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1". Program terminated with signal SIGSEGV, Segmentation fault. The program no longer exists. (gdb) bt No stack. Suggestions? Also, after rpc.nfsd fails, running it again (or, e.g., exportfs -r) hangs and can't be killed. The same failure occurs on another similarly-configured system. Thanks.. Bill
[toc] | [prev] | [next] | [standalone]
| From | Salvatore Bonaccorso <carnil@debian.org> |
|---|---|
| Date | 2024-12-16 21:40 +0100 |
| Message-ID | <JUq2J-2FN-1@gated-at.bofh.it> |
| In reply to | #84839 |
Hi Bill, On Sun, Dec 15, 2024 at 06:46:46PM -0800, Bill Brelsford wrote: > Hi Salvatore, > > On Sun Dec 15 2024 at 02:15 PM +0100, Salvatore Bonaccorso wrote: > > I'm not able to reproduce it here. Can you please Install the dbgsym > > packages as well and get more information by making sure the service > > is stopped and start it by hand under debugger. > > > > This might give some more clue for upstream. > > I apparently need some help with gdb to get useful output. After > running "find-dbgsym-packages /usr/sbin/rpc.nfsd", I installed > libc6-dbg and nfs-kernel-server-dbgsym. Then I stopped > /etc/init.d/nfs-kernel-server just before the call to rpc.nfsd > and ran gdb: > > # gdb /usr/sbin/rpc.nfsd > GNU gdb (Debian 15.2-1) 15.2 > ... > Reading symbols from /usr/sbin/rpc.nfsd... > Reading symbols from /usr/lib/debug/.build-id/4c/400a0c5314bb3884d7adbde7889a7bbc3a0eaa.debug... > (gdb) run > Starting program: /usr/sbin/rpc.nfsd > [Thread debugging using libthread_db enabled] > Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1". > > Program terminated with signal SIGSEGV, Segmentation fault. > The program no longer exists. > (gdb) bt > No stack. > > Suggestions? > > Also, after rpc.nfsd fails, running it again (or, e.g., exportfs -r) > hangs and can't be killed. > > The same failure occurs on another similarly-configured system. Let's try to tackle it from another angle. What is commont on thoe system where you see the failure? Are all not using systemd as init? Additionally can you give some more details on your setup, how the exports look? Can we boild down the setup minimally to trigger the issue (and so report upstream)? Can you additionally please test to downgrade to the 2.8.1-2 version and please report back if you see the problem there was well? For us it should behave actually same as 2.8.1-1 but I would like to double check. the supported way is to run it with systemd, the ship'ed sysvinit scripts contain legacy, for instance with new version we would start the server with nfsdctl: ExecStart=/bin/sh -c '/usr/sbin/nfsdctl autostart || /usr/sbin/rpc.nfsd' If you start it through nfsdctl do you get the exports working? Still, having more information on the underlying setup would be helpful. Regards, Salvatore
[toc] | [prev] | [next] | [standalone]
| From | Bill Brelsford <wb@k2di.net> |
|---|---|
| Date | 2024-12-17 03:40 +0100 |
| Message-ID | <JUvF7-7HE-1@gated-at.bofh.it> |
| In reply to | #84848 |
Hi Salvatore, On Mon Dec 16 2024 at 09:36 PM +0100, Salvatore Bonaccorso wrote: > What is commont on thoe system where you see the failure? Are all not > using systemd as init? Additionally can you give some more details on > your setup, how the exports look? Can we boild down the setup > minimally to trigger the issue (and so report upstream)? The other system is an old one that I maintain as a backup; it also uses sysvinit. The nfs setup for both is simple, with no changes to /etc/nfs.conf or /etc/default/nfs-kernel-server. /etc/exports on the newer system: / 10.20.40.80/29(rw,fsid=8208,no_root_squash,no_subtree_check,sync) /u 10.20.40.80/29(rw,fsid=8202,no_root_squash,no_subtree_check,sync) /u1 10.20.40.80/29(rw,fsid=8203,no_root_squash,no_subtree_check,sync) /n/ws 10.20.40.80/29(rw,fsid=8206,no_root_squash,no_subtree_check,sync) /n/wt 10.20.40.80/29(rw,fsid=8207,no_root_squash,no_subtree_check,sync) It had been working fine on both systems prior to the 2.8.2-1 upgrade. It also works on 5 bookworm systems with the same setup. > Can you additionally please test to downgrade to the 2.8.1-2 version > and please report back if you see the problem there was well? For us > it should behave actually same as 2.8.1-1 but I would like to double > check. I downgraded to 2.8.1-1 because it was available in the trixie repository. Where can I get 2.8.1-2? I also have trixie installed on both systems. Trixie now uses 2.8.2-1 so I upgraded one of them -- it also fails. > the supported way is to run it with systemd, the ship'ed sysvinit > scripts contain legacy, for instance with new version we would start > the server with nfsdctl: > > ExecStart=/bin/sh -c '/usr/sbin/nfsdctl autostart || /usr/sbin/rpc.nfsd' > > If you start it through nfsdctl do you get the exports working? I haven't installed systemd -- nothing I use has required it -- so nfsdctl isn't available. > Still, having more information on the underlying setup would be > helpful. Another possible clue: after upgrading (to 2.8.2-1), there is no problem -- I can start and stop the daemon, export/unexport filesystems, etc. Everything seems normal -- until the system is rebooted and rpc.nfsd is invoked. Regards.. Bill
[toc] | [prev] | [next] | [standalone]
| From | Salvatore Bonaccorso <carnil@debian.org> |
|---|---|
| Date | 2024-12-17 23:20 +0100 |
| Message-ID | <JUO53-jlb-1@gated-at.bofh.it> |
| In reply to | #84852 |
Hi Bill, On Mon, Dec 16, 2024 at 06:30:07PM -0800, Bill Brelsford wrote: > Hi Salvatore, > > On Mon Dec 16 2024 at 09:36 PM +0100, Salvatore Bonaccorso wrote: > > What is commont on thoe system where you see the failure? Are all not > > using systemd as init? Additionally can you give some more details on > > your setup, how the exports look? Can we boild down the setup > > minimally to trigger the issue (and so report upstream)? > > The other system is an old one that I maintain as a backup; it also > uses sysvinit. The nfs setup for both is simple, with no changes to > /etc/nfs.conf or /etc/default/nfs-kernel-server. /etc/exports on > the newer system: > > / 10.20.40.80/29(rw,fsid=8208,no_root_squash,no_subtree_check,sync) > /u 10.20.40.80/29(rw,fsid=8202,no_root_squash,no_subtree_check,sync) > /u1 10.20.40.80/29(rw,fsid=8203,no_root_squash,no_subtree_check,sync) > /n/ws 10.20.40.80/29(rw,fsid=8206,no_root_squash,no_subtree_check,sync) > /n/wt 10.20.40.80/29(rw,fsid=8207,no_root_squash,no_subtree_check,sync) > > It had been working fine on both systems prior to the 2.8.2-1 > upgrade. It also works on 5 bookworm systems with the same setup. Thanks. Unfortunately still no look in a lab setup to trigger your issue. When rpc.segfaults, are there any other nfs related processed and threads running on the system? > > Can you additionally please test to downgrade to the 2.8.1-2 version > > and please report back if you see the problem there was well? For us > > it should behave actually same as 2.8.1-1 but I would like to double > > check. > > I downgraded to 2.8.1-1 because it was available in the trixie > repository. Where can I get 2.8.1-2? It is not anymore available in the archive as it is superseeded. But you can find it on snapshot.d.o: https://snapshot.debian.org/package/nfs-utils/1%3A2.8.1-2/ > > I also have trixie installed on both systems. Trixie now uses > 2.8.2-1 so I upgraded one of them -- it also fails. > > > the supported way is to run it with systemd, the ship'ed sysvinit > > scripts contain legacy, for instance with new version we would start > > the server with nfsdctl: > > > > ExecStart=/bin/sh -c '/usr/sbin/nfsdctl autostart || /usr/sbin/rpc.nfsd' > > > > If you start it through nfsdctl do you get the exports working? > > I haven't installed systemd -- nothing I use has required it -- so > nfsdctl isn't available. nfsdctl is independent of systemd, it is shipped in as the new tool(ing) for starting the nfsd server: /usr/sbin/nfsdctl from nfs-kernel-server package. > > Still, having more information on the underlying setup would be > > helpful. > > Another possible clue: after upgrading (to 2.8.2-1), there is no > problem -- I can start and stop the daemon, export/unexport > filesystems, etc. Everything seems normal -- until the system is > rebooted and rpc.nfsd is invoked. As for the first part: What processes and kernel threads are started on the system at this stage? Regards, Salvatore
[toc] | [prev] | [next] | [standalone]
| From | Bill Brelsford <wb@k2di.net> |
|---|---|
| Date | 2024-12-19 06:50 +0100 |
| Message-ID | <JVhA5-MN8-1@gated-at.bofh.it> |
| In reply to | #84863 |
Hi Salvatore, On Tue Dec 17 2024 at 11:12 PM +0100, Salvatore Bonaccorso wrote: > Thanks. Unfortunately still no look in a lab setup to trigger your > issue. When rpc.segfaults, are there any other nfs related processed > and threads running on the system? ps -eLf shows only one: [kworker/R-nfsiod] > > > Can you additionally please test to downgrade to the 2.8.1-2 version > > > and please report back if you see the problem there was well? For us > > > it should behave actually same as 2.8.1-1 but I would like to double > > > check. Thanks for the pointer to snapshot.d.o. Yes, 2.8.1-2 also has the problem. > > > If you start it through nfsdctl do you get the exports working? Exports/exportfs are already working. Trying to start with nfsdctl instead of /etc/init.d/nfs-kernel-server (in 2.8.2-1) gives # nfsdctl -d autostart nfsdctl> autostart Error: Device or resource busy Error: Operation not permitted > > Another possible clue: after upgrading (to 2.8.2-1), there is no > > problem -- I can start and stop the daemon, export/unexport > > filesystems, etc. Everything seems normal -- until the system is > > rebooted and rpc.nfsd is invoked. > > As for the first part: What processes and kernel threads are started > on the system at this stage? ps -eLf gives 16 [nfsd] as well as the [kworker/R-nfsiod]. And with the daemon started, some nfsdctl commands work: # nfsdctl status # nfsdctl threads gracetime: 90 leasetime: 90 scope: k2ww pool-threads: 16 # nfsdctl listener tcp:[::]:2049 tcp:0.0.0.0:2049 # nfsdctl version +3.0 +4.0 +4.1 +4.2 Hope this helps. Thanks.. Bill
[toc] | [prev] | [next] | [standalone]
| From | Bill Brelsford <wb@k2di.net> |
|---|---|
| Date | 2024-12-19 16:30 +0100 |
| Message-ID | <JVqDn-UiF-11@gated-at.bofh.it> |
| In reply to | #84873 |
On Wed Dec 18 2024 at 09:43 PM -0800, Bill Brelsford wrote: > > > > Can you additionally please test to downgrade to the 2.8.1-2 version > > > > and please report back if you see the problem there was well? For us > > > > it should behave actually same as 2.8.1-1 but I would like to double > > > > check. > > Thanks for the pointer to snapshot.d.o. Yes, 2.8.1-2 also has the > problem. No -- my mistake! Downgrading to 2.8.1-2 works (as does 2.8.1-1). Bill
[toc] | [prev] | [next] | [standalone]
| From | Bill Brelsford <wb@k2di.net> |
|---|---|
| Date | 2025-01-15 03:00 +0100 |
| Message-ID | <K50Rj-910R-5@gated-at.bofh.it> |
| In reply to | #84876 |
Hi Salvatore, The problem has apparently been fixed in unstable by updates to one or more packages since December 19. Version 1:2.8.2-1 now works as expected on both of my machines. But it still fails in testing (trixie). I'll try to determine what future package update fixes it. Regards.. Bill
[toc] | [prev] | [next] | [standalone]
| From | Salvatore Bonaccorso <carnil@debian.org> |
|---|---|
| Date | 2025-01-15 07:20 +0100 |
| Message-ID | <K54UV-949u-1@gated-at.bofh.it> |
| In reply to | #85141 |
Hi Bill, On Tue, Jan 14, 2025 at 05:50:02PM -0800, Bill Brelsford wrote: > Hi Salvatore, > > The problem has apparently been fixed in unstable by updates to one > or more packages since December 19. Version 1:2.8.2-1 now works as > expected on both of my machines. > > But it still fails in testing (trixie). I'll try to determine what > future package update fixes it. I have a suspect what it can be. Can you please post the kernel log / dmesg from the systems which do not work please? Regards, Salvatore
[toc] | [prev] | [next] | [standalone]
| From | "Debian Bug Tracking System" <owner@bugs.debian.org> |
|---|---|
| Date | 2025-01-16 09:40 +0100 |
| Subject | Processed: Re: Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault |
| Message-ID | <K5tzX-9jKC-1@gated-at.bofh.it> |
| In reply to | #84819 |
Processing control commands:
> reassign -1 src:linux
Bug #1089976 [nfs-kernel-server] nfs-kernel-server: Fails to start - segmentation fault
Bug reassigned from package 'nfs-kernel-server' to 'src:linux'.
No longer marked as found in versions nfs-utils/1:2.8.2-1.
Ignoring request to alter fixed versions of bug #1089976 to the same values previously set
> forcemerge 1087900 -1
Bug #1087900 {Done: Bastian Blank <waldi@debian.org>} [src:linux] linux: kernel BUG at fs/nfsd/nfs4recover.c:534
Bug #1091439 {Done: Bastian Blank <waldi@debian.org>} [src:linux] installation of nfs-kernel-server hangs
Bug #1092607 {Done: Bastian Blank <waldi@debian.org>} [src:linux] dracut: upstream-dracut-network-nfs autopkgtest fails on amd64
Bug #1087900 {Done: Bastian Blank <waldi@debian.org>} [src:linux] linux: kernel BUG at fs/nfsd/nfs4recover.c:534
Added tag(s) unreproducible and moreinfo.
Added tag(s) moreinfo and unreproducible.
Added tag(s) unreproducible and moreinfo.
Bug #1089976 [src:linux] nfs-kernel-server: Fails to start - segmentation fault
Set Bug forwarded-to-address to 'https://lore.kernel.org/linux-nfs/Z22DIiV98XBSfPVr@eldamar.lan/'.
Marked Bug as done
Added indication that 1089976 affects src:dracut
Marked as fixed in versions linux/6.12.8-1 and linux/6.13~rc6-1~exp1.
Marked as found in versions linux/6.8.9-1, linux/6.12~rc6-1~exp1, linux/6.11.9-1, linux/6.11.5-1, and linux/6.12.6-1.
Added tag(s) upstream and confirmed.
Bug #1091439 {Done: Bastian Blank <waldi@debian.org>} [src:linux] installation of nfs-kernel-server hangs
Bug #1092607 {Done: Bastian Blank <waldi@debian.org>} [src:linux] dracut: upstream-dracut-network-nfs autopkgtest fails on amd64
Merged 1087900 1089976 1091439 1092607
--
1087900: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1087900
1089976: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1089976
1091439: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1091439
1092607: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1092607
Debian Bug Tracking System
Contact owner@bugs.debian.org with problems
[toc] | [prev] | [next] | [standalone]
| From | Salvatore Bonaccorso <carnil@debian.org> |
|---|---|
| Date | 2025-01-16 09:40 +0100 |
| Message-ID | <K5tzX-9jKC-3@gated-at.bofh.it> |
| In reply to | #84819 |
Control: reassign -1 src:linux Control: forcemerge 1087900 -1 Hi Bill, On Wed, Jan 15, 2025 at 08:02:18AM -0800, Bill Brelsford wrote: > On Wed Jan 15 2025 at 07:08 AM +0100, Salvatore Bonaccorso wrote: > > > But it still fails in testing (trixie). I'll try to determine what > > > future package update fixes it. > > > > I have a suspect what it can be. Can you please post the kernel log / > > dmesg from the systems which do not work please? > > Attached is dmesg from trixie. > > Bill [...] > [ 33.734128] r8169 0000:05:07.0 eth0: Link is Up - 100Mbps/Full - flow control rx/tx > [ 40.873091] NFSD: Using /var/lib/nfs/v4recovery as the NFSv4 state recovery directory > [ 40.967422] NFSD: Using legacy client tracking operations. > [ 40.971774] NFSD: Using /var/lib/nfs/v4recovery as the NFSv4 state recovery directory > [ 40.976080] ------------[ cut here ]------------ > [ 40.980166] kernel BUG at fs/nfsd/nfs4recover.c:534! > [ 40.984266] Oops: invalid opcode: 0000 [#1] PREEMPT SMP PTI > [ 40.988166] CPU: 0 UID: 0 PID: 1935 Comm: rpc.nfsd Not tainted 6.12.6-amd64 #1 Debian 6.12.6-1 > [ 40.988236] Hardware name: To Be Filled By O.E.M. S62E/S62E, BIOS 0303 08/03/2007 > [ 40.988236] RIP: 0010:nfsd4_legacy_tracking_init+0x17d/0x1b0 [nfsd] > [ 40.988236] Code: 19 48 89 de 48 c7 c7 50 99 cb c1 e8 6d fb ff ff 89 c5 85 c0 0f 85 10 61 00 00 48 c7 c7 90 f3 d2 c1 31 ed e8 85 26 3e f8 eb 07 <0f> 0b bd f4 ff ff ff 48 8b 44 24 08 65 48 2b 04 25 28 00 00 00 75 > [ 40.988236] RSP: 0018:ffffbc09c0f87c80 EFLAGS: 00010282 > [ 40.988236] RAX: 0000000000000049 RBX: ffffa0364b4d8000 RCX: 0000000000000003 > [ 40.988236] RDX: 0000000000000000 RSI: 0000000000000003 RDI: 0000000000000001 > [ 40.988236] RBP: ffffffffbbca3600 R08: 0000000000000000 R09: ffffbc09c0f87b10 > [ 40.988236] R10: ffffffffbb0b4348 R11: 0000000000000003 R12: ffffa0364b4d8000 > [ 40.988236] R13: ffffa0364b4d8000 R14: ffffa0364c306b40 R15: ffffa0364b4d8000 > [ 40.988236] FS: 00007faff2288740(0000) GS:ffffa036bd400000(0000) knlGS:0000000000000000 > [ 40.988236] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 > [ 40.988236] CR2: 00007f0d8a7e1320 CR3: 0000000002cfc000 CR4: 00000000000026f0 > [ 40.988236] Call Trace: > [ 40.988236] <TASK> > [ 40.988236] ? __die_body.cold+0x19/0x27 > [ 40.988236] ? die+0x2e/0x50 > [ 40.988236] ? do_trap+0xca/0x110 > [ 40.988236] ? do_error_trap+0x6a/0x90 > [ 40.988236] ? nfsd4_legacy_tracking_init+0x17d/0x1b0 [nfsd] > [ 40.988236] ? exc_invalid_op+0x50/0x70 > [ 40.988236] ? nfsd4_legacy_tracking_init+0x17d/0x1b0 [nfsd] > [ 40.988236] ? asm_exc_invalid_op+0x1a/0x20 > [ 40.988236] ? nfsd4_legacy_tracking_init+0x17d/0x1b0 [nfsd] > [ 40.988236] nfsd4_client_tracking_init+0x57/0x1b0 [nfsd] > [ 40.988236] nfs4_state_start_net+0x2f9/0x3a0 [nfsd] > [ 40.988236] nfsd_svc+0x1ac/0x310 [nfsd] > [ 40.988236] write_threads+0xf9/0x1c0 [nfsd] > [ 40.988236] ? __pfx_write_threads+0x10/0x10 [nfsd] > [ 40.988236] nfsctl_transaction_write+0x4a/0x80 [nfsd] > [ 40.988236] vfs_write+0xf8/0x450 > [ 40.988236] ksys_write+0x6d/0xf0 > [ 40.988236] do_syscall_64+0x82/0x190 > [ 40.988236] ? do_user_addr_fault+0x36c/0x620 > [ 40.988236] ? exc_page_fault+0x7e/0x180 > [ 40.988236] entry_SYSCALL_64_after_hwframe+0x76/0x7e > [ 40.988236] RIP: 0033:0x7faff238f090 > [ 40.988236] Code: 2d 0e 00 64 c7 00 16 00 00 00 b8 ff ff ff ff c3 66 2e 0f 1f 84 00 00 00 00 00 80 3d d9 af 0e 00 00 74 17 b8 01 00 00 00 0f 05 <48> 3d 00 f0 ff ff 77 58 c3 0f 1f 80 00 00 00 00 48 83 ec 28 48 89 > [ 40.988236] RSP: 002b:00007ffeff5e5298 EFLAGS: 00000202 ORIG_RAX: 0000000000000001 > [ 40.988236] RAX: ffffffffffffffda RBX: 0000000000000003 RCX: 00007faff238f090 > [ 40.988236] RDX: 0000000000000003 RSI: 000055744ca80340 RDI: 0000000000000003 > [ 40.988236] RBP: 000055744ca80340 R08: 0000000000000064 R09: 00000000fffffffe > [ 40.988236] R10: 0000000000000000 R11: 0000000000000202 R12: 0000000000020000 > [ 40.988236] R13: 000055744ca7c116 R14: 000055748c50c2a0 R15: 0000000000000000 > [ 40.988236] </TASK> > [ 40.988236] Modules linked in: xt_nat ipt_REJECT nf_reject_ipv4 xt_conntrack xt_tcpudp iptable_mangle iptable_nat nf_nat nf_conntrack nf_defrag_ipv6 nf_defrag_ipv4 libcrc32c iptable_filter ip_tables x_tables nfsd auth_rpcgss nfs_acl nfs lockd grace netfs sunrpc firewire_sbp2 sr_mod at24 cdrom iTCO_wdt intel_pmc_bxt iTCO_vendor_support watchdog coretemp i915 kvm_intel snd_hda_codec_si3054 snd_hda_codec_realtek kvm snd_hda_codec_generic snd_hda_scodec_component snd_hda_intel uvcvideo drm_buddy iwl4965 snd_intel_dspcfg snd_intel_sdw_acpi sha512_ssse3 drm_display_helper snd_hda_codec sha256_ssse3 cec videobuf2_vmalloc iwlegacy uvc snd_hda_core videobuf2_memops rc_core sha1_ssse3 snd_hwdep mac80211 videobuf2_v4l2 snd_pcm_oss ttm r852 snd_mixer_oss videodev snd_pcm sm_common drm_kms_helper firewire_ohci nand r8169 snd_timer pcspkr drm firewire_core libarc4 i2c_i801 nandcore joydev sdhci_pci videobuf2_common snd asus_laptop acpi_cpufreq cfg80211 cqhci realtek sparse_keymap i2c_smbus sdhci mdio_devres bch r592 ata_generic > [ 40.988236] mmc_core mc libphy i2c_algo_bit soundcore mtd crc_itu_t memstick sg ata_piix rfkill video ehci_pci ac lpc_ich wmi button battery ext4 crc16 mbcache jbd2 crc32c_generic xts dm_crypt dm_mod hid_generic usbhid hid sd_mod uhci_hcd ehci_hcd ahci libahci usbcore psmouse libata evdev scsi_mod serio_raw usb_common scsi_common > [ 41.220014] ---[ end trace 0000000000000000 ]--- > [ 41.221857] RIP: 0010:nfsd4_legacy_tracking_init+0x17d/0x1b0 [nfsd] > [ 41.223642] Code: 19 48 89 de 48 c7 c7 50 99 cb c1 e8 6d fb ff ff 89 c5 85 c0 0f 85 10 61 00 00 48 c7 c7 90 f3 d2 c1 31 ed e8 85 26 3e f8 eb 07 <0f> 0b bd f4 ff ff ff 48 8b 44 24 08 65 48 2b 04 25 28 00 00 00 75 > [ 41.227245] RSP: 0018:ffffbc09c0f87c80 EFLAGS: 00010282 > [ 41.229058] RAX: 0000000000000049 RBX: ffffa0364b4d8000 RCX: 0000000000000003 > [ 41.230870] RDX: 0000000000000000 RSI: 0000000000000003 RDI: 0000000000000001 > [ 41.232683] RBP: ffffffffbbca3600 R08: 0000000000000000 R09: ffffbc09c0f87b10 > [ 41.234502] R10: ffffffffbb0b4348 R11: 0000000000000003 R12: ffffa0364b4d8000 > [ 41.236326] R13: ffffa0364b4d8000 R14: ffffa0364c306b40 R15: ffffa0364b4d8000 > [ 41.238171] FS: 00007faff2288740(0000) GS:ffffa036bd400000(0000) knlGS:0000000000000000 > [ 41.240047] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 > [ 41.241942] CR2: 00007f0d8a7e1320 CR3: 0000000002cfc000 CR4: 00000000000026f0 > [ 43.937663] warning: `iwconfig' uses wireless extensions which will stop working for Wi-Fi 7 hardware; use nl80211 > [ 48.021236] Key type dns_resolver registered > [ 48.660788] NFS: Registering the id_resolver key type > [ 48.662980] Key type id_resolver registered > [ 48.665164] Key type id_legacy registered So this is what I suspected. This is #1087900 in src:linux and fixed in 6.12.8-1. I'm reassigning and merging the two bugs. *But* that said, I strongly encourage you to switch to a systemd running system for your nfs server. We still ship the init scripts in the packaging, but for instance you do not start with them the more advanced client tracking daemons. The legacy tracking methods will disapear at some point upstream completely and in fact for the next experimental upload I aim to disable NFSD_LEGACY_CLIENT_TRACKING (cf. https://salsa.debian.org/kernel-team/linux/-/merge_requests/1298). Thanks for the report! Regards, Salvatore
[toc] | [prev] | [next] | [standalone]
| From | Bill Brelsford <wb@k2di.net> |
|---|---|
| Date | 2025-01-17 04:30 +0100 |
| Message-ID | <K5Ldv-9vYk-1@gated-at.bofh.it> |
| In reply to | #85159 |
Hi Salvatore, On Thu Jan 16 2025 at 09:30 AM +0100, Salvatore Bonaccorso wrote: > On Wed, Jan 15, 2025 at 08:02:18AM -0800, Bill Brelsford wrote: > > [ 40.980166] kernel BUG at fs/nfsd/nfs4recover.c:534! > > [ 40.984266] Oops: invalid opcode: 0000 [#1] PREEMPT SMP PTI > So this is what I suspected. This is #1087900 in src:linux and fixed > in 6.12.8-1. I'm embarassed! I thought I had checked dmesg, but obviously hadn't. It's working on trixie now with the latest kernel update (6.12.9-1). > *But* that said, I strongly encourage you to switch to a systemd > running system for your nfs server. We still ship the init scripts in > the packaging, but for instance you do not start with them the more > advanced client tracking daemons. The legacy tracking methods will > disapear at some point upstream completely and in fact for the next > experimental upload I aim to disable NFSD_LEGACY_CLIENT_TRACKING > (cf. > https://salsa.debian.org/kernel-team/linux/-/merge_requests/1298). Yes, it's time for me to switch to systemd. Thanks very much for your help! Bill
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.kernel
csiph-web