Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #84819 > unrolled thread

Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault

Started byBill Brelsford <wb@k2di.net>
First post2024-12-15 02:30 +0100
Last post2025-01-17 04:30 +0100
Articles 14 — 3 participants

Back to article view | Back to linux.debian.kernel


Contents

  Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2024-12-15 02:30 +0100
    Processed: Re: Bug#1089976: nfs-kernel-server: Fails to start -  segmentation fault "Debian Bug Tracking System" <owner@bugs.debian.org> - 2024-12-15 14:20 +0100
    Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Salvatore Bonaccorso <carnil@debian.org> - 2024-12-15 14:20 +0100
      Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2024-12-16 04:30 +0100
        Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Salvatore Bonaccorso <carnil@debian.org> - 2024-12-16 21:40 +0100
          Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2024-12-17 03:40 +0100
            Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Salvatore Bonaccorso <carnil@debian.org> - 2024-12-17 23:20 +0100
              Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2024-12-19 06:50 +0100
                Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2024-12-19 16:30 +0100
                  Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2025-01-15 03:00 +0100
                    Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Salvatore Bonaccorso <carnil@debian.org> - 2025-01-15 07:20 +0100
    Processed: Re: Bug#1089976: nfs-kernel-server: Fails to start -  segmentation fault "Debian Bug Tracking System" <owner@bugs.debian.org> - 2025-01-16 09:40 +0100
    Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Salvatore Bonaccorso <carnil@debian.org> - 2025-01-16 09:40 +0100
      Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault Bill Brelsford <wb@k2di.net> - 2025-01-17 04:30 +0100

#84819 — Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault

FromBill Brelsford <wb@k2di.net>
Date2024-12-15 02:30 +0100
SubjectBug#1089976: nfs-kernel-server: Fails to start - segmentation fault
Message-ID<JTLCh-h5XL-1@gated-at.bofh.it>
Package: nfs-kernel-server
Version: 1:2.8.2-1
Severity: important

Dear Maintainer,

Since upgrading from 1:2.8.1-2, nfs-kernel-server fails to start.
From the bootlog:

  Sat Dec 14 15:35:50 2024: Starting NFS kernel daemon: nfsd/etc/init.d/nfs-kernel-server: line 58:  3169 Segmentation fault      start-stop-daemon --start --oknodo --quiet --nicelevel $RPCNFSDPRIORITY --exec $PREFIX/sbin/rpc.nfsd
  Sat Dec 14 15:35:50 2024:  failed!

Tracing the init.d script shows that the segfault appears to be in
rpc.nfsd:

  + start-stop-daemon --start --oknodo --quiet --nicelevel 0 --exec /usr/sbin/rpc.nfsd
  ./nfs-kernel-server: line 61:  3694 Segmentation fault      start-stop-daemon --start --oknodo --quiet --nicelevel $RPCNFSDPRIORITY --exec $PREFIX/sbin/rpc.nfs

Downgrading to 1:2.8.1-1 (from trixie; also nfs-common and
libnfsidmap1) resolves the problem.

Thanks..  Bill

-- Package-specific info:
-- rpcinfo --
   program vers proto   port  service
    100000    4   tcp    111  portmapper
    100000    3   tcp    111  portmapper
    100000    2   tcp    111  portmapper
    100000    4   udp    111  portmapper
    100000    3   udp    111  portmapper
    100000    2   udp    111  portmapper
    100024    1   udp  34570  status
    100024    1   tcp  46351  status
    100021    1   udp  51436  nlockmgr
    100021    3   udp  51436  nlockmgr
    100021    4   udp  51436  nlockmgr
    100021    1   tcp  33553  nlockmgr
    100021    3   tcp  33553  nlockmgr
    100021    4   tcp  33553  nlockmgr
-- /etc/default/nfs-kernel-server --
RPCNFSDPRIORITY=0
NEED_SVCGSSD=""
-- /etc/nfs.conf --
[general]
pipefs-directory=/run/rpc_pipefs
[nfsrahead]
[exports]
[exportfs]
[gssd]
[lockd]
[exportd]
[mountd]
manage-gids=y
[nfsdcld]
[nfsdcltrack]
[nfsd]
[statd]
[sm-notify]
[svcgssd]
-- /etc/nfs.conf.d/*.conf --

-- System Information:
Debian Release: trixie/sid
  APT prefers unstable
  APT policy: (500, 'unstable')
Architecture: amd64 (x86_64)

Kernel: Linux 6.12.3-amd64 (SMP w/8 CPU threads; PREEMPT)
Locale: LANG=en_US.UTF-8, LC_CTYPE=en_US.UTF-8 (charmap=UTF-8), LANGUAGE not set
Shell: /bin/sh linked to /usr/bin/dash
Init: sysvinit (via /sbin/init)
LSM: AppArmor: enabled

Versions of packages nfs-kernel-server depends on:
ii  keyutils                1.6.3-4
ii  libblkid1               2.40.2-12
ii  libc6                   2.40-4
ii  libcap2                 1:2.66-5+b1
ii  libevent-core-2.1-7t64  2.1.12-stable-10+b1
ii  libnl-3-200             3.7.0-0.3+b1
ii  libnl-genl-3-200        3.7.0-0.3+b1
ii  libreadline8t64         8.2-6
ii  libsqlite3-0            3.46.1-1
ii  libtirpc3t64            1.3.4+ds-1.3+b1
ii  libuuid1                2.40.2-12
ii  libwrap0                7.6.q-34
ii  libxml2                 2.12.7+dfsg+really2.9.14-0.2+b1
ii  netbase                 6.4
ii  nfs-common              1:2.8.2-1
ii  ucf                     3.0045

Versions of packages nfs-kernel-server recommends:
ii  python3       3.12.7-1
pn  python3-yaml  <none>

Versions of packages nfs-kernel-server suggests:
ii  procps  2:4.0.4-6

-- no debconf information

[toc] | [next] | [standalone]


#84823 — Processed: Re: Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault

From"Debian Bug Tracking System" <owner@bugs.debian.org>
Date2024-12-15 14:20 +0100
SubjectProcessed: Re: Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault
Message-ID<JTWHn-hcYL-1@gated-at.bofh.it>
In reply to#84819
Processing control commands:

> tags -1 + unreproducible moreinfo
Bug #1089976 [nfs-kernel-server] nfs-kernel-server: Fails to start - segmentation fault
Added tag(s) moreinfo and unreproducible.

-- 
1089976: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1089976
Debian Bug Tracking System
Contact owner@bugs.debian.org with problems

[toc] | [prev] | [next] | [standalone]


#84824

FromSalvatore Bonaccorso <carnil@debian.org>
Date2024-12-15 14:20 +0100
Message-ID<JTWHn-hcYL-3@gated-at.bofh.it>
In reply to#84819
Control: tags -1 + unreproducible moreinfo

Hi Bill,

On Sat, Dec 14, 2024 at 05:00:34PM -0800, Bill Brelsford wrote:
> Package: nfs-kernel-server
> Version: 1:2.8.2-1
> Severity: important
> 
> Dear Maintainer,
> 
> Since upgrading from 1:2.8.1-2, nfs-kernel-server fails to start.
> >From the bootlog:
> 
>   Sat Dec 14 15:35:50 2024: Starting NFS kernel daemon: nfsd/etc/init.d/nfs-kernel-server: line 58:  3169 Segmentation fault      start-stop-daemon --start --oknodo --quiet --nicelevel $RPCNFSDPRIORITY --exec $PREFIX/sbin/rpc.nfsd
>   Sat Dec 14 15:35:50 2024:  failed!
> 
> Tracing the init.d script shows that the segfault appears to be in
> rpc.nfsd:
> 
>   + start-stop-daemon --start --oknodo --quiet --nicelevel 0 --exec /usr/sbin/rpc.nfsd
>   ./nfs-kernel-server: line 61:  3694 Segmentation fault      start-stop-daemon --start --oknodo --quiet --nicelevel $RPCNFSDPRIORITY --exec $PREFIX/sbin/rpc.nfs
> 
> Downgrading to 1:2.8.1-1 (from trixie; also nfs-common and
> libnfsidmap1) resolves the problem.

I'm not able to reproduce it here. Can you please Install the dbgsym
packages as well and get more information by making sure the service
is stopped and start it by hand under debugger.

This might give some more clue for upstream.

Regards,
Salvatore

[toc] | [prev] | [next] | [standalone]


#84839

FromBill Brelsford <wb@k2di.net>
Date2024-12-16 04:30 +0100
Message-ID<JU9XX-hppY-1@gated-at.bofh.it>
In reply to#84824
Hi Salvatore,

On Sun Dec 15 2024 at 02:15 PM +0100, Salvatore Bonaccorso wrote:
> I'm not able to reproduce it here. Can you please Install the dbgsym
> packages as well and get more information by making sure the service
> is stopped and start it by hand under debugger.
> 
> This might give some more clue for upstream.

I apparently need some help with gdb to get useful output. After
running "find-dbgsym-packages /usr/sbin/rpc.nfsd", I installed
libc6-dbg and nfs-kernel-server-dbgsym. Then I stopped
/etc/init.d/nfs-kernel-server just before the call to rpc.nfsd
and ran gdb:

   # gdb /usr/sbin/rpc.nfsd
   GNU gdb (Debian 15.2-1) 15.2
   ...
   Reading symbols from /usr/sbin/rpc.nfsd...
   Reading symbols from /usr/lib/debug/.build-id/4c/400a0c5314bb3884d7adbde7889a7bbc3a0eaa.debug...
   (gdb) run
   Starting program: /usr/sbin/rpc.nfsd 
   [Thread debugging using libthread_db enabled]
   Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".

   Program terminated with signal SIGSEGV, Segmentation fault.
   The program no longer exists.
   (gdb) bt
   No stack.

Suggestions?

Also, after rpc.nfsd fails, running it again (or, e.g., exportfs -r)
hangs and can't be killed.

The same failure occurs on another similarly-configured system.

Thanks..  Bill

[toc] | [prev] | [next] | [standalone]


#84848

FromSalvatore Bonaccorso <carnil@debian.org>
Date2024-12-16 21:40 +0100
Message-ID<JUq2J-2FN-1@gated-at.bofh.it>
In reply to#84839
Hi Bill,

On Sun, Dec 15, 2024 at 06:46:46PM -0800, Bill Brelsford wrote:
> Hi Salvatore,
> 
> On Sun Dec 15 2024 at 02:15 PM +0100, Salvatore Bonaccorso wrote:
> > I'm not able to reproduce it here. Can you please Install the dbgsym
> > packages as well and get more information by making sure the service
> > is stopped and start it by hand under debugger.
> > 
> > This might give some more clue for upstream.
> 
> I apparently need some help with gdb to get useful output. After
> running "find-dbgsym-packages /usr/sbin/rpc.nfsd", I installed
> libc6-dbg and nfs-kernel-server-dbgsym. Then I stopped
> /etc/init.d/nfs-kernel-server just before the call to rpc.nfsd
> and ran gdb:
> 
>    # gdb /usr/sbin/rpc.nfsd
>    GNU gdb (Debian 15.2-1) 15.2
>    ...
>    Reading symbols from /usr/sbin/rpc.nfsd...
>    Reading symbols from /usr/lib/debug/.build-id/4c/400a0c5314bb3884d7adbde7889a7bbc3a0eaa.debug...
>    (gdb) run
>    Starting program: /usr/sbin/rpc.nfsd 
>    [Thread debugging using libthread_db enabled]
>    Using host libthread_db library "/lib/x86_64-linux-gnu/libthread_db.so.1".
> 
>    Program terminated with signal SIGSEGV, Segmentation fault.
>    The program no longer exists.
>    (gdb) bt
>    No stack.
> 
> Suggestions?
> 
> Also, after rpc.nfsd fails, running it again (or, e.g., exportfs -r)
> hangs and can't be killed.
> 
> The same failure occurs on another similarly-configured system.

Let's try to tackle it from another angle. 

What is commont on thoe system where you see the failure? Are all not
using systemd as init? Additionally can you give some more details on
your setup, how the exports look? Can we boild down the setup
minimally to trigger the issue (and so report upstream)?

Can you additionally please test to downgrade to the 2.8.1-2 version
and please report back if you see the problem there was well? For us
it should behave actually same as 2.8.1-1 but I would like to double
check.

the supported way is to run it with systemd, the ship'ed sysvinit
scripts contain legacy, for instance with new version we would start
the server with nfsdctl:

ExecStart=/bin/sh -c '/usr/sbin/nfsdctl autostart || /usr/sbin/rpc.nfsd'

If you start it through nfsdctl do you get the exports working?

Still, having more information on the underlying setup would be
helpful.

Regards,
Salvatore

[toc] | [prev] | [next] | [standalone]


#84852

FromBill Brelsford <wb@k2di.net>
Date2024-12-17 03:40 +0100
Message-ID<JUvF7-7HE-1@gated-at.bofh.it>
In reply to#84848
Hi Salvatore,

On Mon Dec 16 2024 at 09:36 PM +0100, Salvatore Bonaccorso wrote:
> What is commont on thoe system where you see the failure? Are all not
> using systemd as init? Additionally can you give some more details on
> your setup, how the exports look? Can we boild down the setup
> minimally to trigger the issue (and so report upstream)?

The other system is an old one that I maintain as a backup; it also
uses sysvinit. The nfs setup for both is simple, with no changes to
/etc/nfs.conf or /etc/default/nfs-kernel-server. /etc/exports on
the newer system:

   / 10.20.40.80/29(rw,fsid=8208,no_root_squash,no_subtree_check,sync)
   /u 10.20.40.80/29(rw,fsid=8202,no_root_squash,no_subtree_check,sync)
   /u1 10.20.40.80/29(rw,fsid=8203,no_root_squash,no_subtree_check,sync)
   /n/ws 10.20.40.80/29(rw,fsid=8206,no_root_squash,no_subtree_check,sync)
   /n/wt 10.20.40.80/29(rw,fsid=8207,no_root_squash,no_subtree_check,sync)

It had been working fine on both systems prior to the 2.8.2-1
upgrade. It also works on 5 bookworm systems with the same setup.

> Can you additionally please test to downgrade to the 2.8.1-2 version
> and please report back if you see the problem there was well? For us
> it should behave actually same as 2.8.1-1 but I would like to double
> check.

I downgraded to 2.8.1-1 because it was available in the trixie
repository. Where can I get 2.8.1-2?

I also have trixie installed on both systems. Trixie now uses
2.8.2-1 so I upgraded one of them -- it also fails.

> the supported way is to run it with systemd, the ship'ed sysvinit
> scripts contain legacy, for instance with new version we would start
> the server with nfsdctl:
> 
> ExecStart=/bin/sh -c '/usr/sbin/nfsdctl autostart || /usr/sbin/rpc.nfsd'
> 
> If you start it through nfsdctl do you get the exports working?

I haven't installed systemd -- nothing I use has required it -- so
nfsdctl isn't available.

> Still, having more information on the underlying setup would be
> helpful.

Another possible clue: after upgrading (to 2.8.2-1), there is no
problem -- I can start and stop the daemon, export/unexport
filesystems, etc.  Everything seems normal -- until the system is
rebooted and rpc.nfsd is invoked.

Regards..  Bill

[toc] | [prev] | [next] | [standalone]


#84863

FromSalvatore Bonaccorso <carnil@debian.org>
Date2024-12-17 23:20 +0100
Message-ID<JUO53-jlb-1@gated-at.bofh.it>
In reply to#84852
Hi Bill,

On Mon, Dec 16, 2024 at 06:30:07PM -0800, Bill Brelsford wrote:
> Hi Salvatore,
> 
> On Mon Dec 16 2024 at 09:36 PM +0100, Salvatore Bonaccorso wrote:
> > What is commont on thoe system where you see the failure? Are all not
> > using systemd as init? Additionally can you give some more details on
> > your setup, how the exports look? Can we boild down the setup
> > minimally to trigger the issue (and so report upstream)?
> 
> The other system is an old one that I maintain as a backup; it also
> uses sysvinit. The nfs setup for both is simple, with no changes to
> /etc/nfs.conf or /etc/default/nfs-kernel-server. /etc/exports on
> the newer system:
> 
>    / 10.20.40.80/29(rw,fsid=8208,no_root_squash,no_subtree_check,sync)
>    /u 10.20.40.80/29(rw,fsid=8202,no_root_squash,no_subtree_check,sync)
>    /u1 10.20.40.80/29(rw,fsid=8203,no_root_squash,no_subtree_check,sync)
>    /n/ws 10.20.40.80/29(rw,fsid=8206,no_root_squash,no_subtree_check,sync)
>    /n/wt 10.20.40.80/29(rw,fsid=8207,no_root_squash,no_subtree_check,sync)
> 
> It had been working fine on both systems prior to the 2.8.2-1
> upgrade. It also works on 5 bookworm systems with the same setup.

Thanks. Unfortunately still no look in a lab setup to trigger your
issue. When rpc.segfaults, are there any other nfs related processed
and threads running on the system?

> > Can you additionally please test to downgrade to the 2.8.1-2 version
> > and please report back if you see the problem there was well? For us
> > it should behave actually same as 2.8.1-1 but I would like to double
> > check.
> 
> I downgraded to 2.8.1-1 because it was available in the trixie
> repository. Where can I get 2.8.1-2?

It is not anymore available in the archive as it is superseeded. But
you can find it on snapshot.d.o:

https://snapshot.debian.org/package/nfs-utils/1%3A2.8.1-2/


> 
> I also have trixie installed on both systems. Trixie now uses
> 2.8.2-1 so I upgraded one of them -- it also fails.
> 
> > the supported way is to run it with systemd, the ship'ed sysvinit
> > scripts contain legacy, for instance with new version we would start
> > the server with nfsdctl:
> > 
> > ExecStart=/bin/sh -c '/usr/sbin/nfsdctl autostart || /usr/sbin/rpc.nfsd'
> > 
> > If you start it through nfsdctl do you get the exports working?
> 
> I haven't installed systemd -- nothing I use has required it -- so
> nfsdctl isn't available.

nfsdctl is independent of systemd, it is shipped in as the new
tool(ing) for starting the nfsd server:

/usr/sbin/nfsdctl from nfs-kernel-server package.

> > Still, having more information on the underlying setup would be
> > helpful.
> 
> Another possible clue: after upgrading (to 2.8.2-1), there is no
> problem -- I can start and stop the daemon, export/unexport
> filesystems, etc.  Everything seems normal -- until the system is
> rebooted and rpc.nfsd is invoked.

As for the first part: What processes and kernel threads are started
on the system at this stage?

Regards,
Salvatore

[toc] | [prev] | [next] | [standalone]


#84873

FromBill Brelsford <wb@k2di.net>
Date2024-12-19 06:50 +0100
Message-ID<JVhA5-MN8-1@gated-at.bofh.it>
In reply to#84863
Hi Salvatore,

On Tue Dec 17 2024 at 11:12 PM +0100, Salvatore Bonaccorso wrote:
> Thanks. Unfortunately still no look in a lab setup to trigger your
> issue. When rpc.segfaults, are there any other nfs related processed
> and threads running on the system?

ps -eLf shows only one: [kworker/R-nfsiod]

> > > Can you additionally please test to downgrade to the 2.8.1-2 version
> > > and please report back if you see the problem there was well? For us
> > > it should behave actually same as 2.8.1-1 but I would like to double
> > > check.

Thanks for the pointer to snapshot.d.o.  Yes, 2.8.1-2 also has the
problem.

> > > If you start it through nfsdctl do you get the exports working?

Exports/exportfs are already working.  Trying to start with nfsdctl
instead of /etc/init.d/nfs-kernel-server (in 2.8.2-1) gives

	# nfsdctl -d autostart
	nfsdctl> autostart
	Error: Device or resource busy
	Error: Operation not permitted

> > Another possible clue: after upgrading (to 2.8.2-1), there is no
> > problem -- I can start and stop the daemon, export/unexport
> > filesystems, etc.  Everything seems normal -- until the system is
> > rebooted and rpc.nfsd is invoked.
> 
> As for the first part: What processes and kernel threads are started
> on the system at this stage?

ps -eLf gives 16 [nfsd] as well as the [kworker/R-nfsiod].  And
with the daemon started, some nfsdctl commands work:

	# nfsdctl status                 
	# nfsdctl threads  
	gracetime: 90
	leasetime: 90
	scope: k2ww
	pool-threads: 16   
	# nfsdctl listener 
	tcp:[::]:2049
	tcp:0.0.0.0:2049   
	# nfsdctl version  
	+3.0 +4.0 +4.1 +4.2

Hope this helps.  Thanks..  Bill

[toc] | [prev] | [next] | [standalone]


#84876

FromBill Brelsford <wb@k2di.net>
Date2024-12-19 16:30 +0100
Message-ID<JVqDn-UiF-11@gated-at.bofh.it>
In reply to#84873
On Wed Dec 18 2024 at 09:43 PM -0800, Bill Brelsford wrote:
> > > > Can you additionally please test to downgrade to the 2.8.1-2 version
> > > > and please report back if you see the problem there was well? For us
> > > > it should behave actually same as 2.8.1-1 but I would like to double
> > > > check.
> 
> Thanks for the pointer to snapshot.d.o.  Yes, 2.8.1-2 also has the
> problem.

No -- my mistake!  Downgrading to 2.8.1-2 works (as does 2.8.1-1).

Bill

[toc] | [prev] | [next] | [standalone]


#85141

FromBill Brelsford <wb@k2di.net>
Date2025-01-15 03:00 +0100
Message-ID<K50Rj-910R-5@gated-at.bofh.it>
In reply to#84876
Hi Salvatore,

The problem has apparently been fixed in unstable by updates to one
or more packages since December 19.  Version 1:2.8.2-1 now works as
expected on both of my machines.  

But it still fails in testing (trixie).  I'll try to determine what
future package update fixes it.

Regards..  Bill

[toc] | [prev] | [next] | [standalone]


#85143

FromSalvatore Bonaccorso <carnil@debian.org>
Date2025-01-15 07:20 +0100
Message-ID<K54UV-949u-1@gated-at.bofh.it>
In reply to#85141
Hi Bill,

On Tue, Jan 14, 2025 at 05:50:02PM -0800, Bill Brelsford wrote:
> Hi Salvatore,
> 
> The problem has apparently been fixed in unstable by updates to one
> or more packages since December 19.  Version 1:2.8.2-1 now works as
> expected on both of my machines.  
> 
> But it still fails in testing (trixie).  I'll try to determine what
> future package update fixes it.

I have a suspect what it can be. Can you please post the kernel log /
dmesg from the systems which do not work please?

Regards,
Salvatore

[toc] | [prev] | [next] | [standalone]


#85158 — Processed: Re: Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault

From"Debian Bug Tracking System" <owner@bugs.debian.org>
Date2025-01-16 09:40 +0100
SubjectProcessed: Re: Bug#1089976: nfs-kernel-server: Fails to start - segmentation fault
Message-ID<K5tzX-9jKC-1@gated-at.bofh.it>
In reply to#84819
Processing control commands:

> reassign -1 src:linux
Bug #1089976 [nfs-kernel-server] nfs-kernel-server: Fails to start - segmentation fault
Bug reassigned from package 'nfs-kernel-server' to 'src:linux'.
No longer marked as found in versions nfs-utils/1:2.8.2-1.
Ignoring request to alter fixed versions of bug #1089976 to the same values previously set
> forcemerge 1087900 -1
Bug #1087900 {Done: Bastian Blank <waldi@debian.org>} [src:linux] linux: kernel BUG at fs/nfsd/nfs4recover.c:534
Bug #1091439 {Done: Bastian Blank <waldi@debian.org>} [src:linux] installation of nfs-kernel-server hangs
Bug #1092607 {Done: Bastian Blank <waldi@debian.org>} [src:linux] dracut: upstream-dracut-network-nfs autopkgtest fails on amd64
Bug #1087900 {Done: Bastian Blank <waldi@debian.org>} [src:linux] linux: kernel BUG at fs/nfsd/nfs4recover.c:534
Added tag(s) unreproducible and moreinfo.
Added tag(s) moreinfo and unreproducible.
Added tag(s) unreproducible and moreinfo.
Bug #1089976 [src:linux] nfs-kernel-server: Fails to start - segmentation fault
Set Bug forwarded-to-address to 'https://lore.kernel.org/linux-nfs/Z22DIiV98XBSfPVr@eldamar.lan/'.
Marked Bug as done
Added indication that 1089976 affects src:dracut
Marked as fixed in versions linux/6.12.8-1 and linux/6.13~rc6-1~exp1.
Marked as found in versions linux/6.8.9-1, linux/6.12~rc6-1~exp1, linux/6.11.9-1, linux/6.11.5-1, and linux/6.12.6-1.
Added tag(s) upstream and confirmed.
Bug #1091439 {Done: Bastian Blank <waldi@debian.org>} [src:linux] installation of nfs-kernel-server hangs
Bug #1092607 {Done: Bastian Blank <waldi@debian.org>} [src:linux] dracut: upstream-dracut-network-nfs autopkgtest fails on amd64
Merged 1087900 1089976 1091439 1092607

-- 
1087900: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1087900
1089976: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1089976
1091439: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1091439
1092607: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=1092607
Debian Bug Tracking System
Contact owner@bugs.debian.org with problems

[toc] | [prev] | [next] | [standalone]


#85159

FromSalvatore Bonaccorso <carnil@debian.org>
Date2025-01-16 09:40 +0100
Message-ID<K5tzX-9jKC-3@gated-at.bofh.it>
In reply to#84819
Control: reassign -1 src:linux
Control: forcemerge 1087900 -1

Hi Bill,

On Wed, Jan 15, 2025 at 08:02:18AM -0800, Bill Brelsford wrote:
> On Wed Jan 15 2025 at 07:08 AM +0100, Salvatore Bonaccorso wrote:
> > > But it still fails in testing (trixie).  I'll try to determine what
> > > future package update fixes it.
> > 
> > I have a suspect what it can be. Can you please post the kernel log /
> > dmesg from the systems which do not work please?
> 
> Attached is dmesg from trixie.
> 
> Bill

[...]
> [   33.734128] r8169 0000:05:07.0 eth0: Link is Up - 100Mbps/Full - flow control rx/tx
> [   40.873091] NFSD: Using /var/lib/nfs/v4recovery as the NFSv4 state recovery directory
> [   40.967422] NFSD: Using legacy client tracking operations.
> [   40.971774] NFSD: Using /var/lib/nfs/v4recovery as the NFSv4 state recovery directory
> [   40.976080] ------------[ cut here ]------------
> [   40.980166] kernel BUG at fs/nfsd/nfs4recover.c:534!
> [   40.984266] Oops: invalid opcode: 0000 [#1] PREEMPT SMP PTI
> [   40.988166] CPU: 0 UID: 0 PID: 1935 Comm: rpc.nfsd Not tainted 6.12.6-amd64 #1  Debian 6.12.6-1
> [   40.988236] Hardware name: To Be Filled By O.E.M. S62E/S62E, BIOS 0303    08/03/2007
> [   40.988236] RIP: 0010:nfsd4_legacy_tracking_init+0x17d/0x1b0 [nfsd]
> [   40.988236] Code: 19 48 89 de 48 c7 c7 50 99 cb c1 e8 6d fb ff ff 89 c5 85 c0 0f 85 10 61 00 00 48 c7 c7 90 f3 d2 c1 31 ed e8 85 26 3e f8 eb 07 <0f> 0b bd f4 ff ff ff 48 8b 44 24 08 65 48 2b 04 25 28 00 00 00 75
> [   40.988236] RSP: 0018:ffffbc09c0f87c80 EFLAGS: 00010282
> [   40.988236] RAX: 0000000000000049 RBX: ffffa0364b4d8000 RCX: 0000000000000003
> [   40.988236] RDX: 0000000000000000 RSI: 0000000000000003 RDI: 0000000000000001
> [   40.988236] RBP: ffffffffbbca3600 R08: 0000000000000000 R09: ffffbc09c0f87b10
> [   40.988236] R10: ffffffffbb0b4348 R11: 0000000000000003 R12: ffffa0364b4d8000
> [   40.988236] R13: ffffa0364b4d8000 R14: ffffa0364c306b40 R15: ffffa0364b4d8000
> [   40.988236] FS:  00007faff2288740(0000) GS:ffffa036bd400000(0000) knlGS:0000000000000000
> [   40.988236] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> [   40.988236] CR2: 00007f0d8a7e1320 CR3: 0000000002cfc000 CR4: 00000000000026f0
> [   40.988236] Call Trace:
> [   40.988236]  <TASK>
> [   40.988236]  ? __die_body.cold+0x19/0x27
> [   40.988236]  ? die+0x2e/0x50
> [   40.988236]  ? do_trap+0xca/0x110
> [   40.988236]  ? do_error_trap+0x6a/0x90
> [   40.988236]  ? nfsd4_legacy_tracking_init+0x17d/0x1b0 [nfsd]
> [   40.988236]  ? exc_invalid_op+0x50/0x70
> [   40.988236]  ? nfsd4_legacy_tracking_init+0x17d/0x1b0 [nfsd]
> [   40.988236]  ? asm_exc_invalid_op+0x1a/0x20
> [   40.988236]  ? nfsd4_legacy_tracking_init+0x17d/0x1b0 [nfsd]
> [   40.988236]  nfsd4_client_tracking_init+0x57/0x1b0 [nfsd]
> [   40.988236]  nfs4_state_start_net+0x2f9/0x3a0 [nfsd]
> [   40.988236]  nfsd_svc+0x1ac/0x310 [nfsd]
> [   40.988236]  write_threads+0xf9/0x1c0 [nfsd]
> [   40.988236]  ? __pfx_write_threads+0x10/0x10 [nfsd]
> [   40.988236]  nfsctl_transaction_write+0x4a/0x80 [nfsd]
> [   40.988236]  vfs_write+0xf8/0x450
> [   40.988236]  ksys_write+0x6d/0xf0
> [   40.988236]  do_syscall_64+0x82/0x190
> [   40.988236]  ? do_user_addr_fault+0x36c/0x620
> [   40.988236]  ? exc_page_fault+0x7e/0x180
> [   40.988236]  entry_SYSCALL_64_after_hwframe+0x76/0x7e
> [   40.988236] RIP: 0033:0x7faff238f090
> [   40.988236] Code: 2d 0e 00 64 c7 00 16 00 00 00 b8 ff ff ff ff c3 66 2e 0f 1f 84 00 00 00 00 00 80 3d d9 af 0e 00 00 74 17 b8 01 00 00 00 0f 05 <48> 3d 00 f0 ff ff 77 58 c3 0f 1f 80 00 00 00 00 48 83 ec 28 48 89
> [   40.988236] RSP: 002b:00007ffeff5e5298 EFLAGS: 00000202 ORIG_RAX: 0000000000000001
> [   40.988236] RAX: ffffffffffffffda RBX: 0000000000000003 RCX: 00007faff238f090
> [   40.988236] RDX: 0000000000000003 RSI: 000055744ca80340 RDI: 0000000000000003
> [   40.988236] RBP: 000055744ca80340 R08: 0000000000000064 R09: 00000000fffffffe
> [   40.988236] R10: 0000000000000000 R11: 0000000000000202 R12: 0000000000020000
> [   40.988236] R13: 000055744ca7c116 R14: 000055748c50c2a0 R15: 0000000000000000
> [   40.988236]  </TASK>
> [   40.988236] Modules linked in: xt_nat ipt_REJECT nf_reject_ipv4 xt_conntrack xt_tcpudp iptable_mangle iptable_nat nf_nat nf_conntrack nf_defrag_ipv6 nf_defrag_ipv4 libcrc32c iptable_filter ip_tables x_tables nfsd auth_rpcgss nfs_acl nfs lockd grace netfs sunrpc firewire_sbp2 sr_mod at24 cdrom iTCO_wdt intel_pmc_bxt iTCO_vendor_support watchdog coretemp i915 kvm_intel snd_hda_codec_si3054 snd_hda_codec_realtek kvm snd_hda_codec_generic snd_hda_scodec_component snd_hda_intel uvcvideo drm_buddy iwl4965 snd_intel_dspcfg snd_intel_sdw_acpi sha512_ssse3 drm_display_helper snd_hda_codec sha256_ssse3 cec videobuf2_vmalloc iwlegacy uvc snd_hda_core videobuf2_memops rc_core sha1_ssse3 snd_hwdep mac80211 videobuf2_v4l2 snd_pcm_oss ttm r852 snd_mixer_oss videodev snd_pcm sm_common drm_kms_helper firewire_ohci nand r8169 snd_timer pcspkr drm firewire_core libarc4 i2c_i801 nandcore joydev sdhci_pci videobuf2_common snd asus_laptop acpi_cpufreq cfg80211 cqhci realtek sparse_keymap i2c_smbus sdhci mdio_devres bch r592 ata_generic
> [   40.988236]  mmc_core mc libphy i2c_algo_bit soundcore mtd crc_itu_t memstick sg ata_piix rfkill video ehci_pci ac lpc_ich wmi button battery ext4 crc16 mbcache jbd2 crc32c_generic xts dm_crypt dm_mod hid_generic usbhid hid sd_mod uhci_hcd ehci_hcd ahci libahci usbcore psmouse libata evdev scsi_mod serio_raw usb_common scsi_common
> [   41.220014] ---[ end trace 0000000000000000 ]---
> [   41.221857] RIP: 0010:nfsd4_legacy_tracking_init+0x17d/0x1b0 [nfsd]
> [   41.223642] Code: 19 48 89 de 48 c7 c7 50 99 cb c1 e8 6d fb ff ff 89 c5 85 c0 0f 85 10 61 00 00 48 c7 c7 90 f3 d2 c1 31 ed e8 85 26 3e f8 eb 07 <0f> 0b bd f4 ff ff ff 48 8b 44 24 08 65 48 2b 04 25 28 00 00 00 75
> [   41.227245] RSP: 0018:ffffbc09c0f87c80 EFLAGS: 00010282
> [   41.229058] RAX: 0000000000000049 RBX: ffffa0364b4d8000 RCX: 0000000000000003
> [   41.230870] RDX: 0000000000000000 RSI: 0000000000000003 RDI: 0000000000000001
> [   41.232683] RBP: ffffffffbbca3600 R08: 0000000000000000 R09: ffffbc09c0f87b10
> [   41.234502] R10: ffffffffbb0b4348 R11: 0000000000000003 R12: ffffa0364b4d8000
> [   41.236326] R13: ffffa0364b4d8000 R14: ffffa0364c306b40 R15: ffffa0364b4d8000
> [   41.238171] FS:  00007faff2288740(0000) GS:ffffa036bd400000(0000) knlGS:0000000000000000
> [   41.240047] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> [   41.241942] CR2: 00007f0d8a7e1320 CR3: 0000000002cfc000 CR4: 00000000000026f0
> [   43.937663] warning: `iwconfig' uses wireless extensions which will stop working for Wi-Fi 7 hardware; use nl80211
> [   48.021236] Key type dns_resolver registered
> [   48.660788] NFS: Registering the id_resolver key type
> [   48.662980] Key type id_resolver registered
> [   48.665164] Key type id_legacy registered

So this is what I suspected. This is #1087900 in src:linux and fixed
in 6.12.8-1.

I'm reassigning and merging the two bugs.

*But* that said, I strongly encourage you to switch to a systemd
running system for your nfs server. We still ship the init scripts in
the packaging, but for instance you do not start with them the more
advanced client tracking daemons. The legacy tracking methods will
disapear at some point upstream completely and in fact for the next
experimental upload I aim to disable NFSD_LEGACY_CLIENT_TRACKING
(cf.
https://salsa.debian.org/kernel-team/linux/-/merge_requests/1298).

Thanks for the report!

Regards,
Salvatore

[toc] | [prev] | [next] | [standalone]


#85169

FromBill Brelsford <wb@k2di.net>
Date2025-01-17 04:30 +0100
Message-ID<K5Ldv-9vYk-1@gated-at.bofh.it>
In reply to#85159
Hi Salvatore,

On Thu Jan 16 2025 at 09:30 AM +0100, Salvatore Bonaccorso wrote:
> On Wed, Jan 15, 2025 at 08:02:18AM -0800, Bill Brelsford wrote:
> > [   40.980166] kernel BUG at fs/nfsd/nfs4recover.c:534!
> > [   40.984266] Oops: invalid opcode: 0000 [#1] PREEMPT SMP PTI

> So this is what I suspected. This is #1087900 in src:linux and fixed
> in 6.12.8-1.

I'm embarassed! I thought I had checked dmesg, but obviously hadn't.
It's working on trixie now with the latest kernel update (6.12.9-1).

> *But* that said, I strongly encourage you to switch to a systemd
> running system for your nfs server. We still ship the init scripts in
> the packaging, but for instance you do not start with them the more
> advanced client tracking daemons. The legacy tracking methods will
> disapear at some point upstream completely and in fact for the next
> experimental upload I aim to disable NFSD_LEGACY_CLIENT_TRACKING
> (cf.
> https://salsa.debian.org/kernel-team/linux/-/merge_requests/1298).

Yes, it's time for me to switch to systemd.  

Thanks very much for your help!

Bill

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web