Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1413745

Re: Dcache oops

From Jeff Layton <jlayton@poochiereds.net>
Newsgroups linux.kernel
Subject Re: Dcache oops
Date 2016-06-04 14:30 +0200
Message-ID <rGiSe-A9-5@gated-at.bofh.it> (permalink)
References (6 earlier) <rG5rX-uq-3@gated-at.bofh.it> <rG5Lk-AU-43@gated-at.bofh.it> <rG5UZ-Ei-11@gated-at.bofh.it> <rG7ap-1iQ-7@gated-at.bofh.it> <rG86t-1Sw-7@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


On Sat, 2016-06-04 at 01:56 +0100, Al Viro wrote:
> On Fri, Jun 03, 2016 at 07:58:37PM -0400, Oleg Drokin wrote:
> 
> > 
> > > 
> > > EOPENSTALE, that is...  Oleg, could you check if the following works?
> > Yes, this one lasted for an hour with no crashing, so it must be good.
> > Thanks.
> > (note, I am not equipped to verify correctness of NFS operations, though).
> I suspect that Jeff Layton might have relevant regression tests.  Incidentally,
> we really need a consolidated regression testsuite, including the tests you'd
> been running.  Right now there's some stuff in xfstests, LTP and cthon; if
> anything, this mess shows just why we need all of that and then some in
> a single place.  Lustre stuff has caught a 3 years old NFS bug (missing
> d_drop() in nfs_atomic_open()) and a year-old bug in handling of EOPENSTALE
> retries on the last component of a trailing non-embedded symlink.  Neither
> is hard to trigger; it's just that relevant tests hadn't been run on NFS,
> period.
> 
> Jeff, could you verify that the following does not cause regressions in
> stale fhandles treatment?  I want to rip the damn retry logics out of
> do_last() and if the staleness had only been discovered inside of
> nfs4_file_open() just have the upper-level logics handle it by doing
> a normal LOOKUP_REVAL pass from scratch.  To hell with trying to be clever;
> a few roundtrips it saves us in some cases is not worth the complexity and
> potential for bugs.  I'm fairly sure that the time spent debugging this
> particular turd exceeds the total amount of time it has ever saved,
> and do_last() is in dire need of simplification.  All talk about "enough eyes"
> isn't worth much when the readers of code in question feel like ripping their
> eyes out...
> 

Agreed. I see no need to optimize an error case here. Any performance
hit that we'd get here is almost certainly acceptable in this
situation. The main thing is that we prevent the ESTALE from bubbling
up into userland if we can avoid it by retrying.

No, I didn't have the test for this anymore unfortunately. RHQA might
have one though.

Either way, I cooked one up that does this on the server:

#!/bin/bash

while true; do
	rm -rf foo
	mkdir foo
	echo foo > foo/bar
	usleep 100000
done

...and then this on the client after mounting the fs with
lookupcache=none and noac.

#include <sys/types.h>
#include <sys/stat.h>
#include <fcntl.h>
#include <errno.h>
#include <stdio.h>
#include <unistd.h>

int main(int argc, char **argv)
{
	int fd;

	while(1) {
		fd = open(argv[1], O_RDONLY);
		if (fd < 0) {
			if (errno == ESTALE) {
				printf("ESTALE");
				return 1;
			}
			continue;
		}
		close(fd);
	}
	return 0;
}

I did see some of the OPEN compounds come back with NFS4ERR_STALE on
the PUTFH op but no corresponding ESTALE error in userland. So, this
patch does seem to do the right thing.

Reviewed-and-Tested-by: Jeff Layton <jlayton@poochiereds.net>

Back to linux.kernel | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

NFS/d_splice_alias breakage Oleg Drokin <green@linuxhacker.ru> - 2016-06-03 01:10 +0200
  [PATCH] Allow d_splice_alias to accept hashed dentries green@linuxhacker.ru - 2016-06-03 02:10 +0200
    Re: [PATCH] Allow d_splice_alias to accept hashed dentries Oleg Drokin <green@linuxhacker.ru> - 2016-06-03 02:30 +0200
  Re: NFS/d_splice_alias breakage Trond Myklebust <trondmy@primarydata.com> - 2016-06-03 02:50 +0200
    Re: NFS/d_splice_alias breakage Oleg Drokin <green@linuxhacker.ru> - 2016-06-03 03:00 +0200
      Re: NFS/d_splice_alias breakage Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 05:30 +0200
        Re: NFS/d_splice_alias breakage Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 05:40 +0200
    Re: NFS/d_splice_alias breakage Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 05:30 +0200
  Re: NFS/d_splice_alias breakage Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 05:40 +0200
    Re: NFS/d_splice_alias breakage Oleg Drokin <green@linuxhacker.ru> - 2016-06-03 05:50 +0200
      Re: NFS/d_splice_alias breakage Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 06:30 +0200
        Re: NFS/d_splice_alias breakage Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 06:50 +0200
          Re: NFS/d_splice_alias breakage Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 07:00 +0200
        Re: NFS/d_splice_alias breakage Oleg Drokin <green@linuxhacker.ru> - 2016-06-03 07:00 +0200
          Re: NFS/d_splice_alias breakage Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 08:00 +0200
            Re: NFS/d_splice_alias breakage Oleg Drokin <green@linuxhacker.ru> - 2016-06-07 01:40 +0200
              Re: NFS/d_splice_alias breakage Oleg Drokin <green@linuxhacker.ru> - 2016-06-10 03:40 +0200
                Re: NFS/d_splice_alias breakage Oleg Drokin <green@linuxhacker.ru> - 2016-06-10 18:50 +0200
        Dcache oops Oleg Drokin <green@linuxhacker.ru> - 2016-06-03 18:40 +0200
          Re: Dcache oops Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 20:30 +0200
            Re: Dcache oops Oleg Drokin <green@linuxhacker.ru> - 2016-06-03 20:40 +0200
              Re: Dcache oops Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 22:10 +0200
                Re: Dcache oops Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-03 23:20 +0200
                Re: Dcache oops Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 23:30 +0200
                Re: Dcache oops Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-04 00:10 +0200
                Re: Dcache oops Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-04 00:30 +0200
                Re: Dcache oops Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-04 00:30 +0200
                Re: Dcache oops Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-04 00:40 +0200
                Re: Dcache oops Oleg Drokin <green@linuxhacker.ru> - 2016-06-04 00:50 +0200
                Re: Dcache oops Oleg Drokin <green@linuxhacker.ru> - 2016-06-04 02:00 +0200
                Re: Dcache oops Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-04 03:00 +0200
                Re: Dcache oops Jeff Layton <jlayton@poochiereds.net> - 2016-06-04 14:30 +0200
                Re: Dcache oops Oleg Drokin <green@linuxhacker.ru> - 2016-06-04 18:20 +0200
                [PATCH] nfs4: Fix potential use after free of state in nfs4_do_reclaim. green@linuxhacker.ru - 2016-06-04 18:30 +0200
                Re: [PATCH] nfs4: Fix potential use after free of state in  nfs4_do_reclaim. Jeff Layton <jlayton@poochiereds.net> - 2016-06-04 22:00 +0200
                Re: Dcache oops Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-04 00:40 +0200
                Re: Dcache oops Oleg Drokin <green@linuxhacker.ru> - 2016-06-04 00:50 +0200
                Re: Dcache oops Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-04 00:50 +0200
                Re: Dcache oops Oleg Drokin <green@linuxhacker.ru> - 2016-06-03 23:20 +0200
                Re: Dcache oops Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-03 23:50 +0200
                Re: Dcache oops Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-04 00:20 +0200

csiph-web