Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1415457

Re: performance delta after VFS i_mutex=>i_rwsem conversion

From Al Viro <viro@ZenIV.linux.org.uk>
Newsgroups linux.kernel
Subject Re: performance delta after VFS i_mutex=>i_rwsem conversion
Date 2016-06-06 23:20 +0200
Message-ID <rHa6d-1IG-5@gated-at.bofh.it> (permalink)
References <rH90u-16B-13@gated-at.bofh.it> <rH9Db-1jo-33@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


On Mon, Jun 06, 2016 at 01:46:23PM -0700, Linus Torvalds wrote:

> So my gut feel is that we do want to have the same heuristics for
> rwsems and mutexes (well, modulo possible actual semantic differences
> due to the whole shared-vs-exclusive issues).
> 
> And I also suspect that the mutexes have gotten a lot more performance
> tuning done on them, so it's likely the correct thing to try to make
> the rwsem match the mutex code rather than the other way around.
> 
> I think we had Jason and Davidlohr do mutex work last year, let's see
> if they agree on that "yes, the mutex case is the likely more tuned
> case" feeling.
> 
> The fact that your performance improves when you do that obviously
> then also validates the assumption that the mutex spinning is the
> better optimized one.

FWIW, there's another fun issue on ramfs - dcache_readdir() is doing an
obscene amount of grabbing/releasing ->d_lock and once you take the external
serialization out, parallel getdents load hits contention on *that*.
In spades.  And unlike mutex (or rswem exclusive), contention on ->d_lock
chews a lot of cycles.  The root cause is the use of cursors - we not only
move them more than we ought to (we do that on each entry reported, rather
than once before return from dcache_readdir()), we can't traverse the real
list entries (which remain nice and stable; another low-hanging fruit is
pointless grabbing ->d_lock on those) without ->d_lock on parent.

I think I have a kinda-sorta solution, but it has a problem.  What I want
to do is
	* list_move() only once per dcache_readdir()
	* ->d_lock taken for that and only for that.
	* list_move() itself surrounded with write_seqcount_{begin,end} on
some seqcount
	* traversal to the next real entry done under rcu_read_lock in a
seqretry loop.

The only problem is where to put that seqcount (unsigned int, really).
->i_dir_seq is an obvious candidate, but that'll need careful profiling
on getdents/lookup mixes...

Back to linux.kernel | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

performance delta after VFS i_mutex=>i_rwsem conversion Dave Hansen <dave.hansen@intel.com> - 2016-06-06 22:10 +0200
  Re: performance delta after VFS i_mutex=>i_rwsem conversion Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-06 22:50 +0200
    Re: performance delta after VFS i_mutex=>i_rwsem conversion Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-06 23:20 +0200
      Re: performance delta after VFS i_mutex=>i_rwsem conversion Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-06 23:50 +0200
        Re: performance delta after VFS i_mutex=>i_rwsem conversion Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-07 00:10 +0200
          Re: performance delta after VFS i_mutex=>i_rwsem conversion Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-07 02:00 +0200
            Re: performance delta after VFS i_mutex=>i_rwsem conversion Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-07 02:00 +0200
              Re: performance delta after VFS i_mutex=>i_rwsem conversion Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-07 02:30 +0200
            Re: performance delta after VFS i_mutex=>i_rwsem conversion Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-07 02:50 +0200
              Re: performance delta after VFS i_mutex=>i_rwsem conversion Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-07 02:50 +0200
              Re: performance delta after VFS i_mutex=>i_rwsem conversion Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-07 03:00 +0200
                Re: performance delta after VFS i_mutex=>i_rwsem conversion Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-07 03:30 +0200
              Re: performance delta after VFS i_mutex=>i_rwsem conversion Al Viro <viro@ZenIV.linux.org.uk> - 2016-06-07 03:00 +0200
    Re: performance delta after VFS i_mutex=>i_rwsem conversion Linus Torvalds <torvalds@linux-foundation.org> - 2016-06-06 23:30 +0200
      Re: performance delta after VFS i_mutex=>i_rwsem conversion Valdis.Kletnieks@vt.edu - 2016-06-07 05:30 +0200
    Re: performance delta after VFS i_mutex=>i_rwsem conversion Ingo Molnar <mingo@kernel.org> - 2016-06-08 11:00 +0200
      Re: performance delta after VFS i_mutex=>i_rwsem conversion Ingo Molnar <mingo@kernel.org> - 2016-06-09 12:30 +0200
        Re: performance delta after VFS i_mutex=>i_rwsem conversion Dave Hansen <dave.hansen@intel.com> - 2016-06-09 20:20 +0200
          RE: performance delta after VFS i_mutex=>i_rwsem conversion "Chen, Tim C" <tim.c.chen@intel.com> - 2016-06-09 22:20 +0200

csiph-web