Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1460084 > unrolled thread

Re: [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

Started byDave Chinner <david@fromorbit.com>
First post2016-08-11 01:10 +0200
Last post2016-08-11 02:00 +0200
Articles 20 on this page of 98 — 11 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 01:10 +0200
    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 02:00 +0200
      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 02:20 +0200
        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 02:30 +0200
          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 02:40 +0200
            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 03:10 +0200
              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 06:50 +0200
                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-15 19:30 +0200
                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:30 +0200
              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-11 18:00 +0200
                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 19:00 +0200
                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 20:00 +0200
                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 22:00 +0200
                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-11 22:10 +0200
                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 22:40 +0200
                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Al Viro <viro@ZenIV.linux.org.uk> - 2016-08-12 00:20 +0200
                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 00:40 +0200
                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 23:50 +0200
                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-12 00:10 +0200
                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 03:00 +0200
                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 04:30 +0200
                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 06:00 +0200
                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 20:10 +0200
                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-14 11:00 +0200
                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 02:50 +0200
                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-15 03:40 +0200
                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 04:40 +0200
                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-15 05:00 +0200
                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 07:10 +0200
                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 00:30 +0200
                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 00:50 +0200
                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 01:30 +0200
                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 01:50 +0200
                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:50 +0200
                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-16 17:10 +0200
                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 20:00 +0200
                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Michal Hocko <mhocko@kernel.org> - 2016-08-17 17:50 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Michal Hocko <mhocko@kernel.org> - 2016-08-17 18:50 +0200
                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-17 17:50 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-18 02:50 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-18 09:20 +0200
                                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-18 15:30 +0200
                                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-19 04:10 +0200
                                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-19 04:40 +0200
                                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Michal Hocko <mhocko@kernel.org> - 2016-08-19 11:10 +0200
                                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-19 13:00 +0200
                                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-20 01:50 +0200
                                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-20 03:10 +0200
                                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-20 14:20 +0200
                                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-19 06:10 +0200
                                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-19 17:10 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-24 17:50 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-18 04:50 +0200
                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 02:20 +0200
                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:40 +0200
                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 03:00 +0200
                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:40 +0200
                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 04:00 +0200
                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-17 00:10 +0200
                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-17 01:30 +0200
                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 01:10 +0200
                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:40 +0200
                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 02:50 +0200
                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ingo Molnar <mingo@kernel.org> - 2016-08-15 07:10 +0200
                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Peter Zijlstra <peterz@infradead.org> - 2016-08-17 18:30 +0200
                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-15 15:10 +0200
                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 04:30 +0200
                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 04:40 +0200
                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-12 05:00 +0200
                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 05:30 +0200
                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 06:20 +0200
                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 07:10 +0200
                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 08:10 +0200
                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ye Xiaolong <xiaolong.ye@intel.com> - 2016-08-12 08:40 +0200
                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ye Xiaolong <xiaolong.ye@intel.com> - 2016-08-12 11:00 +0200
                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 12:10 +0200
                                      Re: [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-12 12:50 +0200
                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-13 02:40 +0200
                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-14 10:30 +0200
                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-14 10:40 +0200
                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-14 11:00 +0200
                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-14 11:30 +0200
                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-14 18:20 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 01:50 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-15 02:00 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-15 16:20 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 23:30 +0200
                                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-16 14:30 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-15 22:40 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%     regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-23 00:10 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-16 15:30 +0200
                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-14 12:00 +0200
            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 03:20 +0200
              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 03:40 +0200
                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ye Xiaolong <xiaolong.ye@intel.com> - 2016-08-11 04:50 +0200
                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 05:20 +0200
              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 03:30 +0200
    Re: [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 02:00 +0200

Page 4 of 5 — ← Prev page 1 2 3 [4] 5  Next page →


#1463252 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromLinus Torvalds <torvalds@linux-foundation.org>
Date2016-08-16 01:10 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6zb3-3AC-1@gated-at.bofh.it>
In reply to#1463223
On Mon, Aug 15, 2016 at 3:22 PM, Dave Chinner <david@fromorbit.com> wrote:
>
> Right, but that does not make the profile data useless,

Yes it does. Because it basically hides everything that happens inside
the lock, which is what causes the contention in the first place.

So stop making inane and stupid arguments, Dave.

Your profiles are shit. Deal with it, or accept that nobody is ever
going to bother working on them because your profiles don't give
useful information.

I see that you actually fixed your profiles, but quite frankly, the
amount of pure unadulterated crap you posted in this email is worth
reacting negatively to.

You generally make so much sense that it's shocking to see you then
make these crazy excuses for your completely broken profiles.

                 Linus

[toc] | [prev] | [next] | [standalone]


#1463309 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-16 02:40 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6AAa-4mv-11@gated-at.bofh.it>
In reply to#1463252
On Mon, Aug 15, 2016 at 04:01:00PM -0700, Linus Torvalds wrote:
> On Mon, Aug 15, 2016 at 3:22 PM, Dave Chinner <david@fromorbit.com> wrote:
> >
> > Right, but that does not make the profile data useless,
> 
> Yes it does. Because it basically hides everything that happens inside
> the lock, which is what causes the contention in the first place.

Read the code, Linus?

> So stop making inane and stupid arguments, Dave.

We know what happens inside the lock, and we know exactly how much
it is supposed to cost. And it isn't anywhere near as much as the
profiles indicate the function that contains the lock is costing.

Occam's Razor leads to only one conclusion, like it or not....

> Your profiles are shit. Deal with it, or accept that nobody is ever
> going to bother working on them because your profiles don't give
> useful information.
> 
> I see that you actually fixed your profiles, but quite frankly, the
> amount of pure unadulterated crap you posted in this email is worth
> reacting negatively to.

I'm happy to be told that I'm wrong *when I'm wrong*, but you always
say "read the code to understand a problem" rather than depending on
potentially unreliable tools and debug information that is gathered.

Yet when I do that using partial profile information, your reaction
is to tell me I am "full of shit" because my information isn't 100%
reliable? Really, Linus?

> You generally make so much sense that it's shocking to see you then
> make these crazy excuses for your completely broken profiles.

Except they *aren't broken*. They are simply *less accurate* than
they could be. That does not invalidate the profile nor does it mean
that the insight it gives us into the functioning of the code is
wrong.

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1463311 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromLinus Torvalds <torvalds@linux-foundation.org>
Date2016-08-16 02:50 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6AJP-4pI-9@gated-at.bofh.it>
In reply to#1463309
On Mon, Aug 15, 2016 at 5:17 PM, Dave Chinner <david@fromorbit.com> wrote:
>
> Read the code, Linus?

I am. It's how I came up with my current pet theory.

But I don't actually have enough sane numbers to make it much more
than a cute pet theory. It *might* explain why you see tons of kswap
time and bad lock contention where it didn't use to exist, but ..

I can't recreate the problem, and your old profiles were bad enough
that they aren't really worth looking at.

> Except they *aren't broken*. They are simply *less accurate* than
> they could be.

They are so much less accurate that quite frankly, there's no point in
looking at them outside of "there is contention on the lock".

And considering that the numbers didn't even change when you had
spinlock debugging on, it's not the lock itself that causes this, I'm
pretty sure.

Because when you have normal contention due to the *locking* itself
being the problem, it tends to absolutely _explode_ with the debugging
spinlocks, because the lock itself becomes much more expensive.
Usually super-linearly.

But that wasn't the case here. The numbers stayed constant.

So yeah, I started looking at bigger behavioral issues, which is why I
zeroed in on that zone-vs-node change. But it might be a completely
broken theory. For example, if you still have the contention when
running plain 4.7, that theory was clearly complete BS.

And this is where "less accurate" means that they are almost entirely useless.

More detail needed. It might not be in the profiles themselves, of
course. There might be other much more informative sources if you can
come up with anything...

               Linus

[toc] | [prev] | [next] | [standalone]


#1462551 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromIngo Molnar <mingo@kernel.org>
Date2016-08-15 07:10 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6ijU-1fL-9@gated-at.bofh.it>
In reply to#1462530
* Linus Torvalds <torvalds@linux-foundation.org> wrote:

> Make sure you actually use "perf record -e cycles:pp" or something
> that uses PEBS to get real profiles using CPU performance counters.

Btw., 'perf record -e cycles:pp' is the default now for modern versions
of perf tooling (on most x86 systems) - if you do 'perf record' it will
just use the most precise profiling mode available on that particular
CPU model.

If unsure you can check the event that was used, via:

  triton:~> perf report --stdio 2>&1 | grep '# Samples'
  # Samples: 27K of event 'cycles:pp'

Thanks,

	Ingo

[toc] | [prev] | [next] | [standalone]


#1464648 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromPeter Zijlstra <peterz@infradead.org>
Date2016-08-17 18:30 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s7bT3-3to-3@gated-at.bofh.it>
In reply to#1462551
On Mon, Aug 15, 2016 at 07:03:00AM +0200, Ingo Molnar wrote:
> 
> * Linus Torvalds <torvalds@linux-foundation.org> wrote:
> 
> > Make sure you actually use "perf record -e cycles:pp" or something
> > that uses PEBS to get real profiles using CPU performance counters.
> 
> Btw., 'perf record -e cycles:pp' is the default now for modern versions
> of perf tooling (on most x86 systems) - if you do 'perf record' it will
> just use the most precise profiling mode available on that particular
> CPU model.
> 
> If unsure you can check the event that was used, via:
> 
>   triton:~> perf report --stdio 2>&1 | grep '# Samples'
>   # Samples: 27K of event 'cycles:pp'

Problem here is that Dave is using a KVM thingy. Getting hardware
counters in a guest is somewhat tricky but doable, but PEBS does not
virtualize.

[toc] | [prev] | [next] | [standalone]


#1462766 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromFengguang Wu <fengguang.wu@intel.com>
Date2016-08-15 15:10 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6pOq-60A-17@gated-at.bofh.it>
In reply to#1461361
On Fri, Aug 12, 2016 at 11:03:33AM -0700, Linus Torvalds wrote:
>On Thu, Aug 11, 2016 at 8:56 PM, Dave Chinner <david@fromorbit.com> wrote:
>> On Thu, Aug 11, 2016 at 07:27:52PM -0700, Linus Torvalds wrote:
>>>
>>> I don't recall having ever seen the mapping tree_lock as a contention
>>> point before, but it's not like I've tried that load either. So it
>>> might be a regression (going back long, I suspect), or just an unusual
>>> load that nobody has traditionally tested much.
>>>
>>> Single-threaded big file write one page at a time, was it?
>>
>> Yup. On a 4 node NUMA system.
>
>Ok, I can't see any real contention on my single-node workstation
>(running ext4 too, so there may be filesystem differences), but I
>guess that shouldn't surprise me. The cacheline bouncing just isn't
>expensive enough when it all stays on-die.
>
>I can see the tree_lock in my profiles (just not very high), and at
>least for ext4 the main caller ssems to be
>__set_page_dirty_nobuffers().
>
>And yes, looking at that, the biggest cost by _far_ inside the
>spinlock seems to be the accounting.
>
>Which doesn't even have to be inside the mapping lock, as far as I can
>tell, and as far as comments go.
>
>So a stupid patch to just move the dirty page accounting to outside
>the spinlock might help a lot.
>
>Does this attached patch help your contention numbers?

Hi Linus,

The 1BRD tests finished and there are no conclusive changes.

The overall aim7 jobs-per-min slightly decreases and
the overall fsmark files_per_sec slightly increases,
however both are small enough (less than 1%), which are kind of
expected numbers.

NUMA test results should be available tomorrow.

99091700659f4df9  1b5f2eb4a752e1fa7102f37545  testcase/testparams/testbox
----------------  --------------------------  ---------------------------
         %stddev     %change         %stddev
             \          |                \
     71443                       71286        GEO-MEAN aim7.jobs-per-min
       972                         961        aim7/1BRD_48G-btrfs-creat-clo-4-performance/ivb44
     52205                       51525        aim7/1BRD_48G-btrfs-disk_cp-1500-performance/ivb44
   2184471 ±  4%        -6%    2051740 ±  3%  aim7/1BRD_48G-btrfs-disk_rd-9000-performance/ivb44
     47049                       46630 ±  3%  aim7/1BRD_48G-btrfs-disk_rr-1500-performance/ivb44
     24932              -4%      23812        aim7/1BRD_48G-btrfs-disk_rw-1500-performance/ivb44
      5884                        5856        aim7/1BRD_48G-btrfs-disk_src-500-performance/ivb44
     51430                       51286        aim7/1BRD_48G-btrfs-disk_wrt-1500-performance/ivb44
       218                         220        aim7/1BRD_48G-btrfs-sync_disk_rw-10-performance/ivb44
     22777                       23199        aim7/1BRD_48G-ext4-creat-clo-1000-performance/ivb44
    130085                      128991        aim7/1BRD_48G-ext4-disk_cp-3000-performance/ivb44
   2434088 ±  3%        -8%    2232211 ±  4%  aim7/1BRD_48G-ext4-disk_rd-9000-performance/ivb44
    130351                      128977        aim7/1BRD_48G-ext4-disk_rr-3000-performance/ivb44
     73280                       74044        aim7/1BRD_48G-ext4-disk_rw-3000-performance/ivb44
    277035              -3%     268057        aim7/1BRD_48G-ext4-disk_src-3000-performance/ivb44
    127584               4%     132639        aim7/1BRD_48G-ext4-disk_wrt-3000-performance/ivb44
     10571                       10659        aim7/1BRD_48G-ext4-sync_disk_rw-600-performance/ivb44
     36924 ±  7%                 36327        aim7/1BRD_48G-f2fs-creat-clo-1500-performance/ivb44
    117238                      119130        aim7/1BRD_48G-f2fs-disk_cp-3000-performance/ivb44
   2340512 ±  5%               2352619 ± 10%  aim7/1BRD_48G-f2fs-disk_rd-9000-performance/ivb44
    107506 ±  9%         7%     114869        aim7/1BRD_48G-f2fs-disk_rr-3000-performance/ivb44
    105642                      106835        aim7/1BRD_48G-f2fs-disk_rw-3000-performance/ivb44
     26900 ±  3%                 26442 ±  3%  aim7/1BRD_48G-f2fs-disk_src-3000-performance/ivb44
    117124 ±  3%                117678        aim7/1BRD_48G-f2fs-disk_wrt-3000-performance/ivb44
      3689                        3616        aim7/1BRD_48G-f2fs-sync_disk_rw-600-performance/ivb44
     70897                       72758        aim7/1BRD_48G-xfs-creat-clo-1500-performance/ivb44
    267649 ±  3%                270867        aim7/1BRD_48G-xfs-disk_cp-3000-performance/ivb44
    485217 ±  3%                489403        aim7/1BRD_48G-xfs-disk_rd-9000-performance/ivb44
    360451                      359042        aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
    338114                      336838        aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
     60130 ±  5%         4%      62663        aim7/1BRD_48G-xfs-disk_src-3000-performance/ivb44
    403144                      401476        aim7/1BRD_48G-xfs-disk_wrt-3000-performance/ivb44
     26327                       26513        aim7/1BRD_48G-xfs-sync_disk_rw-600-performance/ivb44

99091700659f4df9  1b5f2eb4a752e1fa7102f37545
----------------  --------------------------
      2117                        2138        GEO-MEAN fsmark.files_per_sec
      4325                        4379        fsmark/1x-1t-1BRD_32G-btrfs-4K-4G-fsyncBeforeClose-1fpd-performance/ivb43
      9466 ±  3%         4%       9804        fsmark/1x-1t-1BRD_32G-ext4-4K-4G-fsyncBeforeClose-1fpd-performance/ivb43
       433 ±  5%                   424        fsmark/1x-1t-1BRD_48G-btrfs-4M-40G-NoSync-performance/ivb44
       185 ±  6%         5%        194        fsmark/1x-1t-1BRD_48G-btrfs-4M-40G-fsyncBeforeClose-performance/ivb44
       368 ±  3%        -4%        355 ±  6%  fsmark/1x-1t-1BRD_48G-ext4-4M-40G-NoSync-performance/ivb44
       191                         191        fsmark/1x-1t-1BRD_48G-ext4-4M-40G-fsyncBeforeClose-performance/ivb44
       393 ±  4%                   397 ±  4%  fsmark/1x-1t-1BRD_48G-xfs-4M-40G-NoSync-performance/ivb44
       200                         201        fsmark/1x-1t-1BRD_48G-xfs-4M-40G-fsyncBeforeClose-performance/ivb44
       924              -3%        896 ±  3%  fsmark/1x-1t-1HDD-xfs-4K-400M-fsyncBeforeClose-1fpd-performance/ivb43
       488 ±  3%         6%        516        fsmark/1x-64t-1BRD_48G-btrfs-4M-40G-NoSync-performance/ivb44
       559                         564        fsmark/1x-64t-1BRD_48G-ext4-4M-40G-NoSync-performance/ivb44
      1130                        1111        fsmark/1x-64t-1BRD_48G-ext4-4M-40G-fsyncBeforeClose-performance/ivb44
       526 ±  7%         6%        557        fsmark/1x-64t-1BRD_48G-xfs-4M-40G-NoSync-performance/ivb44
      1583 ±  3%                  1620        fsmark/1x-64t-1BRD_48G-xfs-4M-40G-fsyncBeforeClose-performance/ivb44
     33202                       33208        fsmark/8-1SSD-16-ext4-8K-75G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
     33889                       33784        fsmark/8-1SSD-16-ext4-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
     25576                       25509        fsmark/8-1SSD-32-xfs-9B-30G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
      9117                        9079        fsmark/8-1SSD-4-btrfs-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
     13288                       13261        fsmark/8-1SSD-4-btrfs-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
     18851 ± 11%        11%      21013        fsmark/8-1SSD-4-f2fs-8K-72G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
     24343              -4%      23473 ±  4%  fsmark/8-1SSD-4-f2fs-9B-40G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4

Thanks,
Fengguang

[toc] | [prev] | [next] | [standalone]


#1460899 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-12 04:30 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5aop-44O-3@gated-at.bofh.it>
In reply to#1460883
On Fri, Aug 12, 2016 at 10:54:42AM +1000, Dave Chinner wrote:
> I'm now going to test Christoph's theory that this is an "overwrite
> doing lots of block mapping" issue. More on that to follow.

Ok, so going back to the profiles, I can say it's not an overwrite
issue, because there is delayed allocation showing up in the
profile. Lots of it. Which lead me to think "maybe the benchmark is
just completely *dumb*".

And, as usual, that's the answer. Here's the reproducer:

# sudo mkfs.xfs -f -m crc=0 /dev/pmem1
# sudo mount -o noatime /dev/pmem1 /mnt/scratch
# sudo xfs_io -f -c "pwrite 0 512m -b 1" /mnt/scratch/fooey

And here's the profile:

   4.50%  [kernel]  [k] xfs_bmapi_read
   3.64%  [kernel]  [k] __block_commit_write.isra.30
   3.55%  [kernel]  [k] __radix_tree_lookup
   3.46%  [kernel]  [k] up_write
   3.43%  [kernel]  [k] ___might_sleep
   3.09%  [kernel]  [k] entry_SYSCALL_64_fastpath
   3.01%  [kernel]  [k] xfs_iext_bno_to_ext
   3.01%  [kernel]  [k] find_get_entry
   2.98%  [kernel]  [k] down_write
   2.71%  [kernel]  [k] mark_buffer_dirty
   2.52%  [kernel]  [k] __mark_inode_dirty
   2.38%  [kernel]  [k] unlock_page
   2.14%  [kernel]  [k] xfs_break_layouts
   2.07%  [kernel]  [k] xfs_bmapi_update_map
   2.06%  [kernel]  [k] xfs_bmap_search_extents
   2.04%  [kernel]  [k] xfs_iomap_write_delay
   2.00%  [kernel]  [k] generic_write_checks
   1.96%  [kernel]  [k] xfs_bmap_search_multi_extents
   1.90%  [kernel]  [k] __xfs_bmbt_get_all
   1.89%  [kernel]  [k] balance_dirty_pages_ratelimited
   1.82%  [kernel]  [k] wait_for_stable_page
   1.76%  [kernel]  [k] xfs_file_write_iter
   1.68%  [kernel]  [k] xfs_iomap_eof_want_preallocate
   1.68%  [kernel]  [k] xfs_bmapi_delay
   1.67%  [kernel]  [k] iomap_write_actor
   1.60%  [kernel]  [k] xfs_file_buffered_aio_write
   1.56%  [kernel]  [k] __might_sleep
   1.48%  [kernel]  [k] do_raw_spin_lock
   1.44%  [kernel]  [k] generic_write_end
   1.41%  [kernel]  [k] pagecache_get_page
   1.38%  [kernel]  [k] xfs_bmapi_trim_map
   1.21%  [kernel]  [k] __block_write_begin_int
   1.17%  [kernel]  [k] vfs_write
   1.17%  [kernel]  [k] xfs_file_iomap_begin
   1.17%  [kernel]  [k] xfs_bmbt_get_startoff
   1.14%  [kernel]  [k] iomap_apply
   1.08%  [kernel]  [k] xfs_iunlock
   1.08%  [kernel]  [k] iov_iter_copy_from_user_atomic
   0.97%  [kernel]  [k] xfs_file_aio_write_checks
   0.96%  [kernel]  [k] xfs_ilock
.....

Yeah, I'm doing a sequential write in *1 byte pwrite() calls*.

Ok, so the benchmark isn't /quite/ that abysmally stupid. It's
still, ah, extremely challenged:

        if (NBUFSIZE != 1024) { /* enforce known block size */
                fprintf(stderr, "NBUFSIZE changed to %d\n", NBUFSIZE);
                exit(1);
        }

i.e. it's hard coded to do all it's "disk" IO in 1k block sizes.
Every read, every write, every file copy, etc are all done with a
1024 byte buffer. There are lots of loops that look like:

	while (--n) {
		write(fd, nbuf, sizeof nbuf)
	}

where n is the file size specified in the job file. Those loops are
what is generating the profile we see: repeated partial page writes
that extend the file.

IOWs, the benchmark is doing exactly what we document in the fstat()
man page *not to do* as it is will cause inefficient IO patterns:

       The st_blksize field gives the "preferred" blocksize for
       efficient filesystem I/O.  (Writing to a file in smaller
       chunks may cause an inefficient  read-modify-rewrite.)

The smallest we ever set st_blksize to is PAGE_SIZE, so the
benchmark is running well known and documented (at least 10 years
ago) slow paths through the IO stack.  I'm very tempted now simply
to say that the aim7 disk benchmark is showing it's age and as such
the results are not actually reflective of what typical applications
will see.

Christoph, maybe there's something we can do to only trigger
speculative prealloc growth checks if the new file size crosses the end of
the currently allocated block at the EOF. That would chop out a fair
chunk of the xfs_bmapi_read calls being done in this workload. I'm
not sure how much effort we should spend optimising this slow path,
though....

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1460901 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromLinus Torvalds <torvalds@linux-foundation.org>
Date2016-08-12 04:40 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5ay5-488-5@gated-at.bofh.it>
In reply to#1460899
On Thu, Aug 11, 2016 at 7:23 PM, Dave Chinner <david@fromorbit.com> wrote:
>
> And, as usual, that's the answer. Here's the reproducer:
>
> # sudo mkfs.xfs -f -m crc=0 /dev/pmem1
> # sudo mount -o noatime /dev/pmem1 /mnt/scratch
> # sudo xfs_io -f -c "pwrite 0 512m -b 1" /mnt/scratch/fooey

Heh. Ok, so 1 byte or 1kB at a time is pretty much the same thing, yeah.

And I guess that also explains why the system call entry showed up so
high in the profiles.

I'l take another look at tree_lock tomorrow, but it sounds like this
particular AIM regression is now effectively a solved (or at least
known) issue. Thanks,

              Linus

[toc] | [prev] | [next] | [standalone]


#1460920 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromChristoph Hellwig <hch@lst.de>
Date2016-08-12 05:00 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5aRr-4fz-5@gated-at.bofh.it>
In reply to#1460899
On Fri, Aug 12, 2016 at 12:23:29PM +1000, Dave Chinner wrote:
> Christoph, maybe there's something we can do to only trigger
> speculative prealloc growth checks if the new file size crosses the end of
> the currently allocated block at the EOF. That would chop out a fair
> chunk of the xfs_bmapi_read calls being done in this workload. I'm
> not sure how much effort we should spend optimising this slow path,
> though....

I can look at that, but indeed optimizing this patch seems a bit
stupid.  The other thing we could do is to optimize xfs_bmapi_read - even
if it shouldn't be called this often it seems like it should waste a whole
lot less CPU cycles.

[toc] | [prev] | [next] | [standalone]


#1460934 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromLinus Torvalds <torvalds@linux-foundation.org>
Date2016-08-12 05:30 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5bkt-4Fk-1@gated-at.bofh.it>
In reply to#1460920
On Thu, Aug 11, 2016 at 7:52 PM, Christoph Hellwig <hch@lst.de> wrote:
>
> I can look at that, but indeed optimizing this patch seems a bit
> stupid.

The "write less than a full block to the end of the file" is actually
a reasonably common case.

It may not make for a great filesystem benchmark, but it also isn't
actually insane. People who do logging in user space do this all the
time, for example. And it is *not* stupid in that context. Not at all.

It's never going to be the *main* thing you do (unless you're AIM),
but I do think it's worth fixing.

And AIM7 remains one of those odd benchmarks that people use. I'm not
quite sure why, but I really do think that the normal "append smaller
chunks to the end of the file" should absolutely not be dismissed as
stupid.

                Linus

[toc] | [prev] | [next] | [standalone]


#1460942 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-12 06:20 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5c6S-5aJ-3@gated-at.bofh.it>
In reply to#1460934
On Thu, Aug 11, 2016 at 08:20:53PM -0700, Linus Torvalds wrote:
> On Thu, Aug 11, 2016 at 7:52 PM, Christoph Hellwig <hch@lst.de> wrote:
> >
> > I can look at that, but indeed optimizing this patch seems a bit
> > stupid.
> 
> The "write less than a full block to the end of the file" is actually
> a reasonably common case.
> 
> It may not make for a great filesystem benchmark, but it also isn't
> actually insane. People who do logging in user space do this all the
> time, for example. And it is *not* stupid in that context. Not at all.
> 
> It's never going to be the *main* thing you do (unless you're AIM),
> but I do think it's worth fixing.
> 
> And AIM7 remains one of those odd benchmarks that people use. I'm not
> quite sure why, but I really do think that the normal "append smaller
> chunks to the end of the file" should absolutely not be dismissed as
> stupid.

Yes, I agree that there are reasons for making sub-block IO work
well (which is why I'm looking to try to fix it), but that does't
mean the benchmark is sane. aim7 is, technically, a "scalability
benchmark". As such, expecting tiny writes to scale to moving large
amounts of data is the "stupid" thing it does. If you scale up the
amount of data you need to move, tehn you ned to scale up the
efficiency of moving that data. Case in point - writing 1GB of data
in 1kb chunks to XFs on a local /dev/pmem1 runs at ~600MB/s, whilst
moving it it in 1MB chunks runs at 1.9GB/s. aim7 doesn't actually
stress the scalability of the hardware, because inefficiencies in
it's implementation prevent it from getting to those limits.

That's what aim7 misses - as speeds and capabilities go up, the way
code needs to be written to make efficient use of the hardware also
changes. e.g. High throughput logging solutions don't write every
incoming log event immediately - they aggregate them into larger
buffers and then write those, knowing that they can support much
higher logging rates by doing this....

That's why running aim7 as your "does the filesystem scale"
benchmark is somewhat irrelevant to scaling applications on high
performance systems these days - users with fast storage will be
expecting to see that 1.9GB/s throughput from their app, not
600MB/s....

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1460947 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromLinus Torvalds <torvalds@linux-foundation.org>
Date2016-08-12 07:10 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5cTf-5LK-1@gated-at.bofh.it>
In reply to#1460942
On Thu, Aug 11, 2016 at 9:16 PM, Dave Chinner <david@fromorbit.com> wrote:
>
> That's why running aim7 as your "does the filesystem scale"
> benchmark is somewhat irrelevant to scaling applications on high
> performance systems these days

Yes, don't get me wrong - I'm not at all trying to say that AIM7 is a
good benchmark. It's just that I think what it happens to test is
still meaningful, even if it's not necessarily in any way some kind of
"high performance IO" thing.

There are probably lots of other more important loads, I just reacted
to Christoph seeming to argue that the AIM7 behavior was _so_ broken
that we shouldn't even care. It's not _that_ broken, it's just not
about high-performance IO streaming, it happens to test something else
entirely.

We've actually had AIM7 occasionally find other issues just because
some of the things it does is so odd.

Iirc it has a fork test that doesn't execve (very unusual - you'd
generally use threads if you care about performance), and that has
shown issues with our anon_vma scaling before anything else did.

I also seem to remember some odd pty open/close/ioctl subtest that
showed problems with some of the last remnants of the old BKL (the
test probably actually tested something else, but ended up choking on
the odd tty things).

So in general, I'm not a fan of AIM as a benchmark, but it actually
_has_ found lots of real issues because it tends to do things that
kernel developers think are insane.

And let's face it, user programs doing odd and not very efficient
things should be considered par for the course. We're never going to
get rid of insane user programs, so we might as well fix the
performance problems even when we say "that's just stupid".

            Linus

[toc] | [prev] | [next] | [standalone]


#1460962 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-12 08:10 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5dPj-6mv-21@gated-at.bofh.it>
In reply to#1460947
On Thu, Aug 11, 2016 at 10:02:39PM -0700, Linus Torvalds wrote:
> On Thu, Aug 11, 2016 at 9:16 PM, Dave Chinner <david@fromorbit.com> wrote:
> >
> > That's why running aim7 as your "does the filesystem scale"
> > benchmark is somewhat irrelevant to scaling applications on high
> > performance systems these days
> 
> Yes, don't get me wrong - I'm not at all trying to say that AIM7 is a
> good benchmark. It's just that I think what it happens to test is
> still meaningful, even if it's not necessarily in any way some kind of
> "high performance IO" thing.
> 
> There are probably lots of other more important loads, I just reacted
> to Christoph seeming to argue that the AIM7 behavior was _so_ broken
> that we shouldn't even care. It's not _that_ broken, it's just not
> about high-performance IO streaming, it happens to test something else
> entirely.

Right - I admit that my first reaction once I worked out what the
problem was is exactly what Christoph said. But after looking at it
further, regardless of how crappy the benchmark it, it is a
regression....

> We've actually had AIM7 occasionally find other issues just because
> some of the things it does is so odd.

*nod*

> And let's face it, user programs doing odd and not very efficient
> things should be considered par for the course. We're never going to
> get rid of insane user programs, so we might as well fix the
> performance problems even when we say "that's just stupid".

Yup, that's what I'm doing :/

It looks like the underlying cause is that the old block mapping
code only fed filesystem block size lengths into
xfs_iomap_write_delay(), whereas the iomap code is feeding the
(capped) write() length into it. Hence xfs_iomap_write_delay() is
not detecting the need for speculative preallocation correctly on
these sub-block writes. The profile looks better for the 1 byte
write - I've combined the old and new for comparison below:

	4.22%		__block_commit_write.isra.30
	3.80%		up_write
	3.74%		xfs_bmapi_read
	3.65%		___might_sleep
	3.55%		down_write
	3.20%		entry_SYSCALL_64_fastpath
	3.02%		mark_buffer_dirty
	2.78%		__mark_inode_dirty
	2.78%		unlock_page
	2.59%		xfs_break_layouts
	2.47%		xfs_iext_bno_to_ext
	2.38%		__block_write_begin_int
	2.22%		find_get_entry
	2.17%		xfs_file_write_iter
	2.16%		__radix_tree_lookup
	2.13%		iomap_write_actor
	2.04%		xfs_bmap_search_extents
	1.98%		__might_sleep
	1.84%		xfs_file_buffered_aio_write
	1.76%		iomap_apply
	1.71%		generic_write_end
	1.68%		vfs_write
	1.66%		iov_iter_copy_from_user_atomic
	1.56%		xfs_bmap_search_multi_extents
	1.55%		__vfs_write
	1.52%		pagecache_get_page
	1.46%		xfs_bmapi_update_map
	1.33%		xfs_iunlock
	1.32%		xfs_iomap_write_delay
	1.29%		xfs_file_iomap_begin
	1.29%		do_raw_spin_lock
	1.29%		__xfs_bmbt_get_all
	1.21%		iov_iter_advance
	1.20%		xfs_file_aio_write_checks
	1.14%		xfs_ilock
	1.11%		balance_dirty_pages_ratelimited
	1.10%		xfs_bmapi_trim_map
	1.06%		xfs_iomap_eof_want_preallocate
	1.00%		xfs_bmapi_delay

Comparison of common functions:

Old	New		function
4.50%	3.74%		xfs_bmapi_read
3.64%	4.22%		__block_commit_write.isra.30
3.55%	2.16%		__radix_tree_lookup
3.46%	3.80%		up_write
3.43%	3.65%		___might_sleep
3.09%	3.20%		entry_SYSCALL_64_fastpath
3.01%	2.47%		xfs_iext_bno_to_ext
3.01%	2.22%		find_get_entry
2.98%	3.55%		down_write
2.71%	3.02%		mark_buffer_dirty
2.52%	2.78%		__mark_inode_dirty
2.38%	2.78%		unlock_page
2.14%	2.59%		xfs_break_layouts
2.07%	1.46%		xfs_bmapi_update_map
2.06%	2.04%		xfs_bmap_search_extents
2.04%	1.32%		xfs_iomap_write_delay
2.00%	0.38%		generic_write_checks
1.96%	1.56%		xfs_bmap_search_multi_extents
1.90%	1.29%		__xfs_bmbt_get_all
1.89%	1.11%		balance_dirty_pages_ratelimited
1.82%	0.28%		wait_for_stable_page
1.76%	2.17%		xfs_file_write_iter
1.68%	1.06%		xfs_iomap_eof_want_preallocate
1.68%	1.00%		xfs_bmapi_delay
1.67%	2.13%		iomap_write_actor
1.60%	1.84%		xfs_file_buffered_aio_write
1.56%	1.98%		__might_sleep
1.48%	1.29%		do_raw_spin_lock
1.44%	1.71%		generic_write_end
1.41%	1.52%		pagecache_get_page
1.38%	1.10%		xfs_bmapi_trim_map
1.21%	2.38%		__block_write_begin_int
1.17%	1.68%		vfs_write
1.17%	1.29%		xfs_file_iomap_begin

This shows more time spent in functions above xfs_file_iomap_begin
(which does the block mapping and allocation) and less time spent
below it. i.e. the generic functions as showing higher CPU usage
and the xfs* functions are showing signficantly reduced CPU usage.
This implies that we're doing a lot less block mapping work....

lkp-folk: the patch I've just tested it attached below - can you
feed that through your test and see if it fixes the regression?

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

xfs: correct speculative prealloc on extending subpage writes

From: Dave Chinner <dchinner@redhat.com>

When a write occurs that extends the file, we check to see if we
need to preallocate more delalloc space.  When we do sub-page
writes, the new iomap write path passes a sub-block write length to
the block mapping code.  xfs_iomap_write_delay does not expect to be
pased byte counts smaller than one filesystem block, so it ends up
checking the BMBT on for blocks beyond EOF on every write,
regardless of whether we need to or not. This causes a regression in
aim7 benchmarks as it is full of sub-page writes.

To fix this, clamp the minimum length of a mapping request coming
through xfs_file_iomap_begin() to one filesystem block. This ensures
we are passing the same length to xfs_iomap_write_delay() as we did
when calling through the get_blocks path. This substantially reduces
the amount of lookup load being placed on the BMBT during sub-block
write loads.

Signed-off-by: Dave Chinner <dchinner@redhat.com>
---
 fs/xfs/xfs_iomap.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
index cf697eb..486b75b 100644
--- a/fs/xfs/xfs_iomap.c
+++ b/fs/xfs/xfs_iomap.c
@@ -1036,10 +1036,15 @@ xfs_file_iomap_begin(
 		 * number pulled out of thin air as a best guess for initial
 		 * testing.
 		 *
+		 * xfs_iomap_write_delay() only works if the length passed in is
+		 * >= one filesystem block. Hence we need to clamp the minimum
+		 * length we map, too.
+		 *
 		 * Note that the values needs to be less than 32-bits wide until
 		 * the lower level functions are updated.
 		 */
 		length = min_t(loff_t, length, 1024 * PAGE_SIZE);
+		length = max_t(loff_t, length, (1 << inode->i_blkbits));
 		if (xfs_get_extsz_hint(ip)) {
 			/*
 			 * xfs_iomap_write_direct() expects the shared lock. It

[toc] | [prev] | [next] | [standalone]


#1460971 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromYe Xiaolong <xiaolong.ye@intel.com>
Date2016-08-12 08:40 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5eil-6wT-7@gated-at.bofh.it>
In reply to#1460962
On 08/12, Dave Chinner wrote:
>On Thu, Aug 11, 2016 at 10:02:39PM -0700, Linus Torvalds wrote:
>> On Thu, Aug 11, 2016 at 9:16 PM, Dave Chinner <david@fromorbit.com> wrote:
>> >
>> > That's why running aim7 as your "does the filesystem scale"
>> > benchmark is somewhat irrelevant to scaling applications on high
>> > performance systems these days
>> 
>> Yes, don't get me wrong - I'm not at all trying to say that AIM7 is a
>> good benchmark. It's just that I think what it happens to test is
>> still meaningful, even if it's not necessarily in any way some kind of
>> "high performance IO" thing.
>> 
>> There are probably lots of other more important loads, I just reacted
>> to Christoph seeming to argue that the AIM7 behavior was _so_ broken
>> that we shouldn't even care. It's not _that_ broken, it's just not
>> about high-performance IO streaming, it happens to test something else
>> entirely.
>
>Right - I admit that my first reaction once I worked out what the
>problem was is exactly what Christoph said. But after looking at it
>further, regardless of how crappy the benchmark it, it is a
>regression....
>
>> We've actually had AIM7 occasionally find other issues just because
>> some of the things it does is so odd.
>
>*nod*
>
>> And let's face it, user programs doing odd and not very efficient
>> things should be considered par for the course. We're never going to
>> get rid of insane user programs, so we might as well fix the
>> performance problems even when we say "that's just stupid".
>
>Yup, that's what I'm doing :/
>
>It looks like the underlying cause is that the old block mapping
>code only fed filesystem block size lengths into
>xfs_iomap_write_delay(), whereas the iomap code is feeding the
>(capped) write() length into it. Hence xfs_iomap_write_delay() is
>not detecting the need for speculative preallocation correctly on
>these sub-block writes. The profile looks better for the 1 byte
>write - I've combined the old and new for comparison below:
>
>	4.22%		__block_commit_write.isra.30
>	3.80%		up_write
>	3.74%		xfs_bmapi_read
>	3.65%		___might_sleep
>	3.55%		down_write
>	3.20%		entry_SYSCALL_64_fastpath
>	3.02%		mark_buffer_dirty
>	2.78%		__mark_inode_dirty
>	2.78%		unlock_page
>	2.59%		xfs_break_layouts
>	2.47%		xfs_iext_bno_to_ext
>	2.38%		__block_write_begin_int
>	2.22%		find_get_entry
>	2.17%		xfs_file_write_iter
>	2.16%		__radix_tree_lookup
>	2.13%		iomap_write_actor
>	2.04%		xfs_bmap_search_extents
>	1.98%		__might_sleep
>	1.84%		xfs_file_buffered_aio_write
>	1.76%		iomap_apply
>	1.71%		generic_write_end
>	1.68%		vfs_write
>	1.66%		iov_iter_copy_from_user_atomic
>	1.56%		xfs_bmap_search_multi_extents
>	1.55%		__vfs_write
>	1.52%		pagecache_get_page
>	1.46%		xfs_bmapi_update_map
>	1.33%		xfs_iunlock
>	1.32%		xfs_iomap_write_delay
>	1.29%		xfs_file_iomap_begin
>	1.29%		do_raw_spin_lock
>	1.29%		__xfs_bmbt_get_all
>	1.21%		iov_iter_advance
>	1.20%		xfs_file_aio_write_checks
>	1.14%		xfs_ilock
>	1.11%		balance_dirty_pages_ratelimited
>	1.10%		xfs_bmapi_trim_map
>	1.06%		xfs_iomap_eof_want_preallocate
>	1.00%		xfs_bmapi_delay
>
>Comparison of common functions:
>
>Old	New		function
>4.50%	3.74%		xfs_bmapi_read
>3.64%	4.22%		__block_commit_write.isra.30
>3.55%	2.16%		__radix_tree_lookup
>3.46%	3.80%		up_write
>3.43%	3.65%		___might_sleep
>3.09%	3.20%		entry_SYSCALL_64_fastpath
>3.01%	2.47%		xfs_iext_bno_to_ext
>3.01%	2.22%		find_get_entry
>2.98%	3.55%		down_write
>2.71%	3.02%		mark_buffer_dirty
>2.52%	2.78%		__mark_inode_dirty
>2.38%	2.78%		unlock_page
>2.14%	2.59%		xfs_break_layouts
>2.07%	1.46%		xfs_bmapi_update_map
>2.06%	2.04%		xfs_bmap_search_extents
>2.04%	1.32%		xfs_iomap_write_delay
>2.00%	0.38%		generic_write_checks
>1.96%	1.56%		xfs_bmap_search_multi_extents
>1.90%	1.29%		__xfs_bmbt_get_all
>1.89%	1.11%		balance_dirty_pages_ratelimited
>1.82%	0.28%		wait_for_stable_page
>1.76%	2.17%		xfs_file_write_iter
>1.68%	1.06%		xfs_iomap_eof_want_preallocate
>1.68%	1.00%		xfs_bmapi_delay
>1.67%	2.13%		iomap_write_actor
>1.60%	1.84%		xfs_file_buffered_aio_write
>1.56%	1.98%		__might_sleep
>1.48%	1.29%		do_raw_spin_lock
>1.44%	1.71%		generic_write_end
>1.41%	1.52%		pagecache_get_page
>1.38%	1.10%		xfs_bmapi_trim_map
>1.21%	2.38%		__block_write_begin_int
>1.17%	1.68%		vfs_write
>1.17%	1.29%		xfs_file_iomap_begin
>
>This shows more time spent in functions above xfs_file_iomap_begin
>(which does the block mapping and allocation) and less time spent
>below it. i.e. the generic functions as showing higher CPU usage
>and the xfs* functions are showing signficantly reduced CPU usage.
>This implies that we're doing a lot less block mapping work....
>
>lkp-folk: the patch I've just tested it attached below - can you
>feed that through your test and see if it fixes the regression?
>

Hi, Dave

I am verifying your fix patch in lkp environment now, will send the
result once I get it.

Thanks,
Xiaolong
>Cheers,
>
>Dave.
>-- 
>Dave Chinner
>david@fromorbit.com
>
>xfs: correct speculative prealloc on extending subpage writes
>
>From: Dave Chinner <dchinner@redhat.com>
>
>When a write occurs that extends the file, we check to see if we
>need to preallocate more delalloc space.  When we do sub-page
>writes, the new iomap write path passes a sub-block write length to
>the block mapping code.  xfs_iomap_write_delay does not expect to be
>pased byte counts smaller than one filesystem block, so it ends up
>checking the BMBT on for blocks beyond EOF on every write,
>regardless of whether we need to or not. This causes a regression in
>aim7 benchmarks as it is full of sub-page writes.
>
>To fix this, clamp the minimum length of a mapping request coming
>through xfs_file_iomap_begin() to one filesystem block. This ensures
>we are passing the same length to xfs_iomap_write_delay() as we did
>when calling through the get_blocks path. This substantially reduces
>the amount of lookup load being placed on the BMBT during sub-block
>write loads.
>
>Signed-off-by: Dave Chinner <dchinner@redhat.com>
>---
> fs/xfs/xfs_iomap.c | 5 +++++
> 1 file changed, 5 insertions(+)
>
>diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
>index cf697eb..486b75b 100644
>--- a/fs/xfs/xfs_iomap.c
>+++ b/fs/xfs/xfs_iomap.c
>@@ -1036,10 +1036,15 @@ xfs_file_iomap_begin(
> 		 * number pulled out of thin air as a best guess for initial
> 		 * testing.
> 		 *
>+		 * xfs_iomap_write_delay() only works if the length passed in is
>+		 * >= one filesystem block. Hence we need to clamp the minimum
>+		 * length we map, too.
>+		 *
> 		 * Note that the values needs to be less than 32-bits wide until
> 		 * the lower level functions are updated.
> 		 */
> 		length = min_t(loff_t, length, 1024 * PAGE_SIZE);
>+		length = max_t(loff_t, length, (1 << inode->i_blkbits));
> 		if (xfs_get_extsz_hint(ip)) {
> 			/*
> 			 * xfs_iomap_write_direct() expects the shared lock. It
>_______________________________________________
>LKP mailing list
>LKP@lists.01.org
>https://lists.01.org/mailman/listinfo/lkp

[toc] | [prev] | [next] | [standalone]


#1461019 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromYe Xiaolong <xiaolong.ye@intel.com>
Date2016-08-12 11:00 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5gtP-7QJ-3@gated-at.bofh.it>
In reply to#1460971
On 08/12, Ye Xiaolong wrote:
>On 08/12, Dave Chinner wrote:

[snip]

>>lkp-folk: the patch I've just tested it attached below - can you
>>feed that through your test and see if it fixes the regression?
>>
>
>Hi, Dave
>
>I am verifying your fix patch in lkp environment now, will send the
>result once I get it.
>

Here is the test result.

commit 636b594f38278080db93f2d67d11d31700924f5d
Author:     Dave Chinner <dchinner@redhat.com>
AuthorDate: Fri Aug 12 14:23:44 2016 +0800
Commit:     Xiaolong Ye <xiaolong.ye@intel.com>
CommitDate: Fri Aug 12 14:23:44 2016 +0800

    When a write occurs that extends the file, we check to see if we
    need to preallocate more delalloc space.  When we do sub-page
    writes, the new iomap write path passes a sub-block write length to
    the block mapping code.  xfs_iomap_write_delay does not expect to be
    pased byte counts smaller than one filesystem block, so it ends up
    checking the BMBT on for blocks beyond EOF on every write,
    regardless of whether we need to or not. This causes a regression in
    aim7 benchmarks as it is full of sub-page writes.
    
    To fix this, clamp the minimum length of a mapping request coming
    through xfs_file_iomap_begin() to one filesystem block. This ensures
    we are passing the same length to xfs_iomap_write_delay() as we did
    when calling through the get_blocks path. This substantially reduces
    the amount of lookup load being placed on the BMBT during sub-block
    write loads.
    
    Signed-off-by: Dave Chinner <dchinner@redhat.com>
---
 fs/xfs/xfs_iomap.c | 5 +++++
 1 file changed, 5 insertions(+)

diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
index 620fc91..5eaace0 100644
--- a/fs/xfs/xfs_iomap.c
+++ b/fs/xfs/xfs_iomap.c
@@ -1015,10 +1015,15 @@ xfs_file_iomap_begin(
                 * number pulled out of thin air as a best guess for initial
                 * testing.
                 *
+                * xfs_iomap_write_delay() only works if the length passed in is
+                * >= one filesystem block. Hence we need to clamp the minimum
+                * length we map, too.
+                *
                 * Note that the values needs to be less than 32-bits wide until
                 * the lower level functions are updated.
                 */
                length = min_t(loff_t, length, 1024 * PAGE_SIZE);
+               length = max_t(loff_t, length, (1 << inode->i_blkbits));
                if (xfs_get_extsz_hint(ip)) {
                        /*
                         * xfs_iomap_write_direct() expects the shared lock. It

f0c6bcba74ac51cb  68a9f5e7007c1afa2cf6830b69  636b594f38278080db93f2d67d  
----------------  --------------------------  --------------------------  
         %stddev     %change         %stddev     %change         %stddev
             \          |                \          |                \  
    484435 ±  0%     -13.3%     420004 ±  0%     -14.0%     416777 ±  0%  aim7.jobs-per-min
      6491 ±  3%     +30.8%       8491 ±  0%     +35.7%       8806 ±  1%  aim7.time.involuntary_context_switches
       376 ±  0%     +28.4%        484 ±  0%     +29.6%        488 ±  0%  aim7.time.system_time
    430512 ±  0%     -20.1%     343838 ±  0%     -19.7%     345708 ±  0%  aim7.time.voluntary_context_switches
     37.37 ±  0%     +15.3%      43.09 ±  0%     +16.1%      43.41 ±  0%  aim7.time.elapsed_time
     37.37 ±  0%     +15.3%      43.09 ±  0%     +16.1%      43.41 ±  0%  aim7.time.elapsed_time.max
    155184 ±  1%      -2.1%     151864 ±  1%      -2.7%     150937 ±  1%  aim7.time.minor_page_faults
         0 ±  0%      +Inf%     215412 ±141%      +Inf%     334416 ± 75%  latency_stats.sum.wait_on_page_bit.__migration_entry_wait.migration_entry_wait.handle_pte_fault.handle_mm_fault.__do_page_fault.do_page_fault.page_fault
     24772 ±  0%     -28.6%      17675 ±  0%     -26.7%      18149 ±  2%  vmstat.system.cs
     26816 ±  8%     +10.2%      29542 ±  1%     +13.3%      30370 ±  1%  interrupts.CAL:Function_call_interrupts
    125122 ± 10%     -10.7%     111758 ± 12%     -11.1%     111223 ± 11%  softirqs.SCHED
      3906 ±  0%     +28.8%       5032 ±  2%     +29.1%       5045 ±  1%  proc-vmstat.nr_active_file
      3444 ±  5%     +41.8%       4884 ±  0%     +25.0%       4304 ± 11%  proc-vmstat.nr_shmem
      4092 ± 14%     +61.2%       6595 ±  1%     +40.0%       5728 ± 15%  proc-vmstat.pgactivate
     15627 ±  0%     +27.7%      19956 ±  1%     +27.4%      19902 ±  0%  meminfo.Active(file)
     16103 ±  3%     +14.3%      18405 ±  8%     +11.2%      17900 ±  1%  meminfo.AnonHugePages
     13777 ±  5%     +43.1%      19709 ±  0%     +25.0%      17220 ± 11%  meminfo.Shmem
   1724300 ± 27%     -40.5%    1025538 ±  1%     -41.3%    1012868 ±  0%  sched_debug.cfs_rq:/.load.max
   1724300 ± 27%     -40.5%    1025538 ±  1%     -41.3%    1012868 ±  0%  sched_debug.cpu.load.max
     37.37 ±  0%     +15.3%      43.09 ±  0%     +16.1%      43.41 ±  0%  time.elapsed_time
     37.37 ±  0%     +15.3%      43.09 ±  0%     +16.1%      43.41 ±  0%  time.elapsed_time.max
      6491 ±  3%     +30.8%       8491 ±  0%     +35.7%       8806 ±  1%  time.involuntary_context_switches
      1037 ±  0%     +10.8%       1148 ±  0%     +10.9%       1149 ±  0%  time.percent_of_cpu_this_job_got
       376 ±  0%     +28.4%        484 ±  0%     +29.6%        488 ±  0%  time.system_time
    430512 ±  0%     -20.1%     343838 ±  0%     -19.7%     345708 ±  0%  time.voluntary_context_switches
    319584 ±  1%     -26.5%     234868 ±  1%     -23.9%     243331 ±  3%  cpuidle.C1-IVT.usage
  52991525 ±  1%     -19.4%   42687208 ±  0%     -20.0%   42368754 ±  0%  cpuidle.C1-IVT.time
     46760 ±  0%     -22.4%      36298 ±  0%     -21.6%      36681 ±  1%  cpuidle.C1E-IVT.usage
   3468808 ±  2%     -19.8%    2783341 ±  3%     -16.9%    2881608 ±  5%  cpuidle.C1E-IVT.time
  12590471 ±  0%     -22.3%    9788585 ±  1%     -21.6%    9866515 ±  1%  cpuidle.C3-IVT.time
     79965 ±  0%     -19.0%      64749 ±  0%     -19.1%      64654 ±  0%  cpuidle.C3-IVT.usage
   1.3e+09 ±  0%     +13.3%  1.473e+09 ±  0%     +13.9%  1.481e+09 ±  0%  cpuidle.C6-IVT.time
     24.18 ±  0%      +9.0%      26.35 ±  0%      +9.6%      26.49 ±  0%  turbostat.%Busy
       686 ±  0%      +9.5%        751 ±  0%      +9.2%        749 ±  1%  turbostat.Avg_MHz
      0.28 ±  0%     -25.0%       0.21 ±  0%     -23.8%       0.21 ±  4%  turbostat.CPU%c3
        79 ±  1%      -0.4%         78 ±  3%     -21.5%         62 ±  2%  turbostat.CoreTmp
        78 ±  0%      +0.4%         79 ±  3%     -21.2%         62 ±  1%  turbostat.PkgTmp
      4.74 ±  0%      -2.7%       4.61 ±  1%     -13.1%       4.12 ±  0%  turbostat.RAMWatt
        51 ±  0%      +0.0%         51 ±  0%    +333.3%        221 ± 10%  slabinfo.dio.active_objs
        51 ±  0%      +0.0%         51 ±  0%    +333.3%        221 ± 10%  slabinfo.dio.num_objs
       876 ±  6%      +2.8%        900 ±  3%     +16.7%       1022 ±  0%  slabinfo.nsproxy.active_objs
       876 ±  6%      +2.8%        900 ±  3%     +16.7%       1022 ±  0%  slabinfo.nsproxy.num_objs
      1975 ± 15%     +63.2%       3224 ± 17%     +45.5%       2874 ± 15%  slabinfo.scsi_data_buffer.active_objs
      1975 ± 15%     +63.2%       3224 ± 17%     +45.5%       2874 ± 15%  slabinfo.scsi_data_buffer.num_objs
       464 ± 15%     +63.3%        758 ± 17%     +46.6%        680 ± 15%  slabinfo.xfs_efd_item.active_objs
       464 ± 15%     +63.3%        758 ± 17%     +46.6%        680 ± 15%  slabinfo.xfs_efd_item.num_objs
      1930 ±  0%     +33.9%       2585 ±  3%     +24.7%       2407 ±  5%  numa-vmstat.node0.nr_active_file
       466 ±  4%     +29.3%        603 ± 14%     +28.9%        601 ± 18%  numa-vmstat.node0.nr_dirty
      1977 ±  1%     +23.6%       2444 ±  1%     +33.6%       2641 ±  7%  numa-vmstat.node1.nr_active_file
     11671 ±  3%     +55.9%      18197 ± 24%     +43.3%      16730 ± 25%  numa-vmstat.node1.nr_anon_pages
      3809 ±  6%     +16.1%       4422 ±  4%     +21.6%       4633 ±  4%  numa-vmstat.node1.nr_alloc_batch
     12026 ±  4%     +64.1%      19734 ± 20%     +43.7%      17276 ± 22%  numa-vmstat.node1.nr_active_anon
      7723 ±  0%     +32.6%      10238 ±  5%     +19.5%       9228 ±  4%  numa-meminfo.node0.Active(file)
      8774 ± 29%      +5.3%       9238 ± 28%     +22.5%      10749 ± 24%  numa-meminfo.node1.Mapped
      7908 ±  1%     +22.9%       9722 ±  3%     +35.8%      10736 ±  3%  numa-meminfo.node1.Active(file)
     46721 ±  3%     +55.9%      72837 ± 24%     +42.8%      66711 ± 26%  numa-meminfo.node1.AnonPages
     56052 ±  3%     +58.2%      88666 ± 17%     +42.2%      79696 ± 19%  numa-meminfo.node1.Active
     48142 ±  4%     +64.0%      78943 ± 19%     +43.2%      68960 ± 22%  numa-meminfo.node1.Active(anon)
 2.658e+11 ±  4%     +24.7%  3.316e+11 ±  2%     +25.9%  3.346e+11 ±  3%  perf-stat.branch-instructions
      0.41 ±  1%      -9.1%       0.37 ±  1%      -9.4%       0.37 ±  1%  perf-stat.branch-miss-rate
  1.09e+09 ±  3%     +13.4%  1.237e+09 ±  1%     +14.1%  1.244e+09 ±  2%  perf-stat.branch-misses
    981138 ±  0%     -18.1%     803696 ±  0%     -16.0%     823913 ±  1%  perf-stat.context-switches
 1.511e+12 ±  5%     +23.4%  1.864e+12 ±  3%     +24.4%   1.88e+12 ±  4%  perf-stat.cpu-cycles
    102600 ±  1%      -7.3%      95075 ±  1%      -5.2%      97261 ±  1%  perf-stat.cpu-migrations
      0.26 ± 12%     -30.8%       0.18 ± 10%     -28.1%       0.19 ± 27%  perf-stat.dTLB-load-miss-rate
 3.164e+11 ±  1%     +39.9%  4.426e+11 ±  4%     +40.0%   4.43e+11 ±  1%  perf-stat.dTLB-loads
      0.03 ± 26%     -41.3%       0.02 ± 13%     -41.8%       0.02 ±  5%  perf-stat.dTLB-store-miss-rate
 2.247e+11 ±  6%     +26.4%  2.839e+11 ±  2%     +29.2%  2.903e+11 ±  5%  perf-stat.dTLB-stores
  34415974 ±  6%      -1.7%   33840719 ± 12%      -6.7%   32119462 ±  2%  perf-stat.iTLB-load-misses
  17863352 ±  4%      +2.1%   18245848 ±  2%      -7.9%   16460161 ±  2%  perf-stat.iTLB-loads
  1.49e+12 ±  4%     +30.1%  1.939e+12 ±  2%     +31.5%  1.959e+12 ±  3%  perf-stat.instructions
     43348 ±  2%     +34.2%      58161 ± 12%     +40.9%      61065 ±  5%  perf-stat.instructions-per-iTLB-miss
      0.99 ±  0%      +5.5%       1.04 ±  0%      +5.7%       1.04 ±  0%  perf-stat.ipc
    262799 ±  0%      +4.4%     274251 ±  1%      +4.3%     274149 ±  0%  perf-stat.minor-faults
     34.12 ±  1%      +2.1%      34.83 ±  0%      +3.5%      35.30 ±  1%  perf-stat.node-load-miss-rate
  46476754 ±  2%      +4.6%   48601269 ±  1%      +6.6%   49534267 ±  0%  perf-stat.node-load-misses
      9.96 ±  0%     +13.4%      11.30 ±  0%     +13.6%      11.31 ±  2%  perf-stat.node-store-miss-rate
  24460859 ±  1%     +14.4%   27971097 ±  1%     +13.8%   27844903 ±  0%  perf-stat.node-store-misses
    262780 ±  0%      +4.4%     274227 ±  1%      +4.3%     274117 ±  0%  perf-stat.page-faults
      0.00 ±  0%      +Inf%      52.94 ±  0%      +Inf%      52.69 ±  0%  perf-profile.cycles-pp.iomap_file_buffered_write.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write
      0.00 ±  0%      +Inf%      52.29 ±  0%      +Inf%      52.11 ±  0%  perf-profile.cycles-pp.iomap_apply.iomap_file_buffered_write.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write
      0.00 ±  0%      +Inf%      34.35 ±  0%      +Inf%      34.05 ±  0%  perf-profile.cycles-pp.iomap_write_actor.iomap_apply.iomap_file_buffered_write.xfs_file_buffered_aio_write.xfs_file_write_iter
      0.00 ±  0%      +Inf%      16.48 ±  0%      +Inf%      16.35 ±  1%  perf-profile.cycles-pp.iomap_write_begin.iomap_write_actor.iomap_apply.iomap_file_buffered_write.xfs_file_buffered_aio_write
      0.00 ±  0%      +Inf%      16.05 ±  0%      +Inf%      16.21 ±  1%  perf-profile.cycles-pp.xfs_file_iomap_begin.iomap_apply.iomap_file_buffered_write.xfs_file_buffered_aio_write.xfs_file_write_iter
      0.00 ±  0%      +Inf%       9.85 ±  0%      +Inf%       9.75 ±  1%  perf-profile.cycles-pp.grab_cache_page_write_begin.iomap_write_begin.iomap_write_actor.iomap_apply.iomap_file_buffered_write
      0.00 ±  0%      +Inf%       9.25 ±  0%      +Inf%       9.18 ±  1%  perf-profile.cycles-pp.pagecache_get_page.grab_cache_page_write_begin.iomap_write_begin.iomap_write_actor.iomap_apply
      0.00 ±  0%      +Inf%       9.08 ±  0%      +Inf%       9.08 ±  1%  perf-profile.cycles-pp.xfs_iomap_write_delay.xfs_file_iomap_begin.iomap_apply.iomap_file_buffered_write.xfs_file_buffered_aio_write
      0.00 ±  0%      +Inf%       7.91 ±  1%      +Inf%       7.90 ±  0%  perf-profile.cycles-pp.generic_write_end.iomap_write_actor.iomap_apply.iomap_file_buffered_write.xfs_file_buffered_aio_write
      0.00 ±  0%      +Inf%       4.69 ±  0%      +Inf%       4.66 ±  0%  perf-profile.cycles-pp.block_write_end.generic_write_end.iomap_write_actor.iomap_apply.iomap_file_buffered_write
      0.00 ±  0%      +Inf%       4.45 ±  1%      +Inf%       4.45 ±  0%  perf-profile.cycles-pp.__block_commit_write.isra.24.block_write_end.generic_write_end.iomap_write_actor.iomap_apply
      0.00 ±  0%      +Inf%       4.14 ±  0%      +Inf%       4.12 ±  1%  perf-profile.cycles-pp.xfs_iomap_eof_want_preallocate.constprop.8.xfs_iomap_write_delay.xfs_file_iomap_begin.iomap_apply.iomap_file_buffered_write
      0.00 ±  0%      +Inf%       3.69 ±  1%      +Inf%       3.69 ±  2%  perf-profile.cycles-pp.add_to_page_cache_lru.pagecache_get_page.grab_cache_page_write_begin.iomap_write_begin.iomap_write_actor
      0.00 ±  0%      +Inf%       3.64 ±  0%      +Inf%       3.62 ±  0%  perf-profile.cycles-pp.__block_write_begin_int.iomap_write_begin.iomap_write_actor.iomap_apply.iomap_file_buffered_write
      0.00 ±  0%      +Inf%       3.44 ±  1%      +Inf%       3.35 ±  2%  perf-profile.cycles-pp.mark_page_accessed.iomap_write_actor.iomap_apply.iomap_file_buffered_write.xfs_file_buffered_aio_write
      0.00 ±  0%      +Inf%       3.04 ±  1%      +Inf%       3.00 ±  3%  perf-profile.cycles-pp.xfs_bmapi_read.xfs_file_iomap_begin.iomap_apply.iomap_file_buffered_write.xfs_file_buffered_aio_write
      0.00 ±  0%      +Inf%       3.22 ±  0%      +Inf%       3.15 ±  1%  perf-profile.cycles-pp.copy_user_enhanced_fast_string.iomap_write_actor.iomap_apply.iomap_file_buffered_write.xfs_file_buffered_aio_write
      0.00 ±  0%      +Inf%       3.06 ±  1%      +Inf%       3.09 ±  0%  perf-profile.cycles-pp.xfs_bmapi_delay.xfs_iomap_write_delay.xfs_file_iomap_begin.iomap_apply.iomap_file_buffered_write
      0.00 ±  0%      +Inf%       3.05 ±  1%      +Inf%       3.05 ±  2%  perf-profile.cycles-pp.xfs_bmapi_read.xfs_iomap_eof_want_preallocate.constprop.8.xfs_iomap_write_delay.xfs_file_iomap_begin.iomap_apply
      0.00 ±  0%      +Inf%       2.78 ±  0%      +Inf%       2.83 ±  1%  perf-profile.cycles-pp.mark_buffer_dirty.__block_commit_write.isra.24.block_write_end.generic_write_end.iomap_write_actor
      0.00 ±  0%      +Inf%       2.68 ±  2%      +Inf%       2.60 ±  1%  perf-profile.cycles-pp.__page_cache_alloc.pagecache_get_page.grab_cache_page_write_begin.iomap_write_begin.iomap_write_actor
      0.00 ±  0%      +Inf%       2.56 ±  2%      +Inf%       2.46 ±  0%  perf-profile.cycles-pp.alloc_pages_current.__page_cache_alloc.pagecache_get_page.grab_cache_page_write_begin.iomap_write_begin
      0.00 ±  0%      +Inf%       2.43 ±  0%      +Inf%       2.42 ±  0%  perf-profile.cycles-pp.memset_erms.iomap_write_begin.iomap_write_actor.iomap_apply.iomap_file_buffered_write
      0.00 ±  0%      +Inf%       1.97 ±  2%      +Inf%       1.90 ±  4%  perf-profile.cycles-pp.xfs_bmap_search_extents.xfs_bmapi_read.xfs_file_iomap_begin.iomap_apply.iomap_file_buffered_write
      0.00 ±  0%      +Inf%       1.55 ±  3%      +Inf%       1.62 ±  2%  perf-profile.cycles-pp.find_get_entry.pagecache_get_page.grab_cache_page_write_begin.iomap_write_begin.iomap_write_actor
      0.00 ±  0%      +Inf%       1.68 ±  1%      +Inf%       1.66 ±  2%  perf-profile.cycles-pp.__add_to_page_cache_locked.add_to_page_cache_lru.pagecache_get_page.grab_cache_page_write_begin.iomap_write_begin
      0.00 ±  0%      +Inf%       1.73 ±  1%      +Inf%       1.71 ±  2%  perf-profile.cycles-pp.xfs_bmap_search_extents.xfs_bmapi_delay.xfs_iomap_write_delay.xfs_file_iomap_begin.iomap_apply
      0.00 ±  0%      +Inf%       1.61 ±  2%      +Inf%       1.64 ±  3%  perf-profile.cycles-pp.xfs_bmap_search_extents.xfs_bmapi_read.xfs_iomap_eof_want_preallocate.constprop.8.xfs_iomap_write_delay.xfs_file_iomap_begin
      0.00 ±  0%      +Inf%       1.52 ±  2%      +Inf%       1.51 ±  4%  perf-profile.cycles-pp.workingset_activation.mark_page_accessed.iomap_write_actor.iomap_apply.iomap_file_buffered_write
      0.00 ±  0%      +Inf%       1.55 ±  1%      +Inf%       1.55 ±  1%  perf-profile.cycles-pp.lru_cache_add.add_to_page_cache_lru.pagecache_get_page.grab_cache_page_write_begin.iomap_write_begin
      0.00 ±  0%      +Inf%       1.53 ±  1%      +Inf%       1.52 ±  1%  perf-profile.cycles-pp.create_page_buffers.__block_write_begin_int.iomap_write_begin.iomap_write_actor.iomap_apply
      0.00 ±  0%      +Inf%       1.46 ±  1%      +Inf%       1.45 ±  3%  perf-profile.cycles-pp.xfs_bmap_search_multi_extents.xfs_bmap_search_extents.xfs_bmapi_read.xfs_file_iomap_begin.iomap_apply
      0.00 ±  0%      +Inf%       1.36 ±  1%      +Inf%       1.39 ±  1%  perf-profile.cycles-pp.unlock_page.generic_write_end.iomap_write_actor.iomap_apply.iomap_file_buffered_write
      0.00 ±  0%      +Inf%       1.18 ±  1%      +Inf%       1.19 ±  1%  perf-profile.cycles-pp.create_empty_buffers.create_page_buffers.__block_write_begin_int.iomap_write_begin.iomap_write_actor
      0.00 ±  0%      +Inf%       1.21 ±  2%      +Inf%       1.23 ±  2%  perf-profile.cycles-pp.xfs_bmap_search_multi_extents.xfs_bmap_search_extents.xfs_bmapi_read.xfs_iomap_eof_want_preallocate.constprop.8.xfs_iomap_write_delay
      0.00 ±  0%      +Inf%       1.24 ±  2%      +Inf%       1.21 ±  2%  perf-profile.cycles-pp.xfs_bmap_search_multi_extents.xfs_bmap_search_extents.xfs_bmapi_delay.xfs_iomap_write_delay.xfs_file_iomap_begin
      0.00 ±  0%      +Inf%       1.14 ±  3%      +Inf%       1.16 ±  3%  perf-profile.cycles-pp.xfs_ilock.xfs_file_iomap_begin.iomap_apply.iomap_file_buffered_write.xfs_file_buffered_aio_write
      0.00 ±  0%      +Inf%       1.09 ±  2%      +Inf%       1.08 ±  1%  perf-profile.cycles-pp.__mark_inode_dirty.generic_write_end.iomap_write_actor.iomap_apply.iomap_file_buffered_write
      0.00 ±  0%      +Inf%       0.95 ±  0%      +Inf%       1.01 ±  3%  perf-profile.cycles-pp.radix_tree_lookup_slot.find_get_entry.pagecache_get_page.grab_cache_page_write_begin.iomap_write_begin
     43.95 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.generic_perform_write.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write
     25.10 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.xfs_vm_write_begin.generic_perform_write.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write
     13.71 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.__block_write_begin.xfs_vm_write_begin.generic_perform_write.xfs_file_buffered_aio_write.xfs_file_write_iter
     11.03 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.xfs_vm_write_end.generic_perform_write.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write
     10.68 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.generic_write_end.xfs_vm_write_end.generic_perform_write.xfs_file_buffered_aio_write.xfs_file_write_iter
     10.96 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.grab_cache_page_write_begin.xfs_vm_write_begin.generic_perform_write.xfs_file_buffered_aio_write.xfs_file_write_iter
     10.36 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.__block_write_begin_int.__block_write_begin.xfs_vm_write_begin.generic_perform_write.xfs_file_buffered_aio_write
     10.37 ±  2%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.pagecache_get_page.grab_cache_page_write_begin.xfs_vm_write_begin.generic_perform_write.xfs_file_buffered_aio_write
      6.46 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.xfs_get_blocks.__block_write_begin_int.__block_write_begin.xfs_vm_write_begin.generic_perform_write
      6.34 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.__xfs_get_blocks.xfs_get_blocks.__block_write_begin_int.__block_write_begin.xfs_vm_write_begin
      6.24 ±  0%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.block_write_end.generic_write_end.xfs_vm_write_end.generic_perform_write.xfs_file_buffered_aio_write
      5.93 ±  0%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.__block_commit_write.isra.24.block_write_end.generic_write_end.xfs_vm_write_end.generic_perform_write
      3.95 ±  2%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.copy_user_enhanced_fast_string.generic_perform_write.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write
      4.02 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.add_to_page_cache_lru.pagecache_get_page.grab_cache_page_write_begin.xfs_vm_write_begin.generic_perform_write
      3.39 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.mark_buffer_dirty.__block_commit_write.isra.24.block_write_end.generic_write_end.xfs_vm_write_end
      3.28 ±  2%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.xfs_iomap_write_delay.__xfs_get_blocks.xfs_get_blocks.__block_write_begin_int.__block_write_begin
      3.03 ±  0%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.memset_erms.__block_write_begin.xfs_vm_write_begin.generic_perform_write.xfs_file_buffered_aio_write
      3.04 ±  3%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.__page_cache_alloc.pagecache_get_page.grab_cache_page_write_begin.xfs_vm_write_begin.generic_perform_write
      2.91 ±  3%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.alloc_pages_current.__page_cache_alloc.pagecache_get_page.grab_cache_page_write_begin.xfs_vm_write_begin
      1.86 ±  2%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.create_page_buffers.__block_write_begin_int.__block_write_begin.xfs_vm_write_begin.generic_perform_write
      1.72 ±  4%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.unlock_page.generic_write_end.xfs_vm_write_end.generic_perform_write.xfs_file_buffered_aio_write
      1.80 ±  1%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.__add_to_page_cache_locked.add_to_page_cache_lru.pagecache_get_page.grab_cache_page_write_begin.xfs_vm_write_begin
      1.83 ±  2%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.find_get_entry.pagecache_get_page.grab_cache_page_write_begin.xfs_vm_write_begin.generic_perform_write
      1.72 ±  2%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.lru_cache_add.add_to_page_cache_lru.pagecache_get_page.grab_cache_page_write_begin.xfs_vm_write_begin
      1.44 ±  3%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.create_empty_buffers.create_page_buffers.__block_write_begin_int.__block_write_begin.xfs_vm_write_begin
      1.32 ±  4%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.__mark_inode_dirty.generic_write_end.xfs_vm_write_end.generic_perform_write.xfs_file_buffered_aio_write
      1.25 ±  0%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.xfs_bmapi_delay.xfs_iomap_write_delay.__xfs_get_blocks.xfs_get_blocks.__block_write_begin_int
      1.23 ±  4%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.xfs_iomap_eof_want_preallocate.constprop.6.xfs_iomap_write_delay.__xfs_get_blocks.xfs_get_blocks.__block_write_begin_int
      1.17 ±  3%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.radix_tree_lookup_slot.find_get_entry.pagecache_get_page.grab_cache_page_write_begin.xfs_vm_write_begin
      1.04 ±  0%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.xfs_bmapi_read.__xfs_get_blocks.xfs_get_blocks.__block_write_begin_int.__block_write_begin
      0.98 ±  5%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.cycles-pp.alloc_page_buffers.create_empty_buffers.create_page_buffers.__block_write_begin_int.__block_write_begin
      1.79 ±  2%     -28.2%       1.28 ±  3%     -27.8%       1.29 ±  4%  perf-profile.cycles-pp.do_unlinkat.sys_unlink.entry_SYSCALL_64_fastpath
      1.79 ±  3%     -27.9%       1.29 ±  3%     -27.7%       1.30 ±  4%  perf-profile.cycles-pp.sys_unlink.entry_SYSCALL_64_fastpath
      1.27 ±  0%     -22.5%       0.99 ±  4%     -24.6%       0.96 ±  4%  perf-profile.cycles-pp.destroy_inode.evict.iput.__dentry_kill.dput
      2.61 ±  1%     -24.3%       1.98 ±  1%     -24.1%       1.98 ±  1%  perf-profile.cycles-pp.do_filp_open.do_sys_open.sys_creat.entry_SYSCALL_64_fastpath
      2.58 ±  1%     -24.1%       1.96 ±  0%     -24.1%       1.96 ±  1%  perf-profile.cycles-pp.path_openat.do_filp_open.do_sys_open.sys_creat.entry_SYSCALL_64_fastpath
      1.07 ±  3%     -23.3%       0.82 ±  3%     -21.1%       0.85 ±  2%  perf-profile.cycles-pp.down_write.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write
      2.66 ±  1%     -24.3%       2.01 ±  1%     -23.7%       2.03 ±  1%  perf-profile.cycles-pp.do_sys_open.sys_creat.entry_SYSCALL_64_fastpath
      2.67 ±  1%     -24.2%       2.02 ±  1%     -23.8%       2.03 ±  1%  perf-profile.cycles-pp.sys_creat.entry_SYSCALL_64_fastpath
      1.24 ±  1%     -23.1%       0.95 ±  4%     -24.2%       0.94 ±  4%  perf-profile.cycles-pp.xfs_fs_destroy_inode.destroy_inode.evict.iput.__dentry_kill
      1.21 ±  1%     -23.4%       0.93 ±  4%     -24.7%       0.91 ±  4%  perf-profile.cycles-pp.xfs_inactive.xfs_fs_destroy_inode.destroy_inode.evict.iput
      0.94 ±  4%     -19.8%       0.76 ±  0%     -21.6%       0.74 ±  4%  perf-profile.cycles-pp.cancel_dirty_page.try_to_free_buffers.xfs_vm_releasepage.try_to_release_page.block_invalidatepage
      1.32 ±  2%     -21.5%       1.04 ±  1%     -22.2%       1.03 ±  3%  perf-profile.cycles-pp.xfs_create.xfs_generic_create.xfs_vn_mknod.xfs_vn_create.path_openat
      1.42 ±  2%     -20.7%       1.13 ±  1%     -22.1%       1.11 ±  3%  perf-profile.cycles-pp.xfs_vn_create.path_openat.do_filp_open.do_sys_open.sys_creat
      2.35 ±  1%     -21.0%       1.86 ±  1%     -20.7%       1.86 ±  1%  perf-profile.cycles-pp.xfs_vm_releasepage.try_to_release_page.block_invalidatepage.xfs_vm_invalidatepage.truncate_inode_page
      1.91 ±  3%     -16.4%       1.59 ±  1%     -19.9%       1.53 ±  1%  perf-profile.cycles-pp.get_page_from_freelist.__alloc_pages_nodemask.alloc_pages_current.__page_cache_alloc.pagecache_get_page
      2.07 ±  1%     -20.4%       1.65 ±  2%     -19.9%       1.66 ±  1%  perf-profile.cycles-pp.try_to_free_buffers.xfs_vm_releasepage.try_to_release_page.block_invalidatepage.xfs_vm_invalidatepage
      1.42 ±  2%     -20.5%       1.13 ±  1%     -22.1%       1.10 ±  2%  perf-profile.cycles-pp.xfs_vn_mknod.xfs_vn_create.path_openat.do_filp_open.do_sys_open
      1.42 ±  2%     -21.2%       1.12 ±  1%     -22.4%       1.10 ±  3%  perf-profile.cycles-pp.xfs_generic_create.xfs_vn_mknod.xfs_vn_create.path_openat.do_filp_open
      1.12 ±  2%     -17.6%       0.92 ±  4%     -22.3%       0.87 ±  4%  perf-profile.cycles-pp.__sb_start_write.vfs_write.sys_write.entry_SYSCALL_64_fastpath
      2.40 ±  1%     -21.0%       1.89 ±  2%     -20.6%       1.90 ±  1%  perf-profile.cycles-pp.try_to_release_page.block_invalidatepage.xfs_vm_invalidatepage.truncate_inode_page.truncate_inode_pages_range
      1.29 ±  3%     -18.9%       1.04 ±  1%     -17.9%       1.06 ±  1%  perf-profile.cycles-pp.xfs_ilock.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write
      3.42 ±  0%     -20.9%       2.71 ±  2%     -20.3%       2.73 ±  2%  perf-profile.cycles-pp.block_invalidatepage.xfs_vm_invalidatepage.truncate_inode_page.truncate_inode_pages_range.truncate_inode_pages_final
      5.96 ±  1%     -20.0%       4.77 ±  0%     -19.4%       4.81 ±  1%  perf-profile.cycles-pp.truncate_inode_page.truncate_inode_pages_range.truncate_inode_pages_final.evict.iput
      3.54 ±  0%     -20.8%       2.81 ±  1%     -20.0%       2.83 ±  2%  perf-profile.cycles-pp.xfs_vm_invalidatepage.truncate_inode_page.truncate_inode_pages_range.truncate_inode_pages_final.evict
      2.55 ±  3%     -14.2%       2.19 ±  2%     -17.5%       2.10 ±  1%  perf-profile.cycles-pp.__alloc_pages_nodemask.alloc_pages_current.__page_cache_alloc.pagecache_get_page.grab_cache_page_write_begin
      1.04 ±  2%     -18.9%       0.84 ±  1%     -19.6%       0.84 ±  0%  perf-profile.cycles-pp.__delete_from_page_cache.delete_from_page_cache.truncate_inode_page.truncate_inode_pages_range.truncate_inode_pages_final
      1.74 ±  2%     -19.9%       1.40 ±  3%     -19.3%       1.41 ±  1%  perf-profile.cycles-pp.delete_from_page_cache.truncate_inode_page.truncate_inode_pages_range.truncate_inode_pages_final.evict
      1.01 ±  3%     -17.9%       0.83 ±  2%     -18.2%       0.82 ±  1%  perf-profile.cycles-pp.down_write.xfs_ilock.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write
     11.21 ±  2%     -18.1%       9.18 ±  0%     -18.4%       9.14 ±  1%  perf-profile.cycles-pp.evict.iput.__dentry_kill.dput.__fput
     11.24 ±  2%     -18.1%       9.21 ±  0%     -18.4%       9.18 ±  1%  perf-profile.cycles-pp.__dentry_kill.dput.__fput.____fput.task_work_run
     11.22 ±  2%     -18.1%       9.19 ±  0%     -18.4%       9.16 ±  1%  perf-profile.cycles-pp.iput.__dentry_kill.dput.__fput.____fput
      1.79 ±  3%     -22.2%       1.39 ±  0%     -18.2%       1.46 ±  0%  perf-profile.cycles-pp.security_file_permission.rw_verify_area.vfs_write.sys_write.entry_SYSCALL_64_fastpath
     11.26 ±  2%     -18.1%       9.23 ±  0%     -18.3%       9.20 ±  1%  perf-profile.cycles-pp.dput.__fput.____fput.task_work_run.exit_to_usermode_loop
     11.31 ±  1%     -18.1%       9.27 ±  0%     -18.2%       9.25 ±  1%  perf-profile.cycles-pp.____fput.task_work_run.exit_to_usermode_loop.syscall_return_slowpath.entry_SYSCALL_64_fastpath
     11.34 ±  2%     -18.1%       9.29 ±  0%     -18.3%       9.27 ±  1%  perf-profile.cycles-pp.exit_to_usermode_loop.syscall_return_slowpath.entry_SYSCALL_64_fastpath
     11.31 ±  2%     -18.1%       9.26 ±  0%     -18.3%       9.24 ±  1%  perf-profile.cycles-pp.__fput.____fput.task_work_run.exit_to_usermode_loop.syscall_return_slowpath
     11.32 ±  1%     -18.0%       9.28 ±  0%     -18.2%       9.26 ±  1%  perf-profile.cycles-pp.task_work_run.exit_to_usermode_loop.syscall_return_slowpath.entry_SYSCALL_64_fastpath
     11.34 ±  1%     -18.1%       9.29 ±  0%     -18.2%       9.27 ±  1%  perf-profile.cycles-pp.syscall_return_slowpath.entry_SYSCALL_64_fastpath
      2.06 ±  3%     -22.5%       1.60 ±  2%     -18.1%       1.69 ±  0%  perf-profile.cycles-pp.rw_verify_area.vfs_write.sys_write.entry_SYSCALL_64_fastpath
      9.87 ±  2%     -17.5%       8.15 ±  0%     -17.6%       8.14 ±  1%  perf-profile.cycles-pp.truncate_inode_pages_range.truncate_inode_pages_final.evict.iput.__dentry_kill
      9.89 ±  2%     -17.4%       8.17 ±  0%     -17.5%       8.16 ±  1%  perf-profile.cycles-pp.truncate_inode_pages_final.evict.iput.__dentry_kill.dput
      1.00 ±  1%     -18.0%       0.82 ±  1%     -14.3%       0.86 ±  3%  perf-profile.cycles-pp.__radix_tree_lookup.radix_tree_lookup_slot.find_get_entry.pagecache_get_page.grab_cache_page_write_begin
     51.83 ±  1%     +14.3%      59.25 ±  0%     +13.8%      58.97 ±  0%  perf-profile.cycles-pp.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write.sys_write
      1.38 ±  2%     -13.3%       1.19 ±  1%      -9.9%       1.24 ±  2%  perf-profile.cycles-pp.__set_page_dirty.mark_buffer_dirty.__block_commit_write.isra.24.block_write_end.generic_write_end
     53.16 ±  1%     +13.6%      60.40 ±  0%     +13.0%      60.10 ±  0%  perf-profile.cycles-pp.xfs_file_write_iter.__vfs_write.vfs_write.sys_write.entry_SYSCALL_64_fastpath
     54.10 ±  1%     +13.1%      61.20 ±  0%     +12.5%      60.86 ±  0%  perf-profile.cycles-pp.__vfs_write.vfs_write.sys_write.entry_SYSCALL_64_fastpath
      1.32 ±  4%     -21.4%       1.04 ±  0%     -14.9%       1.13 ±  1%  perf-profile.cycles-pp.selinux_file_permission.security_file_permission.rw_verify_area.vfs_write.sys_write
     19.79 ±  5%      -9.9%      17.84 ±  0%      -7.5%      18.31 ±  3%  perf-profile.cycles-pp.start_secondary
     19.75 ±  5%      -9.8%      17.81 ±  0%      -7.4%      18.28 ±  3%  perf-profile.cycles-pp.cpu_startup_entry.start_secondary
      2.50 ±  3%     -11.5%       2.21 ±  0%     -13.1%       2.17 ±  0%  perf-profile.cycles-pp.__pagevec_release.truncate_inode_pages_range.truncate_inode_pages_final.evict.iput
      2.39 ±  3%     -11.2%       2.12 ±  0%     -13.0%       2.08 ±  0%  perf-profile.cycles-pp.release_pages.__pagevec_release.truncate_inode_pages_range.truncate_inode_pages_final.evict
     59.63 ±  1%     +10.2%      65.72 ±  0%      +9.7%      65.43 ±  0%  perf-profile.cycles-pp.vfs_write.sys_write.entry_SYSCALL_64_fastpath
      0.00 ±  0%      +Inf%       1.91 ±  1%      +Inf%       1.83 ±  1%  perf-profile.func.cycles-pp.mark_page_accessed
      0.00 ±  0%      +Inf%       1.12 ±  1%      +Inf%       1.12 ±  0%  perf-profile.func.cycles-pp.iomap_write_actor
      0.00 ±  0%      +Inf%       1.10 ±  3%      +Inf%       1.10 ±  2%  perf-profile.func.cycles-pp.xfs_iomap_eof_want_preallocate.constprop.8
      1.30 ±  2%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.func.cycles-pp.generic_perform_write
      1.08 ±  2%    -100.0%       0.00 ±  0%    -100.0%       0.00 ±  0%  perf-profile.func.cycles-pp.__xfs_get_blocks
      0.37 ±  2%    +243.6%       1.26 ±  2%    +236.4%       1.23 ±  0%  perf-profile.func.cycles-pp.xfs_bmap_search_extents
      0.70 ±  5%    +219.5%       2.24 ±  0%    +213.8%       2.20 ±  2%  perf-profile.func.cycles-pp.xfs_bmapi_read
      0.41 ±  1%    +198.4%       1.22 ±  2%    +190.2%       1.19 ±  1%  perf-profile.func.cycles-pp.xfs_bmap_search_multi_extents
      0.64 ±  1%    +182.8%       1.81 ±  4%    +181.3%       1.80 ±  0%  perf-profile.func.cycles-pp.xfs_iext_bno_to_ext
      0.46 ±  4%    +161.6%       1.20 ±  1%    +163.0%       1.21 ±  1%  perf-profile.func.cycles-pp.xfs_iomap_write_delay
      1.31 ±  2%     -46.7%       0.70 ±  0%     -46.9%       0.69 ±  2%  perf-profile.func.cycles-pp.generic_write_end
      2.49 ±  0%     -34.5%       1.63 ±  1%     -36.0%       1.59 ±  1%  perf-profile.func.cycles-pp.__block_commit_write.isra.24
      1.50 ±  1%     -20.9%       1.19 ±  1%     -21.3%       1.18 ±  1%  perf-profile.func.cycles-pp.mark_buffer_dirty
      3.24 ±  0%     -19.8%       2.60 ±  0%     -20.0%       2.59 ±  0%  perf-profile.func.cycles-pp.memset_erms
      3.96 ±  2%     -18.4%       3.23 ±  0%     -20.3%       3.16 ±  1%  perf-profile.func.cycles-pp.copy_user_enhanced_fast_string
      1.79 ±  4%     -16.8%       1.49 ±  1%     -19.6%       1.44 ±  0%  perf-profile.func.cycles-pp.__mark_inode_dirty
      1.41 ±  3%     -20.6%       1.12 ±  3%     -21.3%       1.11 ±  3%  perf-profile.func.cycles-pp.entry_SYSCALL_64_fastpath
      1.16 ±  0%     -18.1%       0.95 ±  1%     -18.1%       0.95 ±  3%  perf-profile.func.cycles-pp._raw_spin_lock
      1.16 ±  1%     -21.6%       0.91 ±  1%     -20.1%       0.93 ±  2%  perf-profile.func.cycles-pp.vfs_write
      1.75 ±  2%     -18.9%       1.42 ±  1%     -17.7%       1.44 ±  2%  perf-profile.func.cycles-pp.unlock_page
      1.32 ±  0%     -16.4%       1.10 ±  1%     -14.1%       1.13 ±  3%  perf-profile.func.cycles-pp.__radix_tree_lookup
      1.51 ±  2%     +15.4%       1.75 ±  1%     +15.9%       1.75 ±  0%  perf-profile.func.cycles-pp.__block_write_begin_int
      1.02 ±  4%      -7.5%       0.94 ±  2%     -12.4%       0.89 ±  2%  perf-profile.func.cycles-pp.pagecache_get_page
      1.05 ±  2%     -15.6%       0.88 ±  3%     -15.6%       0.88 ±  5%  perf-profile.func.cycles-pp.xfs_file_write_iter

raw perf profile data:

  "perf-profile.func.cycles-pp.intel_idle": 17.0,
  "perf-profile.func.cycles-pp.copy_user_enhanced_fast_string": 3.15,
  "perf-profile.func.cycles-pp.memset_erms": 2.59,
  "perf-profile.func.cycles-pp.xfs_bmapi_read": 2.25,
  "perf-profile.func.cycles-pp.___might_sleep": 2.13,
  "perf-profile.func.cycles-pp.mark_page_accessed": 1.8,
  "perf-profile.func.cycles-pp.xfs_iext_bno_to_ext": 1.8,
  "perf-profile.func.cycles-pp.__block_write_begin_int": 1.74,
  "perf-profile.func.cycles-pp.__block_commit_write.isra.24": 1.62,
  "perf-profile.func.cycles-pp.up_write": 1.62,
  "perf-profile.func.cycles-pp.down_write": 1.48,
  "perf-profile.func.cycles-pp.unlock_page": 1.48,
  "perf-profile.func.cycles-pp.__mark_inode_dirty": 1.43,
  "perf-profile.func.cycles-pp.xfs_iomap_write_delay": 1.23,
  "perf-profile.func.cycles-pp.xfs_bmap_search_extents": 1.23,
  "perf-profile.func.cycles-pp.__radix_tree_lookup": 1.19,
  "perf-profile.func.cycles-pp.xfs_bmap_search_multi_extents": 1.18,
  "perf-profile.func.cycles-pp.__might_sleep": 1.17,
  "perf-profile.func.cycles-pp.mark_buffer_dirty": 1.16,
  "perf-profile.func.cycles-pp.entry_SYSCALL_64_fastpath": 1.12,
  "perf-profile.func.cycles-pp.iomap_write_actor": 1.11,
  "perf-profile.func.cycles-pp.xfs_iomap_eof_want_preallocate.constprop.8": 1.07,
  "perf-profile.func.cycles-pp._raw_spin_lock": 1.0,
  "perf-profile.func.cycles-pp.vfs_write": 0.94,
  "perf-profile.func.cycles-pp.pagecache_get_page": 0.92,
  "perf-profile.func.cycles-pp.xfs_bmapi_delay": 0.92,
  "perf-profile.func.cycles-pp.xfs_file_iomap_begin": 0.91,
  "perf-profile.func.cycles-pp.xfs_file_write_iter": 0.9,
  "perf-profile.func.cycles-pp.workingset_activation": 0.82,
  "perf-profile.func.cycles-pp.iomap_apply": 0.77,
  "perf-profile.func.cycles-pp.xfs_bmapi_trim_map.isra.14": 0.75,
  "perf-profile.func.cycles-pp.xfs_file_buffered_aio_write": 0.74,
  "perf-profile.func.cycles-pp.mem_cgroup_zone_lruvec": 0.73,
  "perf-profile.func.cycles-pp.native_queued_spin_lock_slowpath": 0.73,
  "perf-profile.func.cycles-pp.get_page_from_freelist": 0.72,
  "perf-profile.func.cycles-pp.generic_write_end": 0.71,
  "perf-profile.func.cycles-pp.__vfs_write": 0.66,
  "perf-profile.func.cycles-pp.rwsem_spin_on_owner": 0.66,
  "perf-profile.func.cycles-pp.iov_iter_copy_from_user_atomic": 0.66,
  "perf-profile.func.cycles-pp.release_pages": 0.65,
  "perf-profile.func.cycles-pp.find_get_entry": 0.65,

Thanks,
Xiaolong

[toc] | [prev] | [next] | [standalone]


#1461056 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-12 12:10 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5hzA-if-7@gated-at.bofh.it>
In reply to#1461019
On Fri, Aug 12, 2016 at 04:51:24PM +0800, Ye Xiaolong wrote:
> On 08/12, Ye Xiaolong wrote:
> >On 08/12, Dave Chinner wrote:
> 
> [snip]
> 
> >>lkp-folk: the patch I've just tested it attached below - can you
> >>feed that through your test and see if it fixes the regression?
> >>
> >
> >Hi, Dave
> >
> >I am verifying your fix patch in lkp environment now, will send the
> >result once I get it.
> >
> 
> Here is the test result.

Which says "no change". Oh well, back to the drawing board...

Can you send me the aim7 config file and command line you are using
for the test?

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1461073 — Re: [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromFengguang Wu <fengguang.wu@intel.com>
Date2016-08-12 12:50 +0200
SubjectRe: [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5ici-vl-19@gated-at.bofh.it>
In reply to#1461056
Hi Dave,

On Fri, Aug 12, 2016 at 08:02:08PM +1000, Dave Chinner wrote:
>On Fri, Aug 12, 2016 at 04:51:24PM +0800, Ye Xiaolong wrote:
>> On 08/12, Ye Xiaolong wrote:
>> >On 08/12, Dave Chinner wrote:
>>
>> [snip]
>>
>> >>lkp-folk: the patch I've just tested it attached below - can you
>> >>feed that through your test and see if it fixes the regression?
>> >>
>> >
>> >Hi, Dave
>> >
>> >I am verifying your fix patch in lkp environment now, will send the
>> >result once I get it.
>> >
>>
>> Here is the test result.
>
>Which says "no change". Oh well, back to the drawing board...
>
>Can you send me the aim7 config file and command line you are using
>for the test?

The test scripts can be found here:

https://git.kernel.org/cgit/linux/kernel/git/wfg/lkp-tests.git/tree/setup/disk
https://git.kernel.org/cgit/linux/kernel/git/wfg/lkp-tests.git/tree/setup/fs
https://git.kernel.org/cgit/linux/kernel/git/wfg/lkp-tests.git/tree/tests/aim7

Here are the real commands they executed in this test:

modprobe -r brd
modprobe brd rd_nr=1 rd_size=50331648 part_show=1

dmsetup remove_all
wipefs -a --force /dev/ram0

mkfs -t xfs /dev/ram0
mkdir -p /fs/ram0
mount -t xfs -o nobarrier,inode64 /dev/ram0 /fs/ram0

for file in /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
do
        echo performance > $file
done

echo "500 32000 128 512" > /proc/sys/kernel/sem

cat > workfile <<EOF
FILESIZE: 1M
POOLSIZE: 10M
10 disk_wrt
EOF

echo "/fs/ram0" > config

(
        echo ivb44
        echo disk_wrt

        echo 1
        echo 3000
        echo 2
        echo 3000
        echo 1
) | ./multitask -t

Thanks,
Fengguang

[toc] | [prev] | [next] | [standalone]


#1461546 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromChristoph Hellwig <hch@lst.de>
Date2016-08-13 02:40 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5v9v-EX-13@gated-at.bofh.it>
In reply to#1461056
On Fri, Aug 12, 2016 at 08:02:08PM +1000, Dave Chinner wrote:
> Which says "no change". Oh well, back to the drawing board...

I don't see how it would change thing much - for all relevant calculations
we convert to block units first anyway.

But the whole xfs_iomap_write_delay is a giant mess anyway.  For a usual
call we do at least four lookups in the extent btree, which seems rather
costly.  Especially given that the low-level xfs_bmap_search_extents
interface would give us all required information in one single call.

Below is a patch I hacked up this morning to do just that.  It passes
xfstests, but I've not done any real benchmarking with it.  If the
reduced lookup overhead in it doesn't help enough we'll need to some
sort of look aside cache for the information, but I hope that we
can avoid that.  And yes, it's a rather large patch - but the old
path was so entangled that I couldn't come up with something lighter.

diff --git a/fs/xfs/libxfs/xfs_bmap.c b/fs/xfs/libxfs/xfs_bmap.c
index b060bca..614803b 100644
--- a/fs/xfs/libxfs/xfs_bmap.c
+++ b/fs/xfs/libxfs/xfs_bmap.c
@@ -1388,7 +1388,7 @@ xfs_bmap_search_multi_extents(
  * Else, *lastxp will be set to the index of the found
  * entry; *gotp will contain the entry.
  */
-STATIC xfs_bmbt_rec_host_t *                 /* pointer to found extent entry */
+xfs_bmbt_rec_host_t *                 /* pointer to found extent entry */
 xfs_bmap_search_extents(
 	xfs_inode_t     *ip,            /* incore inode pointer */
 	xfs_fileoff_t   bno,            /* block number searched for */
@@ -4074,7 +4074,7 @@ xfs_bmapi_read(
 	return 0;
 }
 
-STATIC int
+int
 xfs_bmapi_reserve_delalloc(
 	struct xfs_inode	*ip,
 	xfs_fileoff_t		aoff,
@@ -4170,91 +4170,6 @@ out_unreserve_quota:
 	return error;
 }
 
-/*
- * Map file blocks to filesystem blocks, adding delayed allocations as needed.
- */
-int
-xfs_bmapi_delay(
-	struct xfs_inode	*ip,	/* incore inode */
-	xfs_fileoff_t		bno,	/* starting file offs. mapped */
-	xfs_filblks_t		len,	/* length to map in file */
-	struct xfs_bmbt_irec	*mval,	/* output: map values */
-	int			*nmap,	/* i/o: mval size/count */
-	int			flags)	/* XFS_BMAPI_... */
-{
-	struct xfs_mount	*mp = ip->i_mount;
-	struct xfs_ifork	*ifp = XFS_IFORK_PTR(ip, XFS_DATA_FORK);
-	struct xfs_bmbt_irec	got;	/* current file extent record */
-	struct xfs_bmbt_irec	prev;	/* previous file extent record */
-	xfs_fileoff_t		obno;	/* old block number (offset) */
-	xfs_fileoff_t		end;	/* end of mapped file region */
-	xfs_extnum_t		lastx;	/* last useful extent number */
-	int			eof;	/* we've hit the end of extents */
-	int			n = 0;	/* current extent index */
-	int			error = 0;
-
-	ASSERT(*nmap >= 1);
-	ASSERT(*nmap <= XFS_BMAP_MAX_NMAP);
-	ASSERT(!(flags & ~XFS_BMAPI_ENTIRE));
-	ASSERT(xfs_isilocked(ip, XFS_ILOCK_EXCL));
-
-	if (unlikely(XFS_TEST_ERROR(
-	    (XFS_IFORK_FORMAT(ip, XFS_DATA_FORK) != XFS_DINODE_FMT_EXTENTS &&
-	     XFS_IFORK_FORMAT(ip, XFS_DATA_FORK) != XFS_DINODE_FMT_BTREE),
-	     mp, XFS_ERRTAG_BMAPIFORMAT, XFS_RANDOM_BMAPIFORMAT))) {
-		XFS_ERROR_REPORT("xfs_bmapi_delay", XFS_ERRLEVEL_LOW, mp);
-		return -EFSCORRUPTED;
-	}
-
-	if (XFS_FORCED_SHUTDOWN(mp))
-		return -EIO;
-
-	XFS_STATS_INC(mp, xs_blk_mapw);
-
-	if (!(ifp->if_flags & XFS_IFEXTENTS)) {
-		error = xfs_iread_extents(NULL, ip, XFS_DATA_FORK);
-		if (error)
-			return error;
-	}
-
-	xfs_bmap_search_extents(ip, bno, XFS_DATA_FORK, &eof, &lastx, &got, &prev);
-	end = bno + len;
-	obno = bno;
-
-	while (bno < end && n < *nmap) {
-		if (eof || got.br_startoff > bno) {
-			error = xfs_bmapi_reserve_delalloc(ip, bno, len, &got,
-							   &prev, &lastx, eof);
-			if (error) {
-				if (n == 0) {
-					*nmap = 0;
-					return error;
-				}
-				break;
-			}
-		}
-
-		/* set up the extent map to return. */
-		xfs_bmapi_trim_map(mval, &got, &bno, len, obno, end, n, flags);
-		xfs_bmapi_update_map(&mval, &bno, &len, obno, end, &n, flags);
-
-		/* If we're done, stop now. */
-		if (bno >= end || n >= *nmap)
-			break;
-
-		/* Else go on to the next record. */
-		prev = got;
-		if (++lastx < ifp->if_bytes / sizeof(xfs_bmbt_rec_t))
-			xfs_bmbt_get_all(xfs_iext_get_ext(ifp, lastx), &got);
-		else
-			eof = 1;
-	}
-
-	*nmap = n;
-	return 0;
-}
-
-
 static int
 xfs_bmapi_allocate(
 	struct xfs_bmalloca	*bma)
diff --git a/fs/xfs/libxfs/xfs_bmap.h b/fs/xfs/libxfs/xfs_bmap.h
index 254034f..d660069 100644
--- a/fs/xfs/libxfs/xfs_bmap.h
+++ b/fs/xfs/libxfs/xfs_bmap.h
@@ -181,9 +181,6 @@ int	xfs_bmap_read_extents(struct xfs_trans *tp, struct xfs_inode *ip,
 int	xfs_bmapi_read(struct xfs_inode *ip, xfs_fileoff_t bno,
 		xfs_filblks_t len, struct xfs_bmbt_irec *mval,
 		int *nmap, int flags);
-int	xfs_bmapi_delay(struct xfs_inode *ip, xfs_fileoff_t bno,
-		xfs_filblks_t len, struct xfs_bmbt_irec *mval,
-		int *nmap, int flags);
 int	xfs_bmapi_write(struct xfs_trans *tp, struct xfs_inode *ip,
 		xfs_fileoff_t bno, xfs_filblks_t len, int flags,
 		xfs_fsblock_t *firstblock, xfs_extlen_t total,
@@ -202,5 +199,12 @@ int	xfs_bmap_shift_extents(struct xfs_trans *tp, struct xfs_inode *ip,
 		struct xfs_defer_ops *dfops, enum shift_direction direction,
 		int num_exts);
 int	xfs_bmap_split_extent(struct xfs_inode *ip, xfs_fileoff_t split_offset);
+struct xfs_bmbt_rec_host *
+	xfs_bmap_search_extents(struct xfs_inode *ip, xfs_fileoff_t bno,
+		int fork, int *eofp, xfs_extnum_t *lastxp,
+		struct xfs_bmbt_irec *gotp, struct xfs_bmbt_irec *prevp);
+int	xfs_bmapi_reserve_delalloc(struct xfs_inode *ip, xfs_fileoff_t aoff,
+		xfs_filblks_t len, struct xfs_bmbt_irec *got,
+		struct xfs_bmbt_irec *prev, xfs_extnum_t *lastx, int eof);
 
 #endif	/* __XFS_BMAP_H__ */
diff --git a/fs/xfs/xfs_iomap.c b/fs/xfs/xfs_iomap.c
index 2114d53..b40e31b 100644
--- a/fs/xfs/xfs_iomap.c
+++ b/fs/xfs/xfs_iomap.c
@@ -42,11 +42,62 @@
 
 #define XFS_WRITEIO_ALIGN(mp,off)	(((off) >> mp->m_writeio_log) \
 						<< mp->m_writeio_log)
-#define XFS_WRITE_IMAPS		XFS_BMAP_MAX_NMAP
+
+void
+xfs_bmbt_to_iomap(
+	struct xfs_inode	*ip,
+	struct iomap		*iomap,
+	struct xfs_bmbt_irec	*imap)
+{
+	struct xfs_mount	*mp = ip->i_mount;
+
+	if (imap->br_startblock == HOLESTARTBLOCK) {
+		iomap->blkno = IOMAP_NULL_BLOCK;
+		iomap->type = IOMAP_HOLE;
+	} else if (imap->br_startblock == DELAYSTARTBLOCK) {
+		iomap->blkno = IOMAP_NULL_BLOCK;
+		iomap->type = IOMAP_DELALLOC;
+	} else {
+		iomap->blkno = xfs_fsb_to_db(ip, imap->br_startblock);
+		if (imap->br_state == XFS_EXT_UNWRITTEN)
+			iomap->type = IOMAP_UNWRITTEN;
+		else
+			iomap->type = IOMAP_MAPPED;
+	}
+	iomap->offset = XFS_FSB_TO_B(mp, imap->br_startoff);
+	iomap->length = XFS_FSB_TO_B(mp, imap->br_blockcount);
+	iomap->bdev = xfs_find_bdev_for_inode(VFS_I(ip));
+}
+
+static xfs_extlen_t
+xfs_align_eof(
+	struct xfs_inode	*ip)
+{
+	struct xfs_mount	*mp = ip->i_mount;
+	xfs_extlen_t		align = 0;
+
+	ASSERT(!XFS_IS_REALTIME_INODE(ip));
+
+	/*
+	 * Round up the allocation request to a stripe unit (m_dalign)
+	 * boundary if the file size is >= stripe unit size, and we are
+	 * allocating past the allocation eof.
+	 *
+	 * If mounted with the "-o swalloc" option the alignment is
+	 * increased from the strip unit size to the stripe width.
+	 */
+	if (mp->m_swidth && (mp->m_flags & XFS_MOUNT_SWALLOC))
+		align = mp->m_swidth;
+	else if (mp->m_dalign)
+		align = mp->m_dalign;
+
+	if (align && XFS_ISIZE(ip) < XFS_FSB_TO_B(mp, align))
+		align = 0;
+	return align;
+}
 
 STATIC int
 xfs_iomap_eof_align_last_fsb(
-	xfs_mount_t	*mp,
 	xfs_inode_t	*ip,
 	xfs_extlen_t	extsize,
 	xfs_fileoff_t	*last_fsb)
@@ -54,23 +105,8 @@ xfs_iomap_eof_align_last_fsb(
 	xfs_extlen_t	align = 0;
 	int		eof, error;
 
-	if (!XFS_IS_REALTIME_INODE(ip)) {
-		/*
-		 * Round up the allocation request to a stripe unit
-		 * (m_dalign) boundary if the file size is >= stripe unit
-		 * size, and we are allocating past the allocation eof.
-		 *
-		 * If mounted with the "-o swalloc" option the alignment is
-		 * increased from the strip unit size to the stripe width.
-		 */
-		if (mp->m_swidth && (mp->m_flags & XFS_MOUNT_SWALLOC))
-			align = mp->m_swidth;
-		else if (mp->m_dalign)
-			align = mp->m_dalign;
-
-		if (align && XFS_ISIZE(ip) < XFS_FSB_TO_B(mp, align))
-			align = 0;
-	}
+	if (!XFS_IS_REALTIME_INODE(ip))
+		align = xfs_align_eof(ip);
 
 	/*
 	 * Always round up the allocation request to an extent boundary
@@ -154,7 +190,7 @@ xfs_iomap_write_direct(
 		 */
 		ASSERT(XFS_IFORK_PTR(ip, XFS_DATA_FORK)->if_flags &
 								XFS_IFEXTENTS);
-		error = xfs_iomap_eof_align_last_fsb(mp, ip, extsz, &last_fsb);
+		error = xfs_iomap_eof_align_last_fsb(ip, extsz, &last_fsb);
 		if (error)
 			goto out_unlock;
 	} else {
@@ -274,130 +310,6 @@ out_trans_cancel:
 	goto out_unlock;
 }
 
-/*
- * If the caller is doing a write at the end of the file, then extend the
- * allocation out to the file system's write iosize.  We clean up any extra
- * space left over when the file is closed in xfs_inactive().
- *
- * If we find we already have delalloc preallocation beyond EOF, don't do more
- * preallocation as it it not needed.
- */
-STATIC int
-xfs_iomap_eof_want_preallocate(
-	xfs_mount_t	*mp,
-	xfs_inode_t	*ip,
-	xfs_off_t	offset,
-	size_t		count,
-	xfs_bmbt_irec_t *imap,
-	int		nimaps,
-	int		*prealloc)
-{
-	xfs_fileoff_t   start_fsb;
-	xfs_filblks_t   count_fsb;
-	int		n, error, imaps;
-	int		found_delalloc = 0;
-
-	*prealloc = 0;
-	if (offset + count <= XFS_ISIZE(ip))
-		return 0;
-
-	/*
-	 * If the file is smaller than the minimum prealloc and we are using
-	 * dynamic preallocation, don't do any preallocation at all as it is
-	 * likely this is the only write to the file that is going to be done.
-	 */
-	if (!(mp->m_flags & XFS_MOUNT_DFLT_IOSIZE) &&
-	    XFS_ISIZE(ip) < XFS_FSB_TO_B(mp, mp->m_writeio_blocks))
-		return 0;
-
-	/*
-	 * If there are any real blocks past eof, then don't
-	 * do any speculative allocation.
-	 */
-	start_fsb = XFS_B_TO_FSBT(mp, ((xfs_ufsize_t)(offset + count - 1)));
-	count_fsb = XFS_B_TO_FSB(mp, mp->m_super->s_maxbytes);
-	while (count_fsb > 0) {
-		imaps = nimaps;
-		error = xfs_bmapi_read(ip, start_fsb, count_fsb, imap, &imaps,
-				       0);
-		if (error)
-			return error;
-		for (n = 0; n < imaps; n++) {
-			if ((imap[n].br_startblock != HOLESTARTBLOCK) &&
-			    (imap[n].br_startblock != DELAYSTARTBLOCK))
-				return 0;
-			start_fsb += imap[n].br_blockcount;
-			count_fsb -= imap[n].br_blockcount;
-
-			if (imap[n].br_startblock == DELAYSTARTBLOCK)
-				found_delalloc = 1;
-		}
-	}
-	if (!found_delalloc)
-		*prealloc = 1;
-	return 0;
-}
-
-/*
- * Determine the initial size of the preallocation. We are beyond the current
- * EOF here, but we need to take into account whether this is a sparse write or
- * an extending write when determining the preallocation size.  Hence we need to
- * look up the extent that ends at the current write offset and use the result
- * to determine the preallocation size.
- *
- * If the extent is a hole, then preallocation is essentially disabled.
- * Otherwise we take the size of the preceeding data extent as the basis for the
- * preallocation size. If the size of the extent is greater than half the
- * maximum extent length, then use the current offset as the basis. This ensures
- * that for large files the preallocation size always extends to MAXEXTLEN
- * rather than falling short due to things like stripe unit/width alignment of
- * real extents.
- */
-STATIC xfs_fsblock_t
-xfs_iomap_eof_prealloc_initial_size(
-	struct xfs_mount	*mp,
-	struct xfs_inode	*ip,
-	xfs_off_t		offset,
-	xfs_bmbt_irec_t		*imap,
-	int			nimaps)
-{
-	xfs_fileoff_t   start_fsb;
-	int		imaps = 1;
-	int		error;
-
-	ASSERT(nimaps >= imaps);
-
-	/* if we are using a specific prealloc size, return now */
-	if (mp->m_flags & XFS_MOUNT_DFLT_IOSIZE)
-		return 0;
-
-	/* If the file is small, then use the minimum prealloc */
-	if (XFS_ISIZE(ip) < XFS_FSB_TO_B(mp, mp->m_dalign))
-		return 0;
-
-	/*
-	 * As we write multiple pages, the offset will always align to the
-	 * start of a page and hence point to a hole at EOF. i.e. if the size is
-	 * 4096 bytes, we only have one block at FSB 0, but XFS_B_TO_FSB(4096)
-	 * will return FSB 1. Hence if there are blocks in the file, we want to
-	 * point to the block prior to the EOF block and not the hole that maps
-	 * directly at @offset.
-	 */
-	start_fsb = XFS_B_TO_FSB(mp, offset);
-	if (start_fsb)
-		start_fsb--;
-	error = xfs_bmapi_read(ip, start_fsb, 1, imap, &imaps, XFS_BMAPI_ENTIRE);
-	if (error)
-		return 0;
-
-	ASSERT(imaps == 1);
-	if (imap[0].br_startblock == HOLESTARTBLOCK)
-		return 0;
-	if (imap[0].br_blockcount <= (MAXEXTLEN >> 1))
-		return imap[0].br_blockcount << 1;
-	return XFS_B_TO_FSB(mp, offset);
-}
-
 STATIC bool
 xfs_quota_need_throttle(
 	struct xfs_inode *ip,
@@ -466,20 +378,37 @@ xfs_quota_calc_throttle(
  */
 STATIC xfs_fsblock_t
 xfs_iomap_prealloc_size(
-	struct xfs_mount	*mp,
 	struct xfs_inode	*ip,
 	xfs_off_t		offset,
-	struct xfs_bmbt_irec	*imap,
-	int			nimaps)
+	struct xfs_bmbt_irec	*prev)
 {
-	xfs_fsblock_t		alloc_blocks = 0;
+	struct xfs_mount	*mp = ip->i_mount;
 	int			shift = 0;
 	int64_t			freesp;
 	xfs_fsblock_t		qblocks;
 	int			qshift = 0;
+	xfs_fsblock_t		alloc_blocks = 0;
 
-	alloc_blocks = xfs_iomap_eof_prealloc_initial_size(mp, ip, offset,
-							   imap, nimaps);
+	/*
+	 * Determine the initial size of the preallocation. We are beyond the
+	 * current EOF here, but we need to take into account whether this is
+	 * a sparse write or an extending write when determining the
+	 * preallocation size.  Hence we need to look up the extent that ends
+	 * at the current write offset and use the result to determine the
+	 * preallocation size.
+	 *
+	 * If the extent is a hole, then preallocation is essentially disabled.
+	 * Otherwise we take the size of the preceding data extent as the basis
+	 * for the preallocation size. If the size of the extent is greater than
+	 * half the maximum extent length, then use the current offset as the
+	 * basis. This ensures that for large files the preallocation size
+	 * always extends to MAXEXTLEN rather than falling short due to things
+	 * like stripe unit/width alignment of real extents.
+	 */
+	if (prev->br_blockcount <= (MAXEXTLEN >> 1))
+		alloc_blocks = prev->br_blockcount << 1;
+	else
+		alloc_blocks = XFS_B_TO_FSB(mp, offset);
 	if (!alloc_blocks)
 		goto check_writeio;
 	qblocks = alloc_blocks;
@@ -550,120 +479,166 @@ xfs_iomap_prealloc_size(
 	 */
 	while (alloc_blocks && alloc_blocks >= freesp)
 		alloc_blocks >>= 4;
-
 check_writeio:
 	if (alloc_blocks < mp->m_writeio_blocks)
 		alloc_blocks = mp->m_writeio_blocks;
-
 	trace_xfs_iomap_prealloc_size(ip, alloc_blocks, shift,
 				      mp->m_writeio_blocks);
-
 	return alloc_blocks;
 }
 
-int
-xfs_iomap_write_delay(
-	xfs_inode_t	*ip,
-	xfs_off_t	offset,
-	size_t		count,
-	xfs_bmbt_irec_t *ret_imap)
+static int
+xfs_file_iomap_begin_delay(
+	struct inode		*inode,
+	loff_t			offset,
+	loff_t			count,
+	unsigned		flags,
+	struct iomap		*iomap)
 {
-	xfs_mount_t	*mp = ip->i_mount;
-	xfs_fileoff_t	offset_fsb;
-	xfs_fileoff_t	last_fsb;
-	xfs_off_t	aligned_offset;
-	xfs_fileoff_t	ioalign;
-	xfs_extlen_t	extsz;
-	int		nimaps;
-	xfs_bmbt_irec_t imap[XFS_WRITE_IMAPS];
-	int		prealloc;
-	int		error;
-
-	ASSERT(xfs_isilocked(ip, XFS_ILOCK_EXCL));
-
-	/*
-	 * Make sure that the dquots are there. This doesn't hold
-	 * the ilock across a disk read.
-	 */
-	error = xfs_qm_dqattach_locked(ip, 0);
-	if (error)
-		return error;
-
-	extsz = xfs_get_extsz_hint(ip);
-	offset_fsb = XFS_B_TO_FSBT(mp, offset);
+	struct xfs_inode	*ip = XFS_I(inode);
+	struct xfs_mount	*mp = ip->i_mount;
+	struct xfs_ifork	*ifp = XFS_IFORK_PTR(ip, XFS_DATA_FORK);
+	xfs_fileoff_t		offset_fsb = XFS_B_TO_FSBT(mp, offset);
+	xfs_fileoff_t		maxbytes_fsb =
+		XFS_B_TO_FSB(mp, mp->m_super->s_maxbytes);
+	xfs_fileoff_t		end_fsb, orig_end_fsb;
+	int			error = 0, eof = 0;
+	struct xfs_bmbt_irec	got;
+	struct xfs_bmbt_irec	prev;
+	xfs_extnum_t		idx;
+
+	ASSERT(!XFS_IS_REALTIME_INODE(ip));
+	ASSERT(!xfs_get_extsz_hint(ip));
 
-	error = xfs_iomap_eof_want_preallocate(mp, ip, offset, count,
-				imap, XFS_WRITE_IMAPS, &prealloc);
-	if (error)
-		return error;
+	xfs_ilock(ip, XFS_ILOCK_EXCL);
 
-retry:
-	if (prealloc) {
-		xfs_fsblock_t	alloc_blocks;
+	if (unlikely(XFS_TEST_ERROR(
+	    (XFS_IFORK_FORMAT(ip, XFS_DATA_FORK) != XFS_DINODE_FMT_EXTENTS &&
+	     XFS_IFORK_FORMAT(ip, XFS_DATA_FORK) != XFS_DINODE_FMT_BTREE),
+	     mp, XFS_ERRTAG_BMAPIFORMAT, XFS_RANDOM_BMAPIFORMAT))) {
+		XFS_ERROR_REPORT(__func__, XFS_ERRLEVEL_LOW, mp);
+		error = -EFSCORRUPTED;
+		goto out_unlock;
+	}
 
-		alloc_blocks = xfs_iomap_prealloc_size(mp, ip, offset, imap,
-						       XFS_WRITE_IMAPS);
+	XFS_STATS_INC(mp, xs_blk_mapw);
 
-		aligned_offset = XFS_WRITEIO_ALIGN(mp, (offset + count - 1));
-		ioalign = XFS_B_TO_FSBT(mp, aligned_offset);
-		last_fsb = ioalign + alloc_blocks;
-	} else {
-		last_fsb = XFS_B_TO_FSB(mp, ((xfs_ufsize_t)(offset + count)));
+	if (!(ifp->if_flags & XFS_IFEXTENTS)) {
+		error = xfs_iread_extents(NULL, ip, XFS_DATA_FORK);
+		if (error)
+			goto out_unlock;
 	}
 
-	if (prealloc || extsz) {
-		error = xfs_iomap_eof_align_last_fsb(mp, ip, extsz, &last_fsb);
-		if (error)
-			return error;
+	xfs_bmap_search_extents(ip, offset_fsb, XFS_DATA_FORK, &eof, &idx,
+			&got, &prev);
+	if (!eof && got.br_startoff <= offset_fsb) {
+		trace_xfs_iomap_found(ip, offset, count, 0, &got);
+		goto done;
 	}
 
+	error = xfs_qm_dqattach_locked(ip, 0);
+	if (error)
+		goto out_unlock;
+
 	/*
-	 * Make sure preallocation does not create extents beyond the range we
-	 * actually support in this filesystem.
+	 * We cap the maximum length we map here to MAX_WRITEBACK_PAGES pages
+	 * to keep the chunks of work done where somewhat symmetric with the
+	 * work writeback does. This is a completely arbitrary number pulled
+	 * out of thin air as a best guess for initial testing.
+	 *
+	 * Note that the values needs to be less than 32-bits wide until
+	 * the lower level functions are updated.
 	 */
-	if (last_fsb > XFS_B_TO_FSB(mp, mp->m_super->s_maxbytes))
-		last_fsb = XFS_B_TO_FSB(mp, mp->m_super->s_maxbytes);
+	count = min_t(loff_t, count, 1024 * PAGE_SIZE);
+	end_fsb = orig_end_fsb =
+		min(XFS_B_TO_FSB(mp, offset + count), maxbytes_fsb);
 
-	ASSERT(last_fsb > offset_fsb);
+	/*
+	 * If we are doing a write at the end of the file and there are no
+	 * allocations past this one, then extend the allocation out to the
+	 * file system's write iosize.
+	 *
+	 * As an exception we don't do any preallocation at all if the file
+	 * is smaller than the minimum preallocation and we are using the
+	 * default dynamic preallocation scheme, as it is likely this is the
+	 * only write to the file that is going to be done.
+	 *
+	 * We clean up any extra space left over when the file is closed in
+	 * xfs_inactive().
+	 */
+	if (eof && offset + count > XFS_ISIZE(ip) &&
+	    ((mp->m_flags & XFS_MOUNT_DFLT_IOSIZE) ||
+	     XFS_ISIZE(ip) >= XFS_FSB_TO_B(mp, mp->m_writeio_blocks))) {
+		xfs_fsblock_t		alloc_blocks;
+		xfs_off_t		aligned_offset;
+		xfs_extlen_t		align;
 
-	nimaps = XFS_WRITE_IMAPS;
-	error = xfs_bmapi_delay(ip, offset_fsb, last_fsb - offset_fsb,
-				imap, &nimaps, XFS_BMAPI_ENTIRE);
+		/*
+		 * If an explicit allocsize is set, the file is small, or we
+		 * are writing behind a hole, then use the minimum prealloc:
+		 */
+		if ((mp->m_flags & XFS_MOUNT_DFLT_IOSIZE) ||
+		    XFS_ISIZE(ip) < XFS_FSB_TO_B(mp, mp->m_dalign) ||
+		    prev.br_startoff + prev.br_blockcount < offset_fsb)
+			alloc_blocks = mp->m_writeio_blocks;
+		else
+			alloc_blocks =
+				xfs_iomap_prealloc_size(ip, offset, &prev);
+
+		aligned_offset = XFS_WRITEIO_ALIGN(mp, offset + count - 1);
+		end_fsb = XFS_B_TO_FSBT(mp, aligned_offset) + alloc_blocks;
+
+		align = xfs_align_eof(ip);
+		if (align)
+			end_fsb = roundup_64(end_fsb, align);
+
+		end_fsb = min(end_fsb, maxbytes_fsb);
+		ASSERT(end_fsb > offset_fsb);
+	}
+
+retry:
+	error = xfs_bmapi_reserve_delalloc(ip, offset_fsb,
+			end_fsb - offset_fsb, &got,
+			&prev, &idx, eof);
 	switch (error) {
 	case 0:
+		break;
 	case -ENOSPC:
 	case -EDQUOT:
-		break;
-	default:
-		return error;
-	}
-
-	/*
-	 * If bmapi returned us nothing, we got either ENOSPC or EDQUOT. Retry
-	 * without EOF preallocation.
-	 */
-	if (nimaps == 0) {
+		/* retry without any preallocation */
 		trace_xfs_delalloc_enospc(ip, offset, count);
-		if (prealloc) {
-			prealloc = 0;
-			error = 0;
+		if (end_fsb != orig_end_fsb) {
+			end_fsb = orig_end_fsb;
 			goto retry;
 		}
-		return error ? error : -ENOSPC;
+		/*FALLTHRU*/
+	default:
+		goto out_unlock;
 	}
 
-	if (!(imap[0].br_startblock || XFS_IS_REALTIME_INODE(ip)))
-		return xfs_alert_fsblock_zero(ip, &imap[0]);
+	trace_xfs_iomap_alloc(ip, offset, count, 0, &got);
+done:
+	if (isnullstartblock(got.br_startblock))
+		got.br_startblock = DELAYSTARTBLOCK;
+
+	if (!got.br_startblock) {
+		error = xfs_alert_fsblock_zero(ip, &got);
+		if (error)
+			goto out_unlock;
+	}
+
+	xfs_bmbt_to_iomap(ip, iomap, &got);
 
 	/*
 	 * Tag the inode as speculatively preallocated so we can reclaim this
 	 * space on demand, if necessary.
 	 */
-	if (prealloc)
+	if (end_fsb != orig_end_fsb)
 		xfs_inode_set_eofblocks_tag(ip);
 
-	*ret_imap = imap[0];
-	return 0;
+out_unlock:
+	xfs_iunlock(ip, XFS_ILOCK_EXCL);
+	return error;
 }
 
 /*
@@ -943,32 +918,6 @@ error_on_bmapi_transaction:
 	return error;
 }
 
-void
-xfs_bmbt_to_iomap(
-	struct xfs_inode	*ip,
-	struct iomap		*iomap,
-	struct xfs_bmbt_irec	*imap)
-{
-	struct xfs_mount	*mp = ip->i_mount;
-
-	if (imap->br_startblock == HOLESTARTBLOCK) {
-		iomap->blkno = IOMAP_NULL_BLOCK;
-		iomap->type = IOMAP_HOLE;
-	} else if (imap->br_startblock == DELAYSTARTBLOCK) {
-		iomap->blkno = IOMAP_NULL_BLOCK;
-		iomap->type = IOMAP_DELALLOC;
-	} else {
-		iomap->blkno = xfs_fsb_to_db(ip, imap->br_startblock);
-		if (imap->br_state == XFS_EXT_UNWRITTEN)
-			iomap->type = IOMAP_UNWRITTEN;
-		else
-			iomap->type = IOMAP_MAPPED;
-	}
-	iomap->offset = XFS_FSB_TO_B(mp, imap->br_startoff);
-	iomap->length = XFS_FSB_TO_B(mp, imap->br_blockcount);
-	iomap->bdev = xfs_find_bdev_for_inode(VFS_I(ip));
-}
-
 static inline bool imap_needs_alloc(struct xfs_bmbt_irec *imap, int nimaps)
 {
 	return !nimaps ||
@@ -993,6 +942,11 @@ xfs_file_iomap_begin(
 	if (XFS_FORCED_SHUTDOWN(mp))
 		return -EIO;
 
+	if ((flags & IOMAP_WRITE) && !xfs_get_extsz_hint(ip)) {
+		return xfs_file_iomap_begin_delay(inode, offset, length, flags,
+				iomap);
+	}
+
 	xfs_ilock(ip, XFS_ILOCK_EXCL);
 
 	ASSERT(offset <= mp->m_super->s_maxbytes);
@@ -1020,19 +974,13 @@ xfs_file_iomap_begin(
 		 * the lower level functions are updated.
 		 */
 		length = min_t(loff_t, length, 1024 * PAGE_SIZE);
-		if (xfs_get_extsz_hint(ip)) {
-			/*
-			 * xfs_iomap_write_direct() expects the shared lock. It
-			 * is unlocked on return.
-			 */
-			xfs_ilock_demote(ip, XFS_ILOCK_EXCL);
-			error = xfs_iomap_write_direct(ip, offset, length, &imap,
-					nimaps);
-		} else {
-			error = xfs_iomap_write_delay(ip, offset, length, &imap);
-			xfs_iunlock(ip, XFS_ILOCK_EXCL);
-		}
-
+		/*
+		 * xfs_iomap_write_direct() expects the shared lock. It
+		 * is unlocked on return.
+		 */
+		xfs_ilock_demote(ip, XFS_ILOCK_EXCL);
+		error = xfs_iomap_write_direct(ip, offset, length, &imap,
+				nimaps);
 		if (error)
 			return error;
 
diff --git a/fs/xfs/xfs_iomap.h b/fs/xfs/xfs_iomap.h
index e066d04..1fdf68d 100644
--- a/fs/xfs/xfs_iomap.h
+++ b/fs/xfs/xfs_iomap.h
@@ -25,8 +25,6 @@ struct xfs_bmbt_irec;
 
 int xfs_iomap_write_direct(struct xfs_inode *, xfs_off_t, size_t,
 			struct xfs_bmbt_irec *, int);
-int xfs_iomap_write_delay(struct xfs_inode *, xfs_off_t, size_t,
-			struct xfs_bmbt_irec *);
 int xfs_iomap_write_allocate(struct xfs_inode *, xfs_off_t,
 			struct xfs_bmbt_irec *);
 int xfs_iomap_write_unwritten(struct xfs_inode *, xfs_off_t, xfs_off_t);

[toc] | [prev] | [next] | [standalone]


#1461692 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromChristoph Hellwig <hch@lst.de>
Date2016-08-14 10:30 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5YXU-5oj-35@gated-at.bofh.it>
In reply to#1461546
On Sat, Aug 13, 2016 at 02:30:54AM +0200, Christoph Hellwig wrote:
> Below is a patch I hacked up this morning to do just that.  It passes
> xfstests, but I've not done any real benchmarking with it.  If the
> reduced lookup overhead in it doesn't help enough we'll need to some
> sort of look aside cache for the information, but I hope that we
> can avoid that.  And yes, it's a rather large patch - but the old
> path was so entangled that I couldn't come up with something lighter.

Hi Fengguang or Xiaolong,

any chance to add this thread to a lkp run?  I've played around with
Dave's simplied xfs_io run, and while the end result for 1k block
size looks pretty similar in terms of execution time and throughput
the profiles look much better.  For 512 byte or 1 byte tests the
tests completes a lot faster too.

Here is the perf report output for a 1k block size run, the first
item directly related to the block mapping shows up is
xfs_file_iomap_begin_delay at .75%.  Although I'm a bit worried
up up_/down_read showing up so much.  While we take a ilock and
iolock a lot they should be mostly uncontended for such a single
threaded write, so the overhead seems a bit worrisome.

(FYI, the tree this was tested on also has the mark_page_accessed
and pagefault_disable fixes applied)

# To display the perf.data header info, please use --header/--header-only options.
#
# Samples: 7K of event 'cpu-clock'
# Event count (approx.): 1909250000
#
# Overhead       Command      Shared Object                                 Symbol
# ........  ............  .................  .....................................
#
    37.71%       swapper  [kernel.kallsyms]  [k] native_safe_halt                 
     9.85%  kworker/u8:5  [kernel.kallsyms]  [k] __copy_user_nocache              
     2.83%        xfs_io  [kernel.kallsyms]  [k] copy_user_generic_string         
     2.33%        xfs_io  [kernel.kallsyms]  [k] __memset                         
     2.23%        xfs_io  [kernel.kallsyms]  [k] __block_commit_write.isra.34     
     1.73%        xfs_io  [kernel.kallsyms]  [k] down_write                       
     1.64%        xfs_io  [kernel.kallsyms]  [k] up_write                         
     1.39%        xfs_io  [kernel.kallsyms]  [k] _raw_spin_unlock_irqrestore      
     1.23%        xfs_io  [kernel.kallsyms]  [k] entry_SYSCALL_64_fastpath        
     1.18%        xfs_io  [kernel.kallsyms]  [k] __mark_inode_dirty               
     1.18%        xfs_io  [kernel.kallsyms]  [k] _raw_spin_lock                   
     1.15%  kworker/u8:5  [kernel.kallsyms]  [k] _raw_spin_unlock_irqrestore      
     1.13%        xfs_io  [kernel.kallsyms]  [k] __block_write_begin_int          
     1.10%        xfs_io  [kernel.kallsyms]  [k] mark_buffer_dirty                
     1.07%        xfs_io  [kernel.kallsyms]  [k] __radix_tree_lookup              
     1.01%   kworker/0:2  [kernel.kallsyms]  [k] end_buffer_async_write           
     0.97%        xfs_io  [kernel.kallsyms]  [k] unlock_page                      
     0.92%   kworker/0:2  [kernel.kallsyms]  [k] _raw_spin_unlock_irqrestore      
     0.89%        xfs_io  [kernel.kallsyms]  [k] iov_iter_copy_from_user_atomic   
     0.84%        xfs_io  [kernel.kallsyms]  [k] generic_write_end                
     0.80%        xfs_io  [kernel.kallsyms]  [k] get_page_from_freelist           
     0.80%        xfs_io  [kernel.kallsyms]  [k] xfs_perag_put                    
     0.79%        xfs_io  [kernel.kallsyms]  [k] __add_to_page_cache_locked       
     0.75%        xfs_io  libc-2.19.so       [.] __libc_pwrite                    
     0.72%        xfs_io  [kernel.kallsyms]  [k] xfs_file_iomap_begin_delay.isra.5
     0.71%        xfs_io  [kernel.kallsyms]  [k] iomap_write_actor                
     0.67%        xfs_io  [kernel.kallsyms]  [k] pagecache_get_page               
     0.64%        xfs_io  [kernel.kallsyms]  [k] balance_dirty_pages_ratelimited  
     0.64%        xfs_io  [kernel.kallsyms]  [k] vfs_write                        
     0.63%  kworker/u8:5  [kernel.kallsyms]  [k] clear_page_dirty_for_io          
     0.62%        xfs_io  [kernel.kallsyms]  [k] xfs_file_write_iter              
     0.60%        xfs_io  [kernel.kallsyms]  [k] __vfs_write                      
     0.55%        xfs_io  [kernel.kallsyms]  [k] page_waitqueue                   
     0.54%        xfs_io  [kernel.kallsyms]  [k] xfs_perag_get                    
     0.52%        xfs_io  [kernel.kallsyms]  [k] __wake_up_bit                    
     0.52%        xfs_io  [kernel.kallsyms]  [k] radix_tree_tag_set               
     0.50%  kworker/u8:5  [kernel.kallsyms]  [k] xfs_do_writepage                 
     0.50%        xfs_io  [kernel.kallsyms]  [k] iov_iter_advance                 
     0.47%        xfs_io  [kernel.kallsyms]  [k] kmem_cache_alloc                 
     0.46%        xfs_io  [kernel.kallsyms]  [k] xfs_file_buffered_aio_write      
     0.46%        xfs_io  [kernel.kallsyms]  [k] xfs_iunlock                      
     0.45%  kworker/u8:5  [kernel.kallsyms]  [k] __wake_up_bit                    
     0.45%        xfs_io  [kernel.kallsyms]  [k] find_get_entry                   
     0.45%        xfs_io  [kernel.kallsyms]  [k] xfs_bmap_search_multi_extents    
     0.41%        xfs_io  [kernel.kallsyms]  [k] xfs_iext_bno_to_ext              
     0.39%        xfs_io  [kernel.kallsyms]  [k] iomap_apply                      
     0.39%        xfs_io  [kernel.kallsyms]  [k] xfs_file_aio_write_checks        
     0.38%        xfs_io  [kernel.kallsyms]  [k] xfs_ilock                        
     0.38%        xfs_io  [kernel.kallsyms]  [k] xfs_inode_set_eofblocks_tag      
     0.37%        xfs_io  [kernel.kallsyms]  [k] xfs_bmap_search_extents          
     0.30%        xfs_io  [kernel.kallsyms]  [k] file_update_time                 
     0.29%        xfs_io  [kernel.kallsyms]  [k] __fget_light                     
     0.27%        xfs_io  [kernel.kallsyms]  [k] rw_verify_area                   
     0.26%  kworker/u8:5  [kernel.kallsyms]  [k] unlock_page                      
     0.26%        xfs_io  xfs_io             [.] pwrite_f                         
     0.25%        xfs_io  [kernel.kallsyms]  [k] iomap_file_buffered_write        
     0.25%        xfs_io  [kernel.kallsyms]  [k] node_dirty_ok                    
     0.25%        xfs_io  [kernel.kallsyms]  [k] xfs_bmbt_to_iomap                
     0.24%   kworker/0:2  [kernel.kallsyms]  [k] xfs_destroy_ioend                
     0.24%        xfs_io  [kernel.kallsyms]  [k] __xfs_bmbt_get_all               
     0.24%        xfs_io  [kernel.kallsyms]  [k] iov_iter_fault_in_readable       
     0.22%  kworker/u8:5  [kernel.kallsyms]  [k] xfs_start_buffer_writeback       
     0.22%        xfs_io  [kernel.kallsyms]  [k] fsnotify                         
     0.22%        xfs_io  [kernel.kallsyms]  [k] sys_pwrite64                     
     0.22%        xfs_io  [kernel.kallsyms]  [k] xfs_file_iomap_begin             
     0.21%  kworker/u8:5  [kernel.kallsyms]  [k] pmem_do_bvec                     
     0.21%  kworker/u8:5  [kernel.kallsyms]  [k] xfs_map_at_offset                
     0.18%  kworker/u8:5  [kernel.kallsyms]  [k] write_cache_pages                
     0.18%  kworker/u8:5  [kernel.kallsyms]  [k] xfs_map_buffer                   
     0.17%  kworker/u8:5  [kernel.kallsyms]  [k] __test_set_page_writeback        
     0.17%        xfs_io  [kernel.kallsyms]  [k] __alloc_pages_nodemask           
     0.17%        xfs_io  [kernel.kallsyms]  [k] __fsnotify_parent                
     0.17%        xfs_io  [kernel.kallsyms]  [k] block_write_end                  
     0.17%        xfs_io  [kernel.kallsyms]  [k] iomap_write_begin                
     0.17%        xfs_io  [kernel.kallsyms]  [k] iov_iter_init                    
     0.17%        xfs_io  [kernel.kallsyms]  [k] percpu_up_read                   
     0.17%        xfs_io  [kernel.kallsyms]  [k] radix_tree_lookup_slot           
     0.16%        xfs_io  [kernel.kallsyms]  [k] create_empty_buffers             
     0.16%        xfs_io  [kernel.kallsyms]  [k] timespec_trunc                   
     0.16%        xfs_io  [kernel.kallsyms]  [k] wait_for_stable_page             
     0.16%        xfs_io  [kernel.kallsyms]  [k] xfs_get_extsz_hint               
     0.14%   kworker/0:2  [kernel.kallsyms]  [k] test_clear_page_writeback        
     0.14%  kworker/u8:5  [kernel.kallsyms]  [k] release_pages                    
     0.13%        xfs_io  [kernel.kallsyms]  [k] iomap_write_end                  
     0.13%        xfs_io  [kernel.kallsyms]  [k] xfs_bmbt_get_startoff            
     0.12%  kworker/u8:5  [kernel.kallsyms]  [k] dec_zone_page_state              
     0.12%        xfs_io  [kernel.kallsyms]  [k] alloc_page_buffers               
     0.12%        xfs_io  [kernel.kallsyms]  [k] generic_write_checks             
     0.12%        xfs_io  [kernel.kallsyms]  [k] percpu_down_read                 
     0.12%        xfs_io  [kernel.kallsyms]  [k] release_pages                    
     0.12%        xfs_io  [kernel.kallsyms]  [k] set_bh_page                      
     0.12%        xfs_io  [kernel.kallsyms]  [k] xfs_find_bdev_for_inode          
     0.12%        xfs_io  xfs_io             [.] do_pwrite                        
     0.10%  kworker/u8:5  [kernel.kallsyms]  [k] mark_buffer_async_write          
     0.10%  kworker/u8:5  [kernel.kallsyms]  [k] page_waitqueue                   
     0.10%        xfs_io  [kernel.kallsyms]  [k] PageHuge                         
     0.10%        xfs_io  [kernel.kallsyms]  [k] add_to_page_cache_lru            
     0.09%   kworker/0:2  [kernel.kallsyms]  [k] end_page_writeback               
     0.09%  kworker/u8:5  [kernel.kallsyms]  [k] find_get_pages_tag               
     0.09%  kworker/u8:5  [kernel.kallsyms]  [k] xfs_start_page_writeback         
     0.09%        xfs_io  [kernel.kallsyms]  [k] create_page_buffers              
     0.09%        xfs_io  [kernel.kallsyms]  [k] page_mapping                     
     0.09%        xfs_io  [kernel.kallsyms]  [k] xfs_bmbt_get_all                 
     0.09%        xfs_io  [kernel.kallsyms]  [k] xfs_file_iomap_end               
     0.08%  kworker/u8:5  [kernel.kallsyms]  [k] inc_node_page_state              
     0.08%  kworker/u8:5  [kernel.kallsyms]  [k] inc_zone_page_state              
     0.08%  kworker/u8:5  [kernel.kallsyms]  [k] page_mapping                     
     0.08%  kworker/u8:5  [kernel.kallsyms]  [k] page_mkclean                     
     0.08%        xfs_io  [kernel.kallsyms]  [k] __sb_start_write                 
     0.08%        xfs_io  [kernel.kallsyms]  [k] current_kernel_time64            
     0.07%   kworker/0:2  [kernel.kallsyms]  [k] __wake_up_bit                    
     0.07%  kworker/u8:5  [kernel.kallsyms]  [k] page_mapped                      
     0.07%        xfs_io  [kernel.kallsyms]  [k] __lru_cache_add                  
     0.07%        xfs_io  [kernel.kallsyms]  [k] current_fs_time                  
     0.07%        xfs_io  [kernel.kallsyms]  [k] grab_cache_page_write_begin      
     0.07%        xfs_io  [kernel.kallsyms]  [k] xfs_iext_get_ext                 
     0.05%  kworker/u8:5  [kernel.kallsyms]  [k] pmem_make_request                
     0.05%  kworker/u8:5  [kernel.kallsyms]  [k] xfs_add_to_ioend                 
     0.05%        xfs_io  [kernel.kallsyms]  [k] __fdget                          
     0.05%        xfs_io  [kernel.kallsyms]  [k] __find_get_block_slow            
     0.05%        xfs_io  [kernel.kallsyms]  [k] __set_page_dirty                 
     0.05%        xfs_io  [kernel.kallsyms]  [k] alloc_buffer_head                
     0.05%        xfs_io  [kernel.kallsyms]  [k] radix_tree_lookup                
     0.05%        xfs_io  [kernel.kallsyms]  [k] radix_tree_tagged                
     0.05%        xfs_io  [kernel.kallsyms]  [k] xfs_bmbt_get_blockcount          
     0.05%        xfs_io  [kernel.kallsyms]  [k] xfs_fsb_to_db                    
     0.04%   kworker/0:2  [kernel.kallsyms]  [k] dec_zone_page_state              
     0.04%  kworker/u8:5  [kernel.kallsyms]  [k] dec_node_page_state              
     0.04%        xfs_io  [kernel.kallsyms]  [k] __radix_tree_preload             
     0.04%        xfs_io  [kernel.kallsyms]  [k] __sb_end_write                   
     0.03%   kworker/0:2  [kernel.kallsyms]  [k] dec_node_page_state              
     0.03%   kworker/0:2  [kernel.kallsyms]  [k] inc_node_page_state              
     0.03%   kworker/0:2  [kernel.kallsyms]  [k] page_mapping                     
     0.03%  kworker/u8:5  [kernel.kallsyms]  [k] radix_tree_next_chunk            
     0.03%  kworker/u8:5  [kernel.kallsyms]  [k] xfs_fsb_to_db                    
     0.03%        xfs_io  [kernel.kallsyms]  [k] _raw_spin_lock_irqsave           
     0.03%        xfs_io  [kernel.kallsyms]  [k] cache_alloc_refill               
     0.03%        xfs_io  [kernel.kallsyms]  [k] lru_cache_add                    
     0.01%   kworker/0:2  [kernel.kallsyms]  [k] cache_reap                       
     0.01%   kworker/0:2  [kernel.kallsyms]  [k] mempool_free                     
     0.01%   kworker/0:2  [kernel.kallsyms]  [k] page_waitqueue                   
     0.01%  kworker/u8:5  [kernel.kallsyms]  [k] bio_add_page                     
     0.01%  kworker/u8:5  [kernel.kallsyms]  [k] kmem_cache_alloc                 
     0.01%  kworker/u8:5  [kernel.kallsyms]  [k] lru_add_drain_cpu                
     0.01%  kworker/u8:5  [kernel.kallsyms]  [k] mempool_alloc                    
     0.01%  kworker/u8:5  [kernel.kallsyms]  [k] pagevec_lookup_tag               
     0.01%  kworker/u8:5  [kernel.kallsyms]  [k] queue_delayed_work_on            
     0.01%  kworker/u8:5  [kernel.kallsyms]  [k] queue_work_on                    
     0.01%  kworker/u8:5  [kernel.kallsyms]  [k] run_timer_softirq                
     0.01%  kworker/u8:5  [kernel.kallsyms]  [k] update_group_capacity            
     0.01%  kworker/u8:5  [kernel.kallsyms]  [k] xfs_trans_reserve                
     0.01%       swapper  [kernel.kallsyms]  [k] _raw_spin_unlock_irqrestore      
     0.01%        xfs_io  [kernel.kallsyms]  [k] _cond_resched                    
     0.01%        xfs_io  [kernel.kallsyms]  [k] mod_zone_page_state              
     0.01%        xfs_io  [kernel.kallsyms]  [k] pagevec_lru_move_fn              
     0.01%        xfs_io  [kernel.kallsyms]  [k] radix_tree_maybe_preload         
     0.01%        xfs_io  [kernel.kallsyms]  [k] unmap_underlying_metadata        
     0.01%        xfs_io  [kernel.kallsyms]  [k] xfs_bmap_worst_indlen            
     0.01%        xfs_io  ld-2.19.so         [.] 0x000000000000d866               
     0.01%        xfs_io  libc-2.19.so       [.] 0x000000000008a8da               
     0.01%        xfs_io  xfs_io             [.] pwrite64@plt                     

[toc] | [prev] | [next] | [standalone]


#1461698 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromFengguang Wu <fengguang.wu@intel.com>
Date2016-08-14 10:40 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5Z7z-5sr-21@gated-at.bofh.it>
In reply to#1461692
Hi Christoph,

On Sat, Aug 13, 2016 at 11:48:25PM +0200, Christoph Hellwig wrote:
>On Sat, Aug 13, 2016 at 02:30:54AM +0200, Christoph Hellwig wrote:
>> Below is a patch I hacked up this morning to do just that.  It passes
>> xfstests, but I've not done any real benchmarking with it.  If the
>> reduced lookup overhead in it doesn't help enough we'll need to some
>> sort of look aside cache for the information, but I hope that we
>> can avoid that.  And yes, it's a rather large patch - but the old
>> path was so entangled that I couldn't come up with something lighter.
>
>Hi Fengguang or Xiaolong,
>
>any chance to add this thread to a lkp run?

Sure. To which base should I apply it? Or if you already pushed the
git tree, I'll test your commit directly.

Thanks,
Fengguang

[toc] | [prev] | [next] | [standalone]


Page 4 of 5 — ← Prev page 1 2 3 [4] 5  Next page →

Back to top | Article view | linux.kernel


csiph-web