Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1459842 > unrolled thread

Re: [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

Started byLinus Torvalds <torvalds@linux-foundation.org>
First post2016-08-10 22:30 +0200
Last post2016-08-11 02:00 +0200
Articles 19 on this page of 99 — 11 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-10 22:30 +0200
    Re: [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 01:10 +0200
      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 02:00 +0200
        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 02:20 +0200
          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 02:30 +0200
            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 02:40 +0200
              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 03:10 +0200
                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 06:50 +0200
                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-15 19:30 +0200
                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:30 +0200
                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-11 18:00 +0200
                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 19:00 +0200
                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 20:00 +0200
                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 22:00 +0200
                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-11 22:10 +0200
                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 22:40 +0200
                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Al Viro <viro@ZenIV.linux.org.uk> - 2016-08-12 00:20 +0200
                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 00:40 +0200
                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 23:50 +0200
                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-12 00:10 +0200
                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 03:00 +0200
                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 04:30 +0200
                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 06:00 +0200
                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 20:10 +0200
                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-14 11:00 +0200
                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 02:50 +0200
                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-15 03:40 +0200
                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 04:40 +0200
                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-15 05:00 +0200
                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 07:10 +0200
                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 00:30 +0200
                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 00:50 +0200
                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 01:30 +0200
                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 01:50 +0200
                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:50 +0200
                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-16 17:10 +0200
                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 20:00 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Michal Hocko <mhocko@kernel.org> - 2016-08-17 17:50 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Michal Hocko <mhocko@kernel.org> - 2016-08-17 18:50 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-17 17:50 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-18 02:50 +0200
                                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-18 09:20 +0200
                                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-18 15:30 +0200
                                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-19 04:10 +0200
                                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-19 04:40 +0200
                                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Michal Hocko <mhocko@kernel.org> - 2016-08-19 11:10 +0200
                                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-19 13:00 +0200
                                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-20 01:50 +0200
                                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-20 03:10 +0200
                                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-20 14:20 +0200
                                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-19 06:10 +0200
                                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-19 17:10 +0200
                                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-24 17:50 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-18 04:50 +0200
                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 02:20 +0200
                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:40 +0200
                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 03:00 +0200
                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:40 +0200
                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 04:00 +0200
                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-17 00:10 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-17 01:30 +0200
                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 01:10 +0200
                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:40 +0200
                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 02:50 +0200
                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ingo Molnar <mingo@kernel.org> - 2016-08-15 07:10 +0200
                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Peter Zijlstra <peterz@infradead.org> - 2016-08-17 18:30 +0200
                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-15 15:10 +0200
                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 04:30 +0200
                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 04:40 +0200
                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-12 05:00 +0200
                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 05:30 +0200
                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 06:20 +0200
                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 07:10 +0200
                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 08:10 +0200
                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ye Xiaolong <xiaolong.ye@intel.com> - 2016-08-12 08:40 +0200
                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ye Xiaolong <xiaolong.ye@intel.com> - 2016-08-12 11:00 +0200
                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 12:10 +0200
                                        Re: [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-12 12:50 +0200
                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-13 02:40 +0200
                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-14 10:30 +0200
                                            Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-14 10:40 +0200
                                              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-14 11:00 +0200
                                                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-14 11:30 +0200
                                                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%  regression Christoph Hellwig <hch@lst.de> - 2016-08-14 18:20 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 01:50 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-15 02:00 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-15 16:20 +0200
                                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 23:30 +0200
                                                        Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-16 14:30 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-15 22:40 +0200
                                                      Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6%     regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-23 00:10 +0200
                                                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-16 15:30 +0200
                                          Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-14 12:00 +0200
              Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 03:20 +0200
                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 03:40 +0200
                  Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ye Xiaolong <xiaolong.ye@intel.com> - 2016-08-11 04:50 +0200
                    Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 05:20 +0200
                Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 03:30 +0200
      Re: [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 02:00 +0200

Page 5 of 5 — ← Prev page 1 2 3 4 [5]


#1461698 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromFengguang Wu <fengguang.wu@intel.com>
Date2016-08-14 10:40 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5Z7z-5sr-21@gated-at.bofh.it>
In reply to#1461692
Hi Christoph,

On Sat, Aug 13, 2016 at 11:48:25PM +0200, Christoph Hellwig wrote:
>On Sat, Aug 13, 2016 at 02:30:54AM +0200, Christoph Hellwig wrote:
>> Below is a patch I hacked up this morning to do just that.  It passes
>> xfstests, but I've not done any real benchmarking with it.  If the
>> reduced lookup overhead in it doesn't help enough we'll need to some
>> sort of look aside cache for the information, but I hope that we
>> can avoid that.  And yes, it's a rather large patch - but the old
>> path was so entangled that I couldn't come up with something lighter.
>
>Hi Fengguang or Xiaolong,
>
>any chance to add this thread to a lkp run?

Sure. To which base should I apply it? Or if you already pushed the
git tree, I'll test your commit directly.

Thanks,
Fengguang

[toc] | [prev] | [next] | [standalone]


#1461731 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromChristoph Hellwig <hch@lst.de>
Date2016-08-14 11:00 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5ZqW-5AD-41@gated-at.bofh.it>
In reply to#1461698
Hi Fengguang,

feel free to try this git tree:

   git://git.infradead.org/users/hch/vfs.git iomap-fixes

[toc] | [prev] | [next] | [standalone]


#1461745 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromFengguang Wu <fengguang.wu@intel.com>
Date2016-08-14 11:30 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s5ZTX-63p-7@gated-at.bofh.it>
In reply to#1461731
Hi Christoph,

On Sun, Aug 14, 2016 at 12:15:08AM +0200, Christoph Hellwig wrote:
>Hi Fengguang,
>
>feel free to try this git tree:
>
>   git://git.infradead.org/users/hch/vfs.git iomap-fixes

I just queued some test jobs for it.

% queue -q vip -t ivb44 -b hch-vfs/iomap-fixes aim7-fs-1brd.yaml fs=xfs -r3 -k fe9c2c81ed073878768785a985295cbacc349e42 -k ca2edab2e1d8f30dda874b7f717c2d4664991e9b -k 99091700659f4df965e138b38b4fa26a29b7eade

That job file can be found here:

        https://git.kernel.org/cgit/linux/kernel/git/wfg/lkp-tests.git/tree/jobs/aim7-fs-1brd.yaml

It specifies a matrix of the below atom tests:

        wfg /c/lkp-tests% split-job jobs/aim7-fs-1brd.yaml -s 'fs: xfs'

        jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_src-3000-performance.yaml
        jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_rr-3000-performance.yaml
        jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_rw-3000-performance.yaml
        jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_cp-3000-performance.yaml
        jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_wrt-3000-performance.yaml
        jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-sync_disk_rw-600-performance.yaml
        jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-creat-clo-1500-performance.yaml
        jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_rd-9000-performance.yaml

If you see other suitable tests for this patch, feel free to drop me a
hint.  I'v queued these jobs to the other machines to make them run in
parallel.

% queue -q vip -t ivb43 -b hch-vfs/iomap-fixes fsmark-stress-journal-1hdd.yaml fsmark-stress-journal-1brd.yaml fs=xfs -r3 -k fe9c2c81ed073878768785a985295cbacc349e42 -k ca2edab2e1d8f30dda874b7f717c2d4664991e9b -k 99091700659f4df965e138b38b4fa26a29b7eade

% queue -q vip -t ivb44 -b hch-vfs/iomap-fixes fsmark-generic-1brd.yaml dd-write-1hdd.yaml  fsmark-generic-1hdd.yaml  fs=xfs -r3 -k fe9c2c81ed073878768785a985295cbacc349e42 -k ca2edab2e1d8f30dda874b7f717c2d4664991e9b -k 99091700659f4df965e138b38b4fa26a29b7eade

% queue -q vip -t lkp-hsx02 -b hch-vfs/iomap-fixes fsmark-generic-brd-raid.yaml fs=xfs -r3 -k fe9c2c81ed073878768785a985[0/1710]349e42 -k ca2edab2e1d8f30dda874b7f717c2d4664991e9b -k 99091700659f4df965e138b38b4fa26a29b7eade

% queue -q vip -t lkp-hsw-ep4 -b hch-vfs/iomap-fixes fsmark-1ssd-nvme-small.yaml fs=xfs -r3 -k fe9c2c81ed073878768785a985295cbacc349e42 -k ca2edab2e1d8f30dda874b7f717c2d4664991e9b -k 99091700659f4df965e138b38b4fa26a29b7eade

Thanks,
Fengguang

[toc] | [prev] | [next] | [standalone]


#1462166 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromChristoph Hellwig <hch@lst.de>
Date2016-08-14 18:20 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s66iJ-1Ra-13@gated-at.bofh.it>
In reply to#1461745
Snipping the long contest:

I think there are three observations here:

 (1) removing the mark_page_accessed (which is the only significant
     change in the parent commit)  hurts the
     aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
     I'd still rather stick to the filemap version and let the
     VM people sort it out.  How do the numbers for this test
     look for XFS vs say ext4 and btrfs?
 (2) lots of additional spinlock contention in the new case.  A quick
     check shows that I fat-fingered my rewrite so that we do
     the xfs_inode_set_eofblocks_tag call now for the pure lookup
     case, and pretty much all new cycles come from that.
 (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
     we're already doing way to many even without my little bug above.

So I've force pushed a new version of the iomap-fixes branch with
(2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
lot less expensive slotted in before that.  Would be good to see
the numbers with that.

[toc] | [prev] | [next] | [standalone]


#1462505 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-15 01:50 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6dkd-6l9-1@gated-at.bofh.it>
In reply to#1462166
On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
> Snipping the long contest:
> 
> I think there are three observations here:
> 
>  (1) removing the mark_page_accessed (which is the only significant
>      change in the parent commit)  hurts the
>      aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
>      I'd still rather stick to the filemap version and let the
>      VM people sort it out.  How do the numbers for this test
>      look for XFS vs say ext4 and btrfs?
>  (2) lots of additional spinlock contention in the new case.  A quick
>      check shows that I fat-fingered my rewrite so that we do
>      the xfs_inode_set_eofblocks_tag call now for the pure lookup
>      case, and pretty much all new cycles come from that.
>  (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
>      we're already doing way to many even without my little bug above.
> 
> So I've force pushed a new version of the iomap-fixes branch with
> (2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
> lot less expensive slotted in before that.  Would be good to see
> the numbers with that.

With this new set of fixes, the 1byte write test runs ~30% faster on
my test machine (130k writes/s vs 100k writes/s), and the 1k write
on the pmem device runs about 10% faster (660MB/s vs 590MB/s).
dbench numbers on the pmem device also go through the roof (they
didn't show any regression to begin with) - 50% faster at 16 clients
on a 16AG filesystem (5700MB/s vs 3800MB/s).

The 10Mx4k file create fsmark workload I run (on the sparse 500TB
XFS filesystem backed by a pair of SSDs)  is giving the highest
throughput *and* the lowest std dev I've ever recorded
(55014.8+/-1.3e+04 files/s) and that shows in the runtime which also
drops from 3m57s to 3m22s.

So regardless of what aim7 results we get from these changes, I'll
be merging them pending review and further testing...

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1462506 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromFengguang Wu <fengguang.wu@intel.com>
Date2016-08-15 02:00 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6dtY-6ox-7@gated-at.bofh.it>
In reply to#1462166
Hi Christoph,

On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
>Snipping the long contest:
>
>I think there are three observations here:
>
> (1) removing the mark_page_accessed (which is the only significant
>     change in the parent commit)  hurts the
>     aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
>     I'd still rather stick to the filemap version and let the
>     VM people sort it out.  How do the numbers for this test
>     look for XFS vs say ext4 and btrfs?

We'll be able to compare between filesystems when the tests for Linus'
patch finish.

> (2) lots of additional spinlock contention in the new case.  A quick
>     check shows that I fat-fingered my rewrite so that we do
>     the xfs_inode_set_eofblocks_tag call now for the pure lookup
>     case, and pretty much all new cycles come from that.
> (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
>     we're already doing way to many even without my little bug above.
>
>So I've force pushed a new version of the iomap-fixes branch with
>(2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
>lot less expensive slotted in before that.  Would be good to see
>the numbers with that.

I just queued these jobs. The comment-out ones will be submitted as
the 2nd stage when the 1st-round quick tests finish.

queue=(
        queue
        -q vip
        --repeat-to 3
        fs=xfs
        perf-profile.delay=1
        -b hch-vfs/iomap-fixes
        -k bf4dc6e4ecc2a3d042029319bc8cd4204c185610
        -k 74a242ad94d13436a1644c0b4586700e39871491
        -k 99091700659f4df965e138b38b4fa26a29b7eade
)

"${queue[@]}" -t ivb44 aim7-fs-1brd.yaml
"${queue[@]}" -t ivb44 fsmark-generic-1brd.yaml
"${queue[@]}" -t ivb43 fsmark-stress-journal-1brd.yaml
"${queue[@]}" -t lkp-hsx02 fsmark-generic-brd-raid.yaml
"${queue[@]}" -t lkp-hsw-ep4 fsmark-1ssd-nvme-small.yaml
#"${queue[@]}" -t ivb43 fsmark-stress-journal-1hdd.yaml
#"${queue[@]}" -t ivb44 dd-write-1hdd.yaml fsmark-generic-1hdd.yaml

Thanks,
Fengguang

[toc] | [prev] | [next] | [standalone]


#1462814 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromFengguang Wu <fengguang.wu@intel.com>
Date2016-08-15 16:20 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6qU9-6EH-17@gated-at.bofh.it>
In reply to#1462166
Hi Christoph,

On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
>Snipping the long contest:
>
>I think there are three observations here:
>
> (1) removing the mark_page_accessed (which is the only significant
>     change in the parent commit)  hurts the
>     aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
>     I'd still rather stick to the filemap version and let the
>     VM people sort it out.  How do the numbers for this test
>     look for XFS vs say ext4 and btrfs?
> (2) lots of additional spinlock contention in the new case.  A quick
>     check shows that I fat-fingered my rewrite so that we do
>     the xfs_inode_set_eofblocks_tag call now for the pure lookup
>     case, and pretty much all new cycles come from that.
> (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
>     we're already doing way to many even without my little bug above.
>
>So I've force pushed a new version of the iomap-fixes branch with
>(2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
>lot less expensive slotted in before that.  Would be good to see
>the numbers with that.

The aim7 1BRD tests finished and there are ups and downs, with overall
performance remain flat.

99091700659f4df9  74a242ad94d13436a1644c0b45  bf4dc6e4ecc2a3d042029319bc  testcase/testparams/testbox
----------------  --------------------------  --------------------------  ---------------------------
         %stddev     %change         %stddev     %change         %stddev
             \          |                \          |                \  
    159926                      157324                      158574        GEO-MEAN aim7.jobs-per-min
     70897               5%      74137               4%      73775        aim7/1BRD_48G-xfs-creat-clo-1500-performance/ivb44
    485217 ±  3%                492431                      477533        aim7/1BRD_48G-xfs-disk_rd-9000-performance/ivb44
    360451             -19%     292980             -17%     299377        aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
    338114                      338410               5%     354078        aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
     60130 ±  5%         4%      62438               5%      62923        aim7/1BRD_48G-xfs-disk_src-3000-performance/ivb44
    403144                      397790                      410648        aim7/1BRD_48G-xfs-disk_wrt-3000-performance/ivb44
     26327                       26534                       26128        aim7/1BRD_48G-xfs-sync_disk_rw-600-performance/ivb44

The new commit bf4dc6e ("xfs: rewrite and optimize the delalloc write
path") improves the aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
case by 5%. Here are the detailed numbers:

aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44

74a242ad94d13436  bf4dc6e4ecc2a3d042029319bc
----------------  --------------------------
         %stddev     %change         %stddev
             \          |                \
    338410               5%     354078        aim7.jobs-per-min
    404390               8%     435117        aim7.time.voluntary_context_switches
      2502              -4%       2396        aim7.time.maximum_resident_set_size
     15018              -9%      13701        aim7.time.involuntary_context_switches
       900             -11%        801        aim7.time.system_time
     17432              11%      19365        vmstat.system.cs
     47736 ± 19%       -24%      36087        interrupts.CAL:Function_call_interrupts
   2129646              31%    2790638        proc-vmstat.pgalloc_dma32
    379503              13%     429384        numa-meminfo.node0.Dirty
     15018              -9%      13701        time.involuntary_context_switches
       900             -11%        801        time.system_time
      1560              10%       1716        slabinfo.mnt_cache.active_objs
      1560              10%       1716        slabinfo.mnt_cache.num_objs
     61.53               -4      57.45 ±  4%  perf-profile.cycles-pp.intel_idle.cpuidle_enter_state.cpuidle_enter.call_cpuidle.cpu_startup_entry
     61.63               -4      57.55 ±  4%  perf-profile.func.cycles-pp.intel_idle
   1007188 ± 16%       156%    2577911 ±  6%  numa-numastat.node0.numa_miss
   9662857 ±  4%       -13%    8420159 ±  3%  numa-numastat.node0.numa_foreign
   1008220 ± 16%       155%    2570630 ±  6%  numa-numastat.node1.numa_foreign
   9664033 ±  4%       -13%    8413184 ±  3%  numa-numastat.node1.numa_miss
  26519887 ±  3%        18%   31322674        cpuidle.C1-IVT.time
    122238              16%     142383        cpuidle.C1-IVT.usage
     46548              11%      51645        cpuidle.C1E-IVT.usage
  17253419              13%   19567582        cpuidle.C3-IVT.time
     86847              13%      98333        cpuidle.C3-IVT.usage
    482033 ± 12%       108%    1000665 ±  8%  numa-vmstat.node0.numa_miss
     94689              14%     107744        numa-vmstat.node0.nr_zone_write_pending
     94677              14%     107718        numa-vmstat.node0.nr_dirty
   3156643 ±  3%       -20%    2527460 ±  3%  numa-vmstat.node0.numa_foreign
    429288 ± 12%       129%     983053 ±  8%  numa-vmstat.node1.numa_foreign
   3104193 ±  3%       -19%    2510128        numa-vmstat.node1.numa_miss
      6.43 ±  5%        51%       9.70 ± 11%  turbostat.Pkg%pc2
      0.30              28%       0.38        turbostat.CPU%c3
      9.71                        9.92        turbostat.RAMWatt
       158                         154        turbostat.PkgWatt
       125              -3%        121        turbostat.CorWatt
      1141              -6%       1078        turbostat.Avg_MHz
     38.70              -6%      36.48        turbostat.%Busy
      5.03 ± 11%       -51%       2.46 ± 40%  turbostat.Pkg%pc6
      8.33 ± 48%        88%      15.67 ± 36%  sched_debug.cfs_rq:/.runnable_load_avg.max
      1947 ±  3%       -12%       1710 ±  7%  sched_debug.cfs_rq:/.spread0.stddev
      1936 ±  3%       -12%       1698 ±  8%  sched_debug.cfs_rq:/.min_vruntime.stddev
      2170 ± 10%       -14%       1863 ±  6%  sched_debug.cfs_rq:/.load_avg.max
    220926 ± 18%        37%     303192 ±  5%  sched_debug.cpu.avg_idle.stddev
      0.06 ± 13%       357%       0.28 ± 23%  sched_debug.rt_rq:/.rt_time.avg
      0.37 ± 10%       240%       1.25 ± 15%  sched_debug.rt_rq:/.rt_time.stddev
      2.54 ± 10%       160%       6.59 ± 10%  sched_debug.rt_rq:/.rt_time.max
      0.32 ± 19%        29%       0.42 ± 10%  perf-stat.dTLB-load-miss-rate
    964727               7%    1028830        perf-stat.context-switches
    176406               4%     184289        perf-stat.cpu-migrations
      0.29               4%       0.30        perf-stat.branch-miss-rate
 1.634e+09                   1.673e+09        perf-stat.node-store-misses
     23.60                       23.99        perf-stat.node-store-miss-rate
     40.01                       40.57        perf-stat.cache-miss-rate
      0.95              -8%       0.87        perf-stat.ipc
 3.203e+12              -9%  2.928e+12        perf-stat.cpu-cycles
 1.506e+09             -11%  1.345e+09        perf-stat.branch-misses
     50.64 ± 13%       -14%      43.45 ±  4%  perf-stat.iTLB-load-miss-rate
 5.285e+11             -14%  4.523e+11        perf-stat.branch-instructions
 3.042e+12             -16%  2.551e+12        perf-stat.instructions
 7.996e+11             -18%  6.584e+11        perf-stat.dTLB-loads
 5.569e+11 ±  4%       -18%  4.578e+11        perf-stat.dTLB-stores


Here are the detailed numbers for the slowed down case:

aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44

99091700659f4df9  bf4dc6e4ecc2a3d042029319bc
----------------  --------------------------
         %stddev      change         %stddev
             \          |                \
    360451             -17%     299377        aim7.jobs-per-min
     12806             481%      74447        aim7.time.involuntary_context_switches
       755              44%       1086        aim7.time.system_time
     50.17              20%      60.36        aim7.time.elapsed_time
     50.17              20%      60.36        aim7.time.elapsed_time.max
    438148                      446012        aim7.time.voluntary_context_switches
     37798 ± 16%       780%     332583 ±  8%  interrupts.CAL:Function_call_interrupts
     78.82 ±  5%        18%      93.35 ±  5%  uptime.boot
      2847 ±  7%        11%       3160 ±  7%  uptime.idle
    147490 ±  8%        34%     197261 ±  3%  softirqs.RCU
    648159              29%     839283        softirqs.TIMER
    160830              10%     177144        softirqs.SCHED
   3845352 ±  4%        91%    7349133        numa-numastat.node0.numa_miss
   4686838 ±  5%        67%    7835640        numa-numastat.node0.numa_foreign
   3848455 ±  4%        91%    7352436        numa-numastat.node1.numa_foreign
   4689920 ±  5%        67%    7838734        numa-numastat.node1.numa_miss
     50.17              20%      60.36        time.elapsed_time.max
     12806             481%      74447        time.involuntary_context_switches
       755              44%       1086        time.system_time
     50.17              20%      60.36        time.elapsed_time
      1563              18%       1846        time.percent_of_cpu_this_job_got
     11699 ± 19%      3738%     449048        vmstat.io.bo
  18836969             -16%   15789996        vmstat.memory.free
        16              19%         19        vmstat.procs.r
     19377             459%     108364        vmstat.system.cs
     48255              11%      53537        vmstat.system.in
   2357299              25%    2951384        meminfo.Inactive(file)
   2366381              25%    2960468        meminfo.Inactive
   1575292              -9%    1429971        meminfo.Cached
  19342499             -17%   16100340        meminfo.MemFree
   1057904             -20%     842987        meminfo.Dirty
      1057              21%       1284        turbostat.Avg_MHz
     35.78              21%      43.24        turbostat.%Busy
      9.95              15%      11.47        turbostat.RAMWatt
        74 ±  5%        10%         81        turbostat.CoreTmp
        74 ±  4%        10%         81        turbostat.PkgTmp
       118               8%        128        turbostat.CorWatt
       151               7%        162        turbostat.PkgWatt
     29.06             -23%      22.39        turbostat.CPU%c6
       487 ± 89%      3e+04      26448 ± 57%  latency_stats.max.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
      1823 ± 82%      2e+06    1913796 ± 38%  latency_stats.sum.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
    208475 ± 43%      1e+06    1409494 ±  5%  latency_stats.sum.wait_on_page_bit.truncate_inode_pages_range.truncate_inode_pages_final.evict.iput.dentry_unlink_inode.__dentry_kill.dput.__fput.____fput.task_work_run.exit_to_usermode_loop
      6884 ± 73%      8e+04      90790 ±  9%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.file_update_time.xfs_file_aio_write_checks.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write.SyS_write
      1598 ± 20%      3e+04      35015 ± 27%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_free_eofblocks.xfs_release.xfs_file_release.__fput.____fput.task_work_run
      2006 ± 25%      3e+04      31143 ± 35%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode.evict.iput
        29 ±101%      1e+04      10214 ± 29%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_defer_trans_roll.xfs_defer_finish.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode
      1206 ± 51%      9e+03       9919 ± 25%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.touch_atime.generic_file_read_iter.xfs_file_buffered_aio_read.xfs_file_read_iter.__vfs_read.vfs_read.SyS_read
  29869205 ±  4%       -10%   26804569        cpuidle.C1-IVT.time
   5737726              39%    7952214        cpuidle.C1E-IVT.time
     51141              17%      59958        cpuidle.C1E-IVT.usage
  18377551              37%   25176426        cpuidle.C3-IVT.time
     96067              17%     112045        cpuidle.C3-IVT.usage
   1806811              12%    2024041        cpuidle.C6-IVT.usage
   1104420 ± 36%       204%    3361085 ± 27%  cpuidle.POLL.time
       281 ± 10%        20%        338        cpuidle.POLL.usage
      5.61 ± 11%       -0.5       5.12 ± 18%  perf-profile.cycles-pp.irq_exit.smp_apic_timer_interrupt.apic_timer_interrupt.cpuidle_enter.call_cpuidle
      5.85 ±  6%       -0.8       5.06 ± 15%  perf-profile.cycles-pp.hrtimer_interrupt.local_apic_timer_interrupt.smp_apic_timer_interrupt.apic_timer_interrupt.cpuidle_enter
      6.32 ±  6%       -0.9       5.42 ± 15%  perf-profile.cycles-pp.local_apic_timer_interrupt.smp_apic_timer_interrupt.apic_timer_interrupt.cpuidle_enter.call_cpuidle
     15.77 ±  8%         -2      13.83 ± 17%  perf-profile.cycles-pp.smp_apic_timer_interrupt.apic_timer_interrupt.cpuidle_enter.call_cpuidle.cpu_startup_entry
     16.04 ±  8%         -2      14.01 ± 15%  perf-profile.cycles-pp.apic_timer_interrupt.cpuidle_enter.call_cpuidle.cpu_startup_entry.start_secondary
     60.25 ±  4%         -7      53.03 ±  7%  perf-profile.cycles-pp.intel_idle.cpuidle_enter_state.cpuidle_enter.call_cpuidle.cpu_startup_entry
     60.41 ±  4%         -7      53.12 ±  7%  perf-profile.func.cycles-pp.intel_idle
   1174104              22%    1436859        numa-meminfo.node0.Inactive
   1167471              22%    1428271        numa-meminfo.node0.Inactive(file)
    770811              -9%     698147        numa-meminfo.node0.FilePages
  20707294             -12%   18281509 ±  6%  numa-meminfo.node0.Active
  20613745             -12%   18180987 ±  6%  numa-meminfo.node0.Active(file)
   9676639             -17%    8003627        numa-meminfo.node0.MemFree
    509906             -22%     396192        numa-meminfo.node0.Dirty
   1189539              28%    1524697        numa-meminfo.node1.Inactive(file)
   1191989              28%    1525194        numa-meminfo.node1.Inactive
    804508             -10%     727067        numa-meminfo.node1.FilePages
   9654540             -16%    8077810        numa-meminfo.node1.MemFree
    547956             -19%     441933        numa-meminfo.node1.Dirty
       396 ± 12%       485%       2320 ± 37%  slabinfo.bio-1.num_objs
       396 ± 12%       481%       2303 ± 37%  slabinfo.bio-1.active_objs
        73             140%        176 ± 14%  slabinfo.kmalloc-128.active_slabs
        73             140%        176 ± 14%  slabinfo.kmalloc-128.num_slabs
      4734              94%       9171 ± 11%  slabinfo.kmalloc-128.num_objs
      4734              88%       8917 ± 13%  slabinfo.kmalloc-128.active_objs
     16238             -10%      14552 ±  3%  slabinfo.kmalloc-256.active_objs
     17189             -13%      15033 ±  3%  slabinfo.kmalloc-256.num_objs
     20651              96%      40387 ± 17%  slabinfo.radix_tree_node.active_objs
       398              91%        761 ± 17%  slabinfo.radix_tree_node.active_slabs
       398              91%        761 ± 17%  slabinfo.radix_tree_node.num_slabs
     22313              91%      42650 ± 17%  slabinfo.radix_tree_node.num_objs
        32             638%        236 ± 28%  slabinfo.xfs_efd_item.active_slabs
        32             638%        236 ± 28%  slabinfo.xfs_efd_item.num_slabs
      1295             281%       4934 ± 23%  slabinfo.xfs_efd_item.num_objs
      1295             280%       4923 ± 23%  slabinfo.xfs_efd_item.active_objs
      1661              81%       3000 ± 42%  slabinfo.xfs_log_ticket.num_objs
      1661              78%       2952 ± 42%  slabinfo.xfs_log_ticket.active_objs
      2617              49%       3905 ± 30%  slabinfo.xfs_trans.num_objs
      2617              48%       3870 ± 31%  slabinfo.xfs_trans.active_objs
   1015933             567%    6779099        perf-stat.context-switches
 4.864e+08             126%  1.101e+09        perf-stat.node-load-misses
 1.179e+09             103%  2.399e+09        perf-stat.node-loads
      0.06 ± 34%        92%       0.12 ± 11%  perf-stat.dTLB-store-miss-rate
 2.985e+08 ± 32%        86%  5.542e+08 ± 11%  perf-stat.dTLB-store-misses
 2.551e+09 ± 15%        81%  4.625e+09 ± 13%  perf-stat.dTLB-load-misses
      0.39 ± 14%        66%       0.65 ± 13%  perf-stat.dTLB-load-miss-rate
  1.26e+09              60%  2.019e+09        perf-stat.node-store-misses
  46072661 ± 27%        49%   68472915        perf-stat.iTLB-loads
 2.738e+12 ±  4%        43%  3.916e+12        perf-stat.cpu-cycles
     21.48              32%      28.35        perf-stat.node-store-miss-rate
 1.612e+10 ±  3%        28%  2.066e+10        perf-stat.cache-references
 1.669e+09 ±  3%        24%  2.063e+09        perf-stat.branch-misses
 6.816e+09 ±  3%        20%  8.179e+09        perf-stat.cache-misses
    177699              18%     209145        perf-stat.cpu-migrations
      0.39              13%       0.44        perf-stat.branch-miss-rate
 4.606e+09              11%  5.102e+09        perf-stat.node-stores
 4.329e+11 ±  4%         9%  4.727e+11        perf-stat.branch-instructions
 6.458e+11               9%  7.046e+11        perf-stat.dTLB-loads
     29.19               8%      31.45        perf-stat.node-load-miss-rate
    286173               8%     308115        perf-stat.page-faults
    286191               8%     308109        perf-stat.minor-faults
  45084934               4%   47073719        perf-stat.iTLB-load-misses
     42.28              -6%      39.58        perf-stat.cache-miss-rate
     50.62 ± 16%       -19%      40.75        perf-stat.iTLB-load-miss-rate
      0.89             -28%       0.64        perf-stat.ipc
         2 ± 36%     4e+07%     970191        proc-vmstat.pgrotated
       150 ± 21%     1e+07%   15356485 ±  3%  proc-vmstat.nr_vmscan_immediate_reclaim
     76823 ± 35%     56899%   43788651        proc-vmstat.pgscan_direct
    153407 ± 19%      4483%    7031431        proc-vmstat.nr_written
    619699 ± 19%      4441%   28139689        proc-vmstat.pgpgout
   5342421            1061%   62050709        proc-vmstat.pgactivate
        47 ± 25%       354%        217        proc-vmstat.nr_pages_scanned
   8542963 ±  3%        78%   15182914        proc-vmstat.numa_miss
   8542963 ±  3%        78%   15182715        proc-vmstat.numa_foreign
   2820568              31%    3699073        proc-vmstat.pgalloc_dma32
    589234              25%     738160        proc-vmstat.nr_zone_inactive_file
    589240              25%     738155        proc-vmstat.nr_inactive_file
  61347830              13%   69522958        proc-vmstat.pgfree
    393711              -9%     356981        proc-vmstat.nr_file_pages
   4831749             -17%    4020131        proc-vmstat.nr_free_pages
  61252784             -18%   50183773        proc-vmstat.pgrefill
  61245420             -18%   50176301        proc-vmstat.pgdeactivate
    264397             -20%     210222        proc-vmstat.nr_zone_write_pending
    264367             -20%     210188        proc-vmstat.nr_dirty
  60420248             -39%   36646178        proc-vmstat.pgscan_kswapd
  60373976             -44%   33735064        proc-vmstat.pgsteal_kswapd
      1753             -98%         43 ± 18%  proc-vmstat.pageoutrun
      1095             -98%         25 ± 17%  proc-vmstat.kswapd_low_wmark_hit_quickly
       656 ±  3%       -98%         15 ± 24%  proc-vmstat.kswapd_high_wmark_hit_quickly
         0                     1136221        numa-vmstat.node0.workingset_refault
         0                     1136221        numa-vmstat.node0.workingset_activate
        23 ± 45%     1e+07%    2756907        numa-vmstat.node0.nr_vmscan_immediate_reclaim
     37618 ± 24%      3234%    1254165        numa-vmstat.node0.nr_written
   1346538 ±  4%       104%    2748439        numa-vmstat.node0.numa_miss
   1577620 ±  5%        80%    2842882        numa-vmstat.node0.numa_foreign
    291242              23%     357407        numa-vmstat.node0.nr_inactive_file
    291237              23%     357390        numa-vmstat.node0.nr_zone_inactive_file
  13961935              12%   15577331        numa-vmstat.node0.numa_local
  13961938              12%   15577332        numa-vmstat.node0.numa_hit
     39831              10%      43768        numa-vmstat.node0.nr_unevictable
     39831              10%      43768        numa-vmstat.node0.nr_zone_unevictable
    193467             -10%     174639        numa-vmstat.node0.nr_file_pages
   5147212             -12%    4542321 ±  6%  numa-vmstat.node0.nr_active_file
   5147237             -12%    4542325 ±  6%  numa-vmstat.node0.nr_zone_active_file
   2426129             -17%    2008637        numa-vmstat.node0.nr_free_pages
    128285             -23%      99206        numa-vmstat.node0.nr_zone_write_pending
    128259             -23%      99183        numa-vmstat.node0.nr_dirty
         0                     1190594        numa-vmstat.node1.workingset_refault
         0                     1190594        numa-vmstat.node1.workingset_activate
        21 ± 36%     1e+07%    3120425 ±  4%  numa-vmstat.node1.nr_vmscan_immediate_reclaim
     38541 ± 26%      3336%    1324185        numa-vmstat.node1.nr_written
   1316819 ±  4%       105%    2699075        numa-vmstat.node1.numa_foreign
   1547929 ±  4%        80%    2793491        numa-vmstat.node1.numa_miss
    296714              28%     381124        numa-vmstat.node1.nr_zone_inactive_file
    296714              28%     381123        numa-vmstat.node1.nr_inactive_file
  14311131              10%   15750908        numa-vmstat.node1.numa_hit
  14311130              10%   15750905        numa-vmstat.node1.numa_local
    201164             -10%     181742        numa-vmstat.node1.nr_file_pages
   2422825             -16%    2027750        numa-vmstat.node1.nr_free_pages
    137069             -19%     110501        numa-vmstat.node1.nr_zone_write_pending
    137069             -19%     110497        numa-vmstat.node1.nr_dirty
       737 ± 29%     27349%     202387        sched_debug.cfs_rq:/.min_vruntime.min
      3637 ± 20%      7919%     291675        sched_debug.cfs_rq:/.min_vruntime.avg
     11.00 ± 44%      4892%     549.17 ±  9%  sched_debug.cfs_rq:/.runnable_load_avg.max
      2.12 ± 36%      4853%     105.12 ±  5%  sched_debug.cfs_rq:/.runnable_load_avg.stddev
      1885 ±  6%      4189%      80870        sched_debug.cfs_rq:/.min_vruntime.stddev
      1896 ±  6%      4166%      80895        sched_debug.cfs_rq:/.spread0.stddev
     10774 ± 13%      4113%     453925        sched_debug.cfs_rq:/.min_vruntime.max
      1.02 ± 19%      2630%      27.72 ±  7%  sched_debug.cfs_rq:/.runnable_load_avg.avg
     63060 ± 45%       776%     552157        sched_debug.cfs_rq:/.load.max
     14442 ± 21%       590%      99615 ± 14%  sched_debug.cfs_rq:/.load.stddev
      8397 ±  9%       309%      34370 ± 12%  sched_debug.cfs_rq:/.load.avg
     46.02 ± 24%       176%     126.96 ±  6%  sched_debug.cfs_rq:/.util_avg.stddev
       817              19%        974 ±  3%  sched_debug.cfs_rq:/.util_avg.max
       721             -17%        600 ±  3%  sched_debug.cfs_rq:/.util_avg.avg
       595 ± 11%       -38%        371 ±  7%  sched_debug.cfs_rq:/.util_avg.min
      1484 ± 20%       -47%        792 ±  5%  sched_debug.cfs_rq:/.load_avg.min
      1798 ±  4%       -50%        903 ±  5%  sched_debug.cfs_rq:/.load_avg.avg
       322 ±  8%      7726%      25239 ±  8%  sched_debug.cpu.nr_switches.min
       969            7238%      71158        sched_debug.cpu.nr_switches.avg
      2.23 ± 40%      4650%     106.14 ±  4%  sched_debug.cpu.cpu_load[0].stddev
       943 ±  4%      3475%      33730 ±  3%  sched_debug.cpu.nr_switches.stddev
      0.87 ± 25%      3057%      27.46 ±  7%  sched_debug.cpu.cpu_load[0].avg
      5.43 ± 13%      2232%     126.61        sched_debug.cpu.nr_uninterruptible.stddev
      6131 ±  3%      2028%     130453        sched_debug.cpu.nr_switches.max
      1.58 ± 29%      1852%      30.90 ±  4%  sched_debug.cpu.cpu_load[4].avg
      2.00 ± 49%      1422%      30.44 ±  5%  sched_debug.cpu.cpu_load[3].avg
     63060 ± 45%      1053%     726920 ± 32%  sched_debug.cpu.load.max
     21.25 ± 44%       777%     186.33 ±  7%  sched_debug.cpu.nr_uninterruptible.max
     14419 ± 21%       731%     119865 ± 31%  sched_debug.cpu.load.stddev
      3586             381%      17262        sched_debug.cpu.nr_load_updates.min
      8286 ±  8%       364%      38414 ± 17%  sched_debug.cpu.load.avg
      5444             303%      21956        sched_debug.cpu.nr_load_updates.avg
      1156             231%       3827        sched_debug.cpu.nr_load_updates.stddev
      8603 ±  4%       222%      27662        sched_debug.cpu.nr_load_updates.max
      1410             165%       3735        sched_debug.cpu.curr->pid.max
     28742 ± 15%       120%      63101 ±  7%  sched_debug.cpu.clock.min
     28742 ± 15%       120%      63101 ±  7%  sched_debug.cpu.clock_task.min
     28748 ± 15%       120%      63107 ±  7%  sched_debug.cpu.clock.avg
     28748 ± 15%       120%      63107 ±  7%  sched_debug.cpu.clock_task.avg
     28751 ± 15%       120%      63113 ±  7%  sched_debug.cpu.clock.max
     28751 ± 15%       120%      63113 ±  7%  sched_debug.cpu.clock_task.max
       442 ± 11%        93%        854 ± 15%  sched_debug.cpu.curr->pid.avg
       618 ±  3%        72%       1065 ±  4%  sched_debug.cpu.curr->pid.stddev
      1.88 ± 11%        50%       2.83 ±  8%  sched_debug.cpu.clock.stddev
      1.88 ± 11%        50%       2.83 ±  8%  sched_debug.cpu.clock_task.stddev
      5.22 ±  9%       -55%       2.34 ± 23%  sched_debug.rt_rq:/.rt_time.max
      0.85             -55%       0.38 ± 28%  sched_debug.rt_rq:/.rt_time.stddev
      0.17             -56%       0.07 ± 33%  sched_debug.rt_rq:/.rt_time.avg
     27633 ± 16%       124%      61980 ±  8%  sched_debug.ktime
     28745 ± 15%       120%      63102 ±  7%  sched_debug.sched_clk
     28745 ± 15%       120%      63102 ±  7%  sched_debug.cpu_clk

Thanks,
Fengguang

[toc] | [prev] | [next] | [standalone]


#1463187 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-15 23:30 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6xCi-2w3-19@gated-at.bofh.it>
In reply to#1462814
On Mon, Aug 15, 2016 at 10:14:55PM +0800, Fengguang Wu wrote:
> Hi Christoph,
> 
> On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
> >Snipping the long contest:
> >
> >I think there are three observations here:
> >
> >(1) removing the mark_page_accessed (which is the only significant
> >    change in the parent commit)  hurts the
> >    aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
> >    I'd still rather stick to the filemap version and let the
> >    VM people sort it out.  How do the numbers for this test
> >    look for XFS vs say ext4 and btrfs?
> >(2) lots of additional spinlock contention in the new case.  A quick
> >    check shows that I fat-fingered my rewrite so that we do
> >    the xfs_inode_set_eofblocks_tag call now for the pure lookup
> >    case, and pretty much all new cycles come from that.
> >(3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
> >    we're already doing way to many even without my little bug above.
> >
> >So I've force pushed a new version of the iomap-fixes branch with
> >(2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
> >lot less expensive slotted in before that.  Would be good to see
> >the numbers with that.
> 
> The aim7 1BRD tests finished and there are ups and downs, with overall
> performance remain flat.
> 
> 99091700659f4df9  74a242ad94d13436a1644c0b45  bf4dc6e4ecc2a3d042029319bc  testcase/testparams/testbox
> ----------------  --------------------------  --------------------------  ---------------------------

What do these commits refer to, please? They mean nothing without
the commit names....

/me goes searching. Ok:

99091700659 is the top of Linus' tree
74a242ad94d is ????
bf4dc6e4ecc is the latest in Christoph's tree (because it's
		mentioned below)

>         %stddev     %change         %stddev     %change         %stddev
>             \          |                \          |
> \     159926                      157324                      158574
> GEO-MEAN aim7.jobs-per-min
>     70897               5%      74137               4%      73775        aim7/1BRD_48G-xfs-creat-clo-1500-performance/ivb44
>    485217 ±  3%                492431                      477533        aim7/1BRD_48G-xfs-disk_rd-9000-performance/ivb44
>    360451             -19%     292980             -17%     299377        aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44

So, why does random read go backwards by 20%? The iomap IO path
patches we are testing only affect the write path, so this
doesn't make a whole lot of sense.

>    338114                      338410               5%     354078        aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
>     60130 ±  5%         4%      62438               5%      62923        aim7/1BRD_48G-xfs-disk_src-3000-performance/ivb44
>    403144                      397790                      410648        aim7/1BRD_48G-xfs-disk_wrt-3000-performance/ivb44

And this is the test the original regression was reported for:

gcc-6/performance/profile/1BRD_48G/xfs/x86_64-rhel/3000/debian-x86_64-2015-02-07.cgz/ivb44/disk_wrt/aim7

And that shows no improvement at all. The orginal regression was:

	484435 ±  0%     -13.3%     420004 ±  0%  aim7.jobs-per-min

So it's still 15% down on the orginal performance which, again,
doesn't make a whole lot of sense given the improvement in so many
other tests I've run....

>     26327                       26534                       26128        aim7/1BRD_48G-xfs-sync_disk_rw-600-performance/ivb44
> 
> The new commit bf4dc6e ("xfs: rewrite and optimize the delalloc write
> path") improves the aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
> case by 5%. Here are the detailed numbers:
> 
> aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44

Not important at all. We need the results for the disk_wrt regression
we are chasing (disk_wrt-3000) so we can see how the code change
affected behaviour.

> Here are the detailed numbers for the slowed down case:
> 
> aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
> 
> 99091700659f4df9  bf4dc6e4ecc2a3d042029319bc
> ----------------  --------------------------
>         %stddev      change         %stddev
>             \          |                \
>    360451             -17%     299377        aim7.jobs-per-min
>     12806             481%      74447        aim7.time.involuntary_context_switches
.....
>     19377             459%     108364        vmstat.system.cs
.....
>       487 ± 89%      3e+04      26448 ± 57%  latency_stats.max.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
>      1823 ± 82%      2e+06    1913796 ± 38%  latency_stats.sum.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
>    208475 ± 43%      1e+06    1409494 ±  5%  latency_stats.sum.wait_on_page_bit.truncate_inode_pages_range.truncate_inode_pages_final.evict.iput.dentry_unlink_inode.__dentry_kill.dput.__fput.____fput.task_work_run.exit_to_usermode_loop
>      6884 ± 73%      8e+04      90790 ±  9%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.file_update_time.xfs_file_aio_write_checks.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write.SyS_write
>      1598 ± 20%      3e+04      35015 ± 27%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_free_eofblocks.xfs_release.xfs_file_release.__fput.____fput.task_work_run
>      2006 ± 25%      3e+04      31143 ± 35%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode.evict.iput
>        29 ±101%      1e+04      10214 ± 29%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_defer_trans_roll.xfs_defer_finish.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode
>      1206 ± 51%      9e+03       9919 ± 25%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.touch_atime.generic_file_read_iter.xfs_file_buffered_aio_read.xfs_file_read_iter.__vfs_read.vfs_read.SyS_read

Significant increase in blocking delays in the journal during atime
updates. There's nothing in Christoph's tree that would affect that
behaviour.  This smells like either a mount option change or
individual tests not being 100% isolated and the previous test run
is affecting this one?

-Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1463755 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromFengguang Wu <fengguang.wu@intel.com>
Date2016-08-16 14:30 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6LFf-30r-17@gated-at.bofh.it>
In reply to#1463187
On Tue, Aug 16, 2016 at 07:22:40AM +1000, Dave Chinner wrote:
>On Mon, Aug 15, 2016 at 10:14:55PM +0800, Fengguang Wu wrote:
>> Hi Christoph,
>>
>> On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
>> >Snipping the long contest:
>> >
>> >I think there are three observations here:
>> >
>> >(1) removing the mark_page_accessed (which is the only significant
>> >    change in the parent commit)  hurts the
>> >    aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
>> >    I'd still rather stick to the filemap version and let the
>> >    VM people sort it out.  How do the numbers for this test
>> >    look for XFS vs say ext4 and btrfs?
>> >(2) lots of additional spinlock contention in the new case.  A quick
>> >    check shows that I fat-fingered my rewrite so that we do
>> >    the xfs_inode_set_eofblocks_tag call now for the pure lookup
>> >    case, and pretty much all new cycles come from that.
>> >(3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
>> >    we're already doing way to many even without my little bug above.
>> >
>> >So I've force pushed a new version of the iomap-fixes branch with
>> >(2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
>> >lot less expensive slotted in before that.  Would be good to see
>> >the numbers with that.
>>
>> The aim7 1BRD tests finished and there are ups and downs, with overall
>> performance remain flat.
>>
>> 99091700659f4df9  74a242ad94d13436a1644c0b45  bf4dc6e4ecc2a3d042029319bc  testcase/testparams/testbox
>> ----------------  --------------------------  --------------------------  ---------------------------
>
>What do these commits refer to, please? They mean nothing without
>the commit names....
>
>/me goes searching. Ok:
>
>99091700659 is the top of Linus' tree
>74a242ad94d is ????

That's the below one's parent commit, 74a242ad94d ("xfs: make
xfs_inode_set_eofblocks_tag cheaper for the common case").

Typically we'll compare a commit with its parent commit, and/or
the branch's base commit, which is normally on mainline kernel.

>bf4dc6e4ecc is the latest in Christoph's tree (because it's
>		mentioned below)
>
>>         %stddev     %change         %stddev     %change         %stddev
>>             \          |                \          |
>> \     159926                      157324                      158574
>> GEO-MEAN aim7.jobs-per-min
>>     70897               5%      74137               4%      73775        aim7/1BRD_48G-xfs-creat-clo-1500-performance/ivb44
>>    485217 ±  3%                492431                      477533        aim7/1BRD_48G-xfs-disk_rd-9000-performance/ivb44
>>    360451             -19%     292980             -17%     299377        aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
>
>So, why does random read go backwards by 20%? The iomap IO path
>patches we are testing only affect the write path, so this
>doesn't make a whole lot of sense.
>
>>    338114                      338410               5%     354078        aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
>>     60130 ±  5%         4%      62438               5%      62923        aim7/1BRD_48G-xfs-disk_src-3000-performance/ivb44
>>    403144                      397790                      410648        aim7/1BRD_48G-xfs-disk_wrt-3000-performance/ivb44
>
>And this is the test the original regression was reported for:
>
>gcc-6/performance/profile/1BRD_48G/xfs/x86_64-rhel/3000/debian-x86_64-2015-02-07.cgz/ivb44/disk_wrt/aim7
>
>And that shows no improvement at all. The orginal regression was:
>
>	484435 ±  0%     -13.3%     420004 ±  0%  aim7.jobs-per-min
>
>So it's still 15% down on the orginal performance which, again,
>doesn't make a whole lot of sense given the improvement in so many
>other tests I've run....

Yes, same performance with 4.8-rc1 means the regression is still not
back comparing to the original reported first-bad-commit's parent
f0c6bcba74ac51cb ("xfs: reorder zeroing and flushing sequence in
truncate") which is on 4.7-rc1. 

>>     26327                       26534                       26128        aim7/1BRD_48G-xfs-sync_disk_rw-600-performance/ivb44
>>
>> The new commit bf4dc6e ("xfs: rewrite and optimize the delalloc write
>> path") improves the aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
>> case by 5%. Here are the detailed numbers:
>>
>> aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
>
>Not important at all. We need the results for the disk_wrt regression
>we are chasing (disk_wrt-3000) so we can see how the code change
>affected behaviour.

Yeah it may not relevant to this case study, however should help
evaluate the patch in a more complete way.

>> Here are the detailed numbers for the slowed down case:
>>
>> aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
>>
>> 99091700659f4df9  bf4dc6e4ecc2a3d042029319bc
>> ----------------  --------------------------
>>         %stddev      change         %stddev
>>             \          |                \
>>    360451             -17%     299377        aim7.jobs-per-min
>>     12806             481%      74447        aim7.time.involuntary_context_switches
>.....
>>     19377             459%     108364        vmstat.system.cs
>.....
>>       487 ± 89%      3e+04      26448 ± 57%  latency_stats.max.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
>>      1823 ± 82%      2e+06    1913796 ± 38%  latency_stats.sum.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
>>    208475 ± 43%      1e+06    1409494 ±  5%  latency_stats.sum.wait_on_page_bit.truncate_inode_pages_range.truncate_inode_pages_final.evict.iput.dentry_unlink_inode.__dentry_kill.dput.__fput.____fput.task_work_run.exit_to_usermode_loop
>>      6884 ± 73%      8e+04      90790 ±  9%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.file_update_time.xfs_file_aio_write_checks.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write.SyS_write
>>      1598 ± 20%      3e+04      35015 ± 27%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_free_eofblocks.xfs_release.xfs_file_release.__fput.____fput.task_work_run
>>      2006 ± 25%      3e+04      31143 ± 35%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode.evict.iput
>>        29 ±101%      1e+04      10214 ± 29%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_defer_trans_roll.xfs_defer_finish.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode
>>      1206 ± 51%      9e+03       9919 ± 25%  latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.touch_atime.generic_file_read_iter.xfs_file_buffered_aio_read.xfs_file_read_iter.__vfs_read.vfs_read.SyS_read
>
>Significant increase in blocking delays in the journal during atime
>updates. There's nothing in Christoph's tree that would affect that
>behaviour.  This smells like either a mount option change or
>individual tests not being 100% isolated and the previous test run
>is affecting this one?

We kexec reboot machines between tests to make sure zero influence
from previous test. The test jobs are queued in a batch and not
likely to change mount option etc. in between (just confirmed).

The kernels are build by a random build server and some builds will
reuse previous .o files (no distclean). To make sure I rebuilt the
kernels in the same build server with distclean. However new tests
still show the same numbers.

Thanks,
Fengguang

[toc] | [prev] | [next] | [standalone]


#1463158 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

From"Huang\, Ying" <ying.huang@intel.com>
Date2016-08-15 22:40 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6wPT-1T2-9@gated-at.bofh.it>
In reply to#1462166
Christoph Hellwig <hch@lst.de> writes:

> Snipping the long contest:
>
> I think there are three observations here:
>
>  (1) removing the mark_page_accessed (which is the only significant
>      change in the parent commit)  hurts the
>      aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
>      I'd still rather stick to the filemap version and let the
>      VM people sort it out.  How do the numbers for this test
>      look for XFS vs say ext4 and btrfs?
>  (2) lots of additional spinlock contention in the new case.  A quick
>      check shows that I fat-fingered my rewrite so that we do
>      the xfs_inode_set_eofblocks_tag call now for the pure lookup
>      case, and pretty much all new cycles come from that.
>  (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
>      we're already doing way to many even without my little bug above.
>
> So I've force pushed a new version of the iomap-fixes branch with
> (2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
> lot less expensive slotted in before that.  Would be good to see
> the numbers with that.

For the original reported regression, the test result is as follow,

=========================================================================================
compiler/cpufreq_governor/debug-setup/disk/fs/kconfig/load/rootfs/tbox_group/test/testcase:
  gcc-6/performance/profile/1BRD_48G/xfs/x86_64-rhel/3000/debian-x86_64-2015-02-07.cgz/ivb44/disk_wrt/aim7

commit: 
  f0c6bcba74ac51cb77aadb33ad35cb2dc1ad1506 (parent of first bad commit)
  68a9f5e7007c1afa2cf6830b690a90d0187c0684 (first bad commit)
  99091700659f4df965e138b38b4fa26a29b7eade (base of your fixes branch)
  bf4dc6e4ecc2a3d042029319bc8cd4204c185610 (head of your fixes branch)

f0c6bcba74ac51cb 68a9f5e7007c1afa2cf6830b69 99091700659f4df965e138b38b bf4dc6e4ecc2a3d042029319bc 
---------------- -------------------------- -------------------------- -------------------------- 
         %stddev     %change         %stddev     %change         %stddev     %change         %stddev
             \          |                \          |                \          |                \  
    484435 ±  0%     -13.3%     420004 ±  0%     -17.0%     402250 ±  0%     -15.6%     408998 ±  0%  aim7.jobs-per-min


And the perf data is as follow,

  "perf-profile.func.cycles-pp.intel_idle": 20.25,
  "perf-profile.func.cycles-pp.memset_erms": 11.72,
  "perf-profile.func.cycles-pp.copy_user_enhanced_fast_string": 8.37,
  "perf-profile.func.cycles-pp.__block_commit_write.isra.21": 3.49,
  "perf-profile.func.cycles-pp.block_write_end": 1.77,
  "perf-profile.func.cycles-pp.native_queued_spin_lock_slowpath": 1.63,
  "perf-profile.func.cycles-pp.unlock_page": 1.58,
  "perf-profile.func.cycles-pp.___might_sleep": 1.56,
  "perf-profile.func.cycles-pp.__block_write_begin_int": 1.33,
  "perf-profile.func.cycles-pp.iov_iter_copy_from_user_atomic": 1.23,
  "perf-profile.func.cycles-pp.up_write": 1.21,
  "perf-profile.func.cycles-pp.__mark_inode_dirty": 1.18,
  "perf-profile.func.cycles-pp.down_write": 1.06,
  "perf-profile.func.cycles-pp.mark_buffer_dirty": 0.94,
  "perf-profile.func.cycles-pp.generic_write_end": 0.92,
  "perf-profile.func.cycles-pp.__radix_tree_lookup": 0.91,
  "perf-profile.func.cycles-pp._raw_spin_lock": 0.81,
  "perf-profile.func.cycles-pp.entry_SYSCALL_64_fastpath": 0.79,
  "perf-profile.func.cycles-pp.__might_sleep": 0.79,
  "perf-profile.func.cycles-pp.xfs_file_iomap_begin_delay.isra.9": 0.7,
  "perf-profile.func.cycles-pp.__list_del_entry": 0.7,
  "perf-profile.func.cycles-pp.vfs_write": 0.69,
  "perf-profile.func.cycles-pp.drop_buffers": 0.68,
  "perf-profile.func.cycles-pp.xfs_file_write_iter": 0.67,
  "perf-profile.func.cycles-pp.rwsem_spin_on_owner": 0.67,

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1468096 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

From"Huang\, Ying" <ying.huang@intel.com>
Date2016-08-23 00:10 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s95zR-3HK-77@gated-at.bofh.it>
In reply to#1463158
Hi, Christoph,

"Huang, Ying" <ying.huang@intel.com> writes:

> Christoph Hellwig <hch@lst.de> writes:
>
>> Snipping the long contest:
>>
>> I think there are three observations here:
>>
>>  (1) removing the mark_page_accessed (which is the only significant
>>      change in the parent commit)  hurts the
>>      aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
>>      I'd still rather stick to the filemap version and let the
>>      VM people sort it out.  How do the numbers for this test
>>      look for XFS vs say ext4 and btrfs?
>>  (2) lots of additional spinlock contention in the new case.  A quick
>>      check shows that I fat-fingered my rewrite so that we do
>>      the xfs_inode_set_eofblocks_tag call now for the pure lookup
>>      case, and pretty much all new cycles come from that.
>>  (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
>>      we're already doing way to many even without my little bug above.
>>
>> So I've force pushed a new version of the iomap-fixes branch with
>> (2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
>> lot less expensive slotted in before that.  Would be good to see
>> the numbers with that.
>
> For the original reported regression, the test result is as follow,
>
> =========================================================================================
> compiler/cpufreq_governor/debug-setup/disk/fs/kconfig/load/rootfs/tbox_group/test/testcase:
>   gcc-6/performance/profile/1BRD_48G/xfs/x86_64-rhel/3000/debian-x86_64-2015-02-07.cgz/ivb44/disk_wrt/aim7
>
> commit: 
>   f0c6bcba74ac51cb77aadb33ad35cb2dc1ad1506 (parent of first bad commit)
>   68a9f5e7007c1afa2cf6830b690a90d0187c0684 (first bad commit)
>   99091700659f4df965e138b38b4fa26a29b7eade (base of your fixes branch)
>   bf4dc6e4ecc2a3d042029319bc8cd4204c185610 (head of your fixes branch)
>
> f0c6bcba74ac51cb 68a9f5e7007c1afa2cf6830b69 99091700659f4df965e138b38b bf4dc6e4ecc2a3d042029319bc 
> ---------------- -------------------------- -------------------------- -------------------------- 
>          %stddev     %change         %stddev     %change         %stddev     %change         %stddev
>              \          |                \          |                \          |                \  
>     484435 ±  0%     -13.3%     420004 ±  0%     -17.0%     402250 ±  0%     -15.6%     408998 ±  0%  aim7.jobs-per-min

It appears the original reported regression hasn't bee resolved by your
commit.  Could you take a look at the test results and the perf data?

Best Regards,
Huang, Ying

>
> And the perf data is as follow,
>
>   "perf-profile.func.cycles-pp.intel_idle": 20.25,
>   "perf-profile.func.cycles-pp.memset_erms": 11.72,
>   "perf-profile.func.cycles-pp.copy_user_enhanced_fast_string": 8.37,
>   "perf-profile.func.cycles-pp.__block_commit_write.isra.21": 3.49,
>   "perf-profile.func.cycles-pp.block_write_end": 1.77,
>   "perf-profile.func.cycles-pp.native_queued_spin_lock_slowpath": 1.63,
>   "perf-profile.func.cycles-pp.unlock_page": 1.58,
>   "perf-profile.func.cycles-pp.___might_sleep": 1.56,
>   "perf-profile.func.cycles-pp.__block_write_begin_int": 1.33,
>   "perf-profile.func.cycles-pp.iov_iter_copy_from_user_atomic": 1.23,
>   "perf-profile.func.cycles-pp.up_write": 1.21,
>   "perf-profile.func.cycles-pp.__mark_inode_dirty": 1.18,
>   "perf-profile.func.cycles-pp.down_write": 1.06,
>   "perf-profile.func.cycles-pp.mark_buffer_dirty": 0.94,
>   "perf-profile.func.cycles-pp.generic_write_end": 0.92,
>   "perf-profile.func.cycles-pp.__radix_tree_lookup": 0.91,
>   "perf-profile.func.cycles-pp._raw_spin_lock": 0.81,
>   "perf-profile.func.cycles-pp.entry_SYSCALL_64_fastpath": 0.79,
>   "perf-profile.func.cycles-pp.__might_sleep": 0.79,
>   "perf-profile.func.cycles-pp.xfs_file_iomap_begin_delay.isra.9": 0.7,
>   "perf-profile.func.cycles-pp.__list_del_entry": 0.7,
>   "perf-profile.func.cycles-pp.vfs_write": 0.69,
>   "perf-profile.func.cycles-pp.drop_buffers": 0.68,
>   "perf-profile.func.cycles-pp.xfs_file_write_iter": 0.67,
>   "perf-profile.func.cycles-pp.rwsem_spin_on_owner": 0.67,
>
> Best Regards,
> Huang, Ying
> _______________________________________________
> LKP mailing list
> LKP@lists.01.org
> https://lists.01.org/mailman/listinfo/lkp

[toc] | [prev] | [next] | [standalone]


#1463786 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromFengguang Wu <fengguang.wu@intel.com>
Date2016-08-16 15:30 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s6MBj-3Dm-13@gated-at.bofh.it>
In reply to#1462166
On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
>Snipping the long contest:
>
>I think there are three observations here:
>
> (1) removing the mark_page_accessed (which is the only significant
>     change in the parent commit)  hurts the
>     aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
>     I'd still rather stick to the filemap version and let the
>     VM people sort it out.  How do the numbers for this test
>     look for XFS vs say ext4 and btrfs?

Here is a basic comparison of the 3 filesystems based on 99091700 ("
Merge tag 'nfs-for-4.8-2' of git://git.linux-nfs.org/projects/trondmy/linux-nfs").

% compare -a -g 99091700659f4df965e138b38b4fa26a29b7eade -d fs xfs ext4 btrfs

             xfs                        ext4                       btrfs  testcase/testparams/testbox
----------------  --------------------------  --------------------------  ---------------------------
         %stddev      change         %stddev      change         %stddev
             \          |                \          |                \
    193335             -27%     141400            -100%          8        GEO-MEAN aim7.jobs-per-min
    267649 ±  3%       -51%     130085                                    aim7/1BRD_48G-disk_cp-3000-performance/ivb44
    485217 ±  3%       402%    2434088 ±  3%       350%    2184471 ±  4%  aim7/1BRD_48G-disk_rd-9000-performance/ivb44
    360286             -64%     130351                                    aim7/1BRD_48G-disk_rr-3000-performance/ivb44
    338114             -78%      73280                                    aim7/1BRD_48G-disk_rw-3000-performance/ivb44
     60130 ±  5%       361%     277035                                    aim7/1BRD_48G-disk_src-3000-performance/ivb44
    403144             -68%     127584                                    aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
     26327             -60%      10571                                    aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
      2652             -96%        118             -82%        468        GEO-MEAN fsmark.files_per_sec
       393 ±  4%        -6%        368 ±  3%        10%        433 ±  5%  fsmark/1x-1t-1BRD_48G-4M-40G-NoSync-performance/ivb44
       200              -4%        191              -7%        185 ±  6%  fsmark/1x-1t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
      1583 ±  3%       -29%       1130             -31%       1088        fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
     21363              59%      33958                                    fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
     11033                                         -17%       9117        fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
     11833                                          12%      13234        fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
   2381976            -100%       6598             -96%     100973        GEO-MEAN fsmark.app_overhead
    564520 ±  7%        21%     681192 ±  3%        63%     919364 ±  3%  fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
    860074 ±  5%       112%    1820590 ± 14%        47%    1262443 ±  3%  fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
  12232633             -18%   10085199                                    fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
   3143334                                         -11%    2784178        fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
   4107347                                         -21%    3248210        fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4

Thanks,
Fengguang
---

Some less important numbers.

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
      1314             222%       4225            -100%          2        GEO-MEAN aim7.time.system_time
      1491 ±  6%       302%       6004                                    aim7/1BRD_48G-disk_cp-3000-performance/ivb44
      4786 ±  3%       -89%        502 ±  7%       -87%        632 ±  7%  aim7/1BRD_48G-disk_rd-9000-performance/ivb44
       756             689%       5971                                    aim7/1BRD_48G-disk_rr-3000-performance/ivb44
       891            1146%      11108                                    aim7/1BRD_48G-disk_rw-3000-performance/ivb44
       940 ±  5%        70%       1598                                    aim7/1BRD_48G-disk_src-3000-performance/ivb44
       599             925%       6148                                    aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
      2496             390%      12225                                    aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
    154597             185%     440025                                    GEO-MEAN aim7.time.minor_page_faults
    156203             144%     381038                                    aim7/1BRD_48G-disk_cp-3000-performance/ivb44
    155952             132%     362294 ±  3%                              aim7/1BRD_48G-disk_rr-3000-performance/ivb44
    157550             266%     577044                                    aim7/1BRD_48G-disk_rw-3000-performance/ivb44
    153880             152%     387212                                    aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
    149531 ±  5%       258%     534809 ±  4%                              aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
     86.94              37%     119.10             -98%       1.59        GEO-MEAN aim7.time.elapsed_time
     67.56 ±  3%       105%     138.61                                    aim7/1BRD_48G-disk_cp-3000-performance/ivb44
    112.15 ±  3%       -80%      22.93 ±  3%       -77%      25.48 ±  4%  aim7/1BRD_48G-disk_rd-9000-performance/ivb44
     50.19             176%     138.32                                    aim7/1BRD_48G-disk_rr-3000-performance/ivb44
     53.46             360%     245.91                                    aim7/1BRD_48G-disk_rw-3000-performance/ivb44
    300.63 ±  5%       -78%      65.30                                    aim7/1BRD_48G-disk_src-3000-performance/ivb44
     44.88             215%     141.35                                    aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
    136.82             149%     340.59                                    aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
     22.55              27%      28.54                                    GEO-MEAN aim7.time.user_time
     18.74              46%      27.27                                    aim7/1BRD_48G-disk_cp-3000-performance/ivb44
     28.71              28%      36.88                                    aim7/1BRD_48G-disk_rr-3000-performance/ivb44
     29.59              38%      40.74                                    aim7/1BRD_48G-disk_rw-3000-performance/ivb44
     41.42 ±  4%       -50%      20.90                                    aim7/1BRD_48G-disk_src-3000-performance/ivb44
     10.93              61%      17.61                                    aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
     18.26              96%      35.85                                    aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
   1171009             -44%     660859            -100%          4        GEO-MEAN aim7.time.voluntary_context_switches
    325355              -7%     303228                                    aim7/1BRD_48G-disk_cp-3000-performance/ivb44
     58321 ±  8%       -48%      30407             -44%      32487 ±  3%  aim7/1BRD_48G-disk_rd-9000-performance/ivb44
    437880             -37%     275709                                    aim7/1BRD_48G-disk_rr-3000-performance/ivb44
    395047              31%     518201                                    aim7/1BRD_48G-disk_rw-3000-performance/ivb44
  31067301             -93%    2034955 ±  5%                              aim7/1BRD_48G-disk_src-3000-performance/ivb44
    506749             -38%     315597                                    aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
  58429810              11%   65070475                                    aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
     41445             658%     314065            -100%          4        GEO-MEAN aim7.time.involuntary_context_switches
     21627 ±  5%      3118%     695989                                    aim7/1BRD_48G-disk_cp-3000-performance/ivb44
    594383 ±  5%       -96%      25928 ±  7%       -94%      38082 ± 11%  aim7/1BRD_48G-disk_rd-9000-performance/ivb44
     12980            5128%     678629                                    aim7/1BRD_48G-disk_rr-3000-performance/ivb44
     14729           10572%    1571856                                    aim7/1BRD_48G-disk_rw-3000-performance/ivb44
      5894 ±  3%       249%      20595 ±  8%                              aim7/1BRD_48G-disk_src-3000-performance/ivb44
      8950            7842%     710852                                    aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
   1620035             -34%    1069492                                    aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
    132117            -100%        216             -99%        990        GEO-MEAN fsmark.time.involuntary_context_switches
     14651             -99%         86 ±  3%       -99%        120 ±  9%  fsmark/1x-1t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
       553             413%       2840 ±  3%      8078%      45278 ± 31%  fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
     19895 ±  3%       -33%      13242             487%     116776 ±  3%  fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
   6206551             -99%      31906                                    fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
   1992225                                        -100%       1236 ±  5%  fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
   2664982                                        -100%       1202 ± 10%  fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
   1755485            -100%       3248             -93%     120902        GEO-MEAN fsmark.time.voluntary_context_switches
    542900             -98%      10270             -98%      10291        fsmark/1x-1t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
     36634 ±  6%       -17%      30317            3893%    1462650 ±  9%  fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
    162300 ± 15%       -55%      72808              89%     306208        fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
  58499255             -11%   51789210                                    fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
  10792062                                         168%   28969245        fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
  14361647                                          63%   23390858        fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
       391             -98%          7             -88%         47        GEO-MEAN fsmark.time.elapsed_time
       591             -37%        371                                    fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
       285                                          21%        345        fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
       354                                         -11%        317        fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
    294.98             -88%      36.23             -51%     145.31        GEO-MEAN fsmark.time.system_time
     45.16 ±  5%       157%     116.15             820%     415.73 ±  7%  fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
    262.79 ±  4%        29%     338.49              21%     316.76 ±  4%  fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
   1419.35              12%    1587.76                                    fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
    320.95                                         124%     719.99        fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
    413.11                                          65%     683.20        fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
       222             -68%         70             -27%        161        GEO-MEAN fsmark.time.percent_of_cpu_this_job_got
        91 ±  4%         9%         99               7%         97        fsmark/1x-1t-1BRD_48G-4M-40G-NoSync-performance/ivb44
        95               3%         98               3%         98        fsmark/1x-1t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
       214 ±  3%       152%        540             807%       1940 ±  9%  fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
      3938              -7%       3668             -17%       3286        fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
       253              75%        443                                    fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
       118                                          80%        213        fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
       123                                          80%        221        fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
     15893             -96%        649              46%      23186        GEO-MEAN fsmark.time.minor_page_faults
     10532              34%      14137             125%      23697 ±  5%  fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
     17053              14%      19388               9%      18557        fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
     22352 ± 34%                                    27%      28346 ± 28%  fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4

             xfs                        ext4                       btrfs
----------------  --------------------------  --------------------------
  30857819            -100%       5089             405%  1.559e+08        GEO-MEAN fsmark.time.file_system_outputs
   6400000 ± 50%        25%    8000000            2051%  1.377e+08        fsmark/1x-1t-1BRD_32G-4K-4G-fsyncBeforeClose-1fpd-performance/ivb43
  83886080                    83886080                    85633962        fsmark/1x-1t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
  50331648                                         352%  2.277e+08        fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
  33554432                                         555%  2.199e+08        fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4

[toc] | [prev] | [next] | [standalone]


#1461759 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-14 12:00 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s60n0-6g3-15@gated-at.bofh.it>
In reply to#1461546
On Sat, Aug 13, 2016 at 02:30:54AM +0200, Christoph Hellwig wrote:
> On Fri, Aug 12, 2016 at 08:02:08PM +1000, Dave Chinner wrote:
> > Which says "no change". Oh well, back to the drawing board...
> 
> I don't see how it would change thing much - for all relevant calculations
> we convert to block units first anyway.

THere was definitely an off-by-one in the code, which meant for
1-byte writes it never triggered speculative prealloc, so it was
doing the past-EOF real block check for every write. With it also
passing less than a block size, when the > XFS_ISIZE check passed
3 out of every 4 want_preallocate checks were landing on an already
allocated block, too, so it was doing 3x as many lookups as needed.
for 1k writes on a 4k block size filesystem. Amongst other things...

> But the whole xfs_iomap_write_delay is a giant mess anyway.  For a usual
> call we do at least four lookups in the extent btree, which seems rather
> costly.  Especially given that the low-level xfs_bmap_search_extents
> interface would give us all required information in one single call.

I noticed, though I was looking for a smaller, targetted fix rather
than rewriting the whole thing. Don't get me wrong, I think it needs
a rewrite to be efficient for the iomap infrastructure, just didn't
want to do that as a regression fix if a 1-liner might be
sufficient...

> Below is a patch I hacked up this morning to do just that.  It passes
> xfstests, but I've not done any real benchmarking with it.  If the
> reduced lookup overhead in it doesn't help enough we'll need to some
> sort of look aside cache for the information, but I hope that we
> can avoid that.  And yes, it's a rather large patch - but the old
> path was so entangled that I couldn't come up with something lighter.

I'll run some tests on it. If it does so;ve the regression, I'm
going to hold it back until we get a decent amount of review and
test coverage on it, though...

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1460127 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-11 03:20 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s4MP7-4rA-9@gated-at.bofh.it>
In reply to#1460115
On Wed, Aug 10, 2016 at 05:33:20PM -0700, Huang, Ying wrote:
> Linus Torvalds <torvalds@linux-foundation.org> writes:
> 
> > On Wed, Aug 10, 2016 at 5:11 PM, Huang, Ying <ying.huang@intel.com> wrote:
> >>
> >> Here is the comparison result with perf-profile data.
> >
> > Heh. The diff is actually harder to read than just showing A/B
> > state.The fact that the call chain shows up as part of the symbol
> > makes it even more so.
> >
> > For example:
> >
> >>       0.00 ± -1%      +Inf%       1.68 ±  1%  perf-profile.cycles-pp.__add_to_page_cache_locked.add_to_page_cache_lru.pagecache_get_page.grab_cache_page_write_begin.iomap_write_begin
> >>       1.80 ±  1%    -100.0%       0.00 ± -1%  perf-profile.cycles-pp.__add_to_page_cache_locked.add_to_page_cache_lru.pagecache_get_page.grab_cache_page_write_begin.xfs_vm_write_begin
> >
> > Ok, so it went from 1.8% to 1.68%, and isn't actually that big of a
> > change, but it shows up as a big change because the caller changed
> > from xfs_vm_write_begin to iomap_write_begin.
> >
> > There's a few other cases of that too.
> >
> > So I think it would actually be easier to just see "what 20 functions
> > were the hottest" (or maybe 50) before and after separately (just
> > sorted by cycles), without the diff part. Because the diff is really
> > hard to read.
> 
> Here it is,
> 
> Before:
> 
>   "perf-profile.func.cycles-pp.intel_idle": 16.88,
>   "perf-profile.func.cycles-pp.copy_user_enhanced_fast_string": 3.94,
>   "perf-profile.func.cycles-pp.memset_erms": 3.26,
>   "perf-profile.func.cycles-pp.__block_commit_write.isra.24": 2.47,
>   "perf-profile.func.cycles-pp.___might_sleep": 2.33,
>   "perf-profile.func.cycles-pp.__mark_inode_dirty": 1.88,
>   "perf-profile.func.cycles-pp.unlock_page": 1.69,
>   "perf-profile.func.cycles-pp.up_write": 1.61,
>   "perf-profile.func.cycles-pp.__block_write_begin_int": 1.56,
>   "perf-profile.func.cycles-pp.down_write": 1.55,
>   "perf-profile.func.cycles-pp.mark_buffer_dirty": 1.53,
>   "perf-profile.func.cycles-pp.entry_SYSCALL_64_fastpath": 1.47,
>   "perf-profile.func.cycles-pp.generic_write_end": 1.36,
>   "perf-profile.func.cycles-pp.generic_perform_write": 1.33,
>   "perf-profile.func.cycles-pp.__radix_tree_lookup": 1.32,
>   "perf-profile.func.cycles-pp.__might_sleep": 1.26,
>   "perf-profile.func.cycles-pp._raw_spin_lock": 1.17,
>   "perf-profile.func.cycles-pp.vfs_write": 1.14,
>   "perf-profile.func.cycles-pp.__xfs_get_blocks": 1.07,

Ok, so that is the old block mapping call in the buffered IO path.
I don't see any of the functions it calls in the profile;
specifically xfs_bmapi_read(), and xfs_iomap_write_delay(), so it
appears the extent mapping and allocation overhead on the old code
totals somewhere under 2-3% of the entire CPU usage.

>   "perf-profile.func.cycles-pp.xfs_file_write_iter": 1.03,
>   "perf-profile.func.cycles-pp.pagecache_get_page": 1.03,
>   "perf-profile.func.cycles-pp.native_queued_spin_lock_slowpath": 0.98,
>   "perf-profile.func.cycles-pp.get_page_from_freelist": 0.94,
>   "perf-profile.func.cycles-pp.rwsem_spin_on_owner": 0.94,
>   "perf-profile.func.cycles-pp.__vfs_write": 0.87,
>   "perf-profile.func.cycles-pp.iov_iter_copy_from_user_atomic": 0.87,
>   "perf-profile.func.cycles-pp.xfs_file_buffered_aio_write": 0.84,
>   "perf-profile.func.cycles-pp.find_get_entry": 0.79,
>   "perf-profile.func.cycles-pp._raw_spin_lock_irqsave": 0.78,
> 
> 
> After:
> 
>   "perf-profile.func.cycles-pp.intel_idle": 16.82,
>   "perf-profile.func.cycles-pp.copy_user_enhanced_fast_string": 3.27,
>   "perf-profile.func.cycles-pp.memset_erms": 2.6,
>   "perf-profile.func.cycles-pp.xfs_bmapi_read": 2.24,

Straight away - thats' at least 3x more overhead block mapping lookups
with the iomap code.

>   "perf-profile.func.cycles-pp.___might_sleep": 2.04,
>   "perf-profile.func.cycles-pp.mark_page_accessed": 1.93,
>   "perf-profile.func.cycles-pp.__block_write_begin_int": 1.78,
>   "perf-profile.func.cycles-pp.up_write": 1.72,
>   "perf-profile.func.cycles-pp.xfs_iext_bno_to_ext": 1.7,

Plus this child.

>   "perf-profile.func.cycles-pp.__block_commit_write.isra.24": 1.65,
>   "perf-profile.func.cycles-pp.down_write": 1.51,
>   "perf-profile.func.cycles-pp.__mark_inode_dirty": 1.51,
>   "perf-profile.func.cycles-pp.unlock_page": 1.43,
>   "perf-profile.func.cycles-pp.xfs_bmap_search_multi_extents": 1.25,
>   "perf-profile.func.cycles-pp.xfs_bmap_search_extents": 1.23,

And these two.

>   "perf-profile.func.cycles-pp.mark_buffer_dirty": 1.21,
>   "perf-profile.func.cycles-pp.xfs_iomap_write_delay": 1.19,
>   "perf-profile.func.cycles-pp.xfs_iomap_eof_want_preallocate.constprop.8": 1.15,

And these two.

So, essentially the old code had maybe 2-3% cpu usage overhead in
the block mapping path on this workload, but the new code is, for
some reason, showing at least 8-9% CPU usage overhead. That, right
now, makes no sense at all to me as we should be doing - at worst -
exactly the same number of block mapping calls as the old code.

We need to know what is happening that is different - there's a good
chance the mapping trace events will tell us. Huang, can you get
a raw event trace from the test?

I need to see these events:

	xfs_file*
	xfs_iomap*
	xfs_get_block*

For both kernels. An example trace from 4.8-rc1 running the command
`xfs_io -f -c 'pwrite 0 512k -b 128k' /mnt/scratch/fooey doing an
overwrite and extend of the existing file ends up looking like:

$ sudo trace-cmd start -e xfs_iomap\* -e xfs_file\* -e xfs_get_blocks\*
$ sudo cat /sys/kernel/tracing/trace_pipe
           <...>-2946  [001] .... 253971.750304: xfs_file_ioctl: dev 253:32 ino 0x84
          xfs_io-2946  [001] .... 253971.750938: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 0x20000
          xfs_io-2946  [001] .... 253971.750961: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
          xfs_io-2946  [001] .... 253971.751114: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 0x20000
          xfs_io-2946  [001] .... 253971.751128: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
          xfs_io-2946  [001] .... 253971.751234: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 0x20000
          xfs_io-2946  [001] .... 253971.751236: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
          xfs_io-2946  [001] .... 253971.751381: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 0x20000
          xfs_io-2946  [001] .... 253971.751415: xfs_iomap_prealloc_size: dev 253:32 ino 0x84 prealloc blocks 128 shift 0 m_writeio_blocks 16
          xfs_io-2946  [001] .... 253971.751425: xfs_iomap_alloc: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 131072 type invalid startoff 0x60 startblock -1 blockcount 0x90

That's the output I need for the complete test - you'll need to use
a better recording mechanism that this (e.g. trace-cmd record,
trace-cmd report) because it will generate a lot of events. Compress
the two report files (they'll be large) and send them to me offlist.

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1460130 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-11 03:40 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s4N8t-4zK-1@gated-at.bofh.it>
In reply to#1460127
On Thu, Aug 11, 2016 at 11:16:12AM +1000, Dave Chinner wrote:
> I need to see these events:
> 
> 	xfs_file*
> 	xfs_iomap*
> 	xfs_get_block*
> 
> For both kernels. An example trace from 4.8-rc1 running the command
> `xfs_io -f -c 'pwrite 0 512k -b 128k' /mnt/scratch/fooey doing an
> overwrite and extend of the existing file ends up looking like:
> 
> $ sudo trace-cmd start -e xfs_iomap\* -e xfs_file\* -e xfs_get_blocks\*
> $ sudo cat /sys/kernel/tracing/trace_pipe
>            <...>-2946  [001] .... 253971.750304: xfs_file_ioctl: dev 253:32 ino 0x84
>           xfs_io-2946  [001] .... 253971.750938: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 0x20000
>           xfs_io-2946  [001] .... 253971.750961: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
>           xfs_io-2946  [001] .... 253971.751114: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 0x20000
>           xfs_io-2946  [001] .... 253971.751128: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
>           xfs_io-2946  [001] .... 253971.751234: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 0x20000
>           xfs_io-2946  [001] .... 253971.751236: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
>           xfs_io-2946  [001] .... 253971.751381: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 0x20000
>           xfs_io-2946  [001] .... 253971.751415: xfs_iomap_prealloc_size: dev 253:32 ino 0x84 prealloc blocks 128 shift 0 m_writeio_blocks 16
>           xfs_io-2946  [001] .... 253971.751425: xfs_iomap_alloc: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 131072 type invalid startoff 0x60 startblock -1 blockcount 0x90
> 
> That's the output I need for the complete test - you'll need to use
> a better recording mechanism that this (e.g. trace-cmd record,
> trace-cmd report) because it will generate a lot of events. Compress
> the two report files (they'll be large) and send them to me offlist.

Can you also send me the output of xfs_info on the filesystem you
are testing?

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1460147 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromYe Xiaolong <xiaolong.ye@intel.com>
Date2016-08-11 04:50 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s4Oee-5jA-3@gated-at.bofh.it>
In reply to#1460130
On 08/11, Dave Chinner wrote:
>On Thu, Aug 11, 2016 at 11:16:12AM +1000, Dave Chinner wrote:
>> I need to see these events:
>> 
>> 	xfs_file*
>> 	xfs_iomap*
>> 	xfs_get_block*
>> 
>> For both kernels. An example trace from 4.8-rc1 running the command
>> `xfs_io -f -c 'pwrite 0 512k -b 128k' /mnt/scratch/fooey doing an
>> overwrite and extend of the existing file ends up looking like:
>> 
>> $ sudo trace-cmd start -e xfs_iomap\* -e xfs_file\* -e xfs_get_blocks\*
>> $ sudo cat /sys/kernel/tracing/trace_pipe
>>            <...>-2946  [001] .... 253971.750304: xfs_file_ioctl: dev 253:32 ino 0x84
>>           xfs_io-2946  [001] .... 253971.750938: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 0x20000
>>           xfs_io-2946  [001] .... 253971.750961: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
>>           xfs_io-2946  [001] .... 253971.751114: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 0x20000
>>           xfs_io-2946  [001] .... 253971.751128: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
>>           xfs_io-2946  [001] .... 253971.751234: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 0x20000
>>           xfs_io-2946  [001] .... 253971.751236: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
>>           xfs_io-2946  [001] .... 253971.751381: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 0x20000
>>           xfs_io-2946  [001] .... 253971.751415: xfs_iomap_prealloc_size: dev 253:32 ino 0x84 prealloc blocks 128 shift 0 m_writeio_blocks 16
>>           xfs_io-2946  [001] .... 253971.751425: xfs_iomap_alloc: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 131072 type invalid startoff 0x60 startblock -1 blockcount 0x90
>> 
>> That's the output I need for the complete test - you'll need to use
>> a better recording mechanism that this (e.g. trace-cmd record,
>> trace-cmd report) because it will generate a lot of events. Compress
>> the two report files (they'll be large) and send them to me offlist.
>
>Can you also send me the output of xfs_info on the filesystem you
>are testing?

Hi, Dave

Here is the xfs_info output:

# xfs_info /fs/ram0/
meta-data=/dev/ram0              isize=256    agcount=4, agsize=3145728 blks
         =                       sectsz=4096  attr=2, projid32bit=1
         =                       crc=0        finobt=0
data     =                       bsize=4096   blocks=12582912, imaxpct=25
         =                       sunit=0      swidth=0 blks
naming   =version 2              bsize=4096   ascii-ci=0 ftype=0
log      =internal               bsize=4096   blocks=6144, version=2
         =                       sectsz=4096  sunit=1 blks, lazy-count=1
realtime =none                   extsz=4096   blocks=0, rtextents=0

Thanks,
Xiaolong
>
>Cheers,
>
>Dave.
>-- 
>Dave Chinner
>david@fromorbit.com
>_______________________________________________
>LKP mailing list
>LKP@lists.01.org
>https://lists.01.org/mailman/listinfo/lkp

[toc] | [prev] | [next] | [standalone]


#1460150 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-11 05:20 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s4OHf-5JF-1@gated-at.bofh.it>
In reply to#1460147
On Thu, Aug 11, 2016 at 10:36:59AM +0800, Ye Xiaolong wrote:
> On 08/11, Dave Chinner wrote:
> >On Thu, Aug 11, 2016 at 11:16:12AM +1000, Dave Chinner wrote:
> >> I need to see these events:
> >> 
> >> 	xfs_file*
> >> 	xfs_iomap*
> >> 	xfs_get_block*
> >> 
> >> For both kernels. An example trace from 4.8-rc1 running the command
> >> `xfs_io -f -c 'pwrite 0 512k -b 128k' /mnt/scratch/fooey doing an
> >> overwrite and extend of the existing file ends up looking like:
> >> 
> >> $ sudo trace-cmd start -e xfs_iomap\* -e xfs_file\* -e xfs_get_blocks\*
> >> $ sudo cat /sys/kernel/tracing/trace_pipe
> >>            <...>-2946  [001] .... 253971.750304: xfs_file_ioctl: dev 253:32 ino 0x84
> >>           xfs_io-2946  [001] .... 253971.750938: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 0x20000
> >>           xfs_io-2946  [001] .... 253971.750961: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
> >>           xfs_io-2946  [001] .... 253971.751114: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 0x20000
> >>           xfs_io-2946  [001] .... 253971.751128: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
> >>           xfs_io-2946  [001] .... 253971.751234: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 0x20000
> >>           xfs_io-2946  [001] .... 253971.751236: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
> >>           xfs_io-2946  [001] .... 253971.751381: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 0x20000
> >>           xfs_io-2946  [001] .... 253971.751415: xfs_iomap_prealloc_size: dev 253:32 ino 0x84 prealloc blocks 128 shift 0 m_writeio_blocks 16
> >>           xfs_io-2946  [001] .... 253971.751425: xfs_iomap_alloc: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 131072 type invalid startoff 0x60 startblock -1 blockcount 0x90
> >> 
> >> That's the output I need for the complete test - you'll need to use
> >> a better recording mechanism that this (e.g. trace-cmd record,
> >> trace-cmd report) because it will generate a lot of events. Compress
> >> the two report files (they'll be large) and send them to me offlist.
> >
> >Can you also send me the output of xfs_info on the filesystem you
> >are testing?
> 
> Hi, Dave
> 
> Here is the xfs_info output:
> 
> # xfs_info /fs/ram0/
> meta-data=/dev/ram0              isize=256    agcount=4, agsize=3145728 blks
>          =                       sectsz=4096  attr=2, projid32bit=1
>          =                       crc=0        finobt=0
> data     =                       bsize=4096   blocks=12582912, imaxpct=25
>          =                       sunit=0      swidth=0 blks
> naming   =version 2              bsize=4096   ascii-ci=0 ftype=0
> log      =internal               bsize=4096   blocks=6144, version=2
>          =                       sectsz=4096  sunit=1 blks, lazy-count=1
> realtime =none                   extsz=4096   blocks=0, rtextents=0

OK, nothing unusual there. One thing that I did just think of - how
close to ENOSPC does this test get? i.e. are we hitting the "we're
almost out of free space" slow paths on this test?

Cheers,

dave.
> 
> Thanks,
> Xiaolong
> >
> >Cheers,
> >
> >Dave.
> >-- 
> >Dave Chinner
> >david@fromorbit.com
> >_______________________________________________
> >LKP mailing list
> >LKP@lists.01.org
> >https://lists.01.org/mailman/listinfo/lkp
> 

-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1460891 — Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression

FromDave Chinner <david@fromorbit.com>
Date2016-08-12 03:30 +0200
SubjectRe: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression
Message-ID<s59sl-3pL-7@gated-at.bofh.it>
In reply to#1460127
On Thu, Aug 11, 2016 at 11:16:12AM +1000, Dave Chinner wrote:
> On Wed, Aug 10, 2016 at 05:33:20PM -0700, Huang, Ying wrote:
> We need to know what is happening that is different - there's a good
> chance the mapping trace events will tell us. Huang, can you get
> a raw event trace from the test?
> 
> I need to see these events:
> 
> 	xfs_file*
> 	xfs_iomap*
> 	xfs_get_block*
> 

lkp-folks, can I please get these traces run and sent to me? I don't
have the time or patience to try to get aim7 running on my machines
- the build is full of hard-coded paths and libraries that aren't
provided by modern distros (e.g. it requires a static libaio.a!) and
it fails at the configure stage complaining that:

configure: error: C compiler cannot create executables

Which is a complete load of BS.

Hence I can't make progress until I have some way of understanding
what the IO pattern is that is generating the profile being
measured. So far I'm unable to do that with any of the tools I been
trying, hence I need the traces to work out what I'm missing...

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1460102

FromLinus Torvalds <torvalds@linux-foundation.org>
Date2016-08-11 02:00 +0200
Message-ID<s4LzH-3vK-9@gated-at.bofh.it>
In reply to#1460084
On Wed, Aug 10, 2016 at 4:08 PM, Dave Chinner <david@fromorbit.com> wrote:
>
> That, to me, says there's a change in lock contention behaviour in
> the workload (which we know aim7 is good at exposing). i.e. the
> iomap change shifted contention from a sleeping lock to a spinning
> lock, or maybe we now trigger optimistic spinning behaviour on a
> lock we previously didn't spin on at all.

Hmm. Possibly. I reacted to the lower cpu load number, but yeah, I
could easily imagine some locking primitive difference too.

> We really need instruction level perf profiles to understand
> this - I don't have a machine with this many cpu cores available
> locally, so I'm not sure I'm going to be able to make any progress
> tracking it down in the short term. Maybe the lkp team has more
> in-depth cpu usage profiles they can share?

Yeah, I've occasionally wanted to see some kind of "top-25 kernel
functions in the profile" thing. That said, when the load isn't all
that familiar, the profiles usually are not all that easy to make
sense of either. But comparing the before and after state might give
us clues.

Fengguang?

           Linus

[toc] | [prev] | [standalone]


Page 5 of 5 — ← Prev page 1 2 3 4 [5]

Back to top | Article view | linux.kernel


csiph-web