Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1459842 > unrolled thread
| Started by | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| First post | 2016-08-10 22:30 +0200 |
| Last post | 2016-08-11 02:00 +0200 |
| Articles | 19 on this page of 99 — 11 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-10 22:30 +0200
Re: [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 01:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 02:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 02:20 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 02:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 02:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 03:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 06:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-15 19:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Christoph Hellwig <hch@lst.de> - 2016-08-11 18:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 19:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-11 20:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 22:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Christoph Hellwig <hch@lst.de> - 2016-08-11 22:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 22:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Al Viro <viro@ZenIV.linux.org.uk> - 2016-08-12 00:20 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 00:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 23:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Christoph Hellwig <hch@lst.de> - 2016-08-12 00:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 03:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 04:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 06:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 20:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-14 11:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 02:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-15 03:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 04:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-15 05:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 07:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 00:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 00:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 01:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 01:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-16 17:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 20:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Michal Hocko <mhocko@kernel.org> - 2016-08-17 17:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Michal Hocko <mhocko@kernel.org> - 2016-08-17 18:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-17 17:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-18 02:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-18 09:20 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-18 15:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-19 04:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-19 04:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Michal Hocko <mhocko@kernel.org> - 2016-08-19 11:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-19 13:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-20 01:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-20 03:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-20 14:20 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-19 06:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Mel Gorman <mgorman@techsingularity.net> - 2016-08-19 17:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-24 17:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-18 04:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 02:20 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 03:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 04:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-17 00:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-17 01:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 01:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-16 02:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-16 02:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ingo Molnar <mingo@kernel.org> - 2016-08-15 07:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Peter Zijlstra <peterz@infradead.org> - 2016-08-17 18:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-15 15:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 04:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 04:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Christoph Hellwig <hch@lst.de> - 2016-08-12 05:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 05:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 06:20 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-12 07:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 08:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ye Xiaolong <xiaolong.ye@intel.com> - 2016-08-12 08:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ye Xiaolong <xiaolong.ye@intel.com> - 2016-08-12 11:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 12:10 +0200
Re: [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-12 12:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Christoph Hellwig <hch@lst.de> - 2016-08-13 02:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Christoph Hellwig <hch@lst.de> - 2016-08-14 10:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-14 10:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Christoph Hellwig <hch@lst.de> - 2016-08-14 11:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-14 11:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Christoph Hellwig <hch@lst.de> - 2016-08-14 18:20 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 01:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-15 02:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-15 16:20 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-15 23:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-16 14:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-15 22:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression "Huang\, Ying" <ying.huang@intel.com> - 2016-08-23 00:10 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Fengguang Wu <fengguang.wu@intel.com> - 2016-08-16 15:30 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-14 12:00 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 03:20 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 03:40 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Ye Xiaolong <xiaolong.ye@intel.com> - 2016-08-11 04:50 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-11 05:20 +0200
Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Dave Chinner <david@fromorbit.com> - 2016-08-12 03:30 +0200
Re: [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-11 02:00 +0200
Page 5 of 5 — ← Prev page 1 2 3 4 [5]
| From | Fengguang Wu <fengguang.wu@intel.com> |
|---|---|
| Date | 2016-08-14 10:40 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s5Z7z-5sr-21@gated-at.bofh.it> |
| In reply to | #1461692 |
Hi Christoph, On Sat, Aug 13, 2016 at 11:48:25PM +0200, Christoph Hellwig wrote: >On Sat, Aug 13, 2016 at 02:30:54AM +0200, Christoph Hellwig wrote: >> Below is a patch I hacked up this morning to do just that. It passes >> xfstests, but I've not done any real benchmarking with it. If the >> reduced lookup overhead in it doesn't help enough we'll need to some >> sort of look aside cache for the information, but I hope that we >> can avoid that. And yes, it's a rather large patch - but the old >> path was so entangled that I couldn't come up with something lighter. > >Hi Fengguang or Xiaolong, > >any chance to add this thread to a lkp run? Sure. To which base should I apply it? Or if you already pushed the git tree, I'll test your commit directly. Thanks, Fengguang
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@lst.de> |
|---|---|
| Date | 2016-08-14 11:00 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s5ZqW-5AD-41@gated-at.bofh.it> |
| In reply to | #1461698 |
Hi Fengguang, feel free to try this git tree: git://git.infradead.org/users/hch/vfs.git iomap-fixes
[toc] | [prev] | [next] | [standalone]
| From | Fengguang Wu <fengguang.wu@intel.com> |
|---|---|
| Date | 2016-08-14 11:30 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s5ZTX-63p-7@gated-at.bofh.it> |
| In reply to | #1461731 |
Hi Christoph,
On Sun, Aug 14, 2016 at 12:15:08AM +0200, Christoph Hellwig wrote:
>Hi Fengguang,
>
>feel free to try this git tree:
>
> git://git.infradead.org/users/hch/vfs.git iomap-fixes
I just queued some test jobs for it.
% queue -q vip -t ivb44 -b hch-vfs/iomap-fixes aim7-fs-1brd.yaml fs=xfs -r3 -k fe9c2c81ed073878768785a985295cbacc349e42 -k ca2edab2e1d8f30dda874b7f717c2d4664991e9b -k 99091700659f4df965e138b38b4fa26a29b7eade
That job file can be found here:
https://git.kernel.org/cgit/linux/kernel/git/wfg/lkp-tests.git/tree/jobs/aim7-fs-1brd.yaml
It specifies a matrix of the below atom tests:
wfg /c/lkp-tests% split-job jobs/aim7-fs-1brd.yaml -s 'fs: xfs'
jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_src-3000-performance.yaml
jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_rr-3000-performance.yaml
jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_rw-3000-performance.yaml
jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_cp-3000-performance.yaml
jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_wrt-3000-performance.yaml
jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-sync_disk_rw-600-performance.yaml
jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-creat-clo-1500-performance.yaml
jobs/aim7-fs-1brd.yaml => ./aim7-fs-1brd-1BRD_48G-xfs-disk_rd-9000-performance.yaml
If you see other suitable tests for this patch, feel free to drop me a
hint. I'v queued these jobs to the other machines to make them run in
parallel.
% queue -q vip -t ivb43 -b hch-vfs/iomap-fixes fsmark-stress-journal-1hdd.yaml fsmark-stress-journal-1brd.yaml fs=xfs -r3 -k fe9c2c81ed073878768785a985295cbacc349e42 -k ca2edab2e1d8f30dda874b7f717c2d4664991e9b -k 99091700659f4df965e138b38b4fa26a29b7eade
% queue -q vip -t ivb44 -b hch-vfs/iomap-fixes fsmark-generic-1brd.yaml dd-write-1hdd.yaml fsmark-generic-1hdd.yaml fs=xfs -r3 -k fe9c2c81ed073878768785a985295cbacc349e42 -k ca2edab2e1d8f30dda874b7f717c2d4664991e9b -k 99091700659f4df965e138b38b4fa26a29b7eade
% queue -q vip -t lkp-hsx02 -b hch-vfs/iomap-fixes fsmark-generic-brd-raid.yaml fs=xfs -r3 -k fe9c2c81ed073878768785a985[0/1710]349e42 -k ca2edab2e1d8f30dda874b7f717c2d4664991e9b -k 99091700659f4df965e138b38b4fa26a29b7eade
% queue -q vip -t lkp-hsw-ep4 -b hch-vfs/iomap-fixes fsmark-1ssd-nvme-small.yaml fs=xfs -r3 -k fe9c2c81ed073878768785a985295cbacc349e42 -k ca2edab2e1d8f30dda874b7f717c2d4664991e9b -k 99091700659f4df965e138b38b4fa26a29b7eade
Thanks,
Fengguang
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@lst.de> |
|---|---|
| Date | 2016-08-14 18:20 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s66iJ-1Ra-13@gated-at.bofh.it> |
| In reply to | #1461745 |
Snipping the long contest:
I think there are three observations here:
(1) removing the mark_page_accessed (which is the only significant
change in the parent commit) hurts the
aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
I'd still rather stick to the filemap version and let the
VM people sort it out. How do the numbers for this test
look for XFS vs say ext4 and btrfs?
(2) lots of additional spinlock contention in the new case. A quick
check shows that I fat-fingered my rewrite so that we do
the xfs_inode_set_eofblocks_tag call now for the pure lookup
case, and pretty much all new cycles come from that.
(3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
we're already doing way to many even without my little bug above.
So I've force pushed a new version of the iomap-fixes branch with
(2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
lot less expensive slotted in before that. Would be good to see
the numbers with that.
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-08-15 01:50 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s6dkd-6l9-1@gated-at.bofh.it> |
| In reply to | #1462166 |
On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote: > Snipping the long contest: > > I think there are three observations here: > > (1) removing the mark_page_accessed (which is the only significant > change in the parent commit) hurts the > aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test. > I'd still rather stick to the filemap version and let the > VM people sort it out. How do the numbers for this test > look for XFS vs say ext4 and btrfs? > (2) lots of additional spinlock contention in the new case. A quick > check shows that I fat-fingered my rewrite so that we do > the xfs_inode_set_eofblocks_tag call now for the pure lookup > case, and pretty much all new cycles come from that. > (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and > we're already doing way to many even without my little bug above. > > So I've force pushed a new version of the iomap-fixes branch with > (2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a > lot less expensive slotted in before that. Would be good to see > the numbers with that. With this new set of fixes, the 1byte write test runs ~30% faster on my test machine (130k writes/s vs 100k writes/s), and the 1k write on the pmem device runs about 10% faster (660MB/s vs 590MB/s). dbench numbers on the pmem device also go through the roof (they didn't show any regression to begin with) - 50% faster at 16 clients on a 16AG filesystem (5700MB/s vs 3800MB/s). The 10Mx4k file create fsmark workload I run (on the sparse 500TB XFS filesystem backed by a pair of SSDs) is giving the highest throughput *and* the lowest std dev I've ever recorded (55014.8+/-1.3e+04 files/s) and that shows in the runtime which also drops from 3m57s to 3m22s. So regardless of what aim7 results we get from these changes, I'll be merging them pending review and further testing... Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Fengguang Wu <fengguang.wu@intel.com> |
|---|---|
| Date | 2016-08-15 02:00 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s6dtY-6ox-7@gated-at.bofh.it> |
| In reply to | #1462166 |
Hi Christoph,
On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
>Snipping the long contest:
>
>I think there are three observations here:
>
> (1) removing the mark_page_accessed (which is the only significant
> change in the parent commit) hurts the
> aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
> I'd still rather stick to the filemap version and let the
> VM people sort it out. How do the numbers for this test
> look for XFS vs say ext4 and btrfs?
We'll be able to compare between filesystems when the tests for Linus'
patch finish.
> (2) lots of additional spinlock contention in the new case. A quick
> check shows that I fat-fingered my rewrite so that we do
> the xfs_inode_set_eofblocks_tag call now for the pure lookup
> case, and pretty much all new cycles come from that.
> (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
> we're already doing way to many even without my little bug above.
>
>So I've force pushed a new version of the iomap-fixes branch with
>(2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
>lot less expensive slotted in before that. Would be good to see
>the numbers with that.
I just queued these jobs. The comment-out ones will be submitted as
the 2nd stage when the 1st-round quick tests finish.
queue=(
queue
-q vip
--repeat-to 3
fs=xfs
perf-profile.delay=1
-b hch-vfs/iomap-fixes
-k bf4dc6e4ecc2a3d042029319bc8cd4204c185610
-k 74a242ad94d13436a1644c0b4586700e39871491
-k 99091700659f4df965e138b38b4fa26a29b7eade
)
"${queue[@]}" -t ivb44 aim7-fs-1brd.yaml
"${queue[@]}" -t ivb44 fsmark-generic-1brd.yaml
"${queue[@]}" -t ivb43 fsmark-stress-journal-1brd.yaml
"${queue[@]}" -t lkp-hsx02 fsmark-generic-brd-raid.yaml
"${queue[@]}" -t lkp-hsw-ep4 fsmark-1ssd-nvme-small.yaml
#"${queue[@]}" -t ivb43 fsmark-stress-journal-1hdd.yaml
#"${queue[@]}" -t ivb44 dd-write-1hdd.yaml fsmark-generic-1hdd.yaml
Thanks,
Fengguang
[toc] | [prev] | [next] | [standalone]
| From | Fengguang Wu <fengguang.wu@intel.com> |
|---|---|
| Date | 2016-08-15 16:20 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s6qU9-6EH-17@gated-at.bofh.it> |
| In reply to | #1462166 |
Hi Christoph,
On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
>Snipping the long contest:
>
>I think there are three observations here:
>
> (1) removing the mark_page_accessed (which is the only significant
> change in the parent commit) hurts the
> aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
> I'd still rather stick to the filemap version and let the
> VM people sort it out. How do the numbers for this test
> look for XFS vs say ext4 and btrfs?
> (2) lots of additional spinlock contention in the new case. A quick
> check shows that I fat-fingered my rewrite so that we do
> the xfs_inode_set_eofblocks_tag call now for the pure lookup
> case, and pretty much all new cycles come from that.
> (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
> we're already doing way to many even without my little bug above.
>
>So I've force pushed a new version of the iomap-fixes branch with
>(2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
>lot less expensive slotted in before that. Would be good to see
>the numbers with that.
The aim7 1BRD tests finished and there are ups and downs, with overall
performance remain flat.
99091700659f4df9 74a242ad94d13436a1644c0b45 bf4dc6e4ecc2a3d042029319bc testcase/testparams/testbox
---------------- -------------------------- -------------------------- ---------------------------
%stddev %change %stddev %change %stddev
\ | \ | \
159926 157324 158574 GEO-MEAN aim7.jobs-per-min
70897 5% 74137 4% 73775 aim7/1BRD_48G-xfs-creat-clo-1500-performance/ivb44
485217 ± 3% 492431 477533 aim7/1BRD_48G-xfs-disk_rd-9000-performance/ivb44
360451 -19% 292980 -17% 299377 aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
338114 338410 5% 354078 aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
60130 ± 5% 4% 62438 5% 62923 aim7/1BRD_48G-xfs-disk_src-3000-performance/ivb44
403144 397790 410648 aim7/1BRD_48G-xfs-disk_wrt-3000-performance/ivb44
26327 26534 26128 aim7/1BRD_48G-xfs-sync_disk_rw-600-performance/ivb44
The new commit bf4dc6e ("xfs: rewrite and optimize the delalloc write
path") improves the aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
case by 5%. Here are the detailed numbers:
aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
74a242ad94d13436 bf4dc6e4ecc2a3d042029319bc
---------------- --------------------------
%stddev %change %stddev
\ | \
338410 5% 354078 aim7.jobs-per-min
404390 8% 435117 aim7.time.voluntary_context_switches
2502 -4% 2396 aim7.time.maximum_resident_set_size
15018 -9% 13701 aim7.time.involuntary_context_switches
900 -11% 801 aim7.time.system_time
17432 11% 19365 vmstat.system.cs
47736 ± 19% -24% 36087 interrupts.CAL:Function_call_interrupts
2129646 31% 2790638 proc-vmstat.pgalloc_dma32
379503 13% 429384 numa-meminfo.node0.Dirty
15018 -9% 13701 time.involuntary_context_switches
900 -11% 801 time.system_time
1560 10% 1716 slabinfo.mnt_cache.active_objs
1560 10% 1716 slabinfo.mnt_cache.num_objs
61.53 -4 57.45 ± 4% perf-profile.cycles-pp.intel_idle.cpuidle_enter_state.cpuidle_enter.call_cpuidle.cpu_startup_entry
61.63 -4 57.55 ± 4% perf-profile.func.cycles-pp.intel_idle
1007188 ± 16% 156% 2577911 ± 6% numa-numastat.node0.numa_miss
9662857 ± 4% -13% 8420159 ± 3% numa-numastat.node0.numa_foreign
1008220 ± 16% 155% 2570630 ± 6% numa-numastat.node1.numa_foreign
9664033 ± 4% -13% 8413184 ± 3% numa-numastat.node1.numa_miss
26519887 ± 3% 18% 31322674 cpuidle.C1-IVT.time
122238 16% 142383 cpuidle.C1-IVT.usage
46548 11% 51645 cpuidle.C1E-IVT.usage
17253419 13% 19567582 cpuidle.C3-IVT.time
86847 13% 98333 cpuidle.C3-IVT.usage
482033 ± 12% 108% 1000665 ± 8% numa-vmstat.node0.numa_miss
94689 14% 107744 numa-vmstat.node0.nr_zone_write_pending
94677 14% 107718 numa-vmstat.node0.nr_dirty
3156643 ± 3% -20% 2527460 ± 3% numa-vmstat.node0.numa_foreign
429288 ± 12% 129% 983053 ± 8% numa-vmstat.node1.numa_foreign
3104193 ± 3% -19% 2510128 numa-vmstat.node1.numa_miss
6.43 ± 5% 51% 9.70 ± 11% turbostat.Pkg%pc2
0.30 28% 0.38 turbostat.CPU%c3
9.71 9.92 turbostat.RAMWatt
158 154 turbostat.PkgWatt
125 -3% 121 turbostat.CorWatt
1141 -6% 1078 turbostat.Avg_MHz
38.70 -6% 36.48 turbostat.%Busy
5.03 ± 11% -51% 2.46 ± 40% turbostat.Pkg%pc6
8.33 ± 48% 88% 15.67 ± 36% sched_debug.cfs_rq:/.runnable_load_avg.max
1947 ± 3% -12% 1710 ± 7% sched_debug.cfs_rq:/.spread0.stddev
1936 ± 3% -12% 1698 ± 8% sched_debug.cfs_rq:/.min_vruntime.stddev
2170 ± 10% -14% 1863 ± 6% sched_debug.cfs_rq:/.load_avg.max
220926 ± 18% 37% 303192 ± 5% sched_debug.cpu.avg_idle.stddev
0.06 ± 13% 357% 0.28 ± 23% sched_debug.rt_rq:/.rt_time.avg
0.37 ± 10% 240% 1.25 ± 15% sched_debug.rt_rq:/.rt_time.stddev
2.54 ± 10% 160% 6.59 ± 10% sched_debug.rt_rq:/.rt_time.max
0.32 ± 19% 29% 0.42 ± 10% perf-stat.dTLB-load-miss-rate
964727 7% 1028830 perf-stat.context-switches
176406 4% 184289 perf-stat.cpu-migrations
0.29 4% 0.30 perf-stat.branch-miss-rate
1.634e+09 1.673e+09 perf-stat.node-store-misses
23.60 23.99 perf-stat.node-store-miss-rate
40.01 40.57 perf-stat.cache-miss-rate
0.95 -8% 0.87 perf-stat.ipc
3.203e+12 -9% 2.928e+12 perf-stat.cpu-cycles
1.506e+09 -11% 1.345e+09 perf-stat.branch-misses
50.64 ± 13% -14% 43.45 ± 4% perf-stat.iTLB-load-miss-rate
5.285e+11 -14% 4.523e+11 perf-stat.branch-instructions
3.042e+12 -16% 2.551e+12 perf-stat.instructions
7.996e+11 -18% 6.584e+11 perf-stat.dTLB-loads
5.569e+11 ± 4% -18% 4.578e+11 perf-stat.dTLB-stores
Here are the detailed numbers for the slowed down case:
aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
99091700659f4df9 bf4dc6e4ecc2a3d042029319bc
---------------- --------------------------
%stddev change %stddev
\ | \
360451 -17% 299377 aim7.jobs-per-min
12806 481% 74447 aim7.time.involuntary_context_switches
755 44% 1086 aim7.time.system_time
50.17 20% 60.36 aim7.time.elapsed_time
50.17 20% 60.36 aim7.time.elapsed_time.max
438148 446012 aim7.time.voluntary_context_switches
37798 ± 16% 780% 332583 ± 8% interrupts.CAL:Function_call_interrupts
78.82 ± 5% 18% 93.35 ± 5% uptime.boot
2847 ± 7% 11% 3160 ± 7% uptime.idle
147490 ± 8% 34% 197261 ± 3% softirqs.RCU
648159 29% 839283 softirqs.TIMER
160830 10% 177144 softirqs.SCHED
3845352 ± 4% 91% 7349133 numa-numastat.node0.numa_miss
4686838 ± 5% 67% 7835640 numa-numastat.node0.numa_foreign
3848455 ± 4% 91% 7352436 numa-numastat.node1.numa_foreign
4689920 ± 5% 67% 7838734 numa-numastat.node1.numa_miss
50.17 20% 60.36 time.elapsed_time.max
12806 481% 74447 time.involuntary_context_switches
755 44% 1086 time.system_time
50.17 20% 60.36 time.elapsed_time
1563 18% 1846 time.percent_of_cpu_this_job_got
11699 ± 19% 3738% 449048 vmstat.io.bo
18836969 -16% 15789996 vmstat.memory.free
16 19% 19 vmstat.procs.r
19377 459% 108364 vmstat.system.cs
48255 11% 53537 vmstat.system.in
2357299 25% 2951384 meminfo.Inactive(file)
2366381 25% 2960468 meminfo.Inactive
1575292 -9% 1429971 meminfo.Cached
19342499 -17% 16100340 meminfo.MemFree
1057904 -20% 842987 meminfo.Dirty
1057 21% 1284 turbostat.Avg_MHz
35.78 21% 43.24 turbostat.%Busy
9.95 15% 11.47 turbostat.RAMWatt
74 ± 5% 10% 81 turbostat.CoreTmp
74 ± 4% 10% 81 turbostat.PkgTmp
118 8% 128 turbostat.CorWatt
151 7% 162 turbostat.PkgWatt
29.06 -23% 22.39 turbostat.CPU%c6
487 ± 89% 3e+04 26448 ± 57% latency_stats.max.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
1823 ± 82% 2e+06 1913796 ± 38% latency_stats.sum.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
208475 ± 43% 1e+06 1409494 ± 5% latency_stats.sum.wait_on_page_bit.truncate_inode_pages_range.truncate_inode_pages_final.evict.iput.dentry_unlink_inode.__dentry_kill.dput.__fput.____fput.task_work_run.exit_to_usermode_loop
6884 ± 73% 8e+04 90790 ± 9% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.file_update_time.xfs_file_aio_write_checks.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write.SyS_write
1598 ± 20% 3e+04 35015 ± 27% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_free_eofblocks.xfs_release.xfs_file_release.__fput.____fput.task_work_run
2006 ± 25% 3e+04 31143 ± 35% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode.evict.iput
29 ±101% 1e+04 10214 ± 29% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_defer_trans_roll.xfs_defer_finish.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode
1206 ± 51% 9e+03 9919 ± 25% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.touch_atime.generic_file_read_iter.xfs_file_buffered_aio_read.xfs_file_read_iter.__vfs_read.vfs_read.SyS_read
29869205 ± 4% -10% 26804569 cpuidle.C1-IVT.time
5737726 39% 7952214 cpuidle.C1E-IVT.time
51141 17% 59958 cpuidle.C1E-IVT.usage
18377551 37% 25176426 cpuidle.C3-IVT.time
96067 17% 112045 cpuidle.C3-IVT.usage
1806811 12% 2024041 cpuidle.C6-IVT.usage
1104420 ± 36% 204% 3361085 ± 27% cpuidle.POLL.time
281 ± 10% 20% 338 cpuidle.POLL.usage
5.61 ± 11% -0.5 5.12 ± 18% perf-profile.cycles-pp.irq_exit.smp_apic_timer_interrupt.apic_timer_interrupt.cpuidle_enter.call_cpuidle
5.85 ± 6% -0.8 5.06 ± 15% perf-profile.cycles-pp.hrtimer_interrupt.local_apic_timer_interrupt.smp_apic_timer_interrupt.apic_timer_interrupt.cpuidle_enter
6.32 ± 6% -0.9 5.42 ± 15% perf-profile.cycles-pp.local_apic_timer_interrupt.smp_apic_timer_interrupt.apic_timer_interrupt.cpuidle_enter.call_cpuidle
15.77 ± 8% -2 13.83 ± 17% perf-profile.cycles-pp.smp_apic_timer_interrupt.apic_timer_interrupt.cpuidle_enter.call_cpuidle.cpu_startup_entry
16.04 ± 8% -2 14.01 ± 15% perf-profile.cycles-pp.apic_timer_interrupt.cpuidle_enter.call_cpuidle.cpu_startup_entry.start_secondary
60.25 ± 4% -7 53.03 ± 7% perf-profile.cycles-pp.intel_idle.cpuidle_enter_state.cpuidle_enter.call_cpuidle.cpu_startup_entry
60.41 ± 4% -7 53.12 ± 7% perf-profile.func.cycles-pp.intel_idle
1174104 22% 1436859 numa-meminfo.node0.Inactive
1167471 22% 1428271 numa-meminfo.node0.Inactive(file)
770811 -9% 698147 numa-meminfo.node0.FilePages
20707294 -12% 18281509 ± 6% numa-meminfo.node0.Active
20613745 -12% 18180987 ± 6% numa-meminfo.node0.Active(file)
9676639 -17% 8003627 numa-meminfo.node0.MemFree
509906 -22% 396192 numa-meminfo.node0.Dirty
1189539 28% 1524697 numa-meminfo.node1.Inactive(file)
1191989 28% 1525194 numa-meminfo.node1.Inactive
804508 -10% 727067 numa-meminfo.node1.FilePages
9654540 -16% 8077810 numa-meminfo.node1.MemFree
547956 -19% 441933 numa-meminfo.node1.Dirty
396 ± 12% 485% 2320 ± 37% slabinfo.bio-1.num_objs
396 ± 12% 481% 2303 ± 37% slabinfo.bio-1.active_objs
73 140% 176 ± 14% slabinfo.kmalloc-128.active_slabs
73 140% 176 ± 14% slabinfo.kmalloc-128.num_slabs
4734 94% 9171 ± 11% slabinfo.kmalloc-128.num_objs
4734 88% 8917 ± 13% slabinfo.kmalloc-128.active_objs
16238 -10% 14552 ± 3% slabinfo.kmalloc-256.active_objs
17189 -13% 15033 ± 3% slabinfo.kmalloc-256.num_objs
20651 96% 40387 ± 17% slabinfo.radix_tree_node.active_objs
398 91% 761 ± 17% slabinfo.radix_tree_node.active_slabs
398 91% 761 ± 17% slabinfo.radix_tree_node.num_slabs
22313 91% 42650 ± 17% slabinfo.radix_tree_node.num_objs
32 638% 236 ± 28% slabinfo.xfs_efd_item.active_slabs
32 638% 236 ± 28% slabinfo.xfs_efd_item.num_slabs
1295 281% 4934 ± 23% slabinfo.xfs_efd_item.num_objs
1295 280% 4923 ± 23% slabinfo.xfs_efd_item.active_objs
1661 81% 3000 ± 42% slabinfo.xfs_log_ticket.num_objs
1661 78% 2952 ± 42% slabinfo.xfs_log_ticket.active_objs
2617 49% 3905 ± 30% slabinfo.xfs_trans.num_objs
2617 48% 3870 ± 31% slabinfo.xfs_trans.active_objs
1015933 567% 6779099 perf-stat.context-switches
4.864e+08 126% 1.101e+09 perf-stat.node-load-misses
1.179e+09 103% 2.399e+09 perf-stat.node-loads
0.06 ± 34% 92% 0.12 ± 11% perf-stat.dTLB-store-miss-rate
2.985e+08 ± 32% 86% 5.542e+08 ± 11% perf-stat.dTLB-store-misses
2.551e+09 ± 15% 81% 4.625e+09 ± 13% perf-stat.dTLB-load-misses
0.39 ± 14% 66% 0.65 ± 13% perf-stat.dTLB-load-miss-rate
1.26e+09 60% 2.019e+09 perf-stat.node-store-misses
46072661 ± 27% 49% 68472915 perf-stat.iTLB-loads
2.738e+12 ± 4% 43% 3.916e+12 perf-stat.cpu-cycles
21.48 32% 28.35 perf-stat.node-store-miss-rate
1.612e+10 ± 3% 28% 2.066e+10 perf-stat.cache-references
1.669e+09 ± 3% 24% 2.063e+09 perf-stat.branch-misses
6.816e+09 ± 3% 20% 8.179e+09 perf-stat.cache-misses
177699 18% 209145 perf-stat.cpu-migrations
0.39 13% 0.44 perf-stat.branch-miss-rate
4.606e+09 11% 5.102e+09 perf-stat.node-stores
4.329e+11 ± 4% 9% 4.727e+11 perf-stat.branch-instructions
6.458e+11 9% 7.046e+11 perf-stat.dTLB-loads
29.19 8% 31.45 perf-stat.node-load-miss-rate
286173 8% 308115 perf-stat.page-faults
286191 8% 308109 perf-stat.minor-faults
45084934 4% 47073719 perf-stat.iTLB-load-misses
42.28 -6% 39.58 perf-stat.cache-miss-rate
50.62 ± 16% -19% 40.75 perf-stat.iTLB-load-miss-rate
0.89 -28% 0.64 perf-stat.ipc
2 ± 36% 4e+07% 970191 proc-vmstat.pgrotated
150 ± 21% 1e+07% 15356485 ± 3% proc-vmstat.nr_vmscan_immediate_reclaim
76823 ± 35% 56899% 43788651 proc-vmstat.pgscan_direct
153407 ± 19% 4483% 7031431 proc-vmstat.nr_written
619699 ± 19% 4441% 28139689 proc-vmstat.pgpgout
5342421 1061% 62050709 proc-vmstat.pgactivate
47 ± 25% 354% 217 proc-vmstat.nr_pages_scanned
8542963 ± 3% 78% 15182914 proc-vmstat.numa_miss
8542963 ± 3% 78% 15182715 proc-vmstat.numa_foreign
2820568 31% 3699073 proc-vmstat.pgalloc_dma32
589234 25% 738160 proc-vmstat.nr_zone_inactive_file
589240 25% 738155 proc-vmstat.nr_inactive_file
61347830 13% 69522958 proc-vmstat.pgfree
393711 -9% 356981 proc-vmstat.nr_file_pages
4831749 -17% 4020131 proc-vmstat.nr_free_pages
61252784 -18% 50183773 proc-vmstat.pgrefill
61245420 -18% 50176301 proc-vmstat.pgdeactivate
264397 -20% 210222 proc-vmstat.nr_zone_write_pending
264367 -20% 210188 proc-vmstat.nr_dirty
60420248 -39% 36646178 proc-vmstat.pgscan_kswapd
60373976 -44% 33735064 proc-vmstat.pgsteal_kswapd
1753 -98% 43 ± 18% proc-vmstat.pageoutrun
1095 -98% 25 ± 17% proc-vmstat.kswapd_low_wmark_hit_quickly
656 ± 3% -98% 15 ± 24% proc-vmstat.kswapd_high_wmark_hit_quickly
0 1136221 numa-vmstat.node0.workingset_refault
0 1136221 numa-vmstat.node0.workingset_activate
23 ± 45% 1e+07% 2756907 numa-vmstat.node0.nr_vmscan_immediate_reclaim
37618 ± 24% 3234% 1254165 numa-vmstat.node0.nr_written
1346538 ± 4% 104% 2748439 numa-vmstat.node0.numa_miss
1577620 ± 5% 80% 2842882 numa-vmstat.node0.numa_foreign
291242 23% 357407 numa-vmstat.node0.nr_inactive_file
291237 23% 357390 numa-vmstat.node0.nr_zone_inactive_file
13961935 12% 15577331 numa-vmstat.node0.numa_local
13961938 12% 15577332 numa-vmstat.node0.numa_hit
39831 10% 43768 numa-vmstat.node0.nr_unevictable
39831 10% 43768 numa-vmstat.node0.nr_zone_unevictable
193467 -10% 174639 numa-vmstat.node0.nr_file_pages
5147212 -12% 4542321 ± 6% numa-vmstat.node0.nr_active_file
5147237 -12% 4542325 ± 6% numa-vmstat.node0.nr_zone_active_file
2426129 -17% 2008637 numa-vmstat.node0.nr_free_pages
128285 -23% 99206 numa-vmstat.node0.nr_zone_write_pending
128259 -23% 99183 numa-vmstat.node0.nr_dirty
0 1190594 numa-vmstat.node1.workingset_refault
0 1190594 numa-vmstat.node1.workingset_activate
21 ± 36% 1e+07% 3120425 ± 4% numa-vmstat.node1.nr_vmscan_immediate_reclaim
38541 ± 26% 3336% 1324185 numa-vmstat.node1.nr_written
1316819 ± 4% 105% 2699075 numa-vmstat.node1.numa_foreign
1547929 ± 4% 80% 2793491 numa-vmstat.node1.numa_miss
296714 28% 381124 numa-vmstat.node1.nr_zone_inactive_file
296714 28% 381123 numa-vmstat.node1.nr_inactive_file
14311131 10% 15750908 numa-vmstat.node1.numa_hit
14311130 10% 15750905 numa-vmstat.node1.numa_local
201164 -10% 181742 numa-vmstat.node1.nr_file_pages
2422825 -16% 2027750 numa-vmstat.node1.nr_free_pages
137069 -19% 110501 numa-vmstat.node1.nr_zone_write_pending
137069 -19% 110497 numa-vmstat.node1.nr_dirty
737 ± 29% 27349% 202387 sched_debug.cfs_rq:/.min_vruntime.min
3637 ± 20% 7919% 291675 sched_debug.cfs_rq:/.min_vruntime.avg
11.00 ± 44% 4892% 549.17 ± 9% sched_debug.cfs_rq:/.runnable_load_avg.max
2.12 ± 36% 4853% 105.12 ± 5% sched_debug.cfs_rq:/.runnable_load_avg.stddev
1885 ± 6% 4189% 80870 sched_debug.cfs_rq:/.min_vruntime.stddev
1896 ± 6% 4166% 80895 sched_debug.cfs_rq:/.spread0.stddev
10774 ± 13% 4113% 453925 sched_debug.cfs_rq:/.min_vruntime.max
1.02 ± 19% 2630% 27.72 ± 7% sched_debug.cfs_rq:/.runnable_load_avg.avg
63060 ± 45% 776% 552157 sched_debug.cfs_rq:/.load.max
14442 ± 21% 590% 99615 ± 14% sched_debug.cfs_rq:/.load.stddev
8397 ± 9% 309% 34370 ± 12% sched_debug.cfs_rq:/.load.avg
46.02 ± 24% 176% 126.96 ± 6% sched_debug.cfs_rq:/.util_avg.stddev
817 19% 974 ± 3% sched_debug.cfs_rq:/.util_avg.max
721 -17% 600 ± 3% sched_debug.cfs_rq:/.util_avg.avg
595 ± 11% -38% 371 ± 7% sched_debug.cfs_rq:/.util_avg.min
1484 ± 20% -47% 792 ± 5% sched_debug.cfs_rq:/.load_avg.min
1798 ± 4% -50% 903 ± 5% sched_debug.cfs_rq:/.load_avg.avg
322 ± 8% 7726% 25239 ± 8% sched_debug.cpu.nr_switches.min
969 7238% 71158 sched_debug.cpu.nr_switches.avg
2.23 ± 40% 4650% 106.14 ± 4% sched_debug.cpu.cpu_load[0].stddev
943 ± 4% 3475% 33730 ± 3% sched_debug.cpu.nr_switches.stddev
0.87 ± 25% 3057% 27.46 ± 7% sched_debug.cpu.cpu_load[0].avg
5.43 ± 13% 2232% 126.61 sched_debug.cpu.nr_uninterruptible.stddev
6131 ± 3% 2028% 130453 sched_debug.cpu.nr_switches.max
1.58 ± 29% 1852% 30.90 ± 4% sched_debug.cpu.cpu_load[4].avg
2.00 ± 49% 1422% 30.44 ± 5% sched_debug.cpu.cpu_load[3].avg
63060 ± 45% 1053% 726920 ± 32% sched_debug.cpu.load.max
21.25 ± 44% 777% 186.33 ± 7% sched_debug.cpu.nr_uninterruptible.max
14419 ± 21% 731% 119865 ± 31% sched_debug.cpu.load.stddev
3586 381% 17262 sched_debug.cpu.nr_load_updates.min
8286 ± 8% 364% 38414 ± 17% sched_debug.cpu.load.avg
5444 303% 21956 sched_debug.cpu.nr_load_updates.avg
1156 231% 3827 sched_debug.cpu.nr_load_updates.stddev
8603 ± 4% 222% 27662 sched_debug.cpu.nr_load_updates.max
1410 165% 3735 sched_debug.cpu.curr->pid.max
28742 ± 15% 120% 63101 ± 7% sched_debug.cpu.clock.min
28742 ± 15% 120% 63101 ± 7% sched_debug.cpu.clock_task.min
28748 ± 15% 120% 63107 ± 7% sched_debug.cpu.clock.avg
28748 ± 15% 120% 63107 ± 7% sched_debug.cpu.clock_task.avg
28751 ± 15% 120% 63113 ± 7% sched_debug.cpu.clock.max
28751 ± 15% 120% 63113 ± 7% sched_debug.cpu.clock_task.max
442 ± 11% 93% 854 ± 15% sched_debug.cpu.curr->pid.avg
618 ± 3% 72% 1065 ± 4% sched_debug.cpu.curr->pid.stddev
1.88 ± 11% 50% 2.83 ± 8% sched_debug.cpu.clock.stddev
1.88 ± 11% 50% 2.83 ± 8% sched_debug.cpu.clock_task.stddev
5.22 ± 9% -55% 2.34 ± 23% sched_debug.rt_rq:/.rt_time.max
0.85 -55% 0.38 ± 28% sched_debug.rt_rq:/.rt_time.stddev
0.17 -56% 0.07 ± 33% sched_debug.rt_rq:/.rt_time.avg
27633 ± 16% 124% 61980 ± 8% sched_debug.ktime
28745 ± 15% 120% 63102 ± 7% sched_debug.sched_clk
28745 ± 15% 120% 63102 ± 7% sched_debug.cpu_clk
Thanks,
Fengguang
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-08-15 23:30 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s6xCi-2w3-19@gated-at.bofh.it> |
| In reply to | #1462814 |
On Mon, Aug 15, 2016 at 10:14:55PM +0800, Fengguang Wu wrote:
> Hi Christoph,
>
> On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
> >Snipping the long contest:
> >
> >I think there are three observations here:
> >
> >(1) removing the mark_page_accessed (which is the only significant
> > change in the parent commit) hurts the
> > aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
> > I'd still rather stick to the filemap version and let the
> > VM people sort it out. How do the numbers for this test
> > look for XFS vs say ext4 and btrfs?
> >(2) lots of additional spinlock contention in the new case. A quick
> > check shows that I fat-fingered my rewrite so that we do
> > the xfs_inode_set_eofblocks_tag call now for the pure lookup
> > case, and pretty much all new cycles come from that.
> >(3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
> > we're already doing way to many even without my little bug above.
> >
> >So I've force pushed a new version of the iomap-fixes branch with
> >(2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
> >lot less expensive slotted in before that. Would be good to see
> >the numbers with that.
>
> The aim7 1BRD tests finished and there are ups and downs, with overall
> performance remain flat.
>
> 99091700659f4df9 74a242ad94d13436a1644c0b45 bf4dc6e4ecc2a3d042029319bc testcase/testparams/testbox
> ---------------- -------------------------- -------------------------- ---------------------------
What do these commits refer to, please? They mean nothing without
the commit names....
/me goes searching. Ok:
99091700659 is the top of Linus' tree
74a242ad94d is ????
bf4dc6e4ecc is the latest in Christoph's tree (because it's
mentioned below)
> %stddev %change %stddev %change %stddev
> \ | \ |
> \ 159926 157324 158574
> GEO-MEAN aim7.jobs-per-min
> 70897 5% 74137 4% 73775 aim7/1BRD_48G-xfs-creat-clo-1500-performance/ivb44
> 485217 ± 3% 492431 477533 aim7/1BRD_48G-xfs-disk_rd-9000-performance/ivb44
> 360451 -19% 292980 -17% 299377 aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
So, why does random read go backwards by 20%? The iomap IO path
patches we are testing only affect the write path, so this
doesn't make a whole lot of sense.
> 338114 338410 5% 354078 aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
> 60130 ± 5% 4% 62438 5% 62923 aim7/1BRD_48G-xfs-disk_src-3000-performance/ivb44
> 403144 397790 410648 aim7/1BRD_48G-xfs-disk_wrt-3000-performance/ivb44
And this is the test the original regression was reported for:
gcc-6/performance/profile/1BRD_48G/xfs/x86_64-rhel/3000/debian-x86_64-2015-02-07.cgz/ivb44/disk_wrt/aim7
And that shows no improvement at all. The orginal regression was:
484435 ± 0% -13.3% 420004 ± 0% aim7.jobs-per-min
So it's still 15% down on the orginal performance which, again,
doesn't make a whole lot of sense given the improvement in so many
other tests I've run....
> 26327 26534 26128 aim7/1BRD_48G-xfs-sync_disk_rw-600-performance/ivb44
>
> The new commit bf4dc6e ("xfs: rewrite and optimize the delalloc write
> path") improves the aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
> case by 5%. Here are the detailed numbers:
>
> aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
Not important at all. We need the results for the disk_wrt regression
we are chasing (disk_wrt-3000) so we can see how the code change
affected behaviour.
> Here are the detailed numbers for the slowed down case:
>
> aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
>
> 99091700659f4df9 bf4dc6e4ecc2a3d042029319bc
> ---------------- --------------------------
> %stddev change %stddev
> \ | \
> 360451 -17% 299377 aim7.jobs-per-min
> 12806 481% 74447 aim7.time.involuntary_context_switches
.....
> 19377 459% 108364 vmstat.system.cs
.....
> 487 ± 89% 3e+04 26448 ± 57% latency_stats.max.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
> 1823 ± 82% 2e+06 1913796 ± 38% latency_stats.sum.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
> 208475 ± 43% 1e+06 1409494 ± 5% latency_stats.sum.wait_on_page_bit.truncate_inode_pages_range.truncate_inode_pages_final.evict.iput.dentry_unlink_inode.__dentry_kill.dput.__fput.____fput.task_work_run.exit_to_usermode_loop
> 6884 ± 73% 8e+04 90790 ± 9% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.file_update_time.xfs_file_aio_write_checks.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write.SyS_write
> 1598 ± 20% 3e+04 35015 ± 27% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_free_eofblocks.xfs_release.xfs_file_release.__fput.____fput.task_work_run
> 2006 ± 25% 3e+04 31143 ± 35% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode.evict.iput
> 29 ±101% 1e+04 10214 ± 29% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_defer_trans_roll.xfs_defer_finish.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode
> 1206 ± 51% 9e+03 9919 ± 25% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.touch_atime.generic_file_read_iter.xfs_file_buffered_aio_read.xfs_file_read_iter.__vfs_read.vfs_read.SyS_read
Significant increase in blocking delays in the journal during atime
updates. There's nothing in Christoph's tree that would affect that
behaviour. This smells like either a mount option change or
individual tests not being 100% isolated and the previous test run
is affecting this one?
-Dave.
--
Dave Chinner
david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Fengguang Wu <fengguang.wu@intel.com> |
|---|---|
| Date | 2016-08-16 14:30 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s6LFf-30r-17@gated-at.bofh.it> |
| In reply to | #1463187 |
On Tue, Aug 16, 2016 at 07:22:40AM +1000, Dave Chinner wrote:
>On Mon, Aug 15, 2016 at 10:14:55PM +0800, Fengguang Wu wrote:
>> Hi Christoph,
>>
>> On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
>> >Snipping the long contest:
>> >
>> >I think there are three observations here:
>> >
>> >(1) removing the mark_page_accessed (which is the only significant
>> > change in the parent commit) hurts the
>> > aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
>> > I'd still rather stick to the filemap version and let the
>> > VM people sort it out. How do the numbers for this test
>> > look for XFS vs say ext4 and btrfs?
>> >(2) lots of additional spinlock contention in the new case. A quick
>> > check shows that I fat-fingered my rewrite so that we do
>> > the xfs_inode_set_eofblocks_tag call now for the pure lookup
>> > case, and pretty much all new cycles come from that.
>> >(3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
>> > we're already doing way to many even without my little bug above.
>> >
>> >So I've force pushed a new version of the iomap-fixes branch with
>> >(2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
>> >lot less expensive slotted in before that. Would be good to see
>> >the numbers with that.
>>
>> The aim7 1BRD tests finished and there are ups and downs, with overall
>> performance remain flat.
>>
>> 99091700659f4df9 74a242ad94d13436a1644c0b45 bf4dc6e4ecc2a3d042029319bc testcase/testparams/testbox
>> ---------------- -------------------------- -------------------------- ---------------------------
>
>What do these commits refer to, please? They mean nothing without
>the commit names....
>
>/me goes searching. Ok:
>
>99091700659 is the top of Linus' tree
>74a242ad94d is ????
That's the below one's parent commit, 74a242ad94d ("xfs: make
xfs_inode_set_eofblocks_tag cheaper for the common case").
Typically we'll compare a commit with its parent commit, and/or
the branch's base commit, which is normally on mainline kernel.
>bf4dc6e4ecc is the latest in Christoph's tree (because it's
> mentioned below)
>
>> %stddev %change %stddev %change %stddev
>> \ | \ |
>> \ 159926 157324 158574
>> GEO-MEAN aim7.jobs-per-min
>> 70897 5% 74137 4% 73775 aim7/1BRD_48G-xfs-creat-clo-1500-performance/ivb44
>> 485217 ± 3% 492431 477533 aim7/1BRD_48G-xfs-disk_rd-9000-performance/ivb44
>> 360451 -19% 292980 -17% 299377 aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
>
>So, why does random read go backwards by 20%? The iomap IO path
>patches we are testing only affect the write path, so this
>doesn't make a whole lot of sense.
>
>> 338114 338410 5% 354078 aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
>> 60130 ± 5% 4% 62438 5% 62923 aim7/1BRD_48G-xfs-disk_src-3000-performance/ivb44
>> 403144 397790 410648 aim7/1BRD_48G-xfs-disk_wrt-3000-performance/ivb44
>
>And this is the test the original regression was reported for:
>
>gcc-6/performance/profile/1BRD_48G/xfs/x86_64-rhel/3000/debian-x86_64-2015-02-07.cgz/ivb44/disk_wrt/aim7
>
>And that shows no improvement at all. The orginal regression was:
>
> 484435 ± 0% -13.3% 420004 ± 0% aim7.jobs-per-min
>
>So it's still 15% down on the orginal performance which, again,
>doesn't make a whole lot of sense given the improvement in so many
>other tests I've run....
Yes, same performance with 4.8-rc1 means the regression is still not
back comparing to the original reported first-bad-commit's parent
f0c6bcba74ac51cb ("xfs: reorder zeroing and flushing sequence in
truncate") which is on 4.7-rc1.
>> 26327 26534 26128 aim7/1BRD_48G-xfs-sync_disk_rw-600-performance/ivb44
>>
>> The new commit bf4dc6e ("xfs: rewrite and optimize the delalloc write
>> path") improves the aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
>> case by 5%. Here are the detailed numbers:
>>
>> aim7/1BRD_48G-xfs-disk_rw-3000-performance/ivb44
>
>Not important at all. We need the results for the disk_wrt regression
>we are chasing (disk_wrt-3000) so we can see how the code change
>affected behaviour.
Yeah it may not relevant to this case study, however should help
evaluate the patch in a more complete way.
>> Here are the detailed numbers for the slowed down case:
>>
>> aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44
>>
>> 99091700659f4df9 bf4dc6e4ecc2a3d042029319bc
>> ---------------- --------------------------
>> %stddev change %stddev
>> \ | \
>> 360451 -17% 299377 aim7.jobs-per-min
>> 12806 481% 74447 aim7.time.involuntary_context_switches
>.....
>> 19377 459% 108364 vmstat.system.cs
>.....
>> 487 ± 89% 3e+04 26448 ± 57% latency_stats.max.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
>> 1823 ± 82% 2e+06 1913796 ± 38% latency_stats.sum.down.xfs_buf_lock._xfs_buf_find.xfs_buf_get_map.xfs_buf_read_map.xfs_trans_read_buf_map.xfs_read_agf.xfs_alloc_read_agf.xfs_alloc_fix_freelist.xfs_free_extent_fix_freelist.xfs_free_extent.xfs_trans_free_extent
>> 208475 ± 43% 1e+06 1409494 ± 5% latency_stats.sum.wait_on_page_bit.truncate_inode_pages_range.truncate_inode_pages_final.evict.iput.dentry_unlink_inode.__dentry_kill.dput.__fput.____fput.task_work_run.exit_to_usermode_loop
>> 6884 ± 73% 8e+04 90790 ± 9% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.file_update_time.xfs_file_aio_write_checks.xfs_file_buffered_aio_write.xfs_file_write_iter.__vfs_write.vfs_write.SyS_write
>> 1598 ± 20% 3e+04 35015 ± 27% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_free_eofblocks.xfs_release.xfs_file_release.__fput.____fput.task_work_run
>> 2006 ± 25% 3e+04 31143 ± 35% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode.evict.iput
>> 29 ±101% 1e+04 10214 ± 29% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.__xfs_trans_roll.xfs_trans_roll.xfs_defer_trans_roll.xfs_defer_finish.xfs_itruncate_extents.xfs_inactive_truncate.xfs_inactive.xfs_fs_destroy_inode.destroy_inode
>> 1206 ± 51% 9e+03 9919 ± 25% latency_stats.sum.call_rwsem_down_read_failed.xfs_log_commit_cil.__xfs_trans_commit.xfs_trans_commit.xfs_vn_update_time.touch_atime.generic_file_read_iter.xfs_file_buffered_aio_read.xfs_file_read_iter.__vfs_read.vfs_read.SyS_read
>
>Significant increase in blocking delays in the journal during atime
>updates. There's nothing in Christoph's tree that would affect that
>behaviour. This smells like either a mount option change or
>individual tests not being 100% isolated and the previous test run
>is affecting this one?
We kexec reboot machines between tests to make sure zero influence
from previous test. The test jobs are queued in a batch and not
likely to change mount option etc. in between (just confirmed).
The kernels are build by a random build server and some builds will
reuse previous .o files (no distclean). To make sure I rebuilt the
kernels in the same build server with distclean. However new tests
still show the same numbers.
Thanks,
Fengguang
[toc] | [prev] | [next] | [standalone]
| From | "Huang\, Ying" <ying.huang@intel.com> |
|---|---|
| Date | 2016-08-15 22:40 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s6wPT-1T2-9@gated-at.bofh.it> |
| In reply to | #1462166 |
Christoph Hellwig <hch@lst.de> writes:
> Snipping the long contest:
>
> I think there are three observations here:
>
> (1) removing the mark_page_accessed (which is the only significant
> change in the parent commit) hurts the
> aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
> I'd still rather stick to the filemap version and let the
> VM people sort it out. How do the numbers for this test
> look for XFS vs say ext4 and btrfs?
> (2) lots of additional spinlock contention in the new case. A quick
> check shows that I fat-fingered my rewrite so that we do
> the xfs_inode_set_eofblocks_tag call now for the pure lookup
> case, and pretty much all new cycles come from that.
> (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and
> we're already doing way to many even without my little bug above.
>
> So I've force pushed a new version of the iomap-fixes branch with
> (2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a
> lot less expensive slotted in before that. Would be good to see
> the numbers with that.
For the original reported regression, the test result is as follow,
=========================================================================================
compiler/cpufreq_governor/debug-setup/disk/fs/kconfig/load/rootfs/tbox_group/test/testcase:
gcc-6/performance/profile/1BRD_48G/xfs/x86_64-rhel/3000/debian-x86_64-2015-02-07.cgz/ivb44/disk_wrt/aim7
commit:
f0c6bcba74ac51cb77aadb33ad35cb2dc1ad1506 (parent of first bad commit)
68a9f5e7007c1afa2cf6830b690a90d0187c0684 (first bad commit)
99091700659f4df965e138b38b4fa26a29b7eade (base of your fixes branch)
bf4dc6e4ecc2a3d042029319bc8cd4204c185610 (head of your fixes branch)
f0c6bcba74ac51cb 68a9f5e7007c1afa2cf6830b69 99091700659f4df965e138b38b bf4dc6e4ecc2a3d042029319bc
---------------- -------------------------- -------------------------- --------------------------
%stddev %change %stddev %change %stddev %change %stddev
\ | \ | \ | \
484435 ± 0% -13.3% 420004 ± 0% -17.0% 402250 ± 0% -15.6% 408998 ± 0% aim7.jobs-per-min
And the perf data is as follow,
"perf-profile.func.cycles-pp.intel_idle": 20.25,
"perf-profile.func.cycles-pp.memset_erms": 11.72,
"perf-profile.func.cycles-pp.copy_user_enhanced_fast_string": 8.37,
"perf-profile.func.cycles-pp.__block_commit_write.isra.21": 3.49,
"perf-profile.func.cycles-pp.block_write_end": 1.77,
"perf-profile.func.cycles-pp.native_queued_spin_lock_slowpath": 1.63,
"perf-profile.func.cycles-pp.unlock_page": 1.58,
"perf-profile.func.cycles-pp.___might_sleep": 1.56,
"perf-profile.func.cycles-pp.__block_write_begin_int": 1.33,
"perf-profile.func.cycles-pp.iov_iter_copy_from_user_atomic": 1.23,
"perf-profile.func.cycles-pp.up_write": 1.21,
"perf-profile.func.cycles-pp.__mark_inode_dirty": 1.18,
"perf-profile.func.cycles-pp.down_write": 1.06,
"perf-profile.func.cycles-pp.mark_buffer_dirty": 0.94,
"perf-profile.func.cycles-pp.generic_write_end": 0.92,
"perf-profile.func.cycles-pp.__radix_tree_lookup": 0.91,
"perf-profile.func.cycles-pp._raw_spin_lock": 0.81,
"perf-profile.func.cycles-pp.entry_SYSCALL_64_fastpath": 0.79,
"perf-profile.func.cycles-pp.__might_sleep": 0.79,
"perf-profile.func.cycles-pp.xfs_file_iomap_begin_delay.isra.9": 0.7,
"perf-profile.func.cycles-pp.__list_del_entry": 0.7,
"perf-profile.func.cycles-pp.vfs_write": 0.69,
"perf-profile.func.cycles-pp.drop_buffers": 0.68,
"perf-profile.func.cycles-pp.xfs_file_write_iter": 0.67,
"perf-profile.func.cycles-pp.rwsem_spin_on_owner": 0.67,
Best Regards,
Huang, Ying
[toc] | [prev] | [next] | [standalone]
| From | "Huang\, Ying" <ying.huang@intel.com> |
|---|---|
| Date | 2016-08-23 00:10 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s95zR-3HK-77@gated-at.bofh.it> |
| In reply to | #1463158 |
Hi, Christoph, "Huang, Ying" <ying.huang@intel.com> writes: > Christoph Hellwig <hch@lst.de> writes: > >> Snipping the long contest: >> >> I think there are three observations here: >> >> (1) removing the mark_page_accessed (which is the only significant >> change in the parent commit) hurts the >> aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test. >> I'd still rather stick to the filemap version and let the >> VM people sort it out. How do the numbers for this test >> look for XFS vs say ext4 and btrfs? >> (2) lots of additional spinlock contention in the new case. A quick >> check shows that I fat-fingered my rewrite so that we do >> the xfs_inode_set_eofblocks_tag call now for the pure lookup >> case, and pretty much all new cycles come from that. >> (3) Boy, are those xfs_inode_set_eofblocks_tag calls expensive, and >> we're already doing way to many even without my little bug above. >> >> So I've force pushed a new version of the iomap-fixes branch with >> (2) fixed, and also a little patch to xfs_inode_set_eofblocks_tag a >> lot less expensive slotted in before that. Would be good to see >> the numbers with that. > > For the original reported regression, the test result is as follow, > > ========================================================================================= > compiler/cpufreq_governor/debug-setup/disk/fs/kconfig/load/rootfs/tbox_group/test/testcase: > gcc-6/performance/profile/1BRD_48G/xfs/x86_64-rhel/3000/debian-x86_64-2015-02-07.cgz/ivb44/disk_wrt/aim7 > > commit: > f0c6bcba74ac51cb77aadb33ad35cb2dc1ad1506 (parent of first bad commit) > 68a9f5e7007c1afa2cf6830b690a90d0187c0684 (first bad commit) > 99091700659f4df965e138b38b4fa26a29b7eade (base of your fixes branch) > bf4dc6e4ecc2a3d042029319bc8cd4204c185610 (head of your fixes branch) > > f0c6bcba74ac51cb 68a9f5e7007c1afa2cf6830b69 99091700659f4df965e138b38b bf4dc6e4ecc2a3d042029319bc > ---------------- -------------------------- -------------------------- -------------------------- > %stddev %change %stddev %change %stddev %change %stddev > \ | \ | \ | \ > 484435 ± 0% -13.3% 420004 ± 0% -17.0% 402250 ± 0% -15.6% 408998 ± 0% aim7.jobs-per-min It appears the original reported regression hasn't bee resolved by your commit. Could you take a look at the test results and the perf data? Best Regards, Huang, Ying > > And the perf data is as follow, > > "perf-profile.func.cycles-pp.intel_idle": 20.25, > "perf-profile.func.cycles-pp.memset_erms": 11.72, > "perf-profile.func.cycles-pp.copy_user_enhanced_fast_string": 8.37, > "perf-profile.func.cycles-pp.__block_commit_write.isra.21": 3.49, > "perf-profile.func.cycles-pp.block_write_end": 1.77, > "perf-profile.func.cycles-pp.native_queued_spin_lock_slowpath": 1.63, > "perf-profile.func.cycles-pp.unlock_page": 1.58, > "perf-profile.func.cycles-pp.___might_sleep": 1.56, > "perf-profile.func.cycles-pp.__block_write_begin_int": 1.33, > "perf-profile.func.cycles-pp.iov_iter_copy_from_user_atomic": 1.23, > "perf-profile.func.cycles-pp.up_write": 1.21, > "perf-profile.func.cycles-pp.__mark_inode_dirty": 1.18, > "perf-profile.func.cycles-pp.down_write": 1.06, > "perf-profile.func.cycles-pp.mark_buffer_dirty": 0.94, > "perf-profile.func.cycles-pp.generic_write_end": 0.92, > "perf-profile.func.cycles-pp.__radix_tree_lookup": 0.91, > "perf-profile.func.cycles-pp._raw_spin_lock": 0.81, > "perf-profile.func.cycles-pp.entry_SYSCALL_64_fastpath": 0.79, > "perf-profile.func.cycles-pp.__might_sleep": 0.79, > "perf-profile.func.cycles-pp.xfs_file_iomap_begin_delay.isra.9": 0.7, > "perf-profile.func.cycles-pp.__list_del_entry": 0.7, > "perf-profile.func.cycles-pp.vfs_write": 0.69, > "perf-profile.func.cycles-pp.drop_buffers": 0.68, > "perf-profile.func.cycles-pp.xfs_file_write_iter": 0.67, > "perf-profile.func.cycles-pp.rwsem_spin_on_owner": 0.67, > > Best Regards, > Huang, Ying > _______________________________________________ > LKP mailing list > LKP@lists.01.org > https://lists.01.org/mailman/listinfo/lkp
[toc] | [prev] | [next] | [standalone]
| From | Fengguang Wu <fengguang.wu@intel.com> |
|---|---|
| Date | 2016-08-16 15:30 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s6MBj-3Dm-13@gated-at.bofh.it> |
| In reply to | #1462166 |
On Sun, Aug 14, 2016 at 06:17:24PM +0200, Christoph Hellwig wrote:
>Snipping the long contest:
>
>I think there are three observations here:
>
> (1) removing the mark_page_accessed (which is the only significant
> change in the parent commit) hurts the
> aim7/1BRD_48G-xfs-disk_rr-3000-performance/ivb44 test.
> I'd still rather stick to the filemap version and let the
> VM people sort it out. How do the numbers for this test
> look for XFS vs say ext4 and btrfs?
Here is a basic comparison of the 3 filesystems based on 99091700 ("
Merge tag 'nfs-for-4.8-2' of git://git.linux-nfs.org/projects/trondmy/linux-nfs").
% compare -a -g 99091700659f4df965e138b38b4fa26a29b7eade -d fs xfs ext4 btrfs
xfs ext4 btrfs testcase/testparams/testbox
---------------- -------------------------- -------------------------- ---------------------------
%stddev change %stddev change %stddev
\ | \ | \
193335 -27% 141400 -100% 8 GEO-MEAN aim7.jobs-per-min
267649 ± 3% -51% 130085 aim7/1BRD_48G-disk_cp-3000-performance/ivb44
485217 ± 3% 402% 2434088 ± 3% 350% 2184471 ± 4% aim7/1BRD_48G-disk_rd-9000-performance/ivb44
360286 -64% 130351 aim7/1BRD_48G-disk_rr-3000-performance/ivb44
338114 -78% 73280 aim7/1BRD_48G-disk_rw-3000-performance/ivb44
60130 ± 5% 361% 277035 aim7/1BRD_48G-disk_src-3000-performance/ivb44
403144 -68% 127584 aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
26327 -60% 10571 aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44
xfs ext4 btrfs
---------------- -------------------------- --------------------------
2652 -96% 118 -82% 468 GEO-MEAN fsmark.files_per_sec
393 ± 4% -6% 368 ± 3% 10% 433 ± 5% fsmark/1x-1t-1BRD_48G-4M-40G-NoSync-performance/ivb44
200 -4% 191 -7% 185 ± 6% fsmark/1x-1t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
1583 ± 3% -29% 1130 -31% 1088 fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
21363 59% 33958 fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
11033 -17% 9117 fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
11833 12% 13234 fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
xfs ext4 btrfs
---------------- -------------------------- --------------------------
2381976 -100% 6598 -96% 100973 GEO-MEAN fsmark.app_overhead
564520 ± 7% 21% 681192 ± 3% 63% 919364 ± 3% fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
860074 ± 5% 112% 1820590 ± 14% 47% 1262443 ± 3% fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
12232633 -18% 10085199 fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
3143334 -11% 2784178 fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
4107347 -21% 3248210 fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
Thanks,
Fengguang
---
Some less important numbers.
xfs ext4 btrfs
---------------- -------------------------- --------------------------
1314 222% 4225 -100% 2 GEO-MEAN aim7.time.system_time
1491 ± 6% 302% 6004 aim7/1BRD_48G-disk_cp-3000-performance/ivb44
4786 ± 3% -89% 502 ± 7% -87% 632 ± 7% aim7/1BRD_48G-disk_rd-9000-performance/ivb44
756 689% 5971 aim7/1BRD_48G-disk_rr-3000-performance/ivb44
891 1146% 11108 aim7/1BRD_48G-disk_rw-3000-performance/ivb44
940 ± 5% 70% 1598 aim7/1BRD_48G-disk_src-3000-performance/ivb44
599 925% 6148 aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
2496 390% 12225 aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44
xfs ext4 btrfs
---------------- -------------------------- --------------------------
154597 185% 440025 GEO-MEAN aim7.time.minor_page_faults
156203 144% 381038 aim7/1BRD_48G-disk_cp-3000-performance/ivb44
155952 132% 362294 ± 3% aim7/1BRD_48G-disk_rr-3000-performance/ivb44
157550 266% 577044 aim7/1BRD_48G-disk_rw-3000-performance/ivb44
153880 152% 387212 aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
149531 ± 5% 258% 534809 ± 4% aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44
xfs ext4 btrfs
---------------- -------------------------- --------------------------
86.94 37% 119.10 -98% 1.59 GEO-MEAN aim7.time.elapsed_time
67.56 ± 3% 105% 138.61 aim7/1BRD_48G-disk_cp-3000-performance/ivb44
112.15 ± 3% -80% 22.93 ± 3% -77% 25.48 ± 4% aim7/1BRD_48G-disk_rd-9000-performance/ivb44
50.19 176% 138.32 aim7/1BRD_48G-disk_rr-3000-performance/ivb44
53.46 360% 245.91 aim7/1BRD_48G-disk_rw-3000-performance/ivb44
300.63 ± 5% -78% 65.30 aim7/1BRD_48G-disk_src-3000-performance/ivb44
44.88 215% 141.35 aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
136.82 149% 340.59 aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44
xfs ext4 btrfs
---------------- -------------------------- --------------------------
22.55 27% 28.54 GEO-MEAN aim7.time.user_time
18.74 46% 27.27 aim7/1BRD_48G-disk_cp-3000-performance/ivb44
28.71 28% 36.88 aim7/1BRD_48G-disk_rr-3000-performance/ivb44
29.59 38% 40.74 aim7/1BRD_48G-disk_rw-3000-performance/ivb44
41.42 ± 4% -50% 20.90 aim7/1BRD_48G-disk_src-3000-performance/ivb44
10.93 61% 17.61 aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
18.26 96% 35.85 aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44
xfs ext4 btrfs
---------------- -------------------------- --------------------------
1171009 -44% 660859 -100% 4 GEO-MEAN aim7.time.voluntary_context_switches
325355 -7% 303228 aim7/1BRD_48G-disk_cp-3000-performance/ivb44
58321 ± 8% -48% 30407 -44% 32487 ± 3% aim7/1BRD_48G-disk_rd-9000-performance/ivb44
437880 -37% 275709 aim7/1BRD_48G-disk_rr-3000-performance/ivb44
395047 31% 518201 aim7/1BRD_48G-disk_rw-3000-performance/ivb44
31067301 -93% 2034955 ± 5% aim7/1BRD_48G-disk_src-3000-performance/ivb44
506749 -38% 315597 aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
58429810 11% 65070475 aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44
xfs ext4 btrfs
---------------- -------------------------- --------------------------
41445 658% 314065 -100% 4 GEO-MEAN aim7.time.involuntary_context_switches
21627 ± 5% 3118% 695989 aim7/1BRD_48G-disk_cp-3000-performance/ivb44
594383 ± 5% -96% 25928 ± 7% -94% 38082 ± 11% aim7/1BRD_48G-disk_rd-9000-performance/ivb44
12980 5128% 678629 aim7/1BRD_48G-disk_rr-3000-performance/ivb44
14729 10572% 1571856 aim7/1BRD_48G-disk_rw-3000-performance/ivb44
5894 ± 3% 249% 20595 ± 8% aim7/1BRD_48G-disk_src-3000-performance/ivb44
8950 7842% 710852 aim7/1BRD_48G-disk_wrt-3000-performance/ivb44
1620035 -34% 1069492 aim7/1BRD_48G-sync_disk_rw-600-performance/ivb44
xfs ext4 btrfs
---------------- -------------------------- --------------------------
132117 -100% 216 -99% 990 GEO-MEAN fsmark.time.involuntary_context_switches
14651 -99% 86 ± 3% -99% 120 ± 9% fsmark/1x-1t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
553 413% 2840 ± 3% 8078% 45278 ± 31% fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
19895 ± 3% -33% 13242 487% 116776 ± 3% fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
6206551 -99% 31906 fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
1992225 -100% 1236 ± 5% fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
2664982 -100% 1202 ± 10% fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
xfs ext4 btrfs
---------------- -------------------------- --------------------------
1755485 -100% 3248 -93% 120902 GEO-MEAN fsmark.time.voluntary_context_switches
542900 -98% 10270 -98% 10291 fsmark/1x-1t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
36634 ± 6% -17% 30317 3893% 1462650 ± 9% fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
162300 ± 15% -55% 72808 89% 306208 fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
58499255 -11% 51789210 fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
10792062 168% 28969245 fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
14361647 63% 23390858 fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
xfs ext4 btrfs
---------------- -------------------------- --------------------------
391 -98% 7 -88% 47 GEO-MEAN fsmark.time.elapsed_time
591 -37% 371 fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
285 21% 345 fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
354 -11% 317 fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
xfs ext4 btrfs
---------------- -------------------------- --------------------------
294.98 -88% 36.23 -51% 145.31 GEO-MEAN fsmark.time.system_time
45.16 ± 5% 157% 116.15 820% 415.73 ± 7% fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
262.79 ± 4% 29% 338.49 21% 316.76 ± 4% fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
1419.35 12% 1587.76 fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
320.95 124% 719.99 fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
413.11 65% 683.20 fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
xfs ext4 btrfs
---------------- -------------------------- --------------------------
222 -68% 70 -27% 161 GEO-MEAN fsmark.time.percent_of_cpu_this_job_got
91 ± 4% 9% 99 7% 97 fsmark/1x-1t-1BRD_48G-4M-40G-NoSync-performance/ivb44
95 3% 98 3% 98 fsmark/1x-1t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
214 ± 3% 152% 540 807% 1940 ± 9% fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
3938 -7% 3668 -17% 3286 fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
253 75% 443 fsmark/8-1SSD-16-9B-48G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
118 80% 213 fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
123 80% 221 fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
xfs ext4 btrfs
---------------- -------------------------- --------------------------
15893 -96% 649 46% 23186 GEO-MEAN fsmark.time.minor_page_faults
10532 34% 14137 125% 23697 ± 5% fsmark/1x-64t-1BRD_48G-4M-40G-NoSync-performance/ivb44
17053 14% 19388 9% 18557 fsmark/1x-64t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
22352 ± 34% 27% 28346 ± 28% fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
xfs ext4 btrfs
---------------- -------------------------- --------------------------
30857819 -100% 5089 405% 1.559e+08 GEO-MEAN fsmark.time.file_system_outputs
6400000 ± 50% 25% 8000000 2051% 1.377e+08 fsmark/1x-1t-1BRD_32G-4K-4G-fsyncBeforeClose-1fpd-performance/ivb43
83886080 83886080 85633962 fsmark/1x-1t-1BRD_48G-4M-40G-fsyncBeforeClose-performance/ivb44
50331648 352% 2.277e+08 fsmark/8-1SSD-4-8K-24G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
33554432 555% 2.199e+08 fsmark/8-1SSD-4-9B-16G-fsyncBeforeClose-16d-256fpd-performance/lkp-hsw-ep4
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-08-14 12:00 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s60n0-6g3-15@gated-at.bofh.it> |
| In reply to | #1461546 |
On Sat, Aug 13, 2016 at 02:30:54AM +0200, Christoph Hellwig wrote: > On Fri, Aug 12, 2016 at 08:02:08PM +1000, Dave Chinner wrote: > > Which says "no change". Oh well, back to the drawing board... > > I don't see how it would change thing much - for all relevant calculations > we convert to block units first anyway. THere was definitely an off-by-one in the code, which meant for 1-byte writes it never triggered speculative prealloc, so it was doing the past-EOF real block check for every write. With it also passing less than a block size, when the > XFS_ISIZE check passed 3 out of every 4 want_preallocate checks were landing on an already allocated block, too, so it was doing 3x as many lookups as needed. for 1k writes on a 4k block size filesystem. Amongst other things... > But the whole xfs_iomap_write_delay is a giant mess anyway. For a usual > call we do at least four lookups in the extent btree, which seems rather > costly. Especially given that the low-level xfs_bmap_search_extents > interface would give us all required information in one single call. I noticed, though I was looking for a smaller, targetted fix rather than rewriting the whole thing. Don't get me wrong, I think it needs a rewrite to be efficient for the iomap infrastructure, just didn't want to do that as a regression fix if a 1-liner might be sufficient... > Below is a patch I hacked up this morning to do just that. It passes > xfstests, but I've not done any real benchmarking with it. If the > reduced lookup overhead in it doesn't help enough we'll need to some > sort of look aside cache for the information, but I hope that we > can avoid that. And yes, it's a rather large patch - but the old > path was so entangled that I couldn't come up with something lighter. I'll run some tests on it. If it does so;ve the regression, I'm going to hold it back until we get a decent amount of review and test coverage on it, though... Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-08-11 03:20 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s4MP7-4rA-9@gated-at.bofh.it> |
| In reply to | #1460115 |
On Wed, Aug 10, 2016 at 05:33:20PM -0700, Huang, Ying wrote:
> Linus Torvalds <torvalds@linux-foundation.org> writes:
>
> > On Wed, Aug 10, 2016 at 5:11 PM, Huang, Ying <ying.huang@intel.com> wrote:
> >>
> >> Here is the comparison result with perf-profile data.
> >
> > Heh. The diff is actually harder to read than just showing A/B
> > state.The fact that the call chain shows up as part of the symbol
> > makes it even more so.
> >
> > For example:
> >
> >> 0.00 ± -1% +Inf% 1.68 ± 1% perf-profile.cycles-pp.__add_to_page_cache_locked.add_to_page_cache_lru.pagecache_get_page.grab_cache_page_write_begin.iomap_write_begin
> >> 1.80 ± 1% -100.0% 0.00 ± -1% perf-profile.cycles-pp.__add_to_page_cache_locked.add_to_page_cache_lru.pagecache_get_page.grab_cache_page_write_begin.xfs_vm_write_begin
> >
> > Ok, so it went from 1.8% to 1.68%, and isn't actually that big of a
> > change, but it shows up as a big change because the caller changed
> > from xfs_vm_write_begin to iomap_write_begin.
> >
> > There's a few other cases of that too.
> >
> > So I think it would actually be easier to just see "what 20 functions
> > were the hottest" (or maybe 50) before and after separately (just
> > sorted by cycles), without the diff part. Because the diff is really
> > hard to read.
>
> Here it is,
>
> Before:
>
> "perf-profile.func.cycles-pp.intel_idle": 16.88,
> "perf-profile.func.cycles-pp.copy_user_enhanced_fast_string": 3.94,
> "perf-profile.func.cycles-pp.memset_erms": 3.26,
> "perf-profile.func.cycles-pp.__block_commit_write.isra.24": 2.47,
> "perf-profile.func.cycles-pp.___might_sleep": 2.33,
> "perf-profile.func.cycles-pp.__mark_inode_dirty": 1.88,
> "perf-profile.func.cycles-pp.unlock_page": 1.69,
> "perf-profile.func.cycles-pp.up_write": 1.61,
> "perf-profile.func.cycles-pp.__block_write_begin_int": 1.56,
> "perf-profile.func.cycles-pp.down_write": 1.55,
> "perf-profile.func.cycles-pp.mark_buffer_dirty": 1.53,
> "perf-profile.func.cycles-pp.entry_SYSCALL_64_fastpath": 1.47,
> "perf-profile.func.cycles-pp.generic_write_end": 1.36,
> "perf-profile.func.cycles-pp.generic_perform_write": 1.33,
> "perf-profile.func.cycles-pp.__radix_tree_lookup": 1.32,
> "perf-profile.func.cycles-pp.__might_sleep": 1.26,
> "perf-profile.func.cycles-pp._raw_spin_lock": 1.17,
> "perf-profile.func.cycles-pp.vfs_write": 1.14,
> "perf-profile.func.cycles-pp.__xfs_get_blocks": 1.07,
Ok, so that is the old block mapping call in the buffered IO path.
I don't see any of the functions it calls in the profile;
specifically xfs_bmapi_read(), and xfs_iomap_write_delay(), so it
appears the extent mapping and allocation overhead on the old code
totals somewhere under 2-3% of the entire CPU usage.
> "perf-profile.func.cycles-pp.xfs_file_write_iter": 1.03,
> "perf-profile.func.cycles-pp.pagecache_get_page": 1.03,
> "perf-profile.func.cycles-pp.native_queued_spin_lock_slowpath": 0.98,
> "perf-profile.func.cycles-pp.get_page_from_freelist": 0.94,
> "perf-profile.func.cycles-pp.rwsem_spin_on_owner": 0.94,
> "perf-profile.func.cycles-pp.__vfs_write": 0.87,
> "perf-profile.func.cycles-pp.iov_iter_copy_from_user_atomic": 0.87,
> "perf-profile.func.cycles-pp.xfs_file_buffered_aio_write": 0.84,
> "perf-profile.func.cycles-pp.find_get_entry": 0.79,
> "perf-profile.func.cycles-pp._raw_spin_lock_irqsave": 0.78,
>
>
> After:
>
> "perf-profile.func.cycles-pp.intel_idle": 16.82,
> "perf-profile.func.cycles-pp.copy_user_enhanced_fast_string": 3.27,
> "perf-profile.func.cycles-pp.memset_erms": 2.6,
> "perf-profile.func.cycles-pp.xfs_bmapi_read": 2.24,
Straight away - thats' at least 3x more overhead block mapping lookups
with the iomap code.
> "perf-profile.func.cycles-pp.___might_sleep": 2.04,
> "perf-profile.func.cycles-pp.mark_page_accessed": 1.93,
> "perf-profile.func.cycles-pp.__block_write_begin_int": 1.78,
> "perf-profile.func.cycles-pp.up_write": 1.72,
> "perf-profile.func.cycles-pp.xfs_iext_bno_to_ext": 1.7,
Plus this child.
> "perf-profile.func.cycles-pp.__block_commit_write.isra.24": 1.65,
> "perf-profile.func.cycles-pp.down_write": 1.51,
> "perf-profile.func.cycles-pp.__mark_inode_dirty": 1.51,
> "perf-profile.func.cycles-pp.unlock_page": 1.43,
> "perf-profile.func.cycles-pp.xfs_bmap_search_multi_extents": 1.25,
> "perf-profile.func.cycles-pp.xfs_bmap_search_extents": 1.23,
And these two.
> "perf-profile.func.cycles-pp.mark_buffer_dirty": 1.21,
> "perf-profile.func.cycles-pp.xfs_iomap_write_delay": 1.19,
> "perf-profile.func.cycles-pp.xfs_iomap_eof_want_preallocate.constprop.8": 1.15,
And these two.
So, essentially the old code had maybe 2-3% cpu usage overhead in
the block mapping path on this workload, but the new code is, for
some reason, showing at least 8-9% CPU usage overhead. That, right
now, makes no sense at all to me as we should be doing - at worst -
exactly the same number of block mapping calls as the old code.
We need to know what is happening that is different - there's a good
chance the mapping trace events will tell us. Huang, can you get
a raw event trace from the test?
I need to see these events:
xfs_file*
xfs_iomap*
xfs_get_block*
For both kernels. An example trace from 4.8-rc1 running the command
`xfs_io -f -c 'pwrite 0 512k -b 128k' /mnt/scratch/fooey doing an
overwrite and extend of the existing file ends up looking like:
$ sudo trace-cmd start -e xfs_iomap\* -e xfs_file\* -e xfs_get_blocks\*
$ sudo cat /sys/kernel/tracing/trace_pipe
<...>-2946 [001] .... 253971.750304: xfs_file_ioctl: dev 253:32 ino 0x84
xfs_io-2946 [001] .... 253971.750938: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 0x20000
xfs_io-2946 [001] .... 253971.750961: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
xfs_io-2946 [001] .... 253971.751114: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 0x20000
xfs_io-2946 [001] .... 253971.751128: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
xfs_io-2946 [001] .... 253971.751234: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 0x20000
xfs_io-2946 [001] .... 253971.751236: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
xfs_io-2946 [001] .... 253971.751381: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 0x20000
xfs_io-2946 [001] .... 253971.751415: xfs_iomap_prealloc_size: dev 253:32 ino 0x84 prealloc blocks 128 shift 0 m_writeio_blocks 16
xfs_io-2946 [001] .... 253971.751425: xfs_iomap_alloc: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 131072 type invalid startoff 0x60 startblock -1 blockcount 0x90
That's the output I need for the complete test - you'll need to use
a better recording mechanism that this (e.g. trace-cmd record,
trace-cmd report) because it will generate a lot of events. Compress
the two report files (they'll be large) and send them to me offlist.
Cheers,
Dave.
--
Dave Chinner
david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-08-11 03:40 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s4N8t-4zK-1@gated-at.bofh.it> |
| In reply to | #1460127 |
On Thu, Aug 11, 2016 at 11:16:12AM +1000, Dave Chinner wrote: > I need to see these events: > > xfs_file* > xfs_iomap* > xfs_get_block* > > For both kernels. An example trace from 4.8-rc1 running the command > `xfs_io -f -c 'pwrite 0 512k -b 128k' /mnt/scratch/fooey doing an > overwrite and extend of the existing file ends up looking like: > > $ sudo trace-cmd start -e xfs_iomap\* -e xfs_file\* -e xfs_get_blocks\* > $ sudo cat /sys/kernel/tracing/trace_pipe > <...>-2946 [001] .... 253971.750304: xfs_file_ioctl: dev 253:32 ino 0x84 > xfs_io-2946 [001] .... 253971.750938: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 0x20000 > xfs_io-2946 [001] .... 253971.750961: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60 > xfs_io-2946 [001] .... 253971.751114: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 0x20000 > xfs_io-2946 [001] .... 253971.751128: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60 > xfs_io-2946 [001] .... 253971.751234: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 0x20000 > xfs_io-2946 [001] .... 253971.751236: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60 > xfs_io-2946 [001] .... 253971.751381: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 0x20000 > xfs_io-2946 [001] .... 253971.751415: xfs_iomap_prealloc_size: dev 253:32 ino 0x84 prealloc blocks 128 shift 0 m_writeio_blocks 16 > xfs_io-2946 [001] .... 253971.751425: xfs_iomap_alloc: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 131072 type invalid startoff 0x60 startblock -1 blockcount 0x90 > > That's the output I need for the complete test - you'll need to use > a better recording mechanism that this (e.g. trace-cmd record, > trace-cmd report) because it will generate a lot of events. Compress > the two report files (they'll be large) and send them to me offlist. Can you also send me the output of xfs_info on the filesystem you are testing? Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Ye Xiaolong <xiaolong.ye@intel.com> |
|---|---|
| Date | 2016-08-11 04:50 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s4Oee-5jA-3@gated-at.bofh.it> |
| In reply to | #1460130 |
On 08/11, Dave Chinner wrote:
>On Thu, Aug 11, 2016 at 11:16:12AM +1000, Dave Chinner wrote:
>> I need to see these events:
>>
>> xfs_file*
>> xfs_iomap*
>> xfs_get_block*
>>
>> For both kernels. An example trace from 4.8-rc1 running the command
>> `xfs_io -f -c 'pwrite 0 512k -b 128k' /mnt/scratch/fooey doing an
>> overwrite and extend of the existing file ends up looking like:
>>
>> $ sudo trace-cmd start -e xfs_iomap\* -e xfs_file\* -e xfs_get_blocks\*
>> $ sudo cat /sys/kernel/tracing/trace_pipe
>> <...>-2946 [001] .... 253971.750304: xfs_file_ioctl: dev 253:32 ino 0x84
>> xfs_io-2946 [001] .... 253971.750938: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 0x20000
>> xfs_io-2946 [001] .... 253971.750961: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
>> xfs_io-2946 [001] .... 253971.751114: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 0x20000
>> xfs_io-2946 [001] .... 253971.751128: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
>> xfs_io-2946 [001] .... 253971.751234: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 0x20000
>> xfs_io-2946 [001] .... 253971.751236: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60
>> xfs_io-2946 [001] .... 253971.751381: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 0x20000
>> xfs_io-2946 [001] .... 253971.751415: xfs_iomap_prealloc_size: dev 253:32 ino 0x84 prealloc blocks 128 shift 0 m_writeio_blocks 16
>> xfs_io-2946 [001] .... 253971.751425: xfs_iomap_alloc: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 131072 type invalid startoff 0x60 startblock -1 blockcount 0x90
>>
>> That's the output I need for the complete test - you'll need to use
>> a better recording mechanism that this (e.g. trace-cmd record,
>> trace-cmd report) because it will generate a lot of events. Compress
>> the two report files (they'll be large) and send them to me offlist.
>
>Can you also send me the output of xfs_info on the filesystem you
>are testing?
Hi, Dave
Here is the xfs_info output:
# xfs_info /fs/ram0/
meta-data=/dev/ram0 isize=256 agcount=4, agsize=3145728 blks
= sectsz=4096 attr=2, projid32bit=1
= crc=0 finobt=0
data = bsize=4096 blocks=12582912, imaxpct=25
= sunit=0 swidth=0 blks
naming =version 2 bsize=4096 ascii-ci=0 ftype=0
log =internal bsize=4096 blocks=6144, version=2
= sectsz=4096 sunit=1 blks, lazy-count=1
realtime =none extsz=4096 blocks=0, rtextents=0
Thanks,
Xiaolong
>
>Cheers,
>
>Dave.
>--
>Dave Chinner
>david@fromorbit.com
>_______________________________________________
>LKP mailing list
>LKP@lists.01.org
>https://lists.01.org/mailman/listinfo/lkp
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-08-11 05:20 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s4OHf-5JF-1@gated-at.bofh.it> |
| In reply to | #1460147 |
On Thu, Aug 11, 2016 at 10:36:59AM +0800, Ye Xiaolong wrote: > On 08/11, Dave Chinner wrote: > >On Thu, Aug 11, 2016 at 11:16:12AM +1000, Dave Chinner wrote: > >> I need to see these events: > >> > >> xfs_file* > >> xfs_iomap* > >> xfs_get_block* > >> > >> For both kernels. An example trace from 4.8-rc1 running the command > >> `xfs_io -f -c 'pwrite 0 512k -b 128k' /mnt/scratch/fooey doing an > >> overwrite and extend of the existing file ends up looking like: > >> > >> $ sudo trace-cmd start -e xfs_iomap\* -e xfs_file\* -e xfs_get_blocks\* > >> $ sudo cat /sys/kernel/tracing/trace_pipe > >> <...>-2946 [001] .... 253971.750304: xfs_file_ioctl: dev 253:32 ino 0x84 > >> xfs_io-2946 [001] .... 253971.750938: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 0x20000 > >> xfs_io-2946 [001] .... 253971.750961: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x0 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60 > >> xfs_io-2946 [001] .... 253971.751114: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 0x20000 > >> xfs_io-2946 [001] .... 253971.751128: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x20000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60 > >> xfs_io-2946 [001] .... 253971.751234: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 0x20000 > >> xfs_io-2946 [001] .... 253971.751236: xfs_iomap_found: dev 253:32 ino 0x84 size 0x40000 offset 0x40000 count 131072 type invalid startoff 0x0 startblock 24 blockcount 0x60 > >> xfs_io-2946 [001] .... 253971.751381: xfs_file_buffered_write: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 0x20000 > >> xfs_io-2946 [001] .... 253971.751415: xfs_iomap_prealloc_size: dev 253:32 ino 0x84 prealloc blocks 128 shift 0 m_writeio_blocks 16 > >> xfs_io-2946 [001] .... 253971.751425: xfs_iomap_alloc: dev 253:32 ino 0x84 size 0x40000 offset 0x60000 count 131072 type invalid startoff 0x60 startblock -1 blockcount 0x90 > >> > >> That's the output I need for the complete test - you'll need to use > >> a better recording mechanism that this (e.g. trace-cmd record, > >> trace-cmd report) because it will generate a lot of events. Compress > >> the two report files (they'll be large) and send them to me offlist. > > > >Can you also send me the output of xfs_info on the filesystem you > >are testing? > > Hi, Dave > > Here is the xfs_info output: > > # xfs_info /fs/ram0/ > meta-data=/dev/ram0 isize=256 agcount=4, agsize=3145728 blks > = sectsz=4096 attr=2, projid32bit=1 > = crc=0 finobt=0 > data = bsize=4096 blocks=12582912, imaxpct=25 > = sunit=0 swidth=0 blks > naming =version 2 bsize=4096 ascii-ci=0 ftype=0 > log =internal bsize=4096 blocks=6144, version=2 > = sectsz=4096 sunit=1 blks, lazy-count=1 > realtime =none extsz=4096 blocks=0, rtextents=0 OK, nothing unusual there. One thing that I did just think of - how close to ENOSPC does this test get? i.e. are we hitting the "we're almost out of free space" slow paths on this test? Cheers, dave. > > Thanks, > Xiaolong > > > >Cheers, > > > >Dave. > >-- > >Dave Chinner > >david@fromorbit.com > >_______________________________________________ > >LKP mailing list > >LKP@lists.01.org > >https://lists.01.org/mailman/listinfo/lkp > -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-08-12 03:30 +0200 |
| Subject | Re: [LKP] [lkp] [xfs] 68a9f5e700: aim7.jobs-per-min -13.6% regression |
| Message-ID | <s59sl-3pL-7@gated-at.bofh.it> |
| In reply to | #1460127 |
On Thu, Aug 11, 2016 at 11:16:12AM +1000, Dave Chinner wrote: > On Wed, Aug 10, 2016 at 05:33:20PM -0700, Huang, Ying wrote: > We need to know what is happening that is different - there's a good > chance the mapping trace events will tell us. Huang, can you get > a raw event trace from the test? > > I need to see these events: > > xfs_file* > xfs_iomap* > xfs_get_block* > lkp-folks, can I please get these traces run and sent to me? I don't have the time or patience to try to get aim7 running on my machines - the build is full of hard-coded paths and libraries that aren't provided by modern distros (e.g. it requires a static libaio.a!) and it fails at the configure stage complaining that: configure: error: C compiler cannot create executables Which is a complete load of BS. Hence I can't make progress until I have some way of understanding what the IO pattern is that is generating the profile being measured. So far I'm unable to do that with any of the tools I been trying, hence I need the traces to work out what I'm missing... Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2016-08-11 02:00 +0200 |
| Message-ID | <s4LzH-3vK-9@gated-at.bofh.it> |
| In reply to | #1460084 |
On Wed, Aug 10, 2016 at 4:08 PM, Dave Chinner <david@fromorbit.com> wrote:
>
> That, to me, says there's a change in lock contention behaviour in
> the workload (which we know aim7 is good at exposing). i.e. the
> iomap change shifted contention from a sleeping lock to a spinning
> lock, or maybe we now trigger optimistic spinning behaviour on a
> lock we previously didn't spin on at all.
Hmm. Possibly. I reacted to the lower cpu load number, but yeah, I
could easily imagine some locking primitive difference too.
> We really need instruction level perf profiles to understand
> this - I don't have a machine with this many cpu cores available
> locally, so I'm not sure I'm going to be able to make any progress
> tracking it down in the short term. Maybe the lkp team has more
> in-depth cpu usage profiles they can share?
Yeah, I've occasionally wanted to see some kind of "top-25 kernel
functions in the profile" thing. That said, when the load isn't all
that familiar, the profiles usually are not all that easy to make
sense of either. But comparing the before and after state might give
us clues.
Fengguang?
Linus
[toc] | [prev] | [standalone]
Page 5 of 5 — ← Prev page 1 2 3 4 [5]
Back to top | Article view | linux.kernel
csiph-web