Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1458902 > unrolled thread

[RFC] mm: Don't use radix tree writeback tags for pages in swap cache

Started by"Huang, Ying" <ying.huang@intel.com>
First post2016-08-09 18:20 +0200
Last post2016-08-09 19:10 +0200
Articles 3 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [RFC] mm: Don't use radix tree writeback tags for pages in swap cache "Huang, Ying" <ying.huang@intel.com> - 2016-08-09 18:20 +0200
    Re: [RFC] mm: Don't use radix tree writeback tags for pages in swap  cache Dave Hansen <dave.hansen@intel.com> - 2016-08-09 18:40 +0200
      Re: [RFC] mm: Don't use radix tree writeback tags for pages in swap cache "Huang\, Ying" <ying.huang@intel.com> - 2016-08-09 19:10 +0200

#1458902 — [RFC] mm: Don't use radix tree writeback tags for pages in swap cache

From"Huang, Ying" <ying.huang@intel.com>
Date2016-08-09 18:20 +0200
Subject[RFC] mm: Don't use radix tree writeback tags for pages in swap cache
Message-ID<s4hV0-1tf-15@gated-at.bofh.it>
From: Huang Ying <ying.huang@intel.com>

File pages uses a set of radix tags (DIRTY, TOWRITE, WRITEBACK) to
accelerate finding the pages with the specific tag in the the radix tree
during writing back an inode.  But for anonymous pages in swap cache,
there are no inode based writeback.  So there is no need to find the
pages with some writeback tags in the radix tree.  It is no necessary to
touch radix tree writeback tags for pages in swap cache.

With this patch, the swap out bandwidth improved 22.3% in vm-scalability
swap-w-seq test case with 8 processes on a Xeon E5 v3 system, because of
reduced contention on swap cache radix tree lock.  To test sequence swap
out, the test case uses 8 processes sequentially allocate and write to
anonymous pages until RAM and part of the swap device is used up.

Details of comparison is as follow,

            base base+patch
---------------- --------------------------
             \          |                \
   2506952 ±  2%     +28.1%    3212076 ±  7%  vm-scalability.throughput
   1207402 ±  7%     +22.3%    1476578 ±  6%  vmstat.swap.so
     10.86 ± 12%     -23.4%       8.31 ± 16%  perf-profile.cycles-pp._raw_spin_lock_irq.__add_to_swap_cache.add_to_swap_cache.add_to_swap.shrink_page_list
     10.82 ± 13%     -33.1%       7.24 ± 14%  perf-profile.cycles-pp._raw_spin_lock_irqsave.__remove_mapping.shrink_page_list.shrink_inactive_list.shrink_zone_memcg
     10.36 ± 11%    -100.0%       0.00 ± -1%  perf-profile.cycles-pp._raw_spin_lock_irqsave.__test_set_page_writeback.bdev_write_page.__swap_writepage.swap_writepage
     10.52 ± 12%    -100.0%       0.00 ± -1%  perf-profile.cycles-pp._raw_spin_lock_irqsave.test_clear_page_writeback.end_page_writeback.page_endio.pmem_rw_page

Cc: Hugh Dickins <hughd@google.com>
Cc: Shaohua Li <shli@kernel.org>
Cc: Minchan Kim <minchan@kernel.org>
Cc: Rik van Riel <riel@redhat.com>
Cc: Mel Gorman <mgorman@techsingularity.net>
Cc: Tejun Heo <tj@kernel.org>
Cc: Wu Fengguang <fengguang.wu@intel.com>
Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
---
 mm/page-writeback.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/mm/page-writeback.c b/mm/page-writeback.c
index f4cd7d8..ebfecb7 100644
--- a/mm/page-writeback.c
+++ b/mm/page-writeback.c
@@ -2758,7 +2758,7 @@ int test_clear_page_writeback(struct page *page)
 	int ret;
 
 	lock_page_memcg(page);
-	if (mapping) {
+	if (mapping && !PageSwapCache(page)) {
 		struct inode *inode = mapping->host;
 		struct backing_dev_info *bdi = inode_to_bdi(inode);
 		unsigned long flags;
@@ -2801,7 +2801,7 @@ int __test_set_page_writeback(struct page *page, bool keep_write)
 	int ret;
 
 	lock_page_memcg(page);
-	if (mapping) {
+	if (mapping && !PageSwapCache(page)) {
 		struct inode *inode = mapping->host;
 		struct backing_dev_info *bdi = inode_to_bdi(inode);
 		unsigned long flags;
-- 
2.8.1

[toc] | [next] | [standalone]


#1458920 — Re: [RFC] mm: Don't use radix tree writeback tags for pages in swap cache

FromDave Hansen <dave.hansen@intel.com>
Date2016-08-09 18:40 +0200
SubjectRe: [RFC] mm: Don't use radix tree writeback tags for pages in swap cache
Message-ID<s4iel-1Aq-19@gated-at.bofh.it>
In reply to#1458902
On 08/09/2016 09:17 AM, Huang, Ying wrote:
> File pages uses a set of radix tags (DIRTY, TOWRITE, WRITEBACK) to
> accelerate finding the pages with the specific tag in the the radix tree
> during writing back an inode.  But for anonymous pages in swap cache,
> there are no inode based writeback.  So there is no need to find the
> pages with some writeback tags in the radix tree.  It is no necessary to
> touch radix tree writeback tags for pages in swap cache.

Seems simple enough.  Do we do any of this unnecessary work for the
other radix tree tags?  If so, maybe we should just fix this once and
for all.  Could we, for instance, WARN_ONCE() in radix_tree_tag_set() if
it sees a swap mapping get handed in there?

In any case, I think the new !PageSwapCache(page) check either needs
commenting, or a common helper for the two sites that you can comment.

> With this patch, the swap out bandwidth improved 22.3% in vm-scalability
> swap-w-seq test case with 8 processes on a Xeon E5 v3 system, because of
> reduced contention on swap cache radix tree lock.  To test sequence swap
> out, the test case uses 8 processes sequentially allocate and write to
> anonymous pages until RAM and part of the swap device is used up.

What was the swap device here, btw?  What is the actual bandwidth
increase you are seeing?  Is it 1MB/s -> 1.223MB/s? :)

[toc] | [prev] | [next] | [standalone]


#1458990

From"Huang\, Ying" <ying.huang@intel.com>
Date2016-08-09 19:10 +0200
Message-ID<s4iHo-20B-21@gated-at.bofh.it>
In reply to#1458920
Hi, Dave,

Dave Hansen <dave.hansen@intel.com> writes:

> On 08/09/2016 09:17 AM, Huang, Ying wrote:
>> File pages uses a set of radix tags (DIRTY, TOWRITE, WRITEBACK) to
>> accelerate finding the pages with the specific tag in the the radix tree
>> during writing back an inode.  But for anonymous pages in swap cache,
>> there are no inode based writeback.  So there is no need to find the
>> pages with some writeback tags in the radix tree.  It is no necessary to
>> touch radix tree writeback tags for pages in swap cache.
>
> Seems simple enough.  Do we do any of this unnecessary work for the
> other radix tree tags?  If so, maybe we should just fix this once and
> for all.  Could we, for instance, WARN_ONCE() in radix_tree_tag_set() if
> it sees a swap mapping get handed in there?

Good idea!  I will do that and try to catch other places if any.

> In any case, I think the new !PageSwapCache(page) check either needs
> commenting, or a common helper for the two sites that you can comment.

Sure.  I will add that.

>> With this patch, the swap out bandwidth improved 22.3% in vm-scalability
>> swap-w-seq test case with 8 processes on a Xeon E5 v3 system, because of
>> reduced contention on swap cache radix tree lock.  To test sequence swap
>> out, the test case uses 8 processes sequentially allocate and write to
>> anonymous pages until RAM and part of the swap device is used up.
>
> What was the swap device here, btw?  What is the actual bandwidth
> increase you are seeing?  Is it 1MB/s -> 1.223MB/s? :)

The swap device here is a DRAM simulated persistent memory block device
(pmem).

   1207402 ±  7%     +22.3%    1476578 ±  6%  vmstat.swap.so

The actual bandwidth increase is from 1.21GB/s -> 1.48 GB/s.  This is
lower than that of NVMe disk, so the bottleneck is in swap subsystem
instead of block subsystem and device.

Best Regards,
Huang, Ying

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web