Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1618508 > unrolled thread

[PATCH -mm -v3] mm, swap: Sort swap entries before free

Started by"Huang, Ying" <ying.huang@intel.com>
First post2017-04-07 08:50 +0200
Last post2017-04-07 23:50 +0200
Articles 3 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH -mm -v3] mm, swap: Sort swap entries before free "Huang, Ying" <ying.huang@intel.com> - 2017-04-07 08:50 +0200
    Re: [PATCH -mm -v3] mm, swap: Sort swap entries before free Rik van Riel <riel@redhat.com> - 2017-04-07 15:10 +0200
    Re: [PATCH -mm -v3] mm, swap: Sort swap entries before free Andrew Morton <akpm@linux-foundation.org> - 2017-04-07 23:50 +0200

#1618508 — [PATCH -mm -v3] mm, swap: Sort swap entries before free

From"Huang, Ying" <ying.huang@intel.com>
Date2017-04-07 08:50 +0200
Subject[PATCH -mm -v3] mm, swap: Sort swap entries before free
Message-ID<ttvSx-3C4-17@gated-at.bofh.it>
From: Huang Ying <ying.huang@intel.com>

To reduce the lock contention of swap_info_struct->lock when freeing
swap entry.  The freed swap entries will be collected in a per-CPU
buffer firstly, and be really freed later in batch.  During the batch
freeing, if the consecutive swap entries in the per-CPU buffer belongs
to same swap device, the swap_info_struct->lock needs to be
acquired/released only once, so that the lock contention could be
reduced greatly.  But if there are multiple swap devices, it is
possible that the lock may be unnecessarily released/acquired because
the swap entries belong to the same swap device are non-consecutive in
the per-CPU buffer.

To solve the issue, the per-CPU buffer is sorted according to the swap
device before freeing the swap entries.  Test shows that the time
spent by swapcache_free_entries() could be reduced after the patch.

Test the patch via measuring the run time of swap_cache_free_entries()
during the exit phase of the applications use much swap space.  The
results shows that the average run time of swap_cache_free_entries()
reduced about 20% after applying the patch.

Signed-off-by: Huang Ying <ying.huang@intel.com>
Acked-by: Tim Chen <tim.c.chen@intel.com>
Cc: Hugh Dickins <hughd@google.com>
Cc: Shaohua Li <shli@kernel.org>
Cc: Minchan Kim <minchan@kernel.org>
Cc: Rik van Riel <riel@redhat.com>

v3:

- Add some comments in code per Rik's suggestion.

v2:

- Avoid sort swap entries if there is only one swap device.
---
 mm/swapfile.c | 12 ++++++++++++
 1 file changed, 12 insertions(+)

diff --git a/mm/swapfile.c b/mm/swapfile.c
index 90054f3c2cdc..f23c56e9be39 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -37,6 +37,7 @@
 #include <linux/swapfile.h>
 #include <linux/export.h>
 #include <linux/swap_slots.h>
+#include <linux/sort.h>
 
 #include <asm/pgtable.h>
 #include <asm/tlbflush.h>
@@ -1065,6 +1066,13 @@ void swapcache_free(swp_entry_t entry)
 	}
 }
 
+static int swp_entry_cmp(const void *ent1, const void *ent2)
+{
+	const swp_entry_t *e1 = ent1, *e2 = ent2;
+
+	return (long)(swp_type(*e1) - swp_type(*e2));
+}
+
 void swapcache_free_entries(swp_entry_t *entries, int n)
 {
 	struct swap_info_struct *p, *prev;
@@ -1075,6 +1083,10 @@ void swapcache_free_entries(swp_entry_t *entries, int n)
 
 	prev = NULL;
 	p = NULL;
+
+	/* Sort swap entries by swap device, so each lock is only taken once. */
+	if (nr_swapfiles > 1)
+		sort(entries, n, sizeof(entries[0]), swp_entry_cmp, NULL);
 	for (i = 0; i < n; ++i) {
 		p = swap_info_get_cont(entries[i], prev);
 		if (p)
-- 
2.11.0

[toc] | [next] | [standalone]


#1618774

FromRik van Riel <riel@redhat.com>
Date2017-04-07 15:10 +0200
Message-ID<ttBOi-7vK-25@gated-at.bofh.it>
In reply to#1618508
On Fri, 2017-04-07 at 14:49 +0800, Huang, Ying wrote:

> To solve the issue, the per-CPU buffer is sorted according to the
> swap
> device before freeing the swap entries.  Test shows that the time
> spent by swapcache_free_entries() could be reduced after the patch.
> 
> Test the patch via measuring the run time of
> swap_cache_free_entries()
> during the exit phase of the applications use much swap space.  The
> results shows that the average run time of swap_cache_free_entries()
> reduced about 20% after applying the patch.
> 
> Signed-off-by: Huang Ying <ying.huang@intel.com>
> Acked-by: Tim Chen <tim.c.chen@intel.com>
> Cc: Hugh Dickins <hughd@google.com>
> Cc: Shaohua Li <shli@kernel.org>
> Cc: Minchan Kim <minchan@kernel.org>
> Cc: Rik van Riel <riel@redhat.com>

Acked-by: Rik van Riel <riel@redhat.com>

[toc] | [prev] | [next] | [standalone]


#1619130

FromAndrew Morton <akpm@linux-foundation.org>
Date2017-04-07 23:50 +0200
Message-ID<ttJVv-4NF-1@gated-at.bofh.it>
In reply to#1618508
On Fri,  7 Apr 2017 14:49:01 +0800 "Huang, Ying" <ying.huang@intel.com> wrote:

> To reduce the lock contention of swap_info_struct->lock when freeing
> swap entry.  The freed swap entries will be collected in a per-CPU
> buffer firstly, and be really freed later in batch.  During the batch
> freeing, if the consecutive swap entries in the per-CPU buffer belongs
> to same swap device, the swap_info_struct->lock needs to be
> acquired/released only once, so that the lock contention could be
> reduced greatly.  But if there are multiple swap devices, it is
> possible that the lock may be unnecessarily released/acquired because
> the swap entries belong to the same swap device are non-consecutive in
> the per-CPU buffer.
> 
> To solve the issue, the per-CPU buffer is sorted according to the swap
> device before freeing the swap entries.  Test shows that the time
> spent by swapcache_free_entries() could be reduced after the patch.
> 
> Test the patch via measuring the run time of swap_cache_free_entries()
> during the exit phase of the applications use much swap space.  The
> results shows that the average run time of swap_cache_free_entries()
> reduced about 20% after applying the patch.

"20%" is useful info, but it is much better to present the absolute
numbers, please.  If it's "20% of one nanosecond" then the patch isn't
very interesting.  If it's "20% of 35 seconds" then we know we have
more work to do.

If there is indeed still a significant problem here then perhaps it
would be better to move the percpu swp_entry_t buffer into the
per-device structure swap_info_struct, so it becomes "per cpu, per
device".  That way we should be able to reduce contention further.

Or maybe we do something else - it all depends upon the significance of
this problem, which is why a full description of your measurements is
useful.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web