Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1625866 > unrolled thread

[PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be split firstly

Started by"Huang, Ying" <ying.huang@intel.com>
First post2017-04-19 09:10 +0200
Last post2017-04-21 02:40 +0200
Articles 5 — 3 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  [PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be split firstly "Huang, Ying" <ying.huang@intel.com> - 2017-04-19 09:10 +0200
    Re: [PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be  split firstly Johannes Weiner <hannes@cmpxchg.org> - 2017-04-19 18:20 +0200
      Re: [PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be split firstly "Huang\, Ying" <ying.huang@intel.com> - 2017-04-20 03:00 +0200
        Re: [PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be  split firstly Johannes Weiner <hannes@cmpxchg.org> - 2017-04-20 23:00 +0200
          Re: [PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be split firstly "Huang\, Ying" <ying.huang@intel.com> - 2017-04-21 02:40 +0200

#1625866 — [PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be split firstly

From"Huang, Ying" <ying.huang@intel.com>
Date2017-04-19 09:10 +0200
Subject[PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be split firstly
Message-ID<txRUu-lm-7@gated-at.bofh.it>
From: Huang Ying <ying.huang@intel.com>

To swap out THP (Transparent Huage Page), before splitting the THP,
the swap cluster will be allocated and the THP will be added into the
swap cache.  But it is possible that the THP cannot be split, so that
we must delete the THP from the swap cache and free the swap cluster.
To avoid that, in this patch, whether the THP can be split is checked
firstly.  The check can only be done racy, but it is good enough for
most cases.

With the patchset, the swap out throughput improves 3.6% (from about
4.16GB/s to about 4.31GB/s) in the vm-scalability swap-w-seq test case
with 8 processes.  The test is done on a Xeon E5 v3 system.  The swap
device used is a RAM simulated PMEM (persistent memory) device.  To
test the sequential swapping out, the test case creates 8 processes,
which sequentially allocate and write to the anonymous pages until the
RAM and part of the swap device is used up.

Cc: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
Acked-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> [for can_split_huge_page()]
---
 include/linux/huge_mm.h |  7 +++++++
 mm/huge_memory.c        | 20 ++++++++++++++++----
 mm/swap_state.c         |  4 ++++
 3 files changed, 27 insertions(+), 4 deletions(-)

diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h
index a3762d49ba39..d3b3e8fcc717 100644
--- a/include/linux/huge_mm.h
+++ b/include/linux/huge_mm.h
@@ -113,6 +113,7 @@ extern unsigned long thp_get_unmapped_area(struct file *filp,
 extern void prep_transhuge_page(struct page *page);
 extern void free_transhuge_page(struct page *page);
 
+bool can_split_huge_page(struct page *page, int *pextra_pins);
 int split_huge_page_to_list(struct page *page, struct list_head *list);
 static inline int split_huge_page(struct page *page)
 {
@@ -231,6 +232,12 @@ static inline void prep_transhuge_page(struct page *page) {}
 
 #define thp_get_unmapped_area	NULL
 
+static inline bool
+can_split_huge_page(struct page *page, int *pextra_pins)
+{
+	BUILD_BUG();
+	return false;
+}
 static inline int
 split_huge_page_to_list(struct page *page, struct list_head *list)
 {
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index b7c06476590e..3a14c77fcce7 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -2384,6 +2384,21 @@ int page_trans_huge_mapcount(struct page *page, int *total_mapcount)
 	return ret;
 }
 
+/* Racy check whether the huge page can be split */
+bool can_split_huge_page(struct page *page, int *pextra_pins)
+{
+	int extra_pins;
+
+	/* Additional pins from radix tree */
+	if (PageAnon(page))
+		extra_pins = PageSwapCache(page) ? HPAGE_PMD_NR : 0;
+	else
+		extra_pins = HPAGE_PMD_NR;
+	if (pextra_pins)
+		*pextra_pins = extra_pins;
+	return total_mapcount(page) == page_count(page) - extra_pins - 1;
+}
+
 /*
  * This function splits huge page into normal pages. @page can point to any
  * subpage of huge page to split. Split doesn't change the position of @page.
@@ -2431,7 +2446,6 @@ int split_huge_page_to_list(struct page *page, struct list_head *list)
 			ret = -EBUSY;
 			goto out;
 		}
-		extra_pins = PageSwapCache(page) ? HPAGE_PMD_NR : 0;
 		mapping = NULL;
 		anon_vma_lock_write(anon_vma);
 	} else {
@@ -2443,8 +2457,6 @@ int split_huge_page_to_list(struct page *page, struct list_head *list)
 			goto out;
 		}
 
-		/* Addidional pins from radix tree */
-		extra_pins = HPAGE_PMD_NR;
 		anon_vma = NULL;
 		i_mmap_lock_read(mapping);
 	}
@@ -2453,7 +2465,7 @@ int split_huge_page_to_list(struct page *page, struct list_head *list)
 	 * Racy check if we can split the page, before freeze_page() will
 	 * split PMDs
 	 */
-	if (total_mapcount(head) != page_count(head) - extra_pins - 1) {
+	if (!can_split_huge_page(head, &extra_pins)) {
 		ret = -EBUSY;
 		goto out_unlock;
 	}
diff --git a/mm/swap_state.c b/mm/swap_state.c
index c3478ee6e633..3a3217f68937 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -192,6 +192,10 @@ int add_to_swap(struct page *page, struct list_head *list)
 	VM_BUG_ON_PAGE(!PageLocked(page), page);
 	VM_BUG_ON_PAGE(!PageUptodate(page), page);
 
+	/* cannot split, skip it */
+	if (unlikely(PageTransHuge(page)) && !can_split_huge_page(page, NULL))
+		return 0;
+
 retry:
 	entry = get_swap_page(page);
 	if (!entry.val)
-- 
2.11.0

[toc] | [next] | [standalone]


#1626438 — Re: [PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be split firstly

FromJohannes Weiner <hannes@cmpxchg.org>
Date2017-04-19 18:20 +0200
SubjectRe: [PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be split firstly
Message-ID<ty0uK-5Cb-29@gated-at.bofh.it>
In reply to#1625866
On Wed, Apr 19, 2017 at 03:06:24PM +0800, Huang, Ying wrote:
> From: Huang Ying <ying.huang@intel.com>
> 
> To swap out THP (Transparent Huage Page), before splitting the THP,
> the swap cluster will be allocated and the THP will be added into the
> swap cache.  But it is possible that the THP cannot be split, so that
> we must delete the THP from the swap cache and free the swap cluster.
> To avoid that, in this patch, whether the THP can be split is checked
> firstly.  The check can only be done racy, but it is good enough for
> most cases.
> 
> With the patchset, the swap out throughput improves 3.6% (from about
> 4.16GB/s to about 4.31GB/s) in the vm-scalability swap-w-seq test case
> with 8 processes.  The test is done on a Xeon E5 v3 system.  The swap
> device used is a RAM simulated PMEM (persistent memory) device.  To
> test the sequential swapping out, the test case creates 8 processes,
> which sequentially allocate and write to the anonymous pages until the
> RAM and part of the swap device is used up.
> 
> Cc: Johannes Weiner <hannes@cmpxchg.org>
> Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
> Acked-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> [for can_split_huge_page()]

How often does this actually happen in practice? Because all that this
protects us from is trying to allocate a swap cluster - which with the
si->free_clusters list really isn't all that expensive - and return it
again. Unless this happens all the time in practice, this optimization
seems misplaced.

It's especially a little strange because in the other email I asked
about the need for unlikely() annotations, yet this patch is adding
branches and checks for what seems to be an unlikely condition into
the THP hot path.

I'd suggest you drop both these optimization attempts unless there is
real data proving that they have a measurable impact.

[toc] | [prev] | [next] | [standalone]


#1626902

From"Huang\, Ying" <ying.huang@intel.com>
Date2017-04-20 03:00 +0200
Message-ID<ty8BY-1Xk-5@gated-at.bofh.it>
In reply to#1626438
Johannes Weiner <hannes@cmpxchg.org> writes:

> On Wed, Apr 19, 2017 at 03:06:24PM +0800, Huang, Ying wrote:
>> From: Huang Ying <ying.huang@intel.com>
>> 
>> To swap out THP (Transparent Huage Page), before splitting the THP,
>> the swap cluster will be allocated and the THP will be added into the
>> swap cache.  But it is possible that the THP cannot be split, so that
>> we must delete the THP from the swap cache and free the swap cluster.
>> To avoid that, in this patch, whether the THP can be split is checked
>> firstly.  The check can only be done racy, but it is good enough for
>> most cases.
>> 
>> With the patchset, the swap out throughput improves 3.6% (from about
>> 4.16GB/s to about 4.31GB/s) in the vm-scalability swap-w-seq test case
>> with 8 processes.  The test is done on a Xeon E5 v3 system.  The swap
>> device used is a RAM simulated PMEM (persistent memory) device.  To
>> test the sequential swapping out, the test case creates 8 processes,
>> which sequentially allocate and write to the anonymous pages until the
>> RAM and part of the swap device is used up.
>> 
>> Cc: Johannes Weiner <hannes@cmpxchg.org>
>> Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
>> Acked-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> [for can_split_huge_page()]
>
> How often does this actually happen in practice? Because all that this
> protects us from is trying to allocate a swap cluster - which with the
> si->free_clusters list really isn't all that expensive - and return it
> again. Unless this happens all the time in practice, this optimization
> seems misplaced.

In addition to allocate/free swap cluster, add/delete to/from the swap
cache will be called too.

> It's especially a little strange because in the other email I asked
> about the need for unlikely() annotations, yet this patch is adding
> branches and checks for what seems to be an unlikely condition into
> the THP hot path.
>
> I'd suggest you drop both these optimization attempts unless there is
> real data proving that they have a measurable impact.

To my surprise too, I found this patch has measurable impact in my
test.  The swap out throughput improves 3.6% in the vm-scalability
swap-w-seq test case with 8 processes.  Details are in the original
patch description.

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1627818 — Re: [PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be split firstly

FromJohannes Weiner <hannes@cmpxchg.org>
Date2017-04-20 23:00 +0200
SubjectRe: [PATCH -mm -v9 2/3] mm, THP, swap: Check whether THP can be split firstly
Message-ID<tyrlf-5cz-1@gated-at.bofh.it>
In reply to#1626902
On Thu, Apr 20, 2017 at 08:50:43AM +0800, Huang, Ying wrote:
> Johannes Weiner <hannes@cmpxchg.org> writes:
> > On Wed, Apr 19, 2017 at 03:06:24PM +0800, Huang, Ying wrote:
> >> With the patchset, the swap out throughput improves 3.6% (from about
> >> 4.16GB/s to about 4.31GB/s) in the vm-scalability swap-w-seq test case
> >> with 8 processes.  The test is done on a Xeon E5 v3 system.  The swap
> >> device used is a RAM simulated PMEM (persistent memory) device.  To
> >> test the sequential swapping out, the test case creates 8 processes,
> >> which sequentially allocate and write to the anonymous pages until the
> >> RAM and part of the swap device is used up.
> >> 
> >> Cc: Johannes Weiner <hannes@cmpxchg.org>
> >> Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
> >> Acked-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> [for can_split_huge_page()]
> >
> > How often does this actually happen in practice? Because all that this
> > protects us from is trying to allocate a swap cluster - which with the
> > si->free_clusters list really isn't all that expensive - and return it
> > again. Unless this happens all the time in practice, this optimization
> > seems misplaced.
>
> To my surprise too, I found this patch has measurable impact in my
> test.  The swap out throughput improves 3.6% in the vm-scalability
> swap-w-seq test case with 8 processes.  Details are in the original
> patch description.

Yeah I think that justifies it.

The changelog says "the patchset", I didn't realize this is the gain
from just this patch alone. Care to update that?

Thanks!

[toc] | [prev] | [next] | [standalone]


#1627878

From"Huang\, Ying" <ying.huang@intel.com>
Date2017-04-21 02:40 +0200
Message-ID<tyuMa-7nU-5@gated-at.bofh.it>
In reply to#1627818
Johannes Weiner <hannes@cmpxchg.org> writes:

> On Thu, Apr 20, 2017 at 08:50:43AM +0800, Huang, Ying wrote:
>> Johannes Weiner <hannes@cmpxchg.org> writes:
>> > On Wed, Apr 19, 2017 at 03:06:24PM +0800, Huang, Ying wrote:
>> >> With the patchset, the swap out throughput improves 3.6% (from about
>> >> 4.16GB/s to about 4.31GB/s) in the vm-scalability swap-w-seq test case
>> >> with 8 processes.  The test is done on a Xeon E5 v3 system.  The swap
>> >> device used is a RAM simulated PMEM (persistent memory) device.  To
>> >> test the sequential swapping out, the test case creates 8 processes,
>> >> which sequentially allocate and write to the anonymous pages until the
>> >> RAM and part of the swap device is used up.
>> >> 
>> >> Cc: Johannes Weiner <hannes@cmpxchg.org>
>> >> Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
>> >> Acked-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> [for can_split_huge_page()]
>> >
>> > How often does this actually happen in practice? Because all that this
>> > protects us from is trying to allocate a swap cluster - which with the
>> > si->free_clusters list really isn't all that expensive - and return it
>> > again. Unless this happens all the time in practice, this optimization
>> > seems misplaced.
>>
>> To my surprise too, I found this patch has measurable impact in my
>> test.  The swap out throughput improves 3.6% in the vm-scalability
>> swap-w-seq test case with 8 processes.  Details are in the original
>> patch description.
>
> Yeah I think that justifies it.
>
> The changelog says "the patchset", I didn't realize this is the gain
> from just this patch alone. Care to update that?

Sorry for confusing, will update it in the next version.

Best Regards,
Huang, Ying

> Thanks!

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web