Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1478488 > unrolled thread

[PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out

Started by"Huang, Ying" <ying.huang@intel.com>
First post2016-09-07 18:50 +0200
Last post2016-09-23 04:20 +0200
Articles 20 on this page of 55 — 13 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out "Huang, Ying" <ying.huang@intel.com> - 2016-09-07 18:50 +0200
    [PATCH -v3 02/10] mm, memcg: Add swap_cgroup_iter iterator "Huang, Ying" <ying.huang@intel.com> - 2016-09-07 18:50 +0200
    [PATCH -v3 08/10] mm, THP: Add can_split_huge_page() "Huang, Ying" <ying.huang@intel.com> - 2016-09-07 18:50 +0200
      Re: [PATCH -v3 08/10] mm, THP: Add can_split_huge_page() "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-09-08 13:20 +0200
        Re: [PATCH -v3 08/10] mm, THP: Add can_split_huge_page() "Huang\, Ying" <ying.huang@intel.com> - 2016-09-08 19:10 +0200
    [PATCH -v3 01/10] mm, swap: Make swap cluster size same of THP size on x86_64 "Huang, Ying" <ying.huang@intel.com> - 2016-09-07 18:50 +0200
      Re: [PATCH -v3 01/10] mm, swap: Make swap cluster size same of THP  size on x86_64 Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-09-08 10:30 +0200
      Re: [PATCH -v3 01/10] mm, swap: Make swap cluster size same of THP  size on x86_64 Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-09-08 11:20 +0200
        Re: [PATCH -v3 01/10] mm, swap: Make swap cluster size same of THP size on x86_64 "Huang\, Ying" <ying.huang@intel.com> - 2016-09-08 20:10 +0200
        Re: [PATCH -v3 01/10] mm, swap: Make swap cluster size same of THP  size on x86_64 Johannes Weiner <hannes@cmpxchg.org> - 2016-09-19 19:20 +0200
          Re: [PATCH -v3 01/10] mm, swap: Make swap cluster size same of THP  size on x86_64 Johannes Weiner <hannes@cmpxchg.org> - 2016-09-22 21:30 +0200
      Re: [PATCH -v3 01/10] mm, swap: Make swap cluster size same of THP  size on x86_64 "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-09-08 13:10 +0200
        Re: [PATCH -v3 01/10] mm, swap: Make swap cluster size same of THP size on x86_64 "Huang\, Ying" <ying.huang@intel.com> - 2016-09-08 19:40 +0200
      Re: [PATCH -v3 01/10] mm, swap: Make swap cluster size same of THP  size on x86_64 "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-09-08 13:10 +0200
        Re: [PATCH -v3 01/10] mm, swap: Make swap cluster size same of THP size on x86_64 "Huang\, Ying" <ying.huang@intel.com> - 2016-09-08 19:30 +0200
    [PATCH -v3 05/10] mm, THP, swap: Add get_huge_swap_page() "Huang, Ying" <ying.huang@intel.com> - 2016-09-07 18:50 +0200
      Re: [PATCH -v3 05/10] mm, THP, swap: Add get_huge_swap_page() "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-09-08 13:20 +0200
        Re: [PATCH -v3 05/10] mm, THP, swap: Add get_huge_swap_page() "Huang\, Ying" <ying.huang@intel.com> - 2016-09-08 19:30 +0200
    [PATCH -v3 09/10] mm, THP, swap: Support to split THP in swap cache "Huang, Ying" <ying.huang@intel.com> - 2016-09-07 18:50 +0200
    [PATCH -v3 06/10] mm, THP, swap: Support to clear SWAP_HAS_CACHE for huge page "Huang, Ying" <ying.huang@intel.com> - 2016-09-07 18:50 +0200
    [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple swap entries "Huang, Ying" <ying.huang@intel.com> - 2016-09-07 18:50 +0200
      Re: [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple  swap entries Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-09-08 10:30 +0200
        Re: [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple swap entries "Huang\, Ying" <ying.huang@intel.com> - 2016-09-08 20:20 +0200
      Re: [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple  swap entries Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-09-08 10:40 +0200
    [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free functions "Huang, Ying" <ying.huang@intel.com> - 2016-09-07 18:50 +0200
      Re: [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free  functions Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-09-08 10:40 +0200
        Re: [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free functions "Huang\, Ying" <ying.huang@intel.com> - 2016-09-08 20:20 +0200
      Re: [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free  functions Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-09-08 11:00 +0200
    [PATCH -v3 07/10] mm, THP, swap: Support to add/delete THP to/from swap cache "Huang, Ying" <ying.huang@intel.com> - 2016-09-07 18:50 +0200
      Re: [PATCH -v3 07/10] mm, THP, swap: Support to add/delete THP to/from  swap cache Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-09-08 11:10 +0200
        Re: [PATCH -v3 07/10] mm, THP, swap: Support to add/delete THP to/from swap cache "Huang\, Ying" <ying.huang@intel.com> - 2016-09-08 20:20 +0200
    Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Minchan Kim <minchan@kernel.org> - 2016-09-09 07:50 +0200
      Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Tim Chen <tim.c.chen@linux.intel.com> - 2016-09-09 18:00 +0200
      Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out "Huang\, Ying" <ying.huang@intel.com> - 2016-09-09 22:40 +0200
        Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Minchan Kim <minchan@kernel.org> - 2016-09-13 08:20 +0200
          Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out "Huang\, Ying" <ying.huang@intel.com> - 2016-09-13 08:50 +0200
            Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Minchan Kim <minchan@kernel.org> - 2016-09-13 09:10 +0200
              Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out "Huang\, Ying" <ying.huang@intel.com> - 2016-09-13 11:00 +0200
                Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Minchan Kim <minchan@kernel.org> - 2016-09-13 11:20 +0200
                  RE: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out "Chen, Tim C" <tim.c.chen@intel.com> - 2016-09-14 02:00 +0200
                    Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Minchan Kim <minchan@kernel.org> - 2016-09-19 09:20 +0200
                      Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Tim Chen <tim.c.chen@linux.intel.com> - 2016-09-19 18:00 +0200
                  Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out "Huang\, Ying" <ying.huang@intel.com> - 2016-09-18 04:00 +0200
                    Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Minchan Kim <minchan@kernel.org> - 2016-09-19 09:10 +0200
                      Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out "Huang\, Ying" <ying.huang@intel.com> - 2016-09-20 05:00 +0200
                        Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out "Huang\, Ying" <ying.huang@intel.com> - 2016-09-20 07:30 +0200
                        Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Minchan Kim <minchan@kernel.org> - 2016-09-20 07:40 +0200
                Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Andrea Arcangeli <aarcange@redhat.com> - 2016-09-13 16:40 +0200
    Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Hugh Dickins <hughd@google.com> - 2016-09-19 19:40 +0200
    Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Shaohua Li <shli@kernel.org> - 2016-09-23 01:00 +0200
      RE: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out "Chen, Tim C" <tim.c.chen@intel.com> - 2016-09-23 01:50 +0200
        Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out Andi Kleen <andi@firstfloor.org> - 2016-09-23 02:00 +0200
      Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping  out Rik van Riel <riel@redhat.com> - 2016-09-23 02:40 +0200
        Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out "Huang\, Ying" <ying.huang@intel.com> - 2016-09-23 04:40 +0200
      Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out "Huang\, Ying" <ying.huang@intel.com> - 2016-09-23 04:20 +0200

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#1478499 — [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple swap entries

From"Huang, Ying" <ying.huang@intel.com>
Date2016-09-07 18:50 +0200
Subject[PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple swap entries
Message-ID<seOcW-2cL-41@gated-at.bofh.it>
In reply to#1478488
From: Huang Ying <ying.huang@intel.com>

This patch make it possible to charge or uncharge a set of continuous
swap entries in the swap cgroup.  The number of swap entries is
specified via an added parameter.

This will be used for the THP (Transparent Huge Page) swap support.
Where a swap cluster backing a THP may be allocated and freed as a
whole.  So a set of continuous swap entries (512 on x86_64) backing one
THP need to be charged or uncharged together.  This will batch the
cgroup operations for the THP swap too.

Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Cc: Vladimir Davydov <vdavydov@virtuozzo.com>
Cc: Johannes Weiner <hannes@cmpxchg.org>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Tejun Heo <tj@kernel.org>
Cc: cgroups@vger.kernel.org
Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
---
 include/linux/swap.h        | 12 ++++++----
 include/linux/swap_cgroup.h |  6 +++--
 mm/memcontrol.c             | 55 +++++++++++++++++++++++++--------------------
 mm/shmem.c                  |  2 +-
 mm/swap_cgroup.c            | 17 ++++++++++----
 mm/swap_state.c             |  2 +-
 mm/swapfile.c               |  2 +-
 7 files changed, 59 insertions(+), 37 deletions(-)

diff --git a/include/linux/swap.h b/include/linux/swap.h
index ed41bec..75aad24 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -550,8 +550,10 @@ static inline int mem_cgroup_swappiness(struct mem_cgroup *mem)
 
 #ifdef CONFIG_MEMCG_SWAP
 extern void mem_cgroup_swapout(struct page *page, swp_entry_t entry);
-extern int mem_cgroup_try_charge_swap(struct page *page, swp_entry_t entry);
-extern void mem_cgroup_uncharge_swap(swp_entry_t entry);
+extern int mem_cgroup_try_charge_swap(struct page *page, swp_entry_t entry,
+				      unsigned int nr_entries);
+extern void mem_cgroup_uncharge_swap(swp_entry_t entry,
+				     unsigned int nr_entries);
 extern long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg);
 extern bool mem_cgroup_swap_full(struct page *page);
 #else
@@ -560,12 +562,14 @@ static inline void mem_cgroup_swapout(struct page *page, swp_entry_t entry)
 }
 
 static inline int mem_cgroup_try_charge_swap(struct page *page,
-					     swp_entry_t entry)
+					     swp_entry_t entry,
+					     unsigned int nr_entries)
 {
 	return 0;
 }
 
-static inline void mem_cgroup_uncharge_swap(swp_entry_t entry)
+static inline void mem_cgroup_uncharge_swap(swp_entry_t entry,
+					    unsigned int nr_entries)
 {
 }
 
diff --git a/include/linux/swap_cgroup.h b/include/linux/swap_cgroup.h
index 145306b..b2b8ec7 100644
--- a/include/linux/swap_cgroup.h
+++ b/include/linux/swap_cgroup.h
@@ -7,7 +7,8 @@
 
 extern unsigned short swap_cgroup_cmpxchg(swp_entry_t ent,
 					unsigned short old, unsigned short new);
-extern unsigned short swap_cgroup_record(swp_entry_t ent, unsigned short id);
+extern unsigned short swap_cgroup_record(swp_entry_t ent, unsigned short id,
+					 unsigned int nr_ents);
 extern unsigned short lookup_swap_cgroup_id(swp_entry_t ent);
 extern int swap_cgroup_swapon(int type, unsigned long max_pages);
 extern void swap_cgroup_swapoff(int type);
@@ -15,7 +16,8 @@ extern void swap_cgroup_swapoff(int type);
 #else
 
 static inline
-unsigned short swap_cgroup_record(swp_entry_t ent, unsigned short id)
+unsigned short swap_cgroup_record(swp_entry_t ent, unsigned short id,
+				  unsigned int nr_ents)
 {
 	return 0;
 }
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index bdb796f..9662fcf 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -2370,10 +2370,9 @@ void mem_cgroup_split_huge_fixup(struct page *head)
 
 #ifdef CONFIG_MEMCG_SWAP
 static void mem_cgroup_swap_statistics(struct mem_cgroup *memcg,
-					 bool charge)
+				       int nr_entries)
 {
-	int val = (charge) ? 1 : -1;
-	this_cpu_add(memcg->stat->count[MEM_CGROUP_STAT_SWAP], val);
+	this_cpu_add(memcg->stat->count[MEM_CGROUP_STAT_SWAP], nr_entries);
 }
 
 /**
@@ -2399,8 +2398,8 @@ static int mem_cgroup_move_swap_account(swp_entry_t entry,
 	new_id = mem_cgroup_id(to);
 
 	if (swap_cgroup_cmpxchg(entry, old_id, new_id) == old_id) {
-		mem_cgroup_swap_statistics(from, false);
-		mem_cgroup_swap_statistics(to, true);
+		mem_cgroup_swap_statistics(from, -1);
+		mem_cgroup_swap_statistics(to, 1);
 		return 0;
 	}
 	return -EINVAL;
@@ -5417,7 +5416,7 @@ void mem_cgroup_commit_charge(struct page *page, struct mem_cgroup *memcg,
 		 * let's not wait for it.  The page already received a
 		 * memory+swap charge, drop the swap entry duplicate.
 		 */
-		mem_cgroup_uncharge_swap(entry);
+		mem_cgroup_uncharge_swap(entry, nr_pages);
 	}
 }
 
@@ -5825,9 +5824,9 @@ void mem_cgroup_swapout(struct page *page, swp_entry_t entry)
 	 * ancestor for the swap instead and transfer the memory+swap charge.
 	 */
 	swap_memcg = mem_cgroup_id_get_online(memcg);
-	oldid = swap_cgroup_record(entry, mem_cgroup_id(swap_memcg));
+	oldid = swap_cgroup_record(entry, mem_cgroup_id(swap_memcg), 1);
 	VM_BUG_ON_PAGE(oldid, page);
-	mem_cgroup_swap_statistics(swap_memcg, true);
+	mem_cgroup_swap_statistics(swap_memcg, 1);
 
 	page->mem_cgroup = NULL;
 
@@ -5854,16 +5853,19 @@ void mem_cgroup_swapout(struct page *page, swp_entry_t entry)
 		css_put(&memcg->css);
 }
 
-/*
- * mem_cgroup_try_charge_swap - try charging a swap entry
+/**
+ * mem_cgroup_try_charge_swap - try charging a set of swap entries
  * @page: page being added to swap
- * @entry: swap entry to charge
+ * @entry: the first swap entry to charge
+ * @nr_entries: the number of swap entries to charge
  *
- * Try to charge @entry to the memcg that @page belongs to.
+ * Try to charge @nr_entries swap entries starting from @entry to the
+ * memcg that @page belongs to.
  *
  * Returns 0 on success, -ENOMEM on failure.
  */
-int mem_cgroup_try_charge_swap(struct page *page, swp_entry_t entry)
+int mem_cgroup_try_charge_swap(struct page *page, swp_entry_t entry,
+			       unsigned int nr_entries)
 {
 	struct mem_cgroup *memcg;
 	struct page_counter *counter;
@@ -5881,25 +5883,29 @@ int mem_cgroup_try_charge_swap(struct page *page, swp_entry_t entry)
 	memcg = mem_cgroup_id_get_online(memcg);
 
 	if (!mem_cgroup_is_root(memcg) &&
-	    !page_counter_try_charge(&memcg->swap, 1, &counter)) {
+	    !page_counter_try_charge(&memcg->swap, nr_entries, &counter)) {
 		mem_cgroup_id_put(memcg);
 		return -ENOMEM;
 	}
 
-	oldid = swap_cgroup_record(entry, mem_cgroup_id(memcg));
+	if (nr_entries > 1)
+		mem_cgroup_id_get_many(memcg, nr_entries - 1);
+	oldid = swap_cgroup_record(entry, mem_cgroup_id(memcg), nr_entries);
 	VM_BUG_ON_PAGE(oldid, page);
-	mem_cgroup_swap_statistics(memcg, true);
+	mem_cgroup_swap_statistics(memcg, nr_entries);
 
 	return 0;
 }
 
 /**
- * mem_cgroup_uncharge_swap - uncharge a swap entry
- * @entry: swap entry to uncharge
+ * mem_cgroup_uncharge_swap - uncharge a set of swap entries
+ * @entry: the first swap entry to uncharge
+ * @nr_entries: the number of swap entries to uncharge
  *
- * Drop the swap charge associated with @entry.
+ * Drop the swap charge associated with @nr_entries swap entries
+ * starting from @entry.
  */
-void mem_cgroup_uncharge_swap(swp_entry_t entry)
+void mem_cgroup_uncharge_swap(swp_entry_t entry, unsigned int nr_entries)
 {
 	struct mem_cgroup *memcg;
 	unsigned short id;
@@ -5907,17 +5913,18 @@ void mem_cgroup_uncharge_swap(swp_entry_t entry)
 	if (!do_swap_account)
 		return;
 
-	id = swap_cgroup_record(entry, 0);
+	id = swap_cgroup_record(entry, 0, nr_entries);
 	rcu_read_lock();
 	memcg = mem_cgroup_from_id(id);
 	if (memcg) {
 		if (!mem_cgroup_is_root(memcg)) {
 			if (cgroup_subsys_on_dfl(memory_cgrp_subsys))
-				page_counter_uncharge(&memcg->swap, 1);
+				page_counter_uncharge(&memcg->swap, nr_entries);
 			else
-				page_counter_uncharge(&memcg->memsw, 1);
+				page_counter_uncharge(&memcg->memsw,
+						      nr_entries);
 		}
-		mem_cgroup_swap_statistics(memcg, false);
+		mem_cgroup_swap_statistics(memcg, -nr_entries);
 		mem_cgroup_id_put(memcg);
 	}
 	rcu_read_unlock();
diff --git a/mm/shmem.c b/mm/shmem.c
index ac35ebd..baeb2f9 100644
--- a/mm/shmem.c
+++ b/mm/shmem.c
@@ -1248,7 +1248,7 @@ static int shmem_writepage(struct page *page, struct writeback_control *wbc)
 	if (!swap.val)
 		goto redirty;
 
-	if (mem_cgroup_try_charge_swap(page, swap))
+	if (mem_cgroup_try_charge_swap(page, swap, 1))
 		goto free_swap;
 
 	/*
diff --git a/mm/swap_cgroup.c b/mm/swap_cgroup.c
index 4ae3e7b..4d3484f 100644
--- a/mm/swap_cgroup.c
+++ b/mm/swap_cgroup.c
@@ -139,14 +139,16 @@ unsigned short swap_cgroup_cmpxchg(swp_entry_t ent,
 }
 
 /**
- * swap_cgroup_record - record mem_cgroup for this swp_entry.
- * @ent: swap entry to be recorded into
+ * swap_cgroup_record - record mem_cgroup for a set of swap entries
+ * @ent: the first swap entry to be recorded into
  * @id: mem_cgroup to be recorded
+ * @nr_ents: number of swap entries to be recorded
  *
  * Returns old value at success, 0 at failure.
  * (Of course, old value can be 0.)
  */
-unsigned short swap_cgroup_record(swp_entry_t ent, unsigned short id)
+unsigned short swap_cgroup_record(swp_entry_t ent, unsigned short id,
+				  unsigned int nr_ents)
 {
 	struct swap_cgroup_iter iter;
 	unsigned short old;
@@ -154,7 +156,14 @@ unsigned short swap_cgroup_record(swp_entry_t ent, unsigned short id)
 	swap_cgroup_iter_init(&iter, ent);
 
 	old = iter.sc->id;
-	iter.sc->id = id;
+	for (;;) {
+		VM_BUG_ON(iter.sc->id != old);
+		iter.sc->id = id;
+		nr_ents--;
+		if (!nr_ents)
+			break;
+		swap_cgroup_iter_advance(&iter);
+	}
 
 	swap_cgroup_iter_exit(&iter);
 	return old;
diff --git a/mm/swap_state.c b/mm/swap_state.c
index 8679c99..c335251 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -172,7 +172,7 @@ int add_to_swap(struct page *page, struct list_head *list)
 	if (!entry.val)
 		return 0;
 
-	if (mem_cgroup_try_charge_swap(page, entry)) {
+	if (mem_cgroup_try_charge_swap(page, entry, 1)) {
 		swapcache_free(entry);
 		return 0;
 	}
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 4b78402..17f25e2 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -806,7 +806,7 @@ static unsigned char swap_entry_free(struct swap_info_struct *p,
 
 	/* free if no reference */
 	if (!usage) {
-		mem_cgroup_uncharge_swap(entry);
+		mem_cgroup_uncharge_swap(entry, 1);
 		dec_cluster_info_page(p, p->cluster_info, offset);
 		if (offset < p->lowest_bit)
 			p->lowest_bit = offset;
-- 
2.8.1

[toc] | [prev] | [next] | [standalone]


#1478915 — Re: [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple swap entries

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-09-08 10:30 +0200
SubjectRe: [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple swap entries
Message-ID<sf2SC-3oQ-13@gated-at.bofh.it>
In reply to#1478499
On 09/07/2016 10:16 PM, Huang, Ying wrote:
> From: Huang Ying <ying.huang@intel.com>
> 
> This patch make it possible to charge or uncharge a set of continuous
> swap entries in the swap cgroup.  The number of swap entries is
> specified via an added parameter.
> 
> This will be used for the THP (Transparent Huge Page) swap support.
> Where a swap cluster backing a THP may be allocated and freed as a
> whole.  So a set of continuous swap entries (512 on x86_64) backing one

Please use HPAGE_SIZE / PAGE_SIZE instead of hard coded number like 512.

[toc] | [prev] | [next] | [standalone]


#1479420 — Re: [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple swap entries

From"Huang\, Ying" <ying.huang@intel.com>
Date2016-09-08 20:20 +0200
SubjectRe: [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple swap entries
Message-ID<sfc5A-OX-9@gated-at.bofh.it>
In reply to#1478915
Anshuman Khandual <khandual@linux.vnet.ibm.com> writes:

> On 09/07/2016 10:16 PM, Huang, Ying wrote:
>> From: Huang Ying <ying.huang@intel.com>
>> 
>> This patch make it possible to charge or uncharge a set of continuous
>> swap entries in the swap cgroup.  The number of swap entries is
>> specified via an added parameter.
>> 
>> This will be used for the THP (Transparent Huge Page) swap support.
>> Where a swap cluster backing a THP may be allocated and freed as a
>> whole.  So a set of continuous swap entries (512 on x86_64) backing one
>
> Please use HPAGE_SIZE / PAGE_SIZE instead of hard coded number like 512.

Sure.  Will change it.

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1478925 — Re: [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple swap entries

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-09-08 10:40 +0200
SubjectRe: [PATCH -v3 03/10] mm, memcg: Support to charge/uncharge multiple swap entries
Message-ID<sf32i-3sy-37@gated-at.bofh.it>
In reply to#1478499
On 09/07/2016 10:16 PM, Huang, Ying wrote:
> From: Huang Ying <ying.huang@intel.com>
> 
> This patch make it possible to charge or uncharge a set of continuous
> swap entries in the swap cgroup.  The number of swap entries is
> specified via an added parameter.
> 
> This will be used for the THP (Transparent Huge Page) swap support.
> Where a swap cluster backing a THP may be allocated and freed as a
> whole.  So a set of continuous swap entries (512 on x86_64) backing one

Please use HPAGE_SIZE / PAGE_SIZE instead of hard coded number like 512.

[toc] | [prev] | [next] | [standalone]


#1478501 — [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free functions

From"Huang, Ying" <ying.huang@intel.com>
Date2016-09-07 18:50 +0200
Subject[PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free functions
Message-ID<seOcW-2cL-47@gated-at.bofh.it>
In reply to#1478488
From: Huang Ying <ying.huang@intel.com>

The swap cluster allocation/free functions are added based on the
existing swap cluster management mechanism for SSD.  These functions
don't work for the rotating hard disks because the existing swap cluster
management mechanism doesn't work for them.  The hard disks support may
be added if someone really need it.  But that needn't be included in
this patchset.

This will be used for the THP (Transparent Huge Page) swap support.
Where one swap cluster will hold the contents of each THP swapped out.

Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Cc: Hugh Dickins <hughd@google.com>
Cc: Shaohua Li <shli@kernel.org>
Cc: Minchan Kim <minchan@kernel.org>
Cc: Rik van Riel <riel@redhat.com>
Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
---
 mm/swapfile.c | 203 +++++++++++++++++++++++++++++++++++++++++-----------------
 1 file changed, 146 insertions(+), 57 deletions(-)

diff --git a/mm/swapfile.c b/mm/swapfile.c
index 17f25e2..0132e8c 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -326,6 +326,14 @@ static void swap_cluster_schedule_discard(struct swap_info_struct *si,
 	schedule_work(&si->discard_work);
 }
 
+static void __free_cluster(struct swap_info_struct *si, unsigned long idx)
+{
+	struct swap_cluster_info *ci = si->cluster_info;
+
+	cluster_set_flag(ci + idx, CLUSTER_FLAG_FREE);
+	cluster_list_add_tail(&si->free_clusters, ci, idx);
+}
+
 /*
  * Doing discard actually. After a cluster discard is finished, the cluster
  * will be added to free cluster list. caller should hold si->lock.
@@ -345,8 +353,7 @@ static void swap_do_scheduled_discard(struct swap_info_struct *si)
 				SWAPFILE_CLUSTER);
 
 		spin_lock(&si->lock);
-		cluster_set_flag(&info[idx], CLUSTER_FLAG_FREE);
-		cluster_list_add_tail(&si->free_clusters, info, idx);
+		__free_cluster(si, idx);
 		memset(si->swap_map + idx * SWAPFILE_CLUSTER,
 				0, SWAPFILE_CLUSTER);
 	}
@@ -363,6 +370,34 @@ static void swap_discard_work(struct work_struct *work)
 	spin_unlock(&si->lock);
 }
 
+static void alloc_cluster(struct swap_info_struct *si, unsigned long idx)
+{
+	struct swap_cluster_info *ci = si->cluster_info;
+
+	VM_BUG_ON(cluster_list_first(&si->free_clusters) != idx);
+	cluster_list_del_first(&si->free_clusters, ci);
+	cluster_set_count_flag(ci + idx, 0, 0);
+}
+
+static void free_cluster(struct swap_info_struct *si, unsigned long idx)
+{
+	struct swap_cluster_info *ci = si->cluster_info + idx;
+
+	VM_BUG_ON(cluster_count(ci) != 0);
+	/*
+	 * If the swap is discardable, prepare discard the cluster
+	 * instead of free it immediately. The cluster will be freed
+	 * after discard.
+	 */
+	if ((si->flags & (SWP_WRITEOK | SWP_PAGE_DISCARD)) ==
+	    (SWP_WRITEOK | SWP_PAGE_DISCARD)) {
+		swap_cluster_schedule_discard(si, idx);
+		return;
+	}
+
+	__free_cluster(si, idx);
+}
+
 /*
  * The cluster corresponding to page_nr will be used. The cluster will be
  * removed from free cluster list and its usage counter will be increased.
@@ -374,11 +409,8 @@ static void inc_cluster_info_page(struct swap_info_struct *p,
 
 	if (!cluster_info)
 		return;
-	if (cluster_is_free(&cluster_info[idx])) {
-		VM_BUG_ON(cluster_list_first(&p->free_clusters) != idx);
-		cluster_list_del_first(&p->free_clusters, cluster_info);
-		cluster_set_count_flag(&cluster_info[idx], 0, 0);
-	}
+	if (cluster_is_free(&cluster_info[idx]))
+		alloc_cluster(p, idx);
 
 	VM_BUG_ON(cluster_count(&cluster_info[idx]) >= SWAPFILE_CLUSTER);
 	cluster_set_count(&cluster_info[idx],
@@ -402,21 +434,8 @@ static void dec_cluster_info_page(struct swap_info_struct *p,
 	cluster_set_count(&cluster_info[idx],
 		cluster_count(&cluster_info[idx]) - 1);
 
-	if (cluster_count(&cluster_info[idx]) == 0) {
-		/*
-		 * If the swap is discardable, prepare discard the cluster
-		 * instead of free it immediately. The cluster will be freed
-		 * after discard.
-		 */
-		if ((p->flags & (SWP_WRITEOK | SWP_PAGE_DISCARD)) ==
-				 (SWP_WRITEOK | SWP_PAGE_DISCARD)) {
-			swap_cluster_schedule_discard(p, idx);
-			return;
-		}
-
-		cluster_set_flag(&cluster_info[idx], CLUSTER_FLAG_FREE);
-		cluster_list_add_tail(&p->free_clusters, cluster_info, idx);
-	}
+	if (cluster_count(&cluster_info[idx]) == 0)
+		free_cluster(p, idx);
 }
 
 /*
@@ -497,6 +516,69 @@ new_cluster:
 	*scan_base = tmp;
 }
 
+#ifdef CONFIG_THP_SWAP_CLUSTER
+static inline unsigned int huge_cluster_nr_entries(bool huge)
+{
+	return huge ? SWAPFILE_CLUSTER : 1;
+}
+#else
+#define huge_cluster_nr_entries(huge)	1
+#endif
+
+static void __swap_entry_alloc(struct swap_info_struct *si,
+			       unsigned long offset, bool huge)
+{
+	unsigned int nr_entries = huge_cluster_nr_entries(huge);
+	unsigned int end = offset + nr_entries - 1;
+
+	if (offset == si->lowest_bit)
+		si->lowest_bit += nr_entries;
+	if (end == si->highest_bit)
+		si->highest_bit -= nr_entries;
+	si->inuse_pages += nr_entries;
+	if (si->inuse_pages == si->pages) {
+		si->lowest_bit = si->max;
+		si->highest_bit = 0;
+		spin_lock(&swap_avail_lock);
+		plist_del(&si->avail_list, &swap_avail_head);
+		spin_unlock(&swap_avail_lock);
+	}
+}
+
+static void __swap_entry_free(struct swap_info_struct *si, unsigned long offset,
+			      bool huge)
+{
+	unsigned int nr_entries = huge_cluster_nr_entries(huge);
+	unsigned long end = offset + nr_entries - 1;
+	void (*swap_slot_free_notify)(struct block_device *, unsigned long);
+
+	if (offset < si->lowest_bit)
+		si->lowest_bit = offset;
+	if (end > si->highest_bit) {
+		bool was_full = !si->highest_bit;
+
+		si->highest_bit = end;
+		if (was_full && (si->flags & SWP_WRITEOK)) {
+			spin_lock(&swap_avail_lock);
+			WARN_ON(!plist_node_empty(&si->avail_list));
+			if (plist_node_empty(&si->avail_list))
+				plist_add(&si->avail_list, &swap_avail_head);
+			spin_unlock(&swap_avail_lock);
+		}
+	}
+	atomic_long_add(nr_entries, &nr_swap_pages);
+	si->inuse_pages -= nr_entries;
+	if (si->flags & SWP_BLKDEV)
+		swap_slot_free_notify =
+			si->bdev->bd_disk->fops->swap_slot_free_notify;
+	while (offset <= end) {
+		frontswap_invalidate_page(si->type, offset);
+		if (swap_slot_free_notify)
+			swap_slot_free_notify(si->bdev, offset);
+		offset++;
+	}
+}
+
 static unsigned long scan_swap_map(struct swap_info_struct *si,
 				   unsigned char usage)
 {
@@ -591,18 +673,7 @@ checks:
 	if (si->swap_map[offset])
 		goto scan;
 
-	if (offset == si->lowest_bit)
-		si->lowest_bit++;
-	if (offset == si->highest_bit)
-		si->highest_bit--;
-	si->inuse_pages++;
-	if (si->inuse_pages == si->pages) {
-		si->lowest_bit = si->max;
-		si->highest_bit = 0;
-		spin_lock(&swap_avail_lock);
-		plist_del(&si->avail_list, &swap_avail_head);
-		spin_unlock(&swap_avail_lock);
-	}
+	__swap_entry_alloc(si, offset, false);
 	si->swap_map[offset] = usage;
 	inc_cluster_info_page(si, si->cluster_info, offset);
 	si->cluster_next = offset + 1;
@@ -649,6 +720,46 @@ no_page:
 	return 0;
 }
 
+#ifdef CONFIG_THP_SWAP_CLUSTER
+static void swap_free_huge_cluster(struct swap_info_struct *si,
+				   unsigned long idx)
+{
+	struct swap_cluster_info *ci = si->cluster_info + idx;
+	unsigned long offset = idx * SWAPFILE_CLUSTER;
+
+	cluster_set_count_flag(ci, 0, 0);
+	free_cluster(si, idx);
+	__swap_entry_free(si, offset, true);
+}
+
+static unsigned long swap_alloc_huge_cluster(struct swap_info_struct *si)
+{
+	unsigned long idx;
+	struct swap_cluster_info *ci;
+	unsigned long offset, i;
+	unsigned char *map;
+
+	if (cluster_list_empty(&si->free_clusters))
+		return 0;
+	idx = cluster_list_first(&si->free_clusters);
+	alloc_cluster(si, idx);
+	ci = si->cluster_info + idx;
+	cluster_set_count_flag(ci, SWAPFILE_CLUSTER, 0);
+
+	offset = idx * SWAPFILE_CLUSTER;
+	__swap_entry_alloc(si, offset, true);
+	map = si->swap_map + offset;
+	for (i = 0; i < SWAPFILE_CLUSTER; i++)
+		map[i] = SWAP_HAS_CACHE;
+	return offset;
+}
+#else
+static inline unsigned long swap_alloc_huge_cluster(struct swap_info_struct *si)
+{
+	return 0;
+}
+#endif
+
 swp_entry_t get_swap_page(void)
 {
 	struct swap_info_struct *si, *next;
@@ -808,29 +919,7 @@ static unsigned char swap_entry_free(struct swap_info_struct *p,
 	if (!usage) {
 		mem_cgroup_uncharge_swap(entry, 1);
 		dec_cluster_info_page(p, p->cluster_info, offset);
-		if (offset < p->lowest_bit)
-			p->lowest_bit = offset;
-		if (offset > p->highest_bit) {
-			bool was_full = !p->highest_bit;
-			p->highest_bit = offset;
-			if (was_full && (p->flags & SWP_WRITEOK)) {
-				spin_lock(&swap_avail_lock);
-				WARN_ON(!plist_node_empty(&p->avail_list));
-				if (plist_node_empty(&p->avail_list))
-					plist_add(&p->avail_list,
-						  &swap_avail_head);
-				spin_unlock(&swap_avail_lock);
-			}
-		}
-		atomic_long_inc(&nr_swap_pages);
-		p->inuse_pages--;
-		frontswap_invalidate_page(p->type, offset);
-		if (p->flags & SWP_BLKDEV) {
-			struct gendisk *disk = p->bdev->bd_disk;
-			if (disk->fops->swap_slot_free_notify)
-				disk->fops->swap_slot_free_notify(p->bdev,
-								  offset);
-		}
+		__swap_entry_free(p, offset, false);
 	}
 
 	return usage;
-- 
2.8.1

[toc] | [prev] | [next] | [standalone]


#1478917 — Re: [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free functions

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-09-08 10:40 +0200
SubjectRe: [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free functions
Message-ID<sf32h-3sy-5@gated-at.bofh.it>
In reply to#1478501
On 09/07/2016 10:16 PM, Huang, Ying wrote:
> From: Huang Ying <ying.huang@intel.com>
> 
> The swap cluster allocation/free functions are added based on the
> existing swap cluster management mechanism for SSD.  These functions
> don't work for the rotating hard disks because the existing swap cluster
> management mechanism doesn't work for them.  The hard disks support may
> be added if someone really need it.  But that needn't be included in
> this patchset.
> 
> This will be used for the THP (Transparent Huge Page) swap support.
> Where one swap cluster will hold the contents of each THP swapped out.

Which tree this series is based against ? This patch does not apply
on the mainline kernel.

[toc] | [prev] | [next] | [standalone]


#1479429 — Re: [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free functions

From"Huang\, Ying" <ying.huang@intel.com>
Date2016-09-08 20:20 +0200
SubjectRe: [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free functions
Message-ID<sfc5A-OX-27@gated-at.bofh.it>
In reply to#1478917
Anshuman Khandual <khandual@linux.vnet.ibm.com> writes:

> On 09/07/2016 10:16 PM, Huang, Ying wrote:
>> From: Huang Ying <ying.huang@intel.com>
>> 
>> The swap cluster allocation/free functions are added based on the
>> existing swap cluster management mechanism for SSD.  These functions
>> don't work for the rotating hard disks because the existing swap cluster
>> management mechanism doesn't work for them.  The hard disks support may
>> be added if someone really need it.  But that needn't be included in
>> this patchset.
>> 
>> This will be used for the THP (Transparent Huge Page) swap support.
>> Where one swap cluster will hold the contents of each THP swapped out.
>
> Which tree this series is based against ? This patch does not apply
> on the mainline kernel.

This series is based on 8/31 head of mmotm/master.  I stated it in
00/10, but I know it is hided inside other text and not obvious at all.
Is there some way to make it obvious?

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1478938 — Re: [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free functions

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-09-08 11:00 +0200
SubjectRe: [PATCH -v3 04/10] mm, THP, swap: Add swap cluster allocate/free functions
Message-ID<sf3lD-3zk-7@gated-at.bofh.it>
In reply to#1478501
On 09/07/2016 10:16 PM, Huang, Ying wrote:
> From: Huang Ying <ying.huang@intel.com>
> 
> The swap cluster allocation/free functions are added based on the
> existing swap cluster management mechanism for SSD.  These functions
> don't work for the rotating hard disks because the existing swap cluster
> management mechanism doesn't work for them.  The hard disks support may
> be added if someone really need it.  But that needn't be included in
> this patchset.
> 
> This will be used for the THP (Transparent Huge Page) swap support.
> Where one swap cluster will hold the contents of each THP swapped out.

Which tree this series is based against ? This patch does not apply
on the mainline kernel today.

[toc] | [prev] | [next] | [standalone]


#1478502 — [PATCH -v3 07/10] mm, THP, swap: Support to add/delete THP to/from swap cache

From"Huang, Ying" <ying.huang@intel.com>
Date2016-09-07 18:50 +0200
Subject[PATCH -v3 07/10] mm, THP, swap: Support to add/delete THP to/from swap cache
Message-ID<seOcX-2cL-51@gated-at.bofh.it>
In reply to#1478488
From: Huang Ying <ying.huang@intel.com>

With this patch, a THP (Transparent Huge Page) can be added/deleted
to/from the swap cache as a set of sub-pages (512 on x86_64).

This will be used for the THP (Transparent Huge Page) swap support.
Where one THP may be added/delted to/from the swap cache.  This will
batch the swap cache operations to reduce the lock acquire/release times
for the THP swap too.

Cc: Hugh Dickins <hughd@google.com>
Cc: Shaohua Li <shli@kernel.org>
Cc: Minchan Kim <minchan@kernel.org>
Cc: Rik van Riel <riel@redhat.com>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
---
 include/linux/page-flags.h |  2 +-
 mm/swap_state.c            | 57 +++++++++++++++++++++++++++++++---------------
 2 files changed, 40 insertions(+), 19 deletions(-)

diff --git a/include/linux/page-flags.h b/include/linux/page-flags.h
index 74e4dda..f5bcbea 100644
--- a/include/linux/page-flags.h
+++ b/include/linux/page-flags.h
@@ -314,7 +314,7 @@ PAGEFLAG_FALSE(HighMem)
 #endif
 
 #ifdef CONFIG_SWAP
-PAGEFLAG(SwapCache, swapcache, PF_NO_COMPOUND)
+PAGEFLAG(SwapCache, swapcache, PF_NO_TAIL)
 #else
 PAGEFLAG_FALSE(SwapCache)
 #endif
diff --git a/mm/swap_state.c b/mm/swap_state.c
index c335251..db2299f 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -43,6 +43,7 @@ struct address_space swapper_spaces[MAX_SWAPFILES] = {
 };
 
 #define INC_CACHE_INFO(x)	do { swap_cache_info.x++; } while (0)
+#define ADD_CACHE_INFO(x, nr)	do { swap_cache_info.x += (nr); } while (0)
 
 static struct {
 	unsigned long add_total;
@@ -80,25 +81,32 @@ void show_swap_cache_info(void)
  */
 int __add_to_swap_cache(struct page *page, swp_entry_t entry)
 {
-	int error;
+	int error, i, nr = hpage_nr_pages(page);
 	struct address_space *address_space;
 
 	VM_BUG_ON_PAGE(!PageLocked(page), page);
 	VM_BUG_ON_PAGE(PageSwapCache(page), page);
 	VM_BUG_ON_PAGE(!PageSwapBacked(page), page);
 
-	get_page(page);
+	page_ref_add(page, nr);
 	SetPageSwapCache(page);
-	set_page_private(page, entry.val);
 
 	address_space = swap_address_space(entry);
 	spin_lock_irq(&address_space->tree_lock);
-	error = radix_tree_insert(&address_space->page_tree,
-					entry.val, page);
+	for (i = 0; i < nr; i++) {
+		struct page *cur_page = page + i;
+		unsigned long index = entry.val + i;
+
+		set_page_private(cur_page, index);
+		error = radix_tree_insert(&address_space->page_tree,
+					  index, cur_page);
+		if (unlikely(error))
+			break;
+	}
 	if (likely(!error)) {
-		address_space->nrpages++;
-		__inc_node_page_state(page, NR_FILE_PAGES);
-		INC_CACHE_INFO(add_total);
+		address_space->nrpages += nr;
+		__mod_node_page_state(page_pgdat(page), NR_FILE_PAGES, nr);
+		ADD_CACHE_INFO(add_total, nr);
 	}
 	spin_unlock_irq(&address_space->tree_lock);
 
@@ -109,9 +117,16 @@ int __add_to_swap_cache(struct page *page, swp_entry_t entry)
 		 * So add_to_swap_cache() doesn't returns -EEXIST.
 		 */
 		VM_BUG_ON(error == -EEXIST);
-		set_page_private(page, 0UL);
 		ClearPageSwapCache(page);
-		put_page(page);
+		set_page_private(page + i, 0UL);
+		while (i--) {
+			struct page *cur_page = page + i;
+			unsigned long index = entry.val + i;
+
+			set_page_private(cur_page, 0UL);
+			radix_tree_delete(&address_space->page_tree, index);
+		}
+		page_ref_sub(page, nr);
 	}
 
 	return error;
@@ -122,7 +137,7 @@ int add_to_swap_cache(struct page *page, swp_entry_t entry, gfp_t gfp_mask)
 {
 	int error;
 
-	error = radix_tree_maybe_preload(gfp_mask);
+	error = radix_tree_maybe_preload_order(gfp_mask, compound_order(page));
 	if (!error) {
 		error = __add_to_swap_cache(page, entry);
 		radix_tree_preload_end();
@@ -138,6 +153,7 @@ void __delete_from_swap_cache(struct page *page)
 {
 	swp_entry_t entry;
 	struct address_space *address_space;
+	int i, nr = hpage_nr_pages(page);
 
 	VM_BUG_ON_PAGE(!PageLocked(page), page);
 	VM_BUG_ON_PAGE(!PageSwapCache(page), page);
@@ -145,12 +161,17 @@ void __delete_from_swap_cache(struct page *page)
 
 	entry.val = page_private(page);
 	address_space = swap_address_space(entry);
-	radix_tree_delete(&address_space->page_tree, page_private(page));
-	set_page_private(page, 0);
 	ClearPageSwapCache(page);
-	address_space->nrpages--;
-	__dec_node_page_state(page, NR_FILE_PAGES);
-	INC_CACHE_INFO(del_total);
+	for (i = 0; i < nr; i++) {
+		struct page *cur_page = page + i;
+
+		radix_tree_delete(&address_space->page_tree,
+				  page_private(cur_page));
+		set_page_private(cur_page, 0);
+	}
+	address_space->nrpages -= nr;
+	__mod_node_page_state(page_pgdat(page), NR_FILE_PAGES, -nr);
+	ADD_CACHE_INFO(del_total, nr);
 }
 
 /**
@@ -227,8 +248,8 @@ void delete_from_swap_cache(struct page *page)
 	__delete_from_swap_cache(page);
 	spin_unlock_irq(&address_space->tree_lock);
 
-	swapcache_free(entry);
-	put_page(page);
+	__swapcache_free(entry, PageTransHuge(page));
+	page_ref_sub(page, hpage_nr_pages(page));
 }
 
 /* 
-- 
2.8.1

[toc] | [prev] | [next] | [standalone]


#1478949 — Re: [PATCH -v3 07/10] mm, THP, swap: Support to add/delete THP to/from swap cache

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-09-08 11:10 +0200
SubjectRe: [PATCH -v3 07/10] mm, THP, swap: Support to add/delete THP to/from swap cache
Message-ID<sf3vk-3RO-27@gated-at.bofh.it>
In reply to#1478502
On 09/07/2016 10:16 PM, Huang, Ying wrote:
> From: Huang Ying <ying.huang@intel.com>
> 
> With this patch, a THP (Transparent Huge Page) can be added/deleted
> to/from the swap cache as a set of sub-pages (512 on x86_64).
> 
> This will be used for the THP (Transparent Huge Page) swap support.
> Where one THP may be added/delted to/from the swap cache.  This will
> batch the swap cache operations to reduce the lock acquire/release times
> for the THP swap too.
> 
> Cc: Hugh Dickins <hughd@google.com>
> Cc: Shaohua Li <shli@kernel.org>
> Cc: Minchan Kim <minchan@kernel.org>
> Cc: Rik van Riel <riel@redhat.com>
> Cc: Andrea Arcangeli <aarcange@redhat.com>
> Cc: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
> Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
> ---
>  include/linux/page-flags.h |  2 +-
>  mm/swap_state.c            | 57 +++++++++++++++++++++++++++++++---------------
>  2 files changed, 40 insertions(+), 19 deletions(-)
> 
> diff --git a/include/linux/page-flags.h b/include/linux/page-flags.h
> index 74e4dda..f5bcbea 100644
> --- a/include/linux/page-flags.h
> +++ b/include/linux/page-flags.h
> @@ -314,7 +314,7 @@ PAGEFLAG_FALSE(HighMem)
>  #endif
>  
>  #ifdef CONFIG_SWAP
> -PAGEFLAG(SwapCache, swapcache, PF_NO_COMPOUND)
> +PAGEFLAG(SwapCache, swapcache, PF_NO_TAIL)

What is the reason for this change ? The commit message does not seem
to explain.

[toc] | [prev] | [next] | [standalone]


#1479421 — Re: [PATCH -v3 07/10] mm, THP, swap: Support to add/delete THP to/from swap cache

From"Huang\, Ying" <ying.huang@intel.com>
Date2016-09-08 20:20 +0200
SubjectRe: [PATCH -v3 07/10] mm, THP, swap: Support to add/delete THP to/from swap cache
Message-ID<sfc5z-OX-1@gated-at.bofh.it>
In reply to#1478949
Hi, Anshuman,

Thanks for comments!

Anshuman Khandual <khandual@linux.vnet.ibm.com> writes:

> On 09/07/2016 10:16 PM, Huang, Ying wrote:
>> From: Huang Ying <ying.huang@intel.com>
>> 
>> With this patch, a THP (Transparent Huge Page) can be added/deleted
>> to/from the swap cache as a set of sub-pages (512 on x86_64).
>> 
>> This will be used for the THP (Transparent Huge Page) swap support.
>> Where one THP may be added/delted to/from the swap cache.  This will
>> batch the swap cache operations to reduce the lock acquire/release times
>> for the THP swap too.
>> 
>> Cc: Hugh Dickins <hughd@google.com>
>> Cc: Shaohua Li <shli@kernel.org>
>> Cc: Minchan Kim <minchan@kernel.org>
>> Cc: Rik van Riel <riel@redhat.com>
>> Cc: Andrea Arcangeli <aarcange@redhat.com>
>> Cc: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
>> Signed-off-by: "Huang, Ying" <ying.huang@intel.com>
>> ---
>>  include/linux/page-flags.h |  2 +-
>>  mm/swap_state.c            | 57 +++++++++++++++++++++++++++++++---------------
>>  2 files changed, 40 insertions(+), 19 deletions(-)
>> 
>> diff --git a/include/linux/page-flags.h b/include/linux/page-flags.h
>> index 74e4dda..f5bcbea 100644
>> --- a/include/linux/page-flags.h
>> +++ b/include/linux/page-flags.h
>> @@ -314,7 +314,7 @@ PAGEFLAG_FALSE(HighMem)
>>  #endif
>>  
>>  #ifdef CONFIG_SWAP
>> -PAGEFLAG(SwapCache, swapcache, PF_NO_COMPOUND)
>> +PAGEFLAG(SwapCache, swapcache, PF_NO_TAIL)
>
> What is the reason for this change ? The commit message does not seem
> to explain.

Before this change, SetPageSwapCache() cannot be called for THP, after
the change, SetPageSwapCache() could be called for the head page of the
THP, but not the tail pages.  Because we will never do that before this
patch series.

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1479661 — Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out

FromMinchan Kim <minchan@kernel.org>
Date2016-09-09 07:50 +0200
SubjectRe: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out
Message-ID<sfmRj-7nj-11@gated-at.bofh.it>
In reply to#1478488
Hi Huang,

On Wed, Sep 07, 2016 at 09:45:59AM -0700, Huang, Ying wrote:
> From: Huang Ying <ying.huang@intel.com>
> 
> This patchset is to optimize the performance of Transparent Huge Page
> (THP) swap.
> 
> Hi, Andrew, could you help me to check whether the overall design is
> reasonable?
> 
> Hi, Hugh, Shaohua, Minchan and Rik, could you help me to review the
> swap part of the patchset?  Especially [01/10], [04/10], [05/10],
> [06/10], [07/10], [10/10].
> 
> Hi, Andrea and Kirill, could you help me to review the THP part of the
> patchset?  Especially [02/10], [03/10], [09/10] and [10/10].
> 
> Hi, Johannes, Michal and Vladimir, I am not very confident about the
> memory cgroup part, especially [02/10] and [03/10].  Could you help me
> to review it?
> 
> And for all, Any comment is welcome!
> 
> 
> Recently, the performance of the storage devices improved so fast that
> we cannot saturate the disk bandwidth when do page swap out even on a
> high-end server machine.  Because the performance of the storage
> device improved faster than that of CPU.  And it seems that the trend
> will not change in the near future.  On the other hand, the THP
> becomes more and more popular because of increased memory size.  So it
> becomes necessary to optimize THP swap performance.
> 
> The advantages of the THP swap support include:
> 
> - Batch the swap operations for the THP to reduce lock
>   acquiring/releasing, including allocating/freeing the swap space,
>   adding/deleting to/from the swap cache, and writing/reading the swap
>   space, etc.  This will help improve the performance of the THP swap.
> 
> - The THP swap space read/write will be 2M sequential IO.  It is
>   particularly helpful for the swap read, which usually are 4k random
>   IO.  This will improve the performance of the THP swap too.
> 
> - It will help the memory fragmentation, especially when the THP is
>   heavily used by the applications.  The 2M continuous pages will be
>   free up after THP swapping out.

I just read patchset right now and still doubt why the all changes
should be coupled with THP tightly. Many parts(e.g., you introduced
or modifying existing functions for making them THP specific) could
just take page_list and the number of pages then would handle them
without THP awareness.

For example, if the nr_pages is larger than SWAPFILE_CLUSTER, we
can try to allocate new cluster. With that, we could allocate new
clusters to meet nr_pages requested or bail out if we fail to allocate
and fallback to 0-order page swapout. With that, swap layer could
support multiple order-0 pages by batch.

IMO, I really want to land Tim Chen's batching swapout work first.
With Tim Chen's work, I expect we can make better refactoring
for batching swap before adding more confuse to the swap layer.
(I expect it would share several pieces of code for or would be base
for batching allocation of swapcache, swapslot)

After that, we could enhance swap for big contiguous batching
like THP and finally we might make it be aware of THP specific to
enhance further.

A thing I remember you aruged: you want to swapin 512 pages
all at once unconditionally. It's really worth to discuss if
your design is going for the way.
I doubt it's generally good idea. Because, currently, we try to
swap in swapped out pages in THP page with conservative approach
but your direction is going to opposite way.

[mm, thp: convert from optimistic swapin collapsing to conservative]

I think general approach(i.e., less effective than targeting
implement for your own specific goal but less hacky and better job
for many cases) is to rely/improve on the swap readahead.
If most of subpages of a THP page are really workingset, swap readahead
could work well.

Yeah, it's fairly vague feedback so sorry if I miss something clear.

[toc] | [prev] | [next] | [standalone]


#1480174 — Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out

FromTim Chen <tim.c.chen@linux.intel.com>
Date2016-09-09 18:00 +0200
SubjectRe: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out
Message-ID<sfwnD-4Sy-1@gated-at.bofh.it>
In reply to#1479661
On Fri, 2016-09-09 at 14:43 +0900, Minchan Kim wrote:
> Hi Huang,
> 
> On Wed, Sep 07, 2016 at 09:45:59AM -0700, Huang, Ying wrote:
> > 
> > From: Huang Ying <ying.huang@intel.com>
> > 
> > This patchset is to optimize the performance of Transparent Huge Page
> > (THP) swap.
> > 
> > Hi, Andrew, could you help me to check whether the overall design is
> > reasonable?
> > 
> > Hi, Hugh, Shaohua, Minchan and Rik, could you help me to review the
> > swap part of the patchset?  Especially [01/10], [04/10], [05/10],
> > [06/10], [07/10], [10/10].
> > 
> > Hi, Andrea and Kirill, could you help me to review the THP part of the
> > patchset?  Especially [02/10], [03/10], [09/10] and [10/10].
> > 
> > Hi, Johannes, Michal and Vladimir, I am not very confident about the
> > memory cgroup part, especially [02/10] and [03/10].  Could you help me
> > to review it?
> > 
> > And for all, Any comment is welcome!
> > 
> > 
> > Recently, the performance of the storage devices improved so fast that
> > we cannot saturate the disk bandwidth when do page swap out even on a
> > high-end server machine.  Because the performance of the storage
> > device improved faster than that of CPU.  And it seems that the trend
> > will not change in the near future.  On the other hand, the THP
> > becomes more and more popular because of increased memory size.  So it
> > becomes necessary to optimize THP swap performance.
> > 
> > The advantages of the THP swap support include:
> > 
> > - Batch the swap operations for the THP to reduce lock
> >   acquiring/releasing, including allocating/freeing the swap space,
> >   adding/deleting to/from the swap cache, and writing/reading the swap
> >   space, etc.  This will help improve the performance of the THP swap.
> > 
> > - The THP swap space read/write will be 2M sequential IO.  It is
> >   particularly helpful for the swap read, which usually are 4k random
> >   IO.  This will improve the performance of the THP swap too.
> > 
> > - It will help the memory fragmentation, especially when the THP is
> >   heavily used by the applications.  The 2M continuous pages will be
> >   free up after THP swapping out.
> I just read patchset right now and still doubt why the all changes
> should be coupled with THP tightly. Many parts(e.g., you introduced
> or modifying existing functions for making them THP specific) could
> just take page_list and the number of pages then would handle them
> without THP awareness.
> 
> For example, if the nr_pages is larger than SWAPFILE_CLUSTER, we
> can try to allocate new cluster. With that, we could allocate new
> clusters to meet nr_pages requested or bail out if we fail to allocate
> and fallback to 0-order page swapout. With that, swap layer could
> support multiple order-0 pages by batch.
> 
> IMO, I really want to land Tim Chen's batching swapout work first.
> With Tim Chen's work, I expect we can make better refactoring
> for batching swap before adding more confuse to the swap layer.
> (I expect it would share several pieces of code for or would be base
> for batching allocation of swapcache, swapslot)

Minchan,

Ying and I do plan to send out a new patch series on batching swapout
and swapin plus a few other optimization on the swapping of 
regular sized pages.

Hopefully we'll be able to do that soon after we fixed up a few
things and retest.

Tim

[toc] | [prev] | [next] | [standalone]


#1480340

From"Huang\, Ying" <ying.huang@intel.com>
Date2016-09-09 22:40 +0200
Message-ID<sfAKB-7F8-7@gated-at.bofh.it>
In reply to#1479661
Hi, Minchan,

Minchan Kim <minchan@kernel.org> writes:
> Hi Huang,
>
> On Wed, Sep 07, 2016 at 09:45:59AM -0700, Huang, Ying wrote:
>> From: Huang Ying <ying.huang@intel.com>
>> 
>> This patchset is to optimize the performance of Transparent Huge Page
>> (THP) swap.
>> 
>> Hi, Andrew, could you help me to check whether the overall design is
>> reasonable?
>> 
>> Hi, Hugh, Shaohua, Minchan and Rik, could you help me to review the
>> swap part of the patchset?  Especially [01/10], [04/10], [05/10],
>> [06/10], [07/10], [10/10].
>> 
>> Hi, Andrea and Kirill, could you help me to review the THP part of the
>> patchset?  Especially [02/10], [03/10], [09/10] and [10/10].
>> 
>> Hi, Johannes, Michal and Vladimir, I am not very confident about the
>> memory cgroup part, especially [02/10] and [03/10].  Could you help me
>> to review it?
>> 
>> And for all, Any comment is welcome!
>> 
>> 
>> Recently, the performance of the storage devices improved so fast that
>> we cannot saturate the disk bandwidth when do page swap out even on a
>> high-end server machine.  Because the performance of the storage
>> device improved faster than that of CPU.  And it seems that the trend
>> will not change in the near future.  On the other hand, the THP
>> becomes more and more popular because of increased memory size.  So it
>> becomes necessary to optimize THP swap performance.
>> 
>> The advantages of the THP swap support include:
>> 
>> - Batch the swap operations for the THP to reduce lock
>>   acquiring/releasing, including allocating/freeing the swap space,
>>   adding/deleting to/from the swap cache, and writing/reading the swap
>>   space, etc.  This will help improve the performance of the THP swap.
>> 
>> - The THP swap space read/write will be 2M sequential IO.  It is
>>   particularly helpful for the swap read, which usually are 4k random
>>   IO.  This will improve the performance of the THP swap too.
>> 
>> - It will help the memory fragmentation, especially when the THP is
>>   heavily used by the applications.  The 2M continuous pages will be
>>   free up after THP swapping out.
>
> I just read patchset right now and still doubt why the all changes
> should be coupled with THP tightly. Many parts(e.g., you introduced
> or modifying existing functions for making them THP specific) could
> just take page_list and the number of pages then would handle them
> without THP awareness.

I am glad if my change could help normal pages swapping too.  And we can
change these functions to work for normal pages when necessary.

> For example, if the nr_pages is larger than SWAPFILE_CLUSTER, we
> can try to allocate new cluster. With that, we could allocate new
> clusters to meet nr_pages requested or bail out if we fail to allocate
> and fallback to 0-order page swapout. With that, swap layer could
> support multiple order-0 pages by batch.
>
> IMO, I really want to land Tim Chen's batching swapout work first.
> With Tim Chen's work, I expect we can make better refactoring
> for batching swap before adding more confuse to the swap layer.
> (I expect it would share several pieces of code for or would be base
> for batching allocation of swapcache, swapslot)

I don't think there is hard conflict between normal pages swapping
optimizing and THP swap optimizing.  Some code may be shared between
them.  That is good for both sides.

> After that, we could enhance swap for big contiguous batching
> like THP and finally we might make it be aware of THP specific to
> enhance further.
>
> A thing I remember you aruged: you want to swapin 512 pages
> all at once unconditionally. It's really worth to discuss if
> your design is going for the way.
> I doubt it's generally good idea. Because, currently, we try to
> swap in swapped out pages in THP page with conservative approach
> but your direction is going to opposite way.
>
> [mm, thp: convert from optimistic swapin collapsing to conservative]
>
> I think general approach(i.e., less effective than targeting
> implement for your own specific goal but less hacky and better job
> for many cases) is to rely/improve on the swap readahead.
> If most of subpages of a THP page are really workingset, swap readahead
> could work well.
>
> Yeah, it's fairly vague feedback so sorry if I miss something clear.

Yes.  I want to go to the direction that to swap in 512 pages together.
And I think it is a good opportunity to discuss that now.  The advantages
of swapping in 512 pages together are:

- Improve the performance of swapping in IO via turning small read size
  into 512 pages big read size.

- Keep THP across swap out/in.  With the memory size become more and
  more large, the 4k pages bring more and more burden to memory
  management.  One solution is to use 2M pages as much as possible, that
  will reduce the management burden greatly, such as much reduced length
  of LRU list, etc.

The disadvantage are:

- Increase the memory pressure when swap in THP.

- Some pages swapped in may not needed in the near future.

Because of the disadvantages, the 512 pages swapping in should be made
optional.  But I don't think we should make it impossible.

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1482160 — Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out

FromMinchan Kim <minchan@kernel.org>
Date2016-09-13 08:20 +0200
SubjectRe: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out
Message-ID<sgPex-6r3-5@gated-at.bofh.it>
In reply to#1480340
Hi Huang,

On Fri, Sep 09, 2016 at 01:35:12PM -0700, Huang, Ying wrote:

< snip >

> >> Recently, the performance of the storage devices improved so fast that
> >> we cannot saturate the disk bandwidth when do page swap out even on a
> >> high-end server machine.  Because the performance of the storage
> >> device improved faster than that of CPU.  And it seems that the trend
> >> will not change in the near future.  On the other hand, the THP
> >> becomes more and more popular because of increased memory size.  So it
> >> becomes necessary to optimize THP swap performance.
> >> 
> >> The advantages of the THP swap support include:
> >> 
> >> - Batch the swap operations for the THP to reduce lock
> >>   acquiring/releasing, including allocating/freeing the swap space,
> >>   adding/deleting to/from the swap cache, and writing/reading the swap
> >>   space, etc.  This will help improve the performance of the THP swap.
> >> 
> >> - The THP swap space read/write will be 2M sequential IO.  It is
> >>   particularly helpful for the swap read, which usually are 4k random
> >>   IO.  This will improve the performance of the THP swap too.
> >> 
> >> - It will help the memory fragmentation, especially when the THP is
> >>   heavily used by the applications.  The 2M continuous pages will be
> >>   free up after THP swapping out.
> >
> > I just read patchset right now and still doubt why the all changes
> > should be coupled with THP tightly. Many parts(e.g., you introduced
> > or modifying existing functions for making them THP specific) could
> > just take page_list and the number of pages then would handle them
> > without THP awareness.
> 
> I am glad if my change could help normal pages swapping too.  And we can
> change these functions to work for normal pages when necessary.

Sure but it would be less painful that THP awareness swapout is
based on multiple normal pages swapout. For exmaple, we don't
touch delay THP split part(i.e., split a THP into 512 pages like
as-is) and enhances swapout further like Tim's suggestion
for mulitple normal pages swapout. With that, it might be enough
for fast-storage without needing THP awareness.

My *point* is let's approach step by step.
First of all, go with batching normal pages swapout and if it's
not enough, dive into further optimization like introducing
THP-aware swapout.

I believe it's natural development process to evolve things
without over-engineering.

> 
> > For example, if the nr_pages is larger than SWAPFILE_CLUSTER, we
> > can try to allocate new cluster. With that, we could allocate new
> > clusters to meet nr_pages requested or bail out if we fail to allocate
> > and fallback to 0-order page swapout. With that, swap layer could
> > support multiple order-0 pages by batch.
> >
> > IMO, I really want to land Tim Chen's batching swapout work first.
> > With Tim Chen's work, I expect we can make better refactoring
> > for batching swap before adding more confuse to the swap layer.
> > (I expect it would share several pieces of code for or would be base
> > for batching allocation of swapcache, swapslot)
> 
> I don't think there is hard conflict between normal pages swapping
> optimizing and THP swap optimizing.  Some code may be shared between
> them.  That is good for both sides.
> 
> > After that, we could enhance swap for big contiguous batching
> > like THP and finally we might make it be aware of THP specific to
> > enhance further.
> >
> > A thing I remember you aruged: you want to swapin 512 pages
> > all at once unconditionally. It's really worth to discuss if
> > your design is going for the way.
> > I doubt it's generally good idea. Because, currently, we try to
> > swap in swapped out pages in THP page with conservative approach
> > but your direction is going to opposite way.
> >
> > [mm, thp: convert from optimistic swapin collapsing to conservative]
> >
> > I think general approach(i.e., less effective than targeting
> > implement for your own specific goal but less hacky and better job
> > for many cases) is to rely/improve on the swap readahead.
> > If most of subpages of a THP page are really workingset, swap readahead
> > could work well.
> >
> > Yeah, it's fairly vague feedback so sorry if I miss something clear.
> 
> Yes.  I want to go to the direction that to swap in 512 pages together.
> And I think it is a good opportunity to discuss that now.  The advantages
> of swapping in 512 pages together are:
> 
> - Improve the performance of swapping in IO via turning small read size
>   into 512 pages big read size.
> 
> - Keep THP across swap out/in.  With the memory size become more and
>   more large, the 4k pages bring more and more burden to memory
>   management.  One solution is to use 2M pages as much as possible, that
>   will reduce the management burden greatly, such as much reduced length
>   of LRU list, etc.
> 
> The disadvantage are:
> 
> - Increase the memory pressure when swap in THP.
> 
> - Some pages swapped in may not needed in the near future.
> 
> Because of the disadvantages, the 512 pages swapping in should be made
> optional.  But I don't think we should make it impossible.

Yeb. No need to make it impossible but your design shouldn't be coupled
with non-existing feature yet.

> 
> Best Regards,
> Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1482175

From"Huang\, Ying" <ying.huang@intel.com>
Date2016-09-13 08:50 +0200
Message-ID<sgPHz-6Bw-15@gated-at.bofh.it>
In reply to#1482160
Minchan Kim <minchan@kernel.org> writes:

> Hi Huang,
>
> On Fri, Sep 09, 2016 at 01:35:12PM -0700, Huang, Ying wrote:
>
> < snip >
>
>> >> Recently, the performance of the storage devices improved so fast that
>> >> we cannot saturate the disk bandwidth when do page swap out even on a
>> >> high-end server machine.  Because the performance of the storage
>> >> device improved faster than that of CPU.  And it seems that the trend
>> >> will not change in the near future.  On the other hand, the THP
>> >> becomes more and more popular because of increased memory size.  So it
>> >> becomes necessary to optimize THP swap performance.
>> >> 
>> >> The advantages of the THP swap support include:
>> >> 
>> >> - Batch the swap operations for the THP to reduce lock
>> >>   acquiring/releasing, including allocating/freeing the swap space,
>> >>   adding/deleting to/from the swap cache, and writing/reading the swap
>> >>   space, etc.  This will help improve the performance of the THP swap.
>> >> 
>> >> - The THP swap space read/write will be 2M sequential IO.  It is
>> >>   particularly helpful for the swap read, which usually are 4k random
>> >>   IO.  This will improve the performance of the THP swap too.
>> >> 
>> >> - It will help the memory fragmentation, especially when the THP is
>> >>   heavily used by the applications.  The 2M continuous pages will be
>> >>   free up after THP swapping out.
>> >
>> > I just read patchset right now and still doubt why the all changes
>> > should be coupled with THP tightly. Many parts(e.g., you introduced
>> > or modifying existing functions for making them THP specific) could
>> > just take page_list and the number of pages then would handle them
>> > without THP awareness.
>> 
>> I am glad if my change could help normal pages swapping too.  And we can
>> change these functions to work for normal pages when necessary.
>
> Sure but it would be less painful that THP awareness swapout is
> based on multiple normal pages swapout. For exmaple, we don't
> touch delay THP split part(i.e., split a THP into 512 pages like
> as-is) and enhances swapout further like Tim's suggestion
> for mulitple normal pages swapout. With that, it might be enough
> for fast-storage without needing THP awareness.
>
> My *point* is let's approach step by step.
> First of all, go with batching normal pages swapout and if it's
> not enough, dive into further optimization like introducing
> THP-aware swapout.
>
> I believe it's natural development process to evolve things
> without over-engineering.

My target is not only the THP swap out acceleration, but also the full
THP swap out/in support without splitting THP.  This patchset is just
the first step of the full THP swap support.

>> > For example, if the nr_pages is larger than SWAPFILE_CLUSTER, we
>> > can try to allocate new cluster. With that, we could allocate new
>> > clusters to meet nr_pages requested or bail out if we fail to allocate
>> > and fallback to 0-order page swapout. With that, swap layer could
>> > support multiple order-0 pages by batch.
>> >
>> > IMO, I really want to land Tim Chen's batching swapout work first.
>> > With Tim Chen's work, I expect we can make better refactoring
>> > for batching swap before adding more confuse to the swap layer.
>> > (I expect it would share several pieces of code for or would be base
>> > for batching allocation of swapcache, swapslot)
>> 
>> I don't think there is hard conflict between normal pages swapping
>> optimizing and THP swap optimizing.  Some code may be shared between
>> them.  That is good for both sides.
>> 
>> > After that, we could enhance swap for big contiguous batching
>> > like THP and finally we might make it be aware of THP specific to
>> > enhance further.
>> >
>> > A thing I remember you aruged: you want to swapin 512 pages
>> > all at once unconditionally. It's really worth to discuss if
>> > your design is going for the way.
>> > I doubt it's generally good idea. Because, currently, we try to
>> > swap in swapped out pages in THP page with conservative approach
>> > but your direction is going to opposite way.
>> >
>> > [mm, thp: convert from optimistic swapin collapsing to conservative]
>> >
>> > I think general approach(i.e., less effective than targeting
>> > implement for your own specific goal but less hacky and better job
>> > for many cases) is to rely/improve on the swap readahead.
>> > If most of subpages of a THP page are really workingset, swap readahead
>> > could work well.
>> >
>> > Yeah, it's fairly vague feedback so sorry if I miss something clear.
>> 
>> Yes.  I want to go to the direction that to swap in 512 pages together.
>> And I think it is a good opportunity to discuss that now.  The advantages
>> of swapping in 512 pages together are:
>> 
>> - Improve the performance of swapping in IO via turning small read size
>>   into 512 pages big read size.
>> 
>> - Keep THP across swap out/in.  With the memory size become more and
>>   more large, the 4k pages bring more and more burden to memory
>>   management.  One solution is to use 2M pages as much as possible, that
>>   will reduce the management burden greatly, such as much reduced length
>>   of LRU list, etc.
>> 
>> The disadvantage are:
>> 
>> - Increase the memory pressure when swap in THP.
>> 
>> - Some pages swapped in may not needed in the near future.
>> 
>> Because of the disadvantages, the 512 pages swapping in should be made
>> optional.  But I don't think we should make it impossible.
>
> Yeb. No need to make it impossible but your design shouldn't be coupled
> with non-existing feature yet.

Sorry, what is the "non-existing feature"?  The full THP swap out/in
support without splitting THP?  If so, this patchset is the just the
first step of that.  I plan to finish the the full THP swap out/in
support in 3 steps:

1. Delay splitting the THP after adding it into swap cache

2. Delay splitting the THP after swapping out being completed

3. Avoid splitting the THP during swap out, and swap in the full THP if
   possible

I plan to do it step by step to make it easier to review the code.

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1482189 — Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out

FromMinchan Kim <minchan@kernel.org>
Date2016-09-13 09:10 +0200
SubjectRe: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out
Message-ID<sgQ0W-6Y0-29@gated-at.bofh.it>
In reply to#1482175
On Tue, Sep 13, 2016 at 02:40:00PM +0800, Huang, Ying wrote:
> Minchan Kim <minchan@kernel.org> writes:
> 
> > Hi Huang,
> >
> > On Fri, Sep 09, 2016 at 01:35:12PM -0700, Huang, Ying wrote:
> >
> > < snip >
> >
> >> >> Recently, the performance of the storage devices improved so fast that
> >> >> we cannot saturate the disk bandwidth when do page swap out even on a
> >> >> high-end server machine.  Because the performance of the storage
> >> >> device improved faster than that of CPU.  And it seems that the trend
> >> >> will not change in the near future.  On the other hand, the THP
> >> >> becomes more and more popular because of increased memory size.  So it
> >> >> becomes necessary to optimize THP swap performance.
> >> >> 
> >> >> The advantages of the THP swap support include:
> >> >> 
> >> >> - Batch the swap operations for the THP to reduce lock
> >> >>   acquiring/releasing, including allocating/freeing the swap space,
> >> >>   adding/deleting to/from the swap cache, and writing/reading the swap
> >> >>   space, etc.  This will help improve the performance of the THP swap.
> >> >> 
> >> >> - The THP swap space read/write will be 2M sequential IO.  It is
> >> >>   particularly helpful for the swap read, which usually are 4k random
> >> >>   IO.  This will improve the performance of the THP swap too.
> >> >> 
> >> >> - It will help the memory fragmentation, especially when the THP is
> >> >>   heavily used by the applications.  The 2M continuous pages will be
> >> >>   free up after THP swapping out.
> >> >
> >> > I just read patchset right now and still doubt why the all changes
> >> > should be coupled with THP tightly. Many parts(e.g., you introduced
> >> > or modifying existing functions for making them THP specific) could
> >> > just take page_list and the number of pages then would handle them
> >> > without THP awareness.
> >> 
> >> I am glad if my change could help normal pages swapping too.  And we can
> >> change these functions to work for normal pages when necessary.
> >
> > Sure but it would be less painful that THP awareness swapout is
> > based on multiple normal pages swapout. For exmaple, we don't
> > touch delay THP split part(i.e., split a THP into 512 pages like
> > as-is) and enhances swapout further like Tim's suggestion
> > for mulitple normal pages swapout. With that, it might be enough
> > for fast-storage without needing THP awareness.
> >
> > My *point* is let's approach step by step.
> > First of all, go with batching normal pages swapout and if it's
> > not enough, dive into further optimization like introducing
> > THP-aware swapout.
> >
> > I believe it's natural development process to evolve things
> > without over-engineering.
> 
> My target is not only the THP swap out acceleration, but also the full
> THP swap out/in support without splitting THP.  This patchset is just
> the first step of the full THP swap support.
> 
> >> > For example, if the nr_pages is larger than SWAPFILE_CLUSTER, we
> >> > can try to allocate new cluster. With that, we could allocate new
> >> > clusters to meet nr_pages requested or bail out if we fail to allocate
> >> > and fallback to 0-order page swapout. With that, swap layer could
> >> > support multiple order-0 pages by batch.
> >> >
> >> > IMO, I really want to land Tim Chen's batching swapout work first.
> >> > With Tim Chen's work, I expect we can make better refactoring
> >> > for batching swap before adding more confuse to the swap layer.
> >> > (I expect it would share several pieces of code for or would be base
> >> > for batching allocation of swapcache, swapslot)
> >> 
> >> I don't think there is hard conflict between normal pages swapping
> >> optimizing and THP swap optimizing.  Some code may be shared between
> >> them.  That is good for both sides.
> >> 
> >> > After that, we could enhance swap for big contiguous batching
> >> > like THP and finally we might make it be aware of THP specific to
> >> > enhance further.
> >> >
> >> > A thing I remember you aruged: you want to swapin 512 pages
> >> > all at once unconditionally. It's really worth to discuss if
> >> > your design is going for the way.
> >> > I doubt it's generally good idea. Because, currently, we try to
> >> > swap in swapped out pages in THP page with conservative approach
> >> > but your direction is going to opposite way.
> >> >
> >> > [mm, thp: convert from optimistic swapin collapsing to conservative]
> >> >
> >> > I think general approach(i.e., less effective than targeting
> >> > implement for your own specific goal but less hacky and better job
> >> > for many cases) is to rely/improve on the swap readahead.
> >> > If most of subpages of a THP page are really workingset, swap readahead
> >> > could work well.
> >> >
> >> > Yeah, it's fairly vague feedback so sorry if I miss something clear.
> >> 
> >> Yes.  I want to go to the direction that to swap in 512 pages together.
> >> And I think it is a good opportunity to discuss that now.  The advantages
> >> of swapping in 512 pages together are:
> >> 
> >> - Improve the performance of swapping in IO via turning small read size
> >>   into 512 pages big read size.
> >> 
> >> - Keep THP across swap out/in.  With the memory size become more and
> >>   more large, the 4k pages bring more and more burden to memory
> >>   management.  One solution is to use 2M pages as much as possible, that
> >>   will reduce the management burden greatly, such as much reduced length
> >>   of LRU list, etc.
> >> 
> >> The disadvantage are:
> >> 
> >> - Increase the memory pressure when swap in THP.
> >> 
> >> - Some pages swapped in may not needed in the near future.
> >> 
> >> Because of the disadvantages, the 512 pages swapping in should be made
> >> optional.  But I don't think we should make it impossible.
> >
> > Yeb. No need to make it impossible but your design shouldn't be coupled
> > with non-existing feature yet.
> 
> Sorry, what is the "non-existing feature"?  The full THP swap out/in

THP swapin.

You said you increased cluster size to fit a THP size for recording
some meta in there for THP swapin.

You gave number about how scale bad current swapout so try to enhance
that path. I agree it alghouth I don't like your approach for first step.
However, you didn't give any clue why we should swap in a THP. How bad
current conservative swapin from khugepagd is really bad and why cannot
enhance that.

> support without splitting THP?  If so, this patchset is the just the
> first step of that.  I plan to finish the the full THP swap out/in
> support in 3 steps:
> 
> 1. Delay splitting the THP after adding it into swap cache
> 
> 2. Delay splitting the THP after swapping out being completed
> 
> 3. Avoid splitting the THP during swap out, and swap in the full THP if
>    possible
> 
> I plan to do it step by step to make it easier to review the code.

1. If we solve batching swapout, then how is THP split for swapout bad?
2. Also, how is current conservatie swapin from khugepaged bad?

I think it's one of decision point for the motivation of your work
and for 1, we need batching swapout feature.

I am saying again that I'm not against your goal but only concern
is approach. If you don't agree, please ignore me.

[toc] | [prev] | [next] | [standalone]


#1482277

From"Huang\, Ying" <ying.huang@intel.com>
Date2016-09-13 11:00 +0200
Message-ID<sgRJp-7Xu-43@gated-at.bofh.it>
In reply to#1482189
Minchan Kim <minchan@kernel.org> writes:
> On Tue, Sep 13, 2016 at 02:40:00PM +0800, Huang, Ying wrote:
>> Minchan Kim <minchan@kernel.org> writes:
>> 
>> > Hi Huang,
>> >
>> > On Fri, Sep 09, 2016 at 01:35:12PM -0700, Huang, Ying wrote:
>> >
>> > < snip >
>> >
>> >> >> Recently, the performance of the storage devices improved so fast that
>> >> >> we cannot saturate the disk bandwidth when do page swap out even on a
>> >> >> high-end server machine.  Because the performance of the storage
>> >> >> device improved faster than that of CPU.  And it seems that the trend
>> >> >> will not change in the near future.  On the other hand, the THP
>> >> >> becomes more and more popular because of increased memory size.  So it
>> >> >> becomes necessary to optimize THP swap performance.
>> >> >> 
>> >> >> The advantages of the THP swap support include:
>> >> >> 
>> >> >> - Batch the swap operations for the THP to reduce lock
>> >> >>   acquiring/releasing, including allocating/freeing the swap space,
>> >> >>   adding/deleting to/from the swap cache, and writing/reading the swap
>> >> >>   space, etc.  This will help improve the performance of the THP swap.
>> >> >> 
>> >> >> - The THP swap space read/write will be 2M sequential IO.  It is
>> >> >>   particularly helpful for the swap read, which usually are 4k random
>> >> >>   IO.  This will improve the performance of the THP swap too.
>> >> >> 
>> >> >> - It will help the memory fragmentation, especially when the THP is
>> >> >>   heavily used by the applications.  The 2M continuous pages will be
>> >> >>   free up after THP swapping out.
>> >> >
>> >> > I just read patchset right now and still doubt why the all changes
>> >> > should be coupled with THP tightly. Many parts(e.g., you introduced
>> >> > or modifying existing functions for making them THP specific) could
>> >> > just take page_list and the number of pages then would handle them
>> >> > without THP awareness.
>> >> 
>> >> I am glad if my change could help normal pages swapping too.  And we can
>> >> change these functions to work for normal pages when necessary.
>> >
>> > Sure but it would be less painful that THP awareness swapout is
>> > based on multiple normal pages swapout. For exmaple, we don't
>> > touch delay THP split part(i.e., split a THP into 512 pages like
>> > as-is) and enhances swapout further like Tim's suggestion
>> > for mulitple normal pages swapout. With that, it might be enough
>> > for fast-storage without needing THP awareness.
>> >
>> > My *point* is let's approach step by step.
>> > First of all, go with batching normal pages swapout and if it's
>> > not enough, dive into further optimization like introducing
>> > THP-aware swapout.
>> >
>> > I believe it's natural development process to evolve things
>> > without over-engineering.
>> 
>> My target is not only the THP swap out acceleration, but also the full
>> THP swap out/in support without splitting THP.  This patchset is just
>> the first step of the full THP swap support.
>> 
>> >> > For example, if the nr_pages is larger than SWAPFILE_CLUSTER, we
>> >> > can try to allocate new cluster. With that, we could allocate new
>> >> > clusters to meet nr_pages requested or bail out if we fail to allocate
>> >> > and fallback to 0-order page swapout. With that, swap layer could
>> >> > support multiple order-0 pages by batch.
>> >> >
>> >> > IMO, I really want to land Tim Chen's batching swapout work first.
>> >> > With Tim Chen's work, I expect we can make better refactoring
>> >> > for batching swap before adding more confuse to the swap layer.
>> >> > (I expect it would share several pieces of code for or would be base
>> >> > for batching allocation of swapcache, swapslot)
>> >> 
>> >> I don't think there is hard conflict between normal pages swapping
>> >> optimizing and THP swap optimizing.  Some code may be shared between
>> >> them.  That is good for both sides.
>> >> 
>> >> > After that, we could enhance swap for big contiguous batching
>> >> > like THP and finally we might make it be aware of THP specific to
>> >> > enhance further.
>> >> >
>> >> > A thing I remember you aruged: you want to swapin 512 pages
>> >> > all at once unconditionally. It's really worth to discuss if
>> >> > your design is going for the way.
>> >> > I doubt it's generally good idea. Because, currently, we try to
>> >> > swap in swapped out pages in THP page with conservative approach
>> >> > but your direction is going to opposite way.
>> >> >
>> >> > [mm, thp: convert from optimistic swapin collapsing to conservative]
>> >> >
>> >> > I think general approach(i.e., less effective than targeting
>> >> > implement for your own specific goal but less hacky and better job
>> >> > for many cases) is to rely/improve on the swap readahead.
>> >> > If most of subpages of a THP page are really workingset, swap readahead
>> >> > could work well.
>> >> >
>> >> > Yeah, it's fairly vague feedback so sorry if I miss something clear.
>> >> 
>> >> Yes.  I want to go to the direction that to swap in 512 pages together.
>> >> And I think it is a good opportunity to discuss that now.  The advantages
>> >> of swapping in 512 pages together are:
>> >> 
>> >> - Improve the performance of swapping in IO via turning small read size
>> >>   into 512 pages big read size.
>> >> 
>> >> - Keep THP across swap out/in.  With the memory size become more and
>> >>   more large, the 4k pages bring more and more burden to memory
>> >>   management.  One solution is to use 2M pages as much as possible, that
>> >>   will reduce the management burden greatly, such as much reduced length
>> >>   of LRU list, etc.
>> >> 
>> >> The disadvantage are:
>> >> 
>> >> - Increase the memory pressure when swap in THP.
>> >> 
>> >> - Some pages swapped in may not needed in the near future.
>> >> 
>> >> Because of the disadvantages, the 512 pages swapping in should be made
>> >> optional.  But I don't think we should make it impossible.
>> >
>> > Yeb. No need to make it impossible but your design shouldn't be coupled
>> > with non-existing feature yet.
>> 
>> Sorry, what is the "non-existing feature"?  The full THP swap out/in
>
> THP swapin.
>
> You said you increased cluster size to fit a THP size for recording
> some meta in there for THP swapin.

And to find the head of the THP to swap in the whole THP when an address
in the middle of a THP is accessed.

> You gave number about how scale bad current swapout so try to enhance
> that path. I agree it alghouth I don't like your approach for first step.
> However, you didn't give any clue why we should swap in a THP. How bad
> current conservative swapin from khugepagd is really bad and why cannot
> enhance that.
>
>> support without splitting THP?  If so, this patchset is the just the
>> first step of that.  I plan to finish the the full THP swap out/in
>> support in 3 steps:
>> 
>> 1. Delay splitting the THP after adding it into swap cache
>> 
>> 2. Delay splitting the THP after swapping out being completed
>> 
>> 3. Avoid splitting the THP during swap out, and swap in the full THP if
>>    possible
>> 
>> I plan to do it step by step to make it easier to review the code.
>
> 1. If we solve batching swapout, then how is THP split for swapout bad?
> 2. Also, how is current conservatie swapin from khugepaged bad?
>
> I think it's one of decision point for the motivation of your work
> and for 1, we need batching swapout feature.
>
> I am saying again that I'm not against your goal but only concern
> is approach. If you don't agree, please ignore me.

I am glad to discuss my final goal, that is, swapping out/in the full
THP without splitting.  Why I want to do that is copied as below,

>> >> The advantages of swapping in 512 pages together are:
>> >> 
>> >> - Improve the performance of swapping in IO via turning small read size
>> >>   into 512 pages big read size.
>> >> 
>> >> - Keep THP across swap out/in.  With the memory size become more and
>> >>   more large, the 4k pages bring more and more burden to memory
>> >>   management.  One solution is to use 2M pages as much as possible, that
>> >>   will reduce the management burden greatly, such as much reduced length
>> >>   of LRU list, etc.

- Avoid CPU time for splitting, collapsing THP across swap out/in.

>> >> 
>> >> The disadvantage are:
>> >> 
>> >> - Increase the memory pressure when swap in THP.
>> >> 
>> >> - Some pages swapped in may not needed in the near future.

I think it is important to use 2M pages as much as possible to deal with
the big memory problem.  Do you agree?

Best Regards,
Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1482290 — Re: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out

FromMinchan Kim <minchan@kernel.org>
Date2016-09-13 11:20 +0200
SubjectRe: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out
Message-ID<sgS2J-8jW-15@gated-at.bofh.it>
In reply to#1482277
On Tue, Sep 13, 2016 at 04:53:49PM +0800, Huang, Ying wrote:
> Minchan Kim <minchan@kernel.org> writes:
> > On Tue, Sep 13, 2016 at 02:40:00PM +0800, Huang, Ying wrote:
> >> Minchan Kim <minchan@kernel.org> writes:
> >> 
> >> > Hi Huang,
> >> >
> >> > On Fri, Sep 09, 2016 at 01:35:12PM -0700, Huang, Ying wrote:
> >> >
> >> > < snip >
> >> >
> >> >> >> Recently, the performance of the storage devices improved so fast that
> >> >> >> we cannot saturate the disk bandwidth when do page swap out even on a
> >> >> >> high-end server machine.  Because the performance of the storage
> >> >> >> device improved faster than that of CPU.  And it seems that the trend
> >> >> >> will not change in the near future.  On the other hand, the THP
> >> >> >> becomes more and more popular because of increased memory size.  So it
> >> >> >> becomes necessary to optimize THP swap performance.
> >> >> >> 
> >> >> >> The advantages of the THP swap support include:
> >> >> >> 
> >> >> >> - Batch the swap operations for the THP to reduce lock
> >> >> >>   acquiring/releasing, including allocating/freeing the swap space,
> >> >> >>   adding/deleting to/from the swap cache, and writing/reading the swap
> >> >> >>   space, etc.  This will help improve the performance of the THP swap.
> >> >> >> 
> >> >> >> - The THP swap space read/write will be 2M sequential IO.  It is
> >> >> >>   particularly helpful for the swap read, which usually are 4k random
> >> >> >>   IO.  This will improve the performance of the THP swap too.
> >> >> >> 
> >> >> >> - It will help the memory fragmentation, especially when the THP is
> >> >> >>   heavily used by the applications.  The 2M continuous pages will be
> >> >> >>   free up after THP swapping out.
> >> >> >
> >> >> > I just read patchset right now and still doubt why the all changes
> >> >> > should be coupled with THP tightly. Many parts(e.g., you introduced
> >> >> > or modifying existing functions for making them THP specific) could
> >> >> > just take page_list and the number of pages then would handle them
> >> >> > without THP awareness.
> >> >> 
> >> >> I am glad if my change could help normal pages swapping too.  And we can
> >> >> change these functions to work for normal pages when necessary.
> >> >
> >> > Sure but it would be less painful that THP awareness swapout is
> >> > based on multiple normal pages swapout. For exmaple, we don't
> >> > touch delay THP split part(i.e., split a THP into 512 pages like
> >> > as-is) and enhances swapout further like Tim's suggestion
> >> > for mulitple normal pages swapout. With that, it might be enough
> >> > for fast-storage without needing THP awareness.
> >> >
> >> > My *point* is let's approach step by step.
> >> > First of all, go with batching normal pages swapout and if it's
> >> > not enough, dive into further optimization like introducing
> >> > THP-aware swapout.
> >> >
> >> > I believe it's natural development process to evolve things
> >> > without over-engineering.
> >> 
> >> My target is not only the THP swap out acceleration, but also the full
> >> THP swap out/in support without splitting THP.  This patchset is just
> >> the first step of the full THP swap support.
> >> 
> >> >> > For example, if the nr_pages is larger than SWAPFILE_CLUSTER, we
> >> >> > can try to allocate new cluster. With that, we could allocate new
> >> >> > clusters to meet nr_pages requested or bail out if we fail to allocate
> >> >> > and fallback to 0-order page swapout. With that, swap layer could
> >> >> > support multiple order-0 pages by batch.
> >> >> >
> >> >> > IMO, I really want to land Tim Chen's batching swapout work first.
> >> >> > With Tim Chen's work, I expect we can make better refactoring
> >> >> > for batching swap before adding more confuse to the swap layer.
> >> >> > (I expect it would share several pieces of code for or would be base
> >> >> > for batching allocation of swapcache, swapslot)
> >> >> 
> >> >> I don't think there is hard conflict between normal pages swapping
> >> >> optimizing and THP swap optimizing.  Some code may be shared between
> >> >> them.  That is good for both sides.
> >> >> 
> >> >> > After that, we could enhance swap for big contiguous batching
> >> >> > like THP and finally we might make it be aware of THP specific to
> >> >> > enhance further.
> >> >> >
> >> >> > A thing I remember you aruged: you want to swapin 512 pages
> >> >> > all at once unconditionally. It's really worth to discuss if
> >> >> > your design is going for the way.
> >> >> > I doubt it's generally good idea. Because, currently, we try to
> >> >> > swap in swapped out pages in THP page with conservative approach
> >> >> > but your direction is going to opposite way.
> >> >> >
> >> >> > [mm, thp: convert from optimistic swapin collapsing to conservative]
> >> >> >
> >> >> > I think general approach(i.e., less effective than targeting
> >> >> > implement for your own specific goal but less hacky and better job
> >> >> > for many cases) is to rely/improve on the swap readahead.
> >> >> > If most of subpages of a THP page are really workingset, swap readahead
> >> >> > could work well.
> >> >> >
> >> >> > Yeah, it's fairly vague feedback so sorry if I miss something clear.
> >> >> 
> >> >> Yes.  I want to go to the direction that to swap in 512 pages together.
> >> >> And I think it is a good opportunity to discuss that now.  The advantages
> >> >> of swapping in 512 pages together are:
> >> >> 
> >> >> - Improve the performance of swapping in IO via turning small read size
> >> >>   into 512 pages big read size.
> >> >> 
> >> >> - Keep THP across swap out/in.  With the memory size become more and
> >> >>   more large, the 4k pages bring more and more burden to memory
> >> >>   management.  One solution is to use 2M pages as much as possible, that
> >> >>   will reduce the management burden greatly, such as much reduced length
> >> >>   of LRU list, etc.
> >> >> 
> >> >> The disadvantage are:
> >> >> 
> >> >> - Increase the memory pressure when swap in THP.
> >> >> 
> >> >> - Some pages swapped in may not needed in the near future.
> >> >> 
> >> >> Because of the disadvantages, the 512 pages swapping in should be made
> >> >> optional.  But I don't think we should make it impossible.
> >> >
> >> > Yeb. No need to make it impossible but your design shouldn't be coupled
> >> > with non-existing feature yet.
> >> 
> >> Sorry, what is the "non-existing feature"?  The full THP swap out/in
> >
> > THP swapin.
> >
> > You said you increased cluster size to fit a THP size for recording
> > some meta in there for THP swapin.
> 
> And to find the head of the THP to swap in the whole THP when an address
> in the middle of a THP is accessed.
> 
> > You gave number about how scale bad current swapout so try to enhance
> > that path. I agree it alghouth I don't like your approach for first step.
> > However, you didn't give any clue why we should swap in a THP. How bad
> > current conservative swapin from khugepagd is really bad and why cannot
> > enhance that.
> >
> >> support without splitting THP?  If so, this patchset is the just the
> >> first step of that.  I plan to finish the the full THP swap out/in
> >> support in 3 steps:
> >> 
> >> 1. Delay splitting the THP after adding it into swap cache
> >> 
> >> 2. Delay splitting the THP after swapping out being completed
> >> 
> >> 3. Avoid splitting the THP during swap out, and swap in the full THP if
> >>    possible
> >> 
> >> I plan to do it step by step to make it easier to review the code.
> >
> > 1. If we solve batching swapout, then how is THP split for swapout bad?
> > 2. Also, how is current conservatie swapin from khugepaged bad?
> >
> > I think it's one of decision point for the motivation of your work
> > and for 1, we need batching swapout feature.
> >
> > I am saying again that I'm not against your goal but only concern
> > is approach. If you don't agree, please ignore me.
> 
> I am glad to discuss my final goal, that is, swapping out/in the full
> THP without splitting.  Why I want to do that is copied as below,

Yes, it's your *final* goal but what if it couldn't be acceptable
on second step you mentioned above, for example?

        Unncessary binded implementation to rejected work.

If you want to achieve your goal step by step, please consider if
one of step you are thinking could be rejected but steps already
merged should be self-contained without side-effect.
If it's hard, send full patchset all at once so reviewers can think
what you want of right direction and implementation is good for it.

> 
> >> >> The advantages of swapping in 512 pages together are:
> >> >> 
> >> >> - Improve the performance of swapping in IO via turning small read size
> >> >>   into 512 pages big read size.
> >> >> 
> >> >> - Keep THP across swap out/in.  With the memory size become more and
> >> >>   more large, the 4k pages bring more and more burden to memory
> >> >>   management.  One solution is to use 2M pages as much as possible, that
> >> >>   will reduce the management burden greatly, such as much reduced length
> >> >>   of LRU list, etc.
> 
> - Avoid CPU time for splitting, collapsing THP across swap out/in.

Yes, if you want, please give us how bad it is.

> 
> >> >> 
> >> >> The disadvantage are:
> >> >> 
> >> >> - Increase the memory pressure when swap in THP.
> >> >> 
> >> >> - Some pages swapped in may not needed in the near future.
> 
> I think it is important to use 2M pages as much as possible to deal with
> the big memory problem.  Do you agree?

There is no number I can think what is current problems and
how it is popular thesedays so I don't agree.

> 
> Best Regards,
> Huang, Ying

[toc] | [prev] | [next] | [standalone]


#1482847 — RE: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out

From"Chen, Tim C" <tim.c.chen@intel.com>
Date2016-09-14 02:00 +0200
SubjectRE: [PATCH -v3 00/10] THP swap: Delay splitting THP during swapping out
Message-ID<sh5Mm-t4-11@gated-at.bofh.it>
In reply to#1482290
>>
>> - Avoid CPU time for splitting, collapsing THP across swap out/in.
>
>Yes, if you want, please give us how bad it is.
>

It could be pretty bad.  In an experiment with THP turned on and we
enter swap, 50% of the cpu are spent in the page compaction path.  
So if we could deal with units of large page for swap, the splitting
and compaction of ordinary pages to large page overhead could be avoided.

   51.89%    51.89%            :1688  [kernel.kallsyms]   [k] pageblock_pfn_to_page                       
                      |
                      --- pageblock_pfn_to_page
                         |          
                         |--64.57%-- compaction_alloc
                         |          migrate_pages
                         |          compact_zone
                         |          compact_zone_order
                         |          try_to_compact_pages
                         |          __alloc_pages_direct_compact
                         |          __alloc_pages_nodemask
                         |          alloc_pages_vma
                         |          do_huge_pmd_anonymous_page
                         |          handle_mm_fault
                         |          __do_page_fault
                         |          do_page_fault
                         |          page_fault
                         |          0x401d9a
                         |          
                         |--34.62%-- compact_zone
                         |          compact_zone_order
                         |          try_to_compact_pages
                         |          __alloc_pages_direct_compact
                         |          __alloc_pages_nodemask
                         |          alloc_pages_vma
                         |          do_huge_pmd_anonymous_page
                         |          handle_mm_fault
                         |          __do_page_fault
                         |          do_page_fault
                         |          page_fault
                         |          0x401d9a
                          --0.81%-- [...]

Tim

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | linux.kernel


csiph-web