Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1506875 > unrolled thread

[RFC 0/8] Define coherent device memory node

Started byAnshuman Khandual <khandual@linux.vnet.ibm.com>
First post2016-10-24 06:40 +0200
Last post2016-10-24 21:40 +0200
Articles 20 on this page of 36 — 6 participants

Back to article view | Back to linux.kernel


Contents

  [RFC 0/8] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:40 +0200
    [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:40 +0200
      Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB  allocation paths Dave Hansen <dave.hansen@intel.com> - 2016-10-24 19:20 +0200
        Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-25 06:20 +0200
          Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB  allocation paths Balbir Singh <bsingharora@gmail.com> - 2016-10-25 09:20 +0200
            Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB  allocation paths Balbir Singh <bsingharora@gmail.com> - 2016-10-25 09:30 +0200
    [RFC 7/8] mm: Add a new migration function migrate_virtual_range() Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:40 +0200
    [DEBUG 06/10] mm: Export definition of 'zone_names' array through mmzone.h Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
    [DEBUG 00/10] Test and debug patches for coherent device memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
      [DEBUG 02/10] powerpc/mm: Create numa nodes for hotplug memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
      [DEBUG 03/10] powerpc/mm: Allow memory hotplug into a memory less node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
      [DEBUG 08/10] powerpc: Enable CONFIG_MOVABLE_NODE for PPC64 platform Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
      [DEBUG 01/10] dt-bindings: Add doc for ibm,hotplug-aperture Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
      [DEBUG 05/10] powerpc/mm: Identify isolation seeking coherent memory nodes during boot Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
      [DEBUG 10/10] test: Add a script to perform random VMA migrations across nodes Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
      [DEBUG 07/10] mm: Add debugfs interface to dump each node's zonelist information Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
      [DEBUG 04/10] mm: Enable CONFIG_MOVABLE_NODE on powerpc Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
      [DEBUG 09/10] drivers: Add two drivers for coherent device memory tests Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-24 06:50 +0200
    Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-24 19:10 +0200
      Re: [RFC 0/8] Define coherent device memory node "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-25 06:30 +0200
        Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-25 17:20 +0200
          Re: [RFC 0/8] Define coherent device memory node "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-26 13:10 +0200
            Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-26 18:10 +0200
      Re: [RFC 0/8] Define coherent device memory node "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-25 07:10 +0200
        Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-25 17:40 +0200
          Re: [RFC 0/8] Define coherent device memory node "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-10-25 19:40 +0200
            Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-25 21:00 +0200
              Re: [RFC 0/8] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-26 13:20 +0200
                Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-26 18:10 +0200
              Re: [RFC 0/8] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-10-26 15:00 +0200
                Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-26 18:30 +0200
      Re: [RFC 0/8] Define coherent device memory node Balbir Singh <bsingharora@gmail.com> - 2016-10-25 14:10 +0200
        Re: [RFC 0/8] Define coherent device memory node Jerome Glisse <j.glisse@gmail.com> - 2016-10-25 17:30 +0200
    Re: [RFC 0/8] Define coherent device memory node Dave Hansen <dave.hansen@intel.com> - 2016-10-24 20:10 +0200
      Re: [RFC 0/8] Define coherent device memory node David Nellans <dnellans@nvidia.com> - 2016-10-24 20:40 +0200
        Re: [RFC 0/8] Define coherent device memory node Dave Hansen <dave.hansen@intel.com> - 2016-10-24 21:40 +0200

Page 1 of 2  [1] 2  Next page →


#1506875 — [RFC 0/8] Define coherent device memory node

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:40 +0200
Subject[RFC 0/8] Define coherent device memory node
Message-ID<svFdf-2Fn-3@gated-at.bofh.it>
	There are certain devices like accelerators, GPU cards, network
cards, FPGA cards, PLD cards etc which might contain on board memory. This
on board memory can be coherent along with system RAM and may be accessible
from either the CPU or from the device. The coherency is usually achieved
through synchronizing the cache accesses from either side. This makes the
device memory appear in the same address space as that of the system RAM.
The on board device memory and system RAM are coherent but have differences
in their properties as explained and elaborated below. Following diagram
explains how the coherent device memory appears in the memory address
space.

                +-----------------+         +-----------------+
                |                 |         |                 |
                |       CPU       |         |     DEVICE      |
                |                 |         |                 |
                +-----------------+         +-----------------+
                         |                           |
                         |   Shared Address Space    |
 +---------------------------------------------------------------------+
 |                                             |                       |
 |                                             |                       |
 |                 System RAM                  |     Coherent Memory   |
 |                                             |                       |
 |                                             |                       |
 +---------------------------------------------------------------------+

	User space applications might be interested in using the coherent
device memory either explicitly or implicitly along with the system RAM
utilizing the basic semantics for memory allocation, access and release.
Basically the user applications should be able to allocate memory any where
(system RAM or coherent memory) and then get it accessed either from the
CPU or from the coherent device for various computation or data
transformation purpose. User space really should not be concerned about
memory placement and their subsequent allocations when the memory really
faults because of the access.

	To achieve seamless integration  between system RAM and coherent
device memory it must be able to utilize core memory kernel features like
anon mapping, file mapping, page cache, driver managed pages, HW poisoning,
migrations, reclaim, compaction, etc. Making the coherent device memory
appear as a distinct memory only NUMA node which will be initialized as any
other node with memory can create this integration with currently available
system RAM memory. Also at the same time there should be a differentiating
mark which indicates that this node is a coherent device memory node not
any other memory only system RAM node.
 
	Coherent device memory invariably isn't available until the driver
for the device has been initialized. It is desirable but not required for
the device to support memory offlining for the purposes such as power
management, link management and hardware errors. Kernel allocation should
not come here as it cannot be moved out. Hence coherent device memory
should go inside ZONE_MOVABLE zone instead. This guarantees that kernel
allocations will never be satisfied from this memory and any process having
un-movable pages on this coherent device memory (likely achieved through
pinning later on after initial allocation) can be killed to free up memory
from page table and eventually hot plugging the node out.

	After similar representation as a NUMA node, the coherent memory
might still need some special consideration while being inside the kernel.
There can be a variety of coherent device memory nodes with different
expectations and special considerations from the core kernel. This RFC
discusses only one such scenario where the coherent device memory requires
just isolation.

	Now let us consider in detail the case of a coherent device memory
node which requires isolation. This kind of coherent device memory is on
board an external device attached to the system through a link where there
is a chance of link errors plugging out the entire memory node with it.
More over the memory might also have higher chances of ECC errors as
compared to the system RAM. These are just some possibilities. But the fact
remains that the coherent device memory can have some other different
properties which might not be desirable for some user space applications.
An application should not be exposed to related risks of a device if its
not taking advantage of special features of that device and it's memory.

	Because of the reasons explained above allocations into isolation
based coherent device memory node should further be regulated apart from
earlier requirement of kernel allocations not coming there. User space
allocations should not come here implicitly without the user application
explicitly knowing about it. This summarizes isolation requirement of
certain kind of a coherent device memory node as an example.

	Some coherent memory devices may not require isolation altogether.
Then there might be other coherent memory devices which require some other
special treatment after being part of core memory representation in kernel.
Though the framework suggested by this RFC has made provisions for them, it
has not considered any other kind of requirement other than isolation for
now.

	Though this RFC series currently attempts to implement one such
isolation seeking coherent device memory example, this framework can be
extended to accommodate any present or future coherent memory devices which
will fit the requirement as explained before even with new requirements
other than isolation. In case of isolation seeking coherent device memory
node, there will be other core VM code paths which need to be taken care
before it can be completely isolated as required.

	Core kernel memory features like reclamation, evictions etc. might
need to be restricted or modified on the coherent device memory node as
they can be performance limiting. The RFC does not propose anything on this
yet but it can be looked into later on. For now it just disables Auto NUMA
for any VMA which has coherent device memory.

	Seamless integration of coherent device memory with system memory
will enable various other features, some of which can be listed as follows.

	a. Seamless migrations between system RAM and the coherent memory
	b. Will have asynchronous and high throughput migrations
	c. Be able to allocate huge order pages from these memory regions
	d. Restrict allocations to a large extent to the tasks using the
	   device for workload acceleration

	Before concluding, will look into the reasons why the existing
solutions don't work. There are two basic requirements which have to be
satisfies before the coherent device memory can be integrated with core
kernel seamlessly.

	a. PFN must have struct page
	b. Struct page must able to be inside standard LRU lists

	The above two basic requirements discard the existing method of
device memory representation approaches like these which then requires the
need of creating a new framework.

(1) Traditional ioremap

	a. Memory is mapped into kernel (linear and virtual) and user space
	b. These PFNs do not have struct pages associated with it
	c. These special PFNs are marked with special flags inside the PTE
	d. Cannot participate in core VM functions much because of this
	e. Cannot do easy user space migrations

(2) Zone ZONE_DEVICE

	a. Memory is mapped into kernel and user space
	b. PFNs do have struct pages associated with it
	c. These struct pages are allocated inside it's own memory range
	d. Unfortunately the struct page's union containing LRU has been
	   used for struct dev_pagemap pointer
	e. Hence it cannot be part of any LRU (like Page cache)
	f. Hence file cached mapping cannot reside on these PFNs
	g. Cannot do easy migrations

	I had also explored non LRU representation of this coherent device
memory where the integration with system RAM in the core VM is limited only
to the following functions. Not being inside LRU is definitely going to
reduce the scope of tight integration with system RAM.

(1) Migration support between system RAM and coherent memory
(2) Migration support between various coherent memory nodes
(3) Isolation of the coherent memory
(4) Mapping the coherent memory into user space through driver's
    struct vm_operations
(5) HW poisoning of the coherent memory

	Allocating the entire memory of the coherent device node right
after hot plug into ZONE_MOVABLE (where the memory is already inside the
buddy system) will still expose a time window where other user space
allocations can come into the coherent device memory node and prevent the
intended isolation. So traditional hot plug is not the solution. Hence
started looking into CMA based non LRU solution but then hit the following
roadblocks.

(1) CMA does not support hot plugging of new memory node
	a. CMA area needs to be marked during boot before buddy is
	   initialized
	b. cma_alloc()/cma_release() can happen on the marked area
	c. Should be able to mark the CMA areas just after memory hot plug
	d. cma_alloc()/cma_release() can happen later after the hot plug
	e. This is not currently supported right now

(2) Mapped non LRU migration of pages
	a. Recent work from Michan Kim makes non LRU page migratable
	b. But it still does not support migration of mapped non LRU pages
	c. With non LRU CMA reserved, again there are some additional
	   challenges

	With hot pluggable CMA and non LRU mapped migration support there
may be an alternate approach to represent coherent device memory. Please
do review this RFC proposal and let me know your comments or suggestions.
Thank you.

Anshuman Khandual (8):
  mm: Define coherent device memory node
  mm: Add specialized fallback zonelist for coherent device memory nodes
  mm: Isolate coherent device memory nodes from HugeTLB allocation paths
  mm: Accommodate coherent device memory nodes in MPOL_BIND implementation
  mm: Add new flag VM_CDM for coherent device memory
  mm: Make VM_CDM marked VMAs non migratable
  mm: Add a new migration function migrate_virtual_range()
  mm: Add N_COHERENT_DEVICE node type into node_states[]

 Documentation/ABI/stable/sysfs-devices-node |  7 +++
 drivers/base/node.c                         |  6 +++
 include/linux/mempolicy.h                   | 24 +++++++++
 include/linux/migrate.h                     |  3 ++
 include/linux/mm.h                          |  5 ++
 include/linux/mmzone.h                      | 29 ++++++++++
 include/linux/nodemask.h                    |  3 ++
 mm/Kconfig                                  | 13 +++++
 mm/hugetlb.c                                | 38 ++++++++++++-
 mm/memory_hotplug.c                         | 10 ++++
 mm/mempolicy.c                              | 70 ++++++++++++++++++++++--
 mm/migrate.c                                | 84 +++++++++++++++++++++++++++++
 mm/page_alloc.c                             | 10 ++++
 13 files changed, 295 insertions(+), 7 deletions(-)

-- 
2.1.0

[toc] | [next] | [standalone]


#1506876 — [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:40 +0200
Subject[RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths
Message-ID<svFdf-2Fn-19@gated-at.bofh.it>
In reply to#1506875
This change is part of the isolation requiring coherent device memory nodes
implementation.

Isolation seeking coherent device memory node requires allocation isolation
from implicit memory allocations from user space. Towards that effect, the
memory should not be used for generic HugeTLB page pool allocations. This
modifies relevant functions to skip all coherent memory nodes present on
the system during allocation, freeing and auditing for HugeTLB pages.

Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 mm/hugetlb.c | 38 ++++++++++++++++++++++++++++++++++++--
 1 file changed, 36 insertions(+), 2 deletions(-)

diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index ec49d9e..466a44c 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -1147,6 +1147,9 @@ static int alloc_fresh_gigantic_page(struct hstate *h,
 	int nr_nodes, node;
 
 	for_each_node_mask_to_alloc(h, nr_nodes, node, nodes_allowed) {
+		if (isolated_cdm_node(node))
+			continue;
+
 		page = alloc_fresh_gigantic_page_node(h, node);
 		if (page)
 			return 1;
@@ -1382,6 +1385,9 @@ static int alloc_fresh_huge_page(struct hstate *h, nodemask_t *nodes_allowed)
 	int ret = 0;
 
 	for_each_node_mask_to_alloc(h, nr_nodes, node, nodes_allowed) {
+		if (isolated_cdm_node(node))
+			continue;
+
 		page = alloc_fresh_huge_page_node(h, node);
 		if (page) {
 			ret = 1;
@@ -1410,6 +1416,9 @@ static int free_pool_huge_page(struct hstate *h, nodemask_t *nodes_allowed,
 	int ret = 0;
 
 	for_each_node_mask_to_free(h, nr_nodes, node, nodes_allowed) {
+		if (isolated_cdm_node(node))
+			continue;
+
 		/*
 		 * If we're returning unused surplus pages, only examine
 		 * nodes with surplus pages.
@@ -2028,6 +2037,9 @@ int __weak alloc_bootmem_huge_page(struct hstate *h)
 	for_each_node_mask_to_alloc(h, nr_nodes, node, &node_states[N_MEMORY]) {
 		void *addr;
 
+		if (isolated_cdm_node(node))
+			continue;
+
 		addr = memblock_virt_alloc_try_nid_nopanic(
 				huge_page_size(h), huge_page_size(h),
 				0, BOOTMEM_ALLOC_ACCESSIBLE, node);
@@ -2156,6 +2168,10 @@ static void try_to_free_low(struct hstate *h, unsigned long count,
 	for_each_node_mask(i, *nodes_allowed) {
 		struct page *page, *next;
 		struct list_head *freel = &h->hugepage_freelists[i];
+
+		if (isolated_cdm_node(i))
+			continue;
+
 		list_for_each_entry_safe(page, next, freel, lru) {
 			if (count >= h->nr_huge_pages)
 				return;
@@ -2189,11 +2205,17 @@ static int adjust_pool_surplus(struct hstate *h, nodemask_t *nodes_allowed,
 
 	if (delta < 0) {
 		for_each_node_mask_to_alloc(h, nr_nodes, node, nodes_allowed) {
+			if (isolated_cdm_node(node))
+				continue;
+
 			if (h->surplus_huge_pages_node[node])
 				goto found;
 		}
 	} else {
 		for_each_node_mask_to_free(h, nr_nodes, node, nodes_allowed) {
+			if (isolated_cdm_node(node))
+				continue;
+
 			if (h->surplus_huge_pages_node[node] <
 					h->nr_huge_pages_node[node])
 				goto found;
@@ -2666,6 +2688,10 @@ static void __init hugetlb_register_all_nodes(void)
 
 	for_each_node_state(nid, N_MEMORY) {
 		struct node *node = node_devices[nid];
+
+		if (isolated_cdm_node(nid))
+			continue;
+
 		if (node->dev.id == nid)
 			hugetlb_register_node(node);
 	}
@@ -2819,8 +2845,12 @@ static unsigned int cpuset_mems_nr(unsigned int *array)
 	int node;
 	unsigned int nr = 0;
 
-	for_each_node_mask(node, cpuset_current_mems_allowed)
+	for_each_node_mask(node, cpuset_current_mems_allowed) {
+		if (isolated_cdm_node(node))
+			continue;
+
 		nr += array[node];
+	}
 
 	return nr;
 }
@@ -2940,7 +2970,10 @@ void hugetlb_show_meminfo(void)
 	if (!hugepages_supported())
 		return;
 
-	for_each_node_state(nid, N_MEMORY)
+	for_each_node_state(nid, N_MEMORY) {
+		if (isolated_cdm_node(nid))
+			continue;
+
 		for_each_hstate(h)
 			pr_info("Node %d hugepages_total=%u hugepages_free=%u hugepages_surp=%u hugepages_size=%lukB\n",
 				nid,
@@ -2948,6 +2981,7 @@ void hugetlb_show_meminfo(void)
 				h->free_huge_pages_node[nid],
 				h->surplus_huge_pages_node[nid],
 				1UL << (huge_page_order(h) + PAGE_SHIFT - 10));
+	}
 }
 
 void hugetlb_report_usage(struct seq_file *m, struct mm_struct *mm)
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1507474 — Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths

FromDave Hansen <dave.hansen@intel.com>
Date2016-10-24 19:20 +0200
SubjectRe: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths
Message-ID<svR4J-2cv-3@gated-at.bofh.it>
In reply to#1506876
On 10/23/2016 09:31 PM, Anshuman Khandual wrote:
> This change is part of the isolation requiring coherent device memory nodes
> implementation.
> 
> Isolation seeking coherent device memory node requires allocation isolation
> from implicit memory allocations from user space. Towards that effect, the
> memory should not be used for generic HugeTLB page pool allocations. This
> modifies relevant functions to skip all coherent memory nodes present on
> the system during allocation, freeing and auditing for HugeTLB pages.

This seems really fragile.  You had to hit, what, 18 call sites?  What
are the odds that this is going to stay working?

> @@ -2666,6 +2688,10 @@ static void __init hugetlb_register_all_nodes(void)
>  
>  	for_each_node_state(nid, N_MEMORY) {
>  		struct node *node = node_devices[nid];
> +
> +		if (isolated_cdm_node(nid))
> +			continue;
> +
>  		if (node->dev.id == nid)
>  			hugetlb_register_node(node);
>  	}

This looks to be completely kneecapping hugetlbfs on these cdm nodes.
Is that really what you want?

> @@ -2819,8 +2845,12 @@ static unsigned int cpuset_mems_nr(unsigned int *array)
>  	int node;
>  	unsigned int nr = 0;
>  
> -	for_each_node_mask(node, cpuset_current_mems_allowed)
> +	for_each_node_mask(node, cpuset_current_mems_allowed) {
> +		if (isolated_cdm_node(node))
> +			continue;
> +
>  		nr += array[node];
> +	}
>  
>  	return nr;
>  }
> @@ -2940,7 +2970,10 @@ void hugetlb_show_meminfo(void)
>  	if (!hugepages_supported())
>  		return;
>  
> -	for_each_node_state(nid, N_MEMORY)
> +	for_each_node_state(nid, N_MEMORY) {
> +		if (isolated_cdm_node(nid))
> +			continue;
> +
>  		for_each_hstate(h)
>  			pr_info("Node %d hugepages_total=%u hugepages_free=%u hugepages_surp=%u hugepages_size=%lukB\n",
>  				nid,
> @@ -2948,6 +2981,7 @@ void hugetlb_show_meminfo(void)
>  				h->free_huge_pages_node[nid],
>  				h->surplus_huge_pages_node[nid],
>  				1UL << (huge_page_order(h) + PAGE_SHIFT - 10));
> +	}
>  }

Your patch description talks about removing *implicit* memory
allocations.  But, this removes even the ability to gather *stats* about
huge pages sitting on one of these nodes.  That's a lot more drastic
than just changing implicit policies.

Is that patch description accurate?

It looks to me like you just went through all the for_each_node*() loops
in hugetlb.c and hacked your node check into them indiscriminately.
This totally removes the ability to *do* hugetlb on this nodes.

Isn't there some simpler way to do all this, like maybe changing the
root cpuset to disallow allocations to these nodes?

[toc] | [prev] | [next] | [standalone]


#1507962 — Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths

From"Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com>
Date2016-10-25 06:20 +0200
SubjectRe: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths
Message-ID<sw1nr-zj-1@gated-at.bofh.it>
In reply to#1507474
Dave Hansen <dave.hansen@intel.com> writes:

> On 10/23/2016 09:31 PM, Anshuman Khandual wrote:
>> This change is part of the isolation requiring coherent device memory nodes
>> implementation.
>> 
>> Isolation seeking coherent device memory node requires allocation isolation
>> from implicit memory allocations from user space. Towards that effect, the
>> memory should not be used for generic HugeTLB page pool allocations. This
>> modifies relevant functions to skip all coherent memory nodes present on
>> the system during allocation, freeing and auditing for HugeTLB pages.
>
> This seems really fragile.  You had to hit, what, 18 call sites?  What
> are the odds that this is going to stay working?


I guess a better approach is to introduce new node_states entry such
that we have one that excludes coherent device memory numa nodes. One
possibility is to add N_SYSTEM_MEMORY and N_MEMORY.

Current N_MEMORY becomes N_SYSTEM_MEMORY and N_MEMORY includes
system and device/any other memory which is coherent.

All the isolation can then be achieved based on the nodemask_t used for
allocation. So for allocations we want to avoid from coherent device we
use N_SYSTEM_MEMORY mask or a derivative of that and where we are ok to
allocate from CDM with fallbacks we use N_MEMORY.

All nodes zonelist will have zones from the coherent device nodes but we
will not end up allocating from coherent device node zone due to the
node mask used.


This will also make sure we end up allocating from the correct coherent
device numa node in the presence of multiple of them based on the
distance of the coherent device node from the current executing numa
node.



>
>> @@ -2666,6 +2688,10 @@ static void __init hugetlb_register_all_nodes(void)
>>  
>>  	for_each_node_state(nid, N_MEMORY) {
>>  		struct node *node = node_devices[nid];
>> +
>> +		if (isolated_cdm_node(nid))
>> +			continue;
>> +
>>  		if (node->dev.id == nid)
>>  			hugetlb_register_node(node);
>>  	}
>
> This looks to be completely kneecapping hugetlbfs on these cdm nodes.
> Is that really what you want?
>
>> @@ -2819,8 +2845,12 @@ static unsigned int cpuset_mems_nr(unsigned int *array)
>>  	int node;
>>  	unsigned int nr = 0;
>>  
>> -	for_each_node_mask(node, cpuset_current_mems_allowed)
>> +	for_each_node_mask(node, cpuset_current_mems_allowed) {
>> +		if (isolated_cdm_node(node))
>> +			continue;
>> +
>>  		nr += array[node];
>> +	}
>>  
>>  	return nr;
>>  }
>> @@ -2940,7 +2970,10 @@ void hugetlb_show_meminfo(void)
>>  	if (!hugepages_supported())
>>  		return;
>>  
>> -	for_each_node_state(nid, N_MEMORY)
>> +	for_each_node_state(nid, N_MEMORY) {
>> +		if (isolated_cdm_node(nid))
>> +			continue;
>> +
>>  		for_each_hstate(h)
>>  			pr_info("Node %d hugepages_total=%u hugepages_free=%u hugepages_surp=%u hugepages_size=%lukB\n",
>>  				nid,
>> @@ -2948,6 +2981,7 @@ void hugetlb_show_meminfo(void)
>>  				h->free_huge_pages_node[nid],
>>  				h->surplus_huge_pages_node[nid],
>>  				1UL << (huge_page_order(h) + PAGE_SHIFT - 10));
>> +	}
>>  }
>
> Your patch description talks about removing *implicit* memory
> allocations.  But, this removes even the ability to gather *stats* about
> huge pages sitting on one of these nodes.  That's a lot more drastic
> than just changing implicit policies.
>
> Is that patch description accurate?
>
> It looks to me like you just went through all the for_each_node*() loops
> in hugetlb.c and hacked your node check into them indiscriminately.
> This totally removes the ability to *do* hugetlb on this nodes.
>
> Isn't there some simpler way to do all this, like maybe changing the
> root cpuset to disallow allocations to these nodes?

-aneesh

[toc] | [prev] | [next] | [standalone]


#1508023 — Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths

FromBalbir Singh <bsingharora@gmail.com>
Date2016-10-25 09:20 +0200
SubjectRe: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths
Message-ID<sw4bE-2p1-15@gated-at.bofh.it>
In reply to#1507962

On 25/10/16 15:15, Aneesh Kumar K.V wrote:
> Dave Hansen <dave.hansen@intel.com> writes:
> 
>> On 10/23/2016 09:31 PM, Anshuman Khandual wrote:
>>> This change is part of the isolation requiring coherent device memory nodes
>>> implementation.
>>>
>>> Isolation seeking coherent device memory node requires allocation isolation
>>> from implicit memory allocations from user space. Towards that effect, the
>>> memory should not be used for generic HugeTLB page pool allocations. This
>>> modifies relevant functions to skip all coherent memory nodes present on
>>> the system during allocation, freeing and auditing for HugeTLB pages.
>>
>> This seems really fragile.  You had to hit, what, 18 call sites?  What
>> are the odds that this is going to stay working?
> 
> 
> I guess a better approach is to introduce new node_states entry such
> that we have one that excludes coherent device memory numa nodes. One
> possibility is to add N_SYSTEM_MEMORY and N_MEMORY.
> 
> Current N_MEMORY becomes N_SYSTEM_MEMORY and N_MEMORY includes
> system and device/any other memory which is coherent.
> 

I thought of this as well, but I would rather see N_COHERENT_MEMORY
as a flag. The idea being that some device memory is a part of
N_MEMORY, but N_COHERENT_MEMORY gives it additional attributes

> All the isolation can then be achieved based on the nodemask_t used for
> allocation. So for allocations we want to avoid from coherent device we
> use N_SYSTEM_MEMORY mask or a derivative of that and where we are ok to
> allocate from CDM with fallbacks we use N_MEMORY.
> 

I suspect its going to be easier to exclude N_COHERENT_MEMORY.

> All nodes zonelist will have zones from the coherent device nodes but we
> will not end up allocating from coherent device node zone due to the
> node mask used.
> 
> 
> This will also make sure we end up allocating from the correct coherent
> device numa node in the presence of multiple of them based on the
> distance of the coherent device node from the current executing numa
> node.
> 

The idea is good overall, but I think its going to be good to document
the exclusions with the flags

Balbir Singh.

[toc] | [prev] | [next] | [standalone]


#1508027 — Re: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths

FromBalbir Singh <bsingharora@gmail.com>
Date2016-10-25 09:30 +0200
SubjectRe: [RFC 3/8] mm: Isolate coherent device memory nodes from HugeTLB allocation paths
Message-ID<sw4lj-2sj-13@gated-at.bofh.it>
In reply to#1508023

On 25/10/16 18:17, Balbir Singh wrote:
> 
> 
> On 25/10/16 15:15, Aneesh Kumar K.V wrote:
>> Dave Hansen <dave.hansen@intel.com> writes:
>>
>>> On 10/23/2016 09:31 PM, Anshuman Khandual wrote:
>>>> This change is part of the isolation requiring coherent device memory nodes
>>>> implementation.
>>>>
>>>> Isolation seeking coherent device memory node requires allocation isolation
>>>> from implicit memory allocations from user space. Towards that effect, the
>>>> memory should not be used for generic HugeTLB page pool allocations. This
>>>> modifies relevant functions to skip all coherent memory nodes present on
>>>> the system during allocation, freeing and auditing for HugeTLB pages.
>>>
>>> This seems really fragile.  You had to hit, what, 18 call sites?  What
>>> are the odds that this is going to stay working?
>>
>>
>> I guess a better approach is to introduce new node_states entry such
>> that we have one that excludes coherent device memory numa nodes. One
>> possibility is to add N_SYSTEM_MEMORY and N_MEMORY.
>>
>> Current N_MEMORY becomes N_SYSTEM_MEMORY and N_MEMORY includes
>> system and device/any other memory which is coherent.
>>
> 
> I thought of this as well, but I would rather see N_COHERENT_MEMORY
> as a flag. The idea being that some device memory is a part of
> N_MEMORY, but N_COHERENT_MEMORY gives it additional attributes
> 
>> All the isolation can then be achieved based on the nodemask_t used for
>> allocation. So for allocations we want to avoid from coherent device we
>> use N_SYSTEM_MEMORY mask or a derivative of that and where we are ok to
>> allocate from CDM with fallbacks we use N_MEMORY.
>>
> 
> I suspect its going to be easier to exclude N_COHERENT_MEMORY.
> 
>> All nodes zonelist will have zones from the coherent device nodes but we
>> will not end up allocating from coherent device node zone due to the
>> node mask used.
>>
>>
>> This will also make sure we end up allocating from the correct coherent
>> device numa node in the presence of multiple of them based on the
>> distance of the coherent device node from the current executing numa
>> node.
>>
> 
> The idea is good overall, but I think its going to be good to document
> the exclusions with the flags
> 

FWIW,, some of this is present in 8/8

Balbir

[toc] | [prev] | [next] | [standalone]


#1506878 — [RFC 7/8] mm: Add a new migration function migrate_virtual_range()

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:40 +0200
Subject[RFC 7/8] mm: Add a new migration function migrate_virtual_range()
Message-ID<svFdf-2Fn-25@gated-at.bofh.it>
In reply to#1506875
This adds a new virtual address range based migration interface which
can migrate all the mapped pages from a virtual range of a process to
a destination node. This also exports this new function symbol.

Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 include/linux/mempolicy.h |  7 ++++
 include/linux/migrate.h   |  3 ++
 mm/mempolicy.c            |  7 ++--
 mm/migrate.c              | 84 +++++++++++++++++++++++++++++++++++++++++++++++
 4 files changed, 96 insertions(+), 5 deletions(-)

diff --git a/include/linux/mempolicy.h b/include/linux/mempolicy.h
index 09d4b70..f18c0ea 100644
--- a/include/linux/mempolicy.h
+++ b/include/linux/mempolicy.h
@@ -152,6 +152,9 @@ extern bool init_nodemask_of_mempolicy(nodemask_t *mask);
 extern bool mempolicy_nodemask_intersects(struct task_struct *tsk,
 				const nodemask_t *mask);
 extern unsigned int mempolicy_slab_node(void);
+extern int queue_pages_range(struct mm_struct *mm, unsigned long start,
+			unsigned long end, nodemask_t *nodes,
+			unsigned long flags, struct list_head *pagelist);
 
 extern enum zone_type policy_zone;
 
@@ -319,4 +322,8 @@ static inline void mpol_put_task_policy(struct task_struct *task)
 {
 }
 #endif /* CONFIG_NUMA */
+
+#define MPOL_MF_DISCONTIG_OK (MPOL_MF_INTERNAL << 0)	/* Skip checks for continuous vmas */
+#define MPOL_MF_INVERT (MPOL_MF_INTERNAL << 1)		/* Invert check for nodemask */
+
 #endif
diff --git a/include/linux/migrate.h b/include/linux/migrate.h
index ae8d475..e2a1af5 100644
--- a/include/linux/migrate.h
+++ b/include/linux/migrate.h
@@ -49,6 +49,9 @@ extern int migrate_page_move_mapping(struct address_space *mapping,
 		struct page *newpage, struct page *page,
 		struct buffer_head *head, enum migrate_mode mode,
 		int extra_count);
+
+extern int migrate_virtual_range(int pid, unsigned long vaddr,
+				unsigned long size, int nid);
 #else
 
 static inline void putback_movable_pages(struct list_head *l) {}
diff --git a/mm/mempolicy.c b/mm/mempolicy.c
index b983cea..aa8479b 100644
--- a/mm/mempolicy.c
+++ b/mm/mempolicy.c
@@ -100,10 +100,6 @@
 
 #include "internal.h"
 
-/* Internal flags */
-#define MPOL_MF_DISCONTIG_OK (MPOL_MF_INTERNAL << 0)	/* Skip checks for continuous vmas */
-#define MPOL_MF_INVERT (MPOL_MF_INTERNAL << 1)		/* Invert check for nodemask */
-
 static struct kmem_cache *policy_cache;
 static struct kmem_cache *sn_cache;
 
@@ -703,7 +699,7 @@ static int queue_pages_test_walk(unsigned long start, unsigned long end,
  * @nodes and @flags,) it's isolated and queued to the pagelist which is
  * passed via @private.)
  */
-static int
+int
 queue_pages_range(struct mm_struct *mm, unsigned long start, unsigned long end,
 		nodemask_t *nodes, unsigned long flags,
 		struct list_head *pagelist)
@@ -724,6 +720,7 @@ queue_pages_range(struct mm_struct *mm, unsigned long start, unsigned long end,
 
 	return walk_page_range(start, end, &queue_pages_walk);
 }
+EXPORT_SYMBOL(queue_pages_range);
 
 /*
  * Apply policy to a single VMA
diff --git a/mm/migrate.c b/mm/migrate.c
index 99250ae..06300bb 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -1367,6 +1367,90 @@ int migrate_pages(struct list_head *from, new_page_t get_new_page,
 	return rc;
 }
 
+static struct page *new_node_page(struct page *page,
+		unsigned long node, int **x)
+{
+	return __alloc_pages_node(node, GFP_HIGHUSER_MOVABLE
+					| __GFP_THISNODE, 0);
+}
+
+#ifdef COHERENT_DEVICE
+static void mark_vma_cdm(struct vm_area_struct *vma)
+{
+	vma->vm_flags |= VM_CDM;
+}
+#else
+static void mark_vma_cdm(struct vm_area_struct *vma) {}
+#endif
+
+/*
+ * migrate_virtual_range - migrate all the pages faulted within a virtual
+ *			address range to a specified node.
+ *
+ * @pid:		PID of the task
+ * @start:		Virtual address range beginning
+ * @end:		Virtual address range end
+ * @nid:		Target migration node
+ *
+ * The function first scans the process VMA list to find out the VMA which
+ * contains the given virtual range. Then validates that the virtual range
+ * is within the given VMA's limits.
+ *
+ * Returns the number of pages that were not migrated or an error code.
+ */
+int migrate_virtual_range(int pid, unsigned long start,
+			unsigned long end, int nid)
+{
+	struct mm_struct *mm;
+	struct vm_area_struct *vma;
+	nodemask_t nmask;
+	int ret = -EINVAL;
+
+	LIST_HEAD(mlist);
+
+	nodes_clear(nmask);
+	nodes_setall(nmask);
+
+	if ((!start) || (!end))
+		return -EINVAL;
+
+	rcu_read_lock();
+	mm = find_task_by_vpid(pid)->mm;
+	rcu_read_unlock();
+
+	start &= PAGE_MASK;
+	end &= PAGE_MASK;
+	down_write(&mm->mmap_sem);
+	for (vma = mm->mmap; vma; vma = vma->vm_next) {
+		if  ((start < vma->vm_start) || (end > vma->vm_end))
+			continue;
+
+		ret = queue_pages_range(mm, start, end, &nmask, MPOL_MF_MOVE_ALL
+						| MPOL_MF_DISCONTIG_OK, &mlist);
+		if (ret) {
+			putback_movable_pages(&mlist);
+			break;
+		}
+
+		if (list_empty(&mlist)) {
+			ret = -ENOMEM;
+			break;
+		}
+
+		ret = migrate_pages(&mlist, new_node_page, NULL, nid,
+					MIGRATE_SYNC, MR_COMPACTION);
+		if (ret) {
+			putback_movable_pages(&mlist);
+		} else {
+			if (isolated_cdm_node(nid))
+				mark_vma_cdm(vma);
+		}
+	}
+	up_write(&mm->mmap_sem);
+	return ret;
+}
+EXPORT_SYMBOL(migrate_virtual_range);
+
 #ifdef CONFIG_NUMA
 /*
  * Move a list of individual pages
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506879 — [DEBUG 06/10] mm: Export definition of 'zone_names' array through mmzone.h

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 06/10] mm: Export definition of 'zone_names' array through mmzone.h
Message-ID<svFmV-2IU-5@gated-at.bofh.it>
In reply to#1506875
zone_names[] is used to identify any zone given it's index which
can be used in many other places. So exporting the definition
through include/linux/mmzone.h header for it's broader access.

Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 include/linux/mmzone.h | 1 +
 mm/page_alloc.c        | 2 +-
 2 files changed, 2 insertions(+), 1 deletion(-)

diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index 821dffb..560bbcd 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -341,6 +341,7 @@ enum zone_type {
 
 };
 
+extern char * const zone_names[];
 #ifndef __GENERATING_BOUNDS_H
 
 struct zone {
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index a2536b4..35c6d2a 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -212,7 +212,7 @@ int sysctl_lowmem_reserve_ratio[MAX_NR_ZONES-1] = {
 
 EXPORT_SYMBOL(totalram_pages);
 
-static char * const zone_names[MAX_NR_ZONES] = {
+char * const zone_names[MAX_NR_ZONES] = {
 #ifdef CONFIG_ZONE_DMA
 	 "DMA",
 #endif
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506881 — [DEBUG 00/10] Test and debug patches for coherent device memory

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 00/10] Test and debug patches for coherent device memory
Message-ID<svFmV-2IU-3@gated-at.bofh.it>
In reply to#1506875
	Coherent device memory support has been experimented around on
POWER platform with simulations and QEMU changes. This series contains
patches which can be classified into three categories.

(1) Memory less node hot plug support
(2) Identifying coherent device nodes during NUMA init
(3) Debug patches to observe zonelists information
(4) Test drivers and scripts

	Patch (2) could have been part of the RFC series but because of the
dependency on patch (1), it goes here. Now lets look at the how all these
components work.

Before Hotplug
==============
NUMACTL Information:
--------------------
available: 5 nodes (0-4)
node 0 cpus: 0 1 2 5 6 20 21 23 27 28 31 32 37 38 39 43 44 48 49 50 51 52
53 54 55 56 57 58 59 60 61 62
node 0 size: 4059 MB
node 0 free: 2956 MB
node 1 cpus: 3 4 7 8 9 10 11 12 13 14 15 16 17 18 19 22 24 25 26 29 30 33
34 35 36 40 41 42 45 46 47 63
node 1 size: 4091 MB
node 1 free: 3920 MB
node 2 cpus:
node 2 size: 0 MB
node 2 free: 0 MB
node 3 cpus:
node 3 size: 0 MB
node 3 free: 0 MB
node 4 cpus:
node 4 size: 0 MB
node 4 free: 0 MB
node distances:
node   0   1   2   3   4 
  0:  10  40  40  40  40 
  1:  40  10  40  40  40 
  2:  40  40  10  40  40 
  3:  40  40  40  10  40 
  4:  40  40  40  40  10 

ZONELIST Information
--------------------
[NODE (0)]
        ZONELIST_FALLBACK (0xc00000000140da00)
                (0) (node 0) (DMA     0xc00000000140c000)
                (1) (node 1) (DMA     0xc000000100000000)
        ZONELIST_NOFALLBACK (0xc000000001411a10)
                (0) (node 0) (DMA     0xc00000000140c000)
[NODE (1)]
        ZONELIST_FALLBACK (0xc000000100001a00)
                (0) (node 1) (DMA     0xc000000100000000)
                (1) (node 0) (DMA     0xc00000000140c000)
        ZONELIST_NOFALLBACK (0xc000000100005a10)
                (0) (node 1) (DMA     0xc000000100000000)
[NODE (2)]
        ZONELIST_FALLBACK (0xc000000001427700)
                (0) (node 0) (DMA     0xc00000000140c000)
                (1) (node 1) (DMA     0xc000000100000000)
        ZONELIST_NOFALLBACK (0xc00000000142b710)
[NODE (3)]
        ZONELIST_FALLBACK (0xc000000001431400)
                (0) (node 0) (DMA     0xc00000000140c000)
                (1) (node 1) (DMA     0xc000000100000000)
        ZONELIST_NOFALLBACK (0xc000000001435410)
[NODE (4)]
        ZONELIST_FALLBACK (0xc00000000143b100)
                (0) (node 0) (DMA     0xc00000000140c000)
                (1) (node 1) (DMA     0xc000000100000000)
        ZONELIST_NOFALLBACK (0xc00000000143f110)

After Hotplug
=============
NUMACTL Information:
--------------------
available: 5 nodes (0-4)
node 0 cpus: 0 1 2 5 6 20 21 23 27 28 31 32 37 38 39 43 44 48 49 50 51 52
53 54 55 56 57 58 59 60 61 62
node 0 size: 4059 MB
node 0 free: 2804 MB
node 1 cpus: 3 4 7 8 9 10 11 12 13 14 15 16 17 18 19 22 24 25 26 29 30 33
34 35 36 40 41 42 45 46 47 63
node 1 size: 4091 MB
node 1 free: 3860 MB
node 2 cpus:
node 2 size: 4096 MB
node 2 free: 4095 MB
node 3 cpus:
node 3 size: 4096 MB
node 3 free: 4095 MB
node 4 cpus:
node 4 size: 4096 MB
node 4 free: 4095 MB
node distances:
node   0   1   2   3   4 
  0:  10  40  40  40  40 
  1:  40  10  40  40  40 
  2:  40  40  10  40  40 
  3:  40  40  40  10  40 
  4:  40  40  40  40  10 

ZONELIST Information:
---------------------
[NODE (0)]
        ZONELIST_FALLBACK (0xc00000000140da00)
                (0) (node 0) (DMA     0xc00000000140c000)
                (1) (node 1) (DMA     0xc000000100000000)
        ZONELIST_NOFALLBACK (0xc000000001411a10)
                (0) (node 0) (DMA     0xc00000000140c000)
[NODE (1)]
        ZONELIST_FALLBACK (0xc000000100001a00)
                (0) (node 1) (DMA     0xc000000100000000)
                (1) (node 0) (DMA     0xc00000000140c000)
        ZONELIST_NOFALLBACK (0xc000000100005a10)
                (0) (node 1) (DMA     0xc000000100000000)
[NODE (2)]
        ZONELIST_FALLBACK (0xc000000001427700)
                (0) (node 2) (Movable 0xc000000001427080)
                (1) (node 0) (DMA     0xc00000000140c000)
                (2) (node 1) (DMA     0xc000000100000000)
        ZONELIST_NOFALLBACK (0xc00000000142b710)
                (0) (node 2) (Movable 0xc000000001427080)
[NODE (3)]
        ZONELIST_FALLBACK (0xc000000001431400)
                (0) (node 3) (Movable 0xc000000001430d80)
                (1) (node 0) (DMA     0xc00000000140c000)
                (2) (node 1) (DMA     0xc000000100000000)
        ZONELIST_NOFALLBACK (0xc000000001435410)
                (0) (node 3) (Movable 0xc000000001430d80)
[NODE (4)]
        ZONELIST_FALLBACK (0xc00000000143b100)
                (0) (node 4) (Movable 0xc00000000143aa80)
                (1) (node 0) (DMA     0xc00000000140c000)
                (2) (node 1) (DMA     0xc000000100000000)
        ZONELIST_NOFALLBACK (0xc00000000143f110)
                (0) (node 4) (Movable 0xc00000000143aa80)

	After the coherent device memory nodes have been hot plugged into
the kernel, did some simple VMA migration tests to verify it's stability.
cdm_migration.sh does the actual test of moving VMAs of ebizzy workload
which results in the following stats and traces.

Results:
-------
passed 13
failed 0
queuef 0
empty 3
missing 0

Traces:
-------
migrate_virtual_range: 55094 10000000 10010000 0: migration_passed
migrate_virtual_range: 55094 10010000 10020000 0: migration_passed
migrate_virtual_range: 55094 10020000 10030000 3: migration_passed
migrate_virtual_range: 55094 3fff3b6a0000 3fff8b3c0000 0: list_empty
migrate_virtual_range: 55094 3fff8b3c0000 3fff8b580000 1: migration_passed
migrate_virtual_range: 55094 3fff8b580000 3fff8b590000 2: migration_passed
migrate_virtual_range: 55094 3fff8b590000 3fff8b5a0000 0: migration_passed
migrate_virtual_range: 55094 3fff8b5a0000 3fff8b5c0000 2: migration_passed
migrate_virtual_range: 55094 3fff8b5c0000 3fff8b5d0000 2: migration_passed
migrate_virtual_range: 55094 3fff8b5d0000 3fff8b5e0000 0: migration_passed
migrate_virtual_range: 55094 3fff8b5e0000 3fff8b5f0000 0: list_empty
migrate_virtual_range: 55094 3fff8b5f0000 3fff8b610000 3: list_empty
migrate_virtual_range: 55094 3fff8b610000 3fff8b640000 3: migration_passed
migrate_virtual_range: 55094 3fff8b640000 3fff8b650000 2: migration_passed
migrate_virtual_range: 55094 3fff8b650000 3fff8b660000 1: migration_passed
migrate_virtual_range: 55094 3ffff25e0000 3ffff2610000 1: migration_passed

Anshuman Khandual (6):
  powerpc/mm: Identify isolation seeking coherent memory nodes during boot
  mm: Export definition of 'zone_names' array through mmzone.h
  mm: Add debugfs interface to dump each node's zonelist information
  powerpc: Enable CONFIG_MOVABLE_NODE for PPC64 platform
  drivers: Add two drivers for coherent device memory tests
  test: Add a script to perform random VMA migrations across nodes

Reza Arbab (4):
  dt-bindings: Add doc for ibm,hotplug-aperture
  powerpc/mm: Create numa nodes for hotplug memory
  powerpc/mm: Allow memory hotplug into a memory less node
  mm: Enable CONFIG_MOVABLE_NODE on powerpc

 .../bindings/powerpc/opal/hotplug-aperture.txt     |  26 ++
 Documentation/kernel-parameters.txt                |   2 +-
 arch/powerpc/Kconfig                               |   4 +
 arch/powerpc/mm/numa.c                             |  43 ++-
 drivers/char/Kconfig                               |  23 ++
 drivers/char/Makefile                              |   2 +
 drivers/char/coherent_hotplug_demo.c               | 133 ++++++++
 drivers/char/coherent_memory_demo.c                | 337 +++++++++++++++++++++
 drivers/char/memory_online_sysfs.h                 | 148 +++++++++
 include/linux/mmzone.h                             |   1 +
 mm/Kconfig                                         |   2 +-
 mm/memory.c                                        |  63 ++++
 mm/migrate.c                                       |  10 +
 mm/page_alloc.c                                    |   2 +-
 tools/testing/selftests/vm/cdm_migration.sh        |  76 +++++
 15 files changed, 855 insertions(+), 17 deletions(-)
 create mode 100644 Documentation/devicetree/bindings/powerpc/opal/hotplug-aperture.txt
 create mode 100644 drivers/char/coherent_hotplug_demo.c
 create mode 100644 drivers/char/coherent_memory_demo.c
 create mode 100644 drivers/char/memory_online_sysfs.h
 create mode 100755 tools/testing/selftests/vm/cdm_migration.sh

-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506882 — [DEBUG 02/10] powerpc/mm: Create numa nodes for hotplug memory

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 02/10] powerpc/mm: Create numa nodes for hotplug memory
Message-ID<svFmW-2IU-13@gated-at.bofh.it>
In reply to#1506881
From: Reza Arbab <arbab@linux.vnet.ibm.com>

When scanning the device tree to initialize the system NUMA topology,
process dt elements with compatible id "ibm,hotplug-aperture" to create
memoryless numa nodes.

These nodes will be filled when hotplug occurs within the associated
address range.

Signed-off-by: Reza Arbab <arbab@linux.vnet.ibm.com>
Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 arch/powerpc/mm/numa.c | 10 ++++++++--
 1 file changed, 8 insertions(+), 2 deletions(-)

diff --git a/arch/powerpc/mm/numa.c b/arch/powerpc/mm/numa.c
index a51c188..42fcc8e 100644
--- a/arch/powerpc/mm/numa.c
+++ b/arch/powerpc/mm/numa.c
@@ -708,6 +708,12 @@ static void __init parse_drconf_memory(struct device_node *memory)
 	}
 }
 
+static const struct of_device_id memory_match[] = {
+	{ .type = "memory" },
+	{ .compatible = "ibm,hotplug-aperture" },
+	{ /* sentinel */ }
+};
+
 static int __init parse_numa_properties(void)
 {
 	struct device_node *memory;
@@ -752,7 +758,7 @@ static int __init parse_numa_properties(void)
 
 	get_n_mem_cells(&n_mem_addr_cells, &n_mem_size_cells);
 
-	for_each_node_by_type(memory, "memory") {
+	for_each_matching_node(memory, memory_match) {
 		unsigned long start;
 		unsigned long size;
 		int nid;
@@ -1044,7 +1050,7 @@ static int hot_add_node_scn_to_nid(unsigned long scn_addr)
 	struct device_node *memory;
 	int nid = -1;
 
-	for_each_node_by_type(memory, "memory") {
+	for_each_matching_node(memory, memory_match) {
 		unsigned long start, size;
 		int ranges;
 		const __be32 *memcell_buf;
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506883 — [DEBUG 03/10] powerpc/mm: Allow memory hotplug into a memory less node

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 03/10] powerpc/mm: Allow memory hotplug into a memory less node
Message-ID<svFmW-2IU-7@gated-at.bofh.it>
In reply to#1506881
From: Reza Arbab <arbab@linux.vnet.ibm.com>

Remove the check which prevents us from hotplugging into an empty node.

Signed-off-by: Reza Arbab <arbab@linux.vnet.ibm.com>
Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 arch/powerpc/mm/numa.c | 13 +------------
 1 file changed, 1 insertion(+), 12 deletions(-)

diff --git a/arch/powerpc/mm/numa.c b/arch/powerpc/mm/numa.c
index 42fcc8e..5010181 100644
--- a/arch/powerpc/mm/numa.c
+++ b/arch/powerpc/mm/numa.c
@@ -1091,7 +1091,7 @@ static int hot_add_node_scn_to_nid(unsigned long scn_addr)
 int hot_add_scn_to_nid(unsigned long scn_addr)
 {
 	struct device_node *memory = NULL;
-	int nid, found = 0;
+	int nid;
 
 	if (!numa_enabled || (min_common_depth < 0))
 		return first_online_node;
@@ -1107,17 +1107,6 @@ int hot_add_scn_to_nid(unsigned long scn_addr)
 	if (nid < 0 || !node_online(nid))
 		nid = first_online_node;
 
-	if (NODE_DATA(nid)->node_spanned_pages)
-		return nid;
-
-	for_each_online_node(nid) {
-		if (NODE_DATA(nid)->node_spanned_pages) {
-			found = 1;
-			break;
-		}
-	}
-
-	BUG_ON(!found);
 	return nid;
 }
 
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506884 — [DEBUG 08/10] powerpc: Enable CONFIG_MOVABLE_NODE for PPC64 platform

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 08/10] powerpc: Enable CONFIG_MOVABLE_NODE for PPC64 platform
Message-ID<svFmW-2IU-9@gated-at.bofh.it>
In reply to#1506881
Just enable MOVABLE_NODE config option for PPC64 platform by default.
This prevents accidentally building the kernel without the required
config option.

Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 arch/powerpc/Kconfig | 4 ++++
 1 file changed, 4 insertions(+)

diff --git a/arch/powerpc/Kconfig b/arch/powerpc/Kconfig
index 65fba4c..3989d89 100644
--- a/arch/powerpc/Kconfig
+++ b/arch/powerpc/Kconfig
@@ -310,6 +310,10 @@ config PGTABLE_LEVELS
 	default 3 if PPC_64K_PAGES && !PPC_BOOK3S_64
 	default 4
 
+config MOVABLE_NODE
+	bool
+	default y if PPC64
+
 source "init/Kconfig"
 
 source "kernel/Kconfig.freezer"
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506885 — [DEBUG 01/10] dt-bindings: Add doc for ibm,hotplug-aperture

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 01/10] dt-bindings: Add doc for ibm,hotplug-aperture
Message-ID<svFmW-2IU-31@gated-at.bofh.it>
In reply to#1506881
From: Reza Arbab <arbab@linux.vnet.ibm.com>

Signed-off-by: Reza Arbab <arbab@linux.vnet.ibm.com>
Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 .../bindings/powerpc/opal/hotplug-aperture.txt     | 26 ++++++++++++++++++++++
 1 file changed, 26 insertions(+)
 create mode 100644 Documentation/devicetree/bindings/powerpc/opal/hotplug-aperture.txt

diff --git a/Documentation/devicetree/bindings/powerpc/opal/hotplug-aperture.txt b/Documentation/devicetree/bindings/powerpc/opal/hotplug-aperture.txt
new file mode 100644
index 0000000..04dde03
--- /dev/null
+++ b/Documentation/devicetree/bindings/powerpc/opal/hotplug-aperture.txt
@@ -0,0 +1,26 @@
+Designated hotplug memory
+-------------------------
+
+This binding describes a region of hotplug memory which is not present
+at boot, allowing its eventual NUMA associativity to be prespecified.
+
+Required properties:
+
+- compatible
+	"ibm,hotplug-aperture"
+
+- reg
+	base address and size of the region (standard definition)
+
+- ibm,associativity
+	NUMA associativity (standard definition)
+
+Example:
+
+A 2 GiB aperture at 0x100000000, to be part of nid 3 when hotplugged:
+
+	hotplug-memory@100000000 {
+		compatible = "ibm,hotplug-aperture";
+		reg = <0x0 0x100000000 0x0 0x80000000>;
+		ibm,associativity = <0x4 0x0 0x0 0x0 0x3>;
+	};
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506886 — [DEBUG 05/10] powerpc/mm: Identify isolation seeking coherent memory nodes during boot

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 05/10] powerpc/mm: Identify isolation seeking coherent memory nodes during boot
Message-ID<svFmW-2IU-11@gated-at.bofh.it>
In reply to#1506881
Isolation seeking coherent memory nodes which wish to be MNODE_ISOLATION
in core VM will have "ibm,hotplug-aperture" as one of the compatible
properties in their respective device nodes in device tree. Detect them
during platform NUMA initialization and mark their respective coherent
mask in pglist_data structure as MNODE_ISOLATION.

Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 arch/powerpc/mm/numa.c | 20 ++++++++++++++++++++
 1 file changed, 20 insertions(+)

diff --git a/arch/powerpc/mm/numa.c b/arch/powerpc/mm/numa.c
index 5010181..89ae64c 100644
--- a/arch/powerpc/mm/numa.c
+++ b/arch/powerpc/mm/numa.c
@@ -64,6 +64,7 @@ static int form1_affinity;
 static int distance_ref_points_depth;
 static const __be32 *distance_ref_points;
 static int distance_lookup_table[MAX_NUMNODES][MAX_DISTANCE_REF_POINTS];
+static int node_to_phys_device_map[MAX_NUMNODES];
 
 /*
  * Allocate node_to_cpumask_map based on number of available nodes
@@ -714,6 +715,17 @@ static const struct of_device_id memory_match[] = {
 	{ /* sentinel */ }
 };
 
+int arch_get_memory_phys_device(unsigned long start_pfn)
+{
+	return node_to_phys_device_map[pfn_to_nid(start_pfn)];
+}
+
+int special_mem_node(int nid)
+{
+	return node_to_phys_device_map[nid];
+}
+EXPORT_SYMBOL(special_mem_node);
+
 static int __init parse_numa_properties(void)
 {
 	struct device_node *memory;
@@ -789,6 +801,9 @@ static int __init parse_numa_properties(void)
 		if (nid < 0)
 			nid = default_nid;
 
+		if (of_device_is_compatible(memory, "ibm,hotplug-aperture"))
+			node_to_phys_device_map[nid] = 1;
+
 		fake_numa_create_new_node(((start + size) >> PAGE_SHIFT), &nid);
 		node_set_online(nid);
 
@@ -908,6 +923,11 @@ static void __init setup_node_data(int nid, u64 start_pfn, u64 end_pfn)
 	NODE_DATA(nid)->node_id = nid;
 	NODE_DATA(nid)->node_start_pfn = start_pfn;
 	NODE_DATA(nid)->node_spanned_pages = spanned_pages;
+
+#ifdef CONFIG_COHERENT_DEVICE
+	if (special_mem_node(nid))
+		set_cdm_isolation(nid);
+#endif
 }
 
 void __init initmem_init(void)
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506888 — [DEBUG 10/10] test: Add a script to perform random VMA migrations across nodes

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 10/10] test: Add a script to perform random VMA migrations across nodes
Message-ID<svFmW-2IU-23@gated-at.bofh.it>
In reply to#1506881
This is a test script which creates a workload (e.g ebizzy) and go through
it's VMAs (/proc/pid/maps) and initiate migration to random nodes which can
be either system memory node or coherent memory node.

Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 tools/testing/selftests/vm/cdm_migration.sh | 76 +++++++++++++++++++++++++++++
 1 file changed, 76 insertions(+)
 create mode 100755 tools/testing/selftests/vm/cdm_migration.sh

diff --git a/tools/testing/selftests/vm/cdm_migration.sh b/tools/testing/selftests/vm/cdm_migration.sh
new file mode 100755
index 0000000..3ab7230
--- /dev/null
+++ b/tools/testing/selftests/vm/cdm_migration.sh
@@ -0,0 +1,76 @@
+#!/usr/bin/bash
+#
+# Should work with any workoad and workload commandline.
+# But for now ebizzy should be installed. Please run it
+# as root.
+#
+# Copyright (C) Anshuman Khandual 2016, IBM Corporation
+#
+# Licensed under GPL V2
+
+# Unload, build and reload modules
+if [ "$1" = "reload" ]
+then
+	rmmod coherent_memory_demo
+	rmmod coherent_hotplug_demo
+	cd ../../../../
+	make -s -j 64 modules
+	insmod drivers/char/coherent_hotplug_demo.ko
+	insmod drivers/char/coherent_memory_demo.ko
+	cd -
+fi
+
+# Workload
+workload=ebizzy
+work_cmd="ebizzy -T -z -m -t 128 -n 100000 -s 32768 -S 10000"
+
+pkill $workload
+$work_cmd &
+
+# File
+if [ -e input_file.txt ]
+then
+	rm input_file.txt
+fi
+
+# Inputs
+pid=`pidof ebizzy`
+cp /proc/$pid/maps input_file.txt
+if [ ! -e input_file.txt ]
+then
+	echo "Input file was not created"
+	exit
+fi
+input=input_file.txt
+
+# Migrations
+dmesg -C
+while read line
+do
+	addr_start=$(echo $line | cut -d '-' -f1)
+	addr_end=$(echo $line | cut -d '-' -f2 | cut -d ' ' -f1)
+	node=`expr $RANDOM % 4`
+
+	echo $pid,0x$addr_start,0x$addr_end,$node > /sys/kernel/debug/coherent_debug
+done < "$input"
+
+# Analyze dmesg output
+passed=`dmesg | grep "migration_passed" | wc -l`
+failed=`dmesg | grep "migration_failed" | wc -l`
+queuef=`dmesg | grep "queue_pages_range_failed" | wc -l`
+empty=`dmesg | grep "list_empty" | wc -l`
+missing=`dmesg | grep "vma_missing" | wc -l`
+
+# Stats
+echo passed	$passed
+echo failed	$failed
+echo queuef	$queuef
+echo empty	$empty
+echo missing	$missing
+
+# Cleanup
+rm input_file.txt
+if pgrep -x $workload > /dev/null
+then
+	pkill $workload
+fi
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506891 — [DEBUG 07/10] mm: Add debugfs interface to dump each node's zonelist information

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 07/10] mm: Add debugfs interface to dump each node's zonelist information
Message-ID<svFmW-2IU-21@gated-at.bofh.it>
In reply to#1506881
Each individual node in the system has a ZONELIST_FALLBACK zonelist
and a ZONELIST_NOFALLBACK zonelist. These zonelists decide fallback
order of zones during memory allocations. Sometimes it helps to dump
these zonelists to see the priority order of various zones in them.

Particularly platforms which support memory hotplug into previously
non existing zones (at boot), this interface helps in visualizing
which all zonelists of the system at what priority level, the new
hot added memory ends up in. POWER is such a platform where all the
memory detected during boot time remains with ZONE_DMA for good but
then hot plug process can actually get new memory into ZONE_MOVABLE.
So having a way to get the snapshot of the zonelists on the system
after memory or node hot[un]plug is desirable. This change adds one
new debugfs interface (/sys/kernel/debug/zonelists) which will fetch
and dump this information.

Example zonelist information from a KVM guest with four NUMA nodes
on a POWER8 platform.

[NODE (0)]
	ZONELIST_FALLBACK
		(0) (Node 0) (DMA)
		(1) (Node 1) (DMA)
		(2) (Node 2) (DMA)
		(3) (Node 3) (DMA)
	ZONELIST_NOFALLBACK
		(0) (Node 0) (DMA)
[NODE (1)]
	ZONELIST_FALLBACK
		(0) (Node 1) (DMA)
		(1) (Node 2) (DMA)
		(2) (Node 3) (DMA)
		(3) (Node 0) (DMA)
	ZONELIST_NOFALLBACK
		(0) (Node 1) (DMA)
[NODE (2)]
	ZONELIST_FALLBACK
		(0) (Node 2) (DMA)
		(1) (Node 3) (DMA)
		(2) (Node 0) (DMA)
		(3) (Node 1) (DMA)
	ZONELIST_NOFALLBACK
		(0) (Node 2) (DMA)
[NODE (3)]
	ZONELIST_FALLBACK
		(0) (Node 3) (DMA)
		(1) (Node 0) (DMA)
		(2) (Node 1) (DMA)
		(3) (Node 2) (DMA)
	ZONELIST_NOFALLBACK
		(0) (Node 3) (DMA)

Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 mm/memory.c | 63 +++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 63 insertions(+)

diff --git a/mm/memory.c b/mm/memory.c
index e18c57b..3be1753 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -64,6 +64,7 @@
 #include <linux/debugfs.h>
 #include <linux/userfaultfd_k.h>
 #include <linux/dax.h>
+#include <linux/mmzone.h>
 
 #include <asm/io.h>
 #include <asm/mmu_context.h>
@@ -3087,6 +3088,68 @@ static int __init fault_around_debugfs(void)
 		pr_warn("Failed to create fault_around_bytes in debugfs");
 	return 0;
 }
+
+#ifdef CONFIG_NUMA
+static void show_zonelist(struct seq_file *m, struct zonelist *zonelist)
+{
+	unsigned int i;
+
+	for (i = 0; zonelist->_zonerefs[i].zone; i++) {
+		seq_printf(m, "\t\t(%d) (Node %d) (%-7s 0x%pK)\n", i,
+			zonelist->_zonerefs[i].zone->zone_pgdat->node_id,
+			zone_names[zonelist->_zonerefs[i].zone_idx],
+			(void *) zonelist->_zonerefs[i].zone);
+	}
+}
+
+static int zonelists_show(struct seq_file *m, void *v)
+{
+	struct zonelist *zonelist;
+	unsigned int node;
+
+	for_each_online_node(node) {
+		zonelist = &(NODE_DATA(node)->
+				node_zonelists[ZONELIST_FALLBACK]);
+		seq_printf(m, "[NODE (%d)]\n", node);
+		seq_puts(m, "\tZONELIST_FALLBACK ");
+		seq_printf(m, "(0x%pK)\n", zonelist);
+		show_zonelist(m, zonelist);
+
+		zonelist = &(NODE_DATA(node)->
+				node_zonelists[ZONELIST_NOFALLBACK]);
+		seq_puts(m, "\tZONELIST_NOFALLBACK ");
+		seq_printf(m, "(0x%pK)\n", zonelist);
+		show_zonelist(m, zonelist);
+	}
+	return 0;
+}
+
+static int zonelists_open(struct inode *inode, struct file *filp)
+{
+	return single_open(filp, zonelists_show, NULL);
+}
+
+static const struct file_operations zonelists_fops = {
+	.open		= zonelists_open,
+	.read		= seq_read,
+	.llseek		= seq_lseek,
+	.release	= single_release,
+};
+
+static int __init zonelists_debugfs(void)
+{
+	void *ret;
+
+	ret = debugfs_create_file("zonelists", 0444, NULL, NULL,
+			&zonelists_fops);
+	if (!ret)
+		pr_warn("Failed to create zonelists in debugfs");
+	return 0;
+}
+
+late_initcall(zonelists_debugfs);
+#endif /* CONFIG_NUMA */
+
 late_initcall(fault_around_debugfs);
 #endif
 
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506892 — [DEBUG 04/10] mm: Enable CONFIG_MOVABLE_NODE on powerpc

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 04/10] mm: Enable CONFIG_MOVABLE_NODE on powerpc
Message-ID<svFmW-2IU-33@gated-at.bofh.it>
In reply to#1506881
From: Reza Arbab <arbab@linux.vnet.ibm.com>

Onlining memory into ZONE_MOVABLE requires CONFIG_MOVABLE_NODE.

Enable the use of this config option on PPC64 platforms.

Signed-off-by: Reza Arbab <arbab@linux.vnet.ibm.com>
Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 Documentation/kernel-parameters.txt | 2 +-
 mm/Kconfig                          | 2 +-
 2 files changed, 2 insertions(+), 2 deletions(-)

diff --git a/Documentation/kernel-parameters.txt b/Documentation/kernel-parameters.txt
index 37babf9..61cfa0b 100644
--- a/Documentation/kernel-parameters.txt
+++ b/Documentation/kernel-parameters.txt
@@ -2401,7 +2401,7 @@ bytes respectively. Such letter suffixes can also be entirely omitted.
 			that the amount of memory usable for all allocations
 			is not too small.
 
-	movable_node	[KNL,X86] Boot-time switch to enable the effects
+	movable_node	[KNL,X86,PPC] Boot-time switch to enable the effects
 			of CONFIG_MOVABLE_NODE=y. See mm/Kconfig for details.
 
 	MTD_Partition=	[MTD]
diff --git a/mm/Kconfig b/mm/Kconfig
index cb50468..a4727fa 100644
--- a/mm/Kconfig
+++ b/mm/Kconfig
@@ -153,7 +153,7 @@ config MOVABLE_NODE
 	bool "Enable to assign a node which has only movable memory"
 	depends on HAVE_MEMBLOCK
 	depends on NO_BOOTMEM
-	depends on X86_64
+	depends on X86_64 || PPC64
 	depends on NUMA
 	default n
 	help
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1506895 — [DEBUG 09/10] drivers: Add two drivers for coherent device memory tests

FromAnshuman Khandual <khandual@linux.vnet.ibm.com>
Date2016-10-24 06:50 +0200
Subject[DEBUG 09/10] drivers: Add two drivers for coherent device memory tests
Message-ID<svFmW-2IU-29@gated-at.bofh.it>
In reply to#1506881
This adds two different drivers inside drivers/char/ directory under two
new kernel config options COHERENT_HOTPLUG_DEMO and COHERENT_MEMORY_DEMO.

1) coherent_hotplug_demo: Detects, hoptlugs the coherent device memory
2) coherent_memory_demo:  Exports debugfs interface for VMA migrations

Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
 drivers/char/Kconfig                 |  23 +++
 drivers/char/Makefile                |   2 +
 drivers/char/coherent_hotplug_demo.c | 133 ++++++++++++++
 drivers/char/coherent_memory_demo.c  | 337 +++++++++++++++++++++++++++++++++++
 drivers/char/memory_online_sysfs.h   | 148 +++++++++++++++
 mm/migrate.c                         |  10 ++
 6 files changed, 653 insertions(+)
 create mode 100644 drivers/char/coherent_hotplug_demo.c
 create mode 100644 drivers/char/coherent_memory_demo.c
 create mode 100644 drivers/char/memory_online_sysfs.h

diff --git a/drivers/char/Kconfig b/drivers/char/Kconfig
index dcc0973..22c538d 100644
--- a/drivers/char/Kconfig
+++ b/drivers/char/Kconfig
@@ -588,6 +588,29 @@ config TILE_SROM
 	  device appear much like a simple EEPROM, and knows
 	  how to partition a single ROM for multiple purposes.
 
+config COHERENT_HOTPLUG_DEMO
+	tristate "Demo driver to test coherent memory node hotplug"
+	depends on PPC64 || COHERENT_DEVICE
+	default n
+	help
+	  Say yes when you want to build a test driver to hotplug all
+	  the coherent memory nodes present on the system. This driver
+	  scans through the device tree, checks on "ibm,memory-device"
+	  property device nodes and onlines its memory. When unloaded,
+	  it goes through the list of memory ranges it onlined before
+	  and oflines them one by one. If not sure, select N.
+
+config COHERENT_MEMORY_DEMO
+	tristate "Demo driver to test coherent memory node functionality"
+	depends on PPC64 || COHERENT_DEVICE
+	default n
+	help
+	  Say yes when you want to build a test driver to demonstrate
+	  the coherent memory functionalities, capabilities and probable
+	  utilizaton. It also exports a debugfs file to accept inputs for
+	  virtual address range migration for any process. If not sure,
+	  select N.
+
 source "drivers/char/xillybus/Kconfig"
 
 endmenu
diff --git a/drivers/char/Makefile b/drivers/char/Makefile
index 6e6c244..92fa338 100644
--- a/drivers/char/Makefile
+++ b/drivers/char/Makefile
@@ -60,3 +60,5 @@ js-rtc-y = rtc.o
 obj-$(CONFIG_TILE_SROM)		+= tile-srom.o
 obj-$(CONFIG_XILLYBUS)		+= xillybus/
 obj-$(CONFIG_POWERNV_OP_PANEL)	+= powernv-op-panel.o
+obj-$(CONFIG_COHERENT_HOTPLUG_DEMO)	+= coherent_hotplug_demo.o
+obj-$(CONFIG_COHERENT_MEMORY_DEMO)	+= coherent_memory_demo.o
diff --git a/drivers/char/coherent_hotplug_demo.c b/drivers/char/coherent_hotplug_demo.c
new file mode 100644
index 0000000..3670081
--- /dev/null
+++ b/drivers/char/coherent_hotplug_demo.c
@@ -0,0 +1,133 @@
+/*
+ * Memory hotplug support for coherent memory nodes in runtime.
+ *
+ * Copyright (C) 2016, Reza Arbab, IBM Corporation.
+ * Copyright (C) 2016, Anshuman Khandual, IBM Corporation.
+ *
+ * This program is free software; you can redistribute it and/or modify
+ * it under the terms of the GNU General Public License as published by
+ * the Free Software Foundation; either version 2 of the License, or
+ * (at your option) any later version.
+ *
+ * This program is distributed in the hope that it will be useful,
+ * but WITHOUT ANY WARRANTY; without even the implied warranty of
+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the
+ * GNU General Public License for more details.
+ */
+#include <linux/of.h>
+#include <linux/export.h>
+#include <linux/spinlock.h>
+#include <linux/init.h>
+#include <linux/memblock.h>
+#include <linux/module.h>
+#include <linux/memory.h>
+#include <linux/sizes.h>
+#include <linux/bitops.h>
+#include <linux/device.h>
+#include <linux/fs.h>
+#include <linux/slab.h>
+#include <linux/mm.h>
+#include <linux/pagemap.h>
+#include <linux/migrate.h>
+#include <linux/memblock.h>
+#include <linux/uaccess.h>
+
+#include <asm/mmu.h>
+#include <asm/pgalloc.h>
+#include "memory_online_sysfs.h"
+
+#define MAX_HOTADD_NODES 100
+phys_addr_t addr[MAX_HOTADD_NODES][2];
+int nr_addr;
+
+/*
+ * extern int memory_failure(unsigned long pfn, int trapno, int flags);
+ * extern int min_free_kbytes;
+ * extern int user_min_free_kbytes;
+ *
+ * extern unsigned long nr_kernel_pages;
+ * extern unsigned long nr_all_pages;
+ * extern unsigned long dma_reserve;
+ */
+
+static void dump_core_vm_tunables(void)
+{
+/*
+ *	printk(":::::::: VM TUNABLES :::::::\n");
+ *	printk("[min_free_kbytes]	%d\n", min_free_kbytes);
+ *	printk("[user_min_free_kbytes]	%d\n", user_min_free_kbytes);
+ *	printk("[nr_kernel_pages]	%ld\n", nr_kernel_pages);
+ *	printk("[nr_all_pages]		%ld\n", nr_all_pages);
+ *	printk("[dma_reserve]		%ld\n", dma_reserve);
+ */
+}
+
+
+
+static int online_coherent_memory(void)
+{
+	struct device_node *memory;
+
+	nr_addr = 0;
+	disable_auto_online();
+	dump_core_vm_tunables();
+	for_each_compatible_node(memory, NULL, "ibm,memory-device") {
+		struct device_node *mem;
+		const __be64 *reg;
+		unsigned int len, ret;
+		phys_addr_t start, size;
+
+		mem = of_parse_phandle(memory, "memory-region", 0);
+		if (!mem) {
+			pr_info("memory-region property not found\n");
+			return -1;
+		}
+
+		reg = of_get_property(mem, "reg", &len);
+		if (!reg || len <= 0) {
+			pr_info("memory-region property not found\n");
+			return -1;
+		}
+		start = be64_to_cpu(*reg);
+		size = be64_to_cpu(*(reg + 1));
+		pr_info("Coherent memory start %llx size %llx\n", start, size);
+		ret = memory_probe_store(start, size);
+		if (ret)
+			pr_info("probe faile\n");
+
+		ret = store_mem_state(start, size, "online_movable");
+		if (ret)
+			pr_info("online_movable failed\n");
+
+		addr[nr_addr][0] = start;
+		addr[nr_addr][1] = size;
+		nr_addr++;
+	}
+	dump_core_vm_tunables();
+	enable_auto_online();
+	return 0;
+}
+
+static int offline_coherent_memory(void)
+{
+	int i;
+
+	for (i = 0; i < nr_addr; i++)
+		store_mem_state(addr[i][0], addr[i][1], "offline");
+	return 0;
+}
+
+static void __exit coherent_hotplug_exit(void)
+{
+	pr_info("%s\n", __func__);
+	offline_coherent_memory();
+}
+
+static int __init coherent_hotplug_init(void)
+{
+	pr_info("%s\n", __func__);
+	return online_coherent_memory();
+}
+module_init(coherent_hotplug_init);
+module_exit(coherent_hotplug_exit);
+MODULE_LICENSE("GPL");
diff --git a/drivers/char/coherent_memory_demo.c b/drivers/char/coherent_memory_demo.c
new file mode 100644
index 0000000..1dcd9f7
--- /dev/null
+++ b/drivers/char/coherent_memory_demo.c
@@ -0,0 +1,337 @@
+/*
+ * Demonstrating various aspects of the coherent memory.
+ *
+ * Copyright (C) 2016, Anshuman Khandual, IBM Corporation.
+ *
+ * This program is free software; you can redistribute it and/or modify
+ * it under the terms of the GNU General Public License as published by
+ * the Free Software Foundation; either version 2 of the License, or
+ * (at your option) any later version.
+ *
+ * This program is distributed in the hope that it will be useful,
+ * but WITHOUT ANY WARRANTY; without even the implied warranty of
+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the
+ * GNU General Public License for more details.
+ */
+#include <linux/of.h>
+#include <linux/export.h>
+#include <linux/spinlock.h>
+#include <linux/init.h>
+#include <linux/memblock.h>
+#include <linux/module.h>
+#include <linux/memory.h>
+#include <linux/sizes.h>
+#include <linux/bitops.h>
+#include <linux/device.h>
+#include <linux/fs.h>
+#include <linux/slab.h>
+#include <linux/mm.h>
+#include <linux/pagemap.h>
+#include <linux/migrate.h>
+#include <linux/memblock.h>
+#include <linux/debugfs.h>
+#include <linux/uaccess.h>
+
+#include <asm/mmu.h>
+#include <asm/pgalloc.h>
+
+#define COHERENT_DEV_MAJOR 89
+#define COHERENT_DEV_NAME  "coherent_memory"
+
+#define CRNT_NODE_NID1 1
+#define CRNT_NODE_NID2 2
+#define CRNT_NODE_NID3 3
+
+#define RAM_CRNT_MIGRATE 1
+#define CRNT_RAM_MIGRATE 2
+
+struct vma_map_info {
+	struct list_head list;
+	unsigned long nr_pages;
+	spinlock_t lock;
+};
+
+static void vma_map_info_init(struct vm_area_struct *vma)
+{
+	struct vma_map_info *info = kmalloc(sizeof(struct vma_map_info),
+								GFP_KERNEL);
+
+	BUG_ON(!info);
+	INIT_LIST_HEAD(&info->list);
+	spin_lock_init(&info->lock);
+	vma->vm_private_data = info;
+	info->nr_pages = 0;
+}
+
+static void coherent_vmops_open(struct vm_area_struct *vma)
+{
+	vma_map_info_init(vma);
+}
+
+static void coherent_vmops_close(struct vm_area_struct *vma)
+{
+	struct vma_map_info *info = vma->vm_private_data;
+
+	BUG_ON(!info);
+again:
+	cond_resched();
+	spin_lock(&info->lock);
+	while (info->nr_pages) {
+		struct page *page, *page2;
+
+		list_for_each_entry_safe(page, page2, &info->list, lru) {
+			if (!trylock_page(page)) {
+				spin_unlock(&info->lock);
+				goto again;
+			}
+
+			list_del_init(&page->lru);
+			info->nr_pages--;
+			unlock_page(page);
+			SetPageReclaim(page);
+			put_page(page);
+		}
+		spin_unlock(&info->lock);
+		cond_resched();
+		spin_lock(&info->lock);
+	}
+	spin_unlock(&info->lock);
+	kfree(info);
+	vma->vm_private_data = NULL;
+}
+
+static int coherent_vmops_fault(struct vm_area_struct *vma,
+					struct vm_fault *vmf)
+{
+	struct vma_map_info *info;
+	struct page *page;
+	static int coherent_node = CRNT_NODE_NID1;
+
+	if (coherent_node == CRNT_NODE_NID1)
+		coherent_node = CRNT_NODE_NID2;
+	else
+		coherent_node = CRNT_NODE_NID1;
+
+	page = alloc_pages_node(coherent_node,
+				GFP_HIGHUSER_MOVABLE | __GFP_THISNODE, 0);
+	if (!page)
+		return VM_FAULT_SIGBUS;
+
+	info = (struct vma_map_info *) vma->vm_private_data;
+	BUG_ON(!info);
+	spin_lock(&info->lock);
+	list_add(&page->lru, &info->list);
+	info->nr_pages++;
+	spin_unlock(&info->lock);
+
+	page->index = vmf->pgoff;
+	get_page(page);
+	vmf->page = page;
+	return 0;
+}
+
+static const struct vm_operations_struct coherent_memory_vmops = {
+	.open = coherent_vmops_open,
+	.close = coherent_vmops_close,
+	.fault = coherent_vmops_fault,
+};
+
+static int coherent_memory_mmap(struct file *file, struct vm_area_struct *vma)
+{
+	pr_info("Mmap opened (file: %lx vma: %lx)\n",
+			(unsigned long) file, (unsigned long) vma);
+	vma->vm_ops = &coherent_memory_vmops;
+	coherent_vmops_open(vma);
+	return 0;
+}
+
+static int coherent_memory_open(struct inode *inode, struct file *file)
+{
+	pr_info("Device opened (inode: %lx file: %lx)\n",
+			(unsigned long) inode, (unsigned long) file);
+	return 0;
+}
+
+static int coherent_memory_close(struct inode *inode, struct file *file)
+{
+	pr_info("Device closed (inode: %lx file: %lx)\n",
+			(unsigned long) inode, (unsigned long) file);
+	return 0;
+}
+
+static void lru_ram_coherent_migrate(unsigned long addr)
+{
+	struct mm_struct *mm = current->mm;
+	struct vm_area_struct *vma;
+	nodemask_t nmask;
+	LIST_HEAD(mlist);
+
+	nodes_clear(nmask);
+	nodes_setall(nmask);
+	down_write(&mm->mmap_sem);
+	for (vma = mm->mmap; vma; vma = vma->vm_next) {
+		if  ((addr < vma->vm_start) || (addr > vma->vm_end))
+			continue;
+		break;
+	}
+	up_write(&mm->mmap_sem);
+	if (!vma) {
+		pr_info("%s: No VMA found\n", __func__);
+		return;
+	}
+	migrate_virtual_range(current->pid, vma->vm_start, vma->vm_end, 2);
+}
+
+static void lru_coherent_ram_migrate(unsigned long addr)
+{
+	struct mm_struct *mm = current->mm;
+	struct vm_area_struct *vma;
+	nodemask_t nmask;
+	LIST_HEAD(mlist);
+
+	nodes_clear(nmask);
+	nodes_setall(nmask);
+	down_write(&mm->mmap_sem);
+	for (vma = mm->mmap; vma; vma = vma->vm_next) {
+		if  ((addr < vma->vm_start) || (addr > vma->vm_end))
+			continue;
+		break;
+	}
+	up_write(&mm->mmap_sem);
+	if (!vma) {
+		pr_info("%s: No VMA found\n", __func__);
+		return;
+	}
+	migrate_virtual_range(current->pid, vma->vm_start, vma->vm_end, 0);
+}
+
+static long coherent_memory_ioctl(struct file *file,
+					unsigned int cmd, unsigned long arg)
+{
+	switch (cmd) {
+	case RAM_CRNT_MIGRATE:
+		lru_ram_coherent_migrate(arg);
+		break;
+
+	case CRNT_RAM_MIGRATE:
+		lru_coherent_ram_migrate(arg);
+		break;
+
+	default:
+		pr_info("%s Invalid ioctl() command: %d\n", __func__, cmd);
+		return -EINVAL;
+	}
+	return 0;
+}
+
+static const struct file_operations fops = {
+	.mmap = coherent_memory_mmap,
+	.open = coherent_memory_open,
+	.release = coherent_memory_close,
+	.unlocked_ioctl = &coherent_memory_ioctl
+};
+
+static char kbuf[100];	/* Will store original user passed buffer */
+static char str[100];	/* Working copy for individual substring */
+
+static u64 args[4];
+static u64 index;
+static void convert_substring(const char *buf)
+{
+	u64 val = 0;
+
+	if (kstrtou64(buf, 0, &val))
+		pr_info("String conversion failed\n");
+
+	args[index] = val;
+	index++;
+}
+
+static ssize_t coherent_debug_write(struct file *file,
+					const char __user *user_buf,
+					size_t count, loff_t *ppos)
+{
+	char *tmp, *tmp1;
+	size_t ret;
+
+	memset(args, 0, sizeof(args));
+	index = 0;
+
+	ret = simple_write_to_buffer(kbuf, sizeof(kbuf), ppos, user_buf, count);
+	if (ret < 0)
+		return ret;
+
+	kbuf[ret] = '\0';
+	tmp = kbuf;
+	do {
+		tmp1 = strchr(tmp, ',');
+		if (tmp1) {
+			*tmp1 = '\0';
+			strncpy(str, (const char *)tmp, strlen(tmp));
+			convert_substring(str);
+		} else {
+			strncpy(str, (const char *)tmp, strlen(tmp));
+			convert_substring(str);
+			break;
+		}
+		tmp = tmp1 + 1;
+		memset(str, 0, sizeof(str));
+	} while (true);
+	migrate_virtual_range(args[0], args[1], args[2], args[3]);
+	return ret;
+}
+
+static int coherent_debug_show(struct seq_file *m, void *v)
+{
+	seq_puts(m, "Expected Value: <pid,vaddr,size,nid>\n");
+	return 0;
+}
+
+static int coherent_debug_open(struct inode *inode, struct file *filp)
+{
+	return single_open(filp, coherent_debug_show, NULL);
+}
+
+static const struct file_operations coherent_debug_fops = {
+	.open		= coherent_debug_open,
+	.write		= coherent_debug_write,
+	.read		= seq_read,
+	.llseek		= seq_lseek,
+	.release	= single_release,
+};
+
+static struct dentry *debugfile;
+
+static void coherent_memory_debugfs(void)
+{
+
+	debugfile = debugfs_create_file("coherent_debug", 0644, NULL, NULL,
+				&coherent_debug_fops);
+	if (!debugfile)
+		pr_warn("Failed to create coherent_memory in debugfs");
+}
+
+static void __exit coherent_memory_exit(void)
+{
+	pr_info("%s\n", __func__);
+	debugfs_remove(debugfile);
+	unregister_chrdev(COHERENT_DEV_MAJOR, COHERENT_DEV_NAME);
+}
+
+static int __init coherent_memory_init(void)
+{
+	int ret;
+
+	pr_info("%s\n", __func__);
+	ret = register_chrdev(COHERENT_DEV_MAJOR, COHERENT_DEV_NAME, &fops);
+	if (ret < 0) {
+		pr_info("%s register_chrdev() failed\n", __func__);
+		return -1;
+	}
+	coherent_memory_debugfs();
+	return 0;
+}
+
+module_init(coherent_memory_init);
+module_exit(coherent_memory_exit);
+MODULE_LICENSE("GPL");
diff --git a/drivers/char/memory_online_sysfs.h b/drivers/char/memory_online_sysfs.h
new file mode 100644
index 0000000..a5f022d
--- /dev/null
+++ b/drivers/char/memory_online_sysfs.h
@@ -0,0 +1,148 @@
+/*
+ * Accessing sysfs interface for memory hotplug operation from
+ * inside the kernel.
+ *
+ * Licensed under GPL V2
+ */
+#ifndef __SYSFS_H
+#define __SYSFS_H
+
+#include <linux/fs.h>
+#include <linux/uaccess.h>
+
+#define AUTO_ONLINE_BLOCKS "/sys/devices/system/memory/auto_online_blocks"
+#define BLOCK_SIZE_BYTES   "/sys/devices/system/memory/block_size_bytes"
+#define MEMORY_PROBE       "/sys/devices/system/memory/probe"
+
+static ssize_t read_buf(char *filename, char *buf, ssize_t count)
+{
+	mm_segment_t old_fs;
+	struct file *filp;
+	loff_t pos = 0;
+
+	if (!count)
+		return 0;
+
+	old_fs = get_fs();
+	set_fs(KERNEL_DS);
+
+	filp = filp_open(filename, O_RDONLY, 0);
+	if (IS_ERR(filp)) {
+		count = PTR_ERR(filp);
+		goto err_open;
+	}
+
+	count = vfs_read(filp, buf, count - 1, &pos);
+	buf[count] = '\0';
+
+	filp_close(filp, NULL);
+
+err_open:
+	set_fs(old_fs);
+
+	return count;
+}
+
+static unsigned long long read_0x(char *filename)
+{
+	unsigned long long ret;
+	char buf[32];
+
+	if (read_buf(filename, buf, 32) <= 0)
+		return 0;
+
+	if (kstrtoull(buf, 16, &ret))
+		return 0;
+
+	return ret;
+}
+
+static ssize_t write_buf(char *filename, char *buf)
+{
+	int ret;
+	mm_segment_t old_fs;
+	struct file *filp;
+	loff_t pos = 0;
+
+	old_fs = get_fs();
+	set_fs(KERNEL_DS);
+
+	filp = filp_open(filename, O_WRONLY, 0);
+	if (IS_ERR(filp)) {
+		ret = PTR_ERR(filp);
+		goto err_open;
+	}
+
+	ret = vfs_write(filp, buf, strlen(buf), &pos);
+
+	filp_close(filp, NULL);
+
+err_open:
+	set_fs(old_fs);
+
+	return ret;
+}
+
+int memory_probe_store(phys_addr_t addr, phys_addr_t size)
+{
+	phys_addr_t block_sz =
+		read_0x(BLOCK_SIZE_BYTES);
+	long i;
+
+	for (i = 0; i < size / block_sz; i++, addr += block_sz) {
+		char s[32];
+		ssize_t count;
+
+		snprintf(s, 32, "0x%llx", addr);
+
+		count = write_buf(MEMORY_PROBE, s);
+		if (count < 0)
+			return count;
+	}
+
+	return 0;
+}
+
+int store_mem_state(phys_addr_t addr, phys_addr_t size, char *state)
+{
+	phys_addr_t block_sz = read_0x(BLOCK_SIZE_BYTES);
+	unsigned long start_block, end_block, i;
+
+	start_block = addr / block_sz;
+	end_block = start_block + size / block_sz;
+
+	for (i = end_block - 1; i >= start_block; i--) {
+		char filename[64];
+		ssize_t count;
+
+		snprintf(filename, 64,
+			 "/sys/devices/system/memory/memory%ld/state", i);
+
+		count = write_buf(filename, state);
+		if (count < 0)
+			return count;
+	}
+
+	return 0;
+}
+
+int disable_auto_online(void)
+{
+	int ret;
+
+	ret = write_buf(AUTO_ONLINE_BLOCKS, "offline");
+	if (ret)
+		return ret;
+	return 0;
+}
+
+int enable_auto_online(void)
+{
+	int ret;
+
+	ret = write_buf(AUTO_ONLINE_BLOCKS, "online");
+	if (ret)
+		return ret;
+	return 0;
+}
+#endif
diff --git a/mm/migrate.c b/mm/migrate.c
index 06300bb..1fb2b19 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -1405,6 +1405,7 @@ int migrate_virtual_range(int pid, unsigned long start,
 	struct vm_area_struct *vma;
 	nodemask_t nmask;
 	int ret = -EINVAL;
+	bool found = false;
 
 	LIST_HEAD(mlist);
 
@@ -1414,6 +1415,7 @@ int migrate_virtual_range(int pid, unsigned long start,
 	if ((!start) || (!end))
 		return -EINVAL;
 
+	pr_info("%s: %d %lx %lx %d: ", __func__, pid, start, end, nid);
 	rcu_read_lock();
 	mm = find_task_by_vpid(pid)->mm;
 	rcu_read_unlock();
@@ -1425,14 +1427,17 @@ int migrate_virtual_range(int pid, unsigned long start,
 		if  ((start < vma->vm_start) || (end > vma->vm_end))
 			continue;
 
+		found = true;
 		ret = queue_pages_range(mm, start, end, &nmask, MPOL_MF_MOVE_ALL
 						| MPOL_MF_DISCONTIG_OK, &mlist);
 		if (ret) {
+			pr_info("queue_pages_range_failed\n");
 			putback_movable_pages(&mlist);
 			break;
 		}
 
 		if (list_empty(&mlist)) {
+			pr_info("list_empty\n");
 			ret = -ENOMEM;
 			break;
 		}
@@ -1440,12 +1445,17 @@ int migrate_virtual_range(int pid, unsigned long start,
 		ret = migrate_pages(&mlist, new_node_page, NULL, nid,
 					MIGRATE_SYNC, MR_COMPACTION);
 		if (ret) {
+			pr_info("migration_failed\n");
 			putback_movable_pages(&mlist);
 		} else {
+			pr_info("migration_passed\n");
 			if (isolated_cdm_node(nid))
 				mark_vma_cdm(vma);
 		}
 	}
+	if (!found)
+		pr_info("vma_missing\n");
+
 	up_write(&mm->mmap_sem);
 	return ret;
 }
-- 
2.1.0

[toc] | [prev] | [next] | [standalone]


#1507467

FromJerome Glisse <j.glisse@gmail.com>
Date2016-10-24 19:10 +0200
Message-ID<svQV4-26N-27@gated-at.bofh.it>
In reply to#1506875
On Mon, Oct 24, 2016 at 10:01:49AM +0530, Anshuman Khandual wrote:

> [...]

> 	Core kernel memory features like reclamation, evictions etc. might
> need to be restricted or modified on the coherent device memory node as
> they can be performance limiting. The RFC does not propose anything on this
> yet but it can be looked into later on. For now it just disables Auto NUMA
> for any VMA which has coherent device memory.
> 
> 	Seamless integration of coherent device memory with system memory
> will enable various other features, some of which can be listed as follows.
> 
> 	a. Seamless migrations between system RAM and the coherent memory
> 	b. Will have asynchronous and high throughput migrations
> 	c. Be able to allocate huge order pages from these memory regions
> 	d. Restrict allocations to a large extent to the tasks using the
> 	   device for workload acceleration
> 
> 	Before concluding, will look into the reasons why the existing
> solutions don't work. There are two basic requirements which have to be
> satisfies before the coherent device memory can be integrated with core
> kernel seamlessly.
> 
> 	a. PFN must have struct page
> 	b. Struct page must able to be inside standard LRU lists
> 
> 	The above two basic requirements discard the existing method of
> device memory representation approaches like these which then requires the
> need of creating a new framework.

I do not believe the LRU list is a hard requirement, yes when faulting in
a page inside the page cache it assumes it needs to be added to lru list.
But i think this can easily be work around.

In HMM i am using ZONE_DEVICE and because memory is not accessible from CPU
(not everyone is bless with decent system bus like CAPI, CCIX, Gen-Z, ...)
so in my case a file back page must always be spawn first from a regular
page and once read from disk then i can migrate to GPU page.

So if you accept this intermediary step you can easily use ZONE_DEVICE for
device memory. This way no lru, no complex dance to make the memory out of
reach from regular memory allocator.

I think we would have much to gain if we pool our effort on a single common
solution for device memory. In my case the device memory is not accessible
by the CPU (because PCIE restrictions), in your case it is. Thus the only
difference is that in my case it can not be map inside the CPU page table
while in yours it can.

> 
> (1) Traditional ioremap
> 
> 	a. Memory is mapped into kernel (linear and virtual) and user space
> 	b. These PFNs do not have struct pages associated with it
> 	c. These special PFNs are marked with special flags inside the PTE
> 	d. Cannot participate in core VM functions much because of this
> 	e. Cannot do easy user space migrations
> 
> (2) Zone ZONE_DEVICE
> 
> 	a. Memory is mapped into kernel and user space
> 	b. PFNs do have struct pages associated with it
> 	c. These struct pages are allocated inside it's own memory range
> 	d. Unfortunately the struct page's union containing LRU has been
> 	   used for struct dev_pagemap pointer
> 	e. Hence it cannot be part of any LRU (like Page cache)
> 	f. Hence file cached mapping cannot reside on these PFNs
> 	g. Cannot do easy migrations
> 
> 	I had also explored non LRU representation of this coherent device
> memory where the integration with system RAM in the core VM is limited only
> to the following functions. Not being inside LRU is definitely going to
> reduce the scope of tight integration with system RAM.
> 
> (1) Migration support between system RAM and coherent memory
> (2) Migration support between various coherent memory nodes
> (3) Isolation of the coherent memory
> (4) Mapping the coherent memory into user space through driver's
>     struct vm_operations
> (5) HW poisoning of the coherent memory
> 
> 	Allocating the entire memory of the coherent device node right
> after hot plug into ZONE_MOVABLE (where the memory is already inside the
> buddy system) will still expose a time window where other user space
> allocations can come into the coherent device memory node and prevent the
> intended isolation. So traditional hot plug is not the solution. Hence
> started looking into CMA based non LRU solution but then hit the following
> roadblocks.
> 
> (1) CMA does not support hot plugging of new memory node
> 	a. CMA area needs to be marked during boot before buddy is
> 	   initialized
> 	b. cma_alloc()/cma_release() can happen on the marked area
> 	c. Should be able to mark the CMA areas just after memory hot plug
> 	d. cma_alloc()/cma_release() can happen later after the hot plug
> 	e. This is not currently supported right now
> 
> (2) Mapped non LRU migration of pages
> 	a. Recent work from Michan Kim makes non LRU page migratable
> 	b. But it still does not support migration of mapped non LRU pages
> 	c. With non LRU CMA reserved, again there are some additional
> 	   challenges
> 
> 	With hot pluggable CMA and non LRU mapped migration support there
> may be an alternate approach to represent coherent device memory. Please
> do review this RFC proposal and let me know your comments or suggestions.
> Thank you.

You can take a look at hmm-v13 if you want to see how i do non LRU page
migration. While i put most of the migration code inside hmm_migrate.c it
could easily be move to migrate.c without hmm_ prefix.

There is 2 missing piece with existing migrate code. First is to put memory
allocation for destination under control of who call the migrate code. Second
is to allow offloading the copy operation to device (ie not use the CPU to
copy data).

I believe same requirement also make sense for platform you are targeting.
Thus same code can be use.

hmm-v13 https://cgit.freedesktop.org/~glisse/linux/log/?h=hmm-v13

I haven't posted this patchset yet because we are doing some modifications
to the device driver API to accomodate some new features. But the ZONE_DEVICE
changes and the overall migration code will stay the same more or less (i have
patches that move it to migrate.c and share more code with existing migrate
code).

If you think i missed anything about lru and page cache please point it to
me. Because when i audited code for that i didn't see any road block with
the few fs i was looking at (ext4, xfs and core page cache code).

> [...]

Cheers,
Jérôme

[toc] | [prev] | [next] | [standalone]


#1507965

From"Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com>
Date2016-10-25 06:30 +0200
Message-ID<sw1x7-Gn-7@gated-at.bofh.it>
In reply to#1507467
Jerome Glisse <j.glisse@gmail.com> writes:

> On Mon, Oct 24, 2016 at 10:01:49AM +0530, Anshuman Khandual wrote:
>
>> [...]
>
>> 	Core kernel memory features like reclamation, evictions etc. might
>> need to be restricted or modified on the coherent device memory node as
>> they can be performance limiting. The RFC does not propose anything on this
>> yet but it can be looked into later on. For now it just disables Auto NUMA
>> for any VMA which has coherent device memory.
>> 
>> 	Seamless integration of coherent device memory with system memory
>> will enable various other features, some of which can be listed as follows.
>> 
>> 	a. Seamless migrations between system RAM and the coherent memory
>> 	b. Will have asynchronous and high throughput migrations
>> 	c. Be able to allocate huge order pages from these memory regions
>> 	d. Restrict allocations to a large extent to the tasks using the
>> 	   device for workload acceleration
>> 
>> 	Before concluding, will look into the reasons why the existing
>> solutions don't work. There are two basic requirements which have to be
>> satisfies before the coherent device memory can be integrated with core
>> kernel seamlessly.
>> 
>> 	a. PFN must have struct page
>> 	b. Struct page must able to be inside standard LRU lists
>> 
>> 	The above two basic requirements discard the existing method of
>> device memory representation approaches like these which then requires the
>> need of creating a new framework.
>
> I do not believe the LRU list is a hard requirement, yes when faulting in
> a page inside the page cache it assumes it needs to be added to lru list.
> But i think this can easily be work around.
>
> In HMM i am using ZONE_DEVICE and because memory is not accessible from CPU
> (not everyone is bless with decent system bus like CAPI, CCIX, Gen-Z, ...)
> so in my case a file back page must always be spawn first from a regular
> page and once read from disk then i can migrate to GPU page.
>
> So if you accept this intermediary step you can easily use ZONE_DEVICE for
> device memory. This way no lru, no complex dance to make the memory out of
> reach from regular memory allocator.

One of the reason to look at this as a NUMA node is to allow things like
over-commit of coherent device memory. The pages backing CDM being part of
lru and considering the coherent device as a numa node makes that really
simpler (we can run kswapd for that node).


>
> I think we would have much to gain if we pool our effort on a single common
> solution for device memory. In my case the device memory is not accessible
> by the CPU (because PCIE restrictions), in your case it is. Thus the only
> difference is that in my case it can not be map inside the CPU page table
> while in yours it can.

IMHO, we should be able to share the HMM migration approach. We
definitely won't need the mirror page table part. That is one of the
reson I requested HMM mirror page table to be a seperate patchset.


>
>> 
>> (1) Traditional ioremap
>> 
>> 	a. Memory is mapped into kernel (linear and virtual) and user space
>> 	b. These PFNs do not have struct pages associated with it
>> 	c. These special PFNs are marked with special flags inside the PTE
>> 	d. Cannot participate in core VM functions much because of this
>> 	e. Cannot do easy user space migrations
>> 
>> (2) Zone ZONE_DEVICE
>> 
>> 	a. Memory is mapped into kernel and user space
>> 	b. PFNs do have struct pages associated with it
>> 	c. These struct pages are allocated inside it's own memory range
>> 	d. Unfortunately the struct page's union containing LRU has been
>> 	   used for struct dev_pagemap pointer
>> 	e. Hence it cannot be part of any LRU (like Page cache)
>> 	f. Hence file cached mapping cannot reside on these PFNs
>> 	g. Cannot do easy migrations
>> 
>> 	I had also explored non LRU representation of this coherent device
>> memory where the integration with system RAM in the core VM is limited only
>> to the following functions. Not being inside LRU is definitely going to
>> reduce the scope of tight integration with system RAM.
>> 
>> (1) Migration support between system RAM and coherent memory
>> (2) Migration support between various coherent memory nodes
>> (3) Isolation of the coherent memory
>> (4) Mapping the coherent memory into user space through driver's
>>     struct vm_operations
>> (5) HW poisoning of the coherent memory
>> 
>> 	Allocating the entire memory of the coherent device node right
>> after hot plug into ZONE_MOVABLE (where the memory is already inside the
>> buddy system) will still expose a time window where other user space
>> allocations can come into the coherent device memory node and prevent the
>> intended isolation. So traditional hot plug is not the solution. Hence
>> started looking into CMA based non LRU solution but then hit the following
>> roadblocks.
>> 
>> (1) CMA does not support hot plugging of new memory node
>> 	a. CMA area needs to be marked during boot before buddy is
>> 	   initialized
>> 	b. cma_alloc()/cma_release() can happen on the marked area
>> 	c. Should be able to mark the CMA areas just after memory hot plug
>> 	d. cma_alloc()/cma_release() can happen later after the hot plug
>> 	e. This is not currently supported right now
>> 
>> (2) Mapped non LRU migration of pages
>> 	a. Recent work from Michan Kim makes non LRU page migratable
>> 	b. But it still does not support migration of mapped non LRU pages
>> 	c. With non LRU CMA reserved, again there are some additional
>> 	   challenges
>> 
>> 	With hot pluggable CMA and non LRU mapped migration support there
>> may be an alternate approach to represent coherent device memory. Please
>> do review this RFC proposal and let me know your comments or suggestions.
>> Thank you.
>
> You can take a look at hmm-v13 if you want to see how i do non LRU page
> migration. While i put most of the migration code inside hmm_migrate.c it
> could easily be move to migrate.c without hmm_ prefix.
>
> There is 2 missing piece with existing migrate code. First is to put memory
> allocation for destination under control of who call the migrate code. Second
> is to allow offloading the copy operation to device (ie not use the CPU to
> copy data).
>
> I believe same requirement also make sense for platform you are targeting.
> Thus same code can be use.
>
> hmm-v13 https://cgit.freedesktop.org/~glisse/linux/log/?h=hmm-v13
>
> I haven't posted this patchset yet because we are doing some modifications
> to the device driver API to accomodate some new features. But the ZONE_DEVICE
> changes and the overall migration code will stay the same more or less (i have
> patches that move it to migrate.c and share more code with existing migrate
> code).
>
> If you think i missed anything about lru and page cache please point it to
> me. Because when i audited code for that i didn't see any road block with
> the few fs i was looking at (ext4, xfs and core page cache code).

I looked at the hmm-v13 w.r.t migration and I guess some form of device
callback/acceleration during migration is something we should definitely
have. I still haven't figured out how non addressable and coherent device
memory can fit together there. I was waiting for the page cache
migration support to be pushed to the repository before I start looking
at this closely.

-aneesh

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.kernel


csiph-web