Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1578347 > unrolled thread
| Started by | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| First post | 2017-02-10 11:10 +0100 |
| Last post | 2017-02-13 16:40 +0100 |
| Articles | 5 — 3 participants |
Back to article view | Back to linux.kernel
[PATCH V2 0/3] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-02-10 11:10 +0100
[PATCH V2 3/3] mm: Enable Buddy allocation isolation for CDM nodes Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-02-10 11:10 +0100
[PATCH V2 1/3] mm: Define coherent device memory (CDM) node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-02-10 11:10 +0100
Re: [PATCH V2 1/3] mm: Define coherent device memory (CDM) node John Hubbard <jhubbard@nvidia.com> - 2017-02-13 05:10 +0100
Re: [PATCH V2 0/3] Define coherent device memory node Vlastimil Babka <vbabka@suse.cz> - 2017-02-13 16:40 +0100
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-02-10 11:10 +0100 |
| Subject | [PATCH V2 0/3] Define coherent device memory node |
| Message-ID | <t9gjn-23D-3@gated-at.bofh.it> |
This three patches define CDM node with HugeTLB & Buddy allocation isolation. Please refer to the last RFC posting mentioned here for details. The series has been split for easier review process. The next part of the work like VM flags, auto NUMA and KSM interactions with tagged VMAs will follow later. https://lkml.org/lkml/2017/1/29/198 Changes in V2: * Removed redundant nodemask_has_cdm() check from zonelist iterator * Dropped the nodemask_had_cdm() function itself * Added node_set/clear_state_cdm() functions and removed bunch of #ifdefs * Moved CDM helper functions into nodemask.h from node.h header file * Fixed the build failure by additional CONFIG_NEED_MULTIPLE_NODES check Previous V1: (https://lkml.org/lkml/2017/2/8/329) Anshuman Khandual (3): mm: Define coherent device memory (CDM) node mm: Enable HugeTLB allocation isolation for CDM nodes mm: Enable Buddy allocation isolation for CDM nodes Documentation/ABI/stable/sysfs-devices-node | 7 ++++ arch/powerpc/Kconfig | 1 + arch/powerpc/mm/numa.c | 7 ++++ drivers/base/node.c | 6 +++ include/linux/nodemask.h | 58 ++++++++++++++++++++++++++++- mm/Kconfig | 4 ++ mm/hugetlb.c | 25 ++++++++----- mm/memory_hotplug.c | 3 ++ mm/page_alloc.c | 24 +++++++++++- 9 files changed, 123 insertions(+), 12 deletions(-) -- 2.9.3
[toc] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-02-10 11:10 +0100 |
| Subject | [PATCH V2 3/3] mm: Enable Buddy allocation isolation for CDM nodes |
| Message-ID | <t9gjo-23D-37@gated-at.bofh.it> |
| In reply to | #1578347 |
This implements allocation isolation for CDM nodes in buddy allocator by
discarding CDM memory zones all the time except in the cases where the gfp
flag has got __GFP_THISNODE or the nodemask contains CDM nodes in cases
where it is non NULL (explicit allocation request in the kernel or user
process MPOL_BIND policy based requests).
Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
mm/page_alloc.c | 16 ++++++++++++++++
1 file changed, 16 insertions(+)
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 84d61bb..392c24a 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -64,6 +64,7 @@
#include <linux/page_owner.h>
#include <linux/kthread.h>
#include <linux/memcontrol.h>
+#include <linux/node.h>
#include <asm/sections.h>
#include <asm/tlbflush.h>
@@ -2908,6 +2909,21 @@ get_page_from_freelist(gfp_t gfp_mask, unsigned int order, int alloc_flags,
struct page *page;
unsigned long mark;
+ /*
+ * CDM nodes get skipped if the requested gfp flag
+ * does not have __GFP_THISNODE set or the nodemask
+ * does not have any CDM nodes in case the nodemask
+ * is non NULL (explicit allocation requests from
+ * kernel or user process MPOL_BIND policy which has
+ * CDM nodes).
+ */
+ if (is_cdm_node(zone->zone_pgdat->node_id)) {
+ if (!(gfp_mask & __GFP_THISNODE)) {
+ if (!ac->nodemask)
+ continue;
+ }
+ }
+
if (cpusets_enabled() &&
(alloc_flags & ALLOC_CPUSET) &&
!__cpuset_zone_allowed(zone, gfp_mask))
--
2.9.3
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-02-10 11:10 +0100 |
| Subject | [PATCH V2 1/3] mm: Define coherent device memory (CDM) node |
| Message-ID | <t9gjo-23D-25@gated-at.bofh.it> |
| In reply to | #1578347 |
There are certain devices like specialized accelerator, GPU cards, network
cards, FPGA cards etc which might contain onboard memory which is coherent
along with the existing system RAM while being accessed either from the CPU
or from the device. They share some similar properties with that of normal
system RAM but at the same time can also be different with respect to
system RAM.
User applications might be interested in using this kind of coherent device
memory explicitly or implicitly along side the system RAM utilizing all
possible core memory functions like anon mapping (LRU), file mapping (LRU),
page cache (LRU), driver managed (non LRU), HW poisoning, NUMA migrations
etc. To achieve this kind of tight integration with core memory subsystem,
the device onboard coherent memory must be represented as a memory only
NUMA node. At the same time arch must export some kind of a function to
identify of this node as a coherent device memory not any other regular
cpu less memory only NUMA node.
After achieving the integration with core memory subsystem coherent device
memory might still need some special consideration inside the kernel. There
can be a variety of coherent memory nodes with different expectations from
the core kernel memory. But right now only one kind of special treatment is
considered which requires certain isolation.
Now consider the case of a coherent device memory node type which requires
isolation. This kind of coherent memory is onboard an external device
attached to the system through a link where there is always a chance of a
link failure taking down the entire memory node with it. More over the
memory might also have higher chance of ECC failure as compared to the
system RAM. Hence allocation into this kind of coherent memory node should
be regulated. Kernel allocations must not come here. Normal user space
allocations too should not come here implicitly (without user application
knowing about it). This summarizes isolation requirement of certain kind of
coherent device memory node as an example. There can be different kinds of
isolation requirement also.
Some coherent memory devices might not require isolation altogether after
all. Then there might be other coherent memory devices which might require
some other special treatment after being part of core memory representation
. For now, will look into isolation seeking coherent device memory node not
the other ones.
To implement the integration as well as isolation, the coherent memory node
must be present in N_MEMORY and a new N_COHERENT_DEVICE node mask inside
the node_states[] array. During memory hotplug operations, the new nodemask
N_COHERENT_DEVICE is updated along with N_MEMORY for these coherent device
memory nodes. This also creates the following new sysfs based interface to
list down all the coherent memory nodes of the system.
/sys/devices/system/node/is_coherent_node
Architectures must export function arch_check_node_cdm() which identifies
any coherent device memory node in case they enable CONFIG_COHERENT_DEVICE.
Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
Documentation/ABI/stable/sysfs-devices-node | 7 ++++
arch/powerpc/Kconfig | 1 +
arch/powerpc/mm/numa.c | 7 ++++
drivers/base/node.c | 6 +++
include/linux/nodemask.h | 58 ++++++++++++++++++++++++++++-
mm/Kconfig | 4 ++
mm/memory_hotplug.c | 3 ++
mm/page_alloc.c | 8 +++-
8 files changed, 91 insertions(+), 3 deletions(-)
diff --git a/Documentation/ABI/stable/sysfs-devices-node b/Documentation/ABI/stable/sysfs-devices-node
index 5b2d0f0..fa2f105 100644
--- a/Documentation/ABI/stable/sysfs-devices-node
+++ b/Documentation/ABI/stable/sysfs-devices-node
@@ -29,6 +29,13 @@ Description:
Nodes that have regular or high memory.
Depends on CONFIG_HIGHMEM.
+What: /sys/devices/system/node/is_coherent_device
+Date: January 2017
+Contact: Linux Memory Management list <linux-mm@kvack.org>
+Description:
+ Lists the nodemask of nodes that have coherent device memory.
+ Depends on CONFIG_COHERENT_DEVICE.
+
What: /sys/devices/system/node/nodeX
Date: October 2002
Contact: Linux Memory Management list <linux-mm@kvack.org>
diff --git a/arch/powerpc/Kconfig b/arch/powerpc/Kconfig
index 281f4f1..1cff239 100644
--- a/arch/powerpc/Kconfig
+++ b/arch/powerpc/Kconfig
@@ -164,6 +164,7 @@ config PPC
select ARCH_HAS_SCALED_CPUTIME if VIRT_CPU_ACCOUNTING_NATIVE
select HAVE_ARCH_HARDENED_USERCOPY
select HAVE_KERNEL_GZIP
+ select COHERENT_DEVICE if PPC_BOOK3S_64 && NEED_MULTIPLE_NODES
config GENERIC_CSUM
def_bool CPU_LITTLE_ENDIAN
diff --git a/arch/powerpc/mm/numa.c b/arch/powerpc/mm/numa.c
index b1099cb..14f0b98 100644
--- a/arch/powerpc/mm/numa.c
+++ b/arch/powerpc/mm/numa.c
@@ -41,6 +41,13 @@
#include <asm/setup.h>
#include <asm/vdso.h>
+#ifdef CONFIG_COHERENT_DEVICE
+inline int arch_check_node_cdm(int nid)
+{
+ return 0;
+}
+#endif
+
static int numa_enabled = 1;
static char *cmdline __initdata;
diff --git a/drivers/base/node.c b/drivers/base/node.c
index 5548f96..5b5dd89 100644
--- a/drivers/base/node.c
+++ b/drivers/base/node.c
@@ -661,6 +661,9 @@ static struct node_attr node_state_attr[] = {
[N_MEMORY] = _NODE_ATTR(has_memory, N_MEMORY),
#endif
[N_CPU] = _NODE_ATTR(has_cpu, N_CPU),
+#ifdef CONFIG_COHERENT_DEVICE
+ [N_COHERENT_DEVICE] = _NODE_ATTR(is_coherent_device, N_COHERENT_DEVICE),
+#endif
};
static struct attribute *node_state_attrs[] = {
@@ -674,6 +677,9 @@ static struct attribute *node_state_attrs[] = {
&node_state_attr[N_MEMORY].attr.attr,
#endif
&node_state_attr[N_CPU].attr.attr,
+#ifdef CONFIG_COHERENT_DEVICE
+ &node_state_attr[N_COHERENT_DEVICE].attr.attr,
+#endif
NULL
};
diff --git a/include/linux/nodemask.h b/include/linux/nodemask.h
index f746e44..175c2d6 100644
--- a/include/linux/nodemask.h
+++ b/include/linux/nodemask.h
@@ -388,11 +388,14 @@ enum node_states {
N_HIGH_MEMORY = N_NORMAL_MEMORY,
#endif
#ifdef CONFIG_MOVABLE_NODE
- N_MEMORY, /* The node has memory(regular, high, movable) */
+ N_MEMORY, /* The node has memory(regular, high, movable, cdm) */
#else
N_MEMORY = N_HIGH_MEMORY,
#endif
N_CPU, /* The node has one or more cpus */
+#ifdef CONFIG_COHERENT_DEVICE
+ N_COHERENT_DEVICE, /* The node has CDM memory */
+#endif
NR_NODE_STATES
};
@@ -496,6 +499,59 @@ static inline int node_random(const nodemask_t *mask)
}
#endif
+#ifdef CONFIG_COHERENT_DEVICE
+extern int arch_check_node_cdm(int nid);
+
+static inline nodemask_t system_mem_nodemask(void)
+{
+ nodemask_t system_mem;
+
+ nodes_clear(system_mem);
+ nodes_andnot(system_mem, node_states[N_MEMORY],
+ node_states[N_COHERENT_DEVICE]);
+ return system_mem;
+}
+
+static inline bool is_cdm_node(int node)
+{
+ return node_isset(node, node_states[N_COHERENT_DEVICE]);
+}
+
+static inline void node_set_state_cdm(int node)
+{
+ if (arch_check_node_cdm(node))
+ node_set_state(node, N_COHERENT_DEVICE);
+}
+
+static inline void node_clear_state_cdm(int node)
+{
+ if (arch_check_node_cdm(node))
+ node_clear_state(node, N_COHERENT_DEVICE);
+}
+
+#else
+
+static inline int arch_check_node_cdm(int nid) { return 0; }
+
+static inline nodemask_t system_mem_nodemask(void)
+{
+ return node_states[N_MEMORY];
+}
+
+static inline bool is_cdm_node(int node)
+{
+ return false;
+}
+
+static inline void node_set_state_cdm(int node)
+{
+}
+
+static inline void node_clear_state_cdm(int node)
+{
+}
+#endif /* CONFIG_COHERENT_DEVICE */
+
#define node_online_map node_states[N_ONLINE]
#define node_possible_map node_states[N_POSSIBLE]
diff --git a/mm/Kconfig b/mm/Kconfig
index 9b8fccb..6263a65 100644
--- a/mm/Kconfig
+++ b/mm/Kconfig
@@ -143,6 +143,10 @@ config HAVE_GENERIC_RCU_GUP
config ARCH_DISCARD_MEMBLOCK
bool
+config COHERENT_DEVICE
+ bool
+ default n
+
config NO_BOOTMEM
bool
diff --git a/mm/memory_hotplug.c b/mm/memory_hotplug.c
index b8c11e0..6bce093 100644
--- a/mm/memory_hotplug.c
+++ b/mm/memory_hotplug.c
@@ -1030,6 +1030,7 @@ static void node_states_set_node(int node, struct memory_notify *arg)
if (arg->status_change_nid_high >= 0)
node_set_state(node, N_HIGH_MEMORY);
+ node_set_state_cdm(node);
node_set_state(node, N_MEMORY);
}
@@ -1843,6 +1844,8 @@ static void node_states_clear_node(int node, struct memory_notify *arg)
if ((N_MEMORY != N_HIGH_MEMORY) &&
(arg->status_change_nid >= 0))
node_clear_state(node, N_MEMORY);
+
+ node_clear_state_cdm(node);
}
static int __ref __offline_pages(unsigned long start_pfn,
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index f3e0c69..84d61bb 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -6080,8 +6080,10 @@ static unsigned long __init early_calculate_totalpages(void)
unsigned long pages = end_pfn - start_pfn;
totalpages += pages;
- if (pages)
+ if (pages) {
+ node_set_state_cdm(nid);
node_set_state(nid, N_MEMORY);
+ }
}
return totalpages;
}
@@ -6392,8 +6394,10 @@ void __init free_area_init_nodes(unsigned long *max_zone_pfn)
find_min_pfn_for_node(nid), NULL);
/* Any memory on that node */
- if (pgdat->node_present_pages)
+ if (pgdat->node_present_pages) {
+ node_set_state_cdm(nid);
node_set_state(nid, N_MEMORY);
+ }
check_for_memory(pgdat, nid);
}
}
--
2.9.3
[toc] | [prev] | [next] | [standalone]
| From | John Hubbard <jhubbard@nvidia.com> |
|---|---|
| Date | 2017-02-13 05:10 +0100 |
| Subject | Re: [PATCH V2 1/3] mm: Define coherent device memory (CDM) node |
| Message-ID | <tag7E-6VH-7@gated-at.bofh.it> |
| In reply to | #1578351 |
On 02/10/2017 02:06 AM, Anshuman Khandual wrote:
> There are certain devices like specialized accelerator, GPU cards, network
> cards, FPGA cards etc which might contain onboard memory which is coherent
> along with the existing system RAM while being accessed either from the CPU
> or from the device. They share some similar properties with that of normal
> system RAM but at the same time can also be different with respect to
> system RAM.
>
> User applications might be interested in using this kind of coherent device
> memory explicitly or implicitly along side the system RAM utilizing all
> possible core memory functions like anon mapping (LRU), file mapping (LRU),
> page cache (LRU), driver managed (non LRU), HW poisoning, NUMA migrations
> etc. To achieve this kind of tight integration with core memory subsystem,
> the device onboard coherent memory must be represented as a memory only
> NUMA node. At the same time arch must export some kind of a function to
> identify of this node as a coherent device memory not any other regular
> cpu less memory only NUMA node.
>
> After achieving the integration with core memory subsystem coherent device
> memory might still need some special consideration inside the kernel. There
> can be a variety of coherent memory nodes with different expectations from
> the core kernel memory. But right now only one kind of special treatment is
> considered which requires certain isolation.
>
> Now consider the case of a coherent device memory node type which requires
> isolation. This kind of coherent memory is onboard an external device
> attached to the system through a link where there is always a chance of a
> link failure taking down the entire memory node with it. More over the
> memory might also have higher chance of ECC failure as compared to the
> system RAM. Hence allocation into this kind of coherent memory node should
> be regulated. Kernel allocations must not come here. Normal user space
> allocations too should not come here implicitly (without user application
> knowing about it). This summarizes isolation requirement of certain kind of
> coherent device memory node as an example. There can be different kinds of
> isolation requirement also.
>
> Some coherent memory devices might not require isolation altogether after
> all. Then there might be other coherent memory devices which might require
> some other special treatment after being part of core memory representation
> . For now, will look into isolation seeking coherent device memory node not
> the other ones.
>
Hi Anshuman,
I'd question the need to avoid kernel allocations in device memory. Maybe we should simply allow
these pages to *potentially* participate in everything that N_MEMORY pages do: huge pages, kernel
allocations, for example.
There is a bit too much emphasis being placed on the idea that these devices are less reliable than
system memory. It's true--they are less reliable. However, they are reliable enough to be allowed
direct (coherent) addressing. And anything that allows that, is, IMHO, good enough to allow all
allocations on it.
On the point of what reliability implies: I've been involved in the development (and debugging) of
similar systems over the years, and what happens is: if the device has a fatal error, you have to
take the computer down, some time in the near future. There are a few reasons for this:
-- sometimes the MCE (machine check) is wired up to fire, if the device has errors, in which
case you are all done very quickly. :)
-- other times, the operating system relied upon now-corrupted data, that came from the device.
So even if you claim "OK, the device has a fatal error, but the OS can continue running just fine",
that's just wrong! You may have corrupted something important.
-- even if the above two didn't get you, you still have a likely expensive computer that cannot
do what you bought it for, so you've got to shut it down and replace the failed device.
Given all that, I think it is not especially worthwhile to design in a lot of constraints and
limitations around coherent device memory.
As for speed, we should be able to put in some hints to help with page placement. I'm still coming
up to speed with what is already there, and I'm sure other people can comment on that.
We should probably just let the allocations happen.
> To implement the integration as well as isolation, the coherent memory node
> must be present in N_MEMORY and a new N_COHERENT_DEVICE node mask inside
> the node_states[] array. During memory hotplug operations, the new nodemask
> N_COHERENT_DEVICE is updated along with N_MEMORY for these coherent device
> memory nodes. This also creates the following new sysfs based interface to
> list down all the coherent memory nodes of the system.
>
> /sys/devices/system/node/is_coherent_node
The naming bothers me: all nodes are coherent already. In fact, the Coherent Device Memory naming is
a little off-base already: what is it *really* trying to say? Less reliable? Slower?
My-special-device? :) Will those things even always be true? Makes me question the whole CDM
concept. Maybe just ZONE_MOVABLE (to handle hotplug) is the way to go.
>
> Architectures must export function arch_check_node_cdm() which identifies
> any coherent device memory node in case they enable CONFIG_COHERENT_DEVICE.
>
> Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
> ---
> Documentation/ABI/stable/sysfs-devices-node | 7 ++++
> arch/powerpc/Kconfig | 1 +
> arch/powerpc/mm/numa.c | 7 ++++
> drivers/base/node.c | 6 +++
> include/linux/nodemask.h | 58 ++++++++++++++++++++++++++++-
> mm/Kconfig | 4 ++
> mm/memory_hotplug.c | 3 ++
> mm/page_alloc.c | 8 +++-
> 8 files changed, 91 insertions(+), 3 deletions(-)
>
> diff --git a/Documentation/ABI/stable/sysfs-devices-node b/Documentation/ABI/stable/sysfs-devices-node
> index 5b2d0f0..fa2f105 100644
> --- a/Documentation/ABI/stable/sysfs-devices-node
> +++ b/Documentation/ABI/stable/sysfs-devices-node
> @@ -29,6 +29,13 @@ Description:
> Nodes that have regular or high memory.
> Depends on CONFIG_HIGHMEM.
>
> +What: /sys/devices/system/node/is_coherent_device
> +Date: January 2017
> +Contact: Linux Memory Management list <linux-mm@kvack.org>
> +Description:
> + Lists the nodemask of nodes that have coherent device memory.
> + Depends on CONFIG_COHERENT_DEVICE.
> +
> What: /sys/devices/system/node/nodeX
> Date: October 2002
> Contact: Linux Memory Management list <linux-mm@kvack.org>
> diff --git a/arch/powerpc/Kconfig b/arch/powerpc/Kconfig
> index 281f4f1..1cff239 100644
> --- a/arch/powerpc/Kconfig
> +++ b/arch/powerpc/Kconfig
> @@ -164,6 +164,7 @@ config PPC
> select ARCH_HAS_SCALED_CPUTIME if VIRT_CPU_ACCOUNTING_NATIVE
> select HAVE_ARCH_HARDENED_USERCOPY
> select HAVE_KERNEL_GZIP
> + select COHERENT_DEVICE if PPC_BOOK3S_64 && NEED_MULTIPLE_NODES
>
> config GENERIC_CSUM
> def_bool CPU_LITTLE_ENDIAN
> diff --git a/arch/powerpc/mm/numa.c b/arch/powerpc/mm/numa.c
> index b1099cb..14f0b98 100644
> --- a/arch/powerpc/mm/numa.c
> +++ b/arch/powerpc/mm/numa.c
> @@ -41,6 +41,13 @@
> #include <asm/setup.h>
> #include <asm/vdso.h>
>
> +#ifdef CONFIG_COHERENT_DEVICE
> +inline int arch_check_node_cdm(int nid)
> +{
> + return 0;
> +}
> +#endif
I'm not sure that we really need this exact sort of arch_ check. Seems like most arches could simply
support the possibility of a CDM node.
But we can probably table that question until we ensure that we want a new NUMA node type (vs.
ZONE_MOVABLE).
> +
> static int numa_enabled = 1;
>
[snip]
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index f3e0c69..84d61bb 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -6080,8 +6080,10 @@ static unsigned long __init early_calculate_totalpages(void)
> unsigned long pages = end_pfn - start_pfn;
>
> totalpages += pages;
> - if (pages)
> + if (pages) {
> + node_set_state_cdm(nid);
> node_set_state(nid, N_MEMORY);
> + }
> }
> return totalpages;
> }
> @@ -6392,8 +6394,10 @@ void __init free_area_init_nodes(unsigned long *max_zone_pfn)
> find_min_pfn_for_node(nid), NULL);
>
> /* Any memory on that node */
> - if (pgdat->node_present_pages)
> + if (pgdat->node_present_pages) {
> + node_set_state_cdm(nid);
> node_set_state(nid, N_MEMORY);
I like that you provide clean wrapper functions, but air-dropping them into all these routines (none
of the other node types have to do this) makes it look like CDM is sort of hacked in. :)
thanks
john h
> + }
> check_for_memory(pgdat, nid);
> }
> }
> --
> 2.9.3
>
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org. For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
>
[toc] | [prev] | [next] | [standalone]
| From | Vlastimil Babka <vbabka@suse.cz> |
|---|---|
| Date | 2017-02-13 16:40 +0100 |
| Message-ID | <taqTq-5xK-71@gated-at.bofh.it> |
| In reply to | #1578347 |
On 02/10/2017 11:06 AM, Anshuman Khandual wrote: > This three patches define CDM node with HugeTLB & Buddy allocation > isolation. Please refer to the last RFC posting mentioned here for details. > The series has been split for easier review process. The next part of the > work like VM flags, auto NUMA and KSM interactions with tagged VMAs will > follow later. Hi, I'm not sure if the splitting to smaller series and focusing on partial implementations is helpful at this point, until there's some consensus about the whole approach from a big picture perspective. Note that it's also confusing that v1 of this partial patchset mentioned some alternative implementations, but only as git branches, and the discussion about their differences is linked elsewhere. That further makes meaningful review harder IMHO. Going back to the bigger picture, I've read the comments on previous postings and I think Jerome makes many good points in this subthread [1] against the idea of representing the device memory as generic memory nodes and expecting userspace to mbind() to them. So if I make a program that uses mbind() to back some mmapped area with memory of "devices like accelerators, GPU cards, network cards, FPGA cards, PLD cards etc which might contain on board memory", then it will get such memory... and then what? How will it benefit from it? I will also need to tell some driver to make the device do some operations with this memory, right? And that most likely won't be a generic operation. In that case I can also ask the driver to give me that memory in the first place, and it can apply whatever policies are best for the device in question? And it's also the driver that can detect if the device memory is being wasted by a process that isn't currently performing the interesting operations, while another process that does them had to fallback its allocations to system memory and thus runs slower. I expect the NUMA balancing can't catch that for device memory (and you also disable it anyway?) So I don't really see how a generic solution would work, without having a full concrete example, and thus it's really hard to say that this approach is the right way to go and should be merged. The only examples I've noticed that don't require any special operations to benefit from placement in the "device memory", were fast memories like MCDRAM, which differentiate by performance of generic CPU operations, so it's not really a "device memory" by your terminology. And I would expect policing access to such performance differentiated memory is already possible with e.g. cpusets? Thanks, Vlastimil [1] https://lkml.kernel.org/r/20161025153256.GB6131@gmail.com > https://lkml.org/lkml/2017/1/29/198 > > Changes in V2: > > * Removed redundant nodemask_has_cdm() check from zonelist iterator > * Dropped the nodemask_had_cdm() function itself > * Added node_set/clear_state_cdm() functions and removed bunch of #ifdefs > * Moved CDM helper functions into nodemask.h from node.h header file > * Fixed the build failure by additional CONFIG_NEED_MULTIPLE_NODES check > > Previous V1: (https://lkml.org/lkml/2017/2/8/329) > > Anshuman Khandual (3): > mm: Define coherent device memory (CDM) node > mm: Enable HugeTLB allocation isolation for CDM nodes > mm: Enable Buddy allocation isolation for CDM nodes > > Documentation/ABI/stable/sysfs-devices-node | 7 ++++ > arch/powerpc/Kconfig | 1 + > arch/powerpc/mm/numa.c | 7 ++++ > drivers/base/node.c | 6 +++ > include/linux/nodemask.h | 58 ++++++++++++++++++++++++++++- > mm/Kconfig | 4 ++ > mm/hugetlb.c | 25 ++++++++----- > mm/memory_hotplug.c | 3 ++ > mm/page_alloc.c | 24 +++++++++++- > 9 files changed, 123 insertions(+), 12 deletions(-) >
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web