Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1569360 > unrolled thread
| Started by | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| First post | 2017-01-30 04:40 +0100 |
| Last post | 2017-01-31 07:20 +0100 |
| Articles | 15 on this page of 35 — 6 participants |
Back to article view | Back to linux.kernel
[RFC V2 00/12] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:40 +0100
[RFC V2 05/12] cpuset: Add cpuset_inc() inside cpuset_init() Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:40 +0100
Re: [RFC V2 05/12] cpuset: Add cpuset_inc() inside cpuset_init() Dave Hansen <dave.hansen@intel.com> - 2017-01-30 19:20 +0100
Re: [RFC V2 05/12] cpuset: Add cpuset_inc() inside cpuset_init() Mel Gorman <mgorman@suse.de> - 2017-01-30 21:50 +0100
[RFC] cpuset: Enable changing of top_cpuset's mems_allowed nodemask Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-31 15:30 +0100
Re: [RFC] cpuset: Enable changing of top_cpuset's mems_allowed nodemask Mel Gorman <mgorman@suse.de> - 2017-01-31 17:10 +0100
Re: [RFC] cpuset: Enable changing of top_cpuset's mems_allowed nodemask Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-02-01 08:40 +0100
Re: [RFC] cpuset: Enable changing of top_cpuset's mems_allowed nodemask Michal Hocko <mhocko@kernel.org> - 2017-02-01 10:00 +0100
Re: [RFC] cpuset: Enable changing of top_cpuset's mems_allowed nodemask Mel Gorman <mgorman@suse.de> - 2017-02-01 10:20 +0100
Re: [RFC V2 05/12] cpuset: Add cpuset_inc() inside cpuset_init() Vlastimil Babka <vbabka@suse.cz> - 2017-01-31 15:40 +0100
Re: [RFC V2 05/12] cpuset: Add cpuset_inc() inside cpuset_init() Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-31 16:40 +0100
[RFC V2 04/12] mm: Change mbind(MPOL_BIND) implementation for CDM nodes Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:40 +0100
[DEBUG 16/21] mm: Enable CONFIG_MOVABLE_NODE on powerpc Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:50 +0100
[DEBUG 18/21] mm: Add debugfs interface to dump each node's zonelist information Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:50 +0100
[DEBUG 19/21] mm: Add migrate_virtual_range migration interface Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:50 +0100
[DEBUG 17/21] mm: Export definition of 'zone_names' array through mmzone.h Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:50 +0100
[DEBUG 15/21] powerpc/mm: Enable CONFIG_MOVABLE_NODE for PPC64 platform Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:50 +0100
[DEBUG 21/21] selftests/powerpc: Add a script to perform random VMA migrations Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:50 +0100
[DEBUG 14/21] powerpc/mm: Create numa nodes for hotplug memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:50 +0100
[DEBUG 13/21] powerpc/mm: Identify coherent device memory nodes during platform init Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:50 +0100
[DEBUG 20/21] drivers: Add two drivers for coherent device memory tests Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 04:50 +0100
[RFC V2 09/12] mm: Exclude CDM marked VMAs from auto NUMA Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 07:10 +0100
[RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 07:20 +0100
Re: [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) Dave Hansen <dave.hansen@intel.com> - 2017-01-30 19:00 +0100
Re: [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-31 05:40 +0100
Re: [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) Dave Hansen <dave.hansen@intel.com> - 2017-02-07 19:10 +0100
Re: [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-02-08 17:50 +0100
Re: [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) Jerome Glisse <jglisse@redhat.com> - 2017-02-08 19:40 +0100
[RFC V2 10/12] mm: Ignore madvise(MADV_MERGEABLE) request for VM_CDM marked VMAs Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 10:30 +0100
[RFC V2 11/12] mm: Tag VMA with VM_CDM flag during page fault Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-30 10:40 +0100
Re: [RFC V2 11/12] mm: Tag VMA with VM_CDM flag during page fault Dave Hansen <dave.hansen@intel.com> - 2017-01-30 19:00 +0100
Re: [RFC V2 11/12] mm: Tag VMA with VM_CDM flag during page fault Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-31 06:20 +0100
Re: [RFC V2 11/12] mm: Tag VMA with VM_CDM flag during page fault Dave Hansen <dave.hansen@intel.com> - 2017-01-31 19:10 +0100
Re: [RFC V2 00/12] Define coherent device memory node Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2017-01-31 06:50 +0100
Re: [RFC V2 00/12] Define coherent device memory node Jerome Glisse <jglisse@redhat.com> - 2017-01-31 07:20 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-01-30 04:50 +0100 |
| Subject | [DEBUG 20/21] drivers: Add two drivers for coherent device memory tests |
| Message-ID | <t5b8C-7pH-19@gated-at.bofh.it> |
| In reply to | #1569360 |
This adds two different drivers inside drivers/char/ directory under two
new kernel config options COHERENT_HOTPLUG_DEMO and COHERENT_MEMORY_DEMO.
1) coherent_hotplug_demo: Detects, hoptlugs the coherent device memory
2) coherent_memory_demo: Exports debugfs interface for VMA migrations
Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
drivers/char/Kconfig | 23 +++
drivers/char/Makefile | 2 +
drivers/char/coherent_hotplug_demo.c | 133 ++++++++++++++
drivers/char/coherent_memory_demo.c | 337 +++++++++++++++++++++++++++++++++++
drivers/char/memory_online_sysfs.h | 148 +++++++++++++++
mm/mempolicy.c | 9 +-
6 files changed, 651 insertions(+), 1 deletion(-)
create mode 100644 drivers/char/coherent_hotplug_demo.c
create mode 100644 drivers/char/coherent_memory_demo.c
create mode 100644 drivers/char/memory_online_sysfs.h
diff --git a/drivers/char/Kconfig b/drivers/char/Kconfig
index fde005e..0a9fb82 100644
--- a/drivers/char/Kconfig
+++ b/drivers/char/Kconfig
@@ -588,6 +588,29 @@ config TILE_SROM
device appear much like a simple EEPROM, and knows
how to partition a single ROM for multiple purposes.
+config COHERENT_HOTPLUG_DEMO
+ tristate "Demo driver to test coherent memory node hotplug"
+ depends on PPC64 || COHERENT_DEVICE
+ default n
+ help
+ Say yes when you want to build a test driver to hotplug all
+ the coherent memory nodes present on the system. This driver
+ scans through the device tree, checks on "ibm,memory-device"
+ property device nodes and onlines its memory. When unloaded,
+ it goes through the list of memory ranges it onlined before
+ and oflines them one by one. If not sure, select N.
+
+config COHERENT_MEMORY_DEMO
+ tristate "Demo driver to test coherent memory node functionality"
+ depends on PPC64 || COHERENT_DEVICE
+ default n
+ help
+ Say yes when you want to build a test driver to demonstrate
+ the coherent memory functionalities, capabilities and probable
+ utilizaton. It also exports a debugfs file to accept inputs for
+ virtual address range migration for any process. If not sure,
+ select N.
+
source "drivers/char/xillybus/Kconfig"
endmenu
diff --git a/drivers/char/Makefile b/drivers/char/Makefile
index 6e6c244..92fa338 100644
--- a/drivers/char/Makefile
+++ b/drivers/char/Makefile
@@ -60,3 +60,5 @@ js-rtc-y = rtc.o
obj-$(CONFIG_TILE_SROM) += tile-srom.o
obj-$(CONFIG_XILLYBUS) += xillybus/
obj-$(CONFIG_POWERNV_OP_PANEL) += powernv-op-panel.o
+obj-$(CONFIG_COHERENT_HOTPLUG_DEMO) += coherent_hotplug_demo.o
+obj-$(CONFIG_COHERENT_MEMORY_DEMO) += coherent_memory_demo.o
diff --git a/drivers/char/coherent_hotplug_demo.c b/drivers/char/coherent_hotplug_demo.c
new file mode 100644
index 0000000..bfc1254
--- /dev/null
+++ b/drivers/char/coherent_hotplug_demo.c
@@ -0,0 +1,133 @@
+/*
+ * Memory hotplug support for coherent memory nodes in runtime.
+ *
+ * Copyright (C) 2016, Reza Arbab, IBM Corporation.
+ * Copyright (C) 2016, Anshuman Khandual, IBM Corporation.
+ *
+ * This program is free software; you can redistribute it and/or modify
+ * it under the terms of the GNU General Public License as published by
+ * the Free Software Foundation; either version 2 of the License, or
+ * (at your option) any later version.
+ *
+ * This program is distributed in the hope that it will be useful,
+ * but WITHOUT ANY WARRANTY; without even the implied warranty of
+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
+ * GNU General Public License for more details.
+ */
+#include <linux/of.h>
+#include <linux/export.h>
+#include <linux/spinlock.h>
+#include <linux/init.h>
+#include <linux/memblock.h>
+#include <linux/module.h>
+#include <linux/memory.h>
+#include <linux/sizes.h>
+#include <linux/bitops.h>
+#include <linux/device.h>
+#include <linux/fs.h>
+#include <linux/slab.h>
+#include <linux/mm.h>
+#include <linux/pagemap.h>
+#include <linux/migrate.h>
+#include <linux/memblock.h>
+#include <linux/uaccess.h>
+
+#include <asm/mmu.h>
+#include <asm/pgalloc.h>
+#include "memory_online_sysfs.h"
+
+#define MAX_HOTADD_NODES 100
+phys_addr_t addr[MAX_HOTADD_NODES][2];
+int nr_addr;
+
+/*
+ * extern int memory_failure(unsigned long pfn, int trapno, int flags);
+ * extern int min_free_kbytes;
+ * extern int user_min_free_kbytes;
+ *
+ * extern unsigned long nr_kernel_pages;
+ * extern unsigned long nr_all_pages;
+ * extern unsigned long dma_reserve;
+ */
+
+static void dump_core_vm_tunables(void)
+{
+/*
+ * printk(":::::::: VM TUNABLES :::::::\n");
+ * printk("[min_free_kbytes] %d\n", min_free_kbytes);
+ * printk("[user_min_free_kbytes] %d\n", user_min_free_kbytes);
+ * printk("[nr_kernel_pages] %ld\n", nr_kernel_pages);
+ * printk("[nr_all_pages] %ld\n", nr_all_pages);
+ * printk("[dma_reserve] %ld\n", dma_reserve);
+ */
+}
+
+
+
+static int online_coherent_memory(void)
+{
+ struct device_node *memory;
+
+ nr_addr = 0;
+ disable_auto_online();
+ dump_core_vm_tunables();
+ for_each_compatible_node(memory, NULL, "ibm,memory-device") {
+ struct device_node *mem;
+ const __be64 *reg;
+ unsigned int len, ret;
+ phys_addr_t start, size;
+
+ mem = of_parse_phandle(memory, "memory-region", 0);
+ if (!mem) {
+ pr_info("memory-region property not found\n");
+ return -1;
+ }
+
+ reg = of_get_property(mem, "reg", &len);
+ if (!reg || len <= 0) {
+ pr_info("memory-region property not found\n");
+ return -1;
+ }
+ start = be64_to_cpu(*reg);
+ size = be64_to_cpu(*(reg + 1));
+ pr_info("Coherent memory start %llx size %llx\n", start, size);
+ ret = memory_probe_store(start, size);
+ if (ret)
+ pr_info("probe failed\n");
+
+ ret = store_mem_state(start, size, "online_movable");
+ if (ret)
+ pr_info("online_movable failed\n");
+
+ addr[nr_addr][0] = start;
+ addr[nr_addr][1] = size;
+ nr_addr++;
+ }
+ dump_core_vm_tunables();
+ enable_auto_online();
+ return 0;
+}
+
+static int offline_coherent_memory(void)
+{
+ int i;
+
+ for (i = 0; i < nr_addr; i++)
+ store_mem_state(addr[i][0], addr[i][1], "offline");
+ return 0;
+}
+
+static void __exit coherent_hotplug_exit(void)
+{
+ pr_info("%s\n", __func__);
+ offline_coherent_memory();
+}
+
+static int __init coherent_hotplug_init(void)
+{
+ pr_info("%s\n", __func__);
+ return online_coherent_memory();
+}
+module_init(coherent_hotplug_init);
+module_exit(coherent_hotplug_exit);
+MODULE_LICENSE("GPL");
diff --git a/drivers/char/coherent_memory_demo.c b/drivers/char/coherent_memory_demo.c
new file mode 100644
index 0000000..e711165
--- /dev/null
+++ b/drivers/char/coherent_memory_demo.c
@@ -0,0 +1,337 @@
+/*
+ * Demonstrating various aspects of the coherent memory.
+ *
+ * Copyright (C) 2016, Anshuman Khandual, IBM Corporation.
+ *
+ * This program is free software; you can redistribute it and/or modify
+ * it under the terms of the GNU General Public License as published by
+ * the Free Software Foundation; either version 2 of the License, or
+ * (at your option) any later version.
+ *
+ * This program is distributed in the hope that it will be useful,
+ * but WITHOUT ANY WARRANTY; without even the implied warranty of
+ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
+ * GNU General Public License for more details.
+ */
+#include <linux/of.h>
+#include <linux/export.h>
+#include <linux/spinlock.h>
+#include <linux/init.h>
+#include <linux/memblock.h>
+#include <linux/module.h>
+#include <linux/memory.h>
+#include <linux/sizes.h>
+#include <linux/bitops.h>
+#include <linux/device.h>
+#include <linux/fs.h>
+#include <linux/slab.h>
+#include <linux/mm.h>
+#include <linux/pagemap.h>
+#include <linux/migrate.h>
+#include <linux/memblock.h>
+#include <linux/debugfs.h>
+#include <linux/uaccess.h>
+
+#include <asm/mmu.h>
+#include <asm/pgalloc.h>
+
+#define COHERENT_DEV_MAJOR 89
+#define COHERENT_DEV_NAME "coherent_memory"
+
+#define CRNT_NODE_NID1 1
+#define CRNT_NODE_NID2 2
+#define CRNT_NODE_NID3 3
+
+#define RAM_CRNT_MIGRATE 1
+#define CRNT_RAM_MIGRATE 2
+
+struct vma_map_info {
+ struct list_head list;
+ unsigned long nr_pages;
+ spinlock_t lock;
+};
+
+static void vma_map_info_init(struct vm_area_struct *vma)
+{
+ struct vma_map_info *info = kmalloc(sizeof(struct vma_map_info),
+ GFP_KERNEL);
+
+ WARN_ON(!info);
+ INIT_LIST_HEAD(&info->list);
+ spin_lock_init(&info->lock);
+ vma->vm_private_data = info;
+ info->nr_pages = 0;
+}
+
+static void coherent_vmops_open(struct vm_area_struct *vma)
+{
+ vma_map_info_init(vma);
+}
+
+static void coherent_vmops_close(struct vm_area_struct *vma)
+{
+ struct vma_map_info *info = vma->vm_private_data;
+
+ WARN_ON(!info);
+again:
+ cond_resched();
+ spin_lock(&info->lock);
+ while (info->nr_pages) {
+ struct page *page, *page2;
+
+ list_for_each_entry_safe(page, page2, &info->list, lru) {
+ if (!trylock_page(page)) {
+ spin_unlock(&info->lock);
+ goto again;
+ }
+
+ list_del_init(&page->lru);
+ info->nr_pages--;
+ unlock_page(page);
+ SetPageReclaim(page);
+ put_page(page);
+ }
+ spin_unlock(&info->lock);
+ cond_resched();
+ spin_lock(&info->lock);
+ }
+ spin_unlock(&info->lock);
+ kfree(info);
+ vma->vm_private_data = NULL;
+}
+
+static int coherent_vmops_fault(struct vm_area_struct *vma,
+ struct vm_fault *vmf)
+{
+ struct vma_map_info *info;
+ struct page *page;
+ static int coherent_node = CRNT_NODE_NID1;
+
+ if (coherent_node == CRNT_NODE_NID1)
+ coherent_node = CRNT_NODE_NID2;
+ else
+ coherent_node = CRNT_NODE_NID1;
+
+ page = alloc_pages_node(coherent_node,
+ GFP_HIGHUSER_MOVABLE | __GFP_THISNODE, 0);
+ if (!page)
+ return VM_FAULT_SIGBUS;
+
+ info = (struct vma_map_info *) vma->vm_private_data;
+ WARN_ON(!info);
+ spin_lock(&info->lock);
+ list_add(&page->lru, &info->list);
+ info->nr_pages++;
+ spin_unlock(&info->lock);
+
+ page->index = vmf->pgoff;
+ get_page(page);
+ vmf->page = page;
+ return 0;
+}
+
+static const struct vm_operations_struct coherent_memory_vmops = {
+ .open = coherent_vmops_open,
+ .close = coherent_vmops_close,
+ .fault = coherent_vmops_fault,
+};
+
+static int coherent_memory_mmap(struct file *file, struct vm_area_struct *vma)
+{
+ pr_info("Mmap opened (file: %lx vma: %lx)\n",
+ (unsigned long) file, (unsigned long) vma);
+ vma->vm_ops = &coherent_memory_vmops;
+ coherent_vmops_open(vma);
+ return 0;
+}
+
+static int coherent_memory_open(struct inode *inode, struct file *file)
+{
+ pr_info("Device opened (inode: %lx file: %lx)\n",
+ (unsigned long) inode, (unsigned long) file);
+ return 0;
+}
+
+static int coherent_memory_close(struct inode *inode, struct file *file)
+{
+ pr_info("Device closed (inode: %lx file: %lx)\n",
+ (unsigned long) inode, (unsigned long) file);
+ return 0;
+}
+
+static void lru_ram_coherent_migrate(unsigned long addr)
+{
+ struct mm_struct *mm = current->mm;
+ struct vm_area_struct *vma;
+ nodemask_t nmask;
+ LIST_HEAD(mlist);
+
+ nodes_clear(nmask);
+ nodes_setall(nmask);
+ down_write(&mm->mmap_sem);
+ for (vma = mm->mmap; vma; vma = vma->vm_next) {
+ if ((addr < vma->vm_start) || (addr > vma->vm_end))
+ continue;
+ break;
+ }
+ up_write(&mm->mmap_sem);
+ if (!vma) {
+ pr_info("%s: No VMA found\n", __func__);
+ return;
+ }
+ migrate_virtual_range(current->pid, vma->vm_start, vma->vm_end, 2);
+}
+
+static void lru_coherent_ram_migrate(unsigned long addr)
+{
+ struct mm_struct *mm = current->mm;
+ struct vm_area_struct *vma;
+ nodemask_t nmask;
+ LIST_HEAD(mlist);
+
+ nodes_clear(nmask);
+ nodes_setall(nmask);
+ down_write(&mm->mmap_sem);
+ for (vma = mm->mmap; vma; vma = vma->vm_next) {
+ if ((addr < vma->vm_start) || (addr > vma->vm_end))
+ continue;
+ break;
+ }
+ up_write(&mm->mmap_sem);
+ if (!vma) {
+ pr_info("%s: No VMA found\n", __func__);
+ return;
+ }
+ migrate_virtual_range(current->pid, vma->vm_start, vma->vm_end, 0);
+}
+
+static long coherent_memory_ioctl(struct file *file,
+ unsigned int cmd, unsigned long arg)
+{
+ switch (cmd) {
+ case RAM_CRNT_MIGRATE:
+ lru_ram_coherent_migrate(arg);
+ break;
+
+ case CRNT_RAM_MIGRATE:
+ lru_coherent_ram_migrate(arg);
+ break;
+
+ default:
+ pr_info("%s Invalid ioctl() command: %d\n", __func__, cmd);
+ return -EINVAL;
+ }
+ return 0;
+}
+
+static const struct file_operations fops = {
+ .mmap = coherent_memory_mmap,
+ .open = coherent_memory_open,
+ .release = coherent_memory_close,
+ .unlocked_ioctl = &coherent_memory_ioctl
+};
+
+static char kbuf[100]; /* Will store original user passed buffer */
+static char str[100]; /* Working copy for individual substring */
+
+static u64 args[4];
+static u64 index;
+static void convert_substring(const char *buf)
+{
+ u64 val = 0;
+
+ if (kstrtou64(buf, 0, &val))
+ pr_info("String conversion failed\n");
+
+ args[index] = val;
+ index++;
+}
+
+static ssize_t coherent_debug_write(struct file *file,
+ const char __user *user_buf,
+ size_t count, loff_t *ppos)
+{
+ char *tmp, *tmp1;
+ size_t ret;
+
+ memset(args, 0, sizeof(args));
+ index = 0;
+
+ ret = simple_write_to_buffer(kbuf, sizeof(kbuf), ppos, user_buf, count);
+ if (ret < 0)
+ return ret;
+
+ kbuf[ret] = '\0';
+ tmp = kbuf;
+ do {
+ tmp1 = strchr(tmp, ',');
+ if (tmp1) {
+ *tmp1 = '\0';
+ strncpy(str, (const char *)tmp, strlen(tmp));
+ convert_substring(str);
+ } else {
+ strncpy(str, (const char *)tmp, strlen(tmp));
+ convert_substring(str);
+ break;
+ }
+ tmp = tmp1 + 1;
+ memset(str, 0, sizeof(str));
+ } while (true);
+ migrate_virtual_range(args[0], args[1], args[2], args[3]);
+ return ret;
+}
+
+static int coherent_debug_show(struct seq_file *m, void *v)
+{
+ seq_puts(m, "Expected Value: <pid,vaddr,size,nid>\n");
+ return 0;
+}
+
+static int coherent_debug_open(struct inode *inode, struct file *filp)
+{
+ return single_open(filp, coherent_debug_show, NULL);
+}
+
+static const struct file_operations coherent_debug_fops = {
+ .open = coherent_debug_open,
+ .write = coherent_debug_write,
+ .read = seq_read,
+ .llseek = seq_lseek,
+ .release = single_release,
+};
+
+static struct dentry *debugfile;
+
+static void coherent_memory_debugfs(void)
+{
+
+ debugfile = debugfs_create_file("coherent_debug", 0644, NULL, NULL,
+ &coherent_debug_fops);
+ if (!debugfile)
+ pr_warn("Failed to create coherent_memory in debugfs");
+}
+
+static void __exit coherent_memory_exit(void)
+{
+ pr_info("%s\n", __func__);
+ debugfs_remove(debugfile);
+ unregister_chrdev(COHERENT_DEV_MAJOR, COHERENT_DEV_NAME);
+}
+
+static int __init coherent_memory_init(void)
+{
+ int ret;
+
+ pr_info("%s\n", __func__);
+ ret = register_chrdev(COHERENT_DEV_MAJOR, COHERENT_DEV_NAME, &fops);
+ if (ret < 0) {
+ pr_info("%s register_chrdev() failed\n", __func__);
+ return -1;
+ }
+ coherent_memory_debugfs();
+ return 0;
+}
+
+module_init(coherent_memory_init);
+module_exit(coherent_memory_exit);
+MODULE_LICENSE("GPL");
diff --git a/drivers/char/memory_online_sysfs.h b/drivers/char/memory_online_sysfs.h
new file mode 100644
index 0000000..a5f022d
--- /dev/null
+++ b/drivers/char/memory_online_sysfs.h
@@ -0,0 +1,148 @@
+/*
+ * Accessing sysfs interface for memory hotplug operation from
+ * inside the kernel.
+ *
+ * Licensed under GPL V2
+ */
+#ifndef __SYSFS_H
+#define __SYSFS_H
+
+#include <linux/fs.h>
+#include <linux/uaccess.h>
+
+#define AUTO_ONLINE_BLOCKS "/sys/devices/system/memory/auto_online_blocks"
+#define BLOCK_SIZE_BYTES "/sys/devices/system/memory/block_size_bytes"
+#define MEMORY_PROBE "/sys/devices/system/memory/probe"
+
+static ssize_t read_buf(char *filename, char *buf, ssize_t count)
+{
+ mm_segment_t old_fs;
+ struct file *filp;
+ loff_t pos = 0;
+
+ if (!count)
+ return 0;
+
+ old_fs = get_fs();
+ set_fs(KERNEL_DS);
+
+ filp = filp_open(filename, O_RDONLY, 0);
+ if (IS_ERR(filp)) {
+ count = PTR_ERR(filp);
+ goto err_open;
+ }
+
+ count = vfs_read(filp, buf, count - 1, &pos);
+ buf[count] = '\0';
+
+ filp_close(filp, NULL);
+
+err_open:
+ set_fs(old_fs);
+
+ return count;
+}
+
+static unsigned long long read_0x(char *filename)
+{
+ unsigned long long ret;
+ char buf[32];
+
+ if (read_buf(filename, buf, 32) <= 0)
+ return 0;
+
+ if (kstrtoull(buf, 16, &ret))
+ return 0;
+
+ return ret;
+}
+
+static ssize_t write_buf(char *filename, char *buf)
+{
+ int ret;
+ mm_segment_t old_fs;
+ struct file *filp;
+ loff_t pos = 0;
+
+ old_fs = get_fs();
+ set_fs(KERNEL_DS);
+
+ filp = filp_open(filename, O_WRONLY, 0);
+ if (IS_ERR(filp)) {
+ ret = PTR_ERR(filp);
+ goto err_open;
+ }
+
+ ret = vfs_write(filp, buf, strlen(buf), &pos);
+
+ filp_close(filp, NULL);
+
+err_open:
+ set_fs(old_fs);
+
+ return ret;
+}
+
+int memory_probe_store(phys_addr_t addr, phys_addr_t size)
+{
+ phys_addr_t block_sz =
+ read_0x(BLOCK_SIZE_BYTES);
+ long i;
+
+ for (i = 0; i < size / block_sz; i++, addr += block_sz) {
+ char s[32];
+ ssize_t count;
+
+ snprintf(s, 32, "0x%llx", addr);
+
+ count = write_buf(MEMORY_PROBE, s);
+ if (count < 0)
+ return count;
+ }
+
+ return 0;
+}
+
+int store_mem_state(phys_addr_t addr, phys_addr_t size, char *state)
+{
+ phys_addr_t block_sz = read_0x(BLOCK_SIZE_BYTES);
+ unsigned long start_block, end_block, i;
+
+ start_block = addr / block_sz;
+ end_block = start_block + size / block_sz;
+
+ for (i = end_block - 1; i >= start_block; i--) {
+ char filename[64];
+ ssize_t count;
+
+ snprintf(filename, 64,
+ "/sys/devices/system/memory/memory%ld/state", i);
+
+ count = write_buf(filename, state);
+ if (count < 0)
+ return count;
+ }
+
+ return 0;
+}
+
+int disable_auto_online(void)
+{
+ int ret;
+
+ ret = write_buf(AUTO_ONLINE_BLOCKS, "offline");
+ if (ret)
+ return ret;
+ return 0;
+}
+
+int enable_auto_online(void)
+{
+ int ret;
+
+ ret = write_buf(AUTO_ONLINE_BLOCKS, "online");
+ if (ret)
+ return ret;
+ return 0;
+}
+#endif
diff --git a/mm/mempolicy.c b/mm/mempolicy.c
index 13cd5eb..f65810a 100644
--- a/mm/mempolicy.c
+++ b/mm/mempolicy.c
@@ -2946,6 +2946,7 @@ int migrate_virtual_range(int pid, unsigned long start,
goto out;
}
+ pr_info("%s: %d %lx %lx %d: ", __func__, pid, start, end, nid);
rcu_read_lock();
mm = find_task_by_vpid(pid)->mm;
rcu_read_unlock();
@@ -2956,8 +2957,14 @@ int migrate_virtual_range(int pid, unsigned long start,
if (!list_empty(&mlist)) {
ret = migrate_pages(&mlist, new_node_page, NULL,
nid, MIGRATE_SYNC, MR_NUMA_MISPLACED);
- if (ret)
+ if (ret) {
+ pr_info("migration_failed for %d pages\n", ret);
putback_movable_pages(&mlist);
+ } else {
+ pr_info("migration_passed\n");
+ }
+ } else {
+ pr_info("list_empty\n");
}
up_write(&mm->mmap_sem);
out:
--
2.9.3
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-01-30 07:10 +0100 |
| Subject | [RFC V2 09/12] mm: Exclude CDM marked VMAs from auto NUMA |
| Message-ID | <t5dk5-se-7@gated-at.bofh.it> |
| In reply to | #1569360 |
Kernel cannot track device memory accesses behind VMAs containing CDM
memory. Hence all the VM_CDM marked VMAs should not be part of the auto
NUMA migration scheme. This patch also adds a new function is_cdm_vma()
to detect any VMA marked with flag VM_CDM.
Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
include/linux/mempolicy.h | 14 ++++++++++++++
kernel/sched/fair.c | 3 ++-
2 files changed, 16 insertions(+), 1 deletion(-)
diff --git a/include/linux/mempolicy.h b/include/linux/mempolicy.h
index 5f4d828..ff0c6bc 100644
--- a/include/linux/mempolicy.h
+++ b/include/linux/mempolicy.h
@@ -172,6 +172,20 @@ extern int mpol_parse_str(char *str, struct mempolicy **mpol);
extern void mpol_to_str(char *buffer, int maxlen, struct mempolicy *pol);
+#ifdef CONFIG_COHERENT_DEVICE
+static inline bool is_cdm_vma(struct vm_area_struct *vma)
+{
+ if (vma->vm_flags & VM_CDM)
+ return true;
+ return false;
+}
+#else
+static inline bool is_cdm_vma(struct vm_area_struct *vma)
+{
+ return false;
+}
+#endif
+
/* Check if a vma is migratable */
static inline bool vma_migratable(struct vm_area_struct *vma)
{
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 6559d19..523508c 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -2482,7 +2482,8 @@ void task_numa_work(struct callback_head *work)
}
for (; vma; vma = vma->vm_next) {
if (!vma_migratable(vma) || !vma_policy_mof(vma) ||
- is_vm_hugetlb_page(vma) || (vma->vm_flags & VM_MIXEDMAP)) {
+ is_vm_hugetlb_page(vma) || is_cdm_vma(vma) ||
+ (vma->vm_flags & VM_MIXEDMAP)) {
continue;
}
--
2.9.3
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-01-30 07:20 +0100 |
| Subject | [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) |
| Message-ID | <t5dtL-vC-7@gated-at.bofh.it> |
| In reply to | #1569360 |
Mark all the applicable VMAs with VM_CDM explicitly during mbind(MPOL_BIND)
call if the user provided nodemask has a CDM node.
Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
mm/mempolicy.c | 17 +++++++++++++++++
1 file changed, 17 insertions(+)
diff --git a/mm/mempolicy.c b/mm/mempolicy.c
index 78e095b..4482140 100644
--- a/mm/mempolicy.c
+++ b/mm/mempolicy.c
@@ -175,6 +175,16 @@ static void mpol_relative_nodemask(nodemask_t *ret, const nodemask_t *orig,
}
#ifdef CONFIG_COHERENT_DEVICE
+static inline void set_vm_cdm(struct vm_area_struct *vma)
+{
+ vma->vm_flags |= VM_CDM;
+}
+
+static inline void clr_vm_cdm(struct vm_area_struct *vma)
+{
+ vma->vm_flags &= ~VM_CDM;
+}
+
static void mark_vma_cdm(nodemask_t *nmask,
struct page *page, struct vm_area_struct *vma)
{
@@ -191,6 +201,9 @@ static void mark_vma_cdm(nodemask_t *nmask,
vma->vm_flags |= VM_CDM;
}
#else
+static inline void set_vm_cdm(struct vm_area_struct *vma) { }
+static inline void clr_vm_cdm(struct vm_area_struct *vma) { }
+
static void mark_vma_cdm(nodemask_t *nmask,
struct page *page, struct vm_area_struct *vma)
{
@@ -770,6 +783,10 @@ static int mbind_range(struct mm_struct *mm, unsigned long start,
vmstart = max(start, vma->vm_start);
vmend = min(end, vma->vm_end);
+ if ((new_pol->mode == MPOL_BIND)
+ && nodemask_has_cdm(new_pol->v.nodes))
+ set_vm_cdm(vma);
+
if (mpol_equal(vma_policy(vma), new_pol))
continue;
--
2.9.3
[toc] | [prev] | [next] | [standalone]
| From | Dave Hansen <dave.hansen@intel.com> |
|---|---|
| Date | 2017-01-30 19:00 +0100 |
| Subject | Re: [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) |
| Message-ID | <t5opd-6Tg-39@gated-at.bofh.it> |
| In reply to | #1569420 |
On 01/29/2017 07:35 PM, Anshuman Khandual wrote: > + if ((new_pol->mode == MPOL_BIND) > + && nodemask_has_cdm(new_pol->v.nodes)) > + set_vm_cdm(vma); So, if you did: mbind(addr, PAGE_SIZE, MPOL_BIND, all_nodes, ...); mbind(addr, PAGE_SIZE, MPOL_BIND, one_non_cdm_node, ...); You end up with a VMA that can never have KSM done on it, etc... Even though there's no good reason for it. I guess /proc/$pid/smaps might be able to help us figure out what was going on here, but that still seems like an awful lot of damage.
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-01-31 05:40 +0100 |
| Subject | Re: [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) |
| Message-ID | <t5yox-4C8-3@gated-at.bofh.it> |
| In reply to | #1569966 |
On 01/30/2017 11:24 PM, Dave Hansen wrote: > On 01/29/2017 07:35 PM, Anshuman Khandual wrote: >> + if ((new_pol->mode == MPOL_BIND) >> + && nodemask_has_cdm(new_pol->v.nodes)) >> + set_vm_cdm(vma); > So, if you did: > > mbind(addr, PAGE_SIZE, MPOL_BIND, all_nodes, ...); > mbind(addr, PAGE_SIZE, MPOL_BIND, one_non_cdm_node, ...); > > You end up with a VMA that can never have KSM done on it, etc... Even > though there's no good reason for it. I guess /proc/$pid/smaps might be > able to help us figure out what was going on here, but that still seems > like an awful lot of damage. Agreed, this VMA should not remain tagged after the second call. It does not make sense. For this kind of scenarios we can re-evaluate the VMA tag every time the nodemask change is attempted. But if we are looking for some runtime re-evaluation then we need to steal some cycles are during general VMA processing opportunity points like merging and split to do the necessary re-evaluation. Should do we do these kind two kinds of re-evaluation to be more optimal ?
[toc] | [prev] | [next] | [standalone]
| From | Dave Hansen <dave.hansen@intel.com> |
|---|---|
| Date | 2017-02-07 19:10 +0100 |
| Subject | Re: [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) |
| Message-ID | <t8ing-6go-11@gated-at.bofh.it> |
| In reply to | #1570313 |
On 01/30/2017 08:36 PM, Anshuman Khandual wrote: > On 01/30/2017 11:24 PM, Dave Hansen wrote: >> On 01/29/2017 07:35 PM, Anshuman Khandual wrote: >>> + if ((new_pol->mode == MPOL_BIND) >>> + && nodemask_has_cdm(new_pol->v.nodes)) >>> + set_vm_cdm(vma); >> So, if you did: >> >> mbind(addr, PAGE_SIZE, MPOL_BIND, all_nodes, ...); >> mbind(addr, PAGE_SIZE, MPOL_BIND, one_non_cdm_node, ...); >> >> You end up with a VMA that can never have KSM done on it, etc... Even >> though there's no good reason for it. I guess /proc/$pid/smaps might be >> able to help us figure out what was going on here, but that still seems >> like an awful lot of damage. > > Agreed, this VMA should not remain tagged after the second call. It does > not make sense. For this kind of scenarios we can re-evaluate the VMA > tag every time the nodemask change is attempted. But if we are looking for > some runtime re-evaluation then we need to steal some cycles are during > general VMA processing opportunity points like merging and split to do > the necessary re-evaluation. Should do we do these kind two kinds of > re-evaluation to be more optimal ? I'm still unconvinced that you *need* detection like this. Scanning big VMAs is going to be really painful. I thought I asked before but I can't find it in this thread. But, we have explicit interfaces for disabling KSM and khugepaged. Why do we need implicit ones like this in addition to those?
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-02-08 17:50 +0100 |
| Subject | Re: [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) |
| Message-ID | <t8DBn-2Hv-3@gated-at.bofh.it> |
| In reply to | #1575932 |
On 02/07/2017 11:37 PM, Dave Hansen wrote:
>> On 01/30/2017 11:24 PM, Dave Hansen wrote:
>>> On 01/29/2017 07:35 PM, Anshuman Khandual wrote:
>>>> + if ((new_pol->mode == MPOL_BIND)
>>>> + && nodemask_has_cdm(new_pol->v.nodes))
>>>> + set_vm_cdm(vma);
>>> So, if you did:
>>>
>>> mbind(addr, PAGE_SIZE, MPOL_BIND, all_nodes, ...);
>>> mbind(addr, PAGE_SIZE, MPOL_BIND, one_non_cdm_node, ...);
>>>
>>> You end up with a VMA that can never have KSM done on it, etc... Even
>>> though there's no good reason for it. I guess /proc/$pid/smaps might be
>>> able to help us figure out what was going on here, but that still seems
>>> like an awful lot of damage.
>> Agreed, this VMA should not remain tagged after the second call. It does
>> not make sense. For this kind of scenarios we can re-evaluate the VMA
>> tag every time the nodemask change is attempted. But if we are looking for
>> some runtime re-evaluation then we need to steal some cycles are during
>> general VMA processing opportunity points like merging and split to do
>> the necessary re-evaluation. Should do we do these kind two kinds of
>> re-evaluation to be more optimal ?
> I'm still unconvinced that you *need* detection like this. Scanning big
> VMAs is going to be really painful.
>
> I thought I asked before but I can't find it in this thread. But, we
> have explicit interfaces for disabling KSM and khugepaged. Why do we
> need implicit ones like this in addition to those?
Missed the discussion we had on this last time around I think. My bad, sorry
about that. IIUC we can disable KSM through madvise() call, in fact I guess
its disabled by default and need to be enabled. We can just have a similar
interface to disable auto NUMA for a specific VMA or we can handle it page
by page basis with something like this.
diff --git a/mm/memory.c b/mm/memory.c
index 1099d35..101dfd9 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -3518,6 +3518,9 @@ static int do_numa_page(struct vm_fault *vmf)
goto out;
}
+ if (is_cdm_node(page_to_nid(page)))
+ goto out;
+
/* Migrate to the requested node */
migrated = migrate_misplaced_page(page, vma, target_nid);
if (migrated) {
I am still looking into these aspects. BTW have posted the minimum set of
CDM patches which defines and isolates CDM node.
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-02-08 19:40 +0100 |
| Subject | Re: [RFC V2 12/12] mm: Tag VMA with VM_CDM flag explicitly during mbind(MPOL_BIND) |
| Message-ID | <t8FjP-3Oi-5@gated-at.bofh.it> |
| In reply to | #1575932 |
> On 01/30/2017 08:36 PM, Anshuman Khandual wrote: > > On 01/30/2017 11:24 PM, Dave Hansen wrote: > >> On 01/29/2017 07:35 PM, Anshuman Khandual wrote: > >>> + if ((new_pol->mode == MPOL_BIND) > >>> + && nodemask_has_cdm(new_pol->v.nodes)) > >>> + set_vm_cdm(vma); > >> So, if you did: > >> > >> mbind(addr, PAGE_SIZE, MPOL_BIND, all_nodes, ...); > >> mbind(addr, PAGE_SIZE, MPOL_BIND, one_non_cdm_node, ...); > >> > >> You end up with a VMA that can never have KSM done on it, etc... Even > >> though there's no good reason for it. I guess /proc/$pid/smaps might be > >> able to help us figure out what was going on here, but that still seems > >> like an awful lot of damage. > > > > Agreed, this VMA should not remain tagged after the second call. It does > > not make sense. For this kind of scenarios we can re-evaluate the VMA > > tag every time the nodemask change is attempted. But if we are looking for > > some runtime re-evaluation then we need to steal some cycles are during > > general VMA processing opportunity points like merging and split to do > > the necessary re-evaluation. Should do we do these kind two kinds of > > re-evaluation to be more optimal ? > > I'm still unconvinced that you *need* detection like this. Scanning big > VMAs is going to be really painful. > > I thought I asked before but I can't find it in this thread. But, we > have explicit interfaces for disabling KSM and khugepaged. Why do we > need implicit ones like this in addition to those? > I said it in other part of the thread i think the vma flag is a no go. Because it try to set something that is orthogonal to vma. That you want some vma to use device memory on new allocation is a valid policy for a vma to have. But to have a flag that say various kernel subsystem hey my memory is special skip me is wrong. The fact that you want to exclude device memory from KSM or autonuma is valid but it should be done at struct page level ie KSM or autonuma should check the type of page before doing anything. For CDM pages they would skip. It could be the flags idea that was discussed. The overhead of doing it at page level is far lower than trying to manage a vma flags with all the issue related to vma merging, splitting and lifetime of such flags. Moreover this flags is an all or nothing, it does not consider the case where you have as much regular page as CDM page in a vma. It would block regular page from under going the usual KSM/autonuma ... I do strongly believe that this vma flag is a bad idea. Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-01-30 10:30 +0100 |
| Subject | [RFC V2 10/12] mm: Ignore madvise(MADV_MERGEABLE) request for VM_CDM marked VMAs |
| Message-ID | <t5grD-2eP-7@gated-at.bofh.it> |
| In reply to | #1569360 |
VMA containing CDM memory should be excluded from KSM merging. This change makes madvise(MADV_MERGEABLE) request on target VMA to be ignored. Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com> --- mm/ksm.c | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/mm/ksm.c b/mm/ksm.c index 9ae6011..2fb8939 100644 --- a/mm/ksm.c +++ b/mm/ksm.c @@ -37,6 +37,7 @@ #include <linux/freezer.h> #include <linux/oom.h> #include <linux/numa.h> +#include <linux/mempolicy.h> #include <asm/tlbflush.h> #include "internal.h" @@ -1751,6 +1752,9 @@ int ksm_madvise(struct vm_area_struct *vma, unsigned long start, VM_HUGETLB | VM_MIXEDMAP)) return 0; /* just ignore the advice */ + if (is_cdm_vma(vma)) + return 0; + #ifdef VM_SAO if (*vm_flags & VM_SAO) return 0; -- 2.9.3
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-01-30 10:40 +0100 |
| Subject | [RFC V2 11/12] mm: Tag VMA with VM_CDM flag during page fault |
| Message-ID | <t5gBl-2i3-63@gated-at.bofh.it> |
| In reply to | #1569360 |
Mark the corresponding VMA with VM_CDM flag if the allocated page happens
to be from a CDM node. This can be expensive from performance stand point.
There are multiple checks to avoid an expensive page_to_nid lookup but it
can be optimized further.
Signed-off-by: Anshuman Khandual <khandual@linux.vnet.ibm.com>
---
mm/mempolicy.c | 24 ++++++++++++++++++++++++
1 file changed, 24 insertions(+)
diff --git a/mm/mempolicy.c b/mm/mempolicy.c
index 6089c711..78e095b 100644
--- a/mm/mempolicy.c
+++ b/mm/mempolicy.c
@@ -174,6 +174,29 @@ static void mpol_relative_nodemask(nodemask_t *ret, const nodemask_t *orig,
nodes_onto(*ret, tmp, *rel);
}
+#ifdef CONFIG_COHERENT_DEVICE
+static void mark_vma_cdm(nodemask_t *nmask,
+ struct page *page, struct vm_area_struct *vma)
+{
+ if (!page)
+ return;
+
+ if (vma->vm_flags & VM_CDM)
+ return;
+
+ if (nmask && !nodemask_has_cdm(*nmask))
+ return;
+
+ if (is_cdm_node(page_to_nid(page)))
+ vma->vm_flags |= VM_CDM;
+}
+#else
+static void mark_vma_cdm(nodemask_t *nmask,
+ struct page *page, struct vm_area_struct *vma)
+{
+}
+#endif
+
static int mpol_new_interleave(struct mempolicy *pol, const nodemask_t *nodes)
{
if (nodes_empty(*nodes))
@@ -2039,6 +2062,7 @@ alloc_pages_vma(gfp_t gfp, int order, struct vm_area_struct *vma,
nmask = policy_nodemask(gfp, pol);
zl = policy_zonelist(gfp, pol, node);
page = __alloc_pages_nodemask(gfp, order, zl, nmask);
+ mark_vma_cdm(nmask, page, vma);
mpol_cond_put(pol);
out:
if (unlikely(!page && read_mems_allowed_retry(cpuset_mems_cookie)))
--
2.9.3
[toc] | [prev] | [next] | [standalone]
| From | Dave Hansen <dave.hansen@intel.com> |
|---|---|
| Date | 2017-01-30 19:00 +0100 |
| Subject | Re: [RFC V2 11/12] mm: Tag VMA with VM_CDM flag during page fault |
| Message-ID | <t5opc-6Tg-23@gated-at.bofh.it> |
| In reply to | #1569493 |
Here's the flag definition:
> +#ifdef CONFIG_COHERENT_DEVICE
> +#define VM_CDM 0x00800000 /* Contains coherent device memory */
> +#endif
But it doesn't match the implementation:
> +#ifdef CONFIG_COHERENT_DEVICE
> +static void mark_vma_cdm(nodemask_t *nmask,
> + struct page *page, struct vm_area_struct *vma)
> +{
> + if (!page)
> + return;
> +
> + if (vma->vm_flags & VM_CDM)
> + return;
> +
> + if (nmask && !nodemask_has_cdm(*nmask))
> + return;
> +
> + if (is_cdm_node(page_to_nid(page)))
> + vma->vm_flags |= VM_CDM;
> +}
That flag is a one-way trip. Any VMA with that flag set on it will keep
it for the life of the VMA, despite whether it has CDM pages in it now
or not. Even if you changed the policy back to one that doesn't allow
CDM and forced all the pages to be migrated out.
This also assumes that the only way to get a page mapped into a VMA is
via alloc_pages_vma(). Do the NUMA migration APIs use this path?
When you *set* this flag, you don't go and turn off KSM merging, for
instance. You keep it from being turned on from this point forward, but
you don't turn it off.
This is happening with mmap_sem held for read. Correct? Is it OK that
you're modifying the VMA? That vm_flags manipulation is non-atomic, so
how can that even be safe?
If you're going to go down this route, I think you need to be very
careful. We need to ensure that when this flag gets set, it's never set
on VMAs that are "normal" and will only be set on VMAs that were
*explicitly* set up for accessing CDM. That means that you'll need to
make sure that there's no possible way to get a CDM page faulted into a
VMA unless it's via an explicitly assigned policy that would have cause
the VMA to be split from any "normal" one in the system.
This all makes me really nervous.
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-01-31 06:20 +0100 |
| Subject | Re: [RFC V2 11/12] mm: Tag VMA with VM_CDM flag during page fault |
| Message-ID | <t5z1g-53O-23@gated-at.bofh.it> |
| In reply to | #1569964 |
On 01/30/2017 11:21 PM, Dave Hansen wrote:
> Here's the flag definition:
>
>> +#ifdef CONFIG_COHERENT_DEVICE
>> +#define VM_CDM 0x00800000 /* Contains coherent device memory */
>> +#endif
>
> But it doesn't match the implementation:
>
>> +#ifdef CONFIG_COHERENT_DEVICE
>> +static void mark_vma_cdm(nodemask_t *nmask,
>> + struct page *page, struct vm_area_struct *vma)
>> +{
>> + if (!page)
>> + return;
>> +
>> + if (vma->vm_flags & VM_CDM)
>> + return;
>> +
>> + if (nmask && !nodemask_has_cdm(*nmask))
>> + return;
>> +
>> + if (is_cdm_node(page_to_nid(page)))
>> + vma->vm_flags |= VM_CDM;
>> +}
>
> That flag is a one-way trip. Any VMA with that flag set on it will keep
> it for the life of the VMA, despite whether it has CDM pages in it now
> or not. Even if you changed the policy back to one that doesn't allow
> CDM and forced all the pages to be migrated out.
Right, we have this limitation right now. But as I have mentioned in the
reply on the other thread, will work towards both static and runtime
re-evaluation of the VMA flag next time around.
>
> This also assumes that the only way to get a page mapped into a VMA is
> via alloc_pages_vma(). Do the NUMA migration APIs use this path?
Right now I have just taken care of these two paths.
* Page fault path
* mbind() path
agreed, will work on the NUMA migration APIs paths next. Wondering if
I need to update for migrate_pages() kernel API also as it will be
used by the driver or should the driver tag the VMA explicitly knowing
what has just happened ? I had also mentioned about this in the cover
letter :) But as you have pointed out will move the documentation
to the patches.
"
VM_CDM tagged VMA:
There are two parts to this problem.
* How to mark a VMA with VM_CDM ?
- During page fault path
- During mbind(MPOL_BIND) call
- Any other paths ?
- Should a driver mark a VMA with VM_CDM explicitly ?
* How VM_CDM marked VMA gets treated ?
- Disabled from auto NUMA migrations
- Disabled from KSM merging
- Anything else ?
"
>
> When you *set* this flag, you don't go and turn off KSM merging, for
> instance. You keep it from being turned on from this point forward, but
> you don't turn it off.
I was in the impression that the KSM merging does not start unless we
do madvise(MADV_MERGEABLE) call on the VMA (where its blocked now). I
might be missing something here if it can start before hand.
>
> This is happening with mmap_sem held for read. Correct? Is it OK that
> you're modifying the VMA? That vm_flags manipulation is non-atomic, so
> how can that even be safe?
Hmm. should it be done with mmap_sem being held for write. Will look
into this further. But intercepting the page faults inside alloc_pages_vma()
for tagging the VMA is okay from over all design perspective ?. Or this
should be moved up or down the call chain in the page fault path ?
>
> If you're going to go down this route, I think you need to be very
> careful. We need to ensure that when this flag gets set, it's never set
> on VMAs that are "normal" and will only be set on VMAs that were
> *explicitly* set up for accessing CDM. That means that you'll need to
> make sure that there's no possible way to get a CDM page faulted into a
> VMA unless it's via an explicitly assigned policy that would have cause
> the VMA to be split from any "normal" one in the system.
>
> This all makes me really nervous.
Got it, will work towards this.
[toc] | [prev] | [next] | [standalone]
| From | Dave Hansen <dave.hansen@intel.com> |
|---|---|
| Date | 2017-01-31 19:10 +0100 |
| Subject | Re: [RFC V2 11/12] mm: Tag VMA with VM_CDM flag during page fault |
| Message-ID | <t5L2q-3UK-23@gated-at.bofh.it> |
| In reply to | #1570325 |
On 01/30/2017 09:10 PM, Anshuman Khandual wrote: >> This is happening with mmap_sem held for read. Correct? Is it OK that >> you're modifying the VMA? That vm_flags manipulation is non-atomic, so >> how can that even be safe? > Hmm. should it be done with mmap_sem being held for write. Will look > into this further. But intercepting the page faults inside alloc_pages_vma() > for tagging the VMA is okay from over all design perspective ?. Or this > should be moved up or down the call chain in the page fault path ? Doing it in the fault path seems wrong to me. Apps have to take *explicit* action to go and get access to device memory. It seems like we should mark the VMA *then*, at the time of the explicit action. I also think _implying_ that we want KSM, etc... turned off just because of the target of an mbind() is a bad idea. Apps have to ask for this stuff *explicitly*, so why not also have them turn KSM off explicitly?
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-01-31 06:50 +0100 |
| Message-ID | <t5zui-5gO-21@gated-at.bofh.it> |
| In reply to | #1569360 |
Hello Dave/Jerome/Mel, Here is the overall layout of the functions I am trying to put together through this patch series. (1) Define CDM from core VM and kernel perspective (2) Isolation/Special consideration for HugeTLB allocations (3) Isolation/Special consideration for buddy allocations (a) Zonelist modification based isolation (proposed) (b) Cpuset modification based isolation (proposed) (c) Buddy modification based isolation (working) (4) Define VMA containing CDM memory with a new flag VM_CDM (5) Special consideration for VM_CDM marked VMAs (a) Special consideration for auto NUMA (b) Special consideration for KSM Is there are any other area which needs to be taken care of before CDM node can be represented completely inside the kernel ? Regards Anshuman
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-01-31 07:20 +0100 |
| Message-ID | <t5zXk-5GO-9@gated-at.bofh.it> |
| In reply to | #1570346 |
On Tue, Jan 31, 2017 at 11:18:49AM +0530, Anshuman Khandual wrote: > Hello Dave/Jerome/Mel, > > Here is the overall layout of the functions I am trying to put together > through this patch series. > > (1) Define CDM from core VM and kernel perspective > > (2) Isolation/Special consideration for HugeTLB allocations > > (3) Isolation/Special consideration for buddy allocations > > (a) Zonelist modification based isolation (proposed) > (b) Cpuset modification based isolation (proposed) > (c) Buddy modification based isolation (working) > > (4) Define VMA containing CDM memory with a new flag VM_CDM > > (5) Special consideration for VM_CDM marked VMAs > > (a) Special consideration for auto NUMA > (b) Special consideration for KSM I believe (5) should not be done on per vma basis but on a page basis. Thus rendering (4) pointless. A vma shouldn't be special because it has some special kind of memory irespective of what the vma points to. > Is there are any other area which needs to be taken care of before CDM > node can be represented completely inside the kernel ? Maybe thing like swap or suspend and resume (i know you are targetting big computer and not laptop :)) but you can't presume what platform CDM might be use latter on. Also userspace might be confuse by looking a /proc/meminfo or any of the sysfs file and see all this device memory without understanding that it is special and might be unwise to be use for regular CPU only task. I would probably want CDM memory be reported separatly from the rest of memory. Which also most likely have repercution with memory cgroup. My expectation is that you only want to use device memory in a process if and only if that process also use the device to some extent. So having new group hierarchy for this memory is probably a better path forward. Cheers, Jérôme
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web