Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1602498 > unrolled thread
| Started by | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| First post | 2017-03-16 16:20 +0100 |
| Last post | 2017-03-19 21:40 +0100 |
| Articles | 12 on this page of 32 — 5 participants |
Back to article view | Back to linux.kernel
[HMM 00/16] HMM (Heterogeneous Memory Management) v18 Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
[HMM 14/16] mm/migrate: allow migrate_vma() to alloc new page on empty entry Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
[HMM 10/16] mm/hmm/mirror: mirror process address space on device with HMM helpers Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
Re: [HMM 10/16] mm/hmm/mirror: mirror process address space on device with HMM helpers Mel Gorman <mgorman@techsingularity.net> - 2017-03-19 21:10 +0100
[HMM 15/16] mm/hmm/devmem: device memory hotplug using ZONE_DEVICE Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
[HMM 09/16] mm/hmm: heterogeneous memory management (HMM for short) Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
Re: [HMM 09/16] mm/hmm: heterogeneous memory management (HMM for short) Mel Gorman <mgorman@techsingularity.net> - 2017-03-19 21:10 +0100
[HMM 11/16] mm/hmm/mirror: helper to snapshot CPU page table v2 Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
Re: [HMM 11/16] mm/hmm/mirror: helper to snapshot CPU page table v2 Mel Gorman <mgorman@techsingularity.net> - 2017-03-19 21:20 +0100
[HMM 03/16] mm/ZONE_DEVICE/free-page: callback when page is freed v3 Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
Re: [HMM 03/16] mm/ZONE_DEVICE/free-page: callback when page is freed v3 Mel Gorman <mgorman@techsingularity.net> - 2017-03-19 21:10 +0100
[HMM 04/16] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory v3 Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
Re: [HMM 04/16] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory v3 Mel Gorman <mgorman@techsingularity.net> - 2017-03-19 21:10 +0100
[HMM 16/16] mm/hmm/devmem: dummy HMM device for ZONE_DEVICE memory v2 Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
Re: [HMM 16/16] mm/hmm/devmem: dummy HMM device for ZONE_DEVICE memory v2 Bob Liu <liubo95@huawei.com> - 2017-03-17 08:10 +0100
Re: [HMM 16/16] mm/hmm/devmem: dummy HMM device for ZONE_DEVICE memory v2 Jerome Glisse <jglisse@redhat.com> - 2017-03-17 18:00 +0100
[HMM 01/16] mm/memory/hotplug: convert device bool to int to allow for more flags v3 Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
Re: [HMM 01/16] mm/memory/hotplug: convert device bool to int to allow for more flags v3 Mel Gorman <mgorman@techsingularity.net> - 2017-03-19 21:10 +0100
[HMM 08/16] mm/migrate: migrate_vma() unmap page from vma while collecting pages Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
[HMM 13/16] mm/hmm/migrate: support un-addressable ZONE_DEVICE page in migration Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:20 +0100
[HMM 05/16] mm/ZONE_DEVICE/x86: add support for un-addressable device memory Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:30 +0100
[HMM 02/16] mm/put_page: move ref decrement to put_zone_device_page() Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:30 +0100
Re: [HMM 02/16] mm/put_page: move ref decrement to put_zone_device_page() Mel Gorman <mgorman@techsingularity.net> - 2017-03-19 21:10 +0100
[HMM 06/16] mm/migrate: add new boolean copy flag to migratepage() callback Jérôme Glisse <jglisse@redhat.com> - 2017-03-16 16:30 +0100
Re: [HMM 06/16] mm/migrate: add new boolean copy flag to migratepage() callback Mel Gorman <mgorman@techsingularity.net> - 2017-03-19 21:10 +0100
Re: [HMM 00/16] HMM (Heterogeneous Memory Management) v18 Andrew Morton <akpm@linux-foundation.org> - 2017-03-16 21:50 +0100
Re: [HMM 00/16] HMM (Heterogeneous Memory Management) v18 Jerome Glisse <jglisse@redhat.com> - 2017-03-17 01:00 +0100
Re: [HMM 00/16] HMM (Heterogeneous Memory Management) v18 Bob Liu <liubo95@huawei.com> - 2017-03-17 09:30 +0100
Re: [HMM 00/16] HMM (Heterogeneous Memory Management) v18 Jerome Glisse <jglisse@redhat.com> - 2017-03-17 17:00 +0100
Re: [HMM 00/16] HMM (Heterogeneous Memory Management) v18 Bob Liu <liubo95@huawei.com> - 2017-03-17 09:50 +0100
Re: [HMM 00/16] HMM (Heterogeneous Memory Management) v18 Jerome Glisse <jglisse@redhat.com> - 2017-03-17 17:20 +0100
Re: [HMM 00/16] HMM (Heterogeneous Memory Management) v18 Mel Gorman <mgorman@techsingularity.net> - 2017-03-19 21:40 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-03-16 16:30 +0100 |
| Subject | [HMM 05/16] mm/ZONE_DEVICE/x86: add support for un-addressable device memory |
| Message-ID | <tlFvI-gv-37@gated-at.bofh.it> |
| In reply to | #1602498 |
It does not need much, just skip populating kernel linear mapping
for range of un-addressable device memory (it is pick so that there
is no physical memory resource overlapping it). All the logic is in
share mm code.
Only support x86-64 as this feature doesn't make much sense with
constrained virtual address space of 32bits architecture.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: "H. Peter Anvin" <hpa@zytor.com>
---
arch/x86/mm/init_64.c | 22 ++++++++++++++++++----
1 file changed, 18 insertions(+), 4 deletions(-)
diff --git a/arch/x86/mm/init_64.c b/arch/x86/mm/init_64.c
index 0098dc9..7c8c91c 100644
--- a/arch/x86/mm/init_64.c
+++ b/arch/x86/mm/init_64.c
@@ -644,7 +644,8 @@ static void update_end_of_memory_vars(u64 start, u64 size)
int arch_add_memory(int nid, u64 start, u64 size, int flags)
{
const int supported_flags = MEMORY_DEVICE |
- MEMORY_DEVICE_ALLOW_MIGRATE;
+ MEMORY_DEVICE_ALLOW_MIGRATE |
+ MEMORY_DEVICE_UNADDRESSABLE;
struct pglist_data *pgdat = NODE_DATA(nid);
struct zone *zone = pgdat->node_zones +
zone_for_memory(nid, start, size, ZONE_NORMAL,
@@ -659,7 +660,17 @@ int arch_add_memory(int nid, u64 start, u64 size, int flags)
return -EINVAL;
}
- init_memory_mapping(start, start + size);
+ /*
+ * We get un-addressable memory when some one is adding a ZONE_DEVICE
+ * to have struct page for a device memory which is not accessible by
+ * the CPU so it is pointless to have a linear kernel mapping of such
+ * memory.
+ *
+ * Core mm should make sure it never set a pte pointing to such fake
+ * physical range.
+ */
+ if (!(flags & MEMORY_DEVICE_UNADDRESSABLE))
+ init_memory_mapping(start, start + size);
ret = __add_pages(nid, zone, start_pfn, nr_pages);
WARN_ON_ONCE(ret);
@@ -958,7 +969,8 @@ kernel_physical_mapping_remove(unsigned long start, unsigned long end)
int __ref arch_remove_memory(u64 start, u64 size, int flags)
{
const int supported_flags = MEMORY_DEVICE |
- MEMORY_DEVICE_ALLOW_MIGRATE;
+ MEMORY_DEVICE_ALLOW_MIGRATE |
+ MEMORY_DEVICE_UNADDRESSABLE;
unsigned long start_pfn = start >> PAGE_SHIFT;
unsigned long nr_pages = size >> PAGE_SHIFT;
struct page *page = pfn_to_page(start_pfn);
@@ -979,7 +991,9 @@ int __ref arch_remove_memory(u64 start, u64 size, int flags)
zone = page_zone(page);
ret = __remove_pages(zone, start_pfn, nr_pages);
WARN_ON_ONCE(ret);
- kernel_physical_mapping_remove(start, start + size);
+
+ if (!(flags & MEMORY_DEVICE_UNADDRESSABLE))
+ kernel_physical_mapping_remove(start, start + size);
return ret;
}
--
2.4.11
[toc] | [prev] | [next] | [standalone]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-03-16 16:30 +0100 |
| Subject | [HMM 02/16] mm/put_page: move ref decrement to put_zone_device_page() |
| Message-ID | <tlFvJ-gv-51@gated-at.bofh.it> |
| In reply to | #1602498 |
This does not affect non ZONE_DEVICE page. In order to allow
ZONE_DEVICE page to be tracked we need to detect when refcount
of a ZONE_DEVICE page reach 1 (not 0 as non ZONE_DEVICE page).
This patch just move put_page_testzero() from put_page() to
put_zone_device_page() and only for ZONE_DEVICE. It does not
add any overhead compare to existing code.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Cc: Dan Williams <dan.j.williams@intel.com>
Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
---
include/linux/mm.h | 8 +++++---
kernel/memremap.c | 2 ++
2 files changed, 7 insertions(+), 3 deletions(-)
diff --git a/include/linux/mm.h b/include/linux/mm.h
index 5f01c88..28e8b28 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -793,11 +793,13 @@ static inline void put_page(struct page *page)
{
page = compound_head(page);
+ if (unlikely(is_zone_device_page(page))) {
+ put_zone_device_page(page);
+ return;
+ }
+
if (put_page_testzero(page))
__put_page(page);
-
- if (unlikely(is_zone_device_page(page)))
- put_zone_device_page(page);
}
#if defined(CONFIG_SPARSEMEM) && !defined(CONFIG_SPARSEMEM_VMEMMAP)
diff --git a/kernel/memremap.c b/kernel/memremap.c
index 40d4af8..c821946 100644
--- a/kernel/memremap.c
+++ b/kernel/memremap.c
@@ -190,6 +190,8 @@ EXPORT_SYMBOL(get_zone_device_page);
void put_zone_device_page(struct page *page)
{
+ page_ref_dec(page);
+
put_dev_pagemap(page->pgmap);
}
EXPORT_SYMBOL(put_zone_device_page);
--
2.4.11
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2017-03-19 21:10 +0100 |
| Subject | Re: [HMM 02/16] mm/put_page: move ref decrement to put_zone_device_page() |
| Message-ID | <tmPjk-2lx-7@gated-at.bofh.it> |
| In reply to | #1602538 |
On Thu, Mar 16, 2017 at 12:05:21PM -0400, J?r?me Glisse wrote:
> This does not affect non ZONE_DEVICE page. In order to allow
> ZONE_DEVICE page to be tracked we need to detect when refcount
> of a ZONE_DEVICE page reach 1 (not 0 as non ZONE_DEVICE page).
>
> This patch just move put_page_testzero() from put_page() to
> put_zone_device_page() and only for ZONE_DEVICE. It does not
> add any overhead compare to existing code.
>
> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
> Cc: Dan Williams <dan.j.williams@intel.com>
> Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
> ---
> include/linux/mm.h | 8 +++++---
> kernel/memremap.c | 2 ++
> 2 files changed, 7 insertions(+), 3 deletions(-)
>
> diff --git a/include/linux/mm.h b/include/linux/mm.h
> index 5f01c88..28e8b28 100644
> --- a/include/linux/mm.h
> +++ b/include/linux/mm.h
> @@ -793,11 +793,13 @@ static inline void put_page(struct page *page)
> {
> page = compound_head(page);
>
> + if (unlikely(is_zone_device_page(page))) {
> + put_zone_device_page(page);
> + return;
> + }
> +
> if (put_page_testzero(page))
> __put_page(page);
> -
> - if (unlikely(is_zone_device_page(page)))
> - put_zone_device_page(page);
> }
>
> #if defined(CONFIG_SPARSEMEM) && !defined(CONFIG_SPARSEMEM_VMEMMAP)
> diff --git a/kernel/memremap.c b/kernel/memremap.c
> index 40d4af8..c821946 100644
> --- a/kernel/memremap.c
> +++ b/kernel/memremap.c
> @@ -190,6 +190,8 @@ EXPORT_SYMBOL(get_zone_device_page);
>
> void put_zone_device_page(struct page *page)
> {
> + page_ref_dec(page);
> +
> put_dev_pagemap(page->pgmap);
> }
> EXPORT_SYMBOL(put_zone_device_page);
So the page refcount goes to zero but where did the __put_page call go? I
haven't read the full series yet but I do note the next patch introduces
a callback. Maybe callbacks free the page but it looks optional. Maybe
it gets fixed later in the series but the changelog should at least say
this is not bisect safe and as this looks like a memory leak.
[toc] | [prev] | [next] | [standalone]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-03-16 16:30 +0100 |
| Subject | [HMM 06/16] mm/migrate: add new boolean copy flag to migratepage() callback |
| Message-ID | <tlFvI-gv-43@gated-at.bofh.it> |
| In reply to | #1602498 |
Allow migration without copy in case destination page already have
source page content. This is usefull for new dma capable migration
where use device dma engine to copy pages.
This feature need carefull audit of filesystem code to make sure
that no one can write to the source page while it is unmapped and
locked. It should be safe for most filesystem but as precaution
return error until support for device migration is added to them.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
---
drivers/staging/lustre/lustre/llite/rw26.c | 8 +++--
fs/aio.c | 7 +++-
fs/btrfs/disk-io.c | 11 ++++--
fs/f2fs/data.c | 8 ++++-
fs/f2fs/f2fs.h | 2 +-
fs/hugetlbfs/inode.c | 9 +++--
fs/nfs/internal.h | 5 +--
fs/nfs/write.c | 9 +++--
fs/ubifs/file.c | 8 ++++-
include/linux/balloon_compaction.h | 3 +-
include/linux/fs.h | 13 ++++---
include/linux/migrate.h | 7 ++--
mm/balloon_compaction.c | 2 +-
mm/migrate.c | 56 +++++++++++++++++++-----------
mm/zsmalloc.c | 12 ++++++-
15 files changed, 114 insertions(+), 46 deletions(-)
diff --git a/drivers/staging/lustre/lustre/llite/rw26.c b/drivers/staging/lustre/lustre/llite/rw26.c
index d89e795..29a59bf 100644
--- a/drivers/staging/lustre/lustre/llite/rw26.c
+++ b/drivers/staging/lustre/lustre/llite/rw26.c
@@ -43,6 +43,7 @@
#include <linux/uaccess.h>
#include <linux/migrate.h>
+#include <linux/memremap.h>
#include <linux/fs.h>
#include <linux/buffer_head.h>
#include <linux/mpage.h>
@@ -642,9 +643,12 @@ static int ll_write_end(struct file *file, struct address_space *mapping,
#ifdef CONFIG_MIGRATION
static int ll_migratepage(struct address_space *mapping,
struct page *newpage, struct page *page,
- enum migrate_mode mode
- )
+ enum migrate_mode mode, bool copy)
{
+ /* Can only migrate addressable memory for now */
+ if (!is_addressable_page(newpage))
+ return -EINVAL;
+
/* Always fail page migration until we have a proper implementation */
return -EIO;
}
diff --git a/fs/aio.c b/fs/aio.c
index f52d925..fa6bb92 100644
--- a/fs/aio.c
+++ b/fs/aio.c
@@ -37,6 +37,7 @@
#include <linux/blkdev.h>
#include <linux/compat.h>
#include <linux/migrate.h>
+#include <linux/memremap.h>
#include <linux/ramfs.h>
#include <linux/percpu-refcount.h>
#include <linux/mount.h>
@@ -366,13 +367,17 @@ static const struct file_operations aio_ring_fops = {
#if IS_ENABLED(CONFIG_MIGRATION)
static int aio_migratepage(struct address_space *mapping, struct page *new,
- struct page *old, enum migrate_mode mode)
+ struct page *old, enum migrate_mode mode, bool copy)
{
struct kioctx *ctx;
unsigned long flags;
pgoff_t idx;
int rc;
+ /* Can only migrate addressable memory for now */
+ if (!is_addressable_page(new))
+ return -EINVAL;
+
rc = 0;
/* mapping->private_lock here protects against the kioctx teardown. */
diff --git a/fs/btrfs/disk-io.c b/fs/btrfs/disk-io.c
index 08b74da..a2b75d6 100644
--- a/fs/btrfs/disk-io.c
+++ b/fs/btrfs/disk-io.c
@@ -27,6 +27,7 @@
#include <linux/kthread.h>
#include <linux/slab.h>
#include <linux/migrate.h>
+#include <linux/memremap.h>
#include <linux/ratelimit.h>
#include <linux/uuid.h>
#include <linux/semaphore.h>
@@ -1061,9 +1062,13 @@ static int btree_submit_bio_hook(struct inode *inode, struct bio *bio,
#ifdef CONFIG_MIGRATION
static int btree_migratepage(struct address_space *mapping,
- struct page *newpage, struct page *page,
- enum migrate_mode mode)
+ struct page *newpage, struct page *page,
+ enum migrate_mode mode, bool copy)
{
+ /* Can only migrate addressable memory for now */
+ if (!is_addressable_page(newpage))
+ return -EINVAL;
+
/*
* we can't safely write a btree page from here,
* we haven't done the locking hook
@@ -1077,7 +1082,7 @@ static int btree_migratepage(struct address_space *mapping,
if (page_has_private(page) &&
!try_to_release_page(page, GFP_KERNEL))
return -EAGAIN;
- return migrate_page(mapping, newpage, page, mode);
+ return migrate_page(mapping, newpage, page, mode, copy);
}
#endif
diff --git a/fs/f2fs/data.c b/fs/f2fs/data.c
index 1602b4b..14208a5 100644
--- a/fs/f2fs/data.c
+++ b/fs/f2fs/data.c
@@ -23,6 +23,7 @@
#include <linux/memcontrol.h>
#include <linux/cleancache.h>
#include <linux/sched/signal.h>
+#include <linux/memremap.h>
#include "f2fs.h"
#include "node.h"
@@ -2049,7 +2050,8 @@ static sector_t f2fs_bmap(struct address_space *mapping, sector_t block)
#include <linux/migrate.h>
int f2fs_migrate_page(struct address_space *mapping,
- struct page *newpage, struct page *page, enum migrate_mode mode)
+ struct page *newpage, struct page *page,
+ enum migrate_mode mode, bool copy)
{
int rc, extra_count;
struct f2fs_inode_info *fi = F2FS_I(mapping->host);
@@ -2057,6 +2059,10 @@ int f2fs_migrate_page(struct address_space *mapping,
BUG_ON(PageWriteback(page));
+ /* Can only migrate addressable memory for now */
+ if (!is_addressable_page(newpage))
+ return -EINVAL;
+
/* migrating an atomic written page is safe with the inmem_lock hold */
if (atomic_written && !mutex_trylock(&fi->inmem_lock))
return -EAGAIN;
diff --git a/fs/f2fs/f2fs.h b/fs/f2fs/f2fs.h
index e849f83..ffa5333 100644
--- a/fs/f2fs/f2fs.h
+++ b/fs/f2fs/f2fs.h
@@ -2299,7 +2299,7 @@ void f2fs_invalidate_page(struct page *page, unsigned int offset,
int f2fs_release_page(struct page *page, gfp_t wait);
#ifdef CONFIG_MIGRATION
int f2fs_migrate_page(struct address_space *mapping, struct page *newpage,
- struct page *page, enum migrate_mode mode);
+ struct page *page, enum migrate_mode mode, bool copy);
#endif
/*
diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c
index 8f96461..13f74d6 100644
--- a/fs/hugetlbfs/inode.c
+++ b/fs/hugetlbfs/inode.c
@@ -35,6 +35,7 @@
#include <linux/security.h>
#include <linux/magic.h>
#include <linux/migrate.h>
+#include <linux/memremap.h>
#include <linux/uio.h>
#include <linux/uaccess.h>
@@ -842,11 +843,15 @@ static int hugetlbfs_set_page_dirty(struct page *page)
}
static int hugetlbfs_migrate_page(struct address_space *mapping,
- struct page *newpage, struct page *page,
- enum migrate_mode mode)
+ struct page *newpage, struct page *page,
+ enum migrate_mode mode, bool copy)
{
int rc;
+ /* Can only migrate addressable memory for now */
+ if (!is_addressable_page(newpage))
+ return -EINVAL;
+
rc = migrate_huge_page_move_mapping(mapping, newpage, page);
if (rc != MIGRATEPAGE_SUCCESS)
return rc;
diff --git a/fs/nfs/internal.h b/fs/nfs/internal.h
index 09ca509..2e23275 100644
--- a/fs/nfs/internal.h
+++ b/fs/nfs/internal.h
@@ -535,8 +535,9 @@ void nfs_clear_pnfs_ds_commit_verifiers(struct pnfs_ds_commit_info *cinfo)
#endif
#ifdef CONFIG_MIGRATION
-extern int nfs_migrate_page(struct address_space *,
- struct page *, struct page *, enum migrate_mode);
+extern int nfs_migrate_page(struct address_space *mapping,
+ struct page *newpage, struct page *page,
+ enum migrate_mode, bool copy);
#endif
static inline int
diff --git a/fs/nfs/write.c b/fs/nfs/write.c
index e75b056..1bc4354 100644
--- a/fs/nfs/write.c
+++ b/fs/nfs/write.c
@@ -14,6 +14,7 @@
#include <linux/writeback.h>
#include <linux/swap.h>
#include <linux/migrate.h>
+#include <linux/memremap.h>
#include <linux/sunrpc/clnt.h>
#include <linux/nfs_fs.h>
@@ -2020,8 +2021,12 @@ int nfs_wb_single_page(struct inode *inode, struct page *page, bool launder)
#ifdef CONFIG_MIGRATION
int nfs_migrate_page(struct address_space *mapping, struct page *newpage,
- struct page *page, enum migrate_mode mode)
+ struct page *page, enum migrate_mode mode, bool copy)
{
+ /* Can only migrate addressable memory for now */
+ if (!is_addressable_page(newpage))
+ return -EINVAL;
+
/*
* If PagePrivate is set, then the page is currently associated with
* an in-progress read or write request. Don't try to migrate it.
@@ -2036,7 +2041,7 @@ int nfs_migrate_page(struct address_space *mapping, struct page *newpage,
if (!nfs_fscache_release_page(page, GFP_KERNEL))
return -EBUSY;
- return migrate_page(mapping, newpage, page, mode);
+ return migrate_page(mapping, newpage, page, mode, copy);
}
#endif
diff --git a/fs/ubifs/file.c b/fs/ubifs/file.c
index d9ae86f..298fbae 100644
--- a/fs/ubifs/file.c
+++ b/fs/ubifs/file.c
@@ -53,6 +53,7 @@
#include <linux/mount.h>
#include <linux/slab.h>
#include <linux/migrate.h>
+#include <linux/memremap.h>
static int read_block(struct inode *inode, void *addr, unsigned int block,
struct ubifs_data_node *dn)
@@ -1469,10 +1470,15 @@ static int ubifs_set_page_dirty(struct page *page)
#ifdef CONFIG_MIGRATION
static int ubifs_migrate_page(struct address_space *mapping,
- struct page *newpage, struct page *page, enum migrate_mode mode)
+ struct page *newpage, struct page *page,
+ enum migrate_mode mode, bool copy)
{
int rc;
+ /* Can only migrate addressable memory for now */
+ if (!is_addressable_page(newpage))
+ return -EINVAL;
+
rc = migrate_page_move_mapping(mapping, newpage, page, NULL, mode, 0);
if (rc != MIGRATEPAGE_SUCCESS)
return rc;
diff --git a/include/linux/balloon_compaction.h b/include/linux/balloon_compaction.h
index 79542b2..27cf3e3 100644
--- a/include/linux/balloon_compaction.h
+++ b/include/linux/balloon_compaction.h
@@ -85,7 +85,8 @@ extern bool balloon_page_isolate(struct page *page,
extern void balloon_page_putback(struct page *page);
extern int balloon_page_migrate(struct address_space *mapping,
struct page *newpage,
- struct page *page, enum migrate_mode mode);
+ struct page *page, enum migrate_mode mode,
+ bool copy);
/*
* balloon_page_insert - insert a page into the balloon's page list and make
diff --git a/include/linux/fs.h b/include/linux/fs.h
index 7251f7b..706a9a9 100644
--- a/include/linux/fs.h
+++ b/include/linux/fs.h
@@ -346,8 +346,9 @@ struct address_space_operations {
* migrate the contents of a page to the specified target. If
* migrate_mode is MIGRATE_ASYNC, it must not block.
*/
- int (*migratepage) (struct address_space *,
- struct page *, struct page *, enum migrate_mode);
+ int (*migratepage)(struct address_space *mapping,
+ struct page *newpage, struct page *page,
+ enum migrate_mode, bool copy);
bool (*isolate_page)(struct page *, isolate_mode_t);
void (*putback_page)(struct page *);
int (*launder_page) (struct page *);
@@ -3013,9 +3014,11 @@ extern int generic_file_fsync(struct file *, loff_t, loff_t, int);
extern int generic_check_addressable(unsigned, u64);
#ifdef CONFIG_MIGRATION
-extern int buffer_migrate_page(struct address_space *,
- struct page *, struct page *,
- enum migrate_mode);
+extern int buffer_migrate_page(struct address_space *mapping,
+ struct page *newpage,
+ struct page *page,
+ enum migrate_mode,
+ bool copy);
#else
#define buffer_migrate_page NULL
#endif
diff --git a/include/linux/migrate.h b/include/linux/migrate.h
index fa76b51..0a66ddd 100644
--- a/include/linux/migrate.h
+++ b/include/linux/migrate.h
@@ -33,8 +33,11 @@ extern char *migrate_reason_names[MR_TYPES];
#ifdef CONFIG_MIGRATION
extern void putback_movable_pages(struct list_head *l);
-extern int migrate_page(struct address_space *,
- struct page *, struct page *, enum migrate_mode);
+extern int migrate_page(struct address_space *mapping,
+ struct page *newpage,
+ struct page *page,
+ enum migrate_mode,
+ bool copy);
extern int migrate_pages(struct list_head *l, new_page_t new, free_page_t free,
unsigned long private, enum migrate_mode mode, int reason);
extern int isolate_movable_page(struct page *page, isolate_mode_t mode);
diff --git a/mm/balloon_compaction.c b/mm/balloon_compaction.c
index da91df5..ed5cacb 100644
--- a/mm/balloon_compaction.c
+++ b/mm/balloon_compaction.c
@@ -135,7 +135,7 @@ void balloon_page_putback(struct page *page)
/* move_to_new_page() counterpart for a ballooned page */
int balloon_page_migrate(struct address_space *mapping,
struct page *newpage, struct page *page,
- enum migrate_mode mode)
+ enum migrate_mode mode, bool copy)
{
struct balloon_dev_info *balloon = balloon_page_device(page);
diff --git a/mm/migrate.c b/mm/migrate.c
index 9a0897a..cb911ce 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -596,18 +596,10 @@ static void copy_huge_page(struct page *dst, struct page *src)
}
}
-/*
- * Copy the page to its new location
- */
-void migrate_page_copy(struct page *newpage, struct page *page)
+static void migrate_page_states(struct page *newpage, struct page *page)
{
int cpupid;
- if (PageHuge(page) || PageTransHuge(page))
- copy_huge_page(newpage, page);
- else
- copy_highpage(newpage, page);
-
if (PageError(page))
SetPageError(newpage);
if (PageReferenced(page))
@@ -661,6 +653,19 @@ void migrate_page_copy(struct page *newpage, struct page *page)
mem_cgroup_migrate(page, newpage);
}
+
+/*
+ * Copy the page to its new location
+ */
+void migrate_page_copy(struct page *newpage, struct page *page)
+{
+ if (PageHuge(page) || PageTransHuge(page))
+ copy_huge_page(newpage, page);
+ else
+ copy_highpage(newpage, page);
+
+ migrate_page_states(newpage, page);
+}
EXPORT_SYMBOL(migrate_page_copy);
/************************************************************
@@ -674,8 +679,8 @@ EXPORT_SYMBOL(migrate_page_copy);
* Pages are locked upon entry and exit.
*/
int migrate_page(struct address_space *mapping,
- struct page *newpage, struct page *page,
- enum migrate_mode mode)
+ struct page *newpage, struct page *page,
+ enum migrate_mode mode, bool copy)
{
int rc;
@@ -686,7 +691,11 @@ int migrate_page(struct address_space *mapping,
if (rc != MIGRATEPAGE_SUCCESS)
return rc;
- migrate_page_copy(newpage, page);
+ if (copy)
+ migrate_page_copy(newpage, page);
+ else
+ migrate_page_states(newpage, page);
+
return MIGRATEPAGE_SUCCESS;
}
EXPORT_SYMBOL(migrate_page);
@@ -698,13 +707,14 @@ EXPORT_SYMBOL(migrate_page);
* exist.
*/
int buffer_migrate_page(struct address_space *mapping,
- struct page *newpage, struct page *page, enum migrate_mode mode)
+ struct page *newpage, struct page *page,
+ enum migrate_mode mode, bool copy)
{
struct buffer_head *bh, *head;
int rc;
if (!page_has_buffers(page))
- return migrate_page(mapping, newpage, page, mode);
+ return migrate_page(mapping, newpage, page, mode, copy);
head = page_buffers(page);
@@ -736,12 +746,15 @@ int buffer_migrate_page(struct address_space *mapping,
SetPagePrivate(newpage);
- migrate_page_copy(newpage, page);
+ if (copy)
+ migrate_page_copy(newpage, page);
+ else
+ migrate_page_states(newpage, page);
bh = head;
do {
unlock_buffer(bh);
- put_bh(bh);
+ put_bh(bh);
bh = bh->b_this_page;
} while (bh != head);
@@ -796,7 +809,8 @@ static int writeout(struct address_space *mapping, struct page *page)
* Default handling if a filesystem does not provide a migration function.
*/
static int fallback_migrate_page(struct address_space *mapping,
- struct page *newpage, struct page *page, enum migrate_mode mode)
+ struct page *newpage, struct page *page,
+ enum migrate_mode mode)
{
if (PageDirty(page)) {
/* Only writeback pages in full synchronous migration */
@@ -813,7 +827,7 @@ static int fallback_migrate_page(struct address_space *mapping,
!try_to_release_page(page, GFP_KERNEL))
return -EAGAIN;
- return migrate_page(mapping, newpage, page, mode);
+ return migrate_page(mapping, newpage, page, mode, true);
}
/*
@@ -841,7 +855,7 @@ static int move_to_new_page(struct page *newpage, struct page *page,
if (likely(is_lru)) {
if (!mapping)
- rc = migrate_page(mapping, newpage, page, mode);
+ rc = migrate_page(mapping, newpage, page, mode, true);
else if (mapping->a_ops->migratepage)
/*
* Most pages have a mapping and most filesystems
@@ -851,7 +865,7 @@ static int move_to_new_page(struct page *newpage, struct page *page,
* for page migration.
*/
rc = mapping->a_ops->migratepage(mapping, newpage,
- page, mode);
+ page, mode, true);
else
rc = fallback_migrate_page(mapping, newpage,
page, mode);
@@ -868,7 +882,7 @@ static int move_to_new_page(struct page *newpage, struct page *page,
}
rc = mapping->a_ops->migratepage(mapping, newpage,
- page, mode);
+ page, mode, true);
WARN_ON_ONCE(rc == MIGRATEPAGE_SUCCESS &&
!PageIsolated(page));
}
diff --git a/mm/zsmalloc.c b/mm/zsmalloc.c
index b7ee9c3..334ff64 100644
--- a/mm/zsmalloc.c
+++ b/mm/zsmalloc.c
@@ -52,6 +52,7 @@
#include <linux/zpool.h>
#include <linux/mount.h>
#include <linux/migrate.h>
+#include <linux/memremap.h>
#include <linux/pagemap.h>
#define ZSPAGE_MAGIC 0x58
@@ -1968,7 +1969,7 @@ bool zs_page_isolate(struct page *page, isolate_mode_t mode)
}
int zs_page_migrate(struct address_space *mapping, struct page *newpage,
- struct page *page, enum migrate_mode mode)
+ struct page *page, enum migrate_mode mode, bool copy)
{
struct zs_pool *pool;
struct size_class *class;
@@ -1986,6 +1987,15 @@ int zs_page_migrate(struct address_space *mapping, struct page *newpage,
VM_BUG_ON_PAGE(!PageMovable(page), page);
VM_BUG_ON_PAGE(!PageIsolated(page), page);
+ /*
+ * Offloading copy operation for zspage require special considerations
+ * due to locking so for now we only support regular migration. I do
+ * not expect we will ever want to support offloading copy. See hmm.h
+ * for more informations on hmm_vma_migrate() and offload copy.
+ */
+ if (!copy || !is_addressable_page(newpage))
+ return -EINVAL;
+
zspage = get_zspage(page);
/* Concurrent compactor cannot migrate any subpage in zspage */
--
2.4.11
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2017-03-19 21:10 +0100 |
| Subject | Re: [HMM 06/16] mm/migrate: add new boolean copy flag to migratepage() callback |
| Message-ID | <tmPjk-2lx-9@gated-at.bofh.it> |
| In reply to | #1602539 |
On Thu, Mar 16, 2017 at 12:05:25PM -0400, J?r?me Glisse wrote:
> Allow migration without copy in case destination page already have
> source page content. This is usefull for new dma capable migration
> where use device dma engine to copy pages.
>
> This feature need carefull audit of filesystem code to make sure
> that no one can write to the source page while it is unmapped and
> locked. It should be safe for most filesystem but as precaution
> return error until support for device migration is added to them.
>
> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
I really dislike the amount of boilerplace code this creates and the fact
that additional headers are needed for that boilerplate. As it's only of
relevance to DMA capable migration, why not simply infer from that if it's
an option instead of updating all supporters of migration?
If that is unsuitable, create a new migreate_mode for a no-copy
migration. You'll need to alter some sites that check the migrate_mode
and it *may* be easier to convert migrate_mode to a bitmask but overall
it would be less boilerplate and confined to just the migration code.
> diff --git a/mm/migrate.c b/mm/migrate.c
> index 9a0897a..cb911ce 100644
> --- a/mm/migrate.c
> +++ b/mm/migrate.c
> @@ -596,18 +596,10 @@ static void copy_huge_page(struct page *dst, struct page *src)
> }
> }
>
> -/*
> - * Copy the page to its new location
> - */
> -void migrate_page_copy(struct page *newpage, struct page *page)
> +static void migrate_page_states(struct page *newpage, struct page *page)
> {
> int cpupid;
>
> - if (PageHuge(page) || PageTransHuge(page))
> - copy_huge_page(newpage, page);
> - else
> - copy_highpage(newpage, page);
> -
> if (PageError(page))
> SetPageError(newpage);
> if (PageReferenced(page))
> @@ -661,6 +653,19 @@ void migrate_page_copy(struct page *newpage, struct page *page)
>
> mem_cgroup_migrate(page, newpage);
> }
> +
> +/*
> + * Copy the page to its new location
> + */
> +void migrate_page_copy(struct page *newpage, struct page *page)
> +{
> + if (PageHuge(page) || PageTransHuge(page))
> + copy_huge_page(newpage, page);
> + else
> + copy_highpage(newpage, page);
> +
> + migrate_page_states(newpage, page);
> +}
> EXPORT_SYMBOL(migrate_page_copy);
>
> /************************************************************
> @@ -674,8 +679,8 @@ EXPORT_SYMBOL(migrate_page_copy);
> * Pages are locked upon entry and exit.
> */
> int migrate_page(struct address_space *mapping,
> - struct page *newpage, struct page *page,
> - enum migrate_mode mode)
> + struct page *newpage, struct page *page,
> + enum migrate_mode mode, bool copy)
> {
> int rc;
>
> @@ -686,7 +691,11 @@ int migrate_page(struct address_space *mapping,
> if (rc != MIGRATEPAGE_SUCCESS)
> return rc;
>
> - migrate_page_copy(newpage, page);
> + if (copy)
> + migrate_page_copy(newpage, page);
> + else
> + migrate_page_states(newpage, page);
> +
> return MIGRATEPAGE_SUCCESS;
> }
> EXPORT_SYMBOL(migrate_page);
Other than some reshuffling, this is the place where the new copy
parameters it used and it has the mode parameter. At worst you end up
creating a helper to check two potential migrate modes to have either
ASYNC, SYNC or SYNC_LIGHT semantics. I expect you want SYNC symantics.
This patch is huge relative to the small thing it acatually requires.
[toc] | [prev] | [next] | [standalone]
| From | Andrew Morton <akpm@linux-foundation.org> |
|---|---|
| Date | 2017-03-16 21:50 +0100 |
| Message-ID | <tlKvo-3HE-9@gated-at.bofh.it> |
| In reply to | #1602498 |
On Thu, 16 Mar 2017 12:05:19 -0400 J__r__me Glisse <jglisse@redhat.com> wrote: > Cliff note: "Cliff's notes" isn't appropriate for a large feature such as this. Where's the long-form description? One which permits readers to fully understand the requirements, design, alternative designs, the implementation, the interface(s), etc? Have you ever spoken about HMM at a conference? If so, the supporting presentation documents might help here. That's the level of detail which should be presented here. > HMM offers 2 things (each standing on its own). First > it allows to use device memory transparently inside any process > without any modifications to process program code. Well. What is "device memory"? That's very vague. What are the characteristics of this memory? Why is it a requirement that userspace code be unaltered? What are the security implications - does the process need particular permissions to access this memory? What is the proposed interface to set up this access? > Second it allows to mirror process address space on a device. Why? Why is this a requirement, how will it be used, what are the use cases, etc? I spent a bit of time trying to locate a decent writeup of this feature but wasn't able to locate one. I'm not seeing a Documentation/ update in this patchset. Perhaps if you were to sit down and write a detailed Documentation/vm/hmm.txt then that would be a good starting point. This stuff is important - it's not really feasible to perform a decent review of this proposal unless the reviewer has access to this high-level conceptual stuff. So I'll take a look at merging this code as-is for testing purposes but I won't be attempting to review it at this stage.
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-03-17 01:00 +0100 |
| Message-ID | <tlNtg-5KZ-11@gated-at.bofh.it> |
| In reply to | #1602829 |
[Multipart message — attachments visible in raw view] — view raw
On Thu, Mar 16, 2017 at 01:43:21PM -0700, Andrew Morton wrote: > On Thu, 16 Mar 2017 12:05:19 -0400 J__r__me Glisse <jglisse@redhat.com> wrote: > > > Cliff note: > > "Cliff's notes" isn't appropriate for a large feature such as this. > Where's the long-form description? One which permits readers to fully > understand the requirements, design, alternative designs, the > implementation, the interface(s), etc? > > Have you ever spoken about HMM at a conference? If so, the supporting > presentation documents might help here. That's the level of detail > which should be presented here. Longer description of patchset rational, motivation and design choices were given in the first few posting of the patchset to which i included a link in my cover letter. Also given that i presented that for last 3 or 4 years to mm summit and kernel summit i thought that by now peoples were familiar about the topic and wanted to spare them the long version. My bad. I attach a patch that is a first stab at a Documentation/hmm.txt that explain the motivation and rational behind HMM. I can probably add a section about how to use HMM from device driver point of view. > > HMM offers 2 things (each standing on its own). First > > it allows to use device memory transparently inside any process > > without any modifications to process program code. > > Well. What is "device memory"? That's very vague. What are the > characteristics of this memory? Why is it a requirement that > userspace code be unaltered? What are the security implications - does > the process need particular permissions to access this memory? What is > the proposed interface to set up this access? Thing like GPU memory, think 16GBytes, 32GBytes with 1TeraBytes/s of bandwidth so something that is just completely in a different category than DDR3/DDR4 or PCIE bandwidth. To allow GPU/FPGA/... to be transparently use by program we need to avoid any requirement to modify any code. Advance in high level langage construct (in C++ but others too) gives opportunities to compiler to leverage GPU transparently without programmer knowledge. But for this to happen we need a share address space ie any pointer in program must be accessible by the device and we must also be able to migrate memory to device memory to benefit from the device memory bandwidth. Moreover if you think about complex software that use a plethora of various library, you want to allow some of the library to leverage GPU or DSP transparently without forcing the library to copy/duplicate its input data which can be highly complex if you think of tree, list, ... Making all this transparent from program/library point of view ease the development of thoses. Quite frankly without that it is border line impossible to efficiently use GPU or other device in many cases. The device memory is treated like regular memory from kernel point of view (except that CPU can not access it) but everything else about page holds (read, write, execution protections ...). So there is no security implications. Device under consideration have page table and works like CPU from process isolation point of view (modulo hardware bug but CPU or main memory have those same issues). There is no propose interface here, nor i see a need for one. When the device starts accessing a range of the process address space the device driver can decide to migrate that range to device memory in order to speed computations. Only the device driver has enough informations on wether or not this is a good idea and this changes continously during run time (depends on what other process are doing ...). So for now like it was discuss in some CDM threads and in some previous HMM threads i believe it is better to let the device driver decide and keep HMM out of any policy choices. Latter down the road once we get more devices and more real world usage we can try to figure out if there is a good way to expose a generic memory placement hint to userspace to allow program to improve performances by helping device driver to make better decissions. > > Second it allows to mirror process address space on a device. > > Why? Why is this a requirement, how will it be used, what are the > use cases, etc? From above, the requirement is that any address the CPU can access could also be access by the device with the same restriction (like read/write protection). This greatly simplify use of such device, either transparently by the compiler without programmer knowledge or through some library again without main program developer knowledge. The whole point is to make it easier to use thing like GPU without having to ask developer to use special memory allocator and to duplicate their dataset. > > I spent a bit of time trying to locate a decent writeup of this feature > but wasn't able to locate one. I'm not seeing a Documentation/ update > in this patchset. Perhaps if you were to sit down and write a detailed > Documentation/vm/hmm.txt then that would be a good starting point. Attach is hmm.txt like i said i thought that all the previous at length description that i have given in the numerous posting of the patchset were enough and that i only needed to refresh peoples memory. > > This stuff is important - it's not really feasible to perform a decent > review of this proposal unless the reviewer has access to this > high-level conceptual stuff. Does the above and the attach documentation answer your questions ? Is there thing i should describe more thouroughly or aspect you feel are missing ? Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Bob Liu <liubo95@huawei.com> |
|---|---|
| Date | 2017-03-17 09:30 +0100 |
| Message-ID | <tlVqN-3v6-1@gated-at.bofh.it> |
| In reply to | #1602922 |
On 2017/3/17 7:49, Jerome Glisse wrote: > On Thu, Mar 16, 2017 at 01:43:21PM -0700, Andrew Morton wrote: >> On Thu, 16 Mar 2017 12:05:19 -0400 J__r__me Glisse <jglisse@redhat.com> wrote: >> >>> Cliff note: >> >> "Cliff's notes" isn't appropriate for a large feature such as this. >> Where's the long-form description? One which permits readers to fully >> understand the requirements, design, alternative designs, the >> implementation, the interface(s), etc? >> >> Have you ever spoken about HMM at a conference? If so, the supporting >> presentation documents might help here. That's the level of detail >> which should be presented here. > > Longer description of patchset rational, motivation and design choices > were given in the first few posting of the patchset to which i included > a link in my cover letter. Also given that i presented that for last 3 > or 4 years to mm summit and kernel summit i thought that by now peoples > were familiar about the topic and wanted to spare them the long version. > My bad. > > I attach a patch that is a first stab at a Documentation/hmm.txt that > explain the motivation and rational behind HMM. I can probably add a > section about how to use HMM from device driver point of view. > Please, that would be very helpful! > +3) Share address space and migration > + > +HMM intends to provide two main features. First one is to share the address > +space by duplication the CPU page table into the device page table so same > +address point to same memory and this for any valid main memory address in > +the process address space. Is this an optional feature? I mean the device don't have to duplicate the CPU page table. But only make use of the second(migration) feature. > +The second mechanism HMM provide is a new kind of ZONE_DEVICE memory that does > +allow to allocate a struct page for each page of the device memory. Those page > +are special because the CPU can not map them. They however allow to migrate > +main memory to device memory using exhisting migration mechanism and everything > +looks like if page was swap out to disk from CPU point of view. Using a struct > +page gives the easiest and cleanest integration with existing mm mechanisms. > +Again here HMM only provide helpers, first to hotplug new ZONE_DEVICE memory > +for the device memory and second to perform migration. Policy decision of what > +and when to migrate things is left to the device driver. > + > +Note that any CPU acess to a device page trigger a page fault which initiate a > +migration back to system memory so that CPU can access it. A bit confused here, do you mean CPU access to a main memory page but that page has been migrated to device memory? Then a page fault will be triggered and initiate a migration back. Thanks, Bob
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-03-17 17:00 +0100 |
| Message-ID | <tm2si-kO-27@gated-at.bofh.it> |
| In reply to | #1603105 |
On Fri, Mar 17, 2017 at 04:29:10PM +0800, Bob Liu wrote: > On 2017/3/17 7:49, Jerome Glisse wrote: > > On Thu, Mar 16, 2017 at 01:43:21PM -0700, Andrew Morton wrote: > >> On Thu, 16 Mar 2017 12:05:19 -0400 J__r__me Glisse <jglisse@redhat.com> wrote: > >> > >>> Cliff note: > >> > >> "Cliff's notes" isn't appropriate for a large feature such as this. > >> Where's the long-form description? One which permits readers to fully > >> understand the requirements, design, alternative designs, the > >> implementation, the interface(s), etc? > >> > >> Have you ever spoken about HMM at a conference? If so, the supporting > >> presentation documents might help here. That's the level of detail > >> which should be presented here. > > > > Longer description of patchset rational, motivation and design choices > > were given in the first few posting of the patchset to which i included > > a link in my cover letter. Also given that i presented that for last 3 > > or 4 years to mm summit and kernel summit i thought that by now peoples > > were familiar about the topic and wanted to spare them the long version. > > My bad. > > > > I attach a patch that is a first stab at a Documentation/hmm.txt that > > explain the motivation and rational behind HMM. I can probably add a > > section about how to use HMM from device driver point of view. > > > > Please, that would be very helpful! > > > +3) Share address space and migration > > + > > +HMM intends to provide two main features. First one is to share the address > > +space by duplication the CPU page table into the device page table so same > > +address point to same memory and this for any valid main memory address in > > +the process address space. > > Is this an optional feature? > I mean the device don't have to duplicate the CPU page table. > But only make use of the second(migration) feature. Correct each feature can be use on its own without the other. > > +The second mechanism HMM provide is a new kind of ZONE_DEVICE memory that does > > +allow to allocate a struct page for each page of the device memory. Those page > > +are special because the CPU can not map them. They however allow to migrate > > +main memory to device memory using exhisting migration mechanism and everything > > +looks like if page was swap out to disk from CPU point of view. Using a struct > > +page gives the easiest and cleanest integration with existing mm mechanisms. > > +Again here HMM only provide helpers, first to hotplug new ZONE_DEVICE memory > > +for the device memory and second to perform migration. Policy decision of what > > +and when to migrate things is left to the device driver. > > + > > +Note that any CPU acess to a device page trigger a page fault which initiate a > > +migration back to system memory so that CPU can access it. > > A bit confused here, do you mean CPU access to a main memory page but that page has > been migrated to device memory? > Then a page fault will be triggered and initiate a migration back. If you migrate the page backing address A from a main memory page to a device page and then CPU try to access address A then you get a page fault because device memory is not accessible by CPU. The page fault is exactly as if the page was swap out to disk from kernel point of view. At any point in time there is only one and one page backing an address either a regular main memory page or device page. There is no change here to this fundamental fact in respect to mm. The only difference is that device page are not accessible by CPU. Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Bob Liu <liubo95@huawei.com> |
|---|---|
| Date | 2017-03-17 09:50 +0100 |
| Message-ID | <tlVKa-3Fq-9@gated-at.bofh.it> |
| In reply to | #1602922 |
On 2017/3/17 7:49, Jerome Glisse wrote: > On Thu, Mar 16, 2017 at 01:43:21PM -0700, Andrew Morton wrote: >> On Thu, 16 Mar 2017 12:05:19 -0400 J__r__me Glisse <jglisse@redhat.com> wrote: >> >>> Cliff note: >> >> "Cliff's notes" isn't appropriate for a large feature such as this. >> Where's the long-form description? One which permits readers to fully >> understand the requirements, design, alternative designs, the >> implementation, the interface(s), etc? >> >> Have you ever spoken about HMM at a conference? If so, the supporting >> presentation documents might help here. That's the level of detail >> which should be presented here. > > Longer description of patchset rational, motivation and design choices > were given in the first few posting of the patchset to which i included > a link in my cover letter. Also given that i presented that for last 3 > or 4 years to mm summit and kernel summit i thought that by now peoples > were familiar about the topic and wanted to spare them the long version. > My bad. > > I attach a patch that is a first stab at a Documentation/hmm.txt that > explain the motivation and rational behind HMM. I can probably add a > section about how to use HMM from device driver point of view. > And a simple example program/pseudo-code make use of the device memory would also very useful for person don't have GPU programming experience :) Regards, Bob
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2017-03-17 17:20 +0100 |
| Message-ID | <tm2LE-KO-23@gated-at.bofh.it> |
| In reply to | #1603114 |
On Fri, Mar 17, 2017 at 04:39:28PM +0800, Bob Liu wrote: > On 2017/3/17 7:49, Jerome Glisse wrote: > > On Thu, Mar 16, 2017 at 01:43:21PM -0700, Andrew Morton wrote: > >> On Thu, 16 Mar 2017 12:05:19 -0400 J__r__me Glisse <jglisse@redhat.com> wrote: > >> > >>> Cliff note: > >> > >> "Cliff's notes" isn't appropriate for a large feature such as this. > >> Where's the long-form description? One which permits readers to fully > >> understand the requirements, design, alternative designs, the > >> implementation, the interface(s), etc? > >> > >> Have you ever spoken about HMM at a conference? If so, the supporting > >> presentation documents might help here. That's the level of detail > >> which should be presented here. > > > > Longer description of patchset rational, motivation and design choices > > were given in the first few posting of the patchset to which i included > > a link in my cover letter. Also given that i presented that for last 3 > > or 4 years to mm summit and kernel summit i thought that by now peoples > > were familiar about the topic and wanted to spare them the long version. > > My bad. > > > > I attach a patch that is a first stab at a Documentation/hmm.txt that > > explain the motivation and rational behind HMM. I can probably add a > > section about how to use HMM from device driver point of view. > > > > And a simple example program/pseudo-code make use of the device memory > would also very useful for person don't have GPU programming experience :) Like i said there is no userspace API to this. Right now it is under driver control what and when to migrate. So this is specific to each driver and without a driver which use this feature nothing happen. Each driver will expose its own API that probably won't be expose to the end user but to the user space driver (OpenCL, Cuda, C++, OpenMP, ...). We are not sure what kind of API we will expose in the nouveau driver this still need to be discuss. Same for the AMD driver. Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2017-03-19 21:40 +0100 |
| Message-ID | <tmPMl-2yV-9@gated-at.bofh.it> |
| In reply to | #1602498 |
On Thu, Mar 16, 2017 at 12:05:19PM -0400, J?r?me Glisse wrote: > Cliff note: HMM offers 2 things (each standing on its own). First > it allows to use device memory transparently inside any process > without any modifications to process program code. Second it allows > to mirror process address space on a device. > > Changes since v17: > - typos > - ZONE_DEVICE page refcount move put_zone_device_page() > > Work is still underway to use this feature inside the upstream > nouveau driver. It has been tested with closed source driver > and test are still underway on top of new kernel. So far we have > found no issues. I expect to get a tested-by soon. Also this > feature is not only useful for NVidia GPU, i expect AMD GPU will > need it too if they want to support some of the new industry API. > I also expect some FPGA company to use it and probably other > hardware. > > That being said I don't expect i will ever get a review-by anyone > for reasons beyond my control. I spent the length of time a battery lasts reading the patches during my flight to LSF/MM showing that you can get people to review anything if you lock them in a metal box for a few hours. I only got as far as patch 13 before running low on time but decided to send what I have anyway so you have the feedback before the LSF/MM topic. The remaining patches are HMM specific and the intent was review how much the core mm is affected and how hard this would be to maintain. I was less concerned with the HMM internals itself but I assume that the authors writing driver support can supply tested-by's. Overall HMM is fairly well isolated. The drivers can cause new and interesting damage through the MMU notifiers and fault handling but that is a driver, not a core, issue. There is new core code but most of it is active only if a driver is so most people won't notice. Fast paths generally remain unaffected except for one major case covered in the review. I also didn't like the migrate_page API update and suggested an alternative. Most of the other overhead is very minor. My expection is that most core code does not have to care about HMM and while there is a risk that a driver can cause damage through the notifiers, that is completely the responsibility of the driver. Maybe some buglets exist in the new core migration code but again, most people won't notice unless a suitable driver is loaded. On that basis, if you address the major aspects of this review, I don't have an objection at the moment to HMM being merged unlike the objections I had to the CDM preparation patches that modified zonelist handling, nodes and the page allocator fast paths. It still leaves the problem of no in-kernel user of the API. The catch-22 has now existed for years that driver support won't exist until it's merged and it won't get merged without drivers. I won't object strongly on that basis any more but others might. Maybe if this passes Andrew's review it could be staged in mmotm until a driver or something like CDM is ready? That would at least give a tree for driver authors to work against with the resonable expectation that both HMM + driver would go in at the same time.
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web