Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1525568 > unrolled thread
| Started by | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| First post | 2016-11-18 18:20 +0100 |
| Last post | 2016-11-19 16:00 +0100 |
| Articles | 14 on this page of 34 — 6 participants |
Back to article view | Back to linux.kernel
[HMM v13 00/18] HMM (Heterogeneous Memory Management) v13 Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:20 +0100
[HMM v13 13/18] mm/hmm/mirror: device page fault handler Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:20 +0100
[HMM v13 12/18] mm/hmm/mirror: helper to snapshot CPU page table Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
[HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-21 09:10 +0100
Re: [HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory Jerome Glisse <jglisse@redhat.com> - 2016-11-21 13:40 +0100
Re: [HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-22 06:20 +0100
[HMM v13 09/18] mm/hmm/mirror: mirror process address space on device with HMM helpers Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 09/18] mm/hmm/mirror: mirror process address space on device with HMM helpers Balbir Singh <bsingharora@gmail.com> - 2016-11-21 03:50 +0100
Re: [HMM v13 09/18] mm/hmm/mirror: mirror process address space on device with HMM helpers Jerome Glisse <jglisse@redhat.com> - 2016-11-21 06:20 +0100
[HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-21 11:40 +0100
Re: [HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory Jerome Glisse <jglisse@redhat.com> - 2016-11-21 13:40 +0100
Re: [HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-22 06:00 +0100
[HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Balbir Singh <bsingharora@gmail.com> - 2016-11-21 03:00 +0100
Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Jerome Glisse <jglisse@redhat.com> - 2016-11-21 06:00 +0100
Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-21 09:30 +0100
Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Jerome Glisse <jglisse@redhat.com> - 2016-11-21 13:40 +0100
Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-22 06:10 +0100
[HMM v13 10/18] mm/hmm/mirror: add range lock helper, prevent CPU page table update for the range Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
[HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Balbir Singh <bsingharora@gmail.com> - 2016-11-21 03:10 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Jerome Glisse <jglisse@redhat.com> - 2016-11-21 06:10 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Balbir Singh <bsingharora@gmail.com> - 2016-11-22 03:30 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Jerome Glisse <jglisse@redhat.com> - 2016-11-22 15:00 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-21 12:20 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-21 12:00 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Jerome Glisse <jglisse@redhat.com> - 2016-11-21 13:50 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-22 05:50 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Jerome Glisse <jglisse@redhat.com> - 2016-11-24 15:00 +0100
[HMM v13 11/18] mm/hmm/mirror: add range monitor helper, to monitor CPU page table update Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 00/18] HMM (Heterogeneous Memory Management) v13 John Hubbard <jhubbard@nvidia.com> - 2016-11-19 01:50 +0100
Re: [HMM v13 00/18] HMM (Heterogeneous Memory Management) v13 "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-11-19 16:00 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-18 18:30 +0100 |
| Subject | [HMM v13 10/18] mm/hmm/mirror: add range lock helper, prevent CPU page table update for the range |
| Message-ID | <sEV98-1Es-33@gated-at.bofh.it> |
| In reply to | #1525568 |
There is two possible strategy when it comes to snapshoting the CPU page table
inside the device page table. First one snapshot the CPU page table and keep
track of active mmu_notifier callback. Once snapshot is done and before updating
the device page table (in an atomic fashion) it check the mmu_notifier sequence.
If sequence is same as the time the CPU page table was snapshot then it means
that no mmu_notifier run in the meantime and hence the snapshot is accurate. If
the sequence is different then one mmu_notifier callback did run and snapshot
might no longer be valid and the whole procedure must be restarted.
Issue with this approach is that it does not garanty forward progress for the
device driver trying to mirror a range of the address space.
The second solution, implemented by this patch, is to serialize CPU snapshot
with mmu_notifier callback and have each waiting on each other according to the
order they happen. This garanty forward progress for driver. The drawback is
that it can stall process waiting on the mmu_notifier callback to finish. So
thing like direct page reclaim (or even indirect one) might stall and this might
increase overall kernel latency.
For now just accept this potential issue and wait to have real world workload to
be affected by it before trying to fix it. Fix is probably to introduce a new
mmu_notifier_try_to_invalidate() that could return failure if it has to wait or
sleep and use it inside reclaim code to decide to skip to next candidate for
reclaimation.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Signed-off-by: Jatin Kumar <jakumar@nvidia.com>
Signed-off-by: John Hubbard <jhubbard@nvidia.com>
Signed-off-by: Mark Hairgrove <mhairgrove@nvidia.com>
Signed-off-by: Sherry Cheung <SCheung@nvidia.com>
Signed-off-by: Subhash Gutti <sgutti@nvidia.com>
---
include/linux/hmm.h | 30 ++++++++++++
mm/hmm.c | 131 +++++++++++++++++++++++++++++++++++++++++++++++++---
2 files changed, 154 insertions(+), 7 deletions(-)
diff --git a/include/linux/hmm.h b/include/linux/hmm.h
index f44e270..c0b1c07 100644
--- a/include/linux/hmm.h
+++ b/include/linux/hmm.h
@@ -224,6 +224,36 @@ int hmm_mirror_register(struct hmm_mirror *mirror, struct mm_struct *mm);
void hmm_mirror_unregister(struct hmm_mirror *mirror);
+/*
+ * struct hmm_range - track invalidation lock on virtual address range
+ *
+ * @hmm: core hmm struct this range is active against
+ * @list: all range lock are on a list
+ * @start: range virtual start address (inclusive)
+ * @end: range virtual end address (exclusive)
+ * @waiting: pointer to range waiting on this one
+ * @wakeup: use to wakeup the range when it was waiting
+ */
+struct hmm_range {
+ struct hmm *hmm;
+ struct list_head list;
+ unsigned long start;
+ unsigned long end;
+ struct hmm_range *waiting;
+ bool wakeup;
+};
+
+/*
+ * Range locking allow to garanty forward progress by blocking CPU page table
+ * invalidation. See functions description in mm/hmm.c for documentation.
+ */
+int hmm_vma_range_lock(struct hmm_range *range,
+ struct vm_area_struct *vma,
+ unsigned long start,
+ unsigned long end);
+void hmm_vma_range_unlock(struct hmm_range *range);
+
+
/* Below are for HMM internal use only ! Not to be use by device driver ! */
void hmm_mm_destroy(struct mm_struct *mm);
diff --git a/mm/hmm.c b/mm/hmm.c
index 3594785..ee05419 100644
--- a/mm/hmm.c
+++ b/mm/hmm.c
@@ -27,7 +27,8 @@
* struct hmm - HMM per mm struct
*
* @mm: mm struct this HMM struct is bound to
- * @lock: lock protecting mirrors list
+ * @lock: lock protecting mirrors and ranges list
+ * @ranges: list of range lock (for snapshot and invalidation serialization)
* @mirrors: list of mirrors for this mm
* @wait_queue: wait queue
* @sequence: we track update to CPU page table with a sequence number
@@ -37,6 +38,7 @@
struct hmm {
struct mm_struct *mm;
spinlock_t lock;
+ struct list_head ranges;
struct list_head mirrors;
atomic_t sequence;
wait_queue_head_t wait_queue;
@@ -66,6 +68,7 @@ static struct hmm *hmm_register(struct mm_struct *mm)
INIT_LIST_HEAD(&hmm->mirrors);
atomic_set(&hmm->sequence, 0);
hmm->mmu_notifier.ops = NULL;
+ INIT_LIST_HEAD(&hmm->ranges);
spin_lock_init(&hmm->lock);
hmm->mm = mm;
}
@@ -104,16 +107,48 @@ void hmm_mm_destroy(struct mm_struct *mm)
kfree(hmm);
}
-
-
static void hmm_invalidate_range(struct hmm *hmm,
enum hmm_update action,
unsigned long start,
unsigned long end)
{
+ struct hmm_range range, *tmp;
struct hmm_mirror *mirror;
/*
+ * Serialize invalidation with CPU snapshot (see hmm_vma_range_lock()).
+ * Need to make change to mmu_notifier so that we can get a struct that
+ * stay alive accross call to mmu_notifier_invalidate_range_start() and
+ * mmu_notifier_invalidate_range_end(). FIXME !
+ */
+ range.waiting = NULL;
+ range.start = start;
+ range.end = end;
+ range.hmm = hmm;
+
+ spin_lock(&hmm->lock);
+ list_for_each_entry (tmp, &hmm->ranges, list) {
+ if (range.start >= tmp->end || range.end <= tmp->start)
+ continue;
+
+ while (tmp->waiting)
+ tmp = tmp->waiting;
+
+ list_add(&range.list, &hmm->ranges);
+ tmp->waiting = ⦥
+ range.wakeup = false;
+ spin_unlock(&hmm->lock);
+
+ wait_event(hmm->wait_queue, range.wakeup);
+ return;
+ }
+ list_add(&range.list, &hmm->ranges);
+ spin_unlock(&hmm->lock);
+
+ atomic_inc(&hmm->notifier_count);
+ atomic_inc(&hmm->sequence);
+
+ /*
* Mirror being added or remove is a rare event so list traversal isn't
* protected by a lock, we rely on simple rules. All list modification
* are done using list_add_rcu() and list_del_rcu() under a spinlock to
@@ -127,6 +162,9 @@ static void hmm_invalidate_range(struct hmm *hmm,
*/
list_for_each_entry (mirror, &hmm->mirrors, list)
mirror->ops->update(mirror, action, start, end);
+
+ /* See above FIXME */
+ hmm_vma_range_unlock(&range);
}
static void hmm_invalidate_page(struct mmu_notifier *mn,
@@ -139,8 +177,6 @@ static void hmm_invalidate_page(struct mmu_notifier *mn,
VM_BUG_ON(!hmm);
- atomic_inc(&hmm->notifier_count);
- atomic_inc(&hmm->sequence);
hmm_invalidate_range(mm->hmm, HMM_UPDATE_INVALIDATE, start, end);
atomic_dec(&hmm->notifier_count);
wake_up(&hmm->wait_queue);
@@ -155,8 +191,6 @@ static void hmm_invalidate_range_start(struct mmu_notifier *mn,
VM_BUG_ON(!hmm);
- atomic_inc(&hmm->notifier_count);
- atomic_inc(&hmm->sequence);
hmm_invalidate_range(mm->hmm, HMM_UPDATE_INVALIDATE, start, end);
}
@@ -244,3 +278,86 @@ void hmm_mirror_unregister(struct hmm_mirror *mirror)
wait_event(hmm->wait_queue, !atomic_read(&hmm->notifier_count));
}
EXPORT_SYMBOL(hmm_mirror_unregister);
+
+
+/*
+ * hmm_vma_range_lock() - lock invalidation of a virtual address range
+ * @range: range lock struct provided by caller to track lock while valid
+ * @vma: virtual memory area containing the virtual address range
+ * @start: range virtual start address (inclusive)
+ * @end: range virtual end address (exclusive)
+ * Returns: -EINVAL or -ENOMEM on error, 0 otherwise
+ *
+ * This will block any invalidation to CPU page table for the range of virtual
+ * address provided as argument. Design pattern is :
+ * hmm_vma_range_lock(vma, start, end, lock);
+ * hmm_vma_range_get_pfns(vma, start, end, pfns);
+ * // Device driver goes over each pfn in the pfns array, snapshot of CPU
+ * // page table and take appropriate actions (use it to populate GPU page
+ * // table, identify address that need faulting, prepare migration, ...)
+ * hmm_vma_range_unlock(&lock);
+ *
+ * DO NOT HOLD THE RANGE LOCK FOR LONGER THAN NECESSARY ! THIS DOES BLOCK CPU
+ * PAGE TABLE INVALIDATION !
+ */
+int hmm_vma_range_lock(struct hmm_range *range,
+ struct vm_area_struct *vma,
+ unsigned long start,
+ unsigned long end)
+{
+ struct hmm *hmm;
+
+ VM_BUG_ON(!vma);
+ VM_BUG_ON(!rwsem_is_locked(&vma->vm_mm->mmap_sem));
+
+ range->hmm = hmm = hmm_register(vma->vm_mm);
+ if (!hmm)
+ return -ENOMEM;
+
+ if (start < vma->vm_start || start >= vma->vm_end)
+ return -EINVAL;
+ if (end < vma->vm_start || end > vma->vm_end)
+ return -EINVAL;
+
+ range->waiting = NULL;
+ range->start = start;
+ range->end = end;
+
+ spin_lock(&hmm->lock);
+ list_add(&range->list, &hmm->ranges);
+ spin_unlock(&hmm->lock);
+
+ /*
+ * Wait for all active mmu_notifier this is because we can not keep an
+ * hmm_range struct around while mmu_notifier is between a start and
+ * end section. This need change to mmu_notifier FIXME !
+ */
+ wait_event(hmm->wait_queue, !atomic_read(&hmm->notifier_count));
+
+ return 0;
+}
+EXPORT_SYMBOL(hmm_vma_range_lock);
+
+/*
+ * hmm_vma_range_unlock() - unlock invalidation of a virtual address range
+ * @lock: lock struct tracking the range lock
+ *
+ * See hmm_vma_range_lock() for usage.
+ */
+void hmm_vma_range_unlock(struct hmm_range *range)
+{
+ struct hmm *hmm = range->hmm;
+ bool wakeup = false;
+
+ spin_lock(&hmm->lock);
+ list_del(&range->list);
+ if (range->waiting) {
+ range->waiting->wakeup = true;
+ wakeup = true;
+ }
+ spin_unlock(&hmm->lock);
+
+ if (wakeup)
+ wake_up(&hmm->wait_queue);
+}
+EXPORT_SYMBOL(hmm_vma_range_unlock);
--
2.4.3
[toc] | [prev] | [next] | [standalone]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-18 18:30 +0100 |
| Subject | [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable |
| Message-ID | <sEV98-1Es-35@gated-at.bofh.it> |
| In reply to | #1525568 |
To allow use of device un-addressable memory inside a process add a
special swap type. Also add a new callback to handle page fault on
such entry.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Cc: Dan Williams <dan.j.williams@intel.com>
Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
---
fs/proc/task_mmu.c | 10 +++++++-
include/linux/memremap.h | 5 ++++
include/linux/swap.h | 18 ++++++++++---
include/linux/swapops.h | 67 ++++++++++++++++++++++++++++++++++++++++++++++++
kernel/memremap.c | 14 ++++++++++
mm/Kconfig | 12 +++++++++
mm/memory.c | 24 +++++++++++++++++
mm/mprotect.c | 12 +++++++++
8 files changed, 158 insertions(+), 4 deletions(-)
diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
index 6909582..0726d39 100644
--- a/fs/proc/task_mmu.c
+++ b/fs/proc/task_mmu.c
@@ -544,8 +544,11 @@ static void smaps_pte_entry(pte_t *pte, unsigned long addr,
} else {
mss->swap_pss += (u64)PAGE_SIZE << PSS_SHIFT;
}
- } else if (is_migration_entry(swpent))
+ } else if (is_migration_entry(swpent)) {
page = migration_entry_to_page(swpent);
+ } else if (is_device_entry(swpent)) {
+ page = device_entry_to_page(swpent);
+ }
} else if (unlikely(IS_ENABLED(CONFIG_SHMEM) && mss->check_shmem_swap
&& pte_none(*pte))) {
page = find_get_entry(vma->vm_file->f_mapping,
@@ -708,6 +711,8 @@ static int smaps_hugetlb_range(pte_t *pte, unsigned long hmask,
if (is_migration_entry(swpent))
page = migration_entry_to_page(swpent);
+ if (is_device_entry(swpent))
+ page = device_entry_to_page(swpent);
}
if (page) {
int mapcount = page_mapcount(page);
@@ -1191,6 +1196,9 @@ static pagemap_entry_t pte_to_pagemap_entry(struct pagemapread *pm,
flags |= PM_SWAP;
if (is_migration_entry(entry))
page = migration_entry_to_page(entry);
+
+ if (is_device_entry(entry))
+ page = device_entry_to_page(entry);
}
if (page && !PageAnon(page))
diff --git a/include/linux/memremap.h b/include/linux/memremap.h
index b6f03e9..d584c74 100644
--- a/include/linux/memremap.h
+++ b/include/linux/memremap.h
@@ -47,6 +47,11 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
*/
struct dev_pagemap {
void (*free_devpage)(struct page *page, void *data);
+ int (*fault)(struct vm_area_struct *vma,
+ unsigned long addr,
+ struct page *page,
+ unsigned flags,
+ pmd_t *pmdp);
struct vmem_altmap *altmap;
const struct resource *res;
struct percpu_ref *ref;
diff --git a/include/linux/swap.h b/include/linux/swap.h
index 7e553e1..599cb54 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -50,6 +50,17 @@ static inline int current_is_kswapd(void)
*/
/*
+ * Un-addressable device memory support
+ */
+#ifdef CONFIG_DEVICE_UNADDRESSABLE
+#define SWP_DEVICE_NUM 2
+#define SWP_DEVICE_WRITE (MAX_SWAPFILES + SWP_HWPOISON_NUM + SWP_MIGRATION_NUM)
+#define SWP_DEVICE (MAX_SWAPFILES + SWP_HWPOISON_NUM + SWP_MIGRATION_NUM + 1)
+#else
+#define SWP_DEVICE_NUM 0
+#endif
+
+/*
* NUMA node memory migration support
*/
#ifdef CONFIG_MIGRATION
@@ -71,7 +82,8 @@ static inline int current_is_kswapd(void)
#endif
#define MAX_SWAPFILES \
- ((1 << MAX_SWAPFILES_SHIFT) - SWP_MIGRATION_NUM - SWP_HWPOISON_NUM)
+ ((1 << MAX_SWAPFILES_SHIFT) - SWP_DEVICE_NUM - \
+ SWP_MIGRATION_NUM - SWP_HWPOISON_NUM)
/*
* Magic header for a swap area. The first part of the union is
@@ -442,8 +454,8 @@ static inline void show_swap_cache_info(void)
{
}
-#define free_swap_and_cache(swp) is_migration_entry(swp)
-#define swapcache_prepare(swp) is_migration_entry(swp)
+#define free_swap_and_cache(e) (is_migration_entry(e) || is_device_entry(e))
+#define swapcache_prepare(e) (is_migration_entry(e) || is_device_entry(e))
static inline int add_swap_count_continuation(swp_entry_t swp, gfp_t gfp_mask)
{
diff --git a/include/linux/swapops.h b/include/linux/swapops.h
index 5c3a5f3..d1aa425 100644
--- a/include/linux/swapops.h
+++ b/include/linux/swapops.h
@@ -100,6 +100,73 @@ static inline void *swp_to_radix_entry(swp_entry_t entry)
return (void *)(value | RADIX_TREE_EXCEPTIONAL_ENTRY);
}
+#ifdef CONFIG_DEVICE_UNADDRESSABLE
+static inline swp_entry_t make_device_entry(struct page *page, bool write)
+{
+ return swp_entry(write?SWP_DEVICE_WRITE:SWP_DEVICE, page_to_pfn(page));
+}
+
+static inline bool is_device_entry(swp_entry_t entry)
+{
+ int type = swp_type(entry);
+ return type == SWP_DEVICE || type == SWP_DEVICE_WRITE;
+}
+
+static inline void make_device_entry_read(swp_entry_t *entry)
+{
+ *entry = swp_entry(SWP_DEVICE, swp_offset(*entry));
+}
+
+static inline bool is_write_device_entry(swp_entry_t entry)
+{
+ return unlikely(swp_type(entry) == SWP_DEVICE_WRITE);
+}
+
+static inline struct page *device_entry_to_page(swp_entry_t entry)
+{
+ return pfn_to_page(swp_offset(entry));
+}
+
+int device_entry_fault(struct vm_area_struct *vma,
+ unsigned long addr,
+ swp_entry_t entry,
+ unsigned flags,
+ pmd_t *pmdp);
+#else /* CONFIG_DEVICE_UNADDRESSABLE */
+static inline swp_entry_t make_device_entry(struct page *page, bool write)
+{
+ return swp_entry(0, 0);
+}
+
+static inline void make_device_entry_read(swp_entry_t *entry)
+{
+}
+
+static inline bool is_device_entry(swp_entry_t entry)
+{
+ return false;
+}
+
+static inline bool is_write_device_entry(swp_entry_t entry)
+{
+ return false;
+}
+
+static inline struct page *device_entry_to_page(swp_entry_t entry)
+{
+ return NULL;
+}
+
+static inline int device_entry_fault(struct vm_area_struct *vma,
+ unsigned long addr,
+ swp_entry_t entry,
+ unsigned flags,
+ pmd_t *pmdp)
+{
+ return VM_FAULT_SIGBUS;
+}
+#endif /* CONFIG_DEVICE_UNADDRESSABLE */
+
#ifdef CONFIG_MIGRATION
static inline swp_entry_t make_migration_entry(struct page *page, int write)
{
diff --git a/kernel/memremap.c b/kernel/memremap.c
index cf83928..0670015 100644
--- a/kernel/memremap.c
+++ b/kernel/memremap.c
@@ -18,6 +18,8 @@
#include <linux/io.h>
#include <linux/mm.h>
#include <linux/memory_hotplug.h>
+#include <linux/swap.h>
+#include <linux/swapops.h>
#ifndef ioremap_cache
/* temporary while we convert existing ioremap_cache users to memremap */
@@ -200,6 +202,18 @@ void put_zone_device_page(struct page *page)
}
EXPORT_SYMBOL(put_zone_device_page);
+int device_entry_fault(struct vm_area_struct *vma,
+ unsigned long addr,
+ swp_entry_t entry,
+ unsigned flags,
+ pmd_t *pmdp)
+{
+ struct page *page = device_entry_to_page(entry);
+
+ return page->pgmap->fault(vma, addr, page, flags, pmdp);
+}
+EXPORT_SYMBOL(device_entry_fault);
+
static void pgmap_radix_release(struct resource *res)
{
resource_size_t key, align_start, align_size, align_end;
diff --git a/mm/Kconfig b/mm/Kconfig
index be0ee11..0a21411 100644
--- a/mm/Kconfig
+++ b/mm/Kconfig
@@ -704,6 +704,18 @@ config ZONE_DEVICE
If FS_DAX is enabled, then say Y.
+config DEVICE_UNADDRESSABLE
+ bool "Un-addressable device memory (GPU memory, ...)"
+ depends on ZONE_DEVICE
+
+ help
+ Allow to create struct page for un-addressable device memory
+ ie memory that is only accessible by the device (or group of
+ devices).
+
+ This allow to migrate chunk of process memory to device memory
+ while that memory is use by the device.
+
config FRAME_VECTOR
bool
diff --git a/mm/memory.c b/mm/memory.c
index 15f2908..a83d690 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -889,6 +889,21 @@ copy_one_pte(struct mm_struct *dst_mm, struct mm_struct *src_mm,
pte = pte_swp_mksoft_dirty(pte);
set_pte_at(src_mm, addr, src_pte, pte);
}
+ } else if (is_device_entry(entry)) {
+ page = device_entry_to_page(entry);
+
+ get_page(page);
+ rss[mm_counter(page)]++;
+ page_dup_rmap(page, false);
+
+ if (is_write_device_entry(entry) &&
+ is_cow_mapping(vm_flags)) {
+ make_device_entry_read(&entry);
+ pte = swp_entry_to_pte(entry);
+ if (pte_swp_soft_dirty(*src_pte))
+ pte = pte_swp_mksoft_dirty(pte);
+ set_pte_at(src_mm, addr, src_pte, pte);
+ }
}
goto out_set_pte;
}
@@ -1191,6 +1206,12 @@ again:
page = migration_entry_to_page(entry);
rss[mm_counter(page)]--;
+ } else if (is_device_entry(entry)) {
+ struct page *page = device_entry_to_page(entry);
+ rss[mm_counter(page)]--;
+
+ page_remove_rmap(page, false);
+ put_page(page);
}
if (unlikely(!free_swap_and_cache(entry)))
print_bad_pte(vma, addr, ptent, NULL);
@@ -2536,6 +2557,9 @@ int do_swap_page(struct fault_env *fe, pte_t orig_pte)
if (unlikely(non_swap_entry(entry))) {
if (is_migration_entry(entry)) {
migration_entry_wait(vma->vm_mm, fe->pmd, fe->address);
+ } else if (is_device_entry(entry)) {
+ ret = device_entry_fault(vma, fe->address, entry,
+ fe->flags, fe->pmd);
} else if (is_hwpoison_entry(entry)) {
ret = VM_FAULT_HWPOISON;
} else {
diff --git a/mm/mprotect.c b/mm/mprotect.c
index 1bc1eb3..70aff3a 100644
--- a/mm/mprotect.c
+++ b/mm/mprotect.c
@@ -139,6 +139,18 @@ static unsigned long change_pte_range(struct vm_area_struct *vma, pmd_t *pmd,
pages++;
}
+
+ if (is_write_device_entry(entry)) {
+ pte_t newpte;
+
+ make_device_entry_read(&entry);
+ newpte = swp_entry_to_pte(entry);
+ if (pte_swp_soft_dirty(oldpte))
+ newpte = pte_swp_mksoft_dirty(newpte);
+ set_pte_at(mm, addr, pte, newpte);
+
+ pages++;
+ }
}
} while (pte++, addr += PAGE_SIZE, addr != end);
arch_leave_lazy_mmu_mode();
--
2.4.3
[toc] | [prev] | [next] | [standalone]
| From | Balbir Singh <bsingharora@gmail.com> |
|---|---|
| Date | 2016-11-21 03:10 +0100 |
| Subject | Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable |
| Message-ID | <sFMdr-3fW-1@gated-at.bofh.it> |
| In reply to | #1525579 |
On 19/11/16 05:18, Jérôme Glisse wrote:
> To allow use of device un-addressable memory inside a process add a
> special swap type. Also add a new callback to handle page fault on
> such entry.
>
> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
> Cc: Dan Williams <dan.j.williams@intel.com>
> Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
> ---
> fs/proc/task_mmu.c | 10 +++++++-
> include/linux/memremap.h | 5 ++++
> include/linux/swap.h | 18 ++++++++++---
> include/linux/swapops.h | 67 ++++++++++++++++++++++++++++++++++++++++++++++++
> kernel/memremap.c | 14 ++++++++++
> mm/Kconfig | 12 +++++++++
> mm/memory.c | 24 +++++++++++++++++
> mm/mprotect.c | 12 +++++++++
> 8 files changed, 158 insertions(+), 4 deletions(-)
>
> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> index 6909582..0726d39 100644
> --- a/fs/proc/task_mmu.c
> +++ b/fs/proc/task_mmu.c
> @@ -544,8 +544,11 @@ static void smaps_pte_entry(pte_t *pte, unsigned long addr,
> } else {
> mss->swap_pss += (u64)PAGE_SIZE << PSS_SHIFT;
> }
> - } else if (is_migration_entry(swpent))
> + } else if (is_migration_entry(swpent)) {
> page = migration_entry_to_page(swpent);
> + } else if (is_device_entry(swpent)) {
> + page = device_entry_to_page(swpent);
> + }
So the reason there is a device swap entry for a page belonging to a user process is
that it is in the middle of migration or is it always that a swap entry represents
unaddressable memory belonging to a GPU device, but its tracked in the page table
entries of the process.
> } else if (unlikely(IS_ENABLED(CONFIG_SHMEM) && mss->check_shmem_swap
> && pte_none(*pte))) {
> page = find_get_entry(vma->vm_file->f_mapping,
> @@ -708,6 +711,8 @@ static int smaps_hugetlb_range(pte_t *pte, unsigned long hmask,
>
> if (is_migration_entry(swpent))
> page = migration_entry_to_page(swpent);
> + if (is_device_entry(swpent))
> + page = device_entry_to_page(swpent);
> }
> if (page) {
> int mapcount = page_mapcount(page);
> @@ -1191,6 +1196,9 @@ static pagemap_entry_t pte_to_pagemap_entry(struct pagemapread *pm,
> flags |= PM_SWAP;
> if (is_migration_entry(entry))
> page = migration_entry_to_page(entry);
> +
> + if (is_device_entry(entry))
> + page = device_entry_to_page(entry);
> }
>
> if (page && !PageAnon(page))
> diff --git a/include/linux/memremap.h b/include/linux/memremap.h
> index b6f03e9..d584c74 100644
> --- a/include/linux/memremap.h
> +++ b/include/linux/memremap.h
> @@ -47,6 +47,11 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
> */
> struct dev_pagemap {
> void (*free_devpage)(struct page *page, void *data);
> + int (*fault)(struct vm_area_struct *vma,
> + unsigned long addr,
> + struct page *page,
> + unsigned flags,
> + pmd_t *pmdp);
> struct vmem_altmap *altmap;
> const struct resource *res;
> struct percpu_ref *ref;
> diff --git a/include/linux/swap.h b/include/linux/swap.h
> index 7e553e1..599cb54 100644
> --- a/include/linux/swap.h
> +++ b/include/linux/swap.h
> @@ -50,6 +50,17 @@ static inline int current_is_kswapd(void)
> */
>
> /*
> + * Un-addressable device memory support
> + */
> +#ifdef CONFIG_DEVICE_UNADDRESSABLE
> +#define SWP_DEVICE_NUM 2
> +#define SWP_DEVICE_WRITE (MAX_SWAPFILES + SWP_HWPOISON_NUM + SWP_MIGRATION_NUM)
> +#define SWP_DEVICE (MAX_SWAPFILES + SWP_HWPOISON_NUM + SWP_MIGRATION_NUM + 1)
> +#else
> +#define SWP_DEVICE_NUM 0
> +#endif
> +
> +/*
> * NUMA node memory migration support
> */
> #ifdef CONFIG_MIGRATION
> @@ -71,7 +82,8 @@ static inline int current_is_kswapd(void)
> #endif
>
> #define MAX_SWAPFILES \
> - ((1 << MAX_SWAPFILES_SHIFT) - SWP_MIGRATION_NUM - SWP_HWPOISON_NUM)
> + ((1 << MAX_SWAPFILES_SHIFT) - SWP_DEVICE_NUM - \
> + SWP_MIGRATION_NUM - SWP_HWPOISON_NUM)
>
> /*
> * Magic header for a swap area. The first part of the union is
> @@ -442,8 +454,8 @@ static inline void show_swap_cache_info(void)
> {
> }
>
> -#define free_swap_and_cache(swp) is_migration_entry(swp)
> -#define swapcache_prepare(swp) is_migration_entry(swp)
> +#define free_swap_and_cache(e) (is_migration_entry(e) || is_device_entry(e))
> +#define swapcache_prepare(e) (is_migration_entry(e) || is_device_entry(e))
>
> static inline int add_swap_count_continuation(swp_entry_t swp, gfp_t gfp_mask)
> {
> diff --git a/include/linux/swapops.h b/include/linux/swapops.h
> index 5c3a5f3..d1aa425 100644
> --- a/include/linux/swapops.h
> +++ b/include/linux/swapops.h
> @@ -100,6 +100,73 @@ static inline void *swp_to_radix_entry(swp_entry_t entry)
> return (void *)(value | RADIX_TREE_EXCEPTIONAL_ENTRY);
> }
>
> +#ifdef CONFIG_DEVICE_UNADDRESSABLE
> +static inline swp_entry_t make_device_entry(struct page *page, bool write)
> +{
> + return swp_entry(write?SWP_DEVICE_WRITE:SWP_DEVICE, page_to_pfn(page));
Code style checks
> +}
> +
> +static inline bool is_device_entry(swp_entry_t entry)
> +{
> + int type = swp_type(entry);
> + return type == SWP_DEVICE || type == SWP_DEVICE_WRITE;
> +}
> +
> +static inline void make_device_entry_read(swp_entry_t *entry)
> +{
> + *entry = swp_entry(SWP_DEVICE, swp_offset(*entry));
> +}
> +
> +static inline bool is_write_device_entry(swp_entry_t entry)
> +{
> + return unlikely(swp_type(entry) == SWP_DEVICE_WRITE);
> +}
> +
> +static inline struct page *device_entry_to_page(swp_entry_t entry)
> +{
> + return pfn_to_page(swp_offset(entry));
> +}
> +
> +int device_entry_fault(struct vm_area_struct *vma,
> + unsigned long addr,
> + swp_entry_t entry,
> + unsigned flags,
> + pmd_t *pmdp);
> +#else /* CONFIG_DEVICE_UNADDRESSABLE */
> +static inline swp_entry_t make_device_entry(struct page *page, bool write)
> +{
> + return swp_entry(0, 0);
> +}
> +
> +static inline void make_device_entry_read(swp_entry_t *entry)
> +{
> +}
> +
> +static inline bool is_device_entry(swp_entry_t entry)
> +{
> + return false;
> +}
> +
> +static inline bool is_write_device_entry(swp_entry_t entry)
> +{
> + return false;
> +}
> +
> +static inline struct page *device_entry_to_page(swp_entry_t entry)
> +{
> + return NULL;
> +}
> +
> +static inline int device_entry_fault(struct vm_area_struct *vma,
> + unsigned long addr,
> + swp_entry_t entry,
> + unsigned flags,
> + pmd_t *pmdp)
> +{
> + return VM_FAULT_SIGBUS;
> +}
> +#endif /* CONFIG_DEVICE_UNADDRESSABLE */
> +
> #ifdef CONFIG_MIGRATION
> static inline swp_entry_t make_migration_entry(struct page *page, int write)
> {
> diff --git a/kernel/memremap.c b/kernel/memremap.c
> index cf83928..0670015 100644
> --- a/kernel/memremap.c
> +++ b/kernel/memremap.c
> @@ -18,6 +18,8 @@
> #include <linux/io.h>
> #include <linux/mm.h>
> #include <linux/memory_hotplug.h>
> +#include <linux/swap.h>
> +#include <linux/swapops.h>
>
> #ifndef ioremap_cache
> /* temporary while we convert existing ioremap_cache users to memremap */
> @@ -200,6 +202,18 @@ void put_zone_device_page(struct page *page)
> }
> EXPORT_SYMBOL(put_zone_device_page);
>
> +int device_entry_fault(struct vm_area_struct *vma,
> + unsigned long addr,
> + swp_entry_t entry,
> + unsigned flags,
> + pmd_t *pmdp)
> +{
> + struct page *page = device_entry_to_page(entry);
> +
> + return page->pgmap->fault(vma, addr, page, flags, pmdp);
> +}
> +EXPORT_SYMBOL(device_entry_fault);
> +
> static void pgmap_radix_release(struct resource *res)
> {
> resource_size_t key, align_start, align_size, align_end;
> diff --git a/mm/Kconfig b/mm/Kconfig
> index be0ee11..0a21411 100644
> --- a/mm/Kconfig
> +++ b/mm/Kconfig
> @@ -704,6 +704,18 @@ config ZONE_DEVICE
>
> If FS_DAX is enabled, then say Y.
>
> +config DEVICE_UNADDRESSABLE
> + bool "Un-addressable device memory (GPU memory, ...)"
> + depends on ZONE_DEVICE
> +
> + help
> + Allow to create struct page for un-addressable device memory
> + ie memory that is only accessible by the device (or group of
> + devices).
> +
> + This allow to migrate chunk of process memory to device memory
> + while that memory is use by the device.
> +
> config FRAME_VECTOR
> bool
>
> diff --git a/mm/memory.c b/mm/memory.c
> index 15f2908..a83d690 100644
> --- a/mm/memory.c
> +++ b/mm/memory.c
> @@ -889,6 +889,21 @@ copy_one_pte(struct mm_struct *dst_mm, struct mm_struct *src_mm,
> pte = pte_swp_mksoft_dirty(pte);
> set_pte_at(src_mm, addr, src_pte, pte);
> }
> + } else if (is_device_entry(entry)) {
> + page = device_entry_to_page(entry);
> +
> + get_page(page);
> + rss[mm_counter(page)]++;
Why does rss count go up?
> + page_dup_rmap(page, false);
> +
> + if (is_write_device_entry(entry) &&
> + is_cow_mapping(vm_flags)) {
> + make_device_entry_read(&entry);
> + pte = swp_entry_to_pte(entry);
> + if (pte_swp_soft_dirty(*src_pte))
> + pte = pte_swp_mksoft_dirty(pte);
> + set_pte_at(src_mm, addr, src_pte, pte);
> + }
> }
> goto out_set_pte;
> }
> @@ -1191,6 +1206,12 @@ again:
>
> page = migration_entry_to_page(entry);
> rss[mm_counter(page)]--;
> + } else if (is_device_entry(entry)) {
> + struct page *page = device_entry_to_page(entry);
> + rss[mm_counter(page)]--;
> +
> + page_remove_rmap(page, false);
> + put_page(page);
> }
> if (unlikely(!free_swap_and_cache(entry)))
> print_bad_pte(vma, addr, ptent, NULL);
> @@ -2536,6 +2557,9 @@ int do_swap_page(struct fault_env *fe, pte_t orig_pte)
> if (unlikely(non_swap_entry(entry))) {
> if (is_migration_entry(entry)) {
> migration_entry_wait(vma->vm_mm, fe->pmd, fe->address);
> + } else if (is_device_entry(entry)) {
> + ret = device_entry_fault(vma, fe->address, entry,
> + fe->flags, fe->pmd);
What does device_entry_fault() actually do here?
> } else if (is_hwpoison_entry(entry)) {
> ret = VM_FAULT_HWPOISON;
> } else {
> diff --git a/mm/mprotect.c b/mm/mprotect.c
> index 1bc1eb3..70aff3a 100644
> --- a/mm/mprotect.c
> +++ b/mm/mprotect.c
> @@ -139,6 +139,18 @@ static unsigned long change_pte_range(struct vm_area_struct *vma, pmd_t *pmd,
>
> pages++;
> }
> +
> + if (is_write_device_entry(entry)) {
> + pte_t newpte;
> +
> + make_device_entry_read(&entry);
> + newpte = swp_entry_to_pte(entry);
> + if (pte_swp_soft_dirty(oldpte))
> + newpte = pte_swp_mksoft_dirty(newpte);
> + set_pte_at(mm, addr, pte, newpte);
> +
> + pages++;
> + }
Does it make sense to call mprotect() on device memory ranges?
> }
> } while (pte++, addr += PAGE_SIZE, addr != end);
> arch_leave_lazy_mmu_mode();
>
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-21 06:10 +0100 |
| Subject | Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable |
| Message-ID | <sFP1D-59B-1@gated-at.bofh.it> |
| In reply to | #1526329 |
On Mon, Nov 21, 2016 at 01:06:45PM +1100, Balbir Singh wrote:
>
>
> On 19/11/16 05:18, Jérôme Glisse wrote:
> > To allow use of device un-addressable memory inside a process add a
> > special swap type. Also add a new callback to handle page fault on
> > such entry.
> >
> > Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
> > Cc: Dan Williams <dan.j.williams@intel.com>
> > Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
> > ---
> > fs/proc/task_mmu.c | 10 +++++++-
> > include/linux/memremap.h | 5 ++++
> > include/linux/swap.h | 18 ++++++++++---
> > include/linux/swapops.h | 67 ++++++++++++++++++++++++++++++++++++++++++++++++
> > kernel/memremap.c | 14 ++++++++++
> > mm/Kconfig | 12 +++++++++
> > mm/memory.c | 24 +++++++++++++++++
> > mm/mprotect.c | 12 +++++++++
> > 8 files changed, 158 insertions(+), 4 deletions(-)
> >
> > diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> > index 6909582..0726d39 100644
> > --- a/fs/proc/task_mmu.c
> > +++ b/fs/proc/task_mmu.c
> > @@ -544,8 +544,11 @@ static void smaps_pte_entry(pte_t *pte, unsigned long addr,
> > } else {
> > mss->swap_pss += (u64)PAGE_SIZE << PSS_SHIFT;
> > }
> > - } else if (is_migration_entry(swpent))
> > + } else if (is_migration_entry(swpent)) {
> > page = migration_entry_to_page(swpent);
> > + } else if (is_device_entry(swpent)) {
> > + page = device_entry_to_page(swpent);
> > + }
>
>
> So the reason there is a device swap entry for a page belonging to a user process is
> that it is in the middle of migration or is it always that a swap entry represents
> unaddressable memory belonging to a GPU device, but its tracked in the page table
> entries of the process.
For page being migrated i use the existing special migration pte entry. This new device
special swap entry is only for unaddressable memory belonging to a device (GPU or any
else). We need to keep track of those inside the CPU page table. Using a new special
swap entry is the easiest way with the minimum amount of change to core mm.
[...]
> > +#ifdef CONFIG_DEVICE_UNADDRESSABLE
> > +static inline swp_entry_t make_device_entry(struct page *page, bool write)
> > +{
> > + return swp_entry(write?SWP_DEVICE_WRITE:SWP_DEVICE, page_to_pfn(page));
>
> Code style checks
I was trying to balance against 79 columns break rule :)
[...]
> > + } else if (is_device_entry(entry)) {
> > + page = device_entry_to_page(entry);
> > +
> > + get_page(page);
> > + rss[mm_counter(page)]++;
>
> Why does rss count go up?
I wanted the device page to be treated like any other page. There is an argument
to be made against and for doing that. Do you have strong argument for not doing
this ?
[...]
> > @@ -2536,6 +2557,9 @@ int do_swap_page(struct fault_env *fe, pte_t orig_pte)
> > if (unlikely(non_swap_entry(entry))) {
> > if (is_migration_entry(entry)) {
> > migration_entry_wait(vma->vm_mm, fe->pmd, fe->address);
> > + } else if (is_device_entry(entry)) {
> > + ret = device_entry_fault(vma, fe->address, entry,
> > + fe->flags, fe->pmd);
>
> What does device_entry_fault() actually do here?
Well it is a special fault handler, it must migrate the memory back to some place
where the CPU can access it. It only matter for unaddressable memory.
> > } else if (is_hwpoison_entry(entry)) {
> > ret = VM_FAULT_HWPOISON;
> > } else {
> > diff --git a/mm/mprotect.c b/mm/mprotect.c
> > index 1bc1eb3..70aff3a 100644
> > --- a/mm/mprotect.c
> > +++ b/mm/mprotect.c
> > @@ -139,6 +139,18 @@ static unsigned long change_pte_range(struct vm_area_struct *vma, pmd_t *pmd,
> >
> > pages++;
> > }
> > +
> > + if (is_write_device_entry(entry)) {
> > + pte_t newpte;
> > +
> > + make_device_entry_read(&entry);
> > + newpte = swp_entry_to_pte(entry);
> > + if (pte_swp_soft_dirty(oldpte))
> > + newpte = pte_swp_mksoft_dirty(newpte);
> > + set_pte_at(mm, addr, pte, newpte);
> > +
> > + pages++;
> > + }
>
> Does it make sense to call mprotect() on device memory ranges?
There is nothing special about vma that containt device memory. They can be
private anonymous, share, file back ... So any existing memory syscall must
behave as expected. This is really just like any other page except that CPU
can not access it.
Cheers,
Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Balbir Singh <bsingharora@gmail.com> |
|---|---|
| Date | 2016-11-22 03:30 +0100 |
| Subject | Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable |
| Message-ID | <sG90l-14h-1@gated-at.bofh.it> |
| In reply to | #1526361 |
On 21/11/16 16:05, Jerome Glisse wrote:
> On Mon, Nov 21, 2016 at 01:06:45PM +1100, Balbir Singh wrote:
>>
>>
>> On 19/11/16 05:18, Jérôme Glisse wrote:
>>> To allow use of device un-addressable memory inside a process add a
>>> special swap type. Also add a new callback to handle page fault on
>>> such entry.
>>>
>>> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
>>> Cc: Dan Williams <dan.j.williams@intel.com>
>>> Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
>>> ---
>>> fs/proc/task_mmu.c | 10 +++++++-
>>> include/linux/memremap.h | 5 ++++
>>> include/linux/swap.h | 18 ++++++++++---
>>> include/linux/swapops.h | 67 ++++++++++++++++++++++++++++++++++++++++++++++++
>>> kernel/memremap.c | 14 ++++++++++
>>> mm/Kconfig | 12 +++++++++
>>> mm/memory.c | 24 +++++++++++++++++
>>> mm/mprotect.c | 12 +++++++++
>>> 8 files changed, 158 insertions(+), 4 deletions(-)
>>>
>>> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
>>> index 6909582..0726d39 100644
>>> --- a/fs/proc/task_mmu.c
>>> +++ b/fs/proc/task_mmu.c
>>> @@ -544,8 +544,11 @@ static void smaps_pte_entry(pte_t *pte, unsigned long addr,
>>> } else {
>>> mss->swap_pss += (u64)PAGE_SIZE << PSS_SHIFT;
>>> }
>>> - } else if (is_migration_entry(swpent))
>>> + } else if (is_migration_entry(swpent)) {
>>> page = migration_entry_to_page(swpent);
>>> + } else if (is_device_entry(swpent)) {
>>> + page = device_entry_to_page(swpent);
>>> + }
>>
>>
>> So the reason there is a device swap entry for a page belonging to a user process is
>> that it is in the middle of migration or is it always that a swap entry represents
>> unaddressable memory belonging to a GPU device, but its tracked in the page table
>> entries of the process.
>
> For page being migrated i use the existing special migration pte entry. This new device
> special swap entry is only for unaddressable memory belonging to a device (GPU or any
> else). We need to keep track of those inside the CPU page table. Using a new special
> swap entry is the easiest way with the minimum amount of change to core mm.
>
Thanks, makes sense
> [...]
>
>>> +#ifdef CONFIG_DEVICE_UNADDRESSABLE
>>> +static inline swp_entry_t make_device_entry(struct page *page, bool write)
>>> +{
>>> + return swp_entry(write?SWP_DEVICE_WRITE:SWP_DEVICE, page_to_pfn(page));
>>
>> Code style checks
>
> I was trying to balance against 79 columns break rule :)
>
> [...]
>
>>> + } else if (is_device_entry(entry)) {
>>> + page = device_entry_to_page(entry);
>>> +
>>> + get_page(page);
>>> + rss[mm_counter(page)]++;
>>
>> Why does rss count go up?
>
> I wanted the device page to be treated like any other page. There is an argument
> to be made against and for doing that. Do you have strong argument for not doing
> this ?
>
Yes, It will end up confusing rss accounting IMHO. If a task is using a lot of
pages on the GPU, should be it a candidate for OOM based on it's RSS for example?
> [...]
>
>>> @@ -2536,6 +2557,9 @@ int do_swap_page(struct fault_env *fe, pte_t orig_pte)
>>> if (unlikely(non_swap_entry(entry))) {
>>> if (is_migration_entry(entry)) {
>>> migration_entry_wait(vma->vm_mm, fe->pmd, fe->address);
>>> + } else if (is_device_entry(entry)) {
>>> + ret = device_entry_fault(vma, fe->address, entry,
>>> + fe->flags, fe->pmd);
>>
>> What does device_entry_fault() actually do here?
>
> Well it is a special fault handler, it must migrate the memory back to some place
> where the CPU can access it. It only matter for unaddressable memory.
So effectively swap the page back in, chances are it can ping pong ...but I was wondering if we can
tell the GPU that the CPU is accessing these pages as well. I presume any operation that causes
memory access - core dump will swap back in things from the HMM side onto the CPU side.
>
>>> } else if (is_hwpoison_entry(entry)) {
>>> ret = VM_FAULT_HWPOISON;
>>> } else {
>>> diff --git a/mm/mprotect.c b/mm/mprotect.c
>>> index 1bc1eb3..70aff3a 100644
>>> --- a/mm/mprotect.c
>>> +++ b/mm/mprotect.c
>>> @@ -139,6 +139,18 @@ static unsigned long change_pte_range(struct vm_area_struct *vma, pmd_t *pmd,
>>>
>>> pages++;
>>> }
>>> +
>>> + if (is_write_device_entry(entry)) {
>>> + pte_t newpte;
>>> +
>>> + make_device_entry_read(&entry);
>>> + newpte = swp_entry_to_pte(entry);
>>> + if (pte_swp_soft_dirty(oldpte))
>>> + newpte = pte_swp_mksoft_dirty(newpte);
>>> + set_pte_at(mm, addr, pte, newpte);
>>> +
>>> + pages++;
>>> + }
>>
>> Does it make sense to call mprotect() on device memory ranges?
>
> There is nothing special about vma that containt device memory. They can be
> private anonymous, share, file back ... So any existing memory syscall must
> behave as expected. This is really just like any other page except that CPU
> can not access it.
I understand that, but what would marking it as R/O when the GPU is in the middle
of write mean? I would also worry about passing "executable" pages over to the
other side.
Balbir Singh.
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-22 15:00 +0100 |
| Subject | Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable |
| Message-ID | <sGjM6-7OO-15@gated-at.bofh.it> |
| In reply to | #1527187 |
On Tue, Nov 22, 2016 at 01:19:42PM +1100, Balbir Singh wrote:
>
>
> On 21/11/16 16:05, Jerome Glisse wrote:
> > On Mon, Nov 21, 2016 at 01:06:45PM +1100, Balbir Singh wrote:
> >>
> >>
> >> On 19/11/16 05:18, Jérôme Glisse wrote:
> >>> To allow use of device un-addressable memory inside a process add a
> >>> special swap type. Also add a new callback to handle page fault on
> >>> such entry.
> >>>
> >>> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
> >>> Cc: Dan Williams <dan.j.williams@intel.com>
> >>> Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
> >>> ---
> >>> fs/proc/task_mmu.c | 10 +++++++-
> >>> include/linux/memremap.h | 5 ++++
> >>> include/linux/swap.h | 18 ++++++++++---
> >>> include/linux/swapops.h | 67 ++++++++++++++++++++++++++++++++++++++++++++++++
> >>> kernel/memremap.c | 14 ++++++++++
> >>> mm/Kconfig | 12 +++++++++
> >>> mm/memory.c | 24 +++++++++++++++++
> >>> mm/mprotect.c | 12 +++++++++
> >>> 8 files changed, 158 insertions(+), 4 deletions(-)
> >>>
> >>> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> >>> index 6909582..0726d39 100644
> >>> --- a/fs/proc/task_mmu.c
> >>> +++ b/fs/proc/task_mmu.c
> >>> @@ -544,8 +544,11 @@ static void smaps_pte_entry(pte_t *pte, unsigned long addr,
> >>> } else {
> >>> mss->swap_pss += (u64)PAGE_SIZE << PSS_SHIFT;
> >>> }
> >>> - } else if (is_migration_entry(swpent))
> >>> + } else if (is_migration_entry(swpent)) {
> >>> page = migration_entry_to_page(swpent);
> >>> + } else if (is_device_entry(swpent)) {
> >>> + page = device_entry_to_page(swpent);
> >>> + }
> >>
> >>
> >> So the reason there is a device swap entry for a page belonging to a user process is
> >> that it is in the middle of migration or is it always that a swap entry represents
> >> unaddressable memory belonging to a GPU device, but its tracked in the page table
> >> entries of the process.
> >
> > For page being migrated i use the existing special migration pte entry. This new device
> > special swap entry is only for unaddressable memory belonging to a device (GPU or any
> > else). We need to keep track of those inside the CPU page table. Using a new special
> > swap entry is the easiest way with the minimum amount of change to core mm.
> >
>
> Thanks, makes sense
>
> > [...]
> >
> >>> +#ifdef CONFIG_DEVICE_UNADDRESSABLE
> >>> +static inline swp_entry_t make_device_entry(struct page *page, bool write)
> >>> +{
> >>> + return swp_entry(write?SWP_DEVICE_WRITE:SWP_DEVICE, page_to_pfn(page));
> >>
> >> Code style checks
> >
> > I was trying to balance against 79 columns break rule :)
> >
> > [...]
> >
> >>> + } else if (is_device_entry(entry)) {
> >>> + page = device_entry_to_page(entry);
> >>> +
> >>> + get_page(page);
> >>> + rss[mm_counter(page)]++;
> >>
> >> Why does rss count go up?
> >
> > I wanted the device page to be treated like any other page. There is an argument
> > to be made against and for doing that. Do you have strong argument for not doing
> > this ?
> >
>
> Yes, It will end up confusing rss accounting IMHO. If a task is using a lot of
> pages on the GPU, should be it a candidate for OOM based on it's RSS for example?
>
> > [...]
> >
> >>> @@ -2536,6 +2557,9 @@ int do_swap_page(struct fault_env *fe, pte_t orig_pte)
> >>> if (unlikely(non_swap_entry(entry))) {
> >>> if (is_migration_entry(entry)) {
> >>> migration_entry_wait(vma->vm_mm, fe->pmd, fe->address);
> >>> + } else if (is_device_entry(entry)) {
> >>> + ret = device_entry_fault(vma, fe->address, entry,
> >>> + fe->flags, fe->pmd);
> >>
> >> What does device_entry_fault() actually do here?
> >
> > Well it is a special fault handler, it must migrate the memory back to some place
> > where the CPU can access it. It only matter for unaddressable memory.
>
> So effectively swap the page back in, chances are it can ping pong ...but I was wondering if we can
> tell the GPU that the CPU is accessing these pages as well. I presume any operation that causes
> memory access - core dump will swap back in things from the HMM side onto the CPU side.
Well it is up to device driver to gather statistic on what can/should be inside device memory.
My expectation is that they will detect ping pong and stop asking to migrate a given address/
range to device memory.
>
> >
> >>> } else if (is_hwpoison_entry(entry)) {
> >>> ret = VM_FAULT_HWPOISON;
> >>> } else {
> >>> diff --git a/mm/mprotect.c b/mm/mprotect.c
> >>> index 1bc1eb3..70aff3a 100644
> >>> --- a/mm/mprotect.c
> >>> +++ b/mm/mprotect.c
> >>> @@ -139,6 +139,18 @@ static unsigned long change_pte_range(struct vm_area_struct *vma, pmd_t *pmd,
> >>>
> >>> pages++;
> >>> }
> >>> +
> >>> + if (is_write_device_entry(entry)) {
> >>> + pte_t newpte;
> >>> +
> >>> + make_device_entry_read(&entry);
> >>> + newpte = swp_entry_to_pte(entry);
> >>> + if (pte_swp_soft_dirty(oldpte))
> >>> + newpte = pte_swp_mksoft_dirty(newpte);
> >>> + set_pte_at(mm, addr, pte, newpte);
> >>> +
> >>> + pages++;
> >>> + }
> >>
> >> Does it make sense to call mprotect() on device memory ranges?
> >
> > There is nothing special about vma that containt device memory. They can be
> > private anonymous, share, file back ... So any existing memory syscall must
> > behave as expected. This is really just like any other page except that CPU
> > can not access it.
>
> I understand that, but what would marking it as R/O when the GPU is in the middle
> of write mean? I would also worry about passing "executable" pages over to the
> other side.
>
Any memory protection change will trigger an mmu_notifier calls which in turn will
update the device page table accordingly. So R/O status will also happen on the GPU.
We assume here that the device driver is not doing evil thing and that device driver
obey memory protection for all range it mirrors. Upstream driver are easy to check.
Close driver might be more problematic, in NVidia case this part if open source and
is easily checkable.
Cheers,
Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-21 12:20 +0100 |
| Subject | Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable |
| Message-ID | <sFUNI-sG-21@gated-at.bofh.it> |
| In reply to | #1526329 |
On 11/21/2016 07:36 AM, Balbir Singh wrote:
>
>
> On 19/11/16 05:18, Jérôme Glisse wrote:
>> To allow use of device un-addressable memory inside a process add a
>> special swap type. Also add a new callback to handle page fault on
>> such entry.
>>
>> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
>> Cc: Dan Williams <dan.j.williams@intel.com>
>> Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
>> ---
>> fs/proc/task_mmu.c | 10 +++++++-
>> include/linux/memremap.h | 5 ++++
>> include/linux/swap.h | 18 ++++++++++---
>> include/linux/swapops.h | 67 ++++++++++++++++++++++++++++++++++++++++++++++++
>> kernel/memremap.c | 14 ++++++++++
>> mm/Kconfig | 12 +++++++++
>> mm/memory.c | 24 +++++++++++++++++
>> mm/mprotect.c | 12 +++++++++
>> 8 files changed, 158 insertions(+), 4 deletions(-)
>>
>> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
>> index 6909582..0726d39 100644
>> --- a/fs/proc/task_mmu.c
>> +++ b/fs/proc/task_mmu.c
>> @@ -544,8 +544,11 @@ static void smaps_pte_entry(pte_t *pte, unsigned long addr,
>> } else {
>> mss->swap_pss += (u64)PAGE_SIZE << PSS_SHIFT;
>> }
>> - } else if (is_migration_entry(swpent))
>> + } else if (is_migration_entry(swpent)) {
>> page = migration_entry_to_page(swpent);
>> + } else if (is_device_entry(swpent)) {
>> + page = device_entry_to_page(swpent);
>> + }
>
>
> So the reason there is a device swap entry for a page belonging to a user process is
> that it is in the middle of migration or is it always that a swap entry represents
> unaddressable memory belonging to a GPU device, but its tracked in the page table
> entries of the process.
I guess the later is the case and its used for the page table mirroring
purpose after intercepting the page faults. But will leave upto Jerome
to explain more on this.
>
>> } else if (unlikely(IS_ENABLED(CONFIG_SHMEM) && mss->check_shmem_swap
>> && pte_none(*pte))) {
>> page = find_get_entry(vma->vm_file->f_mapping,
>> @@ -708,6 +711,8 @@ static int smaps_hugetlb_range(pte_t *pte, unsigned long hmask,
>>
>> if (is_migration_entry(swpent))
>> page = migration_entry_to_page(swpent);
>> + if (is_device_entry(swpent))
>> + page = device_entry_to_page(swpent);
>> }
>> if (page) {
>> int mapcount = page_mapcount(page);
>> @@ -1191,6 +1196,9 @@ static pagemap_entry_t pte_to_pagemap_entry(struct pagemapread *pm,
>> flags |= PM_SWAP;
>> if (is_migration_entry(entry))
>> page = migration_entry_to_page(entry);
>> +
>> + if (is_device_entry(entry))
>> + page = device_entry_to_page(entry);
>> }
>>
>> if (page && !PageAnon(page))
>> diff --git a/include/linux/memremap.h b/include/linux/memremap.h
>> index b6f03e9..d584c74 100644
>> --- a/include/linux/memremap.h
>> +++ b/include/linux/memremap.h
>> @@ -47,6 +47,11 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
>> */
>> struct dev_pagemap {
>> void (*free_devpage)(struct page *page, void *data);
>> + int (*fault)(struct vm_area_struct *vma,
>> + unsigned long addr,
>> + struct page *page,
>> + unsigned flags,
>> + pmd_t *pmdp);
>> struct vmem_altmap *altmap;
>> const struct resource *res;
>> struct percpu_ref *ref;
>> diff --git a/include/linux/swap.h b/include/linux/swap.h
>> index 7e553e1..599cb54 100644
>> --- a/include/linux/swap.h
>> +++ b/include/linux/swap.h
>> @@ -50,6 +50,17 @@ static inline int current_is_kswapd(void)
>> */
>>
>> /*
>> + * Un-addressable device memory support
>> + */
>> +#ifdef CONFIG_DEVICE_UNADDRESSABLE
>> +#define SWP_DEVICE_NUM 2
>> +#define SWP_DEVICE_WRITE (MAX_SWAPFILES + SWP_HWPOISON_NUM + SWP_MIGRATION_NUM)
>> +#define SWP_DEVICE (MAX_SWAPFILES + SWP_HWPOISON_NUM + SWP_MIGRATION_NUM + 1)
>> +#else
>> +#define SWP_DEVICE_NUM 0
>> +#endif
>> +
>> +/*
>> * NUMA node memory migration support
>> */
>> #ifdef CONFIG_MIGRATION
>> @@ -71,7 +82,8 @@ static inline int current_is_kswapd(void)
>> #endif
>>
>> #define MAX_SWAPFILES \
>> - ((1 << MAX_SWAPFILES_SHIFT) - SWP_MIGRATION_NUM - SWP_HWPOISON_NUM)
>> + ((1 << MAX_SWAPFILES_SHIFT) - SWP_DEVICE_NUM - \
>> + SWP_MIGRATION_NUM - SWP_HWPOISON_NUM)
>>
>> /*
>> * Magic header for a swap area. The first part of the union is
>> @@ -442,8 +454,8 @@ static inline void show_swap_cache_info(void)
>> {
>> }
>>
>> -#define free_swap_and_cache(swp) is_migration_entry(swp)
>> -#define swapcache_prepare(swp) is_migration_entry(swp)
>> +#define free_swap_and_cache(e) (is_migration_entry(e) || is_device_entry(e))
>> +#define swapcache_prepare(e) (is_migration_entry(e) || is_device_entry(e))
>>
>> static inline int add_swap_count_continuation(swp_entry_t swp, gfp_t gfp_mask)
>> {
>> diff --git a/include/linux/swapops.h b/include/linux/swapops.h
>> index 5c3a5f3..d1aa425 100644
>> --- a/include/linux/swapops.h
>> +++ b/include/linux/swapops.h
>> @@ -100,6 +100,73 @@ static inline void *swp_to_radix_entry(swp_entry_t entry)
>> return (void *)(value | RADIX_TREE_EXCEPTIONAL_ENTRY);
>> }
>>
>> +#ifdef CONFIG_DEVICE_UNADDRESSABLE
>> +static inline swp_entry_t make_device_entry(struct page *page, bool write)
>> +{
>> + return swp_entry(write?SWP_DEVICE_WRITE:SWP_DEVICE, page_to_pfn(page));
>
> Code style checks
>
>> +}
>> +
>> +static inline bool is_device_entry(swp_entry_t entry)
>> +{
>> + int type = swp_type(entry);
>> + return type == SWP_DEVICE || type == SWP_DEVICE_WRITE;
>> +}
>> +
>> +static inline void make_device_entry_read(swp_entry_t *entry)
>> +{
>> + *entry = swp_entry(SWP_DEVICE, swp_offset(*entry));
>> +}
>> +
>> +static inline bool is_write_device_entry(swp_entry_t entry)
>> +{
>> + return unlikely(swp_type(entry) == SWP_DEVICE_WRITE);
>> +}
>> +
>> +static inline struct page *device_entry_to_page(swp_entry_t entry)
>> +{
>> + return pfn_to_page(swp_offset(entry));
>> +}
>> +
>> +int device_entry_fault(struct vm_area_struct *vma,
>> + unsigned long addr,
>> + swp_entry_t entry,
>> + unsigned flags,
>> + pmd_t *pmdp);
>> +#else /* CONFIG_DEVICE_UNADDRESSABLE */
>> +static inline swp_entry_t make_device_entry(struct page *page, bool write)
>> +{
>> + return swp_entry(0, 0);
>> +}
>> +
>> +static inline void make_device_entry_read(swp_entry_t *entry)
>> +{
>> +}
>> +
>> +static inline bool is_device_entry(swp_entry_t entry)
>> +{
>> + return false;
>> +}
>> +
>> +static inline bool is_write_device_entry(swp_entry_t entry)
>> +{
>> + return false;
>> +}
>> +
>> +static inline struct page *device_entry_to_page(swp_entry_t entry)
>> +{
>> + return NULL;
>> +}
>> +
>> +static inline int device_entry_fault(struct vm_area_struct *vma,
>> + unsigned long addr,
>> + swp_entry_t entry,
>> + unsigned flags,
>> + pmd_t *pmdp)
>> +{
>> + return VM_FAULT_SIGBUS;
>> +}
>> +#endif /* CONFIG_DEVICE_UNADDRESSABLE */
>> +
>> #ifdef CONFIG_MIGRATION
>> static inline swp_entry_t make_migration_entry(struct page *page, int write)
>> {
>> diff --git a/kernel/memremap.c b/kernel/memremap.c
>> index cf83928..0670015 100644
>> --- a/kernel/memremap.c
>> +++ b/kernel/memremap.c
>> @@ -18,6 +18,8 @@
>> #include <linux/io.h>
>> #include <linux/mm.h>
>> #include <linux/memory_hotplug.h>
>> +#include <linux/swap.h>
>> +#include <linux/swapops.h>
>>
>> #ifndef ioremap_cache
>> /* temporary while we convert existing ioremap_cache users to memremap */
>> @@ -200,6 +202,18 @@ void put_zone_device_page(struct page *page)
>> }
>> EXPORT_SYMBOL(put_zone_device_page);
>>
>> +int device_entry_fault(struct vm_area_struct *vma,
>> + unsigned long addr,
>> + swp_entry_t entry,
>> + unsigned flags,
>> + pmd_t *pmdp)
>> +{
>> + struct page *page = device_entry_to_page(entry);
>> +
>> + return page->pgmap->fault(vma, addr, page, flags, pmdp);
>> +}
>> +EXPORT_SYMBOL(device_entry_fault);
>> +
>> static void pgmap_radix_release(struct resource *res)
>> {
>> resource_size_t key, align_start, align_size, align_end;
>> diff --git a/mm/Kconfig b/mm/Kconfig
>> index be0ee11..0a21411 100644
>> --- a/mm/Kconfig
>> +++ b/mm/Kconfig
>> @@ -704,6 +704,18 @@ config ZONE_DEVICE
>>
>> If FS_DAX is enabled, then say Y.
>>
>> +config DEVICE_UNADDRESSABLE
>> + bool "Un-addressable device memory (GPU memory, ...)"
>> + depends on ZONE_DEVICE
>> +
>> + help
>> + Allow to create struct page for un-addressable device memory
>> + ie memory that is only accessible by the device (or group of
>> + devices).
>> +
>> + This allow to migrate chunk of process memory to device memory
>> + while that memory is use by the device.
>> +
>> config FRAME_VECTOR
>> bool
>>
>> diff --git a/mm/memory.c b/mm/memory.c
>> index 15f2908..a83d690 100644
>> --- a/mm/memory.c
>> +++ b/mm/memory.c
>> @@ -889,6 +889,21 @@ copy_one_pte(struct mm_struct *dst_mm, struct mm_struct *src_mm,
>> pte = pte_swp_mksoft_dirty(pte);
>> set_pte_at(src_mm, addr, src_pte, pte);
>> }
>> + } else if (is_device_entry(entry)) {
>> + page = device_entry_to_page(entry);
>> +
>> + get_page(page);
>> + rss[mm_counter(page)]++;
>
> Why does rss count go up?
>
>> + page_dup_rmap(page, false);
>> +
>> + if (is_write_device_entry(entry) &&
>> + is_cow_mapping(vm_flags)) {
>> + make_device_entry_read(&entry);
>> + pte = swp_entry_to_pte(entry);
>> + if (pte_swp_soft_dirty(*src_pte))
>> + pte = pte_swp_mksoft_dirty(pte);
>> + set_pte_at(src_mm, addr, src_pte, pte);
>> + }
>> }
>> goto out_set_pte;
>> }
>> @@ -1191,6 +1206,12 @@ again:
>>
>> page = migration_entry_to_page(entry);
>> rss[mm_counter(page)]--;
>> + } else if (is_device_entry(entry)) {
>> + struct page *page = device_entry_to_page(entry);
>> + rss[mm_counter(page)]--;
>> +
>> + page_remove_rmap(page, false);
>> + put_page(page);
>> }
>> if (unlikely(!free_swap_and_cache(entry)))
>> print_bad_pte(vma, addr, ptent, NULL);
>> @@ -2536,6 +2557,9 @@ int do_swap_page(struct fault_env *fe, pte_t orig_pte)
>> if (unlikely(non_swap_entry(entry))) {
>> if (is_migration_entry(entry)) {
>> migration_entry_wait(vma->vm_mm, fe->pmd, fe->address);
>> + } else if (is_device_entry(entry)) {
>> + ret = device_entry_fault(vma, fe->address, entry,
>> + fe->flags, fe->pmd);
>
> What does device_entry_fault() actually do here?
IIUC it calls page->pgmap->fault() which is device specific page fault for
the page and thats how the control reaches device driver from the core VM.
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-21 12:00 +0100 |
| Subject | Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable |
| Message-ID | <sFUul-7g-7@gated-at.bofh.it> |
| In reply to | #1525579 |
On 11/18/2016 11:48 PM, Jérôme Glisse wrote:
> To allow use of device un-addressable memory inside a process add a
> special swap type. Also add a new callback to handle page fault on
> such entry.
IIUC this swap type is required only for the mirror cases and its
not a requirement for migration. If it's required for mirroring
purpose where we intercept each page fault, the commit message
here should clearly elaborate on that more.
>
> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
> Cc: Dan Williams <dan.j.williams@intel.com>
> Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
> ---
> fs/proc/task_mmu.c | 10 +++++++-
> include/linux/memremap.h | 5 ++++
> include/linux/swap.h | 18 ++++++++++---
> include/linux/swapops.h | 67 ++++++++++++++++++++++++++++++++++++++++++++++++
> kernel/memremap.c | 14 ++++++++++
> mm/Kconfig | 12 +++++++++
> mm/memory.c | 24 +++++++++++++++++
> mm/mprotect.c | 12 +++++++++
> 8 files changed, 158 insertions(+), 4 deletions(-)
>
> diff --git a/fs/proc/task_mmu.c b/fs/proc/task_mmu.c
> index 6909582..0726d39 100644
> --- a/fs/proc/task_mmu.c
> +++ b/fs/proc/task_mmu.c
> @@ -544,8 +544,11 @@ static void smaps_pte_entry(pte_t *pte, unsigned long addr,
> } else {
> mss->swap_pss += (u64)PAGE_SIZE << PSS_SHIFT;
> }
> - } else if (is_migration_entry(swpent))
> + } else if (is_migration_entry(swpent)) {
> page = migration_entry_to_page(swpent);
> + } else if (is_device_entry(swpent)) {
> + page = device_entry_to_page(swpent);
> + }
> } else if (unlikely(IS_ENABLED(CONFIG_SHMEM) && mss->check_shmem_swap
> && pte_none(*pte))) {
> page = find_get_entry(vma->vm_file->f_mapping,
> @@ -708,6 +711,8 @@ static int smaps_hugetlb_range(pte_t *pte, unsigned long hmask,
>
> if (is_migration_entry(swpent))
> page = migration_entry_to_page(swpent);
> + if (is_device_entry(swpent))
> + page = device_entry_to_page(swpent);
> }
> if (page) {
> int mapcount = page_mapcount(page);
> @@ -1191,6 +1196,9 @@ static pagemap_entry_t pte_to_pagemap_entry(struct pagemapread *pm,
> flags |= PM_SWAP;
> if (is_migration_entry(entry))
> page = migration_entry_to_page(entry);
> +
> + if (is_device_entry(entry))
> + page = device_entry_to_page(entry);
> }
>
> if (page && !PageAnon(page))
> diff --git a/include/linux/memremap.h b/include/linux/memremap.h
> index b6f03e9..d584c74 100644
> --- a/include/linux/memremap.h
> +++ b/include/linux/memremap.h
> @@ -47,6 +47,11 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
> */
> struct dev_pagemap {
> void (*free_devpage)(struct page *page, void *data);
> + int (*fault)(struct vm_area_struct *vma,
> + unsigned long addr,
> + struct page *page,
> + unsigned flags,
> + pmd_t *pmdp);
We are extending the dev_pagemap once again to accommodate device driver
specific fault routines for these pages. Wondering if this extension and
the new swap type should be in the same patch.
> struct vmem_altmap *altmap;
> const struct resource *res;
> struct percpu_ref *ref;
> diff --git a/include/linux/swap.h b/include/linux/swap.h
> index 7e553e1..599cb54 100644
> --- a/include/linux/swap.h
> +++ b/include/linux/swap.h
> @@ -50,6 +50,17 @@ static inline int current_is_kswapd(void)
> */
>
> /*
> + * Un-addressable device memory support
> + */
> +#ifdef CONFIG_DEVICE_UNADDRESSABLE
> +#define SWP_DEVICE_NUM 2
> +#define SWP_DEVICE_WRITE (MAX_SWAPFILES + SWP_HWPOISON_NUM + SWP_MIGRATION_NUM)
> +#define SWP_DEVICE (MAX_SWAPFILES + SWP_HWPOISON_NUM + SWP_MIGRATION_NUM + 1)
> +#else
> +#define SWP_DEVICE_NUM 0
> +#endif
> +
> +/*
> * NUMA node memory migration support
> */
> #ifdef CONFIG_MIGRATION
> @@ -71,7 +82,8 @@ static inline int current_is_kswapd(void)
> #endif
>
> #define MAX_SWAPFILES \
> - ((1 << MAX_SWAPFILES_SHIFT) - SWP_MIGRATION_NUM - SWP_HWPOISON_NUM)
> + ((1 << MAX_SWAPFILES_SHIFT) - SWP_DEVICE_NUM - \
> + SWP_MIGRATION_NUM - SWP_HWPOISON_NUM)
>
> /*
> * Magic header for a swap area. The first part of the union is
> @@ -442,8 +454,8 @@ static inline void show_swap_cache_info(void)
> {
> }
>
> -#define free_swap_and_cache(swp) is_migration_entry(swp)
> -#define swapcache_prepare(swp) is_migration_entry(swp)
> +#define free_swap_and_cache(e) (is_migration_entry(e) || is_device_entry(e))
> +#define swapcache_prepare(e) (is_migration_entry(e) || is_device_entry(e))
>
> static inline int add_swap_count_continuation(swp_entry_t swp, gfp_t gfp_mask)
> {
> diff --git a/include/linux/swapops.h b/include/linux/swapops.h
> index 5c3a5f3..d1aa425 100644
> --- a/include/linux/swapops.h
> +++ b/include/linux/swapops.h
> @@ -100,6 +100,73 @@ static inline void *swp_to_radix_entry(swp_entry_t entry)
> return (void *)(value | RADIX_TREE_EXCEPTIONAL_ENTRY);
> }
>
> +#ifdef CONFIG_DEVICE_UNADDRESSABLE
> +static inline swp_entry_t make_device_entry(struct page *page, bool write)
> +{
> + return swp_entry(write?SWP_DEVICE_WRITE:SWP_DEVICE, page_to_pfn(page));
> +}
> +
> +static inline bool is_device_entry(swp_entry_t entry)
> +{
> + int type = swp_type(entry);
> + return type == SWP_DEVICE || type == SWP_DEVICE_WRITE;
> +}
> +
> +static inline void make_device_entry_read(swp_entry_t *entry)
> +{
> + *entry = swp_entry(SWP_DEVICE, swp_offset(*entry));
> +}
> +
> +static inline bool is_write_device_entry(swp_entry_t entry)
> +{
> + return unlikely(swp_type(entry) == SWP_DEVICE_WRITE);
> +}
> +
> +static inline struct page *device_entry_to_page(swp_entry_t entry)
> +{
> + return pfn_to_page(swp_offset(entry));
> +}
> +
> +int device_entry_fault(struct vm_area_struct *vma,
> + unsigned long addr,
> + swp_entry_t entry,
> + unsigned flags,
> + pmd_t *pmdp);
> +#else /* CONFIG_DEVICE_UNADDRESSABLE */
> +static inline swp_entry_t make_device_entry(struct page *page, bool write)
> +{
> + return swp_entry(0, 0);
> +}
> +
> +static inline void make_device_entry_read(swp_entry_t *entry)
> +{
> +}
> +
> +static inline bool is_device_entry(swp_entry_t entry)
> +{
> + return false;
> +}
> +
> +static inline bool is_write_device_entry(swp_entry_t entry)
> +{
> + return false;
> +}
> +
> +static inline struct page *device_entry_to_page(swp_entry_t entry)
> +{
> + return NULL;
> +}
> +
> +static inline int device_entry_fault(struct vm_area_struct *vma,
> + unsigned long addr,
> + swp_entry_t entry,
> + unsigned flags,
> + pmd_t *pmdp)
> +{
> + return VM_FAULT_SIGBUS;
> +}
> +#endif /* CONFIG_DEVICE_UNADDRESSABLE */
> +
> #ifdef CONFIG_MIGRATION
> static inline swp_entry_t make_migration_entry(struct page *page, int write)
> {
> diff --git a/kernel/memremap.c b/kernel/memremap.c
> index cf83928..0670015 100644
> --- a/kernel/memremap.c
> +++ b/kernel/memremap.c
> @@ -18,6 +18,8 @@
> #include <linux/io.h>
> #include <linux/mm.h>
> #include <linux/memory_hotplug.h>
> +#include <linux/swap.h>
> +#include <linux/swapops.h>
>
> #ifndef ioremap_cache
> /* temporary while we convert existing ioremap_cache users to memremap */
> @@ -200,6 +202,18 @@ void put_zone_device_page(struct page *page)
> }
> EXPORT_SYMBOL(put_zone_device_page);
>
> +int device_entry_fault(struct vm_area_struct *vma,
> + unsigned long addr,
> + swp_entry_t entry,
> + unsigned flags,
> + pmd_t *pmdp)
> +{
> + struct page *page = device_entry_to_page(entry);
> +
A BUG_ON() if page->pgmap->fault has not been populated by the driver.
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-21 13:50 +0100 |
| Subject | Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable |
| Message-ID | <sFWcN-1f8-5@gated-at.bofh.it> |
| In reply to | #1526564 |
On Mon, Nov 21, 2016 at 04:28:04PM +0530, Anshuman Khandual wrote:
> On 11/18/2016 11:48 PM, Jérôme Glisse wrote:
> > To allow use of device un-addressable memory inside a process add a
> > special swap type. Also add a new callback to handle page fault on
> > such entry.
>
> IIUC this swap type is required only for the mirror cases and its
> not a requirement for migration. If it's required for mirroring
> purpose where we intercept each page fault, the commit message
> here should clearly elaborate on that more.
It is only require for un-addressable memory. The mirroring has nothing to do
with it. I will clarify commit message.
[...]
> > diff --git a/include/linux/memremap.h b/include/linux/memremap.h
> > index b6f03e9..d584c74 100644
> > --- a/include/linux/memremap.h
> > +++ b/include/linux/memremap.h
> > @@ -47,6 +47,11 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
> > */
> > struct dev_pagemap {
> > void (*free_devpage)(struct page *page, void *data);
> > + int (*fault)(struct vm_area_struct *vma,
> > + unsigned long addr,
> > + struct page *page,
> > + unsigned flags,
> > + pmd_t *pmdp);
>
> We are extending the dev_pagemap once again to accommodate device driver
> specific fault routines for these pages. Wondering if this extension and
> the new swap type should be in the same patch.
It make sense to have it in one single patch as i also change page fault code
path to deal with the new special swap entry and those make use of this new
callback.
> > +int device_entry_fault(struct vm_area_struct *vma,
> > + unsigned long addr,
> > + swp_entry_t entry,
> > + unsigned flags,
> > + pmd_t *pmdp)
> > +{
> > + struct page *page = device_entry_to_page(entry);
> > +
>
> A BUG_ON() if page->pgmap->fault has not been populated by the driver.
>
Ok
Cheers,
Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-22 05:50 +0100 |
| Subject | Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable |
| Message-ID | <sGbbP-2qk-3@gated-at.bofh.it> |
| In reply to | #1526637 |
On 11/21/2016 06:12 PM, Jerome Glisse wrote:
> On Mon, Nov 21, 2016 at 04:28:04PM +0530, Anshuman Khandual wrote:
>> On 11/18/2016 11:48 PM, Jérôme Glisse wrote:
>>> To allow use of device un-addressable memory inside a process add a
>>> special swap type. Also add a new callback to handle page fault on
>>> such entry.
>>
>> IIUC this swap type is required only for the mirror cases and its
>> not a requirement for migration. If it's required for mirroring
>> purpose where we intercept each page fault, the commit message
>> here should clearly elaborate on that more.
>
> It is only require for un-addressable memory. The mirroring has nothing to do
> with it. I will clarify commit message.
One thing though. I dont recall how persistent memory ZONE_DEVICE
pages are handled inside the page tables, point here is it should
be part of the same code block. We should catch that its a device
memory page and then figure out addressable or not and act
accordingly. Because persistent memory are CPU addressable, there
might not been special code block but dealing with device pages
should be handled in a more holistic manner.
>
> [...]
>
>>> diff --git a/include/linux/memremap.h b/include/linux/memremap.h
>>> index b6f03e9..d584c74 100644
>>> --- a/include/linux/memremap.h
>>> +++ b/include/linux/memremap.h
>>> @@ -47,6 +47,11 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
>>> */
>>> struct dev_pagemap {
>>> void (*free_devpage)(struct page *page, void *data);
>>> + int (*fault)(struct vm_area_struct *vma,
>>> + unsigned long addr,
>>> + struct page *page,
>>> + unsigned flags,
>>> + pmd_t *pmdp);
>>
>> We are extending the dev_pagemap once again to accommodate device driver
>> specific fault routines for these pages. Wondering if this extension and
>> the new swap type should be in the same patch.
>
> It make sense to have it in one single patch as i also change page fault code
> path to deal with the new special swap entry and those make use of this new
> callback.
>
Okay.
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-24 15:00 +0100 |
| Subject | Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable |
| Message-ID | <sH2Jc-3tF-29@gated-at.bofh.it> |
| In reply to | #1527222 |
On Tue, Nov 22, 2016 at 10:18:27AM +0530, Anshuman Khandual wrote: > On 11/21/2016 06:12 PM, Jerome Glisse wrote: > > On Mon, Nov 21, 2016 at 04:28:04PM +0530, Anshuman Khandual wrote: > >> On 11/18/2016 11:48 PM, Jérôme Glisse wrote: > >>> To allow use of device un-addressable memory inside a process add a > >>> special swap type. Also add a new callback to handle page fault on > >>> such entry. > >> > >> IIUC this swap type is required only for the mirror cases and its > >> not a requirement for migration. If it's required for mirroring > >> purpose where we intercept each page fault, the commit message > >> here should clearly elaborate on that more. > > > > It is only require for un-addressable memory. The mirroring has nothing to do > > with it. I will clarify commit message. > > One thing though. I dont recall how persistent memory ZONE_DEVICE > pages are handled inside the page tables, point here is it should > be part of the same code block. We should catch that its a device > memory page and then figure out addressable or not and act > accordingly. Because persistent memory are CPU addressable, there > might not been special code block but dealing with device pages > should be handled in a more holistic manner. Before i repost updated patchset i should stress that dealing with un-addressable device page and addressable one in same block is not do-able without re-doing once again the whole mm page fault code path. Because i use special swap entry the logical place for me to handle it is with where swap entry are handled. Regular device page are handle bit simpler that other page because they can't be evicted/swaped so they are always present once faulted. I think right now they are always populated through fs page fault callback (well dax one). So not much reasons to consolidate all device page handling in one place. We are looking at different use case in the end. Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-18 18:30 +0100 |
| Subject | [HMM v13 11/18] mm/hmm/mirror: add range monitor helper, to monitor CPU page table update |
| Message-ID | <sEV98-1Es-39@gated-at.bofh.it> |
| In reply to | #1525568 |
Complement the hmm_vma_range_lock/unlock() mechanism with a range monitor that do
not block CPU page table invalidation and thus do not garanty forward progress. It
is still usefull as in many situations concurrent CPU page table update and CPU
snapshot are taking place in different region of the virtual address space.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Signed-off-by: Jatin Kumar <jakumar@nvidia.com>
Signed-off-by: John Hubbard <jhubbard@nvidia.com>
Signed-off-by: Mark Hairgrove <mhairgrove@nvidia.com>
Signed-off-by: Sherry Cheung <SCheung@nvidia.com>
Signed-off-by: Subhash Gutti <sgutti@nvidia.com>
---
include/linux/hmm.h | 18 ++++++++++
mm/hmm.c | 95 ++++++++++++++++++++++++++++++++++++++++++++++++++++-
2 files changed, 112 insertions(+), 1 deletion(-)
diff --git a/include/linux/hmm.h b/include/linux/hmm.h
index c0b1c07..6571647 100644
--- a/include/linux/hmm.h
+++ b/include/linux/hmm.h
@@ -254,6 +254,24 @@ int hmm_vma_range_lock(struct hmm_range *range,
void hmm_vma_range_unlock(struct hmm_range *range);
+/*
+ * Monitoring a range allow to track any CPU page table modification that can
+ * affect the range. It complements the hmm_vma_range_lock/unlock() mechanism
+ * as a non blocking method for synchronizing device page table with the CPU
+ * page table. See functions description in mm/hmm.c for documentation.
+ *
+ * NOTE AFTER A CALL TO hmm_vma_range_monitor_start() THAT RETURNED TRUE YOU
+ * MUST MAKE A CALL TO hmm_vma_range_monitor_end() BEFORE FREEING THE RANGE
+ * STRUCT OR BAD THING WILL HAPPEN !
+ */
+bool hmm_vma_range_monitor_start(struct hmm_range *range,
+ struct vm_area_struct *vma,
+ unsigned long start,
+ unsigned long end,
+ bool wait);
+bool hmm_vma_range_monitor_end(struct hmm_range *range);
+
+
/* Below are for HMM internal use only ! Not to be use by device driver ! */
void hmm_mm_destroy(struct mm_struct *mm);
diff --git a/mm/hmm.c b/mm/hmm.c
index ee05419..746eb96 100644
--- a/mm/hmm.c
+++ b/mm/hmm.c
@@ -40,6 +40,7 @@ struct hmm {
spinlock_t lock;
struct list_head ranges;
struct list_head mirrors;
+ struct list_head monitors;
atomic_t sequence;
wait_queue_head_t wait_queue;
struct mmu_notifier mmu_notifier;
@@ -65,6 +66,7 @@ static struct hmm *hmm_register(struct mm_struct *mm)
return NULL;
init_waitqueue_head(&hmm->wait_queue);
atomic_set(&hmm->notifier_count, 0);
+ INIT_LIST_HEAD(&hmm->monitors);
INIT_LIST_HEAD(&hmm->mirrors);
atomic_set(&hmm->sequence, 0);
hmm->mmu_notifier.ops = NULL;
@@ -112,7 +114,7 @@ static void hmm_invalidate_range(struct hmm *hmm,
unsigned long start,
unsigned long end)
{
- struct hmm_range range, *tmp;
+ struct hmm_range range, *tmp, *next;
struct hmm_mirror *mirror;
/*
@@ -127,6 +129,13 @@ static void hmm_invalidate_range(struct hmm *hmm,
range.hmm = hmm;
spin_lock(&hmm->lock);
+ /* Remove any range monitors */
+ list_for_each_entry_safe (tmp, next, &hmm->monitors, list) {
+ if (range.start >= tmp->end || range.end <= tmp->start)
+ continue;
+ /* This range is no longer valid */
+ list_del_init(&tmp->list);
+ }
list_for_each_entry (tmp, &hmm->ranges, list) {
if (range.start >= tmp->end || range.end <= tmp->start)
continue;
@@ -361,3 +370,87 @@ void hmm_vma_range_unlock(struct hmm_range *range)
wake_up(&hmm->wait_queue);
}
EXPORT_SYMBOL(hmm_vma_range_unlock);
+
+
+/*
+ * hmm_vma_range_monitor_start() - start monitoring of a range
+ * @range: pointer to hmm_range struct use to monitor
+ * @vma: virtual memory area for the range
+ * @start: start address of the range to monitor (inclusive)
+ * @end: end address of the range to monitor (exclusive)
+ * @wait: wait for any pending CPU page table to finish
+ * Returns: false if there is pendding CPU page table update, true otherwise
+ *
+ * The use pattern of this function is :
+ * retry:
+ * hmm_vma_range_monitor_start(range, vma, start, end, true);
+ * // Do something that rely on stable CPU page table content but do not
+ * // Prepare device page table update transaction
+ * ...
+ * // Take device driver lock that serialize device page table update
+ * driver_lock_device_page_table_update();
+ * if (!hmm_vma_range_monitor_end(range)) {
+ * driver_unlock_device_page_table_update();
+ * // Abort transaction you just build and cleanup anything that need
+ * // to be. Same comment as above, about avoiding busy loop.
+ * goto retry;
+ * }
+ * // Commit device page table update
+ * driver_unlock_device_page_table_update();
+ */
+bool hmm_vma_range_monitor_start(struct hmm_range *range,
+ struct vm_area_struct *vma,
+ unsigned long start,
+ unsigned long end,
+ bool wait)
+{
+ BUG_ON(!vma);
+ BUG_ON(!range);
+
+ INIT_LIST_HEAD(&range->list);
+ range->hmm = hmm_register(vma->vm_mm);
+ if (!range->hmm)
+ return false;
+
+again:
+ spin_lock(&range->hmm->lock);
+ if (atomic_read(&range->hmm->notifier_count)) {
+ spin_unlock(&range->hmm->lock);
+ if (!wait)
+ return false;
+ /*
+ * FIXME: Wait for all active mmu_notifier this is because we
+ * can no keep an hmm_range struct around while waiting for
+ * range invalidation to finish. Need to update mmu_notifier
+ * to make this doable.
+ */
+ wait_event(range->hmm->wait_queue,
+ !atomic_read(&range->hmm->notifier_count));
+ goto again;
+ }
+ list_add_tail(&range->list, &range->hmm->monitors);
+ spin_unlock(&range->hmm->lock);
+ return true;
+}
+EXPORT_SYMBOL(hmm_vma_range_monitor_start);
+
+/*
+ * hmm_vma_range_monitor_end() - end monitoring of a range
+ * @range: range that was being monitored
+ * Returns: true if no invalidation since hmm_vma_range_monitor_start()
+ */
+bool hmm_vma_range_monitor_end(struct hmm_range *range)
+{
+ bool valid;
+
+ if (!range->hmm || list_empty(&range->list))
+ return false;
+
+ spin_lock(&range->hmm->lock);
+ valid = !list_empty(&range->list);
+ list_del_init(&range->list);
+ spin_unlock(&range->hmm->lock);
+
+ return valid;
+}
+EXPORT_SYMBOL(hmm_vma_range_monitor_end);
--
2.4.3
[toc] | [prev] | [next] | [standalone]
| From | John Hubbard <jhubbard@nvidia.com> |
|---|---|
| Date | 2016-11-19 01:50 +0100 |
| Message-ID | <sF20V-65c-9@gated-at.bofh.it> |
| In reply to | #1525568 |
[Multipart message — attachments visible in raw view] — view raw
On Fri, 18 Nov 2016, Jérôme Glisse wrote:
> Cliff note: HMM offers 2 things (each standing on its own). First
> it allows to use device memory transparently inside any process
> without any modifications to process program code. Second it allows
> to mirror process address space on a device.
>
> Change since v12 is the use of struct page for device memory even if
> the device memory is not accessible by the CPU (because of limitation
> impose by the bus between the CPU and the device).
>
> Using struct page means that their are minimal changes to core mm
> code. HMM build on top of ZONE_DEVICE to provide struct page, it
> adds new features to ZONE_DEVICE. The first 7 patches implement
> those changes.
>
> Rest of patchset is divided into 3 features that can each be use
> independently from one another. First is the process address space
> mirroring (patch 9 to 13), this allow to snapshot CPU page table
> and to keep the device page table synchronize with the CPU one.
>
> Second is a new memory migration helper which allow migration of
> a range of virtual address of a process. This memory migration
> also allow device to use their own DMA engine to perform the copy
> between the source memory and destination memory. This can be
> usefull even outside HMM context in many usecase.
>
> Third part of the patchset (patch 17-18) is a set of helper to
> register a ZONE_DEVICE node and manage it. It is meant as a
> convenient helper so that device drivers do not each have to
> reimplement over and over the same boiler plate code.
>
>
> I am hoping that this can now be consider for inclusion upstream.
> Bottom line is that without HMM we can not support some of the new
> hardware features on x86 PCIE. I do believe we need some solution
> to support those features or we won't be able to use such hardware
> in standard like C++17, OpenCL 3.0 and others.
>
> I have been working with NVidia to bring up this feature on their
> Pascal GPU. There are real hardware that you can buy today that
> could benefit from HMM. We also intend to leverage this inside the
> open source nouveau driver.
>
Hi,
We (NVIDIA engineering) have been working closely with Jerome on this for
several years now, and I wanted to mention that NVIDIA is committed to
using HMM. We've done initial testing of this patchset on Pascal GPUs (a
bit more detail below) and it is looking good.
The HMM features are a prerequisite to an important part of NVIDIA's
efforts to make writing code for GPUs (and other page-faulting devices)
easier--by making it more like writing code for CPUs. A big part of that
story involves being able to use malloc'd memory transparently everywhere.
Here's a tiny example (in case it's not obvious from the HMM patchset
documentation) of HMM in action:
int *p = (int*)malloc(SIZE); *p = 5; /* on the CPU */
x = *p; /* on a GPU, or on any page-fault-capable device */
1. A device page fault occurs because the malloc'd memory was never
allocated in the device's page tables.
2. The device driver receives a page fault interrupt, but fails to
recognize the address, so it calls into HMM.
3. HMM knows that p is valid on the CPU, and coordinates with the device
driver to unmap the CPU page, allocate a page on the device, and then
migrate (copy) the data to the device. This allows full device memory
bandwidth to be available, which is critical to getting good performance.
a) Alternatively, leave the page on the CPU, and create a device
PTE to point to that page. This might be done if our performance counters
show that a page is thrashing.
4. The device driver issues a replay-page-fault to the device.
5. The device program continues running, and x == 5 now.
When version 1 of this patchset was created (2.5 years ago! in May, 2014),
one huge concern was that we didn't yet have hardware that could use it.
But now we do: Pascal GPUs, which have been shipping this year, all
support replayable page faults.
Testing:
We have done some testing of this latest patchset on Pascal GPUs using our
nvidia-uvm.ko module (which is open source, separate from the closed
source nvidia.ko). There is still much more testing to do, of course, but
basic page mirroring and page migration (between CPU and GPU), and even
some multi-GPU cases, are all working.
We do think we've found a bug in a corner case that involves invalid GPU
memory (of course, it's always possible that the bug is on our side),
which Jerome is investigating now. If you spot the bug by inspection,
you'll get some major told-you-so points. :)
The performance is looking good on the testing we’ve done so far, too.
thanks,
John Hubbard
NVIDIA Systems Software Engineer
>
> In this patchset i restricted myself to set of core features what
> is missing:
> - force read only on CPU for memory duplication and GPU atomic
> - changes to mmu_notifier for optimization purposes
> - migration of file back page to device memory
>
> I plan to submit a couple more patchset to implement those feature
> once core HMM is upstream.
>
>
> Is there anything blocking HMM inclusion ? Something fundamental ?
>
>
> Previous patchset posting :
> v1 http://lwn.net/Articles/597289/
> v2 https://lkml.org/lkml/2014/6/12/559
> v3 https://lkml.org/lkml/2014/6/13/633
> v4 https://lkml.org/lkml/2014/8/29/423
> v5 https://lkml.org/lkml/2014/11/3/759
> v6 http://lwn.net/Articles/619737/
> v7 http://lwn.net/Articles/627316/
> v8 https://lwn.net/Articles/645515/
> v9 https://lwn.net/Articles/651553/
> v10 https://lwn.net/Articles/654430/
> v11 http://www.gossamer-threads.com/lists/linux/kernel/2286424
> v12 http://www.kernelhub.org/?msg=972982&p=2
>
> Cheers,
> Jérôme
>
> Jérôme Glisse (18):
> mm/memory/hotplug: convert device parameter bool to set of flags
> mm/ZONE_DEVICE/unaddressable: add support for un-addressable device
> memory
> mm/ZONE_DEVICE/free_hot_cold_page: catch ZONE_DEVICE pages
> mm/ZONE_DEVICE/free-page: callback when page is freed
> mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device
> memory
> mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable
> mm/ZONE_DEVICE/x86: add support for un-addressable device memory
> mm/hmm: heterogeneous memory management (HMM for short)
> mm/hmm/mirror: mirror process address space on device with HMM helpers
> mm/hmm/mirror: add range lock helper, prevent CPU page table update
> for the range
> mm/hmm/mirror: add range monitor helper, to monitor CPU page table
> update
> mm/hmm/mirror: helper to snapshot CPU page table
> mm/hmm/mirror: device page fault handler
> mm/hmm/migrate: support un-addressable ZONE_DEVICE page in migration
> mm/hmm/migrate: add new boolean copy flag to migratepage() callback
> mm/hmm/migrate: new memory migration helper for use with device memory
> mm/hmm/devmem: device driver helper to hotplug ZONE_DEVICE memory
> mm/hmm/devmem: dummy HMM device as an helper for ZONE_DEVICE memory
>
> MAINTAINERS | 7 +
> arch/ia64/mm/init.c | 19 +-
> arch/powerpc/mm/mem.c | 18 +-
> arch/s390/mm/init.c | 10 +-
> arch/sh/mm/init.c | 18 +-
> arch/tile/mm/init.c | 10 +-
> arch/x86/mm/init_32.c | 19 +-
> arch/x86/mm/init_64.c | 23 +-
> drivers/dax/pmem.c | 3 +-
> drivers/nvdimm/pmem.c | 5 +-
> drivers/staging/lustre/lustre/llite/rw26.c | 8 +-
> fs/aio.c | 7 +-
> fs/btrfs/disk-io.c | 11 +-
> fs/hugetlbfs/inode.c | 9 +-
> fs/nfs/internal.h | 5 +-
> fs/nfs/write.c | 9 +-
> fs/proc/task_mmu.c | 10 +-
> fs/ubifs/file.c | 8 +-
> include/linux/balloon_compaction.h | 3 +-
> include/linux/fs.h | 13 +-
> include/linux/hmm.h | 516 ++++++++++++
> include/linux/memory_hotplug.h | 17 +-
> include/linux/memremap.h | 39 +-
> include/linux/migrate.h | 7 +-
> include/linux/mm_types.h | 5 +
> include/linux/swap.h | 18 +-
> include/linux/swapops.h | 67 ++
> kernel/fork.c | 2 +
> kernel/memremap.c | 48 +-
> mm/Kconfig | 23 +
> mm/Makefile | 1 +
> mm/balloon_compaction.c | 2 +-
> mm/hmm.c | 1175 ++++++++++++++++++++++++++++
> mm/memory.c | 33 +
> mm/memory_hotplug.c | 4 +-
> mm/migrate.c | 651 ++++++++++++++-
> mm/mprotect.c | 12 +
> mm/page_alloc.c | 10 +
> mm/rmap.c | 47 ++
> tools/testing/nvdimm/test/iomap.c | 2 +-
> 40 files changed, 2811 insertions(+), 83 deletions(-)
> create mode 100644 include/linux/hmm.h
> create mode 100644 mm/hmm.c
>
> --
> 2.4.3
>
>
[toc] | [prev] | [next] | [standalone]
| From | "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-19 16:00 +0100 |
| Message-ID | <sFfhv-6fj-5@gated-at.bofh.it> |
| In reply to | #1525826 |
John Hubbard <jhubbard@nvidia.com> writes: > On Fri, 18 Nov 2016, Jérôme Glisse wrote: > >> Cliff note: HMM offers 2 things (each standing on its own). First >> it allows to use device memory transparently inside any process >> without any modifications to process program code. Second it allows >> to mirror process address space on a device. >> >> Change since v12 is the use of struct page for device memory even if >> the device memory is not accessible by the CPU (because of limitation >> impose by the bus between the CPU and the device). >> >> Using struct page means that their are minimal changes to core mm >> code. HMM build on top of ZONE_DEVICE to provide struct page, it >> adds new features to ZONE_DEVICE. The first 7 patches implement >> those changes. >> >> Rest of patchset is divided into 3 features that can each be use >> independently from one another. First is the process address space >> mirroring (patch 9 to 13), this allow to snapshot CPU page table >> and to keep the device page table synchronize with the CPU one. >> >> Second is a new memory migration helper which allow migration of >> a range of virtual address of a process. This memory migration >> also allow device to use their own DMA engine to perform the copy >> between the source memory and destination memory. This can be >> usefull even outside HMM context in many usecase. >> >> Third part of the patchset (patch 17-18) is a set of helper to >> register a ZONE_DEVICE node and manage it. It is meant as a >> convenient helper so that device drivers do not each have to >> reimplement over and over the same boiler plate code. >> >> >> I am hoping that this can now be consider for inclusion upstream. >> Bottom line is that without HMM we can not support some of the new >> hardware features on x86 PCIE. I do believe we need some solution >> to support those features or we won't be able to use such hardware >> in standard like C++17, OpenCL 3.0 and others. >> >> I have been working with NVidia to bring up this feature on their >> Pascal GPU. There are real hardware that you can buy today that >> could benefit from HMM. We also intend to leverage this inside the >> open source nouveau driver. >> > > Hi, > > We (NVIDIA engineering) have been working closely with Jerome on this for > several years now, and I wanted to mention that NVIDIA is committed to > using HMM. We've done initial testing of this patchset on Pascal GPUs (a > bit more detail below) and it is looking good. > This can also be used on IBM platforms like Minsky ( http://www.tomshardware.com/news/ibm-power8-nvidia-tesla-p100-minsky,32661.html ) There is also discussion around using this for device accelerated page migration. That can help with coherent device memory node work. (https://lkml.kernel.org/r/1477283517-2504-1-git-send-email-khandual@linux.vnet.ibm.com) -aneesh
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web