Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1525568 > unrolled thread
| Started by | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| First post | 2016-11-18 18:20 +0100 |
| Last post | 2016-11-19 16:00 +0100 |
| Articles | 20 on this page of 34 — 6 participants |
Back to article view | Back to linux.kernel
[HMM v13 00/18] HMM (Heterogeneous Memory Management) v13 Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:20 +0100
[HMM v13 13/18] mm/hmm/mirror: device page fault handler Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:20 +0100
[HMM v13 12/18] mm/hmm/mirror: helper to snapshot CPU page table Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
[HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-21 09:10 +0100
Re: [HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory Jerome Glisse <jglisse@redhat.com> - 2016-11-21 13:40 +0100
Re: [HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-22 06:20 +0100
[HMM v13 09/18] mm/hmm/mirror: mirror process address space on device with HMM helpers Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 09/18] mm/hmm/mirror: mirror process address space on device with HMM helpers Balbir Singh <bsingharora@gmail.com> - 2016-11-21 03:50 +0100
Re: [HMM v13 09/18] mm/hmm/mirror: mirror process address space on device with HMM helpers Jerome Glisse <jglisse@redhat.com> - 2016-11-21 06:20 +0100
[HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-21 11:40 +0100
Re: [HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory Jerome Glisse <jglisse@redhat.com> - 2016-11-21 13:40 +0100
Re: [HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-22 06:00 +0100
[HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Balbir Singh <bsingharora@gmail.com> - 2016-11-21 03:00 +0100
Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Jerome Glisse <jglisse@redhat.com> - 2016-11-21 06:00 +0100
Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-21 09:30 +0100
Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Jerome Glisse <jglisse@redhat.com> - 2016-11-21 13:40 +0100
Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-22 06:10 +0100
[HMM v13 10/18] mm/hmm/mirror: add range lock helper, prevent CPU page table update for the range Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
[HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Balbir Singh <bsingharora@gmail.com> - 2016-11-21 03:10 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Jerome Glisse <jglisse@redhat.com> - 2016-11-21 06:10 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Balbir Singh <bsingharora@gmail.com> - 2016-11-22 03:30 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Jerome Glisse <jglisse@redhat.com> - 2016-11-22 15:00 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-21 12:20 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-21 12:00 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Jerome Glisse <jglisse@redhat.com> - 2016-11-21 13:50 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Anshuman Khandual <khandual@linux.vnet.ibm.com> - 2016-11-22 05:50 +0100
Re: [HMM v13 06/18] mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable Jerome Glisse <jglisse@redhat.com> - 2016-11-24 15:00 +0100
[HMM v13 11/18] mm/hmm/mirror: add range monitor helper, to monitor CPU page table update Jérôme Glisse <jglisse@redhat.com> - 2016-11-18 18:30 +0100
Re: [HMM v13 00/18] HMM (Heterogeneous Memory Management) v13 John Hubbard <jhubbard@nvidia.com> - 2016-11-19 01:50 +0100
Re: [HMM v13 00/18] HMM (Heterogeneous Memory Management) v13 "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-11-19 16:00 +0100
Page 1 of 2 [1] 2 Next page →
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-18 18:20 +0100 |
| Subject | [HMM v13 00/18] HMM (Heterogeneous Memory Management) v13 |
| Message-ID | <sEUZr-1B7-3@gated-at.bofh.it> |
Cliff note: HMM offers 2 things (each standing on its own). First
it allows to use device memory transparently inside any process
without any modifications to process program code. Second it allows
to mirror process address space on a device.
Change since v12 is the use of struct page for device memory even if
the device memory is not accessible by the CPU (because of limitation
impose by the bus between the CPU and the device).
Using struct page means that their are minimal changes to core mm
code. HMM build on top of ZONE_DEVICE to provide struct page, it
adds new features to ZONE_DEVICE. The first 7 patches implement
those changes.
Rest of patchset is divided into 3 features that can each be use
independently from one another. First is the process address space
mirroring (patch 9 to 13), this allow to snapshot CPU page table
and to keep the device page table synchronize with the CPU one.
Second is a new memory migration helper which allow migration of
a range of virtual address of a process. This memory migration
also allow device to use their own DMA engine to perform the copy
between the source memory and destination memory. This can be
usefull even outside HMM context in many usecase.
Third part of the patchset (patch 17-18) is a set of helper to
register a ZONE_DEVICE node and manage it. It is meant as a
convenient helper so that device drivers do not each have to
reimplement over and over the same boiler plate code.
I am hoping that this can now be consider for inclusion upstream.
Bottom line is that without HMM we can not support some of the new
hardware features on x86 PCIE. I do believe we need some solution
to support those features or we won't be able to use such hardware
in standard like C++17, OpenCL 3.0 and others.
I have been working with NVidia to bring up this feature on their
Pascal GPU. There are real hardware that you can buy today that
could benefit from HMM. We also intend to leverage this inside the
open source nouveau driver.
In this patchset i restricted myself to set of core features what
is missing:
- force read only on CPU for memory duplication and GPU atomic
- changes to mmu_notifier for optimization purposes
- migration of file back page to device memory
I plan to submit a couple more patchset to implement those feature
once core HMM is upstream.
Is there anything blocking HMM inclusion ? Something fundamental ?
Previous patchset posting :
v1 http://lwn.net/Articles/597289/
v2 https://lkml.org/lkml/2014/6/12/559
v3 https://lkml.org/lkml/2014/6/13/633
v4 https://lkml.org/lkml/2014/8/29/423
v5 https://lkml.org/lkml/2014/11/3/759
v6 http://lwn.net/Articles/619737/
v7 http://lwn.net/Articles/627316/
v8 https://lwn.net/Articles/645515/
v9 https://lwn.net/Articles/651553/
v10 https://lwn.net/Articles/654430/
v11 http://www.gossamer-threads.com/lists/linux/kernel/2286424
v12 http://www.kernelhub.org/?msg=972982&p=2
Cheers,
Jérôme
Jérôme Glisse (18):
mm/memory/hotplug: convert device parameter bool to set of flags
mm/ZONE_DEVICE/unaddressable: add support for un-addressable device
memory
mm/ZONE_DEVICE/free_hot_cold_page: catch ZONE_DEVICE pages
mm/ZONE_DEVICE/free-page: callback when page is freed
mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device
memory
mm/ZONE_DEVICE/unaddressable: add special swap for unaddressable
mm/ZONE_DEVICE/x86: add support for un-addressable device memory
mm/hmm: heterogeneous memory management (HMM for short)
mm/hmm/mirror: mirror process address space on device with HMM helpers
mm/hmm/mirror: add range lock helper, prevent CPU page table update
for the range
mm/hmm/mirror: add range monitor helper, to monitor CPU page table
update
mm/hmm/mirror: helper to snapshot CPU page table
mm/hmm/mirror: device page fault handler
mm/hmm/migrate: support un-addressable ZONE_DEVICE page in migration
mm/hmm/migrate: add new boolean copy flag to migratepage() callback
mm/hmm/migrate: new memory migration helper for use with device memory
mm/hmm/devmem: device driver helper to hotplug ZONE_DEVICE memory
mm/hmm/devmem: dummy HMM device as an helper for ZONE_DEVICE memory
MAINTAINERS | 7 +
arch/ia64/mm/init.c | 19 +-
arch/powerpc/mm/mem.c | 18 +-
arch/s390/mm/init.c | 10 +-
arch/sh/mm/init.c | 18 +-
arch/tile/mm/init.c | 10 +-
arch/x86/mm/init_32.c | 19 +-
arch/x86/mm/init_64.c | 23 +-
drivers/dax/pmem.c | 3 +-
drivers/nvdimm/pmem.c | 5 +-
drivers/staging/lustre/lustre/llite/rw26.c | 8 +-
fs/aio.c | 7 +-
fs/btrfs/disk-io.c | 11 +-
fs/hugetlbfs/inode.c | 9 +-
fs/nfs/internal.h | 5 +-
fs/nfs/write.c | 9 +-
fs/proc/task_mmu.c | 10 +-
fs/ubifs/file.c | 8 +-
include/linux/balloon_compaction.h | 3 +-
include/linux/fs.h | 13 +-
include/linux/hmm.h | 516 ++++++++++++
include/linux/memory_hotplug.h | 17 +-
include/linux/memremap.h | 39 +-
include/linux/migrate.h | 7 +-
include/linux/mm_types.h | 5 +
include/linux/swap.h | 18 +-
include/linux/swapops.h | 67 ++
kernel/fork.c | 2 +
kernel/memremap.c | 48 +-
mm/Kconfig | 23 +
mm/Makefile | 1 +
mm/balloon_compaction.c | 2 +-
mm/hmm.c | 1175 ++++++++++++++++++++++++++++
mm/memory.c | 33 +
mm/memory_hotplug.c | 4 +-
mm/migrate.c | 651 ++++++++++++++-
mm/mprotect.c | 12 +
mm/page_alloc.c | 10 +
mm/rmap.c | 47 ++
tools/testing/nvdimm/test/iomap.c | 2 +-
40 files changed, 2811 insertions(+), 83 deletions(-)
create mode 100644 include/linux/hmm.h
create mode 100644 mm/hmm.c
--
2.4.3
[toc] | [next] | [standalone]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-18 18:20 +0100 |
| Subject | [HMM v13 13/18] mm/hmm/mirror: device page fault handler |
| Message-ID | <sEUZs-1B7-45@gated-at.bofh.it> |
| In reply to | #1525568 |
This handle page fault on behalf of device driver, unlike handle_mm_fault()
it does not trigger migration back to system memory for device memory.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Signed-off-by: Jatin Kumar <jakumar@nvidia.com>
Signed-off-by: John Hubbard <jhubbard@nvidia.com>
Signed-off-by: Mark Hairgrove <mhairgrove@nvidia.com>
Signed-off-by: Sherry Cheung <SCheung@nvidia.com>
Signed-off-by: Subhash Gutti <sgutti@nvidia.com>
---
include/linux/hmm.h | 33 ++++++-
mm/hmm.c | 262 +++++++++++++++++++++++++++++++++++++++++++++++-----
2 files changed, 267 insertions(+), 28 deletions(-)
diff --git a/include/linux/hmm.h b/include/linux/hmm.h
index 9e0f00d..c79abfc 100644
--- a/include/linux/hmm.h
+++ b/include/linux/hmm.h
@@ -99,6 +99,7 @@ struct hmm;
* HMM_PFN_WRITE: CPU page table have the write permission set
* HMM_PFN_ERROR: corresponding CPU page table entry point to poisonous memory
* HMM_PFN_EMPTY: corresponding CPU page table entry is none (pte_none() true)
+ * HMM_PFN_FAULT: use by hmm_vma_fault() to signify which address need faulting
* HMM_PFN_DEVICE: this is device memory (ie a ZONE_DEVICE page)
* HMM_PFN_SPECIAL: corresponding CPU page table entry is special ie result of
* vm_insert_pfn() or vm_insert_page() and thus should not be mirror by a
@@ -113,10 +114,11 @@ typedef unsigned long hmm_pfn_t;
#define HMM_PFN_WRITE (1 << 2)
#define HMM_PFN_ERROR (1 << 3)
#define HMM_PFN_EMPTY (1 << 4)
-#define HMM_PFN_DEVICE (1 << 5)
-#define HMM_PFN_SPECIAL (1 << 6)
-#define HMM_PFN_UNADDRESSABLE (1 << 7)
-#define HMM_PFN_SHIFT 8
+#define HMM_PFN_FAULT (1 << 5)
+#define HMM_PFN_DEVICE (1 << 6)
+#define HMM_PFN_SPECIAL (1 << 7)
+#define HMM_PFN_UNADDRESSABLE (1 << 8)
+#define HMM_PFN_SHIFT 9
static inline struct page *hmm_pfn_to_page(hmm_pfn_t pfn)
{
@@ -298,6 +300,29 @@ int hmm_vma_get_pfns(struct vm_area_struct *vma,
hmm_pfn_t *pfns);
+/*
+ * Fault memory on behalf of device driver unlike handle_mm_fault() it will not
+ * migrate any device memory back to system memory. The hmm_pfn_t array will be
+ * updated with fault result and current snapshot of the CPU page table for the
+ * range. Note that you must use hmm_range_monitor_start/end() to ascertain if
+ * you could use those.
+ *
+ * DO NOT USE hmm_vma_range_lock()/hmm_vma_range_unlock() IT WILL DEADLOCK !
+ *
+ * The mmap_sem must be taken in read mode before entering and it might be drop
+ * by the function if that happen the function return false. Otherwise, if the
+ * mmap_sem is still held it return true. The return value does not reflect if
+ * the fault was successfull or not, you need to inspect the hmm_pfn_t array to
+ * determine fault status.
+ *
+ * See function description in mm/hmm.c for documentation.
+ */
+bool hmm_vma_fault(struct vm_area_struct *vma,
+ unsigned long start,
+ unsigned long end,
+ hmm_pfn_t *pfns);
+
+
/* Below are for HMM internal use only ! Not to be use by device driver ! */
void hmm_mm_destroy(struct mm_struct *mm);
diff --git a/mm/hmm.c b/mm/hmm.c
index f2ea76b..521adfd 100644
--- a/mm/hmm.c
+++ b/mm/hmm.c
@@ -461,6 +461,14 @@ bool hmm_vma_range_monitor_end(struct hmm_range *range)
EXPORT_SYMBOL(hmm_vma_range_monitor_end);
+static void hmm_pfns_error(hmm_pfn_t *pfns,
+ unsigned long addr,
+ unsigned long end)
+{
+ for (; addr < end; addr += PAGE_SIZE, pfns++)
+ *pfns = HMM_PFN_ERROR;
+}
+
static void hmm_pfns_empty(hmm_pfn_t *pfns,
unsigned long addr,
unsigned long end)
@@ -477,10 +485,47 @@ static void hmm_pfns_special(hmm_pfn_t *pfns,
*pfns = HMM_PFN_SPECIAL;
}
-static void hmm_vma_walk(struct vm_area_struct *vma,
+static void hmm_pfns_clear(hmm_pfn_t *pfns,
+ unsigned long addr,
+ unsigned long end)
+{
+ unsigned long npfns = (end - addr) >> PAGE_SHIFT;
+
+ memset(pfns, 0, sizeof(*pfns) * npfns);
+}
+
+static bool hmm_pfns_fault(hmm_pfn_t *pfns,
+ unsigned long addr,
+ unsigned long end)
+{
+ for (; addr < end; addr += PAGE_SIZE, pfns++)
+ if (*pfns & HMM_PFN_FAULT)
+ return true;
+ return false;
+}
+
+static bool hmm_vma_do_fault(struct vm_area_struct *vma,
+ unsigned long addr,
+ hmm_pfn_t *pfn)
+{
+ unsigned flags = FAULT_FLAG_ALLOW_RETRY | FAULT_FLAG_REMOTE;
+ int r;
+
+ flags |= (*pfn & HMM_PFN_WRITE) ? FAULT_FLAG_WRITE : 0;
+ r = handle_mm_fault(vma, addr, flags);
+ if (r & VM_FAULT_RETRY)
+ return false;
+ if (r & VM_FAULT_ERROR)
+ *pfn = HMM_PFN_ERROR;
+
+ return true;
+}
+
+static bool hmm_vma_walk(struct vm_area_struct *vma,
unsigned long start,
unsigned long end,
- hmm_pfn_t *pfns)
+ hmm_pfn_t *pfns,
+ bool fault)
{
unsigned long addr, next;
hmm_pfn_t flag;
@@ -489,6 +534,7 @@ static void hmm_vma_walk(struct vm_area_struct *vma,
for (addr = start; addr < end; addr = next) {
unsigned long i = (addr - start) >> PAGE_SHIFT;
+ bool writefault = false;
pgd_t *pgdp;
pud_t *pudp;
pmd_t *pmdp;
@@ -504,15 +550,37 @@ static void hmm_vma_walk(struct vm_area_struct *vma,
next = pgd_addr_end(addr, end);
pgdp = pgd_offset(vma->vm_mm, addr);
if (pgd_none(*pgdp) || pgd_bad(*pgdp)) {
- hmm_pfns_empty(&pfns[i], addr, next);
- continue;
+ if (!(vma->vm_flags & VM_READ)) {
+ hmm_pfns_empty(&pfns[i], addr, next);
+ continue;
+ }
+ if (!fault || !hmm_pfns_fault(&pfns[i], addr, next)) {
+ hmm_pfns_empty(&pfns[i], addr, next);
+ continue;
+ }
+ pudp = pud_alloc(vma->vm_mm, pgdp, addr);
+ if (!pudp) {
+ hmm_pfns_error(&pfns[i], addr, next);
+ continue;
+ }
}
next = pud_addr_end(addr, end);
pudp = pud_offset(pgdp, addr);
if (pud_none(*pudp) || pud_bad(*pudp)) {
- hmm_pfns_empty(&pfns[i], addr, next);
- continue;
+ if (!(vma->vm_flags & VM_READ)) {
+ hmm_pfns_empty(&pfns[i], addr, next);
+ continue;
+ }
+ if (!fault || !hmm_pfns_fault(&pfns[i], addr, next)) {
+ hmm_pfns_empty(&pfns[i], addr, next);
+ continue;
+ }
+ pmdp = pmd_alloc(vma->vm_mm, pudp, addr);
+ if (!pmdp) {
+ hmm_pfns_error(&pfns[i], addr, next);
+ continue;
+ }
}
next = pmd_addr_end(addr, end);
@@ -520,8 +588,23 @@ static void hmm_vma_walk(struct vm_area_struct *vma,
pmd = pmd_read_atomic(pmdp);
barrier();
if (pmd_none(pmd) || pmd_bad(pmd)) {
- hmm_pfns_empty(&pfns[i], addr, next);
- continue;
+ if (!(vma->vm_flags & VM_READ)) {
+ hmm_pfns_empty(&pfns[i], addr, next);
+ continue;
+ }
+ if (!fault || !hmm_pfns_fault(&pfns[i], addr, next)) {
+ hmm_pfns_empty(&pfns[i], addr, next);
+ continue;
+ }
+ /*
+ * Use pte_alloc() instead of pte_alloc_map, because we
+ * can't run pte_offset_map on the pmd, if an huge pmd
+ * could materialize from under us.
+ */
+ if (unlikely(pte_alloc(vma->vm_mm, pmdp, addr))) {
+ hmm_pfns_error(&pfns[i], addr, next);
+ continue;
+ }
}
if (pmd_trans_huge(pmd) || pmd_devmap(pmd)) {
unsigned long pfn = pmd_pfn(pmd) + pte_index(addr);
@@ -529,12 +612,33 @@ static void hmm_vma_walk(struct vm_area_struct *vma,
if (pmd_protnone(pmd)) {
hmm_pfns_clear(&pfns[i], addr, next);
+ if (!fault || !(vma->vm_flags & VM_READ))
+ continue;
+ if (!hmm_pfns_fault(&pfns[i], addr, next))
+ continue;
+
+ if (!hmm_vma_do_fault(vma, addr, &pfns[i]))
+ return false;
+ /* Start again for current address */
+ next = addr;
continue;
}
flags |= pmd_write(*pmdp) ? HMM_PFN_WRITE : 0;
flags |= pmd_devmap(pmd) ? HMM_PFN_DEVICE : 0;
- for (; addr < next; addr += PAGE_SIZE, i++, pfn++)
+ for (; addr < next; addr += PAGE_SIZE, i++, pfn++) {
+ bool fault = pfns[i] & HMM_PFN_FAULT;
+ bool write = pfns[i] & HMM_PFN_WRITE;
+
pfns[i] = hmm_pfn_from_pfn(pfn) | flags;
+ if (!fault || !write || flags & HMM_PFN_WRITE)
+ continue;
+ pfns[i] = HMM_PFN_FAULT | HMM_PFN_WRITE;
+ if (!hmm_vma_do_fault(vma, addr, &pfns[i]))
+ return false;
+ /* Start again for current address */
+ next = addr;
+ break;
+ }
continue;
}
@@ -543,41 +647,91 @@ static void hmm_vma_walk(struct vm_area_struct *vma,
swp_entry_t entry;
pte_t pte = *ptep;
- pfns[i] = 0;
-
if (pte_none(pte)) {
- pfns[i] = HMM_PFN_EMPTY;
- continue;
+ if (!fault || !(pfns[i] & HMM_PFN_FAULT)) {
+ pfns[i] = HMM_PFN_EMPTY;
+ continue;
+ }
+ if (!(vma->vm_flags & VM_READ)) {
+ pfns[i] = HMM_PFN_EMPTY;
+ continue;
+ }
+ if (!hmm_vma_do_fault(vma, addr, &pfns[i])) {
+ hmm_pfns_clear(&pfns[i], addr, end);
+ pte_unmap(ptep);
+ return false;
+ }
+ pte = *ptep;
}
entry = pte_to_swp_entry(pte);
if (!pte_present(pte) && !non_swap_entry(entry)) {
- continue;
+ if (!fault || !(pfns[i] & HMM_PFN_FAULT)) {
+ pfns[i] = 0;
+ continue;
+ }
+ if (!(vma->vm_flags & VM_READ)) {
+ pfns[i] = 0;
+ continue;
+ }
+ if (!hmm_vma_do_fault(vma, addr, &pfns[i])) {
+ hmm_pfns_clear(&pfns[i], addr, end);
+ pte_unmap(ptep);
+ return false;
+ }
+ pte = *ptep;
}
+ writefault = (pfns[i]&(HMM_PFN_WRITE|HMM_PFN_FAULT)) ==
+ (HMM_PFN_WRITE|HMM_PFN_FAULT) && fault;
+
if (pte_present(pte)) {
pfns[i] = hmm_pfn_from_pfn(pte_pfn(pte))|flag;
pfns[i] |= pte_write(pte) ? HMM_PFN_WRITE : 0;
- continue;
- }
-
- /*
- * This is a special swap entry, ignore migration, use
- * device and report anything else as error.
- */
- if (is_device_entry(entry)) {
+ } else if (is_device_entry(entry)) {
+ /* Do not fault device entry */
pfns[i] = hmm_pfn_from_pfn(swp_offset(entry));
if (is_write_device_entry(entry))
pfns[i] |= HMM_PFN_WRITE;
pfns[i] |= HMM_PFN_DEVICE;
pfns[i] |= HMM_PFN_UNADDRESSABLE;
pfns[i] |= flag;
- } else if (!is_migration_entry(entry)) {
+ } else if (is_migration_entry(entry) && fault) {
+ migration_entry_wait(vma->vm_mm, pmdp, addr);
+ /* Start again for current address */
+ next = addr;
+ ptep++;
+ break;
+ } else {
+ /* Report error for everything else */
pfns[i] = HMM_PFN_ERROR;
}
+ if (!(vma->vm_flags & VM_READ) ||
+ !(vma->vm_flags & VM_WRITE)) {
+ writefault = false;
+ continue;
+ }
+
+ if (writefault && !(pfns[i] & HMM_PFN_WRITE)) {
+ ptep++;
+ break;
+ }
+ writefault = false;
}
pte_unmap(ptep - 1);
+
+ if (writefault && (vma->vm_flags & VM_WRITE)) {
+ pfns[i] = HMM_PFN_WRITE | HMM_PFN_FAULT;
+ if (!hmm_vma_do_fault(vma, addr, &pfns[i])) {
+ return false;
+ }
+ writefault = false;
+ /* Start again for current address */
+ next = addr;
+ }
}
+
+ return true;
}
/*
@@ -613,7 +767,67 @@ int hmm_vma_get_pfns(struct vm_area_struct *vma,
if (end < vma->vm_start || end > vma->vm_end)
return -EINVAL;
- hmm_vma_walk(vma, start, end, pfns);
+ hmm_vma_walk(vma, start, end, pfns, false);
return 0;
}
EXPORT_SYMBOL(hmm_vma_get_pfns);
+
+
+/*
+ * hmm_vma_fault() - try to fault some address in a virtual address range
+ * @vma: virtual memory area containing the virtual address range
+ * @start: fault range virtual start address (inclusive)
+ * @end: fault range virtual end address (exclusive)
+ * @pfns: array of hmm_pfn_t, only entry with fault flag set will be faulted
+ * Returns: true mmap_sem is still held, false mmap_sem have been release
+ *
+ * This is similar to a regular CPU page fault except that it will not trigger
+ * any memory migration if the memory being faulted is not accessible by CPUs.
+ *
+ * Only pfn with fault flag set will be faulted and the hmm_pfn_t write flag
+ * will be use to determine if it is a write fault or not.
+ *
+ * On error, for one virtual address in the range, the function will set the
+ * hmm_pfn_t error flag for the corresponding pfn entry.
+ *
+ * Expected use pattern:
+ * retry:
+ * down_read(&mm->mmap_sem);
+ * // Find vma and address device wants to fault, initialize hmm_pfn_t
+ * // array accordingly
+ * hmm_vma_range_monitor_start(range, vma, start, end);
+ * if (!hmm_vma_fault(vma, start, end, pfns, allow_retry)) {
+ * hmm_vma_range_monitor_end(range);
+ * // You might want to rate limit or yield to play nicely, you may
+ * // also commit any valid pfn in the array assuming that you are
+ * // getting true from hmm_vma_range_monitor_end()
+ * goto retry;
+ * }
+ * // Take device driver lock that serialize device page table update
+ * driver_lock_device_page_table_update();
+ * if (hmm_vma_range_monitor_end(range)) {
+ * // Commit pfns we got from hmm_vma_fault()
+ * }
+ * driver_unlock_device_page_table_update();
+ * up_read(&mm->mmap_sem)
+ */
+bool hmm_vma_fault(struct vm_area_struct *vma,
+ unsigned long start,
+ unsigned long end,
+ hmm_pfn_t *pfns)
+{
+ /* FIXME support hugetlb fs */
+ if (is_vm_hugetlb_page(vma) || (vma->vm_flags & VM_SPECIAL)) {
+ hmm_pfns_special(pfns, start, end);
+ return true;
+ }
+
+ /* Sanity check, this really should not happen ! */
+ if (start < vma->vm_start || start >= vma->vm_end)
+ return true;
+ if (end < vma->vm_start || end > vma->vm_end)
+ return true;
+
+ return hmm_vma_walk(vma, start, end, pfns, true);
+}
+EXPORT_SYMBOL(hmm_vma_fault);
--
2.4.3
[toc] | [prev] | [next] | [standalone]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-18 18:30 +0100 |
| Subject | [HMM v13 12/18] mm/hmm/mirror: helper to snapshot CPU page table |
| Message-ID | <sEV97-1Es-3@gated-at.bofh.it> |
| In reply to | #1525568 |
This does not use existing page table walker because we want to share
same code for our page fault handler.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Signed-off-by: Jatin Kumar <jakumar@nvidia.com>
Signed-off-by: John Hubbard <jhubbard@nvidia.com>
Signed-off-by: Mark Hairgrove <mhairgrove@nvidia.com>
Signed-off-by: Sherry Cheung <SCheung@nvidia.com>
Signed-off-by: Subhash Gutti <sgutti@nvidia.com>
---
include/linux/hmm.h | 30 +++++++++-
mm/hmm.c | 163 ++++++++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 191 insertions(+), 2 deletions(-)
diff --git a/include/linux/hmm.h b/include/linux/hmm.h
index 6571647..9e0f00d 100644
--- a/include/linux/hmm.h
+++ b/include/linux/hmm.h
@@ -95,13 +95,28 @@ struct hmm;
*
* Flags:
* HMM_PFN_VALID: pfn is valid
+ * HMM_PFN_READ: read permission set
* HMM_PFN_WRITE: CPU page table have the write permission set
+ * HMM_PFN_ERROR: corresponding CPU page table entry point to poisonous memory
+ * HMM_PFN_EMPTY: corresponding CPU page table entry is none (pte_none() true)
+ * HMM_PFN_DEVICE: this is device memory (ie a ZONE_DEVICE page)
+ * HMM_PFN_SPECIAL: corresponding CPU page table entry is special ie result of
+ * vm_insert_pfn() or vm_insert_page() and thus should not be mirror by a
+ * device (the entry will never have HMM_PFN_VALID set and the pfn value
+ * is undefine)
+ * HMM_PFN_UNADDRESSABLE: unaddressable device memory (ZONE_DEVICE)
*/
typedef unsigned long hmm_pfn_t;
#define HMM_PFN_VALID (1 << 0)
-#define HMM_PFN_WRITE (1 << 1)
-#define HMM_PFN_SHIFT 2
+#define HMM_PFN_READ (1 << 1)
+#define HMM_PFN_WRITE (1 << 2)
+#define HMM_PFN_ERROR (1 << 3)
+#define HMM_PFN_EMPTY (1 << 4)
+#define HMM_PFN_DEVICE (1 << 5)
+#define HMM_PFN_SPECIAL (1 << 6)
+#define HMM_PFN_UNADDRESSABLE (1 << 7)
+#define HMM_PFN_SHIFT 8
static inline struct page *hmm_pfn_to_page(hmm_pfn_t pfn)
{
@@ -272,6 +287,17 @@ bool hmm_vma_range_monitor_start(struct hmm_range *range,
bool hmm_vma_range_monitor_end(struct hmm_range *range);
+/*
+ * Snapshot CPU page table, the snapshot content validity can be track using
+ * hmm_range_monitor_start/end() or hmm_vma_range_lock()/hmm_vma_range_unlock()
+ * mechanism. See function description in mm/hmm.c for documentation.
+ */
+int hmm_vma_get_pfns(struct vm_area_struct *vma,
+ unsigned long start,
+ unsigned long end,
+ hmm_pfn_t *pfns);
+
+
/* Below are for HMM internal use only ! Not to be use by device driver ! */
void hmm_mm_destroy(struct mm_struct *mm);
diff --git a/mm/hmm.c b/mm/hmm.c
index 746eb96..f2ea76b 100644
--- a/mm/hmm.c
+++ b/mm/hmm.c
@@ -19,10 +19,15 @@
*/
#include <linux/mm.h>
#include <linux/hmm.h>
+#include <linux/rmap.h>
+#include <linux/swap.h>
#include <linux/slab.h>
#include <linux/sched.h>
+#include <linux/swapops.h>
+#include <linux/hugetlb.h>
#include <linux/mmu_notifier.h>
+
/*
* struct hmm - HMM per mm struct
*
@@ -454,3 +459,161 @@ bool hmm_vma_range_monitor_end(struct hmm_range *range)
return valid;
}
EXPORT_SYMBOL(hmm_vma_range_monitor_end);
+
+
+static void hmm_pfns_empty(hmm_pfn_t *pfns,
+ unsigned long addr,
+ unsigned long end)
+{
+ for (; addr < end; addr += PAGE_SIZE, pfns++)
+ *pfns = HMM_PFN_EMPTY;
+}
+
+static void hmm_pfns_special(hmm_pfn_t *pfns,
+ unsigned long addr,
+ unsigned long end)
+{
+ for (; addr < end; addr += PAGE_SIZE, pfns++)
+ *pfns = HMM_PFN_SPECIAL;
+}
+
+static void hmm_vma_walk(struct vm_area_struct *vma,
+ unsigned long start,
+ unsigned long end,
+ hmm_pfn_t *pfns)
+{
+ unsigned long addr, next;
+ hmm_pfn_t flag;
+
+ flag = vma->vm_flags & VM_READ ? HMM_PFN_READ : 0;
+
+ for (addr = start; addr < end; addr = next) {
+ unsigned long i = (addr - start) >> PAGE_SHIFT;
+ pgd_t *pgdp;
+ pud_t *pudp;
+ pmd_t *pmdp;
+ pte_t *ptep;
+ pmd_t pmd;
+
+ /*
+ * We are accessing/faulting for a device from an unknown
+ * thread that might be foreign to the mm we are faulting
+ * against so do not call arch_vma_access_permitted() !
+ */
+
+ next = pgd_addr_end(addr, end);
+ pgdp = pgd_offset(vma->vm_mm, addr);
+ if (pgd_none(*pgdp) || pgd_bad(*pgdp)) {
+ hmm_pfns_empty(&pfns[i], addr, next);
+ continue;
+ }
+
+ next = pud_addr_end(addr, end);
+ pudp = pud_offset(pgdp, addr);
+ if (pud_none(*pudp) || pud_bad(*pudp)) {
+ hmm_pfns_empty(&pfns[i], addr, next);
+ continue;
+ }
+
+ next = pmd_addr_end(addr, end);
+ pmdp = pmd_offset(pudp, addr);
+ pmd = pmd_read_atomic(pmdp);
+ barrier();
+ if (pmd_none(pmd) || pmd_bad(pmd)) {
+ hmm_pfns_empty(&pfns[i], addr, next);
+ continue;
+ }
+ if (pmd_trans_huge(pmd) || pmd_devmap(pmd)) {
+ unsigned long pfn = pmd_pfn(pmd) + pte_index(addr);
+ hmm_pfn_t flags = flag;
+
+ if (pmd_protnone(pmd)) {
+ hmm_pfns_clear(&pfns[i], addr, next);
+ continue;
+ }
+ flags |= pmd_write(*pmdp) ? HMM_PFN_WRITE : 0;
+ flags |= pmd_devmap(pmd) ? HMM_PFN_DEVICE : 0;
+ for (; addr < next; addr += PAGE_SIZE, i++, pfn++)
+ pfns[i] = hmm_pfn_from_pfn(pfn) | flags;
+ continue;
+ }
+
+ ptep = pte_offset_map(pmdp, addr);
+ for (; addr < next; addr += PAGE_SIZE, i++, ptep++) {
+ swp_entry_t entry;
+ pte_t pte = *ptep;
+
+ pfns[i] = 0;
+
+ if (pte_none(pte)) {
+ pfns[i] = HMM_PFN_EMPTY;
+ continue;
+ }
+
+ entry = pte_to_swp_entry(pte);
+ if (!pte_present(pte) && !non_swap_entry(entry)) {
+ continue;
+ }
+
+ if (pte_present(pte)) {
+ pfns[i] = hmm_pfn_from_pfn(pte_pfn(pte))|flag;
+ pfns[i] |= pte_write(pte) ? HMM_PFN_WRITE : 0;
+ continue;
+ }
+
+ /*
+ * This is a special swap entry, ignore migration, use
+ * device and report anything else as error.
+ */
+ if (is_device_entry(entry)) {
+ pfns[i] = hmm_pfn_from_pfn(swp_offset(entry));
+ if (is_write_device_entry(entry))
+ pfns[i] |= HMM_PFN_WRITE;
+ pfns[i] |= HMM_PFN_DEVICE;
+ pfns[i] |= HMM_PFN_UNADDRESSABLE;
+ pfns[i] |= flag;
+ } else if (!is_migration_entry(entry)) {
+ pfns[i] = HMM_PFN_ERROR;
+ }
+ }
+ pte_unmap(ptep - 1);
+ }
+}
+
+/*
+ * hmm_vma_get_pfns() - snapshot CPU page table for a range of virtual address
+ * @vma: virtual memory area containing the virtual address range
+ * @start: range virtual start address (inclusive)
+ * @end: range virtual end address (exclusive)
+ * @entries: array of hmm_pfn_t provided by caller fill by function
+ * Returns: -EINVAL if invalid argument, 0 otherwise
+ *
+ * This snapshot the CPU page table for a range of virtual address, snapshot is
+ * only valid while protected by hmm_vma_range_lock() or if return cookie value
+ * is still valid (see hmm_vma_check_cookie()).
+ *
+ * It will fill the pfns array using CPU pte. Note that any invalid CPU page
+ * table entry, at time of snapshot, can turn into a valid one after this
+ * function return but before calling hmm_vma_range_unlock().
+ */
+int hmm_vma_get_pfns(struct vm_area_struct *vma,
+ unsigned long start,
+ unsigned long end,
+ hmm_pfn_t *pfns)
+{
+ /* FIXME support hugetlb fs */
+ if (is_vm_hugetlb_page(vma) || (vma->vm_flags & VM_SPECIAL)) {
+ hmm_pfns_special(pfns, start, end);
+ return -EINVAL;
+ }
+
+ /* Sanity check, this really should not happen ! */
+ if (start < vma->vm_start || start >= vma->vm_end)
+ return -EINVAL;
+ if (end < vma->vm_start || end > vma->vm_end)
+ return -EINVAL;
+
+ hmm_vma_walk(vma, start, end, pfns);
+ return 0;
+}
+EXPORT_SYMBOL(hmm_vma_get_pfns);
--
2.4.3
[toc] | [prev] | [next] | [standalone]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-18 18:30 +0100 |
| Subject | [HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory |
| Message-ID | <sEV97-1Es-5@gated-at.bofh.it> |
| In reply to | #1525568 |
This add support for un-addressable device memory. Such memory is hotpluged
only so we can have struct page but should never be map. This patch add code
to mm page fault code path to catch any such mapping and SIGBUS on such event.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Cc: Dan Williams <dan.j.williams@intel.com>
Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
---
drivers/dax/pmem.c | 3 ++-
drivers/nvdimm/pmem.c | 5 +++--
include/linux/memremap.h | 23 ++++++++++++++++++++---
kernel/memremap.c | 12 +++++++++---
mm/memory.c | 9 +++++++++
tools/testing/nvdimm/test/iomap.c | 2 +-
6 files changed, 44 insertions(+), 10 deletions(-)
diff --git a/drivers/dax/pmem.c b/drivers/dax/pmem.c
index 1f01e98..1b42aef 100644
--- a/drivers/dax/pmem.c
+++ b/drivers/dax/pmem.c
@@ -107,7 +107,8 @@ static int dax_pmem_probe(struct device *dev)
if (rc)
return rc;
- addr = devm_memremap_pages(dev, &res, &dax_pmem->ref, altmap);
+ addr = devm_memremap_pages(dev, &res, &dax_pmem->ref,
+ altmap, NULL, MEMORY_DEVICE);
if (IS_ERR(addr))
return PTR_ERR(addr);
diff --git a/drivers/nvdimm/pmem.c b/drivers/nvdimm/pmem.c
index 571a6c7..5ffd937 100644
--- a/drivers/nvdimm/pmem.c
+++ b/drivers/nvdimm/pmem.c
@@ -260,7 +260,7 @@ static int pmem_attach_disk(struct device *dev,
pmem->pfn_flags = PFN_DEV;
if (is_nd_pfn(dev)) {
addr = devm_memremap_pages(dev, &pfn_res, &q->q_usage_counter,
- altmap);
+ altmap, NULL, MEMORY_DEVICE);
pfn_sb = nd_pfn->pfn_sb;
pmem->data_offset = le64_to_cpu(pfn_sb->dataoff);
pmem->pfn_pad = resource_size(res) - resource_size(&pfn_res);
@@ -269,7 +269,8 @@ static int pmem_attach_disk(struct device *dev,
res->start += pmem->data_offset;
} else if (pmem_should_map_pages(dev)) {
addr = devm_memremap_pages(dev, &nsio->res,
- &q->q_usage_counter, NULL);
+ &q->q_usage_counter,
+ NULL, NULL, MEMORY_DEVICE);
pmem->pfn_flags |= PFN_MAP;
} else
addr = devm_memremap(dev, pmem->phys_addr,
diff --git a/include/linux/memremap.h b/include/linux/memremap.h
index 9341619..fe61dca 100644
--- a/include/linux/memremap.h
+++ b/include/linux/memremap.h
@@ -41,22 +41,34 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
* @res: physical address range covered by @ref
* @ref: reference count that pins the devm_memremap_pages() mapping
* @dev: host device of the mapping for debug
+ * @flags: memory flags (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
*/
struct dev_pagemap {
struct vmem_altmap *altmap;
const struct resource *res;
struct percpu_ref *ref;
struct device *dev;
+ int flags;
};
#ifdef CONFIG_ZONE_DEVICE
void *devm_memremap_pages(struct device *dev, struct resource *res,
- struct percpu_ref *ref, struct vmem_altmap *altmap);
+ struct percpu_ref *ref, struct vmem_altmap *altmap,
+ struct dev_pagemap **ppgmap, int flags);
struct dev_pagemap *find_dev_pagemap(resource_size_t phys);
+
+static inline bool is_addressable_page(const struct page *page)
+{
+ return ((page_zonenum(page) != ZONE_DEVICE) ||
+ !(page->pgmap->flags & MEMORY_UNADDRESSABLE));
+}
#else
static inline void *devm_memremap_pages(struct device *dev,
- struct resource *res, struct percpu_ref *ref,
- struct vmem_altmap *altmap)
+ struct resource *res,
+ struct percpu_ref *ref,
+ struct vmem_altmap *altmap,
+ struct dev_pagemap **ppgmap,
+ int flags)
{
/*
* Fail attempts to call devm_memremap_pages() without
@@ -71,6 +83,11 @@ static inline struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
{
return NULL;
}
+
+static inline bool is_addressable_page(const struct page *page)
+{
+ return true;
+}
#endif
/**
diff --git a/kernel/memremap.c b/kernel/memremap.c
index 07665eb..438a73aa2 100644
--- a/kernel/memremap.c
+++ b/kernel/memremap.c
@@ -246,7 +246,7 @@ static void devm_memremap_pages_release(struct device *dev, void *data)
/* pages are dead and unused, undo the arch mapping */
align_start = res->start & ~(SECTION_SIZE - 1);
align_size = ALIGN(resource_size(res), SECTION_SIZE);
- arch_remove_memory(align_start, align_size, MEMORY_DEVICE);
+ arch_remove_memory(align_start, align_size, pgmap->flags);
untrack_pfn(NULL, PHYS_PFN(align_start), align_size);
pgmap_radix_release(res);
dev_WARN_ONCE(dev, pgmap->altmap && pgmap->altmap->alloc,
@@ -270,6 +270,8 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
* @res: "host memory" address range
* @ref: a live per-cpu reference count
* @altmap: optional descriptor for allocating the memmap from @res
+ * @ppgmap: pointer set to new page dev_pagemap on success
+ * @flags: flag for memory (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
*
* Notes:
* 1/ @ref must be 'live' on entry and 'dead' before devm_memunmap_pages() time
@@ -280,7 +282,8 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
* this is not enforced.
*/
void *devm_memremap_pages(struct device *dev, struct resource *res,
- struct percpu_ref *ref, struct vmem_altmap *altmap)
+ struct percpu_ref *ref, struct vmem_altmap *altmap,
+ struct dev_pagemap **ppgmap, int flags)
{
resource_size_t key, align_start, align_size, align_end;
pgprot_t pgprot = PAGE_KERNEL;
@@ -322,6 +325,7 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
}
pgmap->ref = ref;
pgmap->res = &page_map->res;
+ pgmap->flags = flags | MEMORY_DEVICE;
mutex_lock(&pgmap_lock);
error = 0;
@@ -358,7 +362,7 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
if (error)
goto err_pfn_remap;
- error = arch_add_memory(nid, align_start, align_size, MEMORY_DEVICE);
+ error = arch_add_memory(nid, align_start, align_size, pgmap->flags);
if (error)
goto err_add_memory;
@@ -375,6 +379,8 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
page->pgmap = pgmap;
}
devres_add(dev, page_map);
+ if (ppgmap)
+ *ppgmap = pgmap;
return __va(res->start);
err_add_memory:
diff --git a/mm/memory.c b/mm/memory.c
index 840adc6..15f2908 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -45,6 +45,7 @@
#include <linux/swap.h>
#include <linux/highmem.h>
#include <linux/pagemap.h>
+#include <linux/memremap.h>
#include <linux/ksm.h>
#include <linux/rmap.h>
#include <linux/export.h>
@@ -3482,6 +3483,7 @@ static inline bool vma_is_accessible(struct vm_area_struct *vma)
static int handle_pte_fault(struct fault_env *fe)
{
pte_t entry;
+ struct page *page;
if (unlikely(pmd_none(*fe->pmd))) {
/*
@@ -3533,6 +3535,13 @@ static int handle_pte_fault(struct fault_env *fe)
if (pte_protnone(entry) && vma_is_accessible(fe->vma))
return do_numa_page(fe, entry);
+ /* Catch mapping of un-addressable memory this should never happen */
+ page = pfn_to_page(pte_pfn(entry));
+ if (!is_addressable_page(page)) {
+ print_bad_pte(fe->vma, fe->address, entry, page);
+ return VM_FAULT_SIGBUS;
+ }
+
fe->ptl = pte_lockptr(fe->vma->vm_mm, fe->pmd);
spin_lock(fe->ptl);
if (unlikely(!pte_same(*fe->pte, entry)))
diff --git a/tools/testing/nvdimm/test/iomap.c b/tools/testing/nvdimm/test/iomap.c
index c29f8dc..899d6a8 100644
--- a/tools/testing/nvdimm/test/iomap.c
+++ b/tools/testing/nvdimm/test/iomap.c
@@ -108,7 +108,7 @@ void *__wrap_devm_memremap_pages(struct device *dev, struct resource *res,
if (nfit_res)
return nfit_res->buf + offset - nfit_res->res->start;
- return devm_memremap_pages(dev, res, ref, altmap);
+ return devm_memremap_pages(dev, res, ref, altmap, MEMORY_DEVICE);
}
EXPORT_SYMBOL(__wrap_devm_memremap_pages);
--
2.4.3
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-21 09:10 +0100 |
| Subject | Re: [HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory |
| Message-ID | <sFRPP-74b-5@gated-at.bofh.it> |
| In reply to | #1525573 |
On 11/18/2016 11:48 PM, Jérôme Glisse wrote:
> This add support for un-addressable device memory. Such memory is hotpluged
> only so we can have struct page but should never be map. This patch add code
struct pages inside the system RAM range unlike the vmem_altmap scheme
where the struct pages can be inside the device memory itself. This
possibility does not arise for un addressable device memory. May be we
will have to block the paths where vmem_altmap is requested along with
un addressable device memory.
> to mm page fault code path to catch any such mapping and SIGBUS on such event.
>
> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
> Cc: Dan Williams <dan.j.williams@intel.com>
> Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
> ---
> drivers/dax/pmem.c | 3 ++-
> drivers/nvdimm/pmem.c | 5 +++--
> include/linux/memremap.h | 23 ++++++++++++++++++++---
> kernel/memremap.c | 12 +++++++++---
> mm/memory.c | 9 +++++++++
> tools/testing/nvdimm/test/iomap.c | 2 +-
> 6 files changed, 44 insertions(+), 10 deletions(-)
>
> diff --git a/drivers/dax/pmem.c b/drivers/dax/pmem.c
> index 1f01e98..1b42aef 100644
> --- a/drivers/dax/pmem.c
> +++ b/drivers/dax/pmem.c
> @@ -107,7 +107,8 @@ static int dax_pmem_probe(struct device *dev)
> if (rc)
> return rc;
>
> - addr = devm_memremap_pages(dev, &res, &dax_pmem->ref, altmap);
> + addr = devm_memremap_pages(dev, &res, &dax_pmem->ref,
> + altmap, NULL, MEMORY_DEVICE);
> if (IS_ERR(addr))
> return PTR_ERR(addr);
>
> diff --git a/drivers/nvdimm/pmem.c b/drivers/nvdimm/pmem.c
> index 571a6c7..5ffd937 100644
> --- a/drivers/nvdimm/pmem.c
> +++ b/drivers/nvdimm/pmem.c
> @@ -260,7 +260,7 @@ static int pmem_attach_disk(struct device *dev,
> pmem->pfn_flags = PFN_DEV;
> if (is_nd_pfn(dev)) {
> addr = devm_memremap_pages(dev, &pfn_res, &q->q_usage_counter,
> - altmap);
> + altmap, NULL, MEMORY_DEVICE);
> pfn_sb = nd_pfn->pfn_sb;
> pmem->data_offset = le64_to_cpu(pfn_sb->dataoff);
> pmem->pfn_pad = resource_size(res) - resource_size(&pfn_res);
> @@ -269,7 +269,8 @@ static int pmem_attach_disk(struct device *dev,
> res->start += pmem->data_offset;
> } else if (pmem_should_map_pages(dev)) {
> addr = devm_memremap_pages(dev, &nsio->res,
> - &q->q_usage_counter, NULL);
> + &q->q_usage_counter,
> + NULL, NULL, MEMORY_DEVICE);
> pmem->pfn_flags |= PFN_MAP;
> } else
> addr = devm_memremap(dev, pmem->phys_addr,
> diff --git a/include/linux/memremap.h b/include/linux/memremap.h
> index 9341619..fe61dca 100644
> --- a/include/linux/memremap.h
> +++ b/include/linux/memremap.h
> @@ -41,22 +41,34 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
> * @res: physical address range covered by @ref
> * @ref: reference count that pins the devm_memremap_pages() mapping
> * @dev: host device of the mapping for debug
> + * @flags: memory flags (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
^^^^^^^^^^^^^ device memory flags instead ?
> */
> struct dev_pagemap {
> struct vmem_altmap *altmap;
> const struct resource *res;
> struct percpu_ref *ref;
> struct device *dev;
> + int flags;
> };
>
> #ifdef CONFIG_ZONE_DEVICE
> void *devm_memremap_pages(struct device *dev, struct resource *res,
> - struct percpu_ref *ref, struct vmem_altmap *altmap);
> + struct percpu_ref *ref, struct vmem_altmap *altmap,
> + struct dev_pagemap **ppgmap, int flags);
> struct dev_pagemap *find_dev_pagemap(resource_size_t phys);
> +
> +static inline bool is_addressable_page(const struct page *page)
> +{
> + return ((page_zonenum(page) != ZONE_DEVICE) ||
> + !(page->pgmap->flags & MEMORY_UNADDRESSABLE));
> +}
> #else
> static inline void *devm_memremap_pages(struct device *dev,
> - struct resource *res, struct percpu_ref *ref,
> - struct vmem_altmap *altmap)
> + struct resource *res,
> + struct percpu_ref *ref,
> + struct vmem_altmap *altmap,
> + struct dev_pagemap **ppgmap,
> + int flags)
As I had mentioned before devm_memremap_pages() should be changed not
to accept a valid altmap along with request for un-addressable memory.
> {
> /*
> * Fail attempts to call devm_memremap_pages() without
> @@ -71,6 +83,11 @@ static inline struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
> {
> return NULL;
> }
> +
> +static inline bool is_addressable_page(const struct page *page)
> +{
> + return true;
> +}
> #endif
>
> /**
> diff --git a/kernel/memremap.c b/kernel/memremap.c
> index 07665eb..438a73aa2 100644
> --- a/kernel/memremap.c
> +++ b/kernel/memremap.c
> @@ -246,7 +246,7 @@ static void devm_memremap_pages_release(struct device *dev, void *data)
> /* pages are dead and unused, undo the arch mapping */
> align_start = res->start & ~(SECTION_SIZE - 1);
> align_size = ALIGN(resource_size(res), SECTION_SIZE);
> - arch_remove_memory(align_start, align_size, MEMORY_DEVICE);
> + arch_remove_memory(align_start, align_size, pgmap->flags);
> untrack_pfn(NULL, PHYS_PFN(align_start), align_size);
> pgmap_radix_release(res);
> dev_WARN_ONCE(dev, pgmap->altmap && pgmap->altmap->alloc,
> @@ -270,6 +270,8 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
> * @res: "host memory" address range
> * @ref: a live per-cpu reference count
> * @altmap: optional descriptor for allocating the memmap from @res
> + * @ppgmap: pointer set to new page dev_pagemap on success
> + * @flags: flag for memory (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
> *
> * Notes:
> * 1/ @ref must be 'live' on entry and 'dead' before devm_memunmap_pages() time
> @@ -280,7 +282,8 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
> * this is not enforced.
> */
> void *devm_memremap_pages(struct device *dev, struct resource *res,
> - struct percpu_ref *ref, struct vmem_altmap *altmap)
> + struct percpu_ref *ref, struct vmem_altmap *altmap,
> + struct dev_pagemap **ppgmap, int flags)
> {
> resource_size_t key, align_start, align_size, align_end;
> pgprot_t pgprot = PAGE_KERNEL;
> @@ -322,6 +325,7 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
> }
> pgmap->ref = ref;
> pgmap->res = &page_map->res;
> + pgmap->flags = flags | MEMORY_DEVICE;
So the caller of devm_memremap_pages() should not have give out MEMORY_DEVICE
in the flag it passed on to this function ? Hmm, else we should just check
that the flags contains all appropriate bits before proceeding.
>
> mutex_lock(&pgmap_lock);
> error = 0;
> @@ -358,7 +362,7 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
> if (error)
> goto err_pfn_remap;
>
> - error = arch_add_memory(nid, align_start, align_size, MEMORY_DEVICE);
> + error = arch_add_memory(nid, align_start, align_size, pgmap->flags);
> if (error)
> goto err_add_memory;
>
> @@ -375,6 +379,8 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
> page->pgmap = pgmap;
> }
> devres_add(dev, page_map);
> + if (ppgmap)
> + *ppgmap = pgmap;
> return __va(res->start);
>
> err_add_memory:
> diff --git a/mm/memory.c b/mm/memory.c
> index 840adc6..15f2908 100644
> --- a/mm/memory.c
> +++ b/mm/memory.c
> @@ -45,6 +45,7 @@
> #include <linux/swap.h>
> #include <linux/highmem.h>
> #include <linux/pagemap.h>
> +#include <linux/memremap.h>
> #include <linux/ksm.h>
> #include <linux/rmap.h>
> #include <linux/export.h>
> @@ -3482,6 +3483,7 @@ static inline bool vma_is_accessible(struct vm_area_struct *vma)
> static int handle_pte_fault(struct fault_env *fe)
> {
> pte_t entry;
> + struct page *page;
>
> if (unlikely(pmd_none(*fe->pmd))) {
> /*
> @@ -3533,6 +3535,13 @@ static int handle_pte_fault(struct fault_env *fe)
> if (pte_protnone(entry) && vma_is_accessible(fe->vma))
> return do_numa_page(fe, entry);
>
> + /* Catch mapping of un-addressable memory this should never happen */
> + page = pfn_to_page(pte_pfn(entry));
> + if (!is_addressable_page(page)) {
> + print_bad_pte(fe->vma, fe->address, entry, page);
> + return VM_FAULT_SIGBUS;
> + }
Right, core VM should never put an un-addressable page in the page table.
> +
> fe->ptl = pte_lockptr(fe->vma->vm_mm, fe->pmd);
> spin_lock(fe->ptl);
> if (unlikely(!pte_same(*fe->pte, entry)))
> diff --git a/tools/testing/nvdimm/test/iomap.c b/tools/testing/nvdimm/test/iomap.c
> index c29f8dc..899d6a8 100644
> --- a/tools/testing/nvdimm/test/iomap.c
> +++ b/tools/testing/nvdimm/test/iomap.c
> @@ -108,7 +108,7 @@ void *__wrap_devm_memremap_pages(struct device *dev, struct resource *res,
>
> if (nfit_res)
> return nfit_res->buf + offset - nfit_res->res->start;
> - return devm_memremap_pages(dev, res, ref, altmap);
> + return devm_memremap_pages(dev, res, ref, altmap, MEMORY_DEVICE);
> }
> EXPORT_SYMBOL(__wrap_devm_memremap_pages);
>
>
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-21 13:40 +0100 |
| Subject | Re: [HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory |
| Message-ID | <sFW37-1bK-1@gated-at.bofh.it> |
| In reply to | #1526430 |
On Mon, Nov 21, 2016 at 01:36:57PM +0530, Anshuman Khandual wrote:
> On 11/18/2016 11:48 PM, Jérôme Glisse wrote:
> > This add support for un-addressable device memory. Such memory is hotpluged
> > only so we can have struct page but should never be map. This patch add code
>
> struct pages inside the system RAM range unlike the vmem_altmap scheme
> where the struct pages can be inside the device memory itself. This
> possibility does not arise for un addressable device memory. May be we
> will have to block the paths where vmem_altmap is requested along with
> un addressable device memory.
I did not think checking for that explicitly was necessary, sounded like shooting
yourself in the foot and that it would be obvious :)
[...]
> > diff --git a/include/linux/memremap.h b/include/linux/memremap.h
> > index 9341619..fe61dca 100644
> > --- a/include/linux/memremap.h
> > +++ b/include/linux/memremap.h
> > @@ -41,22 +41,34 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
> > * @res: physical address range covered by @ref
> > * @ref: reference count that pins the devm_memremap_pages() mapping
> > * @dev: host device of the mapping for debug
> > + * @flags: memory flags (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
>
> ^^^^^^^^^^^^^ device memory flags instead ?
Well maybe it will be use for something else than device memory in the future
but yes for now it is only device memory so i can rename it.
> > */
> > struct dev_pagemap {
> > struct vmem_altmap *altmap;
> > const struct resource *res;
> > struct percpu_ref *ref;
> > struct device *dev;
> > + int flags;
> > };
> >
> > #ifdef CONFIG_ZONE_DEVICE
> > void *devm_memremap_pages(struct device *dev, struct resource *res,
> > - struct percpu_ref *ref, struct vmem_altmap *altmap);
> > + struct percpu_ref *ref, struct vmem_altmap *altmap,
> > + struct dev_pagemap **ppgmap, int flags);
> > struct dev_pagemap *find_dev_pagemap(resource_size_t phys);
> > +
> > +static inline bool is_addressable_page(const struct page *page)
> > +{
> > + return ((page_zonenum(page) != ZONE_DEVICE) ||
> > + !(page->pgmap->flags & MEMORY_UNADDRESSABLE));
> > +}
> > #else
> > static inline void *devm_memremap_pages(struct device *dev,
> > - struct resource *res, struct percpu_ref *ref,
> > - struct vmem_altmap *altmap)
> > + struct resource *res,
> > + struct percpu_ref *ref,
> > + struct vmem_altmap *altmap,
> > + struct dev_pagemap **ppgmap,
> > + int flags)
>
>
> As I had mentioned before devm_memremap_pages() should be changed not
> to accept a valid altmap along with request for un-addressable memory.
If you fear such case yes sure.
[...]
> > diff --git a/kernel/memremap.c b/kernel/memremap.c
> > index 07665eb..438a73aa2 100644
> > --- a/kernel/memremap.c
> > +++ b/kernel/memremap.c
> > @@ -246,7 +246,7 @@ static void devm_memremap_pages_release(struct device *dev, void *data)
> > /* pages are dead and unused, undo the arch mapping */
> > align_start = res->start & ~(SECTION_SIZE - 1);
> > align_size = ALIGN(resource_size(res), SECTION_SIZE);
> > - arch_remove_memory(align_start, align_size, MEMORY_DEVICE);
> > + arch_remove_memory(align_start, align_size, pgmap->flags);
> > untrack_pfn(NULL, PHYS_PFN(align_start), align_size);
> > pgmap_radix_release(res);
> > dev_WARN_ONCE(dev, pgmap->altmap && pgmap->altmap->alloc,
> > @@ -270,6 +270,8 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
> > * @res: "host memory" address range
> > * @ref: a live per-cpu reference count
> > * @altmap: optional descriptor for allocating the memmap from @res
> > + * @ppgmap: pointer set to new page dev_pagemap on success
> > + * @flags: flag for memory (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
> > *
> > * Notes:
> > * 1/ @ref must be 'live' on entry and 'dead' before devm_memunmap_pages() time
> > @@ -280,7 +282,8 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
> > * this is not enforced.
> > */
> > void *devm_memremap_pages(struct device *dev, struct resource *res,
> > - struct percpu_ref *ref, struct vmem_altmap *altmap)
> > + struct percpu_ref *ref, struct vmem_altmap *altmap,
> > + struct dev_pagemap **ppgmap, int flags)
> > {
> > resource_size_t key, align_start, align_size, align_end;
> > pgprot_t pgprot = PAGE_KERNEL;
> > @@ -322,6 +325,7 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
> > }
> > pgmap->ref = ref;
> > pgmap->res = &page_map->res;
> > + pgmap->flags = flags | MEMORY_DEVICE;
>
> So the caller of devm_memremap_pages() should not have give out MEMORY_DEVICE
> in the flag it passed on to this function ? Hmm, else we should just check
> that the flags contains all appropriate bits before proceeding.
Here i was just trying to be on the safe side, yes caller should already have set
the flag but this function is only use for device memory so it did not seem like
it would hurt to be extra safe. I can add a BUG_ON() but it seems people have mix
feeling about BUG_ON()
Cheers,
Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-22 06:20 +0100 |
| Subject | Re: [HMM v13 02/18] mm/ZONE_DEVICE/unaddressable: add support for un-addressable device memory |
| Message-ID | <sGbER-2Ps-1@gated-at.bofh.it> |
| In reply to | #1526632 |
On 11/21/2016 06:03 PM, Jerome Glisse wrote:
> On Mon, Nov 21, 2016 at 01:36:57PM +0530, Anshuman Khandual wrote:
>> On 11/18/2016 11:48 PM, Jérôme Glisse wrote:
>>> This add support for un-addressable device memory. Such memory is hotpluged
>>> only so we can have struct page but should never be map. This patch add code
>>
>> struct pages inside the system RAM range unlike the vmem_altmap scheme
>> where the struct pages can be inside the device memory itself. This
>> possibility does not arise for un addressable device memory. May be we
>> will have to block the paths where vmem_altmap is requested along with
>> un addressable device memory.
>
> I did not think checking for that explicitly was necessary, sounded like shooting
> yourself in the foot and that it would be obvious :)
dev_memremap_pages() is kind of an important interface for getting
device memory into kernel through ZONE_DEVICE. So it should actually
enforce all these checks. Also we should document these things clearly
above the function.
>
> [...]
>
>>> diff --git a/include/linux/memremap.h b/include/linux/memremap.h
>>> index 9341619..fe61dca 100644
>>> --- a/include/linux/memremap.h
>>> +++ b/include/linux/memremap.h
>>> @@ -41,22 +41,34 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
>>> * @res: physical address range covered by @ref
>>> * @ref: reference count that pins the devm_memremap_pages() mapping
>>> * @dev: host device of the mapping for debug
>>> + * @flags: memory flags (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
>>
>> ^^^^^^^^^^^^^ device memory flags instead ?
>
> Well maybe it will be use for something else than device memory in the future
> but yes for now it is only device memory so i can rename it.
>
>>> */
>>> struct dev_pagemap {
>>> struct vmem_altmap *altmap;
>>> const struct resource *res;
>>> struct percpu_ref *ref;
>>> struct device *dev;
>>> + int flags;
>>> };
>>>
>>> #ifdef CONFIG_ZONE_DEVICE
>>> void *devm_memremap_pages(struct device *dev, struct resource *res,
>>> - struct percpu_ref *ref, struct vmem_altmap *altmap);
>>> + struct percpu_ref *ref, struct vmem_altmap *altmap,
>>> + struct dev_pagemap **ppgmap, int flags);
>>> struct dev_pagemap *find_dev_pagemap(resource_size_t phys);
>>> +
>>> +static inline bool is_addressable_page(const struct page *page)
>>> +{
>>> + return ((page_zonenum(page) != ZONE_DEVICE) ||
>>> + !(page->pgmap->flags & MEMORY_UNADDRESSABLE));
>>> +}
>>> #else
>>> static inline void *devm_memremap_pages(struct device *dev,
>>> - struct resource *res, struct percpu_ref *ref,
>>> - struct vmem_altmap *altmap)
>>> + struct resource *res,
>>> + struct percpu_ref *ref,
>>> + struct vmem_altmap *altmap,
>>> + struct dev_pagemap **ppgmap,
>>> + int flags)
>>
>>
>> As I had mentioned before devm_memremap_pages() should be changed not
>> to accept a valid altmap along with request for un-addressable memory.
>
> If you fear such case yes sure.
>
>
> [...]
>
>>> diff --git a/kernel/memremap.c b/kernel/memremap.c
>>> index 07665eb..438a73aa2 100644
>>> --- a/kernel/memremap.c
>>> +++ b/kernel/memremap.c
>>> @@ -246,7 +246,7 @@ static void devm_memremap_pages_release(struct device *dev, void *data)
>>> /* pages are dead and unused, undo the arch mapping */
>>> align_start = res->start & ~(SECTION_SIZE - 1);
>>> align_size = ALIGN(resource_size(res), SECTION_SIZE);
>>> - arch_remove_memory(align_start, align_size, MEMORY_DEVICE);
>>> + arch_remove_memory(align_start, align_size, pgmap->flags);
>>> untrack_pfn(NULL, PHYS_PFN(align_start), align_size);
>>> pgmap_radix_release(res);
>>> dev_WARN_ONCE(dev, pgmap->altmap && pgmap->altmap->alloc,
>>> @@ -270,6 +270,8 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
>>> * @res: "host memory" address range
>>> * @ref: a live per-cpu reference count
>>> * @altmap: optional descriptor for allocating the memmap from @res
>>> + * @ppgmap: pointer set to new page dev_pagemap on success
>>> + * @flags: flag for memory (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
>>> *
>>> * Notes:
>>> * 1/ @ref must be 'live' on entry and 'dead' before devm_memunmap_pages() time
>>> @@ -280,7 +282,8 @@ struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
>>> * this is not enforced.
>>> */
>>> void *devm_memremap_pages(struct device *dev, struct resource *res,
>>> - struct percpu_ref *ref, struct vmem_altmap *altmap)
>>> + struct percpu_ref *ref, struct vmem_altmap *altmap,
>>> + struct dev_pagemap **ppgmap, int flags)
>>> {
>>> resource_size_t key, align_start, align_size, align_end;
>>> pgprot_t pgprot = PAGE_KERNEL;
>>> @@ -322,6 +325,7 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
>>> }
>>> pgmap->ref = ref;
>>> pgmap->res = &page_map->res;
>>> + pgmap->flags = flags | MEMORY_DEVICE;
>>
>> So the caller of devm_memremap_pages() should not have give out MEMORY_DEVICE
>> in the flag it passed on to this function ? Hmm, else we should just check
>> that the flags contains all appropriate bits before proceeding.
>
> Here i was just trying to be on the safe side, yes caller should already have set
> the flag but this function is only use for device memory so it did not seem like
> it would hurt to be extra safe. I can add a BUG_ON() but it seems people have mix
> feeling about BUG_ON()
We dont have to do BUG_ON(), just a check that all expected flags are
in there, else fail the call. Now this function does not return any
value to be checked inside driver, in that case we can just do a error
message print and move on.
[toc] | [prev] | [next] | [standalone]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-18 18:30 +0100 |
| Subject | [HMM v13 09/18] mm/hmm/mirror: mirror process address space on device with HMM helpers |
| Message-ID | <sEV97-1Es-9@gated-at.bofh.it> |
| In reply to | #1525568 |
This is a heterogeneous memory management (HMM) process address space
mirroring. In a nutshell this provide an API to mirror process address
space on a device. This boils down to keeping CPU and device page table
synchronize (we assume that both device and CPU are cache coherent like
PCIe device can be).
This patch provide a simple API for device driver to achieve address
space mirroring thus avoiding each device driver to grow its own CPU
page table walker and its own CPU page table synchronization mechanism.
This is usefull for NVidia GPU >= Pascal, Mellanox IB >= mlx5 and more
hardware in the future.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Signed-off-by: Jatin Kumar <jakumar@nvidia.com>
Signed-off-by: John Hubbard <jhubbard@nvidia.com>
Signed-off-by: Mark Hairgrove <mhairgrove@nvidia.com>
Signed-off-by: Sherry Cheung <SCheung@nvidia.com>
Signed-off-by: Subhash Gutti <sgutti@nvidia.com>
---
include/linux/hmm.h | 97 +++++++++++++++++++++++++++++++
mm/hmm.c | 160 ++++++++++++++++++++++++++++++++++++++++++++++++++++
2 files changed, 257 insertions(+)
diff --git a/include/linux/hmm.h b/include/linux/hmm.h
index 54dd529..f44e270 100644
--- a/include/linux/hmm.h
+++ b/include/linux/hmm.h
@@ -88,6 +88,7 @@
#if IS_ENABLED(CONFIG_HMM)
+struct hmm;
/*
* hmm_pfn_t - HMM use its own pfn type to keep several flags per page
@@ -127,6 +128,102 @@ static inline hmm_pfn_t hmm_pfn_from_pfn(unsigned long pfn)
}
+/*
+ * Mirroring: how to use synchronize device page table with CPU page table ?
+ *
+ * Device driver must always synchronize with CPU page table update, for this
+ * they can either directly use mmu_notifier API or they can use the hmm_mirror
+ * API. Device driver can decide to register one mirror per device per process
+ * or just one mirror per process for a group of device. Pattern is :
+ *
+ * int device_bind_address_space(..., struct mm_struct *mm, ...)
+ * {
+ * struct device_address_space *das;
+ * int ret;
+ * // Device driver specific initialization, and allocation of das
+ * // which contain an hmm_mirror struct as one of its field.
+ * ret = hmm_mirror_register(&das->mirror, mm, &device_mirror_ops);
+ * if (ret) {
+ * // Cleanup on error
+ * return ret;
+ * }
+ * // Other device driver specific initialization
+ * }
+ *
+ * Device driver must not free the struct containing hmm_mirror struct before
+ * calling hmm_mirror_unregister() expected usage is to do that when device
+ * driver is unbinding from an address space.
+ *
+ * void device_unbind_address_space(struct device_address_space *das)
+ * {
+ * // Device driver specific cleanup
+ * hmm_mirror_unregister(&das->mirror);
+ * // Other device driver specific cleanup and now das can be free
+ * }
+ *
+ * Once an hmm_mirror is register for an address space, device driver will get
+ * callback through the update() operation (see hmm_mirror_ops struct).
+ */
+
+struct hmm_mirror;
+
+/*
+ * enum hmm_update - type of update
+ * @HMM_UPDATE_INVALIDATE: invalidate range (no indication as to why)
+ */
+enum hmm_update {
+ HMM_UPDATE_INVALIDATE,
+};
+
+/*
+ * struct hmm_mirror_ops - HMM mirror device operations callback
+ *
+ * @update: callback to update range on a device
+ */
+struct hmm_mirror_ops {
+ /* update() - update virtual address range of memory
+ *
+ * @mirror: pointer to struct hmm_mirror
+ * @update: update's type (turn read only, unmap, ...)
+ * @start: virtual start address of the range to update
+ * @end: virtual end address of the range to update
+ *
+ * This callback is call when the CPU page table is updated, the device
+ * driver must update device page table accordingly to update's action.
+ *
+ * Device driver callback must wait until device have fully updated its
+ * view for the range. Note we plan to make this asynchronous in later
+ * patches. So that multiple devices can schedule update to their page
+ * table and once all device have schedule the update then we wait for
+ * them to propagate.
+ */
+ void (*update)(struct hmm_mirror *mirror,
+ enum hmm_update action,
+ unsigned long start,
+ unsigned long end);
+};
+
+/*
+ * struct hmm_mirror - mirror struct for a device driver
+ *
+ * @hmm: pointer to struct hmm (which is unique per mm_struct)
+ * @ops: device driver callback for HMM mirror operations
+ * @list: for list of mirrors of a given mm
+ *
+ * Each address space (mm_struct) being mirrored by a device must register one
+ * of hmm_mirror struct with HMM. HMM will track list of all mirrors for each
+ * mm_struct (or each process).
+ */
+struct hmm_mirror {
+ struct hmm *hmm;
+ const struct hmm_mirror_ops *ops;
+ struct list_head list;
+};
+
+int hmm_mirror_register(struct hmm_mirror *mirror, struct mm_struct *mm);
+void hmm_mirror_unregister(struct hmm_mirror *mirror);
+
+
/* Below are for HMM internal use only ! Not to be use by device driver ! */
void hmm_mm_destroy(struct mm_struct *mm);
diff --git a/mm/hmm.c b/mm/hmm.c
index 342b596..3594785 100644
--- a/mm/hmm.c
+++ b/mm/hmm.c
@@ -21,14 +21,27 @@
#include <linux/hmm.h>
#include <linux/slab.h>
#include <linux/sched.h>
+#include <linux/mmu_notifier.h>
/*
* struct hmm - HMM per mm struct
*
* @mm: mm struct this HMM struct is bound to
+ * @lock: lock protecting mirrors list
+ * @mirrors: list of mirrors for this mm
+ * @wait_queue: wait queue
+ * @sequence: we track update to CPU page table with a sequence number
+ * @mmu_notifier: mmu notifier to track update to CPU page table
+ * @notifier_count: number of currently active notifier count
*/
struct hmm {
struct mm_struct *mm;
+ spinlock_t lock;
+ struct list_head mirrors;
+ atomic_t sequence;
+ wait_queue_head_t wait_queue;
+ struct mmu_notifier mmu_notifier;
+ atomic_t notifier_count;
};
/*
@@ -48,6 +61,12 @@ static struct hmm *hmm_register(struct mm_struct *mm)
hmm = kmalloc(sizeof(*hmm), GFP_KERNEL);
if (!hmm)
return NULL;
+ init_waitqueue_head(&hmm->wait_queue);
+ atomic_set(&hmm->notifier_count, 0);
+ INIT_LIST_HEAD(&hmm->mirrors);
+ atomic_set(&hmm->sequence, 0);
+ hmm->mmu_notifier.ops = NULL;
+ spin_lock_init(&hmm->lock);
hmm->mm = mm;
}
@@ -84,3 +103,144 @@ void hmm_mm_destroy(struct mm_struct *mm)
kfree(hmm);
}
+
+
+
+static void hmm_invalidate_range(struct hmm *hmm,
+ enum hmm_update action,
+ unsigned long start,
+ unsigned long end)
+{
+ struct hmm_mirror *mirror;
+
+ /*
+ * Mirror being added or remove is a rare event so list traversal isn't
+ * protected by a lock, we rely on simple rules. All list modification
+ * are done using list_add_rcu() and list_del_rcu() under a spinlock to
+ * protect from concurrent addition or removal but not traversal.
+ *
+ * Because hmm_mirror_unregister() wait for all running invalidation to
+ * complete (and thus all list traversal to finish). None of the mirror
+ * struct can be freed from under us while traversing the list and thus
+ * it is safe to dereference their list pointer even if they were just
+ * remove.
+ */
+ list_for_each_entry (mirror, &hmm->mirrors, list)
+ mirror->ops->update(mirror, action, start, end);
+}
+
+static void hmm_invalidate_page(struct mmu_notifier *mn,
+ struct mm_struct *mm,
+ unsigned long addr)
+{
+ unsigned long start = addr & PAGE_MASK;
+ unsigned long end = start + PAGE_SIZE;
+ struct hmm *hmm = mm->hmm;
+
+ VM_BUG_ON(!hmm);
+
+ atomic_inc(&hmm->notifier_count);
+ atomic_inc(&hmm->sequence);
+ hmm_invalidate_range(mm->hmm, HMM_UPDATE_INVALIDATE, start, end);
+ atomic_dec(&hmm->notifier_count);
+ wake_up(&hmm->wait_queue);
+}
+
+static void hmm_invalidate_range_start(struct mmu_notifier *mn,
+ struct mm_struct *mm,
+ unsigned long start,
+ unsigned long end)
+{
+ struct hmm *hmm = mm->hmm;
+
+ VM_BUG_ON(!hmm);
+
+ atomic_inc(&hmm->notifier_count);
+ atomic_inc(&hmm->sequence);
+ hmm_invalidate_range(mm->hmm, HMM_UPDATE_INVALIDATE, start, end);
+}
+
+static void hmm_invalidate_range_end(struct mmu_notifier *mn,
+ struct mm_struct *mm,
+ unsigned long start,
+ unsigned long end)
+{
+ struct hmm *hmm = mm->hmm;
+
+ VM_BUG_ON(!hmm);
+
+ /* Reverse order here because we are getting out of invalidation */
+ atomic_dec(&hmm->notifier_count);
+ wake_up(&hmm->wait_queue);
+}
+
+static const struct mmu_notifier_ops hmm_mmu_notifier_ops = {
+ .invalidate_page = hmm_invalidate_page,
+ .invalidate_range_start = hmm_invalidate_range_start,
+ .invalidate_range_end = hmm_invalidate_range_end,
+};
+
+/*
+ * hmm_mirror_register() - register a mirror against an mm
+ *
+ * @mirror: new mirror struct to register
+ * @mm: mm to register against
+ *
+ * To start mirroring a process address space device driver must register an
+ * HMM mirror struct.
+ */
+int hmm_mirror_register(struct hmm_mirror *mirror, struct mm_struct *mm)
+{
+ /* Sanity check */
+ if (!mm || !mirror || !mirror->ops)
+ return -EINVAL;
+
+ mirror->hmm = hmm_register(mm);
+ if (!mirror->hmm)
+ return -ENOMEM;
+
+ /* Register mmu_notifier if not already, use mmap_sem for locking */
+ if (!mirror->hmm->mmu_notifier.ops) {
+ struct hmm *hmm = mirror->hmm;
+ down_write(&mm->mmap_sem);
+ if (!hmm->mmu_notifier.ops) {
+ hmm->mmu_notifier.ops = &hmm_mmu_notifier_ops;
+ if (__mmu_notifier_register(&hmm->mmu_notifier, mm)) {
+ hmm->mmu_notifier.ops = NULL;
+ up_write(&mm->mmap_sem);
+ return -ENOMEM;
+ }
+ }
+ up_write(&mm->mmap_sem);
+ }
+
+ spin_lock(&mirror->hmm->lock);
+ list_add_rcu(&mirror->list, &mirror->hmm->mirrors);
+ spin_unlock(&mirror->hmm->lock);
+
+ return 0;
+}
+EXPORT_SYMBOL(hmm_mirror_register);
+
+/*
+ * hmm_mirror_unregister() - unregister a mirror
+ *
+ * @mirror: new mirror struct to register
+ *
+ * Stop mirroring a process address space and cleanup.
+ */
+void hmm_mirror_unregister(struct hmm_mirror *mirror)
+{
+ struct hmm *hmm = mirror->hmm;
+
+ spin_lock(&hmm->lock);
+ list_del_rcu(&mirror->list);
+ spin_unlock(&hmm->lock);
+
+ /*
+ * Wait for all active notifier so that it is safe to traverse mirror
+ * list without any lock.
+ */
+ wait_event(hmm->wait_queue, !atomic_read(&hmm->notifier_count));
+}
+EXPORT_SYMBOL(hmm_mirror_unregister);
--
2.4.3
[toc] | [prev] | [next] | [standalone]
| From | Balbir Singh <bsingharora@gmail.com> |
|---|---|
| Date | 2016-11-21 03:50 +0100 |
| Subject | Re: [HMM v13 09/18] mm/hmm/mirror: mirror process address space on device with HMM helpers |
| Message-ID | <sFMQ9-3uc-5@gated-at.bofh.it> |
| In reply to | #1525574 |
On 19/11/16 05:18, Jérôme Glisse wrote:
> This is a heterogeneous memory management (HMM) process address space
> mirroring. In a nutshell this provide an API to mirror process address
> space on a device. This boils down to keeping CPU and device page table
> synchronize (we assume that both device and CPU are cache coherent like
> PCIe device can be).
>
> This patch provide a simple API for device driver to achieve address
> space mirroring thus avoiding each device driver to grow its own CPU
> page table walker and its own CPU page table synchronization mechanism.
>
> This is usefull for NVidia GPU >= Pascal, Mellanox IB >= mlx5 and more
useful
> hardware in the future.
>
> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
> Signed-off-by: Jatin Kumar <jakumar@nvidia.com>
> Signed-off-by: John Hubbard <jhubbard@nvidia.com>
> Signed-off-by: Mark Hairgrove <mhairgrove@nvidia.com>
> Signed-off-by: Sherry Cheung <SCheung@nvidia.com>
> Signed-off-by: Subhash Gutti <sgutti@nvidia.com>
> ---
> include/linux/hmm.h | 97 +++++++++++++++++++++++++++++++
> mm/hmm.c | 160 ++++++++++++++++++++++++++++++++++++++++++++++++++++
> 2 files changed, 257 insertions(+)
>
> diff --git a/include/linux/hmm.h b/include/linux/hmm.h
> index 54dd529..f44e270 100644
> --- a/include/linux/hmm.h
> +++ b/include/linux/hmm.h
> @@ -88,6 +88,7 @@
>
> #if IS_ENABLED(CONFIG_HMM)
>
> +struct hmm;
>
> /*
> * hmm_pfn_t - HMM use its own pfn type to keep several flags per page
> @@ -127,6 +128,102 @@ static inline hmm_pfn_t hmm_pfn_from_pfn(unsigned long pfn)
> }
>
>
> +/*
> + * Mirroring: how to use synchronize device page table with CPU page table ?
> + *
> + * Device driver must always synchronize with CPU page table update, for this
> + * they can either directly use mmu_notifier API or they can use the hmm_mirror
> + * API. Device driver can decide to register one mirror per device per process
> + * or just one mirror per process for a group of device. Pattern is :
> + *
> + * int device_bind_address_space(..., struct mm_struct *mm, ...)
> + * {
> + * struct device_address_space *das;
> + * int ret;
> + * // Device driver specific initialization, and allocation of das
> + * // which contain an hmm_mirror struct as one of its field.
> + * ret = hmm_mirror_register(&das->mirror, mm, &device_mirror_ops);
> + * if (ret) {
> + * // Cleanup on error
> + * return ret;
> + * }
> + * // Other device driver specific initialization
> + * }
> + *
> + * Device driver must not free the struct containing hmm_mirror struct before
> + * calling hmm_mirror_unregister() expected usage is to do that when device
> + * driver is unbinding from an address space.
> + *
> + * void device_unbind_address_space(struct device_address_space *das)
> + * {
> + * // Device driver specific cleanup
> + * hmm_mirror_unregister(&das->mirror);
> + * // Other device driver specific cleanup and now das can be free
> + * }
> + *
> + * Once an hmm_mirror is register for an address space, device driver will get
> + * callback through the update() operation (see hmm_mirror_ops struct).
> + */
> +
> +struct hmm_mirror;
> +
> +/*
> + * enum hmm_update - type of update
> + * @HMM_UPDATE_INVALIDATE: invalidate range (no indication as to why)
> + */
> +enum hmm_update {
> + HMM_UPDATE_INVALIDATE,
> +};
> +
> +/*
> + * struct hmm_mirror_ops - HMM mirror device operations callback
> + *
> + * @update: callback to update range on a device
> + */
> +struct hmm_mirror_ops {
> + /* update() - update virtual address range of memory
> + *
> + * @mirror: pointer to struct hmm_mirror
> + * @update: update's type (turn read only, unmap, ...)
> + * @start: virtual start address of the range to update
> + * @end: virtual end address of the range to update
> + *
> + * This callback is call when the CPU page table is updated, the device
> + * driver must update device page table accordingly to update's action.
> + *
> + * Device driver callback must wait until device have fully updated its
> + * view for the range. Note we plan to make this asynchronous in later
> + * patches. So that multiple devices can schedule update to their page
> + * table and once all device have schedule the update then we wait for
> + * them to propagate.
> + */
> + void (*update)(struct hmm_mirror *mirror,
> + enum hmm_update action,
> + unsigned long start,
> + unsigned long end);
> +};
> +
> +/*
> + * struct hmm_mirror - mirror struct for a device driver
> + *
> + * @hmm: pointer to struct hmm (which is unique per mm_struct)
> + * @ops: device driver callback for HMM mirror operations
> + * @list: for list of mirrors of a given mm
> + *
> + * Each address space (mm_struct) being mirrored by a device must register one
> + * of hmm_mirror struct with HMM. HMM will track list of all mirrors for each
> + * mm_struct (or each process).
> + */
> +struct hmm_mirror {
> + struct hmm *hmm;
> + const struct hmm_mirror_ops *ops;
> + struct list_head list;
> +};
> +
> +int hmm_mirror_register(struct hmm_mirror *mirror, struct mm_struct *mm);
> +void hmm_mirror_unregister(struct hmm_mirror *mirror);
> +
> +
> /* Below are for HMM internal use only ! Not to be use by device driver ! */
> void hmm_mm_destroy(struct mm_struct *mm);
>
> diff --git a/mm/hmm.c b/mm/hmm.c
> index 342b596..3594785 100644
> --- a/mm/hmm.c
> +++ b/mm/hmm.c
> @@ -21,14 +21,27 @@
> #include <linux/hmm.h>
> #include <linux/slab.h>
> #include <linux/sched.h>
> +#include <linux/mmu_notifier.h>
>
> /*
> * struct hmm - HMM per mm struct
> *
> * @mm: mm struct this HMM struct is bound to
> + * @lock: lock protecting mirrors list
> + * @mirrors: list of mirrors for this mm
> + * @wait_queue: wait queue
> + * @sequence: we track update to CPU page table with a sequence number
> + * @mmu_notifier: mmu notifier to track update to CPU page table
> + * @notifier_count: number of currently active notifier count
> */
> struct hmm {
> struct mm_struct *mm;
> + spinlock_t lock;
> + struct list_head mirrors;
> + atomic_t sequence;
> + wait_queue_head_t wait_queue;
> + struct mmu_notifier mmu_notifier;
> + atomic_t notifier_count;
> };
>
> /*
> @@ -48,6 +61,12 @@ static struct hmm *hmm_register(struct mm_struct *mm)
> hmm = kmalloc(sizeof(*hmm), GFP_KERNEL);
> if (!hmm)
> return NULL;
> + init_waitqueue_head(&hmm->wait_queue);
> + atomic_set(&hmm->notifier_count, 0);
> + INIT_LIST_HEAD(&hmm->mirrors);
> + atomic_set(&hmm->sequence, 0);
> + hmm->mmu_notifier.ops = NULL;
> + spin_lock_init(&hmm->lock);
> hmm->mm = mm;
> }
>
> @@ -84,3 +103,144 @@ void hmm_mm_destroy(struct mm_struct *mm)
>
> kfree(hmm);
> }
> +
> +
> +
> +static void hmm_invalidate_range(struct hmm *hmm,
> + enum hmm_update action,
> + unsigned long start,
> + unsigned long end)
> +{
> + struct hmm_mirror *mirror;
> +
> + /*
> + * Mirror being added or remove is a rare event so list traversal isn't
> + * protected by a lock, we rely on simple rules. All list modification
> + * are done using list_add_rcu() and list_del_rcu() under a spinlock to
> + * protect from concurrent addition or removal but not traversal.
> + *
> + * Because hmm_mirror_unregister() wait for all running invalidation to
> + * complete (and thus all list traversal to finish). None of the mirror
> + * struct can be freed from under us while traversing the list and thus
> + * it is safe to dereference their list pointer even if they were just
> + * remove.
> + */
> + list_for_each_entry (mirror, &hmm->mirrors, list)
> + mirror->ops->update(mirror, action, start, end);
> +}
> +
> +static void hmm_invalidate_page(struct mmu_notifier *mn,
> + struct mm_struct *mm,
> + unsigned long addr)
> +{
> + unsigned long start = addr & PAGE_MASK;
> + unsigned long end = start + PAGE_SIZE;
> + struct hmm *hmm = mm->hmm;
> +
> + VM_BUG_ON(!hmm);
> +
> + atomic_inc(&hmm->notifier_count);
> + atomic_inc(&hmm->sequence);
> + hmm_invalidate_range(mm->hmm, HMM_UPDATE_INVALIDATE, start, end);
> + atomic_dec(&hmm->notifier_count);
> + wake_up(&hmm->wait_queue);
> +}
> +
> +static void hmm_invalidate_range_start(struct mmu_notifier *mn,
> + struct mm_struct *mm,
> + unsigned long start,
> + unsigned long end)
> +{
> + struct hmm *hmm = mm->hmm;
> +
> + VM_BUG_ON(!hmm);
> +
> + atomic_inc(&hmm->notifier_count);
> + atomic_inc(&hmm->sequence);
> + hmm_invalidate_range(mm->hmm, HMM_UPDATE_INVALIDATE, start, end);
> +}
> +
> +static void hmm_invalidate_range_end(struct mmu_notifier *mn,
> + struct mm_struct *mm,
> + unsigned long start,
> + unsigned long end)
> +{
> + struct hmm *hmm = mm->hmm;
> +
> + VM_BUG_ON(!hmm);
> +
> + /* Reverse order here because we are getting out of invalidation */
> + atomic_dec(&hmm->notifier_count);
> + wake_up(&hmm->wait_queue);
> +}
> +
> +static const struct mmu_notifier_ops hmm_mmu_notifier_ops = {
> + .invalidate_page = hmm_invalidate_page,
> + .invalidate_range_start = hmm_invalidate_range_start,
> + .invalidate_range_end = hmm_invalidate_range_end,
> +};
> +
> +/*
> + * hmm_mirror_register() - register a mirror against an mm
> + *
> + * @mirror: new mirror struct to register
> + * @mm: mm to register against
> + *
> + * To start mirroring a process address space device driver must register an
> + * HMM mirror struct.
> + */
> +int hmm_mirror_register(struct hmm_mirror *mirror, struct mm_struct *mm)
> +{
> + /* Sanity check */
> + if (!mm || !mirror || !mirror->ops)
> + return -EINVAL;
> +
> + mirror->hmm = hmm_register(mm);
> + if (!mirror->hmm)
> + return -ENOMEM;
> +
> + /* Register mmu_notifier if not already, use mmap_sem for locking */
> + if (!mirror->hmm->mmu_notifier.ops) {
> + struct hmm *hmm = mirror->hmm;
> + down_write(&mm->mmap_sem);
> + if (!hmm->mmu_notifier.ops) {
> + hmm->mmu_notifier.ops = &hmm_mmu_notifier_ops;
> + if (__mmu_notifier_register(&hmm->mmu_notifier, mm)) {
> + hmm->mmu_notifier.ops = NULL;
> + up_write(&mm->mmap_sem);
> + return -ENOMEM;
> + }
> + }
> + up_write(&mm->mmap_sem);
> + }
Does everything get mirrored, every update to the PTE (clear dirty, clear
accessed bit, etc) or does the driver decide?
> +
> + spin_lock(&mirror->hmm->lock);
> + list_add_rcu(&mirror->list, &mirror->hmm->mirrors);
> + spin_unlock(&mirror->hmm->lock);
> +
> + return 0;
> +}
> +EXPORT_SYMBOL(hmm_mirror_register);
> +
> +/*
> + * hmm_mirror_unregister() - unregister a mirror
> + *
> + * @mirror: new mirror struct to register
> + *
> + * Stop mirroring a process address space and cleanup.
> + */
> +void hmm_mirror_unregister(struct hmm_mirror *mirror)
> +{
> + struct hmm *hmm = mirror->hmm;
> +
> + spin_lock(&hmm->lock);
> + list_del_rcu(&mirror->list);
> + spin_unlock(&hmm->lock);
> +
> + /*
> + * Wait for all active notifier so that it is safe to traverse mirror
> + * list without any lock.
> + */
> + wait_event(hmm->wait_queue, !atomic_read(&hmm->notifier_count));
> +}
> +EXPORT_SYMBOL(hmm_mirror_unregister);
>
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-21 06:20 +0100 |
| Subject | Re: [HMM v13 09/18] mm/hmm/mirror: mirror process address space on device with HMM helpers |
| Message-ID | <sFPbj-5cE-5@gated-at.bofh.it> |
| In reply to | #1526338 |
On Mon, Nov 21, 2016 at 01:42:43PM +1100, Balbir Singh wrote:
> On 19/11/16 05:18, Jérôme Glisse wrote:
[...]
> > +/*
> > + * hmm_mirror_register() - register a mirror against an mm
> > + *
> > + * @mirror: new mirror struct to register
> > + * @mm: mm to register against
> > + *
> > + * To start mirroring a process address space device driver must register an
> > + * HMM mirror struct.
> > + */
> > +int hmm_mirror_register(struct hmm_mirror *mirror, struct mm_struct *mm)
> > +{
> > + /* Sanity check */
> > + if (!mm || !mirror || !mirror->ops)
> > + return -EINVAL;
> > +
> > + mirror->hmm = hmm_register(mm);
> > + if (!mirror->hmm)
> > + return -ENOMEM;
> > +
> > + /* Register mmu_notifier if not already, use mmap_sem for locking */
> > + if (!mirror->hmm->mmu_notifier.ops) {
> > + struct hmm *hmm = mirror->hmm;
> > + down_write(&mm->mmap_sem);
> > + if (!hmm->mmu_notifier.ops) {
> > + hmm->mmu_notifier.ops = &hmm_mmu_notifier_ops;
> > + if (__mmu_notifier_register(&hmm->mmu_notifier, mm)) {
> > + hmm->mmu_notifier.ops = NULL;
> > + up_write(&mm->mmap_sem);
> > + return -ENOMEM;
> > + }
> > + }
> > + up_write(&mm->mmap_sem);
> > + }
>
> Does everything get mirrored, every update to the PTE (clear dirty, clear
> accessed bit, etc) or does the driver decide?
Driver decide but only read/write/valid matter for device. Device driver must
report dirtyness on invalidation. Some device do not have access bit and thus
can't provide that information.
The idea here is really to snapshot the CPU page table and duplicate it as
a GPU page table. The only synchronization HMM provide is that each virtual
address point to same memory at that at no point in time the same virtual
address can point to different physical memory on the device and on the CPU.
Cheers,
Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-18 18:30 +0100 |
| Subject | [HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory |
| Message-ID | <sEV98-1Es-13@gated-at.bofh.it> |
| In reply to | #1525568 |
HMM wants to remove device memory early before device tear down so add an
helper to do that.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Cc: Dan Williams <dan.j.williams@intel.com>
Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
---
include/linux/memremap.h | 7 +++++++
kernel/memremap.c | 14 ++++++++++++++
2 files changed, 21 insertions(+)
diff --git a/include/linux/memremap.h b/include/linux/memremap.h
index 469c88d..b6f03e9 100644
--- a/include/linux/memremap.h
+++ b/include/linux/memremap.h
@@ -60,6 +60,7 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
struct percpu_ref *ref, struct vmem_altmap *altmap,
struct dev_pagemap **ppgmap, int flags);
struct dev_pagemap *find_dev_pagemap(resource_size_t phys);
+int devm_memremap_pages_remove(struct device *dev, struct dev_pagemap *pgmap);
static inline bool is_addressable_page(const struct page *page)
{
@@ -88,6 +89,12 @@ static inline struct dev_pagemap *find_dev_pagemap(resource_size_t phys)
return NULL;
}
+static inline int devm_memremap_pages_remove(struct device *dev,
+ struct dev_pagemap *pgmap)
+{
+ return -EINVAL;
+}
+
static inline bool is_addressable_page(const struct page *page)
{
return true;
diff --git a/kernel/memremap.c b/kernel/memremap.c
index 3d28048..cf83928 100644
--- a/kernel/memremap.c
+++ b/kernel/memremap.c
@@ -401,6 +401,20 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
}
EXPORT_SYMBOL(devm_memremap_pages);
+static int devm_page_map_match(struct device *dev, void *data, void *match_data)
+{
+ struct page_map *page_map = data;
+
+ return &page_map->pgmap == match_data;
+}
+
+int devm_memremap_pages_remove(struct device *dev, struct dev_pagemap *pgmap)
+{
+ return devres_release(dev, &devm_memremap_pages_release,
+ &devm_page_map_match, pgmap);
+}
+EXPORT_SYMBOL(devm_memremap_pages_remove);
+
unsigned long vmem_altmap_offset(struct vmem_altmap *altmap)
{
/* number of pfns from base where pfn_to_page() is valid */
--
2.4.3
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-21 11:40 +0100 |
| Subject | Re: [HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory |
| Message-ID | <sFUaZ-8sg-9@gated-at.bofh.it> |
| In reply to | #1525575 |
On 11/18/2016 11:48 PM, Jérôme Glisse wrote: > HMM wants to remove device memory early before device tear down so add an > helper to do that. Could you please explain why HMM wants to remove device memory before device tear down ?
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-21 13:40 +0100 |
| Subject | Re: [HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory |
| Message-ID | <sFW38-1bK-15@gated-at.bofh.it> |
| In reply to | #1526539 |
On Mon, Nov 21, 2016 at 04:07:46PM +0530, Anshuman Khandual wrote: > On 11/18/2016 11:48 PM, Jérôme Glisse wrote: > > HMM wants to remove device memory early before device tear down so add an > > helper to do that. > > Could you please explain why HMM wants to remove device memory before > device tear down ? > Some device driver want to manage memory for several physical devices from a single fake device driver. Because it fits their driver architecture better and those physical devices can have dedicated link between them. Issue is that the fake device driver can outlive any of the real device for a long time so we want to be able to remove device memory before the fake device goes away to free up resources early. Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-22 06:00 +0100 |
| Subject | Re: [HMM v13 05/18] mm/ZONE_DEVICE/devmem_pages_remove: allow early removal of device memory |
| Message-ID | <sGblv-2tj-3@gated-at.bofh.it> |
| In reply to | #1526633 |
On 11/21/2016 06:09 PM, Jerome Glisse wrote: > On Mon, Nov 21, 2016 at 04:07:46PM +0530, Anshuman Khandual wrote: >> On 11/18/2016 11:48 PM, Jérôme Glisse wrote: >>> HMM wants to remove device memory early before device tear down so add an >>> helper to do that. >> >> Could you please explain why HMM wants to remove device memory before >> device tear down ? >> > > Some device driver want to manage memory for several physical devices from a > single fake device driver. Because it fits their driver architecture better > and those physical devices can have dedicated link between them. > > Issue is that the fake device driver can outlive any of the real device for a > long time so we want to be able to remove device memory before the fake device > goes away to free up resources early. Got it.
[toc] | [prev] | [next] | [standalone]
| From | Jérôme Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-18 18:30 +0100 |
| Subject | [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed |
| Message-ID | <sEV98-1Es-25@gated-at.bofh.it> |
| In reply to | #1525568 |
When a ZONE_DEVICE page refcount reach 1 it means it is free and nobody
is holding a reference on it (only device to which the memory belong do).
Add a callback and call it when that happen so device driver can implement
their own free page management.
Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
Cc: Dan Williams <dan.j.williams@intel.com>
Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
---
include/linux/memremap.h | 4 ++++
kernel/memremap.c | 8 ++++++++
2 files changed, 12 insertions(+)
diff --git a/include/linux/memremap.h b/include/linux/memremap.h
index fe61dca..469c88d 100644
--- a/include/linux/memremap.h
+++ b/include/linux/memremap.h
@@ -37,17 +37,21 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
/**
* struct dev_pagemap - metadata for ZONE_DEVICE mappings
+ * @free_devpage: free page callback when page refcount reach 1
* @altmap: pre-allocated/reserved memory for vmemmap allocations
* @res: physical address range covered by @ref
* @ref: reference count that pins the devm_memremap_pages() mapping
* @dev: host device of the mapping for debug
+ * @data: privata data pointer for free_devpage
* @flags: memory flags (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
*/
struct dev_pagemap {
+ void (*free_devpage)(struct page *page, void *data);
struct vmem_altmap *altmap;
const struct resource *res;
struct percpu_ref *ref;
struct device *dev;
+ void *data;
int flags;
};
diff --git a/kernel/memremap.c b/kernel/memremap.c
index 438a73aa2..3d28048 100644
--- a/kernel/memremap.c
+++ b/kernel/memremap.c
@@ -190,6 +190,12 @@ EXPORT_SYMBOL(get_zone_device_page);
void put_zone_device_page(struct page *page)
{
+ /*
+ * If refcount is 1 then page is freed and refcount is stable as nobody
+ * holds a reference on the page.
+ */
+ if (page->pgmap->free_devpage && page_count(page) == 1)
+ page->pgmap->free_devpage(page, page->pgmap->data);
put_dev_pagemap(page->pgmap);
}
EXPORT_SYMBOL(put_zone_device_page);
@@ -326,6 +332,8 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
pgmap->ref = ref;
pgmap->res = &page_map->res;
pgmap->flags = flags | MEMORY_DEVICE;
+ pgmap->free_devpage = NULL;
+ pgmap->data = NULL;
mutex_lock(&pgmap_lock);
error = 0;
--
2.4.3
[toc] | [prev] | [next] | [standalone]
| From | Balbir Singh <bsingharora@gmail.com> |
|---|---|
| Date | 2016-11-21 03:00 +0100 |
| Subject | Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed |
| Message-ID | <sFM3L-2Wa-3@gated-at.bofh.it> |
| In reply to | #1525576 |
On 19/11/16 05:18, Jérôme Glisse wrote: > When a ZONE_DEVICE page refcount reach 1 it means it is free and nobody > is holding a reference on it (only device to which the memory belong do). > Add a callback and call it when that happen so device driver can implement > their own free page management. > Could you give an example of what their own free page management might look like? Balbir Singh.
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-21 06:00 +0100 |
| Subject | Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed |
| Message-ID | <sFORX-4Q0-7@gated-at.bofh.it> |
| In reply to | #1526328 |
On Mon, Nov 21, 2016 at 12:49:55PM +1100, Balbir Singh wrote: > On 19/11/16 05:18, Jérôme Glisse wrote: > > When a ZONE_DEVICE page refcount reach 1 it means it is free and nobody > > is holding a reference on it (only device to which the memory belong do). > > Add a callback and call it when that happen so device driver can implement > > their own free page management. > > > > Could you give an example of what their own free page management might look like? Well hard to do that, the free management is whatever the device driver want to do. So i don't have any example to give. Each device driver (especialy GPU ones) have their own memory management with little commonality. So how the device driver manage that memory is really not important, at least it is not something for which i want to impose a policy onto driver. I want to leave each device driver decide on how to achieve that. Cheers, Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-21 09:30 +0100 |
| Subject | Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed |
| Message-ID | <sFS9c-7d8-9@gated-at.bofh.it> |
| In reply to | #1525576 |
On 11/18/2016 11:48 PM, Jérôme Glisse wrote:
> When a ZONE_DEVICE page refcount reach 1 it means it is free and nobody
> is holding a reference on it (only device to which the memory belong do).
> Add a callback and call it when that happen so device driver can implement
> their own free page management.
>
> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
> Cc: Dan Williams <dan.j.williams@intel.com>
> Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
> ---
> include/linux/memremap.h | 4 ++++
> kernel/memremap.c | 8 ++++++++
> 2 files changed, 12 insertions(+)
>
> diff --git a/include/linux/memremap.h b/include/linux/memremap.h
> index fe61dca..469c88d 100644
> --- a/include/linux/memremap.h
> +++ b/include/linux/memremap.h
> @@ -37,17 +37,21 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
>
> /**
> * struct dev_pagemap - metadata for ZONE_DEVICE mappings
> + * @free_devpage: free page callback when page refcount reach 1
> * @altmap: pre-allocated/reserved memory for vmemmap allocations
> * @res: physical address range covered by @ref
> * @ref: reference count that pins the devm_memremap_pages() mapping
> * @dev: host device of the mapping for debug
> + * @data: privata data pointer for free_devpage
> * @flags: memory flags (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
> */
> struct dev_pagemap {
> + void (*free_devpage)(struct page *page, void *data);
> struct vmem_altmap *altmap;
> const struct resource *res;
> struct percpu_ref *ref;
> struct device *dev;
> + void *data;
> int flags;
> };
>
> diff --git a/kernel/memremap.c b/kernel/memremap.c
> index 438a73aa2..3d28048 100644
> --- a/kernel/memremap.c
> +++ b/kernel/memremap.c
> @@ -190,6 +190,12 @@ EXPORT_SYMBOL(get_zone_device_page);
>
> void put_zone_device_page(struct page *page)
> {
> + /*
> + * If refcount is 1 then page is freed and refcount is stable as nobody
> + * holds a reference on the page.
> + */
> + if (page->pgmap->free_devpage && page_count(page) == 1)
> + page->pgmap->free_devpage(page, page->pgmap->data);
> put_dev_pagemap(page->pgmap);
> }
> EXPORT_SYMBOL(put_zone_device_page);
> @@ -326,6 +332,8 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
> pgmap->ref = ref;
> pgmap->res = &page_map->res;
> pgmap->flags = flags | MEMORY_DEVICE;
> + pgmap->free_devpage = NULL;
> + pgmap->data = NULL;
When is the driver expected to load up pgmap->free_devpage ? I thought
this function is one of the right places. Though as all the pages in
the same hotplug operation point to the same dev_pagemap structure this
loading can be done at later point of time as well.
[toc] | [prev] | [next] | [standalone]
| From | Jerome Glisse <jglisse@redhat.com> |
|---|---|
| Date | 2016-11-21 13:40 +0100 |
| Subject | Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed |
| Message-ID | <sFW38-1bK-19@gated-at.bofh.it> |
| In reply to | #1526447 |
On Mon, Nov 21, 2016 at 01:56:02PM +0530, Anshuman Khandual wrote:
> On 11/18/2016 11:48 PM, Jérôme Glisse wrote:
> > When a ZONE_DEVICE page refcount reach 1 it means it is free and nobody
> > is holding a reference on it (only device to which the memory belong do).
> > Add a callback and call it when that happen so device driver can implement
> > their own free page management.
> >
> > Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
> > Cc: Dan Williams <dan.j.williams@intel.com>
> > Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
> > ---
> > include/linux/memremap.h | 4 ++++
> > kernel/memremap.c | 8 ++++++++
> > 2 files changed, 12 insertions(+)
> >
> > diff --git a/include/linux/memremap.h b/include/linux/memremap.h
> > index fe61dca..469c88d 100644
> > --- a/include/linux/memremap.h
> > +++ b/include/linux/memremap.h
> > @@ -37,17 +37,21 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
> >
> > /**
> > * struct dev_pagemap - metadata for ZONE_DEVICE mappings
> > + * @free_devpage: free page callback when page refcount reach 1
> > * @altmap: pre-allocated/reserved memory for vmemmap allocations
> > * @res: physical address range covered by @ref
> > * @ref: reference count that pins the devm_memremap_pages() mapping
> > * @dev: host device of the mapping for debug
> > + * @data: privata data pointer for free_devpage
> > * @flags: memory flags (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
> > */
> > struct dev_pagemap {
> > + void (*free_devpage)(struct page *page, void *data);
> > struct vmem_altmap *altmap;
> > const struct resource *res;
> > struct percpu_ref *ref;
> > struct device *dev;
> > + void *data;
> > int flags;
> > };
> >
> > diff --git a/kernel/memremap.c b/kernel/memremap.c
> > index 438a73aa2..3d28048 100644
> > --- a/kernel/memremap.c
> > +++ b/kernel/memremap.c
> > @@ -190,6 +190,12 @@ EXPORT_SYMBOL(get_zone_device_page);
> >
> > void put_zone_device_page(struct page *page)
> > {
> > + /*
> > + * If refcount is 1 then page is freed and refcount is stable as nobody
> > + * holds a reference on the page.
> > + */
> > + if (page->pgmap->free_devpage && page_count(page) == 1)
> > + page->pgmap->free_devpage(page, page->pgmap->data);
> > put_dev_pagemap(page->pgmap);
> > }
> > EXPORT_SYMBOL(put_zone_device_page);
> > @@ -326,6 +332,8 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
> > pgmap->ref = ref;
> > pgmap->res = &page_map->res;
> > pgmap->flags = flags | MEMORY_DEVICE;
> > + pgmap->free_devpage = NULL;
> > + pgmap->data = NULL;
>
> When is the driver expected to load up pgmap->free_devpage ? I thought
> this function is one of the right places. Though as all the pages in
> the same hotplug operation point to the same dev_pagemap structure this
> loading can be done at later point of time as well.
>
I wanted to avoid adding more argument to devm_memremap_pages() as it already
has a long list. Hence why i let the caller set those afterward.
Cheers,
Jérôme
[toc] | [prev] | [next] | [standalone]
| From | Anshuman Khandual <khandual@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-22 06:10 +0100 |
| Subject | Re: [HMM v13 04/18] mm/ZONE_DEVICE/free-page: callback when page is freed |
| Message-ID | <sGbvb-2Mm-1@gated-at.bofh.it> |
| In reply to | #1526636 |
On 11/21/2016 06:04 PM, Jerome Glisse wrote:
> On Mon, Nov 21, 2016 at 01:56:02PM +0530, Anshuman Khandual wrote:
>> On 11/18/2016 11:48 PM, Jérôme Glisse wrote:
>>> When a ZONE_DEVICE page refcount reach 1 it means it is free and nobody
>>> is holding a reference on it (only device to which the memory belong do).
>>> Add a callback and call it when that happen so device driver can implement
>>> their own free page management.
>>>
>>> Signed-off-by: Jérôme Glisse <jglisse@redhat.com>
>>> Cc: Dan Williams <dan.j.williams@intel.com>
>>> Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
>>> ---
>>> include/linux/memremap.h | 4 ++++
>>> kernel/memremap.c | 8 ++++++++
>>> 2 files changed, 12 insertions(+)
>>>
>>> diff --git a/include/linux/memremap.h b/include/linux/memremap.h
>>> index fe61dca..469c88d 100644
>>> --- a/include/linux/memremap.h
>>> +++ b/include/linux/memremap.h
>>> @@ -37,17 +37,21 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
>>>
>>> /**
>>> * struct dev_pagemap - metadata for ZONE_DEVICE mappings
>>> + * @free_devpage: free page callback when page refcount reach 1
>>> * @altmap: pre-allocated/reserved memory for vmemmap allocations
>>> * @res: physical address range covered by @ref
>>> * @ref: reference count that pins the devm_memremap_pages() mapping
>>> * @dev: host device of the mapping for debug
>>> + * @data: privata data pointer for free_devpage
>>> * @flags: memory flags (look for MEMORY_FLAGS_NONE in memory_hotplug.h)
>>> */
>>> struct dev_pagemap {
>>> + void (*free_devpage)(struct page *page, void *data);
>>> struct vmem_altmap *altmap;
>>> const struct resource *res;
>>> struct percpu_ref *ref;
>>> struct device *dev;
>>> + void *data;
>>> int flags;
>>> };
>>>
>>> diff --git a/kernel/memremap.c b/kernel/memremap.c
>>> index 438a73aa2..3d28048 100644
>>> --- a/kernel/memremap.c
>>> +++ b/kernel/memremap.c
>>> @@ -190,6 +190,12 @@ EXPORT_SYMBOL(get_zone_device_page);
>>>
>>> void put_zone_device_page(struct page *page)
>>> {
>>> + /*
>>> + * If refcount is 1 then page is freed and refcount is stable as nobody
>>> + * holds a reference on the page.
>>> + */
>>> + if (page->pgmap->free_devpage && page_count(page) == 1)
>>> + page->pgmap->free_devpage(page, page->pgmap->data);
>>> put_dev_pagemap(page->pgmap);
>>> }
>>> EXPORT_SYMBOL(put_zone_device_page);
>>> @@ -326,6 +332,8 @@ void *devm_memremap_pages(struct device *dev, struct resource *res,
>>> pgmap->ref = ref;
>>> pgmap->res = &page_map->res;
>>> pgmap->flags = flags | MEMORY_DEVICE;
>>> + pgmap->free_devpage = NULL;
>>> + pgmap->data = NULL;
>>
>> When is the driver expected to load up pgmap->free_devpage ? I thought
>> this function is one of the right places. Though as all the pages in
>> the same hotplug operation point to the same dev_pagemap structure this
>> loading can be done at later point of time as well.
>>
>
> I wanted to avoid adding more argument to devm_memremap_pages() as it already
> has a long list. Hence why i let the caller set those afterward.
IMHO we should still pass it through this function argument so that
by the time the function returns we will have device memory properly
setup through ZONE_DEVICE with all bells and whistles enabled.
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web