Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1627845 > unrolled thread
| Started by | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| First post | 2017-04-20 23:50 +0200 |
| Last post | 2017-04-24 20:50 +0200 |
| Articles | 11 — 4 participants |
Back to article view | Back to linux.kernel
Re: [tip:x86/mm] x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation Dan Williams <dan.j.williams@intel.com> - 2017-04-20 23:50 +0200
Re: [tip:x86/mm] x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-04-21 16:20 +0200
Re: [tip:x86/mm] x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation Dan Williams <dan.j.williams@intel.com> - 2017-04-21 21:40 +0200
[PATCH] Revert "x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation" Ingo Molnar <mingo@kernel.org> - 2017-04-23 12:00 +0200
get_zone_device_page() in get_page() and page_cache_get_speculative() "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-04-24 01:40 +0200
Re: get_zone_device_page() in get_page() and page_cache_get_speculative() Dan Williams <dan.j.williams@intel.com> - 2017-04-24 19:30 +0200
Re: get_zone_device_page() in get_page() and page_cache_get_speculative() "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-04-24 19:40 +0200
Re: get_zone_device_page() in get_page() and page_cache_get_speculative() Dan Williams <dan.j.williams@intel.com> - 2017-04-24 19:50 +0200
Re: get_zone_device_page() in get_page() and page_cache_get_speculative() "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-04-24 20:10 +0200
Re: get_zone_device_page() in get_page() and page_cache_get_speculative() "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2017-04-24 20:30 +0200
Re: get_zone_device_page() in get_page() and page_cache_get_speculative() Dan Williams <dan.j.williams@intel.com> - 2017-04-24 20:50 +0200
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-20 23:50 +0200 |
| Subject | Re: [tip:x86/mm] x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation |
| Message-ID | <tys7D-5IQ-11@gated-at.bofh.it> |
On Sat, Mar 18, 2017 at 2:52 AM, tip-bot for Kirill A. Shutemov
<tipbot@zytor.com> wrote:
> Commit-ID: 2947ba054a4dabbd82848728d765346886050029
> Gitweb: http://git.kernel.org/tip/2947ba054a4dabbd82848728d765346886050029
> Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
> AuthorDate: Fri, 17 Mar 2017 00:39:06 +0300
> Committer: Ingo Molnar <mingo@kernel.org>
> CommitDate: Sat, 18 Mar 2017 09:48:03 +0100
>
> x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation
>
> This patch provides all required callbacks required by the generic
> get_user_pages_fast() code and switches x86 over - and removes
> the platform specific implementation.
>
> Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
> Cc: Andrew Morton <akpm@linux-foundation.org>
> Cc: Aneesh Kumar K . V <aneesh.kumar@linux.vnet.ibm.com>
> Cc: Borislav Petkov <bp@alien8.de>
> Cc: Catalin Marinas <catalin.marinas@arm.com>
> Cc: Dann Frazier <dann.frazier@canonical.com>
> Cc: Dave Hansen <dave.hansen@intel.com>
> Cc: H. Peter Anvin <hpa@zytor.com>
> Cc: Linus Torvalds <torvalds@linux-foundation.org>
> Cc: Peter Zijlstra <peterz@infradead.org>
> Cc: Rik van Riel <riel@redhat.com>
> Cc: Steve Capper <steve.capper@linaro.org>
> Cc: Thomas Gleixner <tglx@linutronix.de>
> Cc: linux-arch@vger.kernel.org
> Cc: linux-mm@kvack.org
> Link: http://lkml.kernel.org/r/20170316213906.89528-1-kirill.shutemov@linux.intel.com
> [ Minor readability edits. ]
> Signed-off-by: Ingo Molnar <mingo@kernel.org>
I'm still trying to spot the bug, but bisect points to this patch as
the point at which my unit tests start failing with the following
signature:
[ 35.423841] WARNING: CPU: 8 PID: 245 at lib/percpu-refcount.c:155
percpu_ref_switch_to_atomic_rcu+0x1f5/0x200
[ 35.425328] percpu ref (dax_pmem_percpu_release [dax_pmem]) <= 0
(0) after switching to atomic
[ 35.425329] Modules linked in: ip6t_rpfilter ip6t_REJECT
nf_reject_ipv6 xt_conntrack ebtable_nat ebtable_broute bridge stp llc
ip6table_nat nf_conntrack_ipv6 nf_defrag_ipv6 nf_nat_ipv6 ip
6table_mangle ip6table_raw ip6table_security iptable_nat
nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 nf_nat nf_conntrack
iptable_mangle iptable_raw iptable_security ebtable_filter ebtables
ip6table_filter ip6_tables crct10dif_pclmul crc32_pclmul crc32c_intel
ghash_clmulni_intel nd_pmem(O) dax_pmem(O) nd_btt(O) dax(O) serio_raw
nfit(O) nd_e820(O) libnvdimm(O) tpm_tis tpm_tis_co
re tpm nfit_test_iomap(O) nfsd nfs_acl
[ 35.433683] CPU: 8 PID: 245 Comm: rcuos/29 Tainted: G O
4.11.0-rc2+ #55
[ 35.435538] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996),
BIOS 1.9.3-1.fc25 04/01/2014
[ 35.437500] Call Trace:
[ 35.438270] dump_stack+0x86/0xc3
[ 35.439156] __warn+0xcb/0xf0
[ 35.439995] warn_slowpath_fmt+0x5f/0x80
[ 35.440962] ? rcu_nocb_kthread+0x27a/0x500
[ 35.441957] ? dax_pmem_percpu_exit+0x50/0x50 [dax_pmem]
[ 35.443107] percpu_ref_switch_to_atomic_rcu+0x1f5/0x200
[ 35.444251] ? percpu_ref_exit+0x60/0x60
[ 35.445206] rcu_nocb_kthread+0x327/0x500
[ 35.446186] ? rcu_nocb_kthread+0x27a/0x500
[ 35.447188] kthread+0x10c/0x140
[ 35.448058] ? rcu_eqs_enter+0x50/0x50
[ 35.448990] ? kthread_create_on_node+0x60/0x60
[ 35.450038] ret_from_fork+0x31/0x40
[ 35.450976] ---[ end trace eaa40898a09519b5 ]---
This is similar to the backtrace when we were not properly handling
pud faults and was fixed with this commit: 220ced1676c4 "mm: fix
get_user_pages() vs device-dax pud mappings"
I've found some missing _devmap checks in the generic
get_user_pages_fast() path, but this does not fix the regression:
diff --git a/mm/gup.c b/mm/gup.c
index 2559a3987de7..89156cd59cbc 100644
--- a/mm/gup.c
+++ b/mm/gup.c
@@ -1475,7 +1475,8 @@ static int gup_pmd_range(pud_t pud, unsigned
long addr, unsigned long end,
if (pmd_none(pmd))
return 0;
- if (unlikely(pmd_trans_huge(pmd) || pmd_huge(pmd))) {
+ if (unlikely(pmd_trans_huge(pmd) || pmd_huge(pmd)
+ || pmd_devmap(pmd))) {
/*
* NUMA hinting faults need to be handled in the GUP
* slowpath for accounting purposes and so that they
@@ -1516,7 +1517,7 @@ static int gup_pud_range(p4d_t p4d, unsigned
long addr, unsigned long end,
next = pud_addr_end(addr, end);
if (pud_none(pud))
return 0;
- if (unlikely(pud_huge(pud))) {
+ if (unlikely(pud_huge(pud) || pud_devmap(pud))) {
if (!gup_huge_pud(pud, pudp, addr, next, write,
pages, nr))
return 0;
...more hunting tomorrow.
[toc] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2017-04-21 16:20 +0200 |
| Message-ID | <tyHzI-6Pl-25@gated-at.bofh.it> |
| In reply to | #1627845 |
On Thu, Apr 20, 2017 at 02:46:51PM -0700, Dan Williams wrote: > On Sat, Mar 18, 2017 at 2:52 AM, tip-bot for Kirill A. Shutemov > <tipbot@zytor.com> wrote: > > Commit-ID: 2947ba054a4dabbd82848728d765346886050029 > > Gitweb: http://git.kernel.org/tip/2947ba054a4dabbd82848728d765346886050029 > > Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> > > AuthorDate: Fri, 17 Mar 2017 00:39:06 +0300 > > Committer: Ingo Molnar <mingo@kernel.org> > > CommitDate: Sat, 18 Mar 2017 09:48:03 +0100 > > > > x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation > > > > This patch provides all required callbacks required by the generic > > get_user_pages_fast() code and switches x86 over - and removes > > the platform specific implementation. > > > > Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> > > Cc: Andrew Morton <akpm@linux-foundation.org> > > Cc: Aneesh Kumar K . V <aneesh.kumar@linux.vnet.ibm.com> > > Cc: Borislav Petkov <bp@alien8.de> > > Cc: Catalin Marinas <catalin.marinas@arm.com> > > Cc: Dann Frazier <dann.frazier@canonical.com> > > Cc: Dave Hansen <dave.hansen@intel.com> > > Cc: H. Peter Anvin <hpa@zytor.com> > > Cc: Linus Torvalds <torvalds@linux-foundation.org> > > Cc: Peter Zijlstra <peterz@infradead.org> > > Cc: Rik van Riel <riel@redhat.com> > > Cc: Steve Capper <steve.capper@linaro.org> > > Cc: Thomas Gleixner <tglx@linutronix.de> > > Cc: linux-arch@vger.kernel.org > > Cc: linux-mm@kvack.org > > Link: http://lkml.kernel.org/r/20170316213906.89528-1-kirill.shutemov@linux.intel.com > > [ Minor readability edits. ] > > Signed-off-by: Ingo Molnar <mingo@kernel.org> > > I'm still trying to spot the bug, but bisect points to this patch as > the point at which my unit tests start failing with the following > signature: I can't find the issue either. Is it something reproducible without hardware? In KVM? If yes, could you share the test-case? > [ 35.423841] WARNING: CPU: 8 PID: 245 at lib/percpu-refcount.c:155 > percpu_ref_switch_to_atomic_rcu+0x1f5/0x200 > [ 35.425328] percpu ref (dax_pmem_percpu_release [dax_pmem]) <= 0 > (0) after switching to atomic > [ 35.425329] Modules linked in: ip6t_rpfilter ip6t_REJECT > nf_reject_ipv6 xt_conntrack ebtable_nat ebtable_broute bridge stp llc > ip6table_nat nf_conntrack_ipv6 nf_defrag_ipv6 nf_nat_ipv6 ip > 6table_mangle ip6table_raw ip6table_security iptable_nat > nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 nf_nat nf_conntrack > iptable_mangle iptable_raw iptable_security ebtable_filter ebtables > ip6table_filter ip6_tables crct10dif_pclmul crc32_pclmul crc32c_intel > ghash_clmulni_intel nd_pmem(O) dax_pmem(O) nd_btt(O) dax(O) serio_raw > nfit(O) nd_e820(O) libnvdimm(O) tpm_tis tpm_tis_co > re tpm nfit_test_iomap(O) nfsd nfs_acl > [ 35.433683] CPU: 8 PID: 245 Comm: rcuos/29 Tainted: G O > 4.11.0-rc2+ #55 > [ 35.435538] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), > BIOS 1.9.3-1.fc25 04/01/2014 > [ 35.437500] Call Trace: > [ 35.438270] dump_stack+0x86/0xc3 > [ 35.439156] __warn+0xcb/0xf0 > [ 35.439995] warn_slowpath_fmt+0x5f/0x80 > [ 35.440962] ? rcu_nocb_kthread+0x27a/0x500 > [ 35.441957] ? dax_pmem_percpu_exit+0x50/0x50 [dax_pmem] > [ 35.443107] percpu_ref_switch_to_atomic_rcu+0x1f5/0x200 > [ 35.444251] ? percpu_ref_exit+0x60/0x60 > [ 35.445206] rcu_nocb_kthread+0x327/0x500 > [ 35.446186] ? rcu_nocb_kthread+0x27a/0x500 > [ 35.447188] kthread+0x10c/0x140 > [ 35.448058] ? rcu_eqs_enter+0x50/0x50 > [ 35.448990] ? kthread_create_on_node+0x60/0x60 > [ 35.450038] ret_from_fork+0x31/0x40 > [ 35.450976] ---[ end trace eaa40898a09519b5 ]--- > > This is similar to the backtrace when we were not properly handling > pud faults and was fixed with this commit: 220ced1676c4 "mm: fix > get_user_pages() vs device-dax pud mappings" > > I've found some missing _devmap checks in the generic > get_user_pages_fast() path, but this does not fix the regression: I don't see these in x86 GUP. Was the bug there too? -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-21 21:40 +0200 |
| Message-ID | <tyMzo-1ke-13@gated-at.bofh.it> |
| In reply to | #1628308 |
On Fri, Apr 21, 2017 at 7:16 AM, Kirill A. Shutemov
<kirill@shutemov.name> wrote:
> On Thu, Apr 20, 2017 at 02:46:51PM -0700, Dan Williams wrote:
>> On Sat, Mar 18, 2017 at 2:52 AM, tip-bot for Kirill A. Shutemov
>> <tipbot@zytor.com> wrote:
>> > Commit-ID: 2947ba054a4dabbd82848728d765346886050029
>> > Gitweb: http://git.kernel.org/tip/2947ba054a4dabbd82848728d765346886050029
>> > Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
>> > AuthorDate: Fri, 17 Mar 2017 00:39:06 +0300
>> > Committer: Ingo Molnar <mingo@kernel.org>
>> > CommitDate: Sat, 18 Mar 2017 09:48:03 +0100
>> >
>> > x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation
>> >
>> > This patch provides all required callbacks required by the generic
>> > get_user_pages_fast() code and switches x86 over - and removes
>> > the platform specific implementation.
>> >
>> > Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
>> > Cc: Andrew Morton <akpm@linux-foundation.org>
>> > Cc: Aneesh Kumar K . V <aneesh.kumar@linux.vnet.ibm.com>
>> > Cc: Borislav Petkov <bp@alien8.de>
>> > Cc: Catalin Marinas <catalin.marinas@arm.com>
>> > Cc: Dann Frazier <dann.frazier@canonical.com>
>> > Cc: Dave Hansen <dave.hansen@intel.com>
>> > Cc: H. Peter Anvin <hpa@zytor.com>
>> > Cc: Linus Torvalds <torvalds@linux-foundation.org>
>> > Cc: Peter Zijlstra <peterz@infradead.org>
>> > Cc: Rik van Riel <riel@redhat.com>
>> > Cc: Steve Capper <steve.capper@linaro.org>
>> > Cc: Thomas Gleixner <tglx@linutronix.de>
>> > Cc: linux-arch@vger.kernel.org
>> > Cc: linux-mm@kvack.org
>> > Link: http://lkml.kernel.org/r/20170316213906.89528-1-kirill.shutemov@linux.intel.com
>> > [ Minor readability edits. ]
>> > Signed-off-by: Ingo Molnar <mingo@kernel.org>
>>
>> I'm still trying to spot the bug, but bisect points to this patch as
>> the point at which my unit tests start failing with the following
>> signature:
>
> I can't find the issue either.
>
> Is it something reproducible without hardware? In KVM?
You can do it in KVM, just boot with the memmap=ss!nn parameter to
simulate pmem. In this case I'm booting with memmap=4G!8G, you should
also specify "nokaslr".
> If yes, could you share the test-case?
Yes, run:
./autogen.sh
./configure CFLAGS='-g -O0' --prefix=/usr --sysconfdir=/etc
--libdir=/usr/lib64
make TESTS=device-dax check
...from a checkout of the ndctl project:
https://github.com/pmem/ndctl
Let me know if you run into any problems getting the test to build or run.
>
>> [ 35.423841] WARNING: CPU: 8 PID: 245 at lib/percpu-refcount.c:155
>> percpu_ref_switch_to_atomic_rcu+0x1f5/0x200
>> [ 35.425328] percpu ref (dax_pmem_percpu_release [dax_pmem]) <= 0
>> (0) after switching to atomic
>> [ 35.425329] Modules linked in: ip6t_rpfilter ip6t_REJECT
>> nf_reject_ipv6 xt_conntrack ebtable_nat ebtable_broute bridge stp llc
>> ip6table_nat nf_conntrack_ipv6 nf_defrag_ipv6 nf_nat_ipv6 ip
>> 6table_mangle ip6table_raw ip6table_security iptable_nat
>> nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 nf_nat nf_conntrack
>> iptable_mangle iptable_raw iptable_security ebtable_filter ebtables
>> ip6table_filter ip6_tables crct10dif_pclmul crc32_pclmul crc32c_intel
>> ghash_clmulni_intel nd_pmem(O) dax_pmem(O) nd_btt(O) dax(O) serio_raw
>> nfit(O) nd_e820(O) libnvdimm(O) tpm_tis tpm_tis_co
>> re tpm nfit_test_iomap(O) nfsd nfs_acl
>> [ 35.433683] CPU: 8 PID: 245 Comm: rcuos/29 Tainted: G O
>> 4.11.0-rc2+ #55
>> [ 35.435538] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996),
>> BIOS 1.9.3-1.fc25 04/01/2014
>> [ 35.437500] Call Trace:
>> [ 35.438270] dump_stack+0x86/0xc3
>> [ 35.439156] __warn+0xcb/0xf0
>> [ 35.439995] warn_slowpath_fmt+0x5f/0x80
>> [ 35.440962] ? rcu_nocb_kthread+0x27a/0x500
>> [ 35.441957] ? dax_pmem_percpu_exit+0x50/0x50 [dax_pmem]
>> [ 35.443107] percpu_ref_switch_to_atomic_rcu+0x1f5/0x200
>> [ 35.444251] ? percpu_ref_exit+0x60/0x60
>> [ 35.445206] rcu_nocb_kthread+0x327/0x500
>> [ 35.446186] ? rcu_nocb_kthread+0x27a/0x500
>> [ 35.447188] kthread+0x10c/0x140
>> [ 35.448058] ? rcu_eqs_enter+0x50/0x50
>> [ 35.448990] ? kthread_create_on_node+0x60/0x60
>> [ 35.450038] ret_from_fork+0x31/0x40
>> [ 35.450976] ---[ end trace eaa40898a09519b5 ]---
>>
>> This is similar to the backtrace when we were not properly handling
>> pud faults and was fixed with this commit: 220ced1676c4 "mm: fix
>> get_user_pages() vs device-dax pud mappings"
>>
>> I've found some missing _devmap checks in the generic
>> get_user_pages_fast() path, but this does not fix the regression:
>
> I don't see these in x86 GUP. Was the bug there too?
No it wasn't, the test runs fine with v4.11-rc7, so perhaps I'm
looking in the wrong place...
[toc] | [prev] | [next] | [standalone]
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2017-04-23 12:00 +0200 |
| Subject | [PATCH] Revert "x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation" |
| Message-ID | <tzmtc-6QE-17@gated-at.bofh.it> |
| In reply to | #1628549 |
* Dan Williams <dan.j.williams@intel.com> wrote:
> > I can't find the issue either.
> >
> > Is it something reproducible without hardware? In KVM?
>
> You can do it in KVM, just boot with the memmap=ss!nn parameter to
> simulate pmem. In this case I'm booting with memmap=4G!8G, you should
> also specify "nokaslr".
>
> > If yes, could you share the test-case?
>
> Yes, run:
>
> ./autogen.sh
> ./configure CFLAGS='-g -O0' --prefix=/usr --sysconfdir=/etc
> --libdir=/usr/lib64
> make TESTS=device-dax check
>
> ...from a checkout of the ndctl project:
>
> https://github.com/pmem/ndctl
>
> Let me know if you run into any problems getting the test to build or run.
>
> >
> >> [ 35.423841] WARNING: CPU: 8 PID: 245 at lib/percpu-refcount.c:155
> >> percpu_ref_switch_to_atomic_rcu+0x1f5/0x200
> >> [ 35.425328] percpu ref (dax_pmem_percpu_release [dax_pmem]) <= 0
> >> (0) after switching to atomic
Since the bug appears to be pretty severe (GUP race causing possible memory
corruption that could affect a lot of code), and the merge window is awfully
close, plus the reproducer appears to be pretty quick, I've queued up the
revert below for the time being, to not block the rest of the pending
tip:x86/mm changes.
I'd have loved to see this conversion in v4.12, but not at any cost.
Thanks,
Ingo
==================>
From 6dd29b3df975582ef429b5b93c899e6575785940 Mon Sep 17 00:00:00 2001
From: Ingo Molnar <mingo@kernel.org>
Date: Sun, 23 Apr 2017 11:37:17 +0200
Subject: [PATCH] Revert "x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation"
This reverts commit 2947ba054a4dabbd82848728d765346886050029.
Dan Williams reported dax-pmem kernel warnings with the following signature:
WARNING: CPU: 8 PID: 245 at lib/percpu-refcount.c:155 percpu_ref_switch_to_atomic_rcu+0x1f5/0x200
percpu ref (dax_pmem_percpu_release [dax_pmem]) <= 0 (0) after switching to atomic
... and bisected it to this commit, which suggests possible memory corruption
caused by the x86 fast-GUP conversion.
He also pointed out:
"
This is similar to the backtrace when we were not properly handling
pud faults and was fixed with this commit: 220ced1676c4 "mm: fix
get_user_pages() vs device-dax pud mappings"
I've found some missing _devmap checks in the generic
get_user_pages_fast() path, but this does not fix the regression
[...]
"
So given that there are known bugs, and a pretty robust looking bisection
points to this commit suggesting that are unknown bugs in the conversion
as well, revert it for the time being - we'll re-try in v4.13.
Reported-by: Dan Williams <dan.j.williams@intel.com>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Borislav Petkov <bp@alien8.de>
Cc: Catalin Marinas <catalin.marinas@arm.com>
Cc: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Cc: Michal Hocko <mhocko@suse.cz>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Rik van Riel <riel@redhat.com>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: aneesh.kumar@linux.vnet.ibm.com
Cc: dann.frazier@canonical.com
Cc: dave.hansen@intel.com
Cc: steve.capper@linaro.org
Cc: linux-kernel@vger.kernel.org
Signed-off-by: Ingo Molnar <mingo@kernel.org>
---
arch/arm/Kconfig | 2 +-
arch/arm64/Kconfig | 2 +-
arch/powerpc/Kconfig | 2 +-
arch/x86/Kconfig | 3 -
arch/x86/include/asm/mmu_context.h | 12 +
arch/x86/include/asm/pgtable-3level.h | 47 ----
arch/x86/include/asm/pgtable.h | 53 ----
arch/x86/include/asm/pgtable_64.h | 16 +-
arch/x86/mm/Makefile | 2 +-
arch/x86/mm/gup.c | 496 ++++++++++++++++++++++++++++++++++
mm/Kconfig | 2 +-
mm/gup.c | 10 +-
12 files changed, 519 insertions(+), 128 deletions(-)
diff --git a/arch/arm/Kconfig b/arch/arm/Kconfig
index 454fadd077ad..0d4e71b42c77 100644
--- a/arch/arm/Kconfig
+++ b/arch/arm/Kconfig
@@ -1666,7 +1666,7 @@ config ARCH_SELECT_MEMORY_MODEL
config HAVE_ARCH_PFN_VALID
def_bool ARCH_HAS_HOLES_MEMORYMODEL || !SPARSEMEM
-config HAVE_GENERIC_GUP
+config HAVE_GENERIC_RCU_GUP
def_bool y
depends on ARM_LPAE
diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig
index af62bf79721a..3741859765cf 100644
--- a/arch/arm64/Kconfig
+++ b/arch/arm64/Kconfig
@@ -205,7 +205,7 @@ config GENERIC_CALIBRATE_DELAY
config ZONE_DMA
def_bool y
-config HAVE_GENERIC_GUP
+config HAVE_GENERIC_RCU_GUP
def_bool y
config ARCH_DMA_ADDR_T_64BIT
diff --git a/arch/powerpc/Kconfig b/arch/powerpc/Kconfig
index 3a716b2dcde9..97a8bc8a095c 100644
--- a/arch/powerpc/Kconfig
+++ b/arch/powerpc/Kconfig
@@ -135,7 +135,7 @@ config PPC
select HAVE_FUNCTION_GRAPH_TRACER
select HAVE_FUNCTION_TRACER
select HAVE_GCC_PLUGINS
- select HAVE_GENERIC_GUP
+ select HAVE_GENERIC_RCU_GUP
select HAVE_HW_BREAKPOINT if PERF_EVENTS && (PPC_BOOK3S || PPC_8xx)
select HAVE_IDE
select HAVE_IOREMAP_PROT
diff --git a/arch/x86/Kconfig b/arch/x86/Kconfig
index a641b900fc1f..2bde14451e54 100644
--- a/arch/x86/Kconfig
+++ b/arch/x86/Kconfig
@@ -2789,9 +2789,6 @@ config X86_DMA_REMAP
bool
depends on STA2X11
-config HAVE_GENERIC_GUP
- def_bool y
-
source "net/Kconfig"
source "drivers/Kconfig"
diff --git a/arch/x86/include/asm/mmu_context.h b/arch/x86/include/asm/mmu_context.h
index 6e933d2d88d9..68b329d77b3a 100644
--- a/arch/x86/include/asm/mmu_context.h
+++ b/arch/x86/include/asm/mmu_context.h
@@ -220,6 +220,18 @@ static inline int vma_pkey(struct vm_area_struct *vma)
}
#endif
+static inline bool __pkru_allows_pkey(u16 pkey, bool write)
+{
+ u32 pkru = read_pkru();
+
+ if (!__pkru_allows_read(pkru, pkey))
+ return false;
+ if (write && !__pkru_allows_write(pkru, pkey))
+ return false;
+
+ return true;
+}
+
/*
* We only want to enforce protection keys on the current process
* because we effectively have no access to PKRU for other
diff --git a/arch/x86/include/asm/pgtable-3level.h b/arch/x86/include/asm/pgtable-3level.h
index c8821bab938f..50d35e3185f5 100644
--- a/arch/x86/include/asm/pgtable-3level.h
+++ b/arch/x86/include/asm/pgtable-3level.h
@@ -212,51 +212,4 @@ static inline pud_t native_pudp_get_and_clear(pud_t *pudp)
#define __pte_to_swp_entry(pte) ((swp_entry_t){ (pte).pte_high })
#define __swp_entry_to_pte(x) ((pte_t){ { .pte_high = (x).val } })
-#define gup_get_pte gup_get_pte
-/*
- * WARNING: only to be used in the get_user_pages_fast() implementation.
- *
- * With get_user_pages_fast(), we walk down the pagetables without taking
- * any locks. For this we would like to load the pointers atomically,
- * but that is not possible (without expensive cmpxchg8b) on PAE. What
- * we do have is the guarantee that a PTE will only either go from not
- * present to present, or present to not present or both -- it will not
- * switch to a completely different present page without a TLB flush in
- * between; something that we are blocking by holding interrupts off.
- *
- * Setting ptes from not present to present goes:
- *
- * ptep->pte_high = h;
- * smp_wmb();
- * ptep->pte_low = l;
- *
- * And present to not present goes:
- *
- * ptep->pte_low = 0;
- * smp_wmb();
- * ptep->pte_high = 0;
- *
- * We must ensure here that the load of pte_low sees 'l' iff pte_high
- * sees 'h'. We load pte_high *after* loading pte_low, which ensures we
- * don't see an older value of pte_high. *Then* we recheck pte_low,
- * which ensures that we haven't picked up a changed pte high. We might
- * have gotten rubbish values from pte_low and pte_high, but we are
- * guaranteed that pte_low will not have the present bit set *unless*
- * it is 'l'. Because get_user_pages_fast() only operates on present ptes
- * we're safe.
- */
-static inline pte_t gup_get_pte(pte_t *ptep)
-{
- pte_t pte;
-
- do {
- pte.pte_low = ptep->pte_low;
- smp_rmb();
- pte.pte_high = ptep->pte_high;
- smp_rmb();
- } while (unlikely(pte.pte_low != ptep->pte_low));
-
- return pte;
-}
-
#endif /* _ASM_X86_PGTABLE_3LEVEL_H */
diff --git a/arch/x86/include/asm/pgtable.h b/arch/x86/include/asm/pgtable.h
index 942482ac36a8..f5af95a0c6b8 100644
--- a/arch/x86/include/asm/pgtable.h
+++ b/arch/x86/include/asm/pgtable.h
@@ -244,11 +244,6 @@ static inline int pud_devmap(pud_t pud)
return 0;
}
#endif
-
-static inline int pgd_devmap(pgd_t pgd)
-{
- return 0;
-}
#endif
#endif /* CONFIG_TRANSPARENT_HUGEPAGE */
@@ -1190,54 +1185,6 @@ static inline u16 pte_flags_pkey(unsigned long pte_flags)
#endif
}
-static inline bool __pkru_allows_pkey(u16 pkey, bool write)
-{
- u32 pkru = read_pkru();
-
- if (!__pkru_allows_read(pkru, pkey))
- return false;
- if (write && !__pkru_allows_write(pkru, pkey))
- return false;
-
- return true;
-}
-
-/*
- * 'pteval' can come from a PTE, PMD or PUD. We only check
- * _PAGE_PRESENT, _PAGE_USER, and _PAGE_RW in here which are the
- * same value on all 3 types.
- */
-static inline bool __pte_access_permitted(unsigned long pteval, bool write)
-{
- unsigned long need_pte_bits = _PAGE_PRESENT|_PAGE_USER;
-
- if (write)
- need_pte_bits |= _PAGE_RW;
-
- if ((pteval & need_pte_bits) != need_pte_bits)
- return 0;
-
- return __pkru_allows_pkey(pte_flags_pkey(pteval), write);
-}
-
-#define pte_access_permitted pte_access_permitted
-static inline bool pte_access_permitted(pte_t pte, bool write)
-{
- return __pte_access_permitted(pte_val(pte), write);
-}
-
-#define pmd_access_permitted pmd_access_permitted
-static inline bool pmd_access_permitted(pmd_t pmd, bool write)
-{
- return __pte_access_permitted(pmd_val(pmd), write);
-}
-
-#define pud_access_permitted pud_access_permitted
-static inline bool pud_access_permitted(pud_t pud, bool write)
-{
- return __pte_access_permitted(pud_val(pud), write);
-}
-
#include <asm-generic/pgtable.h>
#endif /* __ASSEMBLY__ */
diff --git a/arch/x86/include/asm/pgtable_64.h b/arch/x86/include/asm/pgtable_64.h
index 12ea31274eb6..9991224f6238 100644
--- a/arch/x86/include/asm/pgtable_64.h
+++ b/arch/x86/include/asm/pgtable_64.h
@@ -227,20 +227,6 @@ extern void cleanup_highmap(void);
extern void init_extra_mapping_uc(unsigned long phys, unsigned long size);
extern void init_extra_mapping_wb(unsigned long phys, unsigned long size);
-#define gup_fast_permitted gup_fast_permitted
-static inline bool gup_fast_permitted(unsigned long start, int nr_pages,
- int write)
-{
- unsigned long len, end;
-
- len = (unsigned long)nr_pages << PAGE_SHIFT;
- end = start + len;
- if (end < start)
- return false;
- if (end >> __VIRTUAL_MASK_SHIFT)
- return false;
- return true;
-}
-
#endif /* !__ASSEMBLY__ */
+
#endif /* _ASM_X86_PGTABLE_64_H */
diff --git a/arch/x86/mm/Makefile b/arch/x86/mm/Makefile
index 0fbdcb64f9f8..96d2b847e09e 100644
--- a/arch/x86/mm/Makefile
+++ b/arch/x86/mm/Makefile
@@ -2,7 +2,7 @@
KCOV_INSTRUMENT_tlb.o := n
obj-y := init.o init_$(BITS).o fault.o ioremap.o extable.o pageattr.o mmap.o \
- pat.o pgtable.o physaddr.o setup_nx.o tlb.o
+ pat.o pgtable.o physaddr.o gup.o setup_nx.o tlb.o
# Make sure __phys_addr has no stackprotector
nostackp := $(call cc-option, -fno-stack-protector)
diff --git a/arch/x86/mm/gup.c b/arch/x86/mm/gup.c
new file mode 100644
index 000000000000..456dfdfd2249
--- /dev/null
+++ b/arch/x86/mm/gup.c
@@ -0,0 +1,496 @@
+/*
+ * Lockless get_user_pages_fast for x86
+ *
+ * Copyright (C) 2008 Nick Piggin
+ * Copyright (C) 2008 Novell Inc.
+ */
+#include <linux/sched.h>
+#include <linux/mm.h>
+#include <linux/vmstat.h>
+#include <linux/highmem.h>
+#include <linux/swap.h>
+#include <linux/memremap.h>
+
+#include <asm/mmu_context.h>
+#include <asm/pgtable.h>
+
+static inline pte_t gup_get_pte(pte_t *ptep)
+{
+#ifndef CONFIG_X86_PAE
+ return READ_ONCE(*ptep);
+#else
+ /*
+ * With get_user_pages_fast, we walk down the pagetables without taking
+ * any locks. For this we would like to load the pointers atomically,
+ * but that is not possible (without expensive cmpxchg8b) on PAE. What
+ * we do have is the guarantee that a pte will only either go from not
+ * present to present, or present to not present or both -- it will not
+ * switch to a completely different present page without a TLB flush in
+ * between; something that we are blocking by holding interrupts off.
+ *
+ * Setting ptes from not present to present goes:
+ * ptep->pte_high = h;
+ * smp_wmb();
+ * ptep->pte_low = l;
+ *
+ * And present to not present goes:
+ * ptep->pte_low = 0;
+ * smp_wmb();
+ * ptep->pte_high = 0;
+ *
+ * We must ensure here that the load of pte_low sees l iff pte_high
+ * sees h. We load pte_high *after* loading pte_low, which ensures we
+ * don't see an older value of pte_high. *Then* we recheck pte_low,
+ * which ensures that we haven't picked up a changed pte high. We might
+ * have got rubbish values from pte_low and pte_high, but we are
+ * guaranteed that pte_low will not have the present bit set *unless*
+ * it is 'l'. And get_user_pages_fast only operates on present ptes, so
+ * we're safe.
+ *
+ * gup_get_pte should not be used or copied outside gup.c without being
+ * very careful -- it does not atomically load the pte or anything that
+ * is likely to be useful for you.
+ */
+ pte_t pte;
+
+retry:
+ pte.pte_low = ptep->pte_low;
+ smp_rmb();
+ pte.pte_high = ptep->pte_high;
+ smp_rmb();
+ if (unlikely(pte.pte_low != ptep->pte_low))
+ goto retry;
+
+ return pte;
+#endif
+}
+
+static void undo_dev_pagemap(int *nr, int nr_start, struct page **pages)
+{
+ while ((*nr) - nr_start) {
+ struct page *page = pages[--(*nr)];
+
+ ClearPageReferenced(page);
+ put_page(page);
+ }
+}
+
+/*
+ * 'pteval' can come from a pte, pmd, pud or p4d. We only check
+ * _PAGE_PRESENT, _PAGE_USER, and _PAGE_RW in here which are the
+ * same value on all 4 types.
+ */
+static inline int pte_allows_gup(unsigned long pteval, int write)
+{
+ unsigned long need_pte_bits = _PAGE_PRESENT|_PAGE_USER;
+
+ if (write)
+ need_pte_bits |= _PAGE_RW;
+
+ if ((pteval & need_pte_bits) != need_pte_bits)
+ return 0;
+
+ /* Check memory protection keys permissions. */
+ if (!__pkru_allows_pkey(pte_flags_pkey(pteval), write))
+ return 0;
+
+ return 1;
+}
+
+/*
+ * The performance critical leaf functions are made noinline otherwise gcc
+ * inlines everything into a single function which results in too much
+ * register pressure.
+ */
+static noinline int gup_pte_range(pmd_t pmd, unsigned long addr,
+ unsigned long end, int write, struct page **pages, int *nr)
+{
+ struct dev_pagemap *pgmap = NULL;
+ int nr_start = *nr, ret = 0;
+ pte_t *ptep, *ptem;
+
+ /*
+ * Keep the original mapped PTE value (ptem) around since we
+ * might increment ptep off the end of the page when finishing
+ * our loop iteration.
+ */
+ ptem = ptep = pte_offset_map(&pmd, addr);
+ do {
+ pte_t pte = gup_get_pte(ptep);
+ struct page *page;
+
+ /* Similar to the PMD case, NUMA hinting must take slow path */
+ if (pte_protnone(pte))
+ break;
+
+ if (!pte_allows_gup(pte_val(pte), write))
+ break;
+
+ if (pte_devmap(pte)) {
+ pgmap = get_dev_pagemap(pte_pfn(pte), pgmap);
+ if (unlikely(!pgmap)) {
+ undo_dev_pagemap(nr, nr_start, pages);
+ break;
+ }
+ } else if (pte_special(pte))
+ break;
+
+ VM_BUG_ON(!pfn_valid(pte_pfn(pte)));
+ page = pte_page(pte);
+ get_page(page);
+ put_dev_pagemap(pgmap);
+ SetPageReferenced(page);
+ pages[*nr] = page;
+ (*nr)++;
+
+ } while (ptep++, addr += PAGE_SIZE, addr != end);
+ if (addr == end)
+ ret = 1;
+ pte_unmap(ptem);
+
+ return ret;
+}
+
+static inline void get_head_page_multiple(struct page *page, int nr)
+{
+ VM_BUG_ON_PAGE(page != compound_head(page), page);
+ VM_BUG_ON_PAGE(page_count(page) == 0, page);
+ page_ref_add(page, nr);
+ SetPageReferenced(page);
+}
+
+static int __gup_device_huge(unsigned long pfn, unsigned long addr,
+ unsigned long end, struct page **pages, int *nr)
+{
+ int nr_start = *nr;
+ struct dev_pagemap *pgmap = NULL;
+
+ do {
+ struct page *page = pfn_to_page(pfn);
+
+ pgmap = get_dev_pagemap(pfn, pgmap);
+ if (unlikely(!pgmap)) {
+ undo_dev_pagemap(nr, nr_start, pages);
+ return 0;
+ }
+ SetPageReferenced(page);
+ pages[*nr] = page;
+ get_page(page);
+ put_dev_pagemap(pgmap);
+ (*nr)++;
+ pfn++;
+ } while (addr += PAGE_SIZE, addr != end);
+ return 1;
+}
+
+static int __gup_device_huge_pmd(pmd_t pmd, unsigned long addr,
+ unsigned long end, struct page **pages, int *nr)
+{
+ unsigned long fault_pfn;
+
+ fault_pfn = pmd_pfn(pmd) + ((addr & ~PMD_MASK) >> PAGE_SHIFT);
+ return __gup_device_huge(fault_pfn, addr, end, pages, nr);
+}
+
+static int __gup_device_huge_pud(pud_t pud, unsigned long addr,
+ unsigned long end, struct page **pages, int *nr)
+{
+ unsigned long fault_pfn;
+
+ fault_pfn = pud_pfn(pud) + ((addr & ~PUD_MASK) >> PAGE_SHIFT);
+ return __gup_device_huge(fault_pfn, addr, end, pages, nr);
+}
+
+static noinline int gup_huge_pmd(pmd_t pmd, unsigned long addr,
+ unsigned long end, int write, struct page **pages, int *nr)
+{
+ struct page *head, *page;
+ int refs;
+
+ if (!pte_allows_gup(pmd_val(pmd), write))
+ return 0;
+
+ VM_BUG_ON(!pfn_valid(pmd_pfn(pmd)));
+ if (pmd_devmap(pmd))
+ return __gup_device_huge_pmd(pmd, addr, end, pages, nr);
+
+ /* hugepages are never "special" */
+ VM_BUG_ON(pmd_flags(pmd) & _PAGE_SPECIAL);
+
+ refs = 0;
+ head = pmd_page(pmd);
+ page = head + ((addr & ~PMD_MASK) >> PAGE_SHIFT);
+ do {
+ VM_BUG_ON_PAGE(compound_head(page) != head, page);
+ pages[*nr] = page;
+ (*nr)++;
+ page++;
+ refs++;
+ } while (addr += PAGE_SIZE, addr != end);
+ get_head_page_multiple(head, refs);
+
+ return 1;
+}
+
+static int gup_pmd_range(pud_t pud, unsigned long addr, unsigned long end,
+ int write, struct page **pages, int *nr)
+{
+ unsigned long next;
+ pmd_t *pmdp;
+
+ pmdp = pmd_offset(&pud, addr);
+ do {
+ pmd_t pmd = *pmdp;
+
+ next = pmd_addr_end(addr, end);
+ if (pmd_none(pmd))
+ return 0;
+ if (unlikely(pmd_large(pmd) || !pmd_present(pmd))) {
+ /*
+ * NUMA hinting faults need to be handled in the GUP
+ * slowpath for accounting purposes and so that they
+ * can be serialised against THP migration.
+ */
+ if (pmd_protnone(pmd))
+ return 0;
+ if (!gup_huge_pmd(pmd, addr, next, write, pages, nr))
+ return 0;
+ } else {
+ if (!gup_pte_range(pmd, addr, next, write, pages, nr))
+ return 0;
+ }
+ } while (pmdp++, addr = next, addr != end);
+
+ return 1;
+}
+
+static noinline int gup_huge_pud(pud_t pud, unsigned long addr,
+ unsigned long end, int write, struct page **pages, int *nr)
+{
+ struct page *head, *page;
+ int refs;
+
+ if (!pte_allows_gup(pud_val(pud), write))
+ return 0;
+
+ VM_BUG_ON(!pfn_valid(pud_pfn(pud)));
+ if (pud_devmap(pud))
+ return __gup_device_huge_pud(pud, addr, end, pages, nr);
+
+ /* hugepages are never "special" */
+ VM_BUG_ON(pud_flags(pud) & _PAGE_SPECIAL);
+
+ refs = 0;
+ head = pud_page(pud);
+ page = head + ((addr & ~PUD_MASK) >> PAGE_SHIFT);
+ do {
+ VM_BUG_ON_PAGE(compound_head(page) != head, page);
+ pages[*nr] = page;
+ (*nr)++;
+ page++;
+ refs++;
+ } while (addr += PAGE_SIZE, addr != end);
+ get_head_page_multiple(head, refs);
+
+ return 1;
+}
+
+static int gup_pud_range(p4d_t p4d, unsigned long addr, unsigned long end,
+ int write, struct page **pages, int *nr)
+{
+ unsigned long next;
+ pud_t *pudp;
+
+ pudp = pud_offset(&p4d, addr);
+ do {
+ pud_t pud = *pudp;
+
+ next = pud_addr_end(addr, end);
+ if (pud_none(pud))
+ return 0;
+ if (unlikely(pud_large(pud))) {
+ if (!gup_huge_pud(pud, addr, next, write, pages, nr))
+ return 0;
+ } else {
+ if (!gup_pmd_range(pud, addr, next, write, pages, nr))
+ return 0;
+ }
+ } while (pudp++, addr = next, addr != end);
+
+ return 1;
+}
+
+static int gup_p4d_range(pgd_t pgd, unsigned long addr, unsigned long end,
+ int write, struct page **pages, int *nr)
+{
+ unsigned long next;
+ p4d_t *p4dp;
+
+ p4dp = p4d_offset(&pgd, addr);
+ do {
+ p4d_t p4d = *p4dp;
+
+ next = p4d_addr_end(addr, end);
+ if (p4d_none(p4d))
+ return 0;
+ BUILD_BUG_ON(p4d_large(p4d));
+ if (!gup_pud_range(p4d, addr, next, write, pages, nr))
+ return 0;
+ } while (p4dp++, addr = next, addr != end);
+
+ return 1;
+}
+
+/*
+ * Like get_user_pages_fast() except its IRQ-safe in that it won't fall
+ * back to the regular GUP.
+ */
+int __get_user_pages_fast(unsigned long start, int nr_pages, int write,
+ struct page **pages)
+{
+ struct mm_struct *mm = current->mm;
+ unsigned long addr, len, end;
+ unsigned long next;
+ unsigned long flags;
+ pgd_t *pgdp;
+ int nr = 0;
+
+ start &= PAGE_MASK;
+ addr = start;
+ len = (unsigned long) nr_pages << PAGE_SHIFT;
+ end = start + len;
+ if (unlikely(!access_ok(write ? VERIFY_WRITE : VERIFY_READ,
+ (void __user *)start, len)))
+ return 0;
+
+ /*
+ * XXX: batch / limit 'nr', to avoid large irq off latency
+ * needs some instrumenting to determine the common sizes used by
+ * important workloads (eg. DB2), and whether limiting the batch size
+ * will decrease performance.
+ *
+ * It seems like we're in the clear for the moment. Direct-IO is
+ * the main guy that batches up lots of get_user_pages, and even
+ * they are limited to 64-at-a-time which is not so many.
+ */
+ /*
+ * This doesn't prevent pagetable teardown, but does prevent
+ * the pagetables and pages from being freed on x86.
+ *
+ * So long as we atomically load page table pointers versus teardown
+ * (which we do on x86, with the above PAE exception), we can follow the
+ * address down to the the page and take a ref on it.
+ */
+ local_irq_save(flags);
+ pgdp = pgd_offset(mm, addr);
+ do {
+ pgd_t pgd = *pgdp;
+
+ next = pgd_addr_end(addr, end);
+ if (pgd_none(pgd))
+ break;
+ if (!gup_p4d_range(pgd, addr, next, write, pages, &nr))
+ break;
+ } while (pgdp++, addr = next, addr != end);
+ local_irq_restore(flags);
+
+ return nr;
+}
+
+/**
+ * get_user_pages_fast() - pin user pages in memory
+ * @start: starting user address
+ * @nr_pages: number of pages from start to pin
+ * @write: whether pages will be written to
+ * @pages: array that receives pointers to the pages pinned.
+ * Should be at least nr_pages long.
+ *
+ * Attempt to pin user pages in memory without taking mm->mmap_sem.
+ * If not successful, it will fall back to taking the lock and
+ * calling get_user_pages().
+ *
+ * Returns number of pages pinned. This may be fewer than the number
+ * requested. If nr_pages is 0 or negative, returns 0. If no pages
+ * were pinned, returns -errno.
+ */
+int get_user_pages_fast(unsigned long start, int nr_pages, int write,
+ struct page **pages)
+{
+ struct mm_struct *mm = current->mm;
+ unsigned long addr, len, end;
+ unsigned long next;
+ pgd_t *pgdp;
+ int nr = 0;
+
+ start &= PAGE_MASK;
+ addr = start;
+ len = (unsigned long) nr_pages << PAGE_SHIFT;
+
+ end = start + len;
+ if (end < start)
+ goto slow_irqon;
+
+#ifdef CONFIG_X86_64
+ if (end >> __VIRTUAL_MASK_SHIFT)
+ goto slow_irqon;
+#endif
+
+ /*
+ * XXX: batch / limit 'nr', to avoid large irq off latency
+ * needs some instrumenting to determine the common sizes used by
+ * important workloads (eg. DB2), and whether limiting the batch size
+ * will decrease performance.
+ *
+ * It seems like we're in the clear for the moment. Direct-IO is
+ * the main guy that batches up lots of get_user_pages, and even
+ * they are limited to 64-at-a-time which is not so many.
+ */
+ /*
+ * This doesn't prevent pagetable teardown, but does prevent
+ * the pagetables and pages from being freed on x86.
+ *
+ * So long as we atomically load page table pointers versus teardown
+ * (which we do on x86, with the above PAE exception), we can follow the
+ * address down to the the page and take a ref on it.
+ */
+ local_irq_disable();
+ pgdp = pgd_offset(mm, addr);
+ do {
+ pgd_t pgd = *pgdp;
+
+ next = pgd_addr_end(addr, end);
+ if (pgd_none(pgd))
+ goto slow;
+ if (!gup_p4d_range(pgd, addr, next, write, pages, &nr))
+ goto slow;
+ } while (pgdp++, addr = next, addr != end);
+ local_irq_enable();
+
+ VM_BUG_ON(nr != (end - start) >> PAGE_SHIFT);
+ return nr;
+
+ {
+ int ret;
+
+slow:
+ local_irq_enable();
+slow_irqon:
+ /* Try to get the remaining pages with get_user_pages */
+ start += nr << PAGE_SHIFT;
+ pages += nr;
+
+ ret = get_user_pages_unlocked(start,
+ (end - start) >> PAGE_SHIFT,
+ pages, write ? FOLL_WRITE : 0);
+
+ /* Have to be a bit careful with return values */
+ if (nr > 0) {
+ if (ret < 0)
+ ret = nr;
+ else
+ ret += nr;
+ }
+
+ return ret;
+ }
+}
diff --git a/mm/Kconfig b/mm/Kconfig
index c89f472b658c..9b8fccb969dc 100644
--- a/mm/Kconfig
+++ b/mm/Kconfig
@@ -137,7 +137,7 @@ config HAVE_MEMBLOCK_NODE_MAP
config HAVE_MEMBLOCK_PHYS_MAP
bool
-config HAVE_GENERIC_GUP
+config HAVE_GENERIC_RCU_GUP
bool
config ARCH_DISCARD_MEMBLOCK
diff --git a/mm/gup.c b/mm/gup.c
index 2559a3987de7..527ec2c6cca3 100644
--- a/mm/gup.c
+++ b/mm/gup.c
@@ -1155,7 +1155,7 @@ struct page *get_dump_page(unsigned long addr)
#endif /* CONFIG_ELF_CORE */
/*
- * Generic Fast GUP
+ * Generic RCU Fast GUP
*
* get_user_pages_fast attempts to pin user pages by walking the page
* tables directly and avoids taking locks. Thus the walker needs to be
@@ -1176,8 +1176,8 @@ struct page *get_dump_page(unsigned long addr)
* Before activating this code, please be aware that the following assumptions
* are currently made:
*
- * *) Either HAVE_RCU_TABLE_FREE is enabled, and tlb_remove_table() is used to
- * free pages containing page tables or TLB flushing requires IPI broadcast.
+ * *) HAVE_RCU_TABLE_FREE is enabled, and tlb_remove_table is used to free
+ * pages containing page tables.
*
* *) ptes can be read atomically by the architecture.
*
@@ -1187,7 +1187,7 @@ struct page *get_dump_page(unsigned long addr)
*
* This code is based heavily on the PowerPC implementation by Nick Piggin.
*/
-#ifdef CONFIG_HAVE_GENERIC_GUP
+#ifdef CONFIG_HAVE_GENERIC_RCU_GUP
#ifndef gup_get_pte
/*
@@ -1677,4 +1677,4 @@ int get_user_pages_fast(unsigned long start, int nr_pages, int write,
return ret;
}
-#endif /* CONFIG_HAVE_GENERIC_GUP */
+#endif /* CONFIG_HAVE_GENERIC_RCU_GUP */
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2017-04-24 01:40 +0200 |
| Subject | get_zone_device_page() in get_page() and page_cache_get_speculative() |
| Message-ID | <tzzgJ-6ZY-3@gated-at.bofh.it> |
| In reply to | #1627845 |
On Thu, Apr 20, 2017 at 02:46:51PM -0700, Dan Williams wrote: > On Sat, Mar 18, 2017 at 2:52 AM, tip-bot for Kirill A. Shutemov > <tipbot@zytor.com> wrote: > > Commit-ID: 2947ba054a4dabbd82848728d765346886050029 > > Gitweb: http://git.kernel.org/tip/2947ba054a4dabbd82848728d765346886050029 > > Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> > > AuthorDate: Fri, 17 Mar 2017 00:39:06 +0300 > > Committer: Ingo Molnar <mingo@kernel.org> > > CommitDate: Sat, 18 Mar 2017 09:48:03 +0100 > > > > x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation > > > > This patch provides all required callbacks required by the generic > > get_user_pages_fast() code and switches x86 over - and removes > > the platform specific implementation. > > > > Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> > > Cc: Andrew Morton <akpm@linux-foundation.org> > > Cc: Aneesh Kumar K . V <aneesh.kumar@linux.vnet.ibm.com> > > Cc: Borislav Petkov <bp@alien8.de> > > Cc: Catalin Marinas <catalin.marinas@arm.com> > > Cc: Dann Frazier <dann.frazier@canonical.com> > > Cc: Dave Hansen <dave.hansen@intel.com> > > Cc: H. Peter Anvin <hpa@zytor.com> > > Cc: Linus Torvalds <torvalds@linux-foundation.org> > > Cc: Peter Zijlstra <peterz@infradead.org> > > Cc: Rik van Riel <riel@redhat.com> > > Cc: Steve Capper <steve.capper@linaro.org> > > Cc: Thomas Gleixner <tglx@linutronix.de> > > Cc: linux-arch@vger.kernel.org > > Cc: linux-mm@kvack.org > > Link: http://lkml.kernel.org/r/20170316213906.89528-1-kirill.shutemov@linux.intel.com > > [ Minor readability edits. ] > > Signed-off-by: Ingo Molnar <mingo@kernel.org> > > I'm still trying to spot the bug, but bisect points to this patch as > the point at which my unit tests start failing with the following > signature: > > [ 35.423841] WARNING: CPU: 8 PID: 245 at lib/percpu-refcount.c:155 > percpu_ref_switch_to_atomic_rcu+0x1f5/0x200 Okay, I've tracked it down. The issue is triggered by replacment get_page() with page_cache_get_speculative(). page_cache_get_speculative() doesn't have get_zone_device_page(). :-| And I think it's your bug, Dan: it's wrong to have get_/put_zone_device_page() in get_/put_page(). I must be handled by page_ref_* machinery to catch all cases where we manipulate with page refcount. Back to the big picture: I hate that we need to have such additional code in page refcount primitives. I worked hard to remove compound page ugliness from there and now zone_device creeping in... Is it the only option? -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-24 19:30 +0200 |
| Subject | Re: get_zone_device_page() in get_page() and page_cache_get_speculative() |
| Message-ID | <tzPYd-1pP-7@gated-at.bofh.it> |
| In reply to | #1629091 |
On Sun, Apr 23, 2017 at 4:31 PM, Kirill A. Shutemov <kirill@shutemov.name> wrote: > On Thu, Apr 20, 2017 at 02:46:51PM -0700, Dan Williams wrote: >> On Sat, Mar 18, 2017 at 2:52 AM, tip-bot for Kirill A. Shutemov >> <tipbot@zytor.com> wrote: >> > Commit-ID: 2947ba054a4dabbd82848728d765346886050029 >> > Gitweb: http://git.kernel.org/tip/2947ba054a4dabbd82848728d765346886050029 >> > Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> >> > AuthorDate: Fri, 17 Mar 2017 00:39:06 +0300 >> > Committer: Ingo Molnar <mingo@kernel.org> >> > CommitDate: Sat, 18 Mar 2017 09:48:03 +0100 >> > >> > x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation >> > >> > This patch provides all required callbacks required by the generic >> > get_user_pages_fast() code and switches x86 over - and removes >> > the platform specific implementation. >> > >> > Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> >> > Cc: Andrew Morton <akpm@linux-foundation.org> >> > Cc: Aneesh Kumar K . V <aneesh.kumar@linux.vnet.ibm.com> >> > Cc: Borislav Petkov <bp@alien8.de> >> > Cc: Catalin Marinas <catalin.marinas@arm.com> >> > Cc: Dann Frazier <dann.frazier@canonical.com> >> > Cc: Dave Hansen <dave.hansen@intel.com> >> > Cc: H. Peter Anvin <hpa@zytor.com> >> > Cc: Linus Torvalds <torvalds@linux-foundation.org> >> > Cc: Peter Zijlstra <peterz@infradead.org> >> > Cc: Rik van Riel <riel@redhat.com> >> > Cc: Steve Capper <steve.capper@linaro.org> >> > Cc: Thomas Gleixner <tglx@linutronix.de> >> > Cc: linux-arch@vger.kernel.org >> > Cc: linux-mm@kvack.org >> > Link: http://lkml.kernel.org/r/20170316213906.89528-1-kirill.shutemov@linux.intel.com >> > [ Minor readability edits. ] >> > Signed-off-by: Ingo Molnar <mingo@kernel.org> >> >> I'm still trying to spot the bug, but bisect points to this patch as >> the point at which my unit tests start failing with the following >> signature: >> >> [ 35.423841] WARNING: CPU: 8 PID: 245 at lib/percpu-refcount.c:155 >> percpu_ref_switch_to_atomic_rcu+0x1f5/0x200 > > Okay, I've tracked it down. The issue is triggered by replacment > get_page() with page_cache_get_speculative(). > > page_cache_get_speculative() doesn't have get_zone_device_page(). :-| > > And I think it's your bug, Dan: it's wrong to have > get_/put_zone_device_page() in get_/put_page(). I must be handled by > page_ref_* machinery to catch all cases where we manipulate with page > refcount. The page_ref conversion landed in 4.6 *after* the ZONE_DEVICE implementation that landed in 4.5, so there was a missed conversion of the zone-device reference counting to page_ref. > Back to the big picture: > > I hate that we need to have such additional code in page refcount > primitives. I worked hard to remove compound page ugliness from there and > now zone_device creeping in... > > Is it the only option? Not sure, I need to spend some time to understand what page_ref means to ZONE_DEVICE.
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2017-04-24 19:40 +0200 |
| Subject | Re: get_zone_device_page() in get_page() and page_cache_get_speculative() |
| Message-ID | <tzQ7U-1tV-13@gated-at.bofh.it> |
| In reply to | #1629808 |
On Mon, Apr 24, 2017 at 10:23:59AM -0700, Dan Williams wrote: > On Sun, Apr 23, 2017 at 4:31 PM, Kirill A. Shutemov > <kirill@shutemov.name> wrote: > > On Thu, Apr 20, 2017 at 02:46:51PM -0700, Dan Williams wrote: > >> On Sat, Mar 18, 2017 at 2:52 AM, tip-bot for Kirill A. Shutemov > >> <tipbot@zytor.com> wrote: > >> > Commit-ID: 2947ba054a4dabbd82848728d765346886050029 > >> > Gitweb: http://git.kernel.org/tip/2947ba054a4dabbd82848728d765346886050029 > >> > Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> > >> > AuthorDate: Fri, 17 Mar 2017 00:39:06 +0300 > >> > Committer: Ingo Molnar <mingo@kernel.org> > >> > CommitDate: Sat, 18 Mar 2017 09:48:03 +0100 > >> > > >> > x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation > >> > > >> > This patch provides all required callbacks required by the generic > >> > get_user_pages_fast() code and switches x86 over - and removes > >> > the platform specific implementation. > >> > > >> > Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> > >> > Cc: Andrew Morton <akpm@linux-foundation.org> > >> > Cc: Aneesh Kumar K . V <aneesh.kumar@linux.vnet.ibm.com> > >> > Cc: Borislav Petkov <bp@alien8.de> > >> > Cc: Catalin Marinas <catalin.marinas@arm.com> > >> > Cc: Dann Frazier <dann.frazier@canonical.com> > >> > Cc: Dave Hansen <dave.hansen@intel.com> > >> > Cc: H. Peter Anvin <hpa@zytor.com> > >> > Cc: Linus Torvalds <torvalds@linux-foundation.org> > >> > Cc: Peter Zijlstra <peterz@infradead.org> > >> > Cc: Rik van Riel <riel@redhat.com> > >> > Cc: Steve Capper <steve.capper@linaro.org> > >> > Cc: Thomas Gleixner <tglx@linutronix.de> > >> > Cc: linux-arch@vger.kernel.org > >> > Cc: linux-mm@kvack.org > >> > Link: http://lkml.kernel.org/r/20170316213906.89528-1-kirill.shutemov@linux.intel.com > >> > [ Minor readability edits. ] > >> > Signed-off-by: Ingo Molnar <mingo@kernel.org> > >> > >> I'm still trying to spot the bug, but bisect points to this patch as > >> the point at which my unit tests start failing with the following > >> signature: > >> > >> [ 35.423841] WARNING: CPU: 8 PID: 245 at lib/percpu-refcount.c:155 > >> percpu_ref_switch_to_atomic_rcu+0x1f5/0x200 > > > > Okay, I've tracked it down. The issue is triggered by replacment > > get_page() with page_cache_get_speculative(). > > > > page_cache_get_speculative() doesn't have get_zone_device_page(). :-| > > > > And I think it's your bug, Dan: it's wrong to have > > get_/put_zone_device_page() in get_/put_page(). I must be handled by > > page_ref_* machinery to catch all cases where we manipulate with page > > refcount. > > The page_ref conversion landed in 4.6 *after* the ZONE_DEVICE > implementation that landed in 4.5, so there was a missed conversion of > the zone-device reference counting to page_ref. Fair enough. But get_page_unless_zero() definitely predates ZONE_DEVICE. :) -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-24 19:50 +0200 |
| Subject | Re: get_zone_device_page() in get_page() and page_cache_get_speculative() |
| Message-ID | <tzQhA-1yE-11@gated-at.bofh.it> |
| In reply to | #1629815 |
On Mon, Apr 24, 2017 at 10:30 AM, Kirill A. Shutemov <kirill@shutemov.name> wrote: > On Mon, Apr 24, 2017 at 10:23:59AM -0700, Dan Williams wrote: >> On Sun, Apr 23, 2017 at 4:31 PM, Kirill A. Shutemov >> <kirill@shutemov.name> wrote: >> > On Thu, Apr 20, 2017 at 02:46:51PM -0700, Dan Williams wrote: >> >> On Sat, Mar 18, 2017 at 2:52 AM, tip-bot for Kirill A. Shutemov >> >> <tipbot@zytor.com> wrote: >> >> > Commit-ID: 2947ba054a4dabbd82848728d765346886050029 >> >> > Gitweb: http://git.kernel.org/tip/2947ba054a4dabbd82848728d765346886050029 >> >> > Author: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> >> >> > AuthorDate: Fri, 17 Mar 2017 00:39:06 +0300 >> >> > Committer: Ingo Molnar <mingo@kernel.org> >> >> > CommitDate: Sat, 18 Mar 2017 09:48:03 +0100 >> >> > >> >> > x86/mm/gup: Switch GUP to the generic get_user_page_fast() implementation >> >> > >> >> > This patch provides all required callbacks required by the generic >> >> > get_user_pages_fast() code and switches x86 over - and removes >> >> > the platform specific implementation. >> >> > >> >> > Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> >> >> > Cc: Andrew Morton <akpm@linux-foundation.org> >> >> > Cc: Aneesh Kumar K . V <aneesh.kumar@linux.vnet.ibm.com> >> >> > Cc: Borislav Petkov <bp@alien8.de> >> >> > Cc: Catalin Marinas <catalin.marinas@arm.com> >> >> > Cc: Dann Frazier <dann.frazier@canonical.com> >> >> > Cc: Dave Hansen <dave.hansen@intel.com> >> >> > Cc: H. Peter Anvin <hpa@zytor.com> >> >> > Cc: Linus Torvalds <torvalds@linux-foundation.org> >> >> > Cc: Peter Zijlstra <peterz@infradead.org> >> >> > Cc: Rik van Riel <riel@redhat.com> >> >> > Cc: Steve Capper <steve.capper@linaro.org> >> >> > Cc: Thomas Gleixner <tglx@linutronix.de> >> >> > Cc: linux-arch@vger.kernel.org >> >> > Cc: linux-mm@kvack.org >> >> > Link: http://lkml.kernel.org/r/20170316213906.89528-1-kirill.shutemov@linux.intel.com >> >> > [ Minor readability edits. ] >> >> > Signed-off-by: Ingo Molnar <mingo@kernel.org> >> >> >> >> I'm still trying to spot the bug, but bisect points to this patch as >> >> the point at which my unit tests start failing with the following >> >> signature: >> >> >> >> [ 35.423841] WARNING: CPU: 8 PID: 245 at lib/percpu-refcount.c:155 >> >> percpu_ref_switch_to_atomic_rcu+0x1f5/0x200 >> > >> > Okay, I've tracked it down. The issue is triggered by replacment >> > get_page() with page_cache_get_speculative(). >> > >> > page_cache_get_speculative() doesn't have get_zone_device_page(). :-| >> > >> > And I think it's your bug, Dan: it's wrong to have >> > get_/put_zone_device_page() in get_/put_page(). I must be handled by >> > page_ref_* machinery to catch all cases where we manipulate with page >> > refcount. >> >> The page_ref conversion landed in 4.6 *after* the ZONE_DEVICE >> implementation that landed in 4.5, so there was a missed conversion of >> the zone-device reference counting to page_ref. > > Fair enough. > > But get_page_unless_zero() definitely predates ZONE_DEVICE. :) > It does, but that's deliberate. A ZONE_DEVICE page never has a zero reference count, it's always owned by the device, never by the page allocator. ZONE_DEVICE overrides the ->lru list_head to store private device information and we rely on the behavior that a non-zero reference means the page is not added to any lru or page cache list.
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2017-04-24 20:10 +0200 |
| Subject | Re: get_zone_device_page() in get_page() and page_cache_get_speculative() |
| Message-ID | <tzQAW-1Uk-15@gated-at.bofh.it> |
| In reply to | #1629821 |
On Mon, Apr 24, 2017 at 10:47:43AM -0700, Dan Williams wrote: > On Mon, Apr 24, 2017 at 10:30 AM, Kirill A. Shutemov > >> >> [ 35.423841] WARNING: CPU: 8 PID: 245 at lib/percpu-refcount.c:155 > >> >> percpu_ref_switch_to_atomic_rcu+0x1f5/0x200 > >> > > >> > Okay, I've tracked it down. The issue is triggered by replacment > >> > get_page() with page_cache_get_speculative(). > >> > > >> > page_cache_get_speculative() doesn't have get_zone_device_page(). :-| > >> > > >> > And I think it's your bug, Dan: it's wrong to have > >> > get_/put_zone_device_page() in get_/put_page(). I must be handled by > >> > page_ref_* machinery to catch all cases where we manipulate with page > >> > refcount. > >> > >> The page_ref conversion landed in 4.6 *after* the ZONE_DEVICE > >> implementation that landed in 4.5, so there was a missed conversion of > >> the zone-device reference counting to page_ref. > > > > Fair enough. > > > > But get_page_unless_zero() definitely predates ZONE_DEVICE. :) > > > > It does, but that's deliberate. A ZONE_DEVICE page never has a zero > reference count, it's always owned by the device, never by the page > allocator. ZONE_DEVICE overrides the ->lru list_head to store private > device information and we rely on the behavior that a non-zero > reference means the page is not added to any lru or page cache list. So, what do you propose? Use get_page() instead of page_cache_get_speculative() in GUP_fast() if the page belong to zone device? I don't like it. This situation, when we only can use subset of helpers to manipulate page refcount creates situation waiting to explode. I think it's still better to do it on page_ref_* level. BTW, why do we need to pin pgmap from get_page() in first place? I don't have enough background in ZONE_DEVICE. -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> |
|---|---|
| Date | 2017-04-24 20:30 +0200 |
| Subject | Re: get_zone_device_page() in get_page() and page_cache_get_speculative() |
| Message-ID | <tzQUk-21B-47@gated-at.bofh.it> |
| In reply to | #1629842 |
On Mon, Apr 24, 2017 at 09:01:58PM +0300, Kirill A. Shutemov wrote:
> On Mon, Apr 24, 2017 at 10:47:43AM -0700, Dan Williams wrote:
> I think it's still better to do it on page_ref_* level.
Something like patch below? What do you think?
diff --git a/include/linux/memremap.h b/include/linux/memremap.h
index 93416196ba64..bd1b13af4567 100644
--- a/include/linux/memremap.h
+++ b/include/linux/memremap.h
@@ -35,20 +35,6 @@ static inline struct vmem_altmap *to_vmem_altmap(unsigned long memmap_start)
}
#endif
-/**
- * struct dev_pagemap - metadata for ZONE_DEVICE mappings
- * @altmap: pre-allocated/reserved memory for vmemmap allocations
- * @res: physical address range covered by @ref
- * @ref: reference count that pins the devm_memremap_pages() mapping
- * @dev: host device of the mapping for debug
- */
-struct dev_pagemap {
- struct vmem_altmap *altmap;
- const struct resource *res;
- struct percpu_ref *ref;
- struct device *dev;
-};
-
#ifdef CONFIG_ZONE_DEVICE
void *devm_memremap_pages(struct device *dev, struct resource *res,
struct percpu_ref *ref, struct vmem_altmap *altmap);
diff --git a/include/linux/mm.h b/include/linux/mm.h
index e197d3ca3e8a..c2749b878199 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -760,19 +760,11 @@ static inline enum zone_type page_zonenum(const struct page *page)
}
#ifdef CONFIG_ZONE_DEVICE
-void get_zone_device_page(struct page *page);
-void put_zone_device_page(struct page *page);
static inline bool is_zone_device_page(const struct page *page)
{
return page_zonenum(page) == ZONE_DEVICE;
}
#else
-static inline void get_zone_device_page(struct page *page)
-{
-}
-static inline void put_zone_device_page(struct page *page)
-{
-}
static inline bool is_zone_device_page(const struct page *page)
{
return false;
@@ -788,9 +780,6 @@ static inline void get_page(struct page *page)
*/
VM_BUG_ON_PAGE(page_ref_count(page) <= 0, page);
page_ref_inc(page);
-
- if (unlikely(is_zone_device_page(page)))
- get_zone_device_page(page);
}
static inline void put_page(struct page *page)
@@ -799,9 +788,6 @@ static inline void put_page(struct page *page)
if (put_page_testzero(page))
__put_page(page);
-
- if (unlikely(is_zone_device_page(page)))
- put_zone_device_page(page);
}
#if defined(CONFIG_SPARSEMEM) && !defined(CONFIG_SPARSEMEM_VMEMMAP)
diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h
index 45cdb27791a3..fb7bb60d446b 100644
--- a/include/linux/mm_types.h
+++ b/include/linux/mm_types.h
@@ -601,4 +601,18 @@ typedef struct {
unsigned long val;
} swp_entry_t;
+/**
+ * struct dev_pagemap - metadata for ZONE_DEVICE mappings
+ * @altmap: pre-allocated/reserved memory for vmemmap allocations
+ * @res: physical address range covered by @ref
+ * @ref: reference count that pins the devm_memremap_pages() mapping
+ * @dev: host device of the mapping for debug
+ */
+struct dev_pagemap {
+ struct vmem_altmap *altmap;
+ const struct resource *res;
+ struct percpu_ref *ref;
+ struct device *dev;
+};
+
#endif /* _LINUX_MM_TYPES_H */
diff --git a/include/linux/page_ref.h b/include/linux/page_ref.h
index 610e13271918..d834c68e21fd 100644
--- a/include/linux/page_ref.h
+++ b/include/linux/page_ref.h
@@ -61,6 +61,8 @@ static inline void __page_ref_unfreeze(struct page *page, int v)
#endif
+static inline bool is_zone_device_page(const struct page *page);
+
static inline int page_ref_count(struct page *page)
{
return atomic_read(&page->_refcount);
@@ -92,6 +94,9 @@ static inline void page_ref_add(struct page *page, int nr)
atomic_add(nr, &page->_refcount);
if (page_ref_tracepoint_active(__tracepoint_page_ref_mod))
__page_ref_mod(page, nr);
+
+ if (unlikely(is_zone_device_page(page)))
+ percpu_ref_get_many(page->pgmap->ref, nr);
}
static inline void page_ref_sub(struct page *page, int nr)
@@ -99,6 +104,9 @@ static inline void page_ref_sub(struct page *page, int nr)
atomic_sub(nr, &page->_refcount);
if (page_ref_tracepoint_active(__tracepoint_page_ref_mod))
__page_ref_mod(page, -nr);
+
+ if (unlikely(is_zone_device_page(page)))
+ percpu_ref_put_many(page->pgmap->ref, nr);
}
static inline void page_ref_inc(struct page *page)
@@ -106,6 +114,9 @@ static inline void page_ref_inc(struct page *page)
atomic_inc(&page->_refcount);
if (page_ref_tracepoint_active(__tracepoint_page_ref_mod))
__page_ref_mod(page, 1);
+
+ if (unlikely(is_zone_device_page(page)))
+ percpu_ref_get(page->pgmap->ref);
}
static inline void page_ref_dec(struct page *page)
@@ -113,6 +124,9 @@ static inline void page_ref_dec(struct page *page)
atomic_dec(&page->_refcount);
if (page_ref_tracepoint_active(__tracepoint_page_ref_mod))
__page_ref_mod(page, -1);
+
+ if (unlikely(is_zone_device_page(page)))
+ percpu_ref_put(page->pgmap->ref);
}
static inline int page_ref_sub_and_test(struct page *page, int nr)
@@ -121,6 +135,9 @@ static inline int page_ref_sub_and_test(struct page *page, int nr)
if (page_ref_tracepoint_active(__tracepoint_page_ref_mod_and_test))
__page_ref_mod_and_test(page, -nr, ret);
+
+ if (unlikely(is_zone_device_page(page)))
+ percpu_ref_put_many(page->pgmap->ref, nr);
return ret;
}
@@ -130,6 +147,9 @@ static inline int page_ref_inc_return(struct page *page)
if (page_ref_tracepoint_active(__tracepoint_page_ref_mod_and_return))
__page_ref_mod_and_return(page, 1, ret);
+
+ if (unlikely(is_zone_device_page(page)))
+ percpu_ref_get(page->pgmap->ref);
return ret;
}
@@ -139,6 +159,9 @@ static inline int page_ref_dec_and_test(struct page *page)
if (page_ref_tracepoint_active(__tracepoint_page_ref_mod_and_test))
__page_ref_mod_and_test(page, -1, ret);
+
+ if (unlikely(is_zone_device_page(page)))
+ percpu_ref_put(page->pgmap->ref);
return ret;
}
@@ -148,6 +171,9 @@ static inline int page_ref_dec_return(struct page *page)
if (page_ref_tracepoint_active(__tracepoint_page_ref_mod_and_return))
__page_ref_mod_and_return(page, -1, ret);
+
+ if (unlikely(is_zone_device_page(page)))
+ percpu_ref_put(page->pgmap->ref);
return ret;
}
@@ -157,6 +183,9 @@ static inline int page_ref_add_unless(struct page *page, int nr, int u)
if (page_ref_tracepoint_active(__tracepoint_page_ref_mod_unless))
__page_ref_mod_unless(page, nr, ret);
+
+ if (unlikely(is_zone_device_page(page)) && ret)
+ percpu_ref_get_many(page->pgmap->ref, nr);
return ret;
}
@@ -166,6 +195,9 @@ static inline int page_ref_freeze(struct page *page, int count)
if (page_ref_tracepoint_active(__tracepoint_page_ref_freeze))
__page_ref_freeze(page, count, ret);
+
+ if (unlikely(is_zone_device_page(page)) && ret)
+ percpu_ref_put_many(page->pgmap->ref, count);
return ret;
}
@@ -177,6 +209,9 @@ static inline void page_ref_unfreeze(struct page *page, int count)
atomic_set(&page->_refcount, count);
if (page_ref_tracepoint_active(__tracepoint_page_ref_unfreeze))
__page_ref_unfreeze(page, count);
+
+ if (unlikely(is_zone_device_page(page)))
+ percpu_ref_get_many(page->pgmap->ref, count);
}
#endif
diff --git a/kernel/memremap.c b/kernel/memremap.c
index 06123234f118..936cef79d811 100644
--- a/kernel/memremap.c
+++ b/kernel/memremap.c
@@ -182,18 +182,6 @@ struct page_map {
struct vmem_altmap altmap;
};
-void get_zone_device_page(struct page *page)
-{
- percpu_ref_get(page->pgmap->ref);
-}
-EXPORT_SYMBOL(get_zone_device_page);
-
-void put_zone_device_page(struct page *page)
-{
- put_dev_pagemap(page->pgmap);
-}
-EXPORT_SYMBOL(put_zone_device_page);
-
static void pgmap_radix_release(struct resource *res)
{
resource_size_t key, align_start, align_size, align_end;
--
Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-04-24 20:50 +0200 |
| Subject | Re: get_zone_device_page() in get_page() and page_cache_get_speculative() |
| Message-ID | <tzRdD-29v-15@gated-at.bofh.it> |
| In reply to | #1629882 |
On Mon, Apr 24, 2017 at 11:25 AM, Kirill A. Shutemov <kirill.shutemov@linux.intel.com> wrote: > On Mon, Apr 24, 2017 at 09:01:58PM +0300, Kirill A. Shutemov wrote: >> On Mon, Apr 24, 2017 at 10:47:43AM -0700, Dan Williams wrote: >> I think it's still better to do it on page_ref_* level. > > Something like patch below? What do you think? From a quick glance, I think this looks like the right way to go.
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web