Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1372241

Re: [PATCH 12/31] huge tmpfs: extend get_user_pages_fast to shmem pmd

From Ingo Molnar <mingo@kernel.org>
Newsgroups linux.kernel
Subject Re: [PATCH 12/31] huge tmpfs: extend get_user_pages_fast to shmem pmd
Date 2016-04-06 09:10 +0200
Message-ID <rkPLc-gu-9@gated-at.bofh.it> (permalink)
References <rkGye-1F1-7@gated-at.bofh.it> <rkGRA-1Og-19@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


* Hugh Dickins <hughd@google.com> wrote:

> The arch-specific get_user_pages_fast() has a gup_huge_pmd() designed to
> optimize the refcounting on anonymous THP and hugetlbfs pages, with one
> atomic addition to compound head's common refcount.  That optimization
> must be avoided on huge tmpfs team pages, which use normal separate page
> refcounting.  We could combine the PageTeam and PageCompound cases into
> a single simple loop, but would lose the compound optimization that way.
> 
> One cannot go through these functions without wondering why some arches
> (x86, mips) like to SetPageReferenced, while the rest do not: an x86
> optimization that missed being propagated to the other architectures?
> No, see commit 8ee53820edfd ("thp: mmu_notifier_test_young"): it's a
> KVM GRU EPT thing, maybe not useful beyond x86.  I've just followed
> the established practice in each architecture.
> 
> Signed-off-by: Hugh Dickins <hughd@google.com>
> ---
> Cc'ed to arch maintainers as an FYI: this patch is not expected to
> go into the tree in the next few weeks, and depends upon a PageTeam
> definition not yet available outside this huge tmpfs patchset.
> Please refer to linux-mm or linux-kernel for more context.
> 
>  arch/mips/mm/gup.c  |   15 ++++++++++++++-
>  arch/s390/mm/gup.c  |   19 ++++++++++++++++++-
>  arch/sparc/mm/gup.c |   19 ++++++++++++++++++-
>  arch/x86/mm/gup.c   |   15 ++++++++++++++-
>  mm/gup.c            |   19 ++++++++++++++++++-
>  5 files changed, 82 insertions(+), 5 deletions(-)
> 
> --- a/arch/mips/mm/gup.c
> +++ b/arch/mips/mm/gup.c
> @@ -81,9 +81,22 @@ static int gup_huge_pmd(pmd_t pmd, unsig
>  	VM_BUG_ON(pte_special(pte));
>  	VM_BUG_ON(!pfn_valid(pte_pfn(pte)));
>  
> -	refs = 0;
>  	head = pte_page(pte);
>  	page = head + ((addr & ~PMD_MASK) >> PAGE_SHIFT);
> +
> +	if (PageTeam(head)) {
> +		/* Handle a huge tmpfs team with normal refcounting. */
> +		do {
> +			get_page(page);
> +			SetPageReferenced(page);
> +			pages[*nr] = page;
> +			(*nr)++;
> +			page++;
> +		} while (addr += PAGE_SIZE, addr != end);
> +		return 1;
> +	}
> +
> +	refs = 0;
>  	do {
>  		VM_BUG_ON(compound_head(page) != head);
>  		pages[*nr] = page;
> --- a/arch/s390/mm/gup.c
> +++ b/arch/s390/mm/gup.c
> @@ -66,9 +66,26 @@ static inline int gup_huge_pmd(pmd_t *pm
>  		return 0;
>  	VM_BUG_ON(!pfn_valid(pmd_val(pmd) >> PAGE_SHIFT));
>  
> -	refs = 0;
>  	head = pmd_page(pmd);
>  	page = head + ((addr & ~PMD_MASK) >> PAGE_SHIFT);
> +
> +	if (PageTeam(head)) {
> +		/* Handle a huge tmpfs team with normal refcounting. */
> +		do {
> +			if (!page_cache_get_speculative(page))
> +				return 0;
> +			if (unlikely(pmd_val(pmd) != pmd_val(*pmdp))) {
> +				put_page(page);
> +				return 0;
> +			}
> +			pages[*nr] = page;
> +			(*nr)++;
> +			page++;
> +		} while (addr += PAGE_SIZE, addr != end);
> +		return 1;
> +	}
> +
> +	refs = 0;
>  	do {
>  		VM_BUG_ON(compound_head(page) != head);
>  		pages[*nr] = page;
> --- a/arch/sparc/mm/gup.c
> +++ b/arch/sparc/mm/gup.c
> @@ -77,9 +77,26 @@ static int gup_huge_pmd(pmd_t *pmdp, pmd
>  	if (write && !pmd_write(pmd))
>  		return 0;
>  
> -	refs = 0;
>  	head = pmd_page(pmd);
>  	page = head + ((addr & ~PMD_MASK) >> PAGE_SHIFT);
> +
> +	if (PageTeam(head)) {
> +		/* Handle a huge tmpfs team with normal refcounting. */
> +		do {
> +			if (!page_cache_get_speculative(page))
> +				return 0;
> +			if (unlikely(pmd_val(pmd) != pmd_val(*pmdp))) {
> +				put_page(page);
> +				return 0;
> +			}
> +			pages[*nr] = page;
> +			(*nr)++;
> +			page++;
> +		} while (addr += PAGE_SIZE, addr != end);
> +		return 1;
> +	}
> +
> +	refs = 0;
>  	do {
>  		VM_BUG_ON(compound_head(page) != head);
>  		pages[*nr] = page;
> --- a/arch/x86/mm/gup.c
> +++ b/arch/x86/mm/gup.c
> @@ -196,9 +196,22 @@ static noinline int gup_huge_pmd(pmd_t p
>  	/* hugepages are never "special" */
>  	VM_BUG_ON(pmd_flags(pmd) & _PAGE_SPECIAL);
>  
> -	refs = 0;
>  	head = pmd_page(pmd);
>  	page = head + ((addr & ~PMD_MASK) >> PAGE_SHIFT);
> +
> +	if (PageTeam(head)) {
> +		/* Handle a huge tmpfs team with normal refcounting. */
> +		do {
> +			get_page(page);
> +			SetPageReferenced(page);
> +			pages[*nr] = page;
> +			(*nr)++;
> +			page++;
> +		} while (addr += PAGE_SIZE, addr != end);
> +		return 1;
> +	}
> +
> +	refs = 0;
>  	do {
>  		VM_BUG_ON_PAGE(compound_head(page) != head, page);
>  		pages[*nr] = page;
> --- a/mm/gup.c
> +++ b/mm/gup.c
> @@ -1247,9 +1247,26 @@ static int gup_huge_pmd(pmd_t orig, pmd_
>  	if (write && !pmd_write(orig))
>  		return 0;
>  
> -	refs = 0;
>  	head = pmd_page(orig);
>  	page = head + ((addr & ~PMD_MASK) >> PAGE_SHIFT);
> +
> +	if (PageTeam(head)) {
> +		/* Handle a huge tmpfs team with normal refcounting. */
> +		do {
> +			if (!page_cache_get_speculative(page))
> +				return 0;
> +			if (unlikely(pmd_val(orig) != pmd_val(*pmdp))) {
> +				put_page(page);
> +				return 0;
> +			}
> +			pages[*nr] = page;
> +			(*nr)++;
> +			page++;
> +		} while (addr += PAGE_SIZE, addr != end);
> +		return 1;
> +	}
> +
> +	refs = 0;
>  	do {
>  		VM_BUG_ON_PAGE(compound_head(page) != head, page);
>  		pages[*nr] = page;

Ouch!

Looks like there are two main variants - so these kinds of repetitive patterns 
very much call for some sort of factoring out of common code, right?

Then the fix could be applied to the common portion(s) only, which will cut down 
this gigantic diffstat:

  >  5 files changed, 82 insertions(+), 5 deletions(-)

Thanks,

	Ingo

Back to linux.kernel | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

[PATCH 00/31] huge tmpfs: THPagecache implemented by teams Hugh Dickins <hughd@google.com> - 2016-04-05 23:20 +0200
  [PATCH 03/31] huge tmpfs: huge=N mount option and  /proc/sys/vm/shmem_huge Hugh Dickins <hughd@google.com> - 2016-04-05 23:20 +0200
    Re: [PATCH 03/31] huge tmpfs: huge=N mount option and  /proc/sys/vm/shmem_huge "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-04-11 13:20 +0200
  [PATCH 02/31] huge tmpfs: include shmem freeholes in available  memory Hugh Dickins <hughd@google.com> - 2016-04-05 23:20 +0200
  [PATCH 08/31] huge tmpfs: try_to_unmap_one use  page_check_address_transhuge Hugh Dickins <hughd@google.com> - 2016-04-05 23:30 +0200
  [PATCH 11/31] huge tmpfs: disband split huge pmds on race or memory  failure Hugh Dickins <hughd@google.com> - 2016-04-05 23:30 +0200
  [PATCH 06/31] huge tmpfs: shrinker to migrate and free underused  holes Hugh Dickins <hughd@google.com> - 2016-04-05 23:30 +0200
  [PATCH 10/31] huge tmpfs: map shmem by huge page pmd or by page team  ptes Hugh Dickins <hughd@google.com> - 2016-04-05 23:30 +0200
  [PATCH 07/31] huge tmpfs: get_unmapped_area align & fault supply  huge page Hugh Dickins <hughd@google.com> - 2016-04-05 23:30 +0200
  [PATCH 09/31] huge tmpfs: avoid premature exposure of new  pagetable Hugh Dickins <hughd@google.com> - 2016-04-05 23:30 +0200
    Re: [PATCH 09/31] huge tmpfs: avoid premature exposure of new  pagetable "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-04-11 14:00 +0200
  [PATCH 13/31] huge tmpfs: use Unevictable lru with variable  hpage_nr_pages Hugh Dickins <hughd@google.com> - 2016-04-05 23:40 +0200
  [PATCH 12/31] huge tmpfs: extend get_user_pages_fast to shmem pmd Hugh Dickins <hughd@google.com> - 2016-04-05 23:40 +0200
    Re: [PATCH 12/31] huge tmpfs: extend get_user_pages_fast to shmem pmd Ingo Molnar <mingo@kernel.org> - 2016-04-06 09:10 +0200
      Re: [PATCH 12/31] huge tmpfs: extend get_user_pages_fast to shmem  pmd Hugh Dickins <hughd@google.com> - 2016-04-07 05:00 +0200
        Re: [PATCH 12/31] huge tmpfs: extend get_user_pages_fast to shmem pmd Ingo Molnar <mingo@kernel.org> - 2016-04-13 11:00 +0200
  [PATCH 16/31] kvm: plumb return of hva when resolving page fault. Hugh Dickins <hughd@google.com> - 2016-04-05 23:40 +0200
  [PATCH 14/31] huge tmpfs: fix Mlocked meminfo, track huge & unhuge  mlocks Hugh Dickins <hughd@google.com> - 2016-04-05 23:40 +0200
  [PATCH 15/31] huge tmpfs: fix Mapped meminfo, track huge & unhuge  mappings Hugh Dickins <hughd@google.com> - 2016-04-05 23:40 +0200
  [PATCH 18/31] huge tmpfs: mem_cgroup move charge on shmem huge  pages Hugh Dickins <hughd@google.com> - 2016-04-05 23:50 +0200
  [PATCH 17/31] kvm: teach kvm to map page teams as huge pages. Hugh Dickins <hughd@google.com> - 2016-04-05 23:50 +0200
    Re: [PATCH 17/31] kvm: teach kvm to map page teams as huge pages. Paolo Bonzini <pbonzini@redhat.com> - 2016-04-06 01:40 +0200
      Re: [PATCH 17/31] kvm: teach kvm to map page teams as huge pages. Hugh Dickins <hughd@google.com> - 2016-04-06 03:20 +0200
        Re: [PATCH 17/31] kvm: teach kvm to map page teams as huge pages. Paolo Bonzini <pbonzini@redhat.com> - 2016-04-06 08:50 +0200
  [PATCH 19/31] huge tmpfs: mem_cgroup shmem_pmdmapped accounting Hugh Dickins <hughd@google.com> - 2016-04-05 23:50 +0200
  [PATCH 20/31] huge tmpfs: mem_cgroup shmem_hugepages accounting Hugh Dickins <hughd@google.com> - 2016-04-05 23:50 +0200
  [PATCH 22/31] huge tmpfs: /proc/<pid>/smaps show ShmemHugePages Hugh Dickins <hughd@google.com> - 2016-04-06 00:00 +0200
  [PATCH 24/31] huge tmpfs recovery: shmem_recovery_populate to fill  huge page Hugh Dickins <hughd@google.com> - 2016-04-06 00:00 +0200
  [PATCH 21/31] huge tmpfs: show page team flag in pageflags Hugh Dickins <hughd@google.com> - 2016-04-06 00:00 +0200
  [PATCH 25/31] huge tmpfs recovery: shmem_recovery_remap &  remap_team_by_pmd Hugh Dickins <hughd@google.com> - 2016-04-06 00:00 +0200
  [PATCH 26/31] huge tmpfs recovery: shmem_recovery_swapin to read  from swap Hugh Dickins <hughd@google.com> - 2016-04-06 00:00 +0200
  [PATCH 23/31] huge tmpfs recovery: framework for reconstituting huge  pages Hugh Dickins <hughd@google.com> - 2016-04-06 00:00 +0200
    Re: [PATCH 23/31] huge tmpfs recovery: framework for reconstituting  huge pages Mika Penttilä <mika.penttila@nextfour.com> - 2016-04-06 12:30 +0200
      Re: [PATCH 23/31] huge tmpfs recovery: framework for reconstituting  huge pages Hugh Dickins <hughd@google.com> - 2016-04-07 04:10 +0200
  [PATCH 31/31] huge tmpfs: no kswapd by default on sync allocations Hugh Dickins <hughd@google.com> - 2016-04-06 00:10 +0200
  [PATCH 27/31] huge tmpfs recovery: tweak shmem_getpage_gfp to fill  team Hugh Dickins <hughd@google.com> - 2016-04-06 00:10 +0200
  [PATCH 28/31] huge tmpfs recovery: debugfs stats to complete this  phase Hugh Dickins <hughd@google.com> - 2016-04-06 00:10 +0200
  [PATCH 30/31] huge tmpfs: shmem_huge_gfpmask and  shmem_recovery_gfpmask Hugh Dickins <hughd@google.com> - 2016-04-06 00:10 +0200
  [PATCH 29/31] huge tmpfs recovery: page migration call back into  shmem Hugh Dickins <hughd@google.com> - 2016-04-06 00:10 +0200

csiph-web