Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1330413 > unrolled thread

[PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables

Started by"Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com>
First post2016-02-09 17:20 +0100
Last post2016-02-10 03:40 +0100
Articles 6 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-02-09 17:20 +0100
    Re: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related  values as variables "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-02-09 20:50 +0100
      Re: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-02-10 03:40 +0100
    Re: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related  values as variables Andrew Morton <akpm@linux-foundation.org> - 2016-02-09 22:30 +0100
      Re: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related  values as variables "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-02-10 00:20 +0100
      Re: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-02-10 03:40 +0100

#1330413 — [PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables

From"Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com>
Date2016-02-09 17:20 +0100
Subject[PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables
Message-ID<r0jbc-6vm-15@gated-at.bofh.it>
With next generation power processor, we are having a new mmu model
[1] that require us to maintain a different linux page table format.

Inorder to support both current and future ppc64 systems with a single
kernel we need to make sure kernel can select between different page
table format at runtime. With the new MMU (radix MMU) added, we will
have two different pmd hugepage size 16MB for hash model and 2MB for
Radix model. Hence make HPAGE_PMD related values as a variable.

[1] http://ibm.biz/power-isa3 (Needs registration).

Signed-off-by: Aneesh Kumar K.V <aneesh.kumar@linux.vnet.ibm.com>
---
 arch/powerpc/mm/pgtable_64.c |  7 +++++++
 include/linux/bug.h          |  9 +++++++++
 include/linux/huge_mm.h      |  3 ---
 mm/huge_memory.c             | 17 ++++++++++++++---
 4 files changed, 30 insertions(+), 6 deletions(-)

diff --git a/arch/powerpc/mm/pgtable_64.c b/arch/powerpc/mm/pgtable_64.c
index c8a00da39969..80dd8e0d8322 100644
--- a/arch/powerpc/mm/pgtable_64.c
+++ b/arch/powerpc/mm/pgtable_64.c
@@ -818,6 +818,13 @@ pmd_t pmdp_huge_get_and_clear(struct mm_struct *mm,
 
 int has_transparent_hugepage(void)
 {
+
+	BUILD_BUG_ON_MSG((PMD_SHIFT - PAGE_SHIFT) >= MAX_ORDER,
+		"hugepages can't be allocated by the buddy allocator");
+
+	BUILD_BUG_ON_MSG((PMD_SHIFT - PAGE_SHIFT) < 2,
+			 "We need more than 2 pages to do deferred thp split");
+
 	if (!mmu_has_feature(MMU_FTR_16M_PAGE))
 		return 0;
 	/*
diff --git a/include/linux/bug.h b/include/linux/bug.h
index 7f4818673c41..e51b0709e78d 100644
--- a/include/linux/bug.h
+++ b/include/linux/bug.h
@@ -20,6 +20,7 @@ struct pt_regs;
 #define BUILD_BUG_ON_MSG(cond, msg) (0)
 #define BUILD_BUG_ON(condition) (0)
 #define BUILD_BUG() (0)
+#define MAYBE_BUILD_BUG_ON(cond) (0)
 #else /* __CHECKER__ */
 
 /* Force a compilation error if a constant expression is not a power of 2 */
@@ -83,6 +84,14 @@ struct pt_regs;
  */
 #define BUILD_BUG() BUILD_BUG_ON_MSG(1, "BUILD_BUG failed")
 
+#define MAYBE_BUILD_BUG_ON(cond)			\
+	do {						\
+		if (__builtin_constant_p((cond)))       \
+			BUILD_BUG_ON(cond);             \
+		else                                    \
+			BUG_ON(cond);                   \
+	} while (0)
+
 #endif	/* __CHECKER__ */
 
 #ifdef CONFIG_GENERIC_BUG
diff --git a/include/linux/huge_mm.h b/include/linux/huge_mm.h
index 459fd25b378e..f12513a20a06 100644
--- a/include/linux/huge_mm.h
+++ b/include/linux/huge_mm.h
@@ -111,9 +111,6 @@ void __split_huge_pmd(struct vm_area_struct *vma, pmd_t *pmd,
 			__split_huge_pmd(__vma, __pmd, __address);	\
 	}  while (0)
 
-#if HPAGE_PMD_ORDER >= MAX_ORDER
-#error "hugepages can't be allocated by the buddy allocator"
-#endif
 extern int hugepage_madvise(struct vm_area_struct *vma,
 			    unsigned long *vm_flags, int advice);
 extern void vma_adjust_trans_huge(struct vm_area_struct *vma,
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index cd26f3f14cab..350410e9019e 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -83,7 +83,7 @@ unsigned long transparent_hugepage_flags __read_mostly =
 	(1<<TRANSPARENT_HUGEPAGE_USE_ZERO_PAGE_FLAG);
 
 /* default scan 8*512 pte (or vmas) every 30 second */
-static unsigned int khugepaged_pages_to_scan __read_mostly = HPAGE_PMD_NR*8;
+static unsigned int khugepaged_pages_to_scan __read_mostly;
 static unsigned int khugepaged_pages_collapsed;
 static unsigned int khugepaged_full_scans;
 static unsigned int khugepaged_scan_sleep_millisecs __read_mostly = 10000;
@@ -98,7 +98,7 @@ static DECLARE_WAIT_QUEUE_HEAD(khugepaged_wait);
  * it would have happened if the vma was large enough during page
  * fault.
  */
-static unsigned int khugepaged_max_ptes_none __read_mostly = HPAGE_PMD_NR-1;
+static unsigned int khugepaged_max_ptes_none __read_mostly;
 
 static int khugepaged(void *none);
 static int khugepaged_slab_init(void);
@@ -660,6 +660,18 @@ static int __init hugepage_init(void)
 		return -EINVAL;
 	}
 
+	khugepaged_pages_to_scan = HPAGE_PMD_NR * 8;
+	khugepaged_max_ptes_none = HPAGE_PMD_NR - 1;
+	/*
+	 * hugepages can't be allocated by the buddy allocator
+	 */
+	MAYBE_BUILD_BUG_ON(HPAGE_PMD_ORDER >= MAX_ORDER);
+	/*
+	 * we use page->mapping and page->index in second tail page
+	 * as list_head: assuming THP order >= 2
+	 */
+	MAYBE_BUILD_BUG_ON(HPAGE_PMD_ORDER < 2);
+
 	err = hugepage_init_sysfs(&hugepage_kobj);
 	if (err)
 		goto err_sysfs;
@@ -764,7 +776,6 @@ void prep_transhuge_page(struct page *page)
 	 * we use page->mapping and page->indexlru in second tail page
 	 * as list_head: assuming THP order >= 2
 	 */
-	BUILD_BUG_ON(HPAGE_PMD_ORDER < 2);
 
 	INIT_LIST_HEAD(page_deferred_list(page));
 	set_compound_page_dtor(page, TRANSHUGE_PAGE_DTOR);
-- 
2.5.0

[toc] | [next] | [standalone]


#1330611 — Re: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables

From"Kirill A. Shutemov" <kirill@shutemov.name>
Date2016-02-09 20:50 +0100
SubjectRe: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables
Message-ID<r0msq-5p-13@gated-at.bofh.it>
In reply to#1330413
On Tue, Feb 09, 2016 at 09:41:44PM +0530, Aneesh Kumar K.V wrote:
> With next generation power processor, we are having a new mmu model
> [1] that require us to maintain a different linux page table format.
> 
> Inorder to support both current and future ppc64 systems with a single
> kernel we need to make sure kernel can select between different page
> table format at runtime. With the new MMU (radix MMU) added, we will
> have two different pmd hugepage size 16MB for hash model and 2MB for
> Radix model. Hence make HPAGE_PMD related values as a variable.
> 
> [1] http://ibm.biz/power-isa3 (Needs registration).
> 
> Signed-off-by: Aneesh Kumar K.V <aneesh.kumar@linux.vnet.ibm.com>

I guess it should have my signed-off-by ;)

Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>

-- 
 Kirill A. Shutemov

[toc] | [prev] | [next] | [standalone]


#1330852

From"Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com>
Date2016-02-10 03:40 +0100
Message-ID<r0sRc-4oH-3@gated-at.bofh.it>
In reply to#1330611
"Kirill A. Shutemov" <kirill@shutemov.name> writes:

> On Tue, Feb 09, 2016 at 09:41:44PM +0530, Aneesh Kumar K.V wrote:
>> With next generation power processor, we are having a new mmu model
>> [1] that require us to maintain a different linux page table format.
>> 
>> Inorder to support both current and future ppc64 systems with a single
>> kernel we need to make sure kernel can select between different page
>> table format at runtime. With the new MMU (radix MMU) added, we will
>> have two different pmd hugepage size 16MB for hash model and 2MB for
>> Radix model. Hence make HPAGE_PMD related values as a variable.
>> 
>> [1] http://ibm.biz/power-isa3 (Needs registration).
>> 
>> Signed-off-by: Aneesh Kumar K.V <aneesh.kumar@linux.vnet.ibm.com>
>
> I guess it should have my signed-off-by ;)
>
> Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>

Thanks will update. I will also update the From:

-aneesh

[toc] | [prev] | [next] | [standalone]


#1330661 — Re: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables

FromAndrew Morton <akpm@linux-foundation.org>
Date2016-02-09 22:30 +0100
SubjectRe: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables
Message-ID<r0o1c-1jK-11@gated-at.bofh.it>
In reply to#1330413
On Tue,  9 Feb 2016 21:41:44 +0530 "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> wrote:

> With next generation power processor, we are having a new mmu model
> [1] that require us to maintain a different linux page table format.
> 
> Inorder to support both current and future ppc64 systems with a single
> kernel we need to make sure kernel can select between different page
> table format at runtime. With the new MMU (radix MMU) added, we will
> have two different pmd hugepage size 16MB for hash model and 2MB for
> Radix model. Hence make HPAGE_PMD related values as a variable.
> 
> [1] http://ibm.biz/power-isa3 (Needs registration).
> 
> ...
>
> --- a/include/linux/bug.h
> +++ b/include/linux/bug.h
> @@ -20,6 +20,7 @@ struct pt_regs;
>  #define BUILD_BUG_ON_MSG(cond, msg) (0)
>  #define BUILD_BUG_ON(condition) (0)
>  #define BUILD_BUG() (0)
> +#define MAYBE_BUILD_BUG_ON(cond) (0)
>  #else /* __CHECKER__ */
>  
>  /* Force a compilation error if a constant expression is not a power of 2 */
> @@ -83,6 +84,14 @@ struct pt_regs;
>   */
>  #define BUILD_BUG() BUILD_BUG_ON_MSG(1, "BUILD_BUG failed")
>  
> +#define MAYBE_BUILD_BUG_ON(cond)			\
> +	do {						\
> +		if (__builtin_constant_p((cond)))       \
> +			BUILD_BUG_ON(cond);             \
> +		else                                    \
> +			BUG_ON(cond);                   \
> +	} while (0)
> +

hm.  I suppose so.

> --- a/include/linux/huge_mm.h
> +++ b/include/linux/huge_mm.h
> @@ -111,9 +111,6 @@ void __split_huge_pmd(struct vm_area_struct *vma, pmd_t *pmd,
>  			__split_huge_pmd(__vma, __pmd, __address);	\
>  	}  while (0)
>  
> -#if HPAGE_PMD_ORDER >= MAX_ORDER
> -#error "hugepages can't be allocated by the buddy allocator"
> -#endif
>  extern int hugepage_madvise(struct vm_area_struct *vma,
>  			    unsigned long *vm_flags, int advice);
>  extern void vma_adjust_trans_huge(struct vm_area_struct *vma,
> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> index cd26f3f14cab..350410e9019e 100644
> --- a/mm/huge_memory.c
> +++ b/mm/huge_memory.c
> @@ -83,7 +83,7 @@ unsigned long transparent_hugepage_flags __read_mostly =
>  	(1<<TRANSPARENT_HUGEPAGE_USE_ZERO_PAGE_FLAG);
>  
>  /* default scan 8*512 pte (or vmas) every 30 second */
> -static unsigned int khugepaged_pages_to_scan __read_mostly = HPAGE_PMD_NR*8;
> +static unsigned int khugepaged_pages_to_scan __read_mostly;
>  static unsigned int khugepaged_pages_collapsed;
>  static unsigned int khugepaged_full_scans;
>  static unsigned int khugepaged_scan_sleep_millisecs __read_mostly = 10000;
> @@ -98,7 +98,7 @@ static DECLARE_WAIT_QUEUE_HEAD(khugepaged_wait);
>   * it would have happened if the vma was large enough during page
>   * fault.
>   */
> -static unsigned int khugepaged_max_ptes_none __read_mostly = HPAGE_PMD_NR-1;
> +static unsigned int khugepaged_max_ptes_none __read_mostly;
>  
>  static int khugepaged(void *none);
>  static int khugepaged_slab_init(void);
> @@ -660,6 +660,18 @@ static int __init hugepage_init(void)
>  		return -EINVAL;
>  	}
>  
> +	khugepaged_pages_to_scan = HPAGE_PMD_NR * 8;
> +	khugepaged_max_ptes_none = HPAGE_PMD_NR - 1;

I don't understand this change.  We change the initialization from
at-compile-time to at-run-time, but nothing useful appears to have been
done.

> +	/*
> +	 * hugepages can't be allocated by the buddy allocator
> +	 */
> +	MAYBE_BUILD_BUG_ON(HPAGE_PMD_ORDER >= MAX_ORDER);
> +	/*
> +	 * we use page->mapping and page->index in second tail page
> +	 * as list_head: assuming THP order >= 2
> +	 */
> +	MAYBE_BUILD_BUG_ON(HPAGE_PMD_ORDER < 2);
> +
>  	err = hugepage_init_sysfs(&hugepage_kobj);
>  	if (err)
>  		goto err_sysfs;
>
> ...
>

[toc] | [prev] | [next] | [standalone]


#1330766 — Re: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables

From"Kirill A. Shutemov" <kirill@shutemov.name>
Date2016-02-10 00:20 +0100
SubjectRe: [PATCH V2] mm: Some arch may want to use HPAGE_PMD related values as variables
Message-ID<r0pJD-2qT-15@gated-at.bofh.it>
In reply to#1330661
On Tue, Feb 09, 2016 at 01:26:08PM -0800, Andrew Morton wrote:
> On Tue,  9 Feb 2016 21:41:44 +0530 "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> wrote:
> > @@ -660,6 +660,18 @@ static int __init hugepage_init(void)
> >  		return -EINVAL;
> >  	}
> >  
> > +	khugepaged_pages_to_scan = HPAGE_PMD_NR * 8;
> > +	khugepaged_max_ptes_none = HPAGE_PMD_NR - 1;
> 
> I don't understand this change.  We change the initialization from
> at-compile-time to at-run-time, but nothing useful appears to have been
> done.

It's preparation patch. HPAGE_PMD_NR is going to be based on variable on
Power soon. Compile-time is not an option.

-- 
 Kirill A. Shutemov

[toc] | [prev] | [next] | [standalone]


#1330851

From"Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com>
Date2016-02-10 03:40 +0100
Message-ID<r0sRc-4oH-1@gated-at.bofh.it>
In reply to#1330661
Andrew Morton <akpm@linux-foundation.org> writes:

> On Tue,  9 Feb 2016 21:41:44 +0530 "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> wrote:
>
>> With next generation power processor, we are having a new mmu model
>> [1] that require us to maintain a different linux page table format.
>> 
>> Inorder to support both current and future ppc64 systems with a single
>> kernel we need to make sure kernel can select between different page
>> table format at runtime. With the new MMU (radix MMU) added, we will
>> have two different pmd hugepage size 16MB for hash model and 2MB for
>> Radix model. Hence make HPAGE_PMD related values as a variable.
>> 
>> [1] http://ibm.biz/power-isa3 (Needs registration).
>> 
>> ...
>>
>> --- a/include/linux/bug.h
>> +++ b/include/linux/bug.h
>> @@ -20,6 +20,7 @@ struct pt_regs;
>>  #define BUILD_BUG_ON_MSG(cond, msg) (0)
>>  #define BUILD_BUG_ON(condition) (0)
>>  #define BUILD_BUG() (0)
>> +#define MAYBE_BUILD_BUG_ON(cond) (0)
>>  #else /* __CHECKER__ */
>>  
>>  /* Force a compilation error if a constant expression is not a power of 2 */
>> @@ -83,6 +84,14 @@ struct pt_regs;
>>   */
>>  #define BUILD_BUG() BUILD_BUG_ON_MSG(1, "BUILD_BUG failed")
>>  
>> +#define MAYBE_BUILD_BUG_ON(cond)			\
>> +	do {						\
>> +		if (__builtin_constant_p((cond)))       \
>> +			BUILD_BUG_ON(cond);             \
>> +		else                                    \
>> +			BUG_ON(cond);                   \
>> +	} while (0)
>> +
>
> hm.  I suppose so.
>
>> --- a/include/linux/huge_mm.h
>> +++ b/include/linux/huge_mm.h
>> @@ -111,9 +111,6 @@ void __split_huge_pmd(struct vm_area_struct *vma, pmd_t *pmd,
>>  			__split_huge_pmd(__vma, __pmd, __address);	\
>>  	}  while (0)
>>  
>> -#if HPAGE_PMD_ORDER >= MAX_ORDER
>> -#error "hugepages can't be allocated by the buddy allocator"
>> -#endif
>>  extern int hugepage_madvise(struct vm_area_struct *vma,
>>  			    unsigned long *vm_flags, int advice);
>>  extern void vma_adjust_trans_huge(struct vm_area_struct *vma,
>> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
>> index cd26f3f14cab..350410e9019e 100644
>> --- a/mm/huge_memory.c
>> +++ b/mm/huge_memory.c
>> @@ -83,7 +83,7 @@ unsigned long transparent_hugepage_flags __read_mostly =
>>  	(1<<TRANSPARENT_HUGEPAGE_USE_ZERO_PAGE_FLAG);
>>  
>>  /* default scan 8*512 pte (or vmas) every 30 second */
>> -static unsigned int khugepaged_pages_to_scan __read_mostly = HPAGE_PMD_NR*8;
>> +static unsigned int khugepaged_pages_to_scan __read_mostly;
>>  static unsigned int khugepaged_pages_collapsed;
>>  static unsigned int khugepaged_full_scans;
>>  static unsigned int khugepaged_scan_sleep_millisecs __read_mostly = 10000;
>> @@ -98,7 +98,7 @@ static DECLARE_WAIT_QUEUE_HEAD(khugepaged_wait);
>>   * it would have happened if the vma was large enough during page
>>   * fault.
>>   */
>> -static unsigned int khugepaged_max_ptes_none __read_mostly = HPAGE_PMD_NR-1;
>> +static unsigned int khugepaged_max_ptes_none __read_mostly;
>>  
>>  static int khugepaged(void *none);
>>  static int khugepaged_slab_init(void);
>> @@ -660,6 +660,18 @@ static int __init hugepage_init(void)
>>  		return -EINVAL;
>>  	}
>>  
>> +	khugepaged_pages_to_scan = HPAGE_PMD_NR * 8;
>> +	khugepaged_max_ptes_none = HPAGE_PMD_NR - 1;
>
> I don't understand this change.  We change the initialization from
> at-compile-time to at-run-time, but nothing useful appears to have been
> done.
>

The related changes are in another series, 
https://lists.ozlabs.org/pipermail/linuxppc-dev/2016-February/thread.html#138948

I would also like to keep the two core mm patches[1] in that series so that
it can go with the other related changes via powerpc tree. The reason to
send out them as separate patches is to get the correct feedback so that
it won't get lost in the large series.

Let me know if that is ok with you.

[1] https://lists.ozlabs.org/pipermail/linuxppc-dev/2016-February/138955.html
    https://lists.ozlabs.org/pipermail/linuxppc-dev/2016-February/138964.html

-aneesh

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web