Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1742908 > unrolled thread
| Started by | Alexandru Moise <00moses.alexander00@gmail.com> |
|---|---|
| First post | 2017-10-02 01:00 +0200 |
| Last post | 2017-10-02 22:50 +0200 |
| Articles | 3 — 2 participants |
Back to article view | Back to linux.kernel
[PATCH] mm,hugetlb,migration: don't migrate kernelcore hugepages Alexandru Moise <00moses.alexander00@gmail.com> - 2017-10-02 01:00 +0200
Re: [PATCH] mm,hugetlb,migration: don't migrate kernelcore hugepages Michal Hocko <mhocko@kernel.org> - 2017-10-02 15:00 +0200
Re: [PATCH] mm,hugetlb,migration: don't migrate kernelcore hugepages Michal Hocko <mhocko@kernel.org> - 2017-10-02 22:50 +0200
| From | Alexandru Moise <00moses.alexander00@gmail.com> |
|---|---|
| Date | 2017-10-02 01:00 +0200 |
| Subject | [PATCH] mm,hugetlb,migration: don't migrate kernelcore hugepages |
| Message-ID | <uvVnk-7UN-9@gated-at.bofh.it> |
This attempts to bring more flexibility to how hugepages are allocated
by making it possible to decide whether we want the hugepages to be
allocated from ZONE_MOVABLE or to the zone allocated by the "kernelcore="
boot parameter for non-movable allocations.
A new boot parameter is introduced, "hugepages_movable=", this sets the
default value for the "hugepages_treat_as_movable" sysctl. This allows
us to determine the zone for hugepages allocated at boot time. It only
affects 2M hugepages allocated at boot time for now because 1G
hugepages are allocated much earlier in the boot process and ignore
this sysctl completely.
The "hugepages_treat_as_movable" sysctl is also turned into a mandatory
setting that all hugepage allocations at runtime must respect (both
2M and 1G sized hugepages). The default value is changed to "1" to
preserve the existing behavior that if hugepage migration is supported,
then the pages will be allocated from ZONE_MOVABLE.
Note however if not enough contiguous memory is present in ZONE_MOVABLE
then the allocation will fallback to the non-movable zone and those
pages will not be migratable.
The implementation is a bit dirty so obviously I'm open to suggestions
for a better way to implement this behavior, or comments whether the whole
idea is fundamentally __wrong__.
Signed-off-by: Alexandru Moise <00moses.alexander00@gmail.com>
---
Documentation/admin-guide/kernel-parameters.txt | 8 ++++++++
Documentation/sysctl/vm.txt | 3 +++
mm/hugetlb.c | 15 +++++++++++++--
mm/migrate.c | 8 +++++++-
4 files changed, 31 insertions(+), 3 deletions(-)
diff --git a/Documentation/admin-guide/kernel-parameters.txt b/Documentation/admin-guide/kernel-parameters.txt
index 05496622b4ef..25116d32d59e 100644
--- a/Documentation/admin-guide/kernel-parameters.txt
+++ b/Documentation/admin-guide/kernel-parameters.txt
@@ -1318,6 +1318,14 @@
x86-64 are 2M (when the CPU supports "pse") and 1G
(when the CPU supports the "pdpe1gb" cpuinfo flag).
+ hugepages_movable=
+ [HW,IA-64,PPC,X86-64] Default value for the
+ hugepages_treat_as_movable sysctl (default is 1).
+ When 1 this will attempt to allocate hugepages from
+ ZONE_MOVABLE, if 0 it will attempt to allocate hugepages
+ from the non-movable zone created with the "kernelcore="
+ kernel parameter.
+
hvc_iucv= [S390] Number of z/VM IUCV hypervisor console (HVC)
terminal devices. Valid values: 0..8
hvc_iucv_allow= [S390] Comma-separated list of z/VM user IDs.
diff --git a/Documentation/sysctl/vm.txt b/Documentation/sysctl/vm.txt
index 9baf66a9ef4e..4c5755a1cf9f 100644
--- a/Documentation/sysctl/vm.txt
+++ b/Documentation/sysctl/vm.txt
@@ -267,6 +267,9 @@ or not. If set to non-zero, hugepages can be allocated from ZONE_MOVABLE.
ZONE_MOVABLE is created when kernel boot parameter kernelcore= is specified,
so this parameter has no effect if used without kernelcore=.
+The default value for this sysctl can also be set via the hugepages_movable=
+kernel boot parameter (to 0 or 1), default is 1.
+
Hugepage migration is now available in some situations which depend on the
architecture and/or the hugepage size. If a hugepage supports migration,
allocation from ZONE_MOVABLE is always enabled for the hugepage regardless
diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index 424b0ef08a60..5d4efdadbd56 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -36,7 +36,7 @@
#include <linux/userfaultfd_k.h>
#include "internal.h"
-int hugepages_treat_as_movable;
+int hugepages_treat_as_movable = 1;
int hugetlb_max_hstate __read_mostly;
unsigned int default_hstate_idx;
@@ -926,7 +926,7 @@ static struct page *dequeue_huge_page_nodemask(struct hstate *h, gfp_t gfp_mask,
/* Movability of hugepages depends on migration support. */
static inline gfp_t htlb_alloc_mask(struct hstate *h)
{
- if (hugepages_treat_as_movable || hugepage_migration_supported(h))
+ if (hugepages_treat_as_movable && hugepage_migration_supported(h))
return GFP_HIGHUSER_MOVABLE;
else
return GFP_HIGHUSER;
@@ -2805,6 +2805,17 @@ static int __init hugetlb_init(void)
}
subsys_initcall(hugetlb_init);
+static int __init hugepages_movable(char *str)
+{
+ if (!strncmp(str, "0", 1))
+ hugepages_treat_as_movable = 0;
+ else if (!strncmp(str, "1", 1))
+ hugepages_treat_as_movable = 1;
+
+ return 1;
+}
+__setup("hugepages_movable=", hugepages_movable);
+
/* Should be called on processing a hugepagesz=... option */
void __init hugetlb_bad_size(void)
{
diff --git a/mm/migrate.c b/mm/migrate.c
index 6954c1435833..23946d88e533 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -1266,6 +1266,7 @@ static int unmap_and_move_huge_page(new_page_t get_new_page,
int page_was_mapped = 0;
struct page *new_hpage;
struct anon_vma *anon_vma = NULL;
+ bool zone_movable_present;
/*
* Movability of hugepages depends on architectures and hugepage size.
@@ -1274,7 +1275,12 @@ static int unmap_and_move_huge_page(new_page_t get_new_page,
* tables or check whether the hugepage is pmd-based or not before
* kicking migration.
*/
- if (!hugepage_migration_supported(page_hstate(hpage))) {
+ zone_movable_present = (NODE_DATA(page_to_nid(hpage))->node_zones[ZONE_MOVABLE].spanned_pages > 0);
+
+ if (!hugepage_migration_supported(page_hstate(hpage)) ||
+ zone_movable_present ?
+ !(zone_idx(page_zone(hpage)) == ZONE_MOVABLE) :
+ false) {
putback_active_hugepage(hpage);
return -ENOSYS;
}
--
2.14.2
[toc] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-10-02 15:00 +0200 |
| Message-ID | <uw8ud-7od-7@gated-at.bofh.it> |
| In reply to | #1742908 |
On Mon 02-10-17 00:51:11, Alexandru Moise wrote: > This attempts to bring more flexibility to how hugepages are allocated > by making it possible to decide whether we want the hugepages to be > allocated from ZONE_MOVABLE or to the zone allocated by the "kernelcore=" > boot parameter for non-movable allocations. > > A new boot parameter is introduced, "hugepages_movable=", this sets the > default value for the "hugepages_treat_as_movable" sysctl. This allows > us to determine the zone for hugepages allocated at boot time. It only > affects 2M hugepages allocated at boot time for now because 1G > hugepages are allocated much earlier in the boot process and ignore > this sysctl completely. > > The "hugepages_treat_as_movable" sysctl is also turned into a mandatory > setting that all hugepage allocations at runtime must respect (both > 2M and 1G sized hugepages). The default value is changed to "1" to > preserve the existing behavior that if hugepage migration is supported, > then the pages will be allocated from ZONE_MOVABLE. > > Note however if not enough contiguous memory is present in ZONE_MOVABLE > then the allocation will fallback to the non-movable zone and those > pages will not be migratable. This changelog doesn't explain _why_ we would need something like that. > The implementation is a bit dirty so obviously I'm open to suggestions > for a better way to implement this behavior, or comments whether the whole > idea is fundamentally __wrong__. To be honest I think this is just a wrong approach. hugepages_treat_as_movable is quite questionable to be honest because it breaks the basic semantic of the movable zone if the hugetlb pages are not really migratable which should be the only criterion. Hugetlb pages are no different from other migratable pages in that regards. -- Michal Hocko SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-10-02 22:50 +0200 |
| Message-ID | <uwfPh-ur-239@gated-at.bofh.it> |
| In reply to | #1743197 |
On Mon 02-10-17 17:06:38, Alexandru Moise wrote: > On Mon, Oct 02, 2017 at 04:27:17PM +0200, Michal Hocko wrote: > > On Mon 02-10-17 16:06:33, Alexandru Moise wrote: > > > On Mon, Oct 02, 2017 at 02:54:32PM +0200, Michal Hocko wrote: > > > > On Mon 02-10-17 00:51:11, Alexandru Moise wrote: > > > > > This attempts to bring more flexibility to how hugepages are allocated > > > > > by making it possible to decide whether we want the hugepages to be > > > > > allocated from ZONE_MOVABLE or to the zone allocated by the "kernelcore=" > > > > > boot parameter for non-movable allocations. > > > > > > > > > > A new boot parameter is introduced, "hugepages_movable=", this sets the > > > > > default value for the "hugepages_treat_as_movable" sysctl. This allows > > > > > us to determine the zone for hugepages allocated at boot time. It only > > > > > affects 2M hugepages allocated at boot time for now because 1G > > > > > hugepages are allocated much earlier in the boot process and ignore > > > > > this sysctl completely. > > > > > > > > > > The "hugepages_treat_as_movable" sysctl is also turned into a mandatory > > > > > setting that all hugepage allocations at runtime must respect (both > > > > > 2M and 1G sized hugepages). The default value is changed to "1" to > > > > > preserve the existing behavior that if hugepage migration is supported, > > > > > then the pages will be allocated from ZONE_MOVABLE. > > > > > > > > > > Note however if not enough contiguous memory is present in ZONE_MOVABLE > > > > > then the allocation will fallback to the non-movable zone and those > > > > > pages will not be migratable. > > > > > > > > This changelog doesn't explain _why_ we would need something like that. > > > > > > > > > > So people shouldn't be able to choose whether their hugepages should be > > > migratable or not? > > > > How are hugetlb pages any different from THP wrt. migrateability POV? Or > > any other mapped memory to the userspace in general? > > THP shares more with regular userspace mapped memory than with hugetlbfs pages. > They have separate codepaths in migrate_pages(). That is a mere implementation detail. You are right that THP shares more with regular userspace memory because it is transparent from the configuration POV but that has nothing to do with page migration AFAICS. > And no one ever sets the movable > flag on a hugetlbfs mapping, so even though __PageMovable(hpage) on a hugetlbfs > page returns false, it will still move. __PageMovable is a completely unrelated thing. It is for pages which are !LRU but still movable. > > > > > > Maybe they consider some of their applications more important than > > > others. > > > > I do not understand this part. > > > > > Say: > > > You have a large number of correctable errors on a subpage of a compound > > > page. So you copy the contents of the page to another hugepage, break the > > > original page and offline the subpage. > > > > I suspect you have HWPoisoning in mind right? > > No, rather soft offlining. I thought this is the same thing. > > > But maybe you'd rather that some of > > > your hugepages not be broken and moved because you're not that worried about > > > memory corruption, but more about availability. > > > > Could you be more specific please? > > You can have a platform with reliable DIMM modules and a platform with less reliable > DIMM modules. So you would prefer to inhibit hugepage migration on the platform with > reliable DIMM modules that you know will behave ok even under a high number of > correctable memory errors. tools like mcelog however are not hugepage aware and > cannot be told "if this PFN is part of a hugepage, don't try to soft offline it", > rather deciding which PFNs should be unmovable should be done in the kernel, > but it should still be controllable by the administrator. This sounds like a userspace policy that should be handled outside of the kernel. > For hugetlbfs pages in particular, this behavior is not present, without this patch. > > > > > > Without this patch even if hugepages are in the non-movable zone, they move. > > > > which is ok. This is very same with any other movable allocations. > > So you can have movable pages in the non-movable kernel zone? yes. Most configuration even do not have any movable zone unless explicitly configured. > > > > > The implementation is a bit dirty so obviously I'm open to suggestions > > > > > for a better way to implement this behavior, or comments whether the whole > > > > > idea is fundamentally __wrong__. > > > > > > > > To be honest I think this is just a wrong approach. hugepages_treat_as_movable > > > > is quite questionable to be honest because it breaks the basic semantic > > > > of the movable zone if the hugetlb pages are not really migratable which > > > > should be the only criterion. Hugetlb pages are no different from other > > > > migratable pages in that regards. > > > > > > Shouldn't hugepages allocated to unmovable zone, by definition, not be able > > > to be migrated? With this patch, hugepages in the movable zone do move, but > > > hugepages in the non-movable zone don't. Or am I misunderstanding the semantics > > > completely? > > > > yes. movable zone is only about a guarantee to move memory around. > > Movable allocations are still allowed to use kernel zones (aka > > non-movable). The main reason for the movable zone these days is memory > > hotplug which needs a semi-guarantee that the memory used can be > > migrated elsewhere to free up the offlined memory. > > But isn't kernel-zone memory guaranteed not to migrate? No. > I agree that movable allocations are allowed to fallback to kernel zones. > i.e. This is behavior is correct: > Page A is in ZONE_MOVABLE, page B is in kernel zone. > Page A gets soft-offlined, the contents are moved to page B. > > This behavior is not correct: > Page C is in kernel zone, page D is also in kernel zone. > Page C gets soft offlined, contents of page C get moved to page D. Why is this incorrect? > With hugepages, there is no check for whereto the migration goes because > the pages are pre-allocated and simply dequeued from the hstate freelist. true > Thus hugepages will end up being unreserved and moved to a different > reserved hugepage, and the administrator has no control over this behavior, > even if they're kernel zone pages. I really fail to see why kernel vs. movable zones play any role here. Zones should be mostly an implementation detail which userspace shouldn't really care about. -- Michal Hocko SUSE Labs
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web