Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1185305 > unrolled thread

[PATCH v1 0/4] hwpoison: fixes on v4.2-rc2

Started byNaoya Horiguchi <n-horiguchi@ah.jp.nec.com>
First post2015-07-16 03:50 +0200
Last post2015-07-16 03:50 +0200
Articles 3 — 1 participant

Back to article view | Back to linux.kernel


Contents

  [PATCH v1 0/4] hwpoison: fixes on v4.2-rc2 Naoya Horiguchi <n-horiguchi@ah.jp.nec.com> - 2015-07-16 03:50 +0200
    [PATCH v1 2/4] mm/memory-failure: fix race in counting  num_poisoned_pages Naoya Horiguchi <n-horiguchi@ah.jp.nec.com> - 2015-07-16 03:50 +0200
    [PATCH v1 1/4] mm/memory-failure: unlock_page before put_page Naoya Horiguchi <n-horiguchi@ah.jp.nec.com> - 2015-07-16 03:50 +0200

#1185305 — [PATCH v1 0/4] hwpoison: fixes on v4.2-rc2

FromNaoya Horiguchi <n-horiguchi@ah.jp.nec.com>
Date2015-07-16 03:50 +0200
Subject[PATCH v1 0/4] hwpoison: fixes on v4.2-rc2
Message-ID<pMGtc-pn-27@gated-at.bofh.it>
Recently I addressed a few of hwpoison race problems and the patches are merged
on v4.2-rc1. It made progress, but unfortunately some problems still remain due
to less coverage of my testing. So I'm trying to fix or avoid them in this series.

One point I'm expecting to discuss is that patch 4/4 changes the page flag set
to be checked on free time. In current behavior, __PG_HWPOISON is not supposed
to be set when the page is freed. I think that there is no strong reason for this
behavior, and it causes a problem hard to fix only in error handler side (because
__PG_HWPOISON could be set at arbitrary timing.) So I suggest to change it.

With this patchset, the stress testing in official mce-test testsuite passes.

Thanks,
Naoya Horiguchi
---
Tree: https://github.com/Naoya-Horiguchi/linux/tree/v4.2-rc2/hwpoison.v1
---
Summary:

Naoya Horiguchi (4):
      mm/memory-failure: unlock_page before put_page
      mm/memory-failure: fix race in counting num_poisoned_pages
      mm/memory-failure: give up error handling for non-tail-refcounted thp
      mm/memory-failure: check __PG_HWPOISON separately from PAGE_FLAGS_CHECK_AT_*

 include/linux/page-flags.h | 10 +++++++---
 mm/huge_memory.c           |  7 +------
 mm/memory-failure.c        | 32 ++++++++++++++++++--------------
 mm/migrate.c               |  9 +++------
 mm/page_alloc.c            |  4 ++++
 5 files changed, 33 insertions(+), 29 deletions(-)--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1185306 — [PATCH v1 2/4] mm/memory-failure: fix race in counting num_poisoned_pages

FromNaoya Horiguchi <n-horiguchi@ah.jp.nec.com>
Date2015-07-16 03:50 +0200
Subject[PATCH v1 2/4] mm/memory-failure: fix race in counting num_poisoned_pages
Message-ID<pMGtd-pn-59@gated-at.bofh.it>
In reply to#1185305
When memory_failure() is called on a page which are just freed after page
migration from soft offlining, the counter num_poisoned_pages is raised twice.
So let's fix it with using TestSetPageHWPoison.

Signed-off-by: Naoya Horiguchi <n-horiguchi@ah.jp.nec.com>
---
 mm/memory-failure.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git v4.2-rc2.orig/mm/memory-failure.c v4.2-rc2/mm/memory-failure.c
index 04d677048af7..f72d2fad0b90 100644
--- v4.2-rc2.orig/mm/memory-failure.c
+++ v4.2-rc2/mm/memory-failure.c
@@ -1671,8 +1671,8 @@ static int __soft_offline_page(struct page *page, int flags)
 			if (ret > 0)
 				ret = -EIO;
 		} else {
-			SetPageHWPoison(page);
-			atomic_long_inc(&num_poisoned_pages);
+			if (!TestSetPageHWPoison(page))
+				atomic_long_inc(&num_poisoned_pages);
 		}
 	} else {
 		pr_info("soft offline: %#lx: isolation failed: %d, page count %d, type %lx\n",
-- 
2.4.3
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1185316 — [PATCH v1 1/4] mm/memory-failure: unlock_page before put_page

FromNaoya Horiguchi <n-horiguchi@ah.jp.nec.com>
Date2015-07-16 03:50 +0200
Subject[PATCH v1 1/4] mm/memory-failure: unlock_page before put_page
Message-ID<pMGte-pn-67@gated-at.bofh.it>
In reply to#1185305
In "just unpoisoned" path, we do put_page and then unlock_page, which is a
wrong order and causes "freeing locked page" bug. So let's fix it.

Signed-off-by: Naoya Horiguchi <n-horiguchi@ah.jp.nec.com>
---
 mm/memory-failure.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git v4.2-rc2.orig/mm/memory-failure.c v4.2-rc2/mm/memory-failure.c
index c53543d89282..04d677048af7 100644
--- v4.2-rc2.orig/mm/memory-failure.c
+++ v4.2-rc2/mm/memory-failure.c
@@ -1209,9 +1209,9 @@ int memory_failure(unsigned long pfn, int trapno, int flags)
 	if (!PageHWPoison(p)) {
 		printk(KERN_ERR "MCE %#lx: just unpoisoned\n", pfn);
 		atomic_long_sub(nr_pages, &num_poisoned_pages);
+		unlock_page(hpage);
 		put_page(hpage);
-		res = 0;
-		goto out;
+		return 0;
 	}
 	if (hwpoison_filter(p)) {
 		if (TestClearPageHWPoison(p))
-- 
2.4.3
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web