Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1561257 > unrolled thread

[RFC] HWPOISON: soft offlining for non-lru movable page

Started byYisheng Xie <xieyisheng1@huawei.com>
First post2017-01-18 05:10 +0100
Last post2017-01-19 02:30 +0100
Articles 6 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [RFC] HWPOISON: soft offlining for non-lru movable page Yisheng Xie <xieyisheng1@huawei.com> - 2017-01-18 05:10 +0100
    Re: [RFC] HWPOISON: soft offlining for non-lru movable page Naoya Horiguchi <n-horiguchi@ah.jp.nec.com> - 2017-01-18 11:00 +0100
      Re: [RFC] HWPOISON: soft offlining for non-lru movable page Yisheng Xie <xieyisheng1@huawei.com> - 2017-01-20 11:00 +0100
        Re: [RFC] HWPOISON: soft offlining for non-lru movable page Naoya Horiguchi <n-horiguchi@ah.jp.nec.com> - 2017-01-23 05:50 +0100
    Re: [RFC] HWPOISON: soft offlining for non-lru movable page Michal Hocko <mhocko@kernel.org> - 2017-01-18 11:10 +0100
      Re: [RFC] HWPOISON: soft offlining for non-lru movable page Yisheng Xie <xieyisheng1@huawei.com> - 2017-01-19 02:30 +0100

#1561257 — [RFC] HWPOISON: soft offlining for non-lru movable page

FromYisheng Xie <xieyisheng1@huawei.com>
Date2017-01-18 05:10 +0100
Subject[RFC] HWPOISON: soft offlining for non-lru movable page
Message-ID<t0PJn-7Ew-1@gated-at.bofh.it>
This patch is to extends soft offlining framework to support
non-lru page, which already support migration after
commit bda807d44454 ("mm: migrate: support non-lru movable page
migration")

When memory corrected errors occur on a non-lru movable page,
we can choose to stop using it by migrating data onto another
page and disable the original (maybe half-broken) one.

Signed-off-by: Yisheng Xie <xieyisheng1@huawei.com>
---
 mm/memory-failure.c | 55 +++++++++++++++++++++++++++++++++++++++++++++++++++--
 1 file changed, 53 insertions(+), 2 deletions(-)

diff --git a/mm/memory-failure.c b/mm/memory-failure.c
index f283c7e..10043a4 100644
--- a/mm/memory-failure.c
+++ b/mm/memory-failure.c
@@ -1527,7 +1527,8 @@ static int get_any_page(struct page *page, unsigned long pfn, int flags)
 {
 	int ret = __get_any_page(page, pfn, flags);
 
-	if (ret == 1 && !PageHuge(page) && !PageLRU(page)) {
+	if (ret == 1 && !PageHuge(page) &&
+	    !PageLRU(page) && !__PageMovable(page)) {
 		/*
 		 * Try to free it.
 		 */
@@ -1549,6 +1550,54 @@ static int get_any_page(struct page *page, unsigned long pfn, int flags)
 	return ret;
 }
 
+static int soft_offline_movable_page(struct page *page, int flags)
+{
+	int ret;
+	unsigned long pfn = page_to_pfn(page);
+	LIST_HEAD(pagelist);
+
+	/*
+	 * This double-check of PageHWPoison is to avoid the race with
+	 * memory_failure(). See also comment in __soft_offline_page().
+	 */
+	lock_page(page);
+	if (PageHWPoison(page)) {
+		unlock_page(page);
+		put_hwpoison_page(page);
+		pr_info("soft offline: %#lx movable page already poisoned\n",
+			pfn);
+		return -EBUSY;
+	}
+	unlock_page(page);
+
+	ret = isolate_movable_page(page, ISOLATE_UNEVICTABLE);
+	/*
+	 * get_any_page() and isolate_movable_page() takes a refcount each,
+	 * so need to drop one here.
+	 */
+	put_hwpoison_page(page);
+	if (!ret) {
+		pr_info("soft offline: %#lx movable page failed to isolate\n",
+			pfn);
+		return -EBUSY;
+	}
+
+	list_add(&page->lru, &pagelist);
+	ret = migrate_pages(&pagelist, new_page, NULL, MPOL_MF_MOVE_ALL,
+			    MIGRATE_SYNC, MR_MEMORY_FAILURE);
+	if (ret) {
+		if (!list_empty(&pagelist))
+			putback_movable_pages(&pagelist);
+
+		pr_info("soft offline: %#lx: migration failed %d, type %lx\n",
+			pfn, ret, page->flags);
+		if (ret > 0)
+			ret = -EIO;
+	}
+
+	return ret;
+}
+
 static int soft_offline_huge_page(struct page *page, int flags)
 {
 	int ret;
@@ -1705,8 +1754,10 @@ static int soft_offline_in_use_page(struct page *page, int flags)
 
 	if (PageHuge(page))
 		ret = soft_offline_huge_page(page, flags);
-	else
+	else if (PageLRU(page))
 		ret = __soft_offline_page(page, flags);
+	else
+		ret = soft_offline_movable_page(page, flags);
 
 	return ret;
 }
-- 
1.7.12.4

[toc] | [next] | [standalone]


#1561424

FromNaoya Horiguchi <n-horiguchi@ah.jp.nec.com>
Date2017-01-18 11:00 +0100
Message-ID<t0Vc6-2mj-23@gated-at.bofh.it>
In reply to#1561257
On Wed, Jan 18, 2017 at 12:00:54PM +0800, Yisheng Xie wrote:
> This patch is to extends soft offlining framework to support
> non-lru page, which already support migration after
> commit bda807d44454 ("mm: migrate: support non-lru movable page
> migration")
> 
> When memory corrected errors occur on a non-lru movable page,
> we can choose to stop using it by migrating data onto another
> page and disable the original (maybe half-broken) one.
> 
> Signed-off-by: Yisheng Xie <xieyisheng1@huawei.com>

It looks OK in my quick glance. I'll do some testing more tomorrow.

Thanks,
Naoya Horiguchi

> ---
>  mm/memory-failure.c | 55 +++++++++++++++++++++++++++++++++++++++++++++++++++--
>  1 file changed, 53 insertions(+), 2 deletions(-)
> 
> diff --git a/mm/memory-failure.c b/mm/memory-failure.c
> index f283c7e..10043a4 100644
> --- a/mm/memory-failure.c
> +++ b/mm/memory-failure.c
> @@ -1527,7 +1527,8 @@ static int get_any_page(struct page *page, unsigned long pfn, int flags)
>  {
>  	int ret = __get_any_page(page, pfn, flags);
>  
> -	if (ret == 1 && !PageHuge(page) && !PageLRU(page)) {
> +	if (ret == 1 && !PageHuge(page) &&
> +	    !PageLRU(page) && !__PageMovable(page)) {
>  		/*
>  		 * Try to free it.
>  		 */
> @@ -1549,6 +1550,54 @@ static int get_any_page(struct page *page, unsigned long pfn, int flags)
>  	return ret;
>  }
>  
> +static int soft_offline_movable_page(struct page *page, int flags)
> +{
> +	int ret;
> +	unsigned long pfn = page_to_pfn(page);
> +	LIST_HEAD(pagelist);
> +
> +	/*
> +	 * This double-check of PageHWPoison is to avoid the race with
> +	 * memory_failure(). See also comment in __soft_offline_page().
> +	 */
> +	lock_page(page);
> +	if (PageHWPoison(page)) {
> +		unlock_page(page);
> +		put_hwpoison_page(page);
> +		pr_info("soft offline: %#lx movable page already poisoned\n",
> +			pfn);
> +		return -EBUSY;
> +	}
> +	unlock_page(page);
> +
> +	ret = isolate_movable_page(page, ISOLATE_UNEVICTABLE);
> +	/*
> +	 * get_any_page() and isolate_movable_page() takes a refcount each,
> +	 * so need to drop one here.
> +	 */
> +	put_hwpoison_page(page);
> +	if (!ret) {
> +		pr_info("soft offline: %#lx movable page failed to isolate\n",
> +			pfn);
> +		return -EBUSY;
> +	}
> +
> +	list_add(&page->lru, &pagelist);
> +	ret = migrate_pages(&pagelist, new_page, NULL, MPOL_MF_MOVE_ALL,
> +			    MIGRATE_SYNC, MR_MEMORY_FAILURE);
> +	if (ret) {
> +		if (!list_empty(&pagelist))
> +			putback_movable_pages(&pagelist);
> +
> +		pr_info("soft offline: %#lx: migration failed %d, type %lx\n",
> +			pfn, ret, page->flags);
> +		if (ret > 0)
> +			ret = -EIO;
> +	}
> +
> +	return ret;
> +}
> +
>  static int soft_offline_huge_page(struct page *page, int flags)
>  {
>  	int ret;
> @@ -1705,8 +1754,10 @@ static int soft_offline_in_use_page(struct page *page, int flags)
>  
>  	if (PageHuge(page))
>  		ret = soft_offline_huge_page(page, flags);
> -	else
> +	else if (PageLRU(page))
>  		ret = __soft_offline_page(page, flags);
> +	else
> +		ret = soft_offline_movable_page(page, flags);
>  
>  	return ret;
>  }
> -- 
> 1.7.12.4
> 

[toc] | [prev] | [next] | [standalone]


#1563428

FromYisheng Xie <xieyisheng1@huawei.com>
Date2017-01-20 11:00 +0100
Message-ID<t1E9b-5CQ-9@gated-at.bofh.it>
In reply to#1561424
Hi Naoya,

On 2017/1/18 17:45, Naoya Horiguchi wrote:
> On Wed, Jan 18, 2017 at 12:00:54PM +0800, Yisheng Xie wrote:
>> This patch is to extends soft offlining framework to support
>> non-lru page, which already support migration after
>> commit bda807d44454 ("mm: migrate: support non-lru movable page
>> migration")
>>
>> When memory corrected errors occur on a non-lru movable page,
>> we can choose to stop using it by migrating data onto another
>> page and disable the original (maybe half-broken) one.
>>
>> Signed-off-by: Yisheng Xie <xieyisheng1@huawei.com>
> 
> It looks OK in my quick glance. I'll do some testing more tomorrow.
> 
Thanks for reviewing.
I have do some basic test like offline movable page and unpoison it.
Do you have some test suit or test suggestion? So I can do some more
test of it for double check? Very thanks for that.

Thanks
Yisheng Xie.

> 

[toc] | [prev] | [next] | [standalone]


#1564668

FromNaoya Horiguchi <n-horiguchi@ah.jp.nec.com>
Date2017-01-23 05:50 +0100
Message-ID<t2EJQ-1Od-23@gated-at.bofh.it>
In reply to#1563428
On Fri, Jan 20, 2017 at 05:52:13PM +0800, Yisheng Xie wrote:
> Hi Naoya,
> 
> On 2017/1/18 17:45, Naoya Horiguchi wrote:
> > On Wed, Jan 18, 2017 at 12:00:54PM +0800, Yisheng Xie wrote:
> >> This patch is to extends soft offlining framework to support
> >> non-lru page, which already support migration after
> >> commit bda807d44454 ("mm: migrate: support non-lru movable page
> >> migration")
> >>
> >> When memory corrected errors occur on a non-lru movable page,
> >> we can choose to stop using it by migrating data onto another
> >> page and disable the original (maybe half-broken) one.
> >>
> >> Signed-off-by: Yisheng Xie <xieyisheng1@huawei.com>
> > 
> > It looks OK in my quick glance. I'll do some testing more tomorrow.
> > 
> Thanks for reviewing.
> I have do some basic test like offline movable page and unpoison it.
> Do you have some test suit or test suggestion? So I can do some more
> test of it for double check? Very thanks for that.

I've tried soft offline on zram pages with your v2 patch, and it works fine.
I have no specific suggestion about other testcases.

Thanks,
Naoya Horiguchi

[toc] | [prev] | [next] | [standalone]


#1561430

FromMichal Hocko <mhocko@kernel.org>
Date2017-01-18 11:10 +0100
Message-ID<t0VlM-2Ex-17@gated-at.bofh.it>
In reply to#1561257
On Wed 18-01-17 12:00:54, Yisheng Xie wrote:
> This patch is to extends soft offlining framework to support
> non-lru page, which already support migration after
> commit bda807d44454 ("mm: migrate: support non-lru movable page
> migration")
> 
> When memory corrected errors occur on a non-lru movable page,
> we can choose to stop using it by migrating data onto another
> page and disable the original (maybe half-broken) one.

soft_offline_movable_page duplicates quite a lot from
__soft_offline_page. Would it be better to handle both cases in
__soft_offline_page?
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1562357

FromYisheng Xie <xieyisheng1@huawei.com>
Date2017-01-19 02:30 +0100
Message-ID<t19I5-3fA-5@gated-at.bofh.it>
In reply to#1561430

On 2017/1/18 17:51, Michal Hocko wrote:
> On Wed 18-01-17 12:00:54, Yisheng Xie wrote:
>> This patch is to extends soft offlining framework to support
>> non-lru page, which already support migration after
>> commit bda807d44454 ("mm: migrate: support non-lru movable page
>> migration")
>>
>> When memory corrected errors occur on a non-lru movable page,
>> we can choose to stop using it by migrating data onto another
>> page and disable the original (maybe half-broken) one.
> 
> soft_offline_movable_page duplicates quite a lot from
> __soft_offline_page. Would it be better to handle both cases in
> __soft_offline_page?
> 
Hi Michal,
Thanks for reviewing.
Yes, the most code of soft_offline_movable_page is duplicates with
__soft_offline_page, I use a single function to make code looks clear,
just as what soft_offline_hugetlb_page do.

I will try to make a v2 as your suggestion.

Thanks
Yisheng Xie.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web