Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1277184 > unrolled thread

[PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves

Started byMichal Hocko <mhocko@kernel.org>
First post2015-11-25 11:50 +0100
Last post2015-12-03 01:10 +0100
Articles 9 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves Michal Hocko <mhocko@kernel.org> - 2015-11-25 11:50 +0100
    Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to  memory reserves David Rientjes <rientjes@google.com> - 2015-11-25 12:00 +0100
      Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to  memory reserves Michal Hocko <mhocko@kernel.org> - 2015-11-25 12:20 +0100
        Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to  memory reserves David Rientjes <rientjes@google.com> - 2015-11-25 22:00 +0100
          Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to  memory reserves Michal Hocko <mhocko@kernel.org> - 2015-11-26 10:40 +0100
            Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to  memory reserves David Rientjes <rientjes@google.com> - 2015-11-30 23:20 +0100
              Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to  memory reserves Michal Hocko <mhocko@kernel.org> - 2015-12-02 16:10 +0100
    [PATCH v2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves Michal Hocko <mhocko@kernel.org> - 2015-12-02 16:20 +0100
      Re: [PATCH v2] mm, oom: Give __GFP_NOFAIL allocations access to  memory reserves David Rientjes <rientjes@google.com> - 2015-12-03 01:10 +0100

#1277184 — [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves

FromMichal Hocko <mhocko@kernel.org>
Date2015-11-25 11:50 +0100
Subject[PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves
Message-ID<qyFO9-56R-1@gated-at.bofh.it>
From: Michal Hocko <mhocko@suse.com>

__GFP_NOFAIL is a big hammer used to ensure that the allocation
request can never fail. This is a strong requirement and as such
it also deserves a special treatment when the system is OOM. The
primary problem here is that the allocation request might have
come with some locks held and the oom victim might be blocked
on the same locks. This is basically an OOM deadlock situation.

This patch tries to reduce the risk of such a deadlocks by giving
__GFP_NOFAIL allocations a special treatment and let them dive into
memory reserves after oom killer invocation. This should help them
to make a progress and release resources they are holding. The OOM
victim should compensate for the reserves consumption.

Suggested-by: Andrea Arcangeli <aarcange@redhat.com>
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
 mm/page_alloc.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 8034909faad2..70db11c27046 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -2766,8 +2766,13 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
 			goto out;
 	}
 	/* Exhausted what can be done so it's blamo time */
-	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL))
+	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL)) {
 		*did_some_progress = 1;
+
+		if (gfp_mask & __GFP_NOFAIL)
+			page = get_page_from_freelist(gfp_mask, order,
+					ALLOC_NO_WATERMARKS|ALLOC_CPUSET, ac);
+	}
 out:
 	mutex_unlock(&oom_lock);
 	return page;
-- 
2.6.2

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1277197 — Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves

FromDavid Rientjes <rientjes@google.com>
Date2015-11-25 12:00 +0100
SubjectRe: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves
Message-ID<qyFXP-5aJ-9@gated-at.bofh.it>
In reply to#1277184
On Wed, 25 Nov 2015, Michal Hocko wrote:

> From: Michal Hocko <mhocko@suse.com>
> 
> __GFP_NOFAIL is a big hammer used to ensure that the allocation
> request can never fail. This is a strong requirement and as such
> it also deserves a special treatment when the system is OOM. The
> primary problem here is that the allocation request might have
> come with some locks held and the oom victim might be blocked
> on the same locks. This is basically an OOM deadlock situation.
> 
> This patch tries to reduce the risk of such a deadlocks by giving
> __GFP_NOFAIL allocations a special treatment and let them dive into
> memory reserves after oom killer invocation. This should help them
> to make a progress and release resources they are holding. The OOM
> victim should compensate for the reserves consumption.
> 
> Suggested-by: Andrea Arcangeli <aarcange@redhat.com>
> Signed-off-by: Michal Hocko <mhocko@suse.com>
> ---
>  mm/page_alloc.c | 7 ++++++-
>  1 file changed, 6 insertions(+), 1 deletion(-)
> 
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index 8034909faad2..70db11c27046 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -2766,8 +2766,13 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
>  			goto out;
>  	}
>  	/* Exhausted what can be done so it's blamo time */
> -	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL))
> +	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL)) {
>  		*did_some_progress = 1;
> +
> +		if (gfp_mask & __GFP_NOFAIL)
> +			page = get_page_from_freelist(gfp_mask, order,
> +					ALLOC_NO_WATERMARKS|ALLOC_CPUSET, ac);
> +	}
>  out:
>  	mutex_unlock(&oom_lock);
>  	return page;

I don't understand why you're setting ALLOC_CPUSET if you're giving them 
"special treatment".  If you want to allow access to memory reserves to 
prevent an oom livelock, then why not also allow it access to allocate 
outside its cpuset?
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1277233 — Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves

FromMichal Hocko <mhocko@kernel.org>
Date2015-11-25 12:20 +0100
SubjectRe: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves
Message-ID<qyGhb-5xi-7@gated-at.bofh.it>
In reply to#1277197
On Wed 25-11-15 02:51:38, David Rientjes wrote:
> On Wed, 25 Nov 2015, Michal Hocko wrote:
> 
> > From: Michal Hocko <mhocko@suse.com>
> > 
> > __GFP_NOFAIL is a big hammer used to ensure that the allocation
> > request can never fail. This is a strong requirement and as such
> > it also deserves a special treatment when the system is OOM. The
> > primary problem here is that the allocation request might have
> > come with some locks held and the oom victim might be blocked
> > on the same locks. This is basically an OOM deadlock situation.
> > 
> > This patch tries to reduce the risk of such a deadlocks by giving
> > __GFP_NOFAIL allocations a special treatment and let them dive into
> > memory reserves after oom killer invocation. This should help them
> > to make a progress and release resources they are holding. The OOM
> > victim should compensate for the reserves consumption.
> > 
> > Suggested-by: Andrea Arcangeli <aarcange@redhat.com>
> > Signed-off-by: Michal Hocko <mhocko@suse.com>
> > ---
> >  mm/page_alloc.c | 7 ++++++-
> >  1 file changed, 6 insertions(+), 1 deletion(-)
> > 
> > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > index 8034909faad2..70db11c27046 100644
> > --- a/mm/page_alloc.c
> > +++ b/mm/page_alloc.c
> > @@ -2766,8 +2766,13 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
> >  			goto out;
> >  	}
> >  	/* Exhausted what can be done so it's blamo time */
> > -	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL))
> > +	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL)) {
> >  		*did_some_progress = 1;
> > +
> > +		if (gfp_mask & __GFP_NOFAIL)
> > +			page = get_page_from_freelist(gfp_mask, order,
> > +					ALLOC_NO_WATERMARKS|ALLOC_CPUSET, ac);
> > +	}
> >  out:
> >  	mutex_unlock(&oom_lock);
> >  	return page;
> 
> I don't understand why you're setting ALLOC_CPUSET if you're giving them 
> "special treatment".  If you want to allow access to memory reserves to 
> prevent an oom livelock, then why not also allow it access to allocate 
> outside its cpuset?

Good question. My thinking was that __GFP_NOFAIL allocations might be
done on behalf on a process so they are not necessarily system wide. We
do the same before we actually go to out_of_memory. On the other hand
__GFP_NOFAIL should be used really rarely and so breaking the cpuset
restriction shouldn't be a big deal if that helps to break out from the
potential OOM deadlock. I will drop it.

Thanks!
---
From d89d17e72e5e1c03539f8c81fc6e120bccd2b460 Mon Sep 17 00:00:00 2001
From: Michal Hocko <mhocko@suse.com>
Date: Tue, 23 Jun 2015 09:15:00 +0200
Subject: [PATCH] mm, oom: Give __GFP_NOFAIL allocations access to memory
 reserves

__GFP_NOFAIL is a big hammer used to ensure that the allocation
request can never fail. This is a strong requirement and as such
it also deserves a special treatment when the system is OOM. The
primary problem here is that the allocation request might have
come with some locks held and the oom victim might be blocked
on the same locks. This is basically an OOM deadlock situation.

This patch tries to reduce the risk of such a deadlocks by giving
__GFP_NOFAIL allocations a special treatment and let them dive into
memory reserves after oom killer invocation. This should help them
to make a progress and release resources they are holding. The OOM
victim should compensate for the reserves consumption.

[rientjes@google.com: do not use ALLOC_CPUSET]
Suggested-by: Andrea Arcangeli <aarcange@redhat.com>
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
 mm/page_alloc.c | 7 ++++++-
 1 file changed, 6 insertions(+), 1 deletion(-)

diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 8034909faad2..94b04c1e894a 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -2766,8 +2766,13 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
 			goto out;
 	}
 	/* Exhausted what can be done so it's blamo time */
-	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL))
+	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL)) {
 		*did_some_progress = 1;
+
+		if (gfp_mask & __GFP_NOFAIL)
+			page = get_page_from_freelist(gfp_mask, order,
+					ALLOC_NO_WATERMARKS, ac);
+	}
 out:
 	mutex_unlock(&oom_lock);
 	return page;
-- 
2.6.2


-- 
Michal Hocko
SUSE Labs
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1277793 — Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves

FromDavid Rientjes <rientjes@google.com>
Date2015-11-25 22:00 +0100
SubjectRe: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves
Message-ID<qyPku-2Rv-3@gated-at.bofh.it>
In reply to#1277233
On Wed, 25 Nov 2015, Michal Hocko wrote:

> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index 8034909faad2..94b04c1e894a 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -2766,8 +2766,13 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
>  			goto out;
>  	}
>  	/* Exhausted what can be done so it's blamo time */
> -	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL))
> +	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL)) {
>  		*did_some_progress = 1;
> +
> +		if (gfp_mask & __GFP_NOFAIL)
> +			page = get_page_from_freelist(gfp_mask, order,
> +					ALLOC_NO_WATERMARKS, ac);
> +	}
>  out:
>  	mutex_unlock(&oom_lock);
>  	return page;

Well, sure, that's one way to do it, but for cpuset users, wouldn't this 
lead to a depletion of the first system zone since you've dropped 
ALLOC_CPUSET and are doing ALLOC_NO_WATERMARKS in the same call?  
get_page_from_freelist() shouldn't be doing any balancing over the set of 
allowed zones.  Can you justify depleting memory reserves on a zone 
outside of the set of allowed cpuset mems rather than trying to drop 
ALLOC_CPUSET first?
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1278116 — Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves

FromMichal Hocko <mhocko@kernel.org>
Date2015-11-26 10:40 +0100
SubjectRe: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves
Message-ID<qz1bY-3hr-11@gated-at.bofh.it>
In reply to#1277793
On Wed 25-11-15 12:57:08, David Rientjes wrote:
> On Wed, 25 Nov 2015, Michal Hocko wrote:
> 
> > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > index 8034909faad2..94b04c1e894a 100644
> > --- a/mm/page_alloc.c
> > +++ b/mm/page_alloc.c
> > @@ -2766,8 +2766,13 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
> >  			goto out;
> >  	}
> >  	/* Exhausted what can be done so it's blamo time */
> > -	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL))
> > +	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL)) {
> >  		*did_some_progress = 1;
> > +
> > +		if (gfp_mask & __GFP_NOFAIL)
> > +			page = get_page_from_freelist(gfp_mask, order,
> > +					ALLOC_NO_WATERMARKS, ac);
> > +	}
> >  out:
> >  	mutex_unlock(&oom_lock);
> >  	return page;
> 
> Well, sure, that's one way to do it, but for cpuset users, wouldn't this 
> lead to a depletion of the first system zone since you've dropped 
> ALLOC_CPUSET and are doing ALLOC_NO_WATERMARKS in the same call?  

Are you suggesting to do?
		if (gfp_mask & __GFP_NOFAIL) {
			page = get_page_from_freelist(gfp_mask, order,
					ALLOC_NO_WATERMARKS|ALLOC_CPUSET, ac);
			/*
			 * fallback to ignore cpuset if our nodes are
			 * depleted
			 */
			if (!page)
				get_page_from_freelist(gfp_mask, order,
					ALLOC_NO_WATERMARKS, ac);
		}

I am not really sure this worth complication. __GFP_NOFAIL should be
relatively rare and nodes are rarely depeleted so much that
ALLOC_NO_WATERMARKS wouldn't be able to allocate from the first zone in
the zone list. I mean I have no problem to do the above it just sounds
overcomplicating the situation without making practical difference.
If you and others insist I can resping the patch though.
-- 
Michal Hocko
SUSE Labs
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1280360 — Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves

FromDavid Rientjes <rientjes@google.com>
Date2015-11-30 23:20 +0100
SubjectRe: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves
Message-ID<qAEXF-11R-33@gated-at.bofh.it>
In reply to#1278116
On Thu, 26 Nov 2015, Michal Hocko wrote:

> > > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > > index 8034909faad2..94b04c1e894a 100644
> > > --- a/mm/page_alloc.c
> > > +++ b/mm/page_alloc.c
> > > @@ -2766,8 +2766,13 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
> > >  			goto out;
> > >  	}
> > >  	/* Exhausted what can be done so it's blamo time */
> > > -	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL))
> > > +	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL)) {
> > >  		*did_some_progress = 1;
> > > +
> > > +		if (gfp_mask & __GFP_NOFAIL)
> > > +			page = get_page_from_freelist(gfp_mask, order,
> > > +					ALLOC_NO_WATERMARKS, ac);
> > > +	}
> > >  out:
> > >  	mutex_unlock(&oom_lock);
> > >  	return page;
> > 
> > Well, sure, that's one way to do it, but for cpuset users, wouldn't this 
> > lead to a depletion of the first system zone since you've dropped 
> > ALLOC_CPUSET and are doing ALLOC_NO_WATERMARKS in the same call?  
> 
> Are you suggesting to do?
> 		if (gfp_mask & __GFP_NOFAIL) {
> 			page = get_page_from_freelist(gfp_mask, order,
> 					ALLOC_NO_WATERMARKS|ALLOC_CPUSET, ac);
> 			/*
> 			 * fallback to ignore cpuset if our nodes are
> 			 * depleted
> 			 */
> 			if (!page)
> 				get_page_from_freelist(gfp_mask, order,
> 					ALLOC_NO_WATERMARKS, ac);
> 		}
> 
> I am not really sure this worth complication.

I'm objecting to the ability of a process that is doing a __GFP_NOFAIL 
allocation, which has been disallowed access from allocating on certain 
mems through cpusets, to cause an oom condition on those disallowed nodes, 
yes.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1281886 — Re: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves

FromMichal Hocko <mhocko@kernel.org>
Date2015-12-02 16:10 +0100
SubjectRe: [PATCH 1/2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves
Message-ID<qBhcB-nj-17@gated-at.bofh.it>
In reply to#1280360
On Mon 30-11-15 14:17:03, David Rientjes wrote:
> On Thu, 26 Nov 2015, Michal Hocko wrote:
> 
> > > > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > > > index 8034909faad2..94b04c1e894a 100644
> > > > --- a/mm/page_alloc.c
> > > > +++ b/mm/page_alloc.c
> > > > @@ -2766,8 +2766,13 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
> > > >  			goto out;
> > > >  	}
> > > >  	/* Exhausted what can be done so it's blamo time */
> > > > -	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL))
> > > > +	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL)) {
> > > >  		*did_some_progress = 1;
> > > > +
> > > > +		if (gfp_mask & __GFP_NOFAIL)
> > > > +			page = get_page_from_freelist(gfp_mask, order,
> > > > +					ALLOC_NO_WATERMARKS, ac);
> > > > +	}
> > > >  out:
> > > >  	mutex_unlock(&oom_lock);
> > > >  	return page;
> > > 
> > > Well, sure, that's one way to do it, but for cpuset users, wouldn't this 
> > > lead to a depletion of the first system zone since you've dropped 
> > > ALLOC_CPUSET and are doing ALLOC_NO_WATERMARKS in the same call?  
> > 
> > Are you suggesting to do?
> > 		if (gfp_mask & __GFP_NOFAIL) {
> > 			page = get_page_from_freelist(gfp_mask, order,
> > 					ALLOC_NO_WATERMARKS|ALLOC_CPUSET, ac);
> > 			/*
> > 			 * fallback to ignore cpuset if our nodes are
> > 			 * depleted
> > 			 */
> > 			if (!page)
> > 				get_page_from_freelist(gfp_mask, order,
> > 					ALLOC_NO_WATERMARKS, ac);
> > 		}
> > 
> > I am not really sure this worth complication.
> 
> I'm objecting to the ability of a process that is doing a __GFP_NOFAIL 
> allocation, which has been disallowed access from allocating on certain 
> mems through cpusets, to cause an oom condition on those disallowed nodes, 
> yes.

That ability will be there even with the fallback mechanism. My primary
objections was that the fallback is unnecessarily complex without any
evidence that such a situation would happen in the real life often
enought to bother about it. __GFP_NOFAIL allocations are and should be
rare and any runaway triggerable from the userspace is a kernel bug.

Anyway, as you seem to feel really strongly about this I will post v2
with the above fallback. This is a superslow path anyway...

-- 
Michal Hocko
SUSE Labs
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1281893 — [PATCH v2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves

FromMichal Hocko <mhocko@kernel.org>
Date2015-12-02 16:20 +0100
Subject[PATCH v2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves
Message-ID<qBhmh-sK-1@gated-at.bofh.it>
In reply to#1277184
From: Michal Hocko <mhocko@suse.com>

__GFP_NOFAIL is a big hammer used to ensure that the allocation
request can never fail. This is a strong requirement and as such
it also deserves a special treatment when the system is OOM. The
primary problem here is that the allocation request might have
come with some locks held and the oom victim might be blocked
on the same locks. This is basically an OOM deadlock situation.

This patch tries to reduce the risk of such a deadlocks by giving
__GFP_NOFAIL allocations a special treatment and let them dive into
memory reserves after oom killer invocation. This should help them
to make a progress and release resources they are holding. The OOM
victim should compensate for the reserves consumption.

Suggested-by: Andrea Arcangeli <aarcange@redhat.com>
Signed-off-by: Michal Hocko <mhocko@suse.com>
---
 mm/page_alloc.c | 15 ++++++++++++++-
 1 file changed, 14 insertions(+), 1 deletion(-)

diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 8034909faad2..367523b2948b 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -2766,8 +2766,21 @@ __alloc_pages_may_oom(gfp_t gfp_mask, unsigned int order,
 			goto out;
 	}
 	/* Exhausted what can be done so it's blamo time */
-	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL))
+	if (out_of_memory(&oc) || WARN_ON_ONCE(gfp_mask & __GFP_NOFAIL)) {
 		*did_some_progress = 1;
+
+		if (gfp_mask & __GFP_NOFAIL) {
+			page = get_page_from_freelist(gfp_mask, order,
+					ALLOC_NO_WATERMARKS|ALLOC_CPUSET, ac);
+			/*
+			 * fallback to ignore cpuset restriction if our nodes
+			 * are depleted
+			 */
+			if (!page)
+				page = get_page_from_freelist(gfp_mask, order,
+					ALLOC_NO_WATERMARKS, ac);
+		}
+	}
 out:
 	mutex_unlock(&oom_lock);
 	return page;
-- 
2.6.2

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1282610 — Re: [PATCH v2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves

FromDavid Rientjes <rientjes@google.com>
Date2015-12-03 01:10 +0100
SubjectRe: [PATCH v2] mm, oom: Give __GFP_NOFAIL allocations access to memory reserves
Message-ID<qBpDb-5UD-11@gated-at.bofh.it>
In reply to#1281893
On Wed, 2 Dec 2015, Michal Hocko wrote:

> From: Michal Hocko <mhocko@suse.com>
> 
> __GFP_NOFAIL is a big hammer used to ensure that the allocation
> request can never fail. This is a strong requirement and as such
> it also deserves a special treatment when the system is OOM. The
> primary problem here is that the allocation request might have
> come with some locks held and the oom victim might be blocked
> on the same locks. This is basically an OOM deadlock situation.
> 
> This patch tries to reduce the risk of such a deadlocks by giving
> __GFP_NOFAIL allocations a special treatment and let them dive into
> memory reserves after oom killer invocation. This should help them
> to make a progress and release resources they are holding. The OOM
> victim should compensate for the reserves consumption.
> 
> Suggested-by: Andrea Arcangeli <aarcange@redhat.com>
> Signed-off-by: Michal Hocko <mhocko@suse.com>

Acked-by: David Rientjes <rientjes@google.com>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web