Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1448297

Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the reclaim path

From NeilBrown <neilb@suse.com>
Newsgroups linux.kernel
Subject Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the reclaim path
Date 2016-07-22 03:50 +0200
Message-ID <rXxLb-7m5-7@gated-at.bofh.it> (permalink)
References (4 earlier) <rWUTw-7mO-19@gated-at.bofh.it> <rX6UG-6zY-5@gated-at.bofh.it> <rXhZL-57r-15@gated-at.bofh.it> <rXl7j-7jq-5@gated-at.bofh.it> <rXnCa-kD-23@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


[Multipart message — attachments visible in raw view] - view raw

On Fri, Jul 22 2016, Michal Hocko wrote:

> On Thu 21-07-16 08:13:00, Johannes Weiner wrote:
>> On Thu, Jul 21, 2016 at 10:52:03AM +0200, Michal Hocko wrote:
>> > Look, there are
>> > $ git grep mempool_alloc | wc -l
>> > 304
>> > 
>> > many users of this API and we do not want to flip the default behavior
>> > which is there for more than 10 years. So far you have been arguing
>> > about potential deadlocks and haven't shown any particular path which
>> > would have a direct or indirect dependency between mempool and normal
>> > allocator and it wouldn't be a bug. As the matter of fact the change
>> > we are discussing here causes a regression. If you want to change the
>> > semantic of mempool allocator then you are absolutely free to do so. In
>> > a separate patch which would be discussed with IO people and other
>> > users, though. But we _absolutely_ want to fix the regression first
>> > and have a simple fix for 4.6 and 4.7 backports. At this moment there
>> > are revert and patch 1 on the table.  The later one should make your
>> > backtrace happy and should be only as a temporal fix until we find out
>> > what is actually misbehaving on your systems. If you are not interested
>> > to pursue that way I will simply go with the revert.
>> 
>> +1
>> 
>> It's very unlikely that decade-old mempool semantics are suddenly a
>> fundamental livelock problem, when all the evidence we have is one
>> hang and vague speculation. Given that the patch causes regressions,
>> and that the bug is most likely elsewhere anyway, a full revert rather
>> than merely-less-invasive mempool changes makes the most sense to me.
>
> OK, fair enough. What do you think about the following then? Mikulas, I
> have dropped your Tested-by and Reviewed-by because the patch is
> different but unless you have hit the OOM killer then the testing
> results should be same.
> ---
> From d64815758c212643cc1750774e2751721685059a Mon Sep 17 00:00:00 2001
> From: Michal Hocko <mhocko@suse.com>
> Date: Thu, 21 Jul 2016 16:40:59 +0200
> Subject: [PATCH] Revert "mm, mempool: only set __GFP_NOMEMALLOC if there are
>  free elements"
>
> This reverts commit f9054c70d28bc214b2857cf8db8269f4f45a5e23.
>
> There has been a report about OOM killer invoked when swapping out to
> a dm-crypt device. The primary reason seems to be that the swapout
> out IO managed to completely deplete memory reserves. Ondrej was
> able to bisect and explained the issue by pointing to f9054c70d28b
> ("mm, mempool: only set __GFP_NOMEMALLOC if there are free elements").
>
> The reason is that the swapout path is not throttled properly because
> the md-raid layer needs to allocate from the generic_make_request path
> which means it allocates from the PF_MEMALLOC context. dm layer uses
> mempool_alloc in order to guarantee a forward progress which used to
> inhibit access to memory reserves when using page allocator. This has
> changed by f9054c70d28b ("mm, mempool: only set __GFP_NOMEMALLOC if
> there are free elements") which has dropped the __GFP_NOMEMALLOC
> protection when the memory pool is depleted.
>
> If we are running out of memory and the only way forward to free memory
> is to perform swapout we just keep consuming memory reserves rather than
> throttling the mempool allocations and allowing the pending IO to
> complete up to a moment when the memory is depleted completely and there
> is no way forward but invoking the OOM killer. This is less than
> optimal.
>
> The original intention of f9054c70d28b was to help with the OOM
> situations where the oom victim depends on mempool allocation to make a
> forward progress. David has mentioned the following backtrace:
>
> schedule
> schedule_timeout
> io_schedule_timeout
> mempool_alloc
> __split_and_process_bio
> dm_request
> generic_make_request
> submit_bio
> mpage_readpages
> ext4_readpages
> __do_page_cache_readahead
> ra_submit
> filemap_fault
> handle_mm_fault
> __do_page_fault
> do_page_fault
> page_fault
>
> We do not know more about why the mempool is depleted without being
> replenished in time, though. In any case the dm layer shouldn't depend
> on any allocations outside of the dedicated pools so a forward progress
> should be guaranteed. If this is not the case then the dm should be
> fixed rather than papering over the problem and postponing it to later
> by accessing more memory reserves.
>
> mempools are a mechanism to maintain dedicated memory reserves to guaratee
> forward progress. Allowing them an unbounded access to the page allocator
> memory reserves is going against the whole purpose of this mechanism.
>
> Bisected-by: Ondrej Kozina <okozina@redhat.com>
> Signed-off-by: Michal Hocko <mhocko@suse.com>
> ---
>  mm/mempool.c | 20 ++++----------------
>  1 file changed, 4 insertions(+), 16 deletions(-)
>
> diff --git a/mm/mempool.c b/mm/mempool.c
> index 8f65464da5de..5ba6c8b3b814 100644
> --- a/mm/mempool.c
> +++ b/mm/mempool.c
> @@ -306,36 +306,25 @@ EXPORT_SYMBOL(mempool_resize);
>   * returns NULL. Note that due to preallocation, this function
>   * *never* fails when called from process contexts. (it might
>   * fail if called from an IRQ context.)
> - * Note: neither __GFP_NOMEMALLOC nor __GFP_ZERO are supported.
> + * Note: using __GFP_ZERO is not supported.
>   */
> -void *mempool_alloc(mempool_t *pool, gfp_t gfp_mask)
> +void * mempool_alloc(mempool_t *pool, gfp_t gfp_mask)
>  {
>  	void *element;
>  	unsigned long flags;
>  	wait_queue_t wait;
>  	gfp_t gfp_temp;
>  
> -	/* If oom killed, memory reserves are essential to prevent livelock */
> -	VM_WARN_ON_ONCE(gfp_mask & __GFP_NOMEMALLOC);
> -	/* No element size to zero on allocation */
>  	VM_WARN_ON_ONCE(gfp_mask & __GFP_ZERO);
> -
>  	might_sleep_if(gfp_mask & __GFP_DIRECT_RECLAIM);
>  
> +	gfp_mask |= __GFP_NOMEMALLOC;	/* don't allocate emergency reserves */
>  	gfp_mask |= __GFP_NORETRY;	/* don't loop in __alloc_pages */
>  	gfp_mask |= __GFP_NOWARN;	/* failures are OK */

As I was reading through this thread I kept thinking "Surely
mempool_alloc() should never ever allocate from emergency reserves.
Ever."
Then I saw this patch.  It made me happy.

Thanks.

Acked-by: NeilBrown <neilb@suse.com>
(if you want it)

NeilBrown

Back to linux.kernel | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

[RFC PATCH 0/2] mempool vs. page allocator interaction Michal Hocko <mhocko@kernel.org> - 2016-07-18 10:40 +0200
  [RFC PATCH 1/2] mempool: do not consume memory reserves from the reclaim path Michal Hocko <mhocko@kernel.org> - 2016-07-18 10:50 +0200
    [RFC PATCH 2/2] mm, mempool: do not throttle PF_LESS_THROTTLE tasks Michal Hocko <mhocko@kernel.org> - 2016-07-18 10:50 +0200
      Re: [RFC PATCH 2/2] mm, mempool: do not throttle PF_LESS_THROTTLE  tasks Mikulas Patocka <mpatocka@redhat.com> - 2016-07-20 00:00 +0200
      Re: [RFC PATCH 2/2] mm, mempool: do not throttle PF_LESS_THROTTLE tasks NeilBrown <neilb@suse.de> - 2016-07-22 10:50 +0200
        Re: [RFC PATCH 2/2] mm, mempool: do not throttle PF_LESS_THROTTLE tasks NeilBrown <neilb@suse.com> - 2016-07-22 11:10 +0200
        Re: [RFC PATCH 2/2] mm, mempool: do not throttle PF_LESS_THROTTLE  tasks Michal Hocko <mhocko@kernel.org> - 2016-07-22 11:20 +0200
          Re: [RFC PATCH 2/2] mm, mempool: do not throttle PF_LESS_THROTTLE tasks NeilBrown <neilb@suse.com> - 2016-07-23 02:20 +0200
            Re: [RFC PATCH 2/2] mm, mempool: do not throttle PF_LESS_THROTTLE  tasks Michal Hocko <mhocko@kernel.org> - 2016-07-25 10:40 +0200
              Re: [RFC PATCH 2/2] mm, mempool: do not throttle PF_LESS_THROTTLE  tasks Michal Hocko <mhocko@kernel.org> - 2016-07-25 21:30 +0200
    Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from  the reclaim path David Rientjes <rientjes@google.com> - 2016-07-19 04:10 +0200
      Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Michal Hocko <mhocko@kernel.org> - 2016-07-19 09:50 +0200
    Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Johannes Weiner <hannes@cmpxchg.org> - 2016-07-19 16:00 +0200
      Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Michal Hocko <mhocko@kernel.org> - 2016-07-19 16:30 +0200
        Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from  the reclaim path Mikulas Patocka <mpatocka@redhat.com> - 2016-07-20 00:10 +0200
      Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from  the reclaim path David Rientjes <rientjes@google.com> - 2016-07-19 22:50 +0200
        Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Michal Hocko <mhocko@kernel.org> - 2016-07-20 10:20 +0200
          Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from  the reclaim path David Rientjes <rientjes@google.com> - 2016-07-20 23:10 +0200
            Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Michal Hocko <mhocko@kernel.org> - 2016-07-21 11:00 +0200
              Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Johannes Weiner <hannes@cmpxchg.org> - 2016-07-21 14:20 +0200
                Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Michal Hocko <mhocko@kernel.org> - 2016-07-21 17:00 +0200
                Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Johannes Weiner <hannes@cmpxchg.org> - 2016-07-21 17:30 +0200
                Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the reclaim path NeilBrown <neilb@suse.com> - 2016-07-22 03:50 +0200
                Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Michal Hocko <mhocko@kernel.org> - 2016-07-22 08:40 +0200
                Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Vlastimil Babka <vbabka@suse.cz> - 2016-07-22 14:30 +0200
                Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from  the reclaim path Andrew Morton <akpm@linux-foundation.org> - 2016-07-22 21:50 +0200
                Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Vlastimil Babka <vbabka@suse.cz> - 2016-07-23 21:00 +0200
    Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from  the reclaim path Mikulas Patocka <mpatocka@redhat.com> - 2016-07-20 00:00 +0200
      Re: [RFC PATCH 1/2] mempool: do not consume memory reserves from the  reclaim path Michal Hocko <mhocko@kernel.org> - 2016-07-20 08:50 +0200

csiph-web