Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1215957 > unrolled thread
| Started by | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| First post | 2015-08-30 21:10 +0200 |
| Last post | 2015-09-04 16:40 +0200 |
| Articles | 20 on this page of 31 — 4 participants |
Back to article view | Back to linux.kernel
[PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-08-30 21:10 +0200
[PATCH 2/2] mm/slub: do not bypass memcg reclaim for high-order page allocation Vladimir Davydov <vdavydov@parallels.com> - 2015-08-30 21:10 +0200
[PATCH 1/2] mm/slab: skip memcg reclaim only if in atomic context Vladimir Davydov <vdavydov@parallels.com> - 2015-08-30 21:10 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Michal Hocko <mhocko@kernel.org> - 2015-08-31 15:30 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Tejun Heo <tj@kernel.org> - 2015-08-31 15:50 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Tejun Heo <tj@kernel.org> - 2015-08-31 16:40 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-08-31 17:20 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Tejun Heo <tj@kernel.org> - 2015-08-31 17:50 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-08-31 19:00 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Tejun Heo <tj@kernel.org> - 2015-08-31 19:10 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-08-31 21:30 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Christoph Lameter <cl@linux.com> - 2015-08-31 22:30 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-09-01 11:30 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-08-31 16:40 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-08-31 16:30 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Tejun Heo <tj@kernel.org> - 2015-08-31 16:50 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-08-31 17:30 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Michal Hocko <mhocko@kernel.org> - 2015-09-01 14:40 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-09-01 15:50 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Michal Hocko <mhocko@kernel.org> - 2015-09-01 17:10 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-09-01 19:00 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Michal Hocko <mhocko@kernel.org> - 2015-09-01 20:40 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-09-02 11:40 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Christoph Lameter <cl@linux.com> - 2015-09-02 20:20 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-09-03 11:40 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Tejun Heo <tj@kernel.org> - 2015-09-03 18:40 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-09-04 13:20 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Tejun Heo <tj@kernel.org> - 2015-09-04 17:50 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Vladimir Davydov <vdavydov@parallels.com> - 2015-09-04 20:30 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Tejun Heo <tj@kernel.org> - 2015-09-04 21:40 +0200
Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled Michal Hocko <mhocko@kernel.org> - 2015-09-04 16:40 +0200
Page 1 of 2 [1] 2 Next page →
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-08-30 21:10 +0200 |
| Subject | [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3g9j-17i-3@gated-at.bofh.it> |
Hi, Tejun reported that sometimes memcg/memory.high threshold seems to be silently ignored if kmem accounting is enabled: http://www.spinics.net/lists/linux-mm/msg93613.html It turned out that both SLAB and SLUB try to allocate without __GFP_WAIT first. As a result, if there is enough free pages, memcg reclaim will not get invoked on kmem allocations, which will lead to uncontrollable growth of memory usage no matter what memory.high is set to. This patch set attempts to fix this issue. For more details please see comments to individual patches. Thanks, Vladimir Davydov (2): mm/slab: skip memcg reclaim only if in atomic context mm/slub: do not bypass memcg reclaim for high-order page allocation mm/slab.c | 32 +++++++++++--------------------- mm/slub.c | 24 +++++++++++------------- 2 files changed, 22 insertions(+), 34 deletions(-) -- 2.1.4 -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-08-30 21:10 +0200 |
| Subject | [PATCH 2/2] mm/slub: do not bypass memcg reclaim for high-order page allocation |
| Message-ID | <q3g9k-17i-17@gated-at.bofh.it> |
| In reply to | #1215957 |
Commit 6af3142bed1f52 ("mm/slub: don't wait for high-order page
allocation") made allocate_slab() try to allocate high order slab pages
without __GFP_WAIT in order to avoid invoking reclaim/compaction when we
can fall back on low order pages. However, it broke memcg/memory.high
logic in case kmem accounting is enabled. The memory.high threshold
works as a soft limit: an allocation does not fail if it is breached,
but we call direct reclaim to compensate for the excess. Without
__GFP_WAIT we cannot invoke reclaimer and therefore we will go on
exceeding memory.high more and more until a normal __GFP_WAIT allocation
is issued.
Since memcg reclaim never triggers compaction, we can pass __GFP_WAIT to
memcg_charge_slab() even on high order page allocations w/o any
performance impact. So let us fix this problem by excluding __GFP_WAIT
only from alloc_pages() while still forwarding it to memcg_charge_slab()
if the context allows.
Reported-by: Tejun Heo <tj@kernel.org>
Signed-off-by: Vladimir Davydov <vdavydov@parallels.com>
---
mm/slub.c | 24 +++++++++++-------------
1 file changed, 11 insertions(+), 13 deletions(-)
diff --git a/mm/slub.c b/mm/slub.c
index e180f8dcd06d..416a332277cb 100644
--- a/mm/slub.c
+++ b/mm/slub.c
@@ -1333,6 +1333,14 @@ static inline struct page *alloc_slab_page(struct kmem_cache *s,
if (memcg_charge_slab(s, flags, order))
return NULL;
+ /*
+ * Let the initial higher-order allocation fail under memory pressure
+ * so we fall-back to the minimum order allocation.
+ */
+ if (oo_order(oo) > oo_order(s->min))
+ flags = (flags | __GFP_NOWARN | __GFP_NOMEMALLOC) &
+ ~(__GFP_NOFAIL | __GFP_WAIT);
+
if (node == NUMA_NO_NODE)
page = alloc_pages(flags, order);
else
@@ -1348,7 +1356,6 @@ static struct page *allocate_slab(struct kmem_cache *s, gfp_t flags, int node)
{
struct page *page;
struct kmem_cache_order_objects oo = s->oo;
- gfp_t alloc_gfp;
void *start, *p;
int idx, order;
@@ -1359,23 +1366,14 @@ static struct page *allocate_slab(struct kmem_cache *s, gfp_t flags, int node)
flags |= s->allocflags;
- /*
- * Let the initial higher-order allocation fail under memory pressure
- * so we fall-back to the minimum order allocation.
- */
- alloc_gfp = (flags | __GFP_NOWARN | __GFP_NORETRY) & ~__GFP_NOFAIL;
- if ((alloc_gfp & __GFP_WAIT) && oo_order(oo) > oo_order(s->min))
- alloc_gfp = (alloc_gfp | __GFP_NOMEMALLOC) & ~__GFP_WAIT;
-
- page = alloc_slab_page(s, alloc_gfp, node, oo);
+ page = alloc_slab_page(s, flags, node, oo);
if (unlikely(!page)) {
oo = s->min;
- alloc_gfp = flags;
/*
* Allocation may have failed due to fragmentation.
* Try a lower order alloc if possible
*/
- page = alloc_slab_page(s, alloc_gfp, node, oo);
+ page = alloc_slab_page(s, flags, node, oo);
if (unlikely(!page))
goto out;
stat(s, ORDER_FALLBACK);
@@ -1385,7 +1383,7 @@ static struct page *allocate_slab(struct kmem_cache *s, gfp_t flags, int node)
!(s->flags & (SLAB_NOTRACK | DEBUG_DEFAULT_FLAGS))) {
int pages = 1 << oo_order(oo);
- kmemcheck_alloc_shadow(page, oo_order(oo), alloc_gfp, node);
+ kmemcheck_alloc_shadow(page, oo_order(oo), flags, node);
/*
* Objects from caches that have a constructor don't get
--
2.1.4
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-08-30 21:10 +0200 |
| Subject | [PATCH 1/2] mm/slab: skip memcg reclaim only if in atomic context |
| Message-ID | <q3g9j-17i-7@gated-at.bofh.it> |
| In reply to | #1215957 |
SLAB's implementation of kmem_cache_alloc() works as follows:
1. First, it tries to allocate from the preferred NUMA node without
issuing reclaim.
2. If step 1 fails, it tries all nodes in the order of preference,
again without invoking reclaimer
3. Only if steps 1 and 2 fails, it falls back on allocation from any
allowed node with reclaim enabled.
Before commit 4167e9b2cf10f ("mm: remove GFP_THISNODE"), GFP_THISNODE
combination, which equaled __GFP_THISNODE|__GFP_NOWARN|__GFP_NORETRY on
NUMA enabled builds, was used in order to avoid reclaim during steps 1
and 2. If __alloc_pages_slowpath() saw this combination in gfp flags, it
aborted immediately even if __GFP_WAIT flag was set. So there was no
need in clearing __GFP_WAIT flag while performing steps 1 and 2 and
hence we could invoke memcg reclaim when allocating a slab page if the
context allowed.
Commit 4167e9b2cf10f zapped GFP_THISNODE combination. Instead of OR-ing
the gfp mask with GFP_THISNODE, gfp_exact_node() helper should now be
used. The latter sets __GFP_THISNODE and __GFP_NOWARN flags and clears
__GFP_WAIT on the current gfp mask. As a result, it effectively
prohibits invoking memcg reclaim on steps 1 and 2. This breaks
memcg/memory.high logic when kmem accounting is enabled. The memory.high
threshold is supposed to work as a soft limit, i.e. it does not fail an
allocation on breaching it, but it still forces the caller to invoke
direct reclaim to compensate for the excess. Without __GFP_WAIT flag
direct reclaim is impossible so the caller will go on without being
pushed back to the threshold.
To fix this issue, we get rid of gfp_exact_node() helper and move gfp
flags filtering to kmem_getpages() after memcg_charge_slab() is called.
To understand the patch, note that:
- In fallback_alloc() the only effect of using gfp_exact_node() is
preventing recursion fallback_alloc() -> ____cache_alloc_node() ->
fallback_alloc().
- Aside from fallback_alloc(), gfp_exact_node() is only used along with
cache_grow(). Moreover, the only place where cache_grow() is used
without it is fallback_alloc(), which, in contrast to other
cache_grow() users, preallocates a page and passes it to cache_grow()
so that the latter does not need to invoke kmem_getpages() by itself.
Reported-by: Tejun Heo <tj@kernel.org>
Signed-off-by: Vladimir Davydov <vdavydov@parallels.com>
---
mm/slab.c | 32 +++++++++++---------------------
1 file changed, 11 insertions(+), 21 deletions(-)
diff --git a/mm/slab.c b/mm/slab.c
index d890750ec31e..9ee809d2ed8b 100644
--- a/mm/slab.c
+++ b/mm/slab.c
@@ -857,11 +857,6 @@ static inline void *____cache_alloc_node(struct kmem_cache *cachep,
return NULL;
}
-static inline gfp_t gfp_exact_node(gfp_t flags)
-{
- return flags;
-}
-
#else /* CONFIG_NUMA */
static void *____cache_alloc_node(struct kmem_cache *, gfp_t, int);
@@ -1028,15 +1023,6 @@ static inline int cache_free_alien(struct kmem_cache *cachep, void *objp)
return __cache_free_alien(cachep, objp, node, page_node);
}
-
-/*
- * Construct gfp mask to allocate from a specific node but do not invoke reclaim
- * or warn about failures.
- */
-static inline gfp_t gfp_exact_node(gfp_t flags)
-{
- return (flags | __GFP_THISNODE | __GFP_NOWARN) & ~__GFP_WAIT;
-}
#endif
/*
@@ -1583,7 +1569,7 @@ slab_out_of_memory(struct kmem_cache *cachep, gfp_t gfpflags, int nodeid)
* would be relatively rare and ignorable.
*/
static struct page *kmem_getpages(struct kmem_cache *cachep, gfp_t flags,
- int nodeid)
+ int nodeid, bool fallback)
{
struct page *page;
int nr_pages;
@@ -1595,6 +1581,9 @@ static struct page *kmem_getpages(struct kmem_cache *cachep, gfp_t flags,
if (memcg_charge_slab(cachep, flags, cachep->gfporder))
return NULL;
+ if (!fallback)
+ flags = (flags | __GFP_THISNODE | __GFP_NOWARN) & ~__GFP_WAIT;
+
page = __alloc_pages_node(nodeid, flags | __GFP_NOTRACK, cachep->gfporder);
if (!page) {
memcg_uncharge_slab(cachep, cachep->gfporder);
@@ -2641,7 +2630,8 @@ static int cache_grow(struct kmem_cache *cachep,
* 'nodeid'.
*/
if (!page)
- page = kmem_getpages(cachep, local_flags, nodeid);
+ page = kmem_getpages(cachep, local_flags, nodeid,
+ !IS_ENABLED(CONFIG_NUMA));
if (!page)
goto failed;
@@ -2840,7 +2830,7 @@ alloc_done:
if (unlikely(!ac->avail)) {
int x;
force_grow:
- x = cache_grow(cachep, gfp_exact_node(flags), node, NULL);
+ x = cache_grow(cachep, flags, node, NULL);
/* cache_grow can reenable interrupts, then ac could change. */
ac = cpu_cache_get(cachep);
@@ -3034,7 +3024,7 @@ retry:
get_node(cache, nid) &&
get_node(cache, nid)->free_objects) {
obj = ____cache_alloc_node(cache,
- gfp_exact_node(flags), nid);
+ flags | __GFP_THISNODE, nid);
if (obj)
break;
}
@@ -3052,7 +3042,7 @@ retry:
if (local_flags & __GFP_WAIT)
local_irq_enable();
kmem_flagcheck(cache, flags);
- page = kmem_getpages(cache, local_flags, numa_mem_id());
+ page = kmem_getpages(cache, local_flags, numa_mem_id(), true);
if (local_flags & __GFP_WAIT)
local_irq_disable();
if (page) {
@@ -3062,7 +3052,7 @@ retry:
nid = page_to_nid(page);
if (cache_grow(cache, flags, nid, page)) {
obj = ____cache_alloc_node(cache,
- gfp_exact_node(flags), nid);
+ flags | __GFP_THISNODE, nid);
if (!obj)
/*
* Another processor may allocate the
@@ -3133,7 +3123,7 @@ retry:
must_grow:
spin_unlock(&n->list_lock);
- x = cache_grow(cachep, gfp_exact_node(flags), nodeid, NULL);
+ x = cache_grow(cachep, flags, nodeid, NULL);
if (x)
goto retry;
--
2.1.4
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2015-08-31 15:30 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3xjQ-v6-13@gated-at.bofh.it> |
| In reply to | #1215957 |
On Sun 30-08-15 22:02:16, Vladimir Davydov wrote: > Hi, > > Tejun reported that sometimes memcg/memory.high threshold seems to be > silently ignored if kmem accounting is enabled: > > http://www.spinics.net/lists/linux-mm/msg93613.html > > It turned out that both SLAB and SLUB try to allocate without __GFP_WAIT > first. As a result, if there is enough free pages, memcg reclaim will > not get invoked on kmem allocations, which will lead to uncontrollable > growth of memory usage no matter what memory.high is set to. Right but isn't that what the caller explicitly asked for? Why should we ignore that for kmem accounting? It seems like a fix at a wrong layer to me. Either we should start failing GFP_NOWAIT charges when we are above high wmark or deploy an additional catchup mechanism as suggested by Tejun. I like the later more because it allows to better handle GFP_NOFS requests as well and there are many sources of these from kmem paths. > This patch set attempts to fix this issue. For more details please see > comments to individual patches. > > Thanks, > > Vladimir Davydov (2): > mm/slab: skip memcg reclaim only if in atomic context > mm/slub: do not bypass memcg reclaim for high-order page allocation > > mm/slab.c | 32 +++++++++++--------------------- > mm/slub.c | 24 +++++++++++------------- > 2 files changed, 22 insertions(+), 34 deletions(-) > > -- > 2.1.4 -- Michal Hocko SUSE Labs -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-08-31 15:50 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3xDb-RS-11@gated-at.bofh.it> |
| In reply to | #1216177 |
Hello, On Mon, Aug 31, 2015 at 03:24:15PM +0200, Michal Hocko wrote: > Right but isn't that what the caller explicitly asked for? Why should we > ignore that for kmem accounting? It seems like a fix at a wrong layer to > me. Either we should start failing GFP_NOWAIT charges when we are above > high wmark or deploy an additional catchup mechanism as suggested by > Tejun. I like the later more because it allows to better handle GFP_NOFS > requests as well and there are many sources of these from kmem paths. Yeah, this is beginning to look like we're trying to solve the problem at the wrong layer. slab/slub or whatever else should be able to use GFP_NOWAIT in whatever frequency they want for speculative allocations. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-08-31 16:40 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3ypA-259-9@gated-at.bofh.it> |
| In reply to | #1216187 |
On Mon, Aug 31, 2015 at 05:30:08PM +0300, Vladimir Davydov wrote: > slab/slub can issue alloc_pages() any time with any flags they want and > it won't be accounted to memcg, because kmem is accounted at slab/slub > layer, not in buddy. Hmmm? I meant the eventual calling into try_charge w/ GFP_NOWAIT. Speculative usage of GFP_NOWAIT is bound to increase and we don't want to put on extra restrictions from memcg side. For memory.high, punting to the return path is a pretty stright-forward solution which should make the problem go away almost entirely. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-08-31 17:20 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3z2i-33U-5@gated-at.bofh.it> |
| In reply to | #1216224 |
On Mon, Aug 31, 2015 at 10:39:39AM -0400, Tejun Heo wrote: > On Mon, Aug 31, 2015 at 05:30:08PM +0300, Vladimir Davydov wrote: > > slab/slub can issue alloc_pages() any time with any flags they want and > > it won't be accounted to memcg, because kmem is accounted at slab/slub > > layer, not in buddy. > > Hmmm? I meant the eventual calling into try_charge w/ GFP_NOWAIT. > Speculative usage of GFP_NOWAIT is bound to increase and we don't want > to put on extra restrictions from memcg side. We already put restrictions on slab/slub from memcg side, because kmem accounting is a part of slab/slub. They have to cooperate in order to get things working. If slab/slub wants to make a speculative allocation for some reason, it should just put memcg_charge out of this speculative alloc section. This is what this patch set does. We have to be cautious about placing memcg_charge in slab/slub. To understand why, consider SLAB case, which first tries to allocate from all nodes in the order of preference w/o __GFP_WAIT and only if it fails falls back on an allocation from any node w/ __GFP_WAIT. This is its internal algorithm. If we blindly put memcg_charge to alloc_slab method, then, when we are near the memcg limit, we will go over all NUMA nodes in vain, then finally fall back to __GFP_WAIT allocation, which will get a slab from a random node. Not only we do more work than necessary due to walking over all NUMA nodes for nothing, but we also break SLAB internal logic! And you just can't fix it in memcg, because memcg knows nothing about the internal logic of SLAB, how it handles NUMA nodes. SLUB has a different problem. It tries to avoid high-order allocations if there is a risk of invoking costly memory compactor. It has nothing to do with memcg, because memcg does not care if the charge is for a high order page or not. Thanks, Vladimir > For memory.high, > punting to the return path is a pretty stright-forward solution which > should make the problem go away almost entirely. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-08-31 17:50 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3zvj-3E1-5@gated-at.bofh.it> |
| In reply to | #1216251 |
Hello, On Mon, Aug 31, 2015 at 06:18:14PM +0300, Vladimir Davydov wrote: > We have to be cautious about placing memcg_charge in slab/slub. To > understand why, consider SLAB case, which first tries to allocate from > all nodes in the order of preference w/o __GFP_WAIT and only if it fails > falls back on an allocation from any node w/ __GFP_WAIT. This is its > internal algorithm. If we blindly put memcg_charge to alloc_slab method, > then, when we are near the memcg limit, we will go over all NUMA nodes > in vain, then finally fall back to __GFP_WAIT allocation, which will get > a slab from a random node. Not only we do more work than necessary due > to walking over all NUMA nodes for nothing, but we also break SLAB > internal logic! And you just can't fix it in memcg, because memcg knows > nothing about the internal logic of SLAB, how it handles NUMA nodes. > > SLUB has a different problem. It tries to avoid high-order allocations > if there is a risk of invoking costly memory compactor. It has nothing > to do with memcg, because memcg does not care if the charge is for a > high order page or not. Maybe I'm missing something but aren't both issues caused by memcg failing to provide headroom for NOWAIT allocations when the consumption gets close to the max limit? Regardless of the specific usage, !__GFP_WAIT means "give me memory if it can be spared w/o inducing direct time-consuming maintenance work" and the contract around it is that such requests will mostly succeed under nominal conditions. Also, slab/slub might not stay as the only user of try_charge(). I still think solving this from memcg side is the right direction. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-08-31 19:00 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3AB6-5ck-43@gated-at.bofh.it> |
| In reply to | #1216277 |
On Mon, Aug 31, 2015 at 11:47:56AM -0400, Tejun Heo wrote: > On Mon, Aug 31, 2015 at 06:18:14PM +0300, Vladimir Davydov wrote: > > We have to be cautious about placing memcg_charge in slab/slub. To > > understand why, consider SLAB case, which first tries to allocate from > > all nodes in the order of preference w/o __GFP_WAIT and only if it fails > > falls back on an allocation from any node w/ __GFP_WAIT. This is its > > internal algorithm. If we blindly put memcg_charge to alloc_slab method, > > then, when we are near the memcg limit, we will go over all NUMA nodes > > in vain, then finally fall back to __GFP_WAIT allocation, which will get > > a slab from a random node. Not only we do more work than necessary due > > to walking over all NUMA nodes for nothing, but we also break SLAB > > internal logic! And you just can't fix it in memcg, because memcg knows > > nothing about the internal logic of SLAB, how it handles NUMA nodes. > > > > SLUB has a different problem. It tries to avoid high-order allocations > > if there is a risk of invoking costly memory compactor. It has nothing > > to do with memcg, because memcg does not care if the charge is for a > > high order page or not. > > Maybe I'm missing something but aren't both issues caused by memcg > failing to provide headroom for NOWAIT allocations when the > consumption gets close to the max limit? That's correct. > Regardless of the specific usage, !__GFP_WAIT means "give me memory if > it can be spared w/o inducing direct time-consuming maintenance work" > and the contract around it is that such requests will mostly succeed > under nominal conditions. Also, slab/slub might not stay as the only > user of try_charge(). Indeed, there might be other users trying GFP_NOWAIT before falling back to GFP_KERNEL, but they are not doing that constantly and hence cause no problems. If SLAB/SLUB plays such tricks, the problem becomes massive: under certain conditions *every* try_charge may be invoked w/o __GFP_WAIT, resulting in memory.high breaching and hitting memory.max. Generally speaking, handing over reclaim responsibility to task_work won't help, because there might be cases when a process spends quite a lot of time in kernel invoking lots of GFP_KERNEL allocations before returning to userspace. Without fixing slab/slub, such a process will charge w/o __GFP_WAIT and therefore can exceed memory.high and reach memory.max. If there are no other active processes in the cgroup, the cgroup can stay with memory.high excess for a relatively long time (suppose the process was throttled in kernel), possibly hurting the rest of the system. What is worse, if the process happens to invoke a real GFP_NOWAIT allocation when it's about to hit the limit, it will fail. If we want to allow slab/slub implementation to invoke try_charge wherever it wants, we need to introduce an asynchronous thread doing reclaim when a memcg is approaching its limit (or teach kswapd do that). That's a way to go, but what's the point to complicate things prematurely while it seems we can fix the problem by using the technique similar to the one behind memory.high? Nevertheless, even if we introduced such a thread, it'd be just insane to allow slab/slub blindly insert try_charge. Let me repeat the examples of SLAB/SLUB sub-optimal behavior caused by thoughtless usage of try_charge I gave above: - memcg knows nothing about NUMA nodes, so what's the point in failing !__GFP_WAIT allocations used by SLAB while inspecting NUMA nodes? - memcg knows nothing about high order pages, so what's the point in failing !__GFP_WAIT allocations used by SLUB to try to allocate a high order page? Thanks, Vladimir > I still think solving this from memcg side is the right direction. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-08-31 19:10 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3AKL-5CZ-29@gated-at.bofh.it> |
| In reply to | #1216323 |
Hello, On Mon, Aug 31, 2015 at 07:51:32PM +0300, Vladimir Davydov wrote: ... > If we want to allow slab/slub implementation to invoke try_charge > wherever it wants, we need to introduce an asynchronous thread doing > reclaim when a memcg is approaching its limit (or teach kswapd do that). In the long term, I think this is the way to go. > That's a way to go, but what's the point to complicate things > prematurely while it seems we can fix the problem by using the technique > similar to the one behind memory.high? Cuz we're now scattering workarounds to multiple places and I'm sure we'll add more try_charge() users (e.g. we want to fold in tcp memcg under the same knobs) and we'll have to worry about the same problem all over again and will inevitably miss some cases leading to subtle failures. > Nevertheless, even if we introduced such a thread, it'd be just insane > to allow slab/slub blindly insert try_charge. Let me repeat the examples > of SLAB/SLUB sub-optimal behavior caused by thoughtless usage of > try_charge I gave above: > > - memcg knows nothing about NUMA nodes, so what's the point in failing > !__GFP_WAIT allocations used by SLAB while inspecting NUMA nodes? > - memcg knows nothing about high order pages, so what's the point in > failing !__GFP_WAIT allocations used by SLUB to try to allocate a > high order page? Both are optimistic speculative actions and as long as memcg can guarantee that those requests will succeed under normal circumstances, as does the system-wide mm does, it isn't a problem. In general, we want to make sure inside-cgroup behaviors as close to system-wide behaviors as possible, scoped but equivalent in kind. Doing things differently, while inevitable in certain cases, is likely to get messy in the long term. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-08-31 21:30 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3CWf-ej-33@gated-at.bofh.it> |
| In reply to | #1216328 |
On Mon, Aug 31, 2015 at 01:03:09PM -0400, Tejun Heo wrote:
> On Mon, Aug 31, 2015 at 07:51:32PM +0300, Vladimir Davydov wrote:
> ...
> > If we want to allow slab/slub implementation to invoke try_charge
> > wherever it wants, we need to introduce an asynchronous thread doing
> > reclaim when a memcg is approaching its limit (or teach kswapd do that).
>
> In the long term, I think this is the way to go.
Quite probably, or we can use task_work, or direct reclaim instead. It's
not that obvious to me yet which one is the best.
>
> > That's a way to go, but what's the point to complicate things
> > prematurely while it seems we can fix the problem by using the technique
> > similar to the one behind memory.high?
>
> Cuz we're now scattering workarounds to multiple places and I'm sure
> we'll add more try_charge() users (e.g. we want to fold in tcp memcg
> under the same knobs) and we'll have to worry about the same problem
> all over again and will inevitably miss some cases leading to subtle
> failures.
I don't think we will need to insert try_charge_kmem anywhere else,
because all kmem users either allocate memory using kmalloc and friends
or using alloc_pages. kmalloc is accounted. For those who prefer
alloc_pages, there is alloc_kmem_pages helper.
>
> > Nevertheless, even if we introduced such a thread, it'd be just insane
> > to allow slab/slub blindly insert try_charge. Let me repeat the examples
> > of SLAB/SLUB sub-optimal behavior caused by thoughtless usage of
> > try_charge I gave above:
> >
> > - memcg knows nothing about NUMA nodes, so what's the point in failing
> > !__GFP_WAIT allocations used by SLAB while inspecting NUMA nodes?
> > - memcg knows nothing about high order pages, so what's the point in
> > failing !__GFP_WAIT allocations used by SLUB to try to allocate a
> > high order page?
>
> Both are optimistic speculative actions and as long as memcg can
> guarantee that those requests will succeed under normal circumstances,
> as does the system-wide mm does, it isn't a problem.
>
> In general, we want to make sure inside-cgroup behaviors as close to
> system-wide behaviors as possible, scoped but equivalent in kind.
> Doing things differently, while inevitable in certain cases, is likely
> to get messy in the long term.
I totally agree that we should strive to make a kmem user feel roughly
the same in memcg as if it were running on a host with equal amount of
RAM. There are two ways to achieve that:
1. Make the API functions, i.e. kmalloc and friends, behave inside
memcg roughly the same way as they do in the root cgroup.
2. Make the internal memcg functions, i.e. try_charge and friends,
behave roughly the same way as alloc_pages.
I find way 1 more flexible, because we don't have to blindly follow
heuristics used on global memory reclaim and therefore have more
opportunities to achieve the same goal.
Thanks,
Vladimir
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2015-08-31 22:30 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3DSj-1zN-11@gated-at.bofh.it> |
| In reply to | #1216396 |
On Mon, 31 Aug 2015, Vladimir Davydov wrote: > I totally agree that we should strive to make a kmem user feel roughly > the same in memcg as if it were running on a host with equal amount of > RAM. There are two ways to achieve that: > > 1. Make the API functions, i.e. kmalloc and friends, behave inside > memcg roughly the same way as they do in the root cgroup. > 2. Make the internal memcg functions, i.e. try_charge and friends, > behave roughly the same way as alloc_pages. > > I find way 1 more flexible, because we don't have to blindly follow > heuristics used on global memory reclaim and therefore have more > opportunities to achieve the same goal. The heuristics need to integrate well if its in a cgroup or not. In general make use of cgroups as transparent as possible to the rest of the code. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-09-01 11:30 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3Q37-2rk-3@gated-at.bofh.it> |
| In reply to | #1216426 |
On Mon, Aug 31, 2015 at 03:22:22PM -0500, Christoph Lameter wrote: > On Mon, 31 Aug 2015, Vladimir Davydov wrote: > > > I totally agree that we should strive to make a kmem user feel roughly > > the same in memcg as if it were running on a host with equal amount of > > RAM. There are two ways to achieve that: > > > > 1. Make the API functions, i.e. kmalloc and friends, behave inside > > memcg roughly the same way as they do in the root cgroup. > > 2. Make the internal memcg functions, i.e. try_charge and friends, > > behave roughly the same way as alloc_pages. > > > > I find way 1 more flexible, because we don't have to blindly follow > > heuristics used on global memory reclaim and therefore have more > > opportunities to achieve the same goal. > > The heuristics need to integrate well if its in a cgroup or not. In > general make use of cgroups as transparent as possible to the rest of the > code. Half of kmem accounting implementation resides in SLAB/SLUB. We can't just make use of cgroups there transparent. For the rest of the code using kmalloc, cgroups are transparent. Indeed, we can make memcg_charge_slab behave exactly like alloc_pages, we can even put it to alloc_pages (where it used to be), but why if the only user of memcg_charge_slab is SLAB/SLUB core? I think we'd have more space to manoeuvre if we just taught SLAB/SLUB to use memcg_charge_slab wisely (as it used to until recently), because memcg charge/reclaim is quite different from global alloc/reclaim: - it isn't aware of NUMA nodes, so trying to charge w/o __GFP_WAIT while inspecting nodes, like in case of SLAB, is meaningless - it isn't aware of high order page allocations, so trying to charge w/o __GFP_WAIT while trying optimistically to get a high order page, like in case of SLUB, is meaningless too - it can always let a high prio allocation go unaccounted, so IMO there is no point in introducing emergency reserves (__GFP_MEMALLOC handling) - it can always charge a GFP_NOWAIT allocation even if it exceeds the limit, issuing direct reclaim when a GFP_KERNEL allocation comes or from a task work, because there is no risk of depleting memory reserves; so it isn't obvious to me whether we really need an aync thread handling memcg reclaim like kswapd Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-08-31 16:40 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3ypA-259-11@gated-at.bofh.it> |
| In reply to | #1216187 |
On Mon, Aug 31, 2015 at 09:43:35AM -0400, Tejun Heo wrote: > On Mon, Aug 31, 2015 at 03:24:15PM +0200, Michal Hocko wrote: > > Right but isn't that what the caller explicitly asked for? Why should we > > ignore that for kmem accounting? It seems like a fix at a wrong layer to > > me. Either we should start failing GFP_NOWAIT charges when we are above > > high wmark or deploy an additional catchup mechanism as suggested by > > Tejun. I like the later more because it allows to better handle GFP_NOFS > > requests as well and there are many sources of these from kmem paths. > > Yeah, this is beginning to look like we're trying to solve the problem > at the wrong layer. slab/slub or whatever else should be able to use > GFP_NOWAIT in whatever frequency they want for speculative > allocations. slab/slub can issue alloc_pages() any time with any flags they want and it won't be accounted to memcg, because kmem is accounted at slab/slub layer, not in buddy. Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-08-31 16:30 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3yfW-1QK-45@gated-at.bofh.it> |
| In reply to | #1216177 |
On Mon, Aug 31, 2015 at 03:24:15PM +0200, Michal Hocko wrote:
> On Sun 30-08-15 22:02:16, Vladimir Davydov wrote:
> > Tejun reported that sometimes memcg/memory.high threshold seems to be
> > silently ignored if kmem accounting is enabled:
> >
> > http://www.spinics.net/lists/linux-mm/msg93613.html
> >
> > It turned out that both SLAB and SLUB try to allocate without __GFP_WAIT
> > first. As a result, if there is enough free pages, memcg reclaim will
> > not get invoked on kmem allocations, which will lead to uncontrollable
> > growth of memory usage no matter what memory.high is set to.
>
> Right but isn't that what the caller explicitly asked for?
No. If the caller of kmalloc() asked for a __GFP_WAIT allocation, we
might ignore that and charge memcg w/o __GFP_WAIT.
> Why should we ignore that for kmem accounting? It seems like a fix at
> a wrong layer to me.
Let's forget about memory.high for a minute.
1. SLAB. Suppose someone calls kmalloc_node and there is enough free
memory on the preferred node. W/o memcg limit set, the allocation
will happen from the preferred node, which is OK. If there is memcg
limit, we can currently fail to allocate from the preferred node if
we are near the limit. We issue memcg reclaim and go to fallback
alloc then, which will most probably allocate from a different node,
although there is no reason for that. This is a bug.
2. SLUB. Someone calls kmalloc and there is enough free high order
pages. If there is no memcg limit, we will allocate a high order
slab page, which is in accordance with SLUB internal logic. With
memcg limit set, we are likely to fail to charge high order page
(because we currently try to charge high order pages w/o __GFP_WAIT)
and fallback on a low order page. The latter is unexpected and
unjustified.
That being said, this is the fix at the right layer.
> Either we should start failing GFP_NOWAIT charges when we are above
> high wmark or deploy an additional catchup mechanism as suggested by
> Tejun.
The mechanism proposed by Tejun won't help us to avoid allocation
failures if we are hitting memory.max w/o __GFP_WAIT or __GFP_FS.
To fix GFP_NOFS/GFP_NOWAIT failures we just need to start reclaim when
the gap between limit and usage is getting too small. It may be done
from a workqueue or from task_work, but currently I don't see any reason
why complicate and not just start reclaim directly, just like
memory.high does.
I mean, currently you can protect against GFP_NOWAIT failures by setting
memory.high to be 1-2 MB lower than memory.high and this *will* work,
because GFP_NOWAIT/GFP_NOFS allocations can't go on infinitely - they
will alternate with normal GFP_KERNEL allocations sooner or later. It
does not mean we should encourage users to set memory.high to protect
against such failures, because, as pointed out by Tejun, logic behind
memory.high is currently opaque and can change, but we can introduce
memcg-internal watermarks that would work exactly as memory.high and
hence help us against GFP_NOWAIT/GFP_NOFS failures.
Thanks,
Vladimir
> I like the later more because it allows to better handle GFP_NOFS
> requests as well and there are many sources of these from kmem paths.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-08-31 16:50 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3yzf-2gs-19@gated-at.bofh.it> |
| In reply to | #1216223 |
Hello, Vladimir. On Mon, Aug 31, 2015 at 05:20:49PM +0300, Vladimir Davydov wrote: ... > That being said, this is the fix at the right layer. While this *might* be a necessary workaround for the hard limit case right now, this is by no means the fix at the right layer. The expectation is that mm keeps a reasonable amount of memory available for allocations which can't block. These allocations may fail from time to time depending on luck and under extreme memory pressure but the caller should be able to depend on it as a speculative allocation mechanism which doesn't fail willy-nilly. Hardlimit breaking GFP_NOWAIT behavior is a bug on memcg side, not slab or slub. Thanks. -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-08-31 17:30 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3zbY-3fy-13@gated-at.bofh.it> |
| In reply to | #1216231 |
On Mon, Aug 31, 2015 at 10:46:04AM -0400, Tejun Heo wrote: > Hello, Vladimir. > > On Mon, Aug 31, 2015 at 05:20:49PM +0300, Vladimir Davydov wrote: > ... > > That being said, this is the fix at the right layer. > > While this *might* be a necessary workaround for the hard limit case > right now, this is by no means the fix at the right layer. The > expectation is that mm keeps a reasonable amount of memory available > for allocations which can't block. These allocations may fail from > time to time depending on luck and under extreme memory pressure but > the caller should be able to depend on it as a speculative allocation > mechanism which doesn't fail willy-nilly. > > Hardlimit breaking GFP_NOWAIT behavior is a bug on memcg side, not > slab or slub. I never denied that there is GFP_NOWAIT/GFP_NOFS problem in memcg. I even proposed ways to cope with it in one of the previous e-mails. Nevertheless, we just can't allow slab/slub internals call memcg_charge whenever they want as I pointed out in a parallel thread. Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2015-09-01 14:40 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3T11-6Dv-17@gated-at.bofh.it> |
| In reply to | #1216223 |
On Mon 31-08-15 17:20:49, Vladimir Davydov wrote: > On Mon, Aug 31, 2015 at 03:24:15PM +0200, Michal Hocko wrote: > > On Sun 30-08-15 22:02:16, Vladimir Davydov wrote: > > > > Tejun reported that sometimes memcg/memory.high threshold seems to be > > > silently ignored if kmem accounting is enabled: > > > > > > http://www.spinics.net/lists/linux-mm/msg93613.html > > > > > > It turned out that both SLAB and SLUB try to allocate without __GFP_WAIT > > > first. As a result, if there is enough free pages, memcg reclaim will > > > not get invoked on kmem allocations, which will lead to uncontrollable > > > growth of memory usage no matter what memory.high is set to. > > > > Right but isn't that what the caller explicitly asked for? > > No. If the caller of kmalloc() asked for a __GFP_WAIT allocation, we > might ignore that and charge memcg w/o __GFP_WAIT. I was referring to the slab allocator as the caller. Sorry for not being clear about that. > > Why should we ignore that for kmem accounting? It seems like a fix at > > a wrong layer to me. > > Let's forget about memory.high for a minute. > > 1. SLAB. Suppose someone calls kmalloc_node and there is enough free > memory on the preferred node. W/o memcg limit set, the allocation > will happen from the preferred node, which is OK. If there is memcg > limit, we can currently fail to allocate from the preferred node if > we are near the limit. We issue memcg reclaim and go to fallback > alloc then, which will most probably allocate from a different node, > although there is no reason for that. This is a bug. I am not familiar with the SLAB internals much but how is it different from the global case. If the preferred node is full then __GFP_THISNODE request will make it fail early even without giving GFP_NOWAIT additional access to atomic memory reserves. The fact that memcg case fails earlier is perfectly expected because the restriction is tighter than the global case. How the fallback is implemented and whether trying other node before reclaiming from the preferred one is reasonable I dunno. This is for SLAB to decide. But ignoring GFP_NOWAIT for this path makes the behavior for memcg enabled setups subtly different. And that is bad. > 2. SLUB. Someone calls kmalloc and there is enough free high order > pages. If there is no memcg limit, we will allocate a high order > slab page, which is in accordance with SLUB internal logic. With > memcg limit set, we are likely to fail to charge high order page > (because we currently try to charge high order pages w/o __GFP_WAIT) > and fallback on a low order page. The latter is unexpected and > unjustified. And this case very similar and I even argue that it shows more brokenness with your patch. The SLUB allocator has _explicitly_ asked for an allocation _without_ reclaim because that would be unnecessarily too costly and there is other less expensive fallback. But memcg would be ignoring this with your patch AFAIU and break the optimization. There are other cases like that. E.g. THP pages are allocated without GFP_WAIT when defrag is disabled. > That being said, this is the fix at the right layer. > > > Either we should start failing GFP_NOWAIT charges when we are above > > high wmark or deploy an additional catchup mechanism as suggested by > > Tejun. > > The mechanism proposed by Tejun won't help us to avoid allocation > failures if we are hitting memory.max w/o __GFP_WAIT or __GFP_FS. Why would be that a problem. The _hard_ limit is reached and reclaim cannot make any progress. An allocation failure is to be expected. GFP_NOWAIT will fail normally and GFP_NOFS will attempt to reclaim before failing. > To fix GFP_NOFS/GFP_NOWAIT failures we just need to start reclaim when > the gap between limit and usage is getting too small. It may be done > from a workqueue or from task_work, but currently I don't see any reason > why complicate and not just start reclaim directly, just like > memory.high does. Yes we can do better than we do right now. But that doesn't mean we should put hacks all over the place and lie about the allocation context. > I mean, currently you can protect against GFP_NOWAIT failures by setting > memory.high to be 1-2 MB lower than memory.high and this *will* work, > because GFP_NOWAIT/GFP_NOFS allocations can't go on infinitely - they > will alternate with normal GFP_KERNEL allocations sooner or later. It > does not mean we should encourage users to set memory.high to protect > against such failures, because, as pointed out by Tejun, logic behind > memory.high is currently opaque and can change, but we can introduce > memcg-internal watermarks that would work exactly as memory.high and > hence help us against GFP_NOWAIT/GFP_NOFS failures. I am not against something like watermarks and doing more pro-active reclaim but this is far from easy to do - which is one of the reason we do not have it yet. The idea from Tejun about the return to userspace reclaim is nice in that regards that it happens from a well defined context and helps to keep memory.high behavior much saner. -- Michal Hocko SUSE Labs -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@parallels.com> |
|---|---|
| Date | 2015-09-01 15:50 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3U6K-8ac-19@gated-at.bofh.it> |
| In reply to | #1216815 |
On Tue, Sep 01, 2015 at 02:36:12PM +0200, Michal Hocko wrote: > On Mon 31-08-15 17:20:49, Vladimir Davydov wrote: > > On Mon, Aug 31, 2015 at 03:24:15PM +0200, Michal Hocko wrote: > > > On Sun 30-08-15 22:02:16, Vladimir Davydov wrote: > > > > > > Tejun reported that sometimes memcg/memory.high threshold seems to be > > > > silently ignored if kmem accounting is enabled: > > > > > > > > http://www.spinics.net/lists/linux-mm/msg93613.html > > > > > > > > It turned out that both SLAB and SLUB try to allocate without __GFP_WAIT > > > > first. As a result, if there is enough free pages, memcg reclaim will > > > > not get invoked on kmem allocations, which will lead to uncontrollable > > > > growth of memory usage no matter what memory.high is set to. > > > > > > Right but isn't that what the caller explicitly asked for? > > > > No. If the caller of kmalloc() asked for a __GFP_WAIT allocation, we > > might ignore that and charge memcg w/o __GFP_WAIT. > > I was referring to the slab allocator as the caller. Sorry for not being > clear about that. > > > > Why should we ignore that for kmem accounting? It seems like a fix at > > > a wrong layer to me. > > > > Let's forget about memory.high for a minute. > > > > 1. SLAB. Suppose someone calls kmalloc_node and there is enough free > > memory on the preferred node. W/o memcg limit set, the allocation > > will happen from the preferred node, which is OK. If there is memcg > > limit, we can currently fail to allocate from the preferred node if > > we are near the limit. We issue memcg reclaim and go to fallback > > alloc then, which will most probably allocate from a different node, > > although there is no reason for that. This is a bug. > > I am not familiar with the SLAB internals much but how is it different > from the global case. If the preferred node is full then __GFP_THISNODE > request will make it fail early even without giving GFP_NOWAIT > additional access to atomic memory reserves. The fact that memcg case > fails earlier is perfectly expected because the restriction is tighter > than the global case. memcg restrictions are orthogonal to NUMA: failing an allocation from a particular node does not mean failing memcg charge and vice versa. > > How the fallback is implemented and whether trying other node before > reclaiming from the preferred one is reasonable I dunno. This is for > SLAB to decide. But ignoring GFP_NOWAIT for this path makes the behavior > for memcg enabled setups subtly different. And that is bad. Quite the contrary. Trying to charge memcg w/o __GFP_WAIT while inspecting if a NUMA node has free pages makes SLAB behaviour subtly differently: SLAB will walk over all NUMA nodes for nothing instead of invoking memcg reclaim once a free page is found. You are talking about memcg/kmem accounting as if it were done in the buddy allocator on top of which the slab layer is built knowing nothing about memcg accounting on the lower layer. That's not true and that simply can't be true. Kmem accounting is implemented at the slab layer. Memcg provides its memcg_charge_slab/uncharge methods solely for slab core, so it's OK to have some calling conventions between them. What we are really obliged to do is to preserve behavior of slab's external API, i.e. kmalloc and friends. > > > 2. SLUB. Someone calls kmalloc and there is enough free high order > > pages. If there is no memcg limit, we will allocate a high order > > slab page, which is in accordance with SLUB internal logic. With > > memcg limit set, we are likely to fail to charge high order page > > (because we currently try to charge high order pages w/o __GFP_WAIT) > > and fallback on a low order page. The latter is unexpected and > > unjustified. > > And this case very similar and I even argue that it shows more > brokenness with your patch. The SLUB allocator has _explicitly_ asked > for an allocation _without_ reclaim because that would be unnecessarily > too costly and there is other less expensive fallback. But memcg would You are ignoring the fact that, in contrast to alloc_pages, for memcg there is practically no difference between charging a 4-order page or a 1-order page. OTOH, using 1-order pages where we could go with 4-order pages increases page fragmentation at the global level. This subtly breaks internal SLUB optimization. Once again, kmem accounting is not something staying aside from slab core, it's a part of slab core. > be ignoring this with your patch AFAIU and break the optimization. There > are other cases like that. E.g. THP pages are allocated without GFP_WAIT > when defrag is disabled. It might be wrong. If we can't find a continuous 2Mb page, we should probably give up instead of calling compactor. For memcg it might be better to reclaim some space for 2Mb page right now and map a 2Mb page instead of reclaiming space for 512 4Kb pages a moment later, because in memcg case there is absolutely no difference between reclaiming 2Mb for a huge page and 2Mb for 512 4Kb pages. > > > That being said, this is the fix at the right layer. > > > > > Either we should start failing GFP_NOWAIT charges when we are above > > > high wmark or deploy an additional catchup mechanism as suggested by > > > Tejun. > > > > The mechanism proposed by Tejun won't help us to avoid allocation > > failures if we are hitting memory.max w/o __GFP_WAIT or __GFP_FS. > > Why would be that a problem. The _hard_ limit is reached and reclaim > cannot make any progress. An allocation failure is to be expected. > GFP_NOWAIT will fail normally and GFP_NOFS will attempt to reclaim > before failing. Quoting my e-mail to Tejun explaining why using task_work won't help if we don't fix SLAB/SLUB: : Generally speaking, handing over reclaim responsibility to task_work : won't help, because there might be cases when a process spends quite a : lot of time in kernel invoking lots of GFP_KERNEL allocations before : returning to userspace. Without fixing slab/slub, such a process will : charge w/o __GFP_WAIT and therefore can exceed memory.high and reach : memory.max. If there are no other active processes in the cgroup, the : cgroup can stay with memory.high excess for a relatively long time : (suppose the process was throttled in kernel), possibly hurting the rest : of the system. What is worse, if the process happens to invoke a real : GFP_NOWAIT allocation when it's about to hit the limit, it will fail. For a kmalloc user that's completely unexpected. > > > To fix GFP_NOFS/GFP_NOWAIT failures we just need to start reclaim when > > the gap between limit and usage is getting too small. It may be done > > from a workqueue or from task_work, but currently I don't see any reason > > why complicate and not just start reclaim directly, just like > > memory.high does. > > Yes we can do better than we do right now. But that doesn't mean we > should put hacks all over the place and lie about the allocation > context. What do you mean by saying "all over the place"? It's a fix for kmem implementation, to be more exact for the part of it residing in the slab core. Everyone else, except a couple of kmem users issuing alloc_page directly like threadinfo, will use kmalloc and know nothing what's going on there and how all this accounting stuff is handled - they will just use plain old convenient kmalloc, which works exactly as it does in the root cgroup. > > > I mean, currently you can protect against GFP_NOWAIT failures by setting > > memory.high to be 1-2 MB lower than memory.high and this *will* work, > > because GFP_NOWAIT/GFP_NOFS allocations can't go on infinitely - they > > will alternate with normal GFP_KERNEL allocations sooner or later. It > > does not mean we should encourage users to set memory.high to protect > > against such failures, because, as pointed out by Tejun, logic behind > > memory.high is currently opaque and can change, but we can introduce > > memcg-internal watermarks that would work exactly as memory.high and > > hence help us against GFP_NOWAIT/GFP_NOFS failures. > > I am not against something like watermarks and doing more pro-active > reclaim but this is far from easy to do - which is one of the reason we > do not have it yet. The idea from Tejun about the return to userspace > reclaim is nice in that regards that it happens from a well defined > context and helps to keep memory.high behavior much saner. I don't say what Tejun proposed is a crap. It might be a very good lightweight alternative to per memcg kswapd. However, w/o fixing SLAB/SLUB it's useless. Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2015-09-01 17:10 +0200 |
| Subject | Re: [PATCH 0/2] Fix memcg/memory.high in case kmem accounting is enabled |
| Message-ID | <q3Vma-1Gq-1@gated-at.bofh.it> |
| In reply to | #1216849 |
On Tue 01-09-15 16:40:03, Vladimir Davydov wrote:
> On Tue, Sep 01, 2015 at 02:36:12PM +0200, Michal Hocko wrote:
> > On Mon 31-08-15 17:20:49, Vladimir Davydov wrote:
{...}
> > > 1. SLAB. Suppose someone calls kmalloc_node and there is enough free
> > > memory on the preferred node. W/o memcg limit set, the allocation
> > > will happen from the preferred node, which is OK. If there is memcg
> > > limit, we can currently fail to allocate from the preferred node if
> > > we are near the limit. We issue memcg reclaim and go to fallback
> > > alloc then, which will most probably allocate from a different node,
> > > although there is no reason for that. This is a bug.
> >
> > I am not familiar with the SLAB internals much but how is it different
> > from the global case. If the preferred node is full then __GFP_THISNODE
> > request will make it fail early even without giving GFP_NOWAIT
> > additional access to atomic memory reserves. The fact that memcg case
> > fails earlier is perfectly expected because the restriction is tighter
> > than the global case.
>
> memcg restrictions are orthogonal to NUMA: failing an allocation from a
> particular node does not mean failing memcg charge and vice versa.
Sure memcg doesn't care about NUMA it just puts an additional constrain
on top of all existing ones. The point I've tried to make is that the
logic is currently same whether it is page allocator (with the node
restriction) or memcg (cumulative amount restriction) are behaving
consistently. Neither of them try to reclaim in order to achieve its
goals. How conservative is memcg about allowing GFP_NOWAIT allocation
is a separate issue and all those details belong to memcg proper same
as the allocation strategy for these allocations belongs to the page
allocator.
> > How the fallback is implemented and whether trying other node before
> > reclaiming from the preferred one is reasonable I dunno. This is for
> > SLAB to decide. But ignoring GFP_NOWAIT for this path makes the behavior
> > for memcg enabled setups subtly different. And that is bad.
>
> Quite the contrary. Trying to charge memcg w/o __GFP_WAIT while
> inspecting if a NUMA node has free pages makes SLAB behaviour subtly
> differently: SLAB will walk over all NUMA nodes for nothing instead of
> invoking memcg reclaim once a free page is found.
So you are saying that the SLAB kmem accounting in this particular path
is suboptimal because the fallback mode doesn't retry local node with
the reclaim enabled before falling back to other nodes?
I would consider it quite surprising as well even for the global case
because __GFP_THISNODE doesn't wake up kswapd to make room on that node.
> You are talking about memcg/kmem accounting as if it were done in the
> buddy allocator on top of which the slab layer is built knowing nothing
> about memcg accounting on the lower layer. That's not true and that
> simply can't be true. Kmem accounting is implemented at the slab layer.
> Memcg provides its memcg_charge_slab/uncharge methods solely for
> slab core, so it's OK to have some calling conventions between them.
> What we are really obliged to do is to preserve behavior of slab's
> external API, i.e. kmalloc and friends.
I guess I understand what you are saying here but it sounds like special
casing which tries to be clever because the current code understands
both the lower level allocator and kmem charge paths to decide how to
juggle with them. This is imho bad and hard to maintain long term.
> > > 2. SLUB. Someone calls kmalloc and there is enough free high order
> > > pages. If there is no memcg limit, we will allocate a high order
> > > slab page, which is in accordance with SLUB internal logic. With
> > > memcg limit set, we are likely to fail to charge high order page
> > > (because we currently try to charge high order pages w/o __GFP_WAIT)
> > > and fallback on a low order page. The latter is unexpected and
> > > unjustified.
> >
> > And this case very similar and I even argue that it shows more
> > brokenness with your patch. The SLUB allocator has _explicitly_ asked
> > for an allocation _without_ reclaim because that would be unnecessarily
> > too costly and there is other less expensive fallback. But memcg would
>
> You are ignoring the fact that, in contrast to alloc_pages, for memcg
> there is practically no difference between charging a 4-order page or a
> 1-order page.
But this is an implementation details which might change anytime in
future.
> OTOH, using 1-order pages where we could go with 4-order
> pages increases page fragmentation at the global level. This subtly
> breaks internal SLUB optimization. Once again, kmem accounting is not
> something staying aside from slab core, it's a part of slab core.
This is certainly true and it is what you get when you put an additional
constrain on top of an existing one. You simply cannot get both the
great performance _and_ a local memory restriction.
> > be ignoring this with your patch AFAIU and break the optimization. There
> > are other cases like that. E.g. THP pages are allocated without GFP_WAIT
> > when defrag is disabled.
>
> It might be wrong. If we can't find a continuous 2Mb page, we should
> probably give up instead of calling compactor. For memcg it might be
> better to reclaim some space for 2Mb page right now and map a 2Mb page
> instead of reclaiming space for 512 4Kb pages a moment later, because in
> memcg case there is absolutely no difference between reclaiming 2Mb for
> a huge page and 2Mb for 512 4Kb pages.
Or maybe the whole reclaim just doesn't pay off because the TLB savings
will never compensate for the reclaim. The defrag knob basically says
that we shouldn't try to opportunistically prepare a room for the THP
page.
> > > That being said, this is the fix at the right layer.
> > >
> > > > Either we should start failing GFP_NOWAIT charges when we are above
> > > > high wmark or deploy an additional catchup mechanism as suggested by
> > > > Tejun.
> > >
> > > The mechanism proposed by Tejun won't help us to avoid allocation
> > > failures if we are hitting memory.max w/o __GFP_WAIT or __GFP_FS.
> >
> > Why would be that a problem. The _hard_ limit is reached and reclaim
> > cannot make any progress. An allocation failure is to be expected.
> > GFP_NOWAIT will fail normally and GFP_NOFS will attempt to reclaim
> > before failing.
>
> Quoting my e-mail to Tejun explaining why using task_work won't help if
> we don't fix SLAB/SLUB:
>
> : Generally speaking, handing over reclaim responsibility to task_work
> : won't help, because there might be cases when a process spends quite a
> : lot of time in kernel invoking lots of GFP_KERNEL allocations before
> : returning to userspace. Without fixing slab/slub, such a process will
> : charge w/o __GFP_WAIT and therefore can exceed memory.high and reach
> : memory.max. If there are no other active processes in the cgroup, the
> : cgroup can stay with memory.high excess for a relatively long time
> : (suppose the process was throttled in kernel), possibly hurting the rest
> : of the system. What is worse, if the process happens to invoke a real
> : GFP_NOWAIT allocation when it's about to hit the limit, it will fail.
>
> For a kmalloc user that's completely unexpected.
We have the global reclaim which handles the global memory pressure. And
until the hard limit is enforced I do not see what is the huge problem
here. Sure we can have high limit in excess but that is to be expected.
Same as failing allocations for the hard limit enforcement.
Maybe moving whole high limit reclaim to the delayed context is not what
we will end up with and reduce this only for GFP_NOWAIT or other weak
reclaim contexts. This is to be discussed of course.
> > > To fix GFP_NOFS/GFP_NOWAIT failures we just need to start reclaim when
> > > the gap between limit and usage is getting too small. It may be done
> > > from a workqueue or from task_work, but currently I don't see any reason
> > > why complicate and not just start reclaim directly, just like
> > > memory.high does.
> >
> > Yes we can do better than we do right now. But that doesn't mean we
> > should put hacks all over the place and lie about the allocation
> > context.
>
> What do you mean by saying "all over the place"? It's a fix for kmem
> implementation, to be more exact for the part of it residing in the slab
> core.
I meant into two slab allocators currently because of the implementation
details which are spread into three different places - page allocator,
memcg charging code and the respective slab allocator specific details.
> Everyone else, except a couple of kmem users issuing alloc_page
> directly like threadinfo, will use kmalloc and know nothing what's going
> on there and how all this accounting stuff is handled - they will just
> use plain old convenient kmalloc, which works exactly as it does in the
> root cgroup.
If we ever grow more users and charge more kernel memory then they might
be doing similar assumptions and tweak allocation/charge context and we
would end up in a bigger mess. It makes much more sense to have
allocation and charge context consistent.
--
Michal Hocko
SUSE Labs
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web