Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1253466 > unrolled thread
| Started by | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| First post | 2015-10-22 06:30 +0200 |
| Last post | 2015-10-29 19:00 +0100 |
| Articles | 20 on this page of 41 — 4 participants |
Back to article view | Back to linux.kernel
[PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-22 06:30 +0200
[PATCH 1/8] mm: page_counter: let page_counter_try_charge() return bool Johannes Weiner <hannes@cmpxchg.org> - 2015-10-22 06:30 +0200
Re: [PATCH 1/8] mm: page_counter: let page_counter_try_charge() return bool Michal Hocko <mhocko@kernel.org> - 2015-10-23 13:40 +0200
[PATCH 6/8] mm: vmscan: simplify memcg vs. global shrinker invocation Johannes Weiner <hannes@cmpxchg.org> - 2015-10-22 06:30 +0200
Re: [PATCH 6/8] mm: vmscan: simplify memcg vs. global shrinker invocation Michal Hocko <mhocko@kernel.org> - 2015-10-23 15:30 +0200
[PATCH 4/8] mm: memcontrol: prepare for unified hierarchy socket accounting Johannes Weiner <hannes@cmpxchg.org> - 2015-10-22 06:30 +0200
Re: [PATCH 4/8] mm: memcontrol: prepare for unified hierarchy socket accounting Michal Hocko <mhocko@kernel.org> - 2015-10-23 14:40 +0200
[PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-22 06:30 +0200
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-10-22 20:50 +0200
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Michal Hocko <mhocko@kernel.org> - 2015-10-23 15:30 +0200
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy David Miller <davem@davemloft.net> - 2015-10-23 15:50 +0200
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-26 18:00 +0100
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Michal Hocko <mhocko@kernel.org> - 2015-10-27 13:30 +0100
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy David Miller <davem@davemloft.net> - 2015-10-27 14:40 +0100
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-27 16:50 +0100
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Michal Hocko <mhocko@kernel.org> - 2015-10-27 17:20 +0100
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-27 17:50 +0100
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy David Miller <davem@davemloft.net> - 2015-10-28 01:30 +0100
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-28 04:10 +0100
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Michal Hocko <mhocko@kernel.org> - 2015-10-29 16:30 +0100
Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-29 17:20 +0100
[PATCH 2/8] mm: memcontrol: export root_mem_cgroup Johannes Weiner <hannes@cmpxchg.org> - 2015-10-22 06:30 +0200
Re: [PATCH 2/8] mm: memcontrol: export root_mem_cgroup Michal Hocko <mhocko@kernel.org> - 2015-10-23 13:40 +0200
[PATCH 7/8] mm: vmscan: report vmpressure at the level of reclaim activity Johannes Weiner <hannes@cmpxchg.org> - 2015-10-22 06:30 +0200
Re: [PATCH 7/8] mm: vmscan: report vmpressure at the level of reclaim activity Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-10-22 20:50 +0200
Re: [PATCH 7/8] mm: vmscan: report vmpressure at the level of reclaim activity Michal Hocko <mhocko@kernel.org> - 2015-10-23 16:00 +0200
[PATCH 8/8] mm: memcontrol: hook up vmpressure to socket pressure Johannes Weiner <hannes@cmpxchg.org> - 2015-10-22 06:30 +0200
Re: [PATCH 8/8] mm: memcontrol: hook up vmpressure to socket pressure Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-10-22 21:00 +0200
[PATCH 3/8] net: consolidate memcg socket buffer tracking and accounting Johannes Weiner <hannes@cmpxchg.org> - 2015-10-22 06:30 +0200
Re: [PATCH 3/8] net: consolidate memcg socket buffer tracking and accounting Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-10-22 20:50 +0200
Re: [PATCH 3/8] net: consolidate memcg socket buffer tracking and accounting Johannes Weiner <hannes@cmpxchg.org> - 2015-10-22 21:20 +0200
Re: [PATCH 3/8] net: consolidate memcg socket buffer tracking and accounting Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-10-23 15:50 +0200
Re: [PATCH 3/8] net: consolidate memcg socket buffer tracking and accounting Michal Hocko <mhocko@kernel.org> - 2015-10-23 14:40 +0200
Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-10-22 20:50 +0200
Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-26 18:30 +0100
Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-10-27 09:50 +0100
Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-27 17:10 +0100
Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-10-28 09:30 +0100
Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-28 20:00 +0100
Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-10-29 10:30 +0100
Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy Johannes Weiner <hannes@cmpxchg.org> - 2015-10-29 19:00 +0100
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2015-10-29 17:20 +0100 |
| Subject | Re: [PATCH 5/8] mm: memcontrol: account socket memory on unified hierarchy |
| Message-ID | <qoY5I-2tW-11@gated-at.bofh.it> |
| In reply to | #1258867 |
On Thu, Oct 29, 2015 at 04:25:46PM +0100, Michal Hocko wrote: > On Tue 27-10-15 09:42:27, Johannes Weiner wrote: > > On Tue, Oct 27, 2015 at 05:15:54PM +0100, Michal Hocko wrote: > > > On Tue 27-10-15 11:41:38, Johannes Weiner wrote: > > > > IMO that's an implementation detail and a historical artifact that > > > > should not be exposed to the user. And that's the thing I hate about > > > > the current opt-out knob. > > > > You carefully skipped over this part. We can ignore it for socket > > memory but it's something we need to figure out when it comes to slab > > accounting and tracking. > > I am sorry, I didn't mean to skip this part, I though it would be clear > from the previous text. I think kmem accounting falls into the same > category. Have a sane default and a global boottime knob to override it > for those that think differently - for whatever reason they might have. Yes, that makes sense to me. Like cgroup.memory=nosocket, would you think it makes sense to include slab in the default for functional/semantical completeness and provide a cgroup.memory=noslab for powerusers? -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2015-10-22 06:30 +0200 |
| Subject | [PATCH 2/8] mm: memcontrol: export root_mem_cgroup |
| Message-ID | <qmfFL-87d-11@gated-at.bofh.it> |
| In reply to | #1253466 |
A later patch will need this symbol in files other than memcontrol.c,
so export it now and replace mem_cgroup_root_css at the same time.
Signed-off-by: Johannes Weiner <hannes@cmpxchg.org>
---
include/linux/memcontrol.h | 3 ++-
mm/backing-dev.c | 2 +-
mm/memcontrol.c | 5 ++---
3 files changed, 5 insertions(+), 5 deletions(-)
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index 805da1f..19ff87b 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -275,7 +275,8 @@ struct mem_cgroup {
struct mem_cgroup_per_node *nodeinfo[0];
/* WARNING: nodeinfo must be the last member here */
};
-extern struct cgroup_subsys_state *mem_cgroup_root_css;
+
+extern struct mem_cgroup *root_mem_cgroup;
/**
* mem_cgroup_events - count memory events against a cgroup
diff --git a/mm/backing-dev.c b/mm/backing-dev.c
index 095b23b..73ab967 100644
--- a/mm/backing-dev.c
+++ b/mm/backing-dev.c
@@ -702,7 +702,7 @@ static int cgwb_bdi_init(struct backing_dev_info *bdi)
ret = wb_init(&bdi->wb, bdi, 1, GFP_KERNEL);
if (!ret) {
- bdi->wb.memcg_css = mem_cgroup_root_css;
+ bdi->wb.memcg_css = &root_mem_cgroup->css;
bdi->wb.blkcg_css = blkcg_root_css;
}
return ret;
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index a8ccdbc..e54f434 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -76,9 +76,9 @@
struct cgroup_subsys memory_cgrp_subsys __read_mostly;
EXPORT_SYMBOL(memory_cgrp_subsys);
+struct mem_cgroup *root_mem_cgroup __read_mostly;
+
#define MEM_CGROUP_RECLAIM_RETRIES 5
-static struct mem_cgroup *root_mem_cgroup __read_mostly;
-struct cgroup_subsys_state *mem_cgroup_root_css __read_mostly;
/* Whether the swap controller is active */
#ifdef CONFIG_MEMCG_SWAP
@@ -4213,7 +4213,6 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *parent_css)
/* root ? */
if (parent_css == NULL) {
root_mem_cgroup = memcg;
- mem_cgroup_root_css = &memcg->css;
page_counter_init(&memcg->memory, NULL);
memcg->high = PAGE_COUNTER_MAX;
memcg->soft_limit = PAGE_COUNTER_MAX;
--
2.6.1
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2015-10-23 13:40 +0200 |
| Subject | Re: [PATCH 2/8] mm: memcontrol: export root_mem_cgroup |
| Message-ID | <qmIRt-9n-33@gated-at.bofh.it> |
| In reply to | #1253471 |
On Thu 22-10-15 00:21:30, Johannes Weiner wrote:
> A later patch will need this symbol in files other than memcontrol.c,
> so export it now and replace mem_cgroup_root_css at the same time.
>
> Signed-off-by: Johannes Weiner <hannes@cmpxchg.org>
Acked-by: Michal Hocko <mhocko@suse.com>
> ---
> include/linux/memcontrol.h | 3 ++-
> mm/backing-dev.c | 2 +-
> mm/memcontrol.c | 5 ++---
> 3 files changed, 5 insertions(+), 5 deletions(-)
>
> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
> index 805da1f..19ff87b 100644
> --- a/include/linux/memcontrol.h
> +++ b/include/linux/memcontrol.h
> @@ -275,7 +275,8 @@ struct mem_cgroup {
> struct mem_cgroup_per_node *nodeinfo[0];
> /* WARNING: nodeinfo must be the last member here */
> };
> -extern struct cgroup_subsys_state *mem_cgroup_root_css;
> +
> +extern struct mem_cgroup *root_mem_cgroup;
>
> /**
> * mem_cgroup_events - count memory events against a cgroup
> diff --git a/mm/backing-dev.c b/mm/backing-dev.c
> index 095b23b..73ab967 100644
> --- a/mm/backing-dev.c
> +++ b/mm/backing-dev.c
> @@ -702,7 +702,7 @@ static int cgwb_bdi_init(struct backing_dev_info *bdi)
>
> ret = wb_init(&bdi->wb, bdi, 1, GFP_KERNEL);
> if (!ret) {
> - bdi->wb.memcg_css = mem_cgroup_root_css;
> + bdi->wb.memcg_css = &root_mem_cgroup->css;
> bdi->wb.blkcg_css = blkcg_root_css;
> }
> return ret;
> diff --git a/mm/memcontrol.c b/mm/memcontrol.c
> index a8ccdbc..e54f434 100644
> --- a/mm/memcontrol.c
> +++ b/mm/memcontrol.c
> @@ -76,9 +76,9 @@
> struct cgroup_subsys memory_cgrp_subsys __read_mostly;
> EXPORT_SYMBOL(memory_cgrp_subsys);
>
> +struct mem_cgroup *root_mem_cgroup __read_mostly;
> +
> #define MEM_CGROUP_RECLAIM_RETRIES 5
> -static struct mem_cgroup *root_mem_cgroup __read_mostly;
> -struct cgroup_subsys_state *mem_cgroup_root_css __read_mostly;
>
> /* Whether the swap controller is active */
> #ifdef CONFIG_MEMCG_SWAP
> @@ -4213,7 +4213,6 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *parent_css)
> /* root ? */
> if (parent_css == NULL) {
> root_mem_cgroup = memcg;
> - mem_cgroup_root_css = &memcg->css;
> page_counter_init(&memcg->memory, NULL);
> memcg->high = PAGE_COUNTER_MAX;
> memcg->soft_limit = PAGE_COUNTER_MAX;
> --
> 2.6.1
--
Michal Hocko
SUSE Labs
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2015-10-22 06:30 +0200 |
| Subject | [PATCH 7/8] mm: vmscan: report vmpressure at the level of reclaim activity |
| Message-ID | <qmfFM-87d-13@gated-at.bofh.it> |
| In reply to | #1253466 |
The vmpressure metric is based on reclaim efficiency, which in turn is
an attribute of the LRU. However, vmpressure events are currently
reported at the source of pressure rather than at the reclaim level.
Switch the reporting to the reclaim level to allow finer-grained
analysis of which memcg is having trouble reclaiming its pages.
As far as memory.pressure_level interface semantics go, events are
escalated up the hierarchy until a listener is found, so this won't
affect existing users that listen at higher levels.
This also prepares vmpressure for hooking it up to the networking
stack's memory pressure code.
Signed-off-by: Johannes Weiner <hannes@cmpxchg.org>
---
mm/vmscan.c | 10 ++++++----
1 file changed, 6 insertions(+), 4 deletions(-)
diff --git a/mm/vmscan.c b/mm/vmscan.c
index ecc2125..50630c8 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -2404,6 +2404,7 @@ static bool shrink_zone(struct zone *zone, struct scan_control *sc,
memcg = mem_cgroup_iter(root, NULL, &reclaim);
do {
unsigned long lru_pages;
+ unsigned long reclaimed;
unsigned long scanned;
struct lruvec *lruvec;
int swappiness;
@@ -2416,6 +2417,7 @@ static bool shrink_zone(struct zone *zone, struct scan_control *sc,
lruvec = mem_cgroup_zone_lruvec(zone, memcg);
swappiness = mem_cgroup_swappiness(memcg);
+ reclaimed = sc->nr_reclaimed;
scanned = sc->nr_scanned;
shrink_lruvec(lruvec, swappiness, sc, &lru_pages);
@@ -2437,6 +2439,10 @@ static bool shrink_zone(struct zone *zone, struct scan_control *sc,
}
}
+ vmpressure(sc->gfp_mask, memcg,
+ sc->nr_scanned - scanned,
+ sc->nr_reclaimed - reclaimed);
+
/*
* Direct reclaim and kswapd have to scan all memory
* cgroups to fulfill the overall scan target for the
@@ -2454,10 +2460,6 @@ static bool shrink_zone(struct zone *zone, struct scan_control *sc,
}
} while ((memcg = mem_cgroup_iter(root, memcg, &reclaim)));
- vmpressure(sc->gfp_mask, sc->target_mem_cgroup,
- sc->nr_scanned - nr_scanned,
- sc->nr_reclaimed - nr_reclaimed);
-
if (sc->nr_reclaimed - nr_reclaimed)
reclaimable = true;
--
2.6.1
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@virtuozzo.com> |
|---|---|
| Date | 2015-10-22 20:50 +0200 |
| Subject | Re: [PATCH 7/8] mm: vmscan: report vmpressure at the level of reclaim activity |
| Message-ID | <qmt62-2wa-21@gated-at.bofh.it> |
| In reply to | #1253472 |
On Thu, Oct 22, 2015 at 12:21:35AM -0400, Johannes Weiner wrote: ... > @@ -2437,6 +2439,10 @@ static bool shrink_zone(struct zone *zone, struct scan_control *sc, > } > } > > + vmpressure(sc->gfp_mask, memcg, > + sc->nr_scanned - scanned, > + sc->nr_reclaimed - reclaimed); > + > /* > * Direct reclaim and kswapd have to scan all memory > * cgroups to fulfill the overall scan target for the > @@ -2454,10 +2460,6 @@ static bool shrink_zone(struct zone *zone, struct scan_control *sc, > } > } while ((memcg = mem_cgroup_iter(root, memcg, &reclaim))); > > - vmpressure(sc->gfp_mask, sc->target_mem_cgroup, > - sc->nr_scanned - nr_scanned, > - sc->nr_reclaimed - nr_reclaimed); > - > if (sc->nr_reclaimed - nr_reclaimed) > reclaimable = true; > I may be mistaken, but AFAIU this patch subtly changes the behavior of vmpressure visible from the userspace: w/o this patch a userspace process will only receive a notification for a memory cgroup only if *this* memory cgroup calls reclaimer; with this patch userspace notification will be issued even if reclaimer is invoked by any cgroup up the hierarchy. Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2015-10-23 16:00 +0200 |
| Subject | Re: [PATCH 7/8] mm: vmscan: report vmpressure at the level of reclaim activity |
| Message-ID | <qmL2W-3e7-13@gated-at.bofh.it> |
| In reply to | #1253472 |
On Thu 22-10-15 00:21:35, Johannes Weiner wrote: > The vmpressure metric is based on reclaim efficiency, which in turn is > an attribute of the LRU. However, vmpressure events are currently > reported at the source of pressure rather than at the reclaim level. > > Switch the reporting to the reclaim level to allow finer-grained > analysis of which memcg is having trouble reclaiming its pages. I can see how this can be useful. > As far as memory.pressure_level interface semantics go, events are > escalated up the hierarchy until a listener is found, so this won't > affect existing users that listen at higher levels. This is true but the parent will not see cumulative events anymore. One memcg might be fighting and barely reclaim anything so it would report high pressure while other would be doing just fine. The parent will just see conflicting events in a short time period and cannot match them the source memcg. This sounds really confusing. Even more confusing than the current semantic which allows the same behavior under certain configurations. I dunno, have to think about it some more. Maybe we need to rethink the way how the pressure is signaled. If we want the breakdown of the particular memcgs then we should be able to identify them for this to be useful. [...] -- Michal Hocko SUSE Labs -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2015-10-22 06:30 +0200 |
| Subject | [PATCH 8/8] mm: memcontrol: hook up vmpressure to socket pressure |
| Message-ID | <qmfFM-87d-15@gated-at.bofh.it> |
| In reply to | #1253466 |
Let the networking stack know when a memcg is under reclaim pressure,
so it can shrink its transmit windows accordingly.
Whenever the reclaim efficiency of a memcg's LRU lists drops low
enough for a MEDIUM or HIGH vmpressure event to occur, assert a
pressure state in the socket and tcp memory code that tells it to
reduce memory usage in sockets associated with said memory cgroup.
vmpressure events are edge triggered, so for hysteresis assert socket
pressure for a second to allow for subsequent vmpressure events to
occur before letting the socket code return to normal.
Signed-off-by: Johannes Weiner <hannes@cmpxchg.org>
---
include/linux/memcontrol.h | 9 +++++++++
include/net/sock.h | 4 ++++
include/net/tcp.h | 4 ++++
mm/memcontrol.c | 1 +
mm/vmpressure.c | 29 ++++++++++++++++++++++++-----
5 files changed, 42 insertions(+), 5 deletions(-)
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index d66ae18..b9990f7 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -246,6 +246,7 @@ struct mem_cgroup {
#ifdef CONFIG_INET
struct work_struct socket_work;
+ unsigned long socket_pressure;
#endif
/* List of events which userspace want to receive */
@@ -696,6 +697,10 @@ void sock_update_memcg(struct sock *sk);
void sock_release_memcg(struct sock *sk);
bool mem_cgroup_charge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages);
void mem_cgroup_uncharge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages);
+static inline bool mem_cgroup_socket_pressure(struct mem_cgroup *memcg)
+{
+ return time_before(jiffies, memcg->socket_pressure);
+}
#else
static inline bool mem_cgroup_do_sockets(void)
{
@@ -716,6 +721,10 @@ static inline void mem_cgroup_uncharge_skmem(struct mem_cgroup *memcg,
unsigned int nr_pages)
{
}
+static inline bool mem_cgroup_socket_pressure(struct mem_cgroup *memcg)
+{
+ return false;
+}
#endif /* CONFIG_INET */
#ifdef CONFIG_MEMCG_KMEM
diff --git a/include/net/sock.h b/include/net/sock.h
index 67795fc..22bfb9c 100644
--- a/include/net/sock.h
+++ b/include/net/sock.h
@@ -1087,6 +1087,10 @@ static inline bool sk_has_memory_pressure(const struct sock *sk)
static inline bool sk_under_memory_pressure(const struct sock *sk)
{
+ if (mem_cgroup_do_sockets() && sk->sk_memcg &&
+ mem_cgroup_socket_pressure(sk->sk_memcg))
+ return true;
+
if (!sk->sk_prot->memory_pressure)
return false;
diff --git a/include/net/tcp.h b/include/net/tcp.h
index 77b6c7e..c7d342c 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -291,6 +291,10 @@ extern int tcp_memory_pressure;
/* optimized version of sk_under_memory_pressure() for TCP sockets */
static inline bool tcp_under_memory_pressure(const struct sock *sk)
{
+ if (mem_cgroup_do_sockets() && sk->sk_memcg &&
+ mem_cgroup_socket_pressure(sk->sk_memcg))
+ return true;
+
return tcp_memory_pressure;
}
/*
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index cb1d6aa..2e09def 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -4178,6 +4178,7 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *parent_css)
#endif
#ifdef CONFIG_INET
INIT_WORK(&memcg->socket_work, socket_work_func);
+ memcg->socket_pressure = jiffies;
#endif
return &memcg->css;
diff --git a/mm/vmpressure.c b/mm/vmpressure.c
index 4c25e62..f64c0e1 100644
--- a/mm/vmpressure.c
+++ b/mm/vmpressure.c
@@ -137,14 +137,11 @@ struct vmpressure_event {
};
static bool vmpressure_event(struct vmpressure *vmpr,
- unsigned long scanned, unsigned long reclaimed)
+ enum vmpressure_levels level)
{
struct vmpressure_event *ev;
- enum vmpressure_levels level;
bool signalled = false;
- level = vmpressure_calc_level(scanned, reclaimed);
-
mutex_lock(&vmpr->events_lock);
list_for_each_entry(ev, &vmpr->events, node) {
@@ -162,6 +159,7 @@ static bool vmpressure_event(struct vmpressure *vmpr,
static void vmpressure_work_fn(struct work_struct *work)
{
struct vmpressure *vmpr = work_to_vmpressure(work);
+ enum vmpressure_levels level;
unsigned long scanned;
unsigned long reclaimed;
@@ -185,8 +183,29 @@ static void vmpressure_work_fn(struct work_struct *work)
vmpr->reclaimed = 0;
spin_unlock(&vmpr->sr_lock);
+ level = vmpressure_calc_level(scanned, reclaimed);
+
+ if (level > VMPRESSURE_LOW) {
+ struct mem_cgroup *memcg;
+ /*
+ * Let the socket buffer allocator know that we are
+ * having trouble reclaiming LRU pages.
+ *
+ * For hysteresis, keep the pressure state asserted
+ * for a second in which subsequent pressure events
+ * can occur.
+ *
+ * XXX: is vmpressure a global feature or part of
+ * memcg? There shouldn't be anything memcg-specific
+ * about exporting reclaim success ratios from the VM.
+ */
+ memcg = container_of(vmpr, struct mem_cgroup, vmpressure);
+ if (memcg != root_mem_cgroup)
+ memcg->socket_pressure = jiffies + HZ;
+ }
+
do {
- if (vmpressure_event(vmpr, scanned, reclaimed))
+ if (vmpressure_event(vmpr, level))
break;
/*
* If not handled, propagate the event upward into the
--
2.6.1
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@virtuozzo.com> |
|---|---|
| Date | 2015-10-22 21:00 +0200 |
| Subject | Re: [PATCH 8/8] mm: memcontrol: hook up vmpressure to socket pressure |
| Message-ID | <qmtfI-2HW-23@gated-at.bofh.it> |
| In reply to | #1253473 |
On Thu, Oct 22, 2015 at 12:21:36AM -0400, Johannes Weiner wrote:
...
> @@ -185,8 +183,29 @@ static void vmpressure_work_fn(struct work_struct *work)
> vmpr->reclaimed = 0;
> spin_unlock(&vmpr->sr_lock);
>
> + level = vmpressure_calc_level(scanned, reclaimed);
> +
> + if (level > VMPRESSURE_LOW) {
So we start socket_pressure at MEDIUM. Why not at LOW or CRITICAL?
> + struct mem_cgroup *memcg;
> + /*
> + * Let the socket buffer allocator know that we are
> + * having trouble reclaiming LRU pages.
> + *
> + * For hysteresis, keep the pressure state asserted
> + * for a second in which subsequent pressure events
> + * can occur.
> + *
> + * XXX: is vmpressure a global feature or part of
> + * memcg? There shouldn't be anything memcg-specific
> + * about exporting reclaim success ratios from the VM.
> + */
> + memcg = container_of(vmpr, struct mem_cgroup, vmpressure);
> + if (memcg != root_mem_cgroup)
> + memcg->socket_pressure = jiffies + HZ;
Why 1 second?
Thanks,
Vladimir
> + }
> +
> do {
> - if (vmpressure_event(vmpr, scanned, reclaimed))
> + if (vmpressure_event(vmpr, level))
> break;
> /*
> * If not handled, propagate the event upward into the
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2015-10-22 06:30 +0200 |
| Subject | [PATCH 3/8] net: consolidate memcg socket buffer tracking and accounting |
| Message-ID | <qmfFM-87d-21@gated-at.bofh.it> |
| In reply to | #1253466 |
The tcp memory controller has extensive provisions for future memory
accounting interfaces that won't materialize after all. Cut the code
base down to what's actually used, now and in the likely future.
- There won't be any different protocol counters in the future, so a
direct sock->sk_memcg linkage is enough. This eliminates a lot of
callback maze and boilerplate code, and restores most of the socket
allocation code to pre-tcp_memcontrol state.
- There won't be a tcp control soft limit, so integrating the memcg
code into the global skmem limiting scheme complicates things
unnecessarily. Replace all that with simple and clear charge and
uncharge calls--hidden behind a jump label--to account skb memory.
- The previous jump label code was an elaborate state machine that
tracked the number of cgroups with an active socket limit in order
to enable the skmem tracking and accounting code only when actively
necessary. But this is overengineered: it was meant to protect the
people who never use this feature in the first place. Simply enable
the branches once when the first limit is set until the next reboot.
Signed-off-by: Johannes Weiner <hannes@cmpxchg.org>
---
include/linux/memcontrol.h | 64 ++++++++-----------
include/net/sock.h | 135 +++------------------------------------
include/net/tcp.h | 3 -
include/net/tcp_memcontrol.h | 7 ---
mm/memcontrol.c | 101 +++++++++++++++--------------
net/core/sock.c | 78 ++++++-----------------
net/ipv4/sysctl_net_ipv4.c | 1 -
net/ipv4/tcp.c | 3 +-
net/ipv4/tcp_ipv4.c | 9 +--
net/ipv4/tcp_memcontrol.c | 147 +++++++------------------------------------
net/ipv4/tcp_output.c | 6 +-
net/ipv6/tcp_ipv6.c | 3 -
12 files changed, 136 insertions(+), 421 deletions(-)
delete mode 100644 include/net/tcp_memcontrol.h
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index 19ff87b..5b72f83 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -85,34 +85,6 @@ enum mem_cgroup_events_target {
MEM_CGROUP_NTARGETS,
};
-/*
- * Bits in struct cg_proto.flags
- */
-enum cg_proto_flags {
- /* Currently active and new sockets should be assigned to cgroups */
- MEMCG_SOCK_ACTIVE,
- /* It was ever activated; we must disarm static keys on destruction */
- MEMCG_SOCK_ACTIVATED,
-};
-
-struct cg_proto {
- struct page_counter memory_allocated; /* Current allocated memory. */
- struct percpu_counter sockets_allocated; /* Current number of sockets. */
- int memory_pressure;
- long sysctl_mem[3];
- unsigned long flags;
- /*
- * memcg field is used to find which memcg we belong directly
- * Each memcg struct can hold more than one cg_proto, so container_of
- * won't really cut.
- *
- * The elegant solution would be having an inverse function to
- * proto_cgroup in struct proto, but that means polluting the structure
- * for everybody, instead of just for memcg users.
- */
- struct mem_cgroup *memcg;
-};
-
#ifdef CONFIG_MEMCG
struct mem_cgroup_stat_cpu {
long count[MEM_CGROUP_STAT_NSTATS];
@@ -185,8 +157,15 @@ struct mem_cgroup {
/* Accounted resources */
struct page_counter memory;
+
+ /*
+ * Legacy non-resource counters. In unified hierarchy, all
+ * memory is accounted and limited through memcg->memory.
+ * Consumer breakdown happens in the statistics.
+ */
struct page_counter memsw;
struct page_counter kmem;
+ struct page_counter skmem;
/* Normal memory consumption range */
unsigned long low;
@@ -246,9 +225,6 @@ struct mem_cgroup {
*/
struct mem_cgroup_stat_cpu __percpu *stat;
-#if defined(CONFIG_MEMCG_KMEM) && defined(CONFIG_INET)
- struct cg_proto tcp_mem;
-#endif
#if defined(CONFIG_MEMCG_KMEM)
/* Index in the kmem_cache->memcg_params.memcg_caches array */
int kmemcg_id;
@@ -676,12 +652,6 @@ void mem_cgroup_count_vm_event(struct mm_struct *mm, enum vm_event_item idx)
}
#endif /* CONFIG_MEMCG */
-enum {
- UNDER_LIMIT,
- SOFT_LIMIT,
- OVER_LIMIT,
-};
-
#ifdef CONFIG_CGROUP_WRITEBACK
struct list_head *mem_cgroup_cgwb_list(struct mem_cgroup *memcg);
@@ -707,15 +677,35 @@ static inline void mem_cgroup_wb_stats(struct bdi_writeback *wb,
struct sock;
#if defined(CONFIG_INET) && defined(CONFIG_MEMCG_KMEM)
+extern struct static_key_false mem_cgroup_sockets;
+static inline bool mem_cgroup_do_sockets(void)
+{
+ return static_branch_unlikely(&mem_cgroup_sockets);
+}
void sock_update_memcg(struct sock *sk);
void sock_release_memcg(struct sock *sk);
+bool mem_cgroup_charge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages);
+void mem_cgroup_uncharge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages);
#else
+static inline bool mem_cgroup_do_sockets(void)
+{
+ return false;
+}
static inline void sock_update_memcg(struct sock *sk)
{
}
static inline void sock_release_memcg(struct sock *sk)
{
}
+static inline bool mem_cgroup_charge_skmem(struct mem_cgroup *memcg,
+ unsigned int nr_pages)
+{
+ return true;
+}
+static inline void mem_cgroup_uncharge_skmem(struct mem_cgroup *memcg,
+ unsigned int nr_pages)
+{
+}
#endif /* CONFIG_INET && CONFIG_MEMCG_KMEM */
#ifdef CONFIG_MEMCG_KMEM
diff --git a/include/net/sock.h b/include/net/sock.h
index 59a7196..67795fc 100644
--- a/include/net/sock.h
+++ b/include/net/sock.h
@@ -69,22 +69,6 @@
#include <net/tcp_states.h>
#include <linux/net_tstamp.h>
-struct cgroup;
-struct cgroup_subsys;
-#ifdef CONFIG_NET
-int mem_cgroup_sockets_init(struct mem_cgroup *memcg, struct cgroup_subsys *ss);
-void mem_cgroup_sockets_destroy(struct mem_cgroup *memcg);
-#else
-static inline
-int mem_cgroup_sockets_init(struct mem_cgroup *memcg, struct cgroup_subsys *ss)
-{
- return 0;
-}
-static inline
-void mem_cgroup_sockets_destroy(struct mem_cgroup *memcg)
-{
-}
-#endif
/*
* This structure really needs to be cleaned up.
* Most of it is for TCP, and not used by any of
@@ -243,7 +227,6 @@ struct sock_common {
/* public: */
};
-struct cg_proto;
/**
* struct sock - network layer representation of sockets
* @__sk_common: shared layout with inet_timewait_sock
@@ -310,7 +293,7 @@ struct cg_proto;
* @sk_security: used by security modules
* @sk_mark: generic packet mark
* @sk_classid: this socket's cgroup classid
- * @sk_cgrp: this socket's cgroup-specific proto data
+ * @sk_memcg: this socket's memcg association
* @sk_write_pending: a write to stream socket waits to start
* @sk_state_change: callback to indicate change in the state of the sock
* @sk_data_ready: callback to indicate there is data to be processed
@@ -447,7 +430,7 @@ struct sock {
#ifdef CONFIG_CGROUP_NET_CLASSID
u32 sk_classid;
#endif
- struct cg_proto *sk_cgrp;
+ struct mem_cgroup *sk_memcg;
void (*sk_state_change)(struct sock *sk);
void (*sk_data_ready)(struct sock *sk);
void (*sk_write_space)(struct sock *sk);
@@ -1051,18 +1034,6 @@ struct proto {
#ifdef SOCK_REFCNT_DEBUG
atomic_t socks;
#endif
-#ifdef CONFIG_MEMCG_KMEM
- /*
- * cgroup specific init/deinit functions. Called once for all
- * protocols that implement it, from cgroups populate function.
- * This function has to setup any files the protocol want to
- * appear in the kmem cgroup filesystem.
- */
- int (*init_cgroup)(struct mem_cgroup *memcg,
- struct cgroup_subsys *ss);
- void (*destroy_cgroup)(struct mem_cgroup *memcg);
- struct cg_proto *(*proto_cgroup)(struct mem_cgroup *memcg);
-#endif
};
int proto_register(struct proto *prot, int alloc_slab);
@@ -1093,23 +1064,6 @@ static inline void sk_refcnt_debug_release(const struct sock *sk)
#define sk_refcnt_debug_release(sk) do { } while (0)
#endif /* SOCK_REFCNT_DEBUG */
-#if defined(CONFIG_MEMCG_KMEM) && defined(CONFIG_NET)
-extern struct static_key memcg_socket_limit_enabled;
-static inline struct cg_proto *parent_cg_proto(struct proto *proto,
- struct cg_proto *cg_proto)
-{
- return proto->proto_cgroup(parent_mem_cgroup(cg_proto->memcg));
-}
-#define mem_cgroup_sockets_enabled static_key_false(&memcg_socket_limit_enabled)
-#else
-#define mem_cgroup_sockets_enabled 0
-static inline struct cg_proto *parent_cg_proto(struct proto *proto,
- struct cg_proto *cg_proto)
-{
- return NULL;
-}
-#endif
-
static inline bool sk_stream_memory_free(const struct sock *sk)
{
if (sk->sk_wmem_queued >= sk->sk_sndbuf)
@@ -1136,9 +1090,6 @@ static inline bool sk_under_memory_pressure(const struct sock *sk)
if (!sk->sk_prot->memory_pressure)
return false;
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
- return !!sk->sk_cgrp->memory_pressure;
-
return !!*sk->sk_prot->memory_pressure;
}
@@ -1146,61 +1097,19 @@ static inline void sk_leave_memory_pressure(struct sock *sk)
{
int *memory_pressure = sk->sk_prot->memory_pressure;
- if (!memory_pressure)
- return;
-
- if (*memory_pressure)
+ if (memory_pressure && *memory_pressure)
*memory_pressure = 0;
-
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
- struct cg_proto *cg_proto = sk->sk_cgrp;
- struct proto *prot = sk->sk_prot;
-
- for (; cg_proto; cg_proto = parent_cg_proto(prot, cg_proto))
- cg_proto->memory_pressure = 0;
- }
-
}
static inline void sk_enter_memory_pressure(struct sock *sk)
{
- if (!sk->sk_prot->enter_memory_pressure)
- return;
-
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
- struct cg_proto *cg_proto = sk->sk_cgrp;
- struct proto *prot = sk->sk_prot;
-
- for (; cg_proto; cg_proto = parent_cg_proto(prot, cg_proto))
- cg_proto->memory_pressure = 1;
- }
-
- sk->sk_prot->enter_memory_pressure(sk);
+ if (sk->sk_prot->enter_memory_pressure)
+ sk->sk_prot->enter_memory_pressure(sk);
}
static inline long sk_prot_mem_limits(const struct sock *sk, int index)
{
- long *prot = sk->sk_prot->sysctl_mem;
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
- prot = sk->sk_cgrp->sysctl_mem;
- return prot[index];
-}
-
-static inline void memcg_memory_allocated_add(struct cg_proto *prot,
- unsigned long amt,
- int *parent_status)
-{
- page_counter_charge(&prot->memory_allocated, amt);
-
- if (page_counter_read(&prot->memory_allocated) >
- prot->memory_allocated.limit)
- *parent_status = OVER_LIMIT;
-}
-
-static inline void memcg_memory_allocated_sub(struct cg_proto *prot,
- unsigned long amt)
-{
- page_counter_uncharge(&prot->memory_allocated, amt);
+ return sk->sk_prot->sysctl_mem[index];
}
static inline long
@@ -1208,24 +1117,14 @@ sk_memory_allocated(const struct sock *sk)
{
struct proto *prot = sk->sk_prot;
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
- return page_counter_read(&sk->sk_cgrp->memory_allocated);
-
return atomic_long_read(prot->memory_allocated);
}
static inline long
-sk_memory_allocated_add(struct sock *sk, int amt, int *parent_status)
+sk_memory_allocated_add(struct sock *sk, int amt)
{
struct proto *prot = sk->sk_prot;
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
- memcg_memory_allocated_add(sk->sk_cgrp, amt, parent_status);
- /* update the root cgroup regardless */
- atomic_long_add_return(amt, prot->memory_allocated);
- return page_counter_read(&sk->sk_cgrp->memory_allocated);
- }
-
return atomic_long_add_return(amt, prot->memory_allocated);
}
@@ -1234,9 +1133,6 @@ sk_memory_allocated_sub(struct sock *sk, int amt)
{
struct proto *prot = sk->sk_prot;
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
- memcg_memory_allocated_sub(sk->sk_cgrp, amt);
-
atomic_long_sub(amt, prot->memory_allocated);
}
@@ -1244,13 +1140,6 @@ static inline void sk_sockets_allocated_dec(struct sock *sk)
{
struct proto *prot = sk->sk_prot;
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
- struct cg_proto *cg_proto = sk->sk_cgrp;
-
- for (; cg_proto; cg_proto = parent_cg_proto(prot, cg_proto))
- percpu_counter_dec(&cg_proto->sockets_allocated);
- }
-
percpu_counter_dec(prot->sockets_allocated);
}
@@ -1258,13 +1147,6 @@ static inline void sk_sockets_allocated_inc(struct sock *sk)
{
struct proto *prot = sk->sk_prot;
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
- struct cg_proto *cg_proto = sk->sk_cgrp;
-
- for (; cg_proto; cg_proto = parent_cg_proto(prot, cg_proto))
- percpu_counter_inc(&cg_proto->sockets_allocated);
- }
-
percpu_counter_inc(prot->sockets_allocated);
}
@@ -1273,9 +1155,6 @@ sk_sockets_allocated_read_positive(struct sock *sk)
{
struct proto *prot = sk->sk_prot;
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
- return percpu_counter_read_positive(&sk->sk_cgrp->sockets_allocated);
-
return percpu_counter_read_positive(prot->sockets_allocated);
}
diff --git a/include/net/tcp.h b/include/net/tcp.h
index eed94fc..77b6c7e 100644
--- a/include/net/tcp.h
+++ b/include/net/tcp.h
@@ -291,9 +291,6 @@ extern int tcp_memory_pressure;
/* optimized version of sk_under_memory_pressure() for TCP sockets */
static inline bool tcp_under_memory_pressure(const struct sock *sk)
{
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
- return !!sk->sk_cgrp->memory_pressure;
-
return tcp_memory_pressure;
}
/*
diff --git a/include/net/tcp_memcontrol.h b/include/net/tcp_memcontrol.h
deleted file mode 100644
index 05b94d9..0000000
--- a/include/net/tcp_memcontrol.h
+++ /dev/null
@@ -1,7 +0,0 @@
-#ifndef _TCP_MEMCG_H
-#define _TCP_MEMCG_H
-
-struct cg_proto *tcp_proto_cgroup(struct mem_cgroup *memcg);
-int tcp_init_cgroup(struct mem_cgroup *memcg, struct cgroup_subsys *ss);
-void tcp_destroy_cgroup(struct mem_cgroup *memcg);
-#endif /* _TCP_MEMCG_H */
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index e54f434..c41e6d7 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -66,7 +66,6 @@
#include "internal.h"
#include <net/sock.h>
#include <net/ip.h>
-#include <net/tcp_memcontrol.h>
#include "slab.h"
#include <asm/uaccess.h>
@@ -291,58 +290,68 @@ static inline struct mem_cgroup *mem_cgroup_from_id(unsigned short id)
/* Writing them here to avoid exposing memcg's inner layout */
#if defined(CONFIG_INET) && defined(CONFIG_MEMCG_KMEM)
+DEFINE_STATIC_KEY_FALSE(mem_cgroup_sockets);
+
void sock_update_memcg(struct sock *sk)
{
- if (mem_cgroup_sockets_enabled) {
- struct mem_cgroup *memcg;
- struct cg_proto *cg_proto;
-
- BUG_ON(!sk->sk_prot->proto_cgroup);
-
- /* Socket cloning can throw us here with sk_cgrp already
- * filled. It won't however, necessarily happen from
- * process context. So the test for root memcg given
- * the current task's memcg won't help us in this case.
- *
- * Respecting the original socket's memcg is a better
- * decision in this case.
- */
- if (sk->sk_cgrp) {
- BUG_ON(mem_cgroup_is_root(sk->sk_cgrp->memcg));
- css_get(&sk->sk_cgrp->memcg->css);
- return;
- }
-
- rcu_read_lock();
- memcg = mem_cgroup_from_task(current);
- cg_proto = sk->sk_prot->proto_cgroup(memcg);
- if (cg_proto && test_bit(MEMCG_SOCK_ACTIVE, &cg_proto->flags) &&
- css_tryget_online(&memcg->css)) {
- sk->sk_cgrp = cg_proto;
- }
- rcu_read_unlock();
+ struct mem_cgroup *memcg;
+ /*
+ * Socket cloning can throw us here with sk_cgrp already
+ * filled. It won't however, necessarily happen from
+ * process context. So the test for root memcg given
+ * the current task's memcg won't help us in this case.
+ *
+ * Respecting the original socket's memcg is a better
+ * decision in this case.
+ */
+ if (sk->sk_memcg) {
+ BUG_ON(mem_cgroup_is_root(sk->sk_memcg));
+ css_get(&sk->sk_memcg->css);
+ return;
}
+
+ rcu_read_lock();
+ memcg = mem_cgroup_from_task(current);
+ if (css_tryget_online(&memcg->css))
+ sk->sk_memcg = memcg;
+ rcu_read_unlock();
}
EXPORT_SYMBOL(sock_update_memcg);
void sock_release_memcg(struct sock *sk)
{
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
- struct mem_cgroup *memcg;
- WARN_ON(!sk->sk_cgrp->memcg);
- memcg = sk->sk_cgrp->memcg;
- css_put(&sk->sk_cgrp->memcg->css);
- }
+ if (sk->sk_memcg)
+ css_put(&sk->sk_memcg->css);
}
-struct cg_proto *tcp_proto_cgroup(struct mem_cgroup *memcg)
+/**
+ * mem_cgroup_charge_skmem - charge socket memory
+ * @memcg: memcg to charge
+ * @nr_pages: number of pages to charge
+ *
+ * Charges @nr_pages to @memcg. Returns %true if the charge fit within
+ * the memcg's configured limit, %false if the charge had to be forced.
+ */
+bool mem_cgroup_charge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages)
{
- if (!memcg || mem_cgroup_is_root(memcg))
- return NULL;
+ struct page_counter *counter;
+
+ if (page_counter_try_charge(&memcg->skmem, nr_pages, &counter))
+ return true;
- return &memcg->tcp_mem;
+ page_counter_charge(&memcg->skmem, nr_pages);
+ return false;
+}
+
+/**
+ * mem_cgroup_uncharge_skmem - uncharge socket memory
+ * @memcg: memcg to uncharge
+ * @nr_pages: number of pages to uncharge
+ */
+void mem_cgroup_uncharge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages)
+{
+ page_counter_uncharge(&memcg->skmem, nr_pages);
}
-EXPORT_SYMBOL(tcp_proto_cgroup);
#endif
@@ -3592,13 +3601,7 @@ static int mem_cgroup_oom_control_write(struct cgroup_subsys_state *css,
#ifdef CONFIG_MEMCG_KMEM
static int memcg_init_kmem(struct mem_cgroup *memcg, struct cgroup_subsys *ss)
{
- int ret;
-
- ret = memcg_propagate_kmem(memcg);
- if (ret)
- return ret;
-
- return mem_cgroup_sockets_init(memcg, ss);
+ return memcg_propagate_kmem(memcg);
}
static void memcg_deactivate_kmem(struct mem_cgroup *memcg)
@@ -3654,7 +3657,6 @@ static void memcg_destroy_kmem(struct mem_cgroup *memcg)
static_key_slow_dec(&memcg_kmem_enabled_key);
WARN_ON(page_counter_read(&memcg->kmem));
}
- mem_cgroup_sockets_destroy(memcg);
}
#else
static int memcg_init_kmem(struct mem_cgroup *memcg, struct cgroup_subsys *ss)
@@ -4218,6 +4220,7 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *parent_css)
memcg->soft_limit = PAGE_COUNTER_MAX;
page_counter_init(&memcg->memsw, NULL);
page_counter_init(&memcg->kmem, NULL);
+ page_counter_init(&memcg->skmem, NULL);
}
memcg->last_scanned_node = MAX_NUMNODES;
@@ -4266,6 +4269,7 @@ mem_cgroup_css_online(struct cgroup_subsys_state *css)
memcg->soft_limit = PAGE_COUNTER_MAX;
page_counter_init(&memcg->memsw, &parent->memsw);
page_counter_init(&memcg->kmem, &parent->kmem);
+ page_counter_init(&memcg->skmem, &parent->skmem);
/*
* No need to take a reference to the parent because cgroup
@@ -4277,6 +4281,7 @@ mem_cgroup_css_online(struct cgroup_subsys_state *css)
memcg->soft_limit = PAGE_COUNTER_MAX;
page_counter_init(&memcg->memsw, NULL);
page_counter_init(&memcg->kmem, NULL);
+ page_counter_init(&memcg->skmem, NULL);
/*
* Deeper hierachy with use_hierarchy == false doesn't make
* much sense so let cgroup subsystem know about this
diff --git a/net/core/sock.c b/net/core/sock.c
index 0fafd27..0debff5 100644
--- a/net/core/sock.c
+++ b/net/core/sock.c
@@ -194,44 +194,6 @@ bool sk_net_capable(const struct sock *sk, int cap)
}
EXPORT_SYMBOL(sk_net_capable);
-
-#ifdef CONFIG_MEMCG_KMEM
-int mem_cgroup_sockets_init(struct mem_cgroup *memcg, struct cgroup_subsys *ss)
-{
- struct proto *proto;
- int ret = 0;
-
- mutex_lock(&proto_list_mutex);
- list_for_each_entry(proto, &proto_list, node) {
- if (proto->init_cgroup) {
- ret = proto->init_cgroup(memcg, ss);
- if (ret)
- goto out;
- }
- }
-
- mutex_unlock(&proto_list_mutex);
- return ret;
-out:
- list_for_each_entry_continue_reverse(proto, &proto_list, node)
- if (proto->destroy_cgroup)
- proto->destroy_cgroup(memcg);
- mutex_unlock(&proto_list_mutex);
- return ret;
-}
-
-void mem_cgroup_sockets_destroy(struct mem_cgroup *memcg)
-{
- struct proto *proto;
-
- mutex_lock(&proto_list_mutex);
- list_for_each_entry_reverse(proto, &proto_list, node)
- if (proto->destroy_cgroup)
- proto->destroy_cgroup(memcg);
- mutex_unlock(&proto_list_mutex);
-}
-#endif
-
/*
* Each address family might have different locking rules, so we have
* one slock key per address family:
@@ -239,11 +201,6 @@ void mem_cgroup_sockets_destroy(struct mem_cgroup *memcg)
static struct lock_class_key af_family_keys[AF_MAX];
static struct lock_class_key af_family_slock_keys[AF_MAX];
-#if defined(CONFIG_MEMCG_KMEM)
-struct static_key memcg_socket_limit_enabled;
-EXPORT_SYMBOL(memcg_socket_limit_enabled);
-#endif
-
/*
* Make lock validator output more readable. (we pre-construct these
* strings build-time, so that runtime initialization of socket
@@ -1476,12 +1433,6 @@ void sk_free(struct sock *sk)
}
EXPORT_SYMBOL(sk_free);
-static void sk_update_clone(const struct sock *sk, struct sock *newsk)
-{
- if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
- sock_update_memcg(newsk);
-}
-
/**
* sk_clone_lock - clone a socket, and lock its clone
* @sk: the socket to clone
@@ -1577,7 +1528,8 @@ struct sock *sk_clone_lock(const struct sock *sk, const gfp_t priority)
sk_set_socket(newsk, NULL);
newsk->sk_wq = NULL;
- sk_update_clone(sk, newsk);
+ if (mem_cgroup_do_sockets())
+ sock_update_memcg(newsk);
if (newsk->sk_prot->sockets_allocated)
sk_sockets_allocated_inc(newsk);
@@ -2036,27 +1988,27 @@ int __sk_mem_schedule(struct sock *sk, int size, int kind)
struct proto *prot = sk->sk_prot;
int amt = sk_mem_pages(size);
long allocated;
- int parent_status = UNDER_LIMIT;
sk->sk_forward_alloc += amt * SK_MEM_QUANTUM;
- allocated = sk_memory_allocated_add(sk, amt, &parent_status);
+ allocated = sk_memory_allocated_add(sk, amt);
+
+ if (mem_cgroup_do_sockets() && sk->sk_memcg &&
+ !mem_cgroup_charge_skmem(sk->sk_memcg, amt))
+ goto suppress_allocation;
/* Under limit. */
- if (parent_status == UNDER_LIMIT &&
- allocated <= sk_prot_mem_limits(sk, 0)) {
+ if (allocated <= sk_prot_mem_limits(sk, 0)) {
sk_leave_memory_pressure(sk);
return 1;
}
- /* Under pressure. (we or our parents) */
- if ((parent_status > SOFT_LIMIT) ||
- allocated > sk_prot_mem_limits(sk, 1))
+ /* Under pressure. */
+ if (allocated > sk_prot_mem_limits(sk, 1))
sk_enter_memory_pressure(sk);
- /* Over hard limit (we or our parents) */
- if ((parent_status == OVER_LIMIT) ||
- (allocated > sk_prot_mem_limits(sk, 2)))
+ /* Over hard limit. */
+ if (allocated > sk_prot_mem_limits(sk, 2))
goto suppress_allocation;
/* guarantee minimum buffer size under pressure */
@@ -2105,6 +2057,9 @@ suppress_allocation:
sk_memory_allocated_sub(sk, amt);
+ if (mem_cgroup_do_sockets() && sk->sk_memcg)
+ mem_cgroup_uncharge_skmem(sk->sk_memcg, amt);
+
return 0;
}
EXPORT_SYMBOL(__sk_mem_schedule);
@@ -2120,6 +2075,9 @@ void __sk_mem_reclaim(struct sock *sk, int amount)
sk_memory_allocated_sub(sk, amount);
sk->sk_forward_alloc -= amount << SK_MEM_QUANTUM_SHIFT;
+ if (mem_cgroup_do_sockets() && sk->sk_memcg)
+ mem_cgroup_uncharge_skmem(sk->sk_memcg, amount);
+
if (sk_under_memory_pressure(sk) &&
(sk_memory_allocated(sk) < sk_prot_mem_limits(sk, 0)))
sk_leave_memory_pressure(sk);
diff --git a/net/ipv4/sysctl_net_ipv4.c b/net/ipv4/sysctl_net_ipv4.c
index 894da3a..1f00819 100644
--- a/net/ipv4/sysctl_net_ipv4.c
+++ b/net/ipv4/sysctl_net_ipv4.c
@@ -24,7 +24,6 @@
#include <net/cipso_ipv4.h>
#include <net/inet_frag.h>
#include <net/ping.h>
-#include <net/tcp_memcontrol.h>
static int zero;
static int one = 1;
diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
index ac1bdbb..ec931c0 100644
--- a/net/ipv4/tcp.c
+++ b/net/ipv4/tcp.c
@@ -421,7 +421,8 @@ void tcp_init_sock(struct sock *sk)
sk->sk_rcvbuf = sysctl_tcp_rmem[1];
local_bh_disable();
- sock_update_memcg(sk);
+ if (mem_cgroup_do_sockets())
+ sock_update_memcg(sk);
sk_sockets_allocated_inc(sk);
local_bh_enable();
}
diff --git a/net/ipv4/tcp_ipv4.c b/net/ipv4/tcp_ipv4.c
index 30dd45c..bb5f4f2 100644
--- a/net/ipv4/tcp_ipv4.c
+++ b/net/ipv4/tcp_ipv4.c
@@ -73,7 +73,6 @@
#include <net/timewait_sock.h>
#include <net/xfrm.h>
#include <net/secure_seq.h>
-#include <net/tcp_memcontrol.h>
#include <net/busy_poll.h>
#include <linux/inet.h>
@@ -1808,7 +1807,8 @@ void tcp_v4_destroy_sock(struct sock *sk)
tcp_saved_syn_free(tp);
sk_sockets_allocated_dec(sk);
- sock_release_memcg(sk);
+ if (mem_cgroup_do_sockets())
+ sock_release_memcg(sk);
}
EXPORT_SYMBOL(tcp_v4_destroy_sock);
@@ -2330,11 +2330,6 @@ struct proto tcp_prot = {
.compat_setsockopt = compat_tcp_setsockopt,
.compat_getsockopt = compat_tcp_getsockopt,
#endif
-#ifdef CONFIG_MEMCG_KMEM
- .init_cgroup = tcp_init_cgroup,
- .destroy_cgroup = tcp_destroy_cgroup,
- .proto_cgroup = tcp_proto_cgroup,
-#endif
};
EXPORT_SYMBOL(tcp_prot);
diff --git a/net/ipv4/tcp_memcontrol.c b/net/ipv4/tcp_memcontrol.c
index 2379c1b..09a37eb 100644
--- a/net/ipv4/tcp_memcontrol.c
+++ b/net/ipv4/tcp_memcontrol.c
@@ -1,107 +1,10 @@
-#include <net/tcp.h>
-#include <net/tcp_memcontrol.h>
-#include <net/sock.h>
-#include <net/ip.h>
-#include <linux/nsproxy.h>
+#include <linux/page_counter.h>
#include <linux/memcontrol.h>
+#include <linux/cgroup.h>
#include <linux/module.h>
-
-int tcp_init_cgroup(struct mem_cgroup *memcg, struct cgroup_subsys *ss)
-{
- /*
- * The root cgroup does not use page_counters, but rather,
- * rely on the data already collected by the network
- * subsystem
- */
- struct mem_cgroup *parent = parent_mem_cgroup(memcg);
- struct page_counter *counter_parent = NULL;
- struct cg_proto *cg_proto, *parent_cg;
-
- cg_proto = tcp_prot.proto_cgroup(memcg);
- if (!cg_proto)
- return 0;
-
- cg_proto->sysctl_mem[0] = sysctl_tcp_mem[0];
- cg_proto->sysctl_mem[1] = sysctl_tcp_mem[1];
- cg_proto->sysctl_mem[2] = sysctl_tcp_mem[2];
- cg_proto->memory_pressure = 0;
- cg_proto->memcg = memcg;
-
- parent_cg = tcp_prot.proto_cgroup(parent);
- if (parent_cg)
- counter_parent = &parent_cg->memory_allocated;
-
- page_counter_init(&cg_proto->memory_allocated, counter_parent);
- percpu_counter_init(&cg_proto->sockets_allocated, 0, GFP_KERNEL);
-
- return 0;
-}
-EXPORT_SYMBOL(tcp_init_cgroup);
-
-void tcp_destroy_cgroup(struct mem_cgroup *memcg)
-{
- struct cg_proto *cg_proto;
-
- cg_proto = tcp_prot.proto_cgroup(memcg);
- if (!cg_proto)
- return;
-
- percpu_counter_destroy(&cg_proto->sockets_allocated);
-
- if (test_bit(MEMCG_SOCK_ACTIVATED, &cg_proto->flags))
- static_key_slow_dec(&memcg_socket_limit_enabled);
-
-}
-EXPORT_SYMBOL(tcp_destroy_cgroup);
-
-static int tcp_update_limit(struct mem_cgroup *memcg, unsigned long nr_pages)
-{
- struct cg_proto *cg_proto;
- int i;
- int ret;
-
- cg_proto = tcp_prot.proto_cgroup(memcg);
- if (!cg_proto)
- return -EINVAL;
-
- ret = page_counter_limit(&cg_proto->memory_allocated, nr_pages);
- if (ret)
- return ret;
-
- for (i = 0; i < 3; i++)
- cg_proto->sysctl_mem[i] = min_t(long, nr_pages,
- sysctl_tcp_mem[i]);
-
- if (nr_pages == PAGE_COUNTER_MAX)
- clear_bit(MEMCG_SOCK_ACTIVE, &cg_proto->flags);
- else {
- /*
- * The active bit needs to be written after the static_key
- * update. This is what guarantees that the socket activation
- * function is the last one to run. See sock_update_memcg() for
- * details, and note that we don't mark any socket as belonging
- * to this memcg until that flag is up.
- *
- * We need to do this, because static_keys will span multiple
- * sites, but we can't control their order. If we mark a socket
- * as accounted, but the accounting functions are not patched in
- * yet, we'll lose accounting.
- *
- * We never race with the readers in sock_update_memcg(),
- * because when this value change, the code to process it is not
- * patched in yet.
- *
- * The activated bit is used to guarantee that no two writers
- * will do the update in the same memcg. Without that, we can't
- * properly shutdown the static key.
- */
- if (!test_and_set_bit(MEMCG_SOCK_ACTIVATED, &cg_proto->flags))
- static_key_slow_inc(&memcg_socket_limit_enabled);
- set_bit(MEMCG_SOCK_ACTIVE, &cg_proto->flags);
- }
-
- return 0;
-}
+#include <linux/kernfs.h>
+#include <linux/mutex.h>
+#include <net/tcp.h>
enum {
RES_USAGE,
@@ -124,11 +27,17 @@ static ssize_t tcp_cgroup_write(struct kernfs_open_file *of,
switch (of_cft(of)->private) {
case RES_LIMIT:
/* see memcontrol.c */
+ if (memcg == root_mem_cgroup) {
+ ret = -EINVAL;
+ break;
+ }
ret = page_counter_memparse(buf, "-1", &nr_pages);
if (ret)
break;
mutex_lock(&tcp_limit_mutex);
- ret = tcp_update_limit(memcg, nr_pages);
+ ret = page_counter_limit(&memcg->skmem, nr_pages);
+ if (!ret)
+ static_branch_enable(&mem_cgroup_sockets);
mutex_unlock(&tcp_limit_mutex);
break;
default:
@@ -141,32 +50,28 @@ static ssize_t tcp_cgroup_write(struct kernfs_open_file *of,
static u64 tcp_cgroup_read(struct cgroup_subsys_state *css, struct cftype *cft)
{
struct mem_cgroup *memcg = mem_cgroup_from_css(css);
- struct cg_proto *cg_proto = tcp_prot.proto_cgroup(memcg);
u64 val;
switch (cft->private) {
case RES_LIMIT:
- if (!cg_proto)
- return PAGE_COUNTER_MAX;
- val = cg_proto->memory_allocated.limit;
+ val = memcg->skmem.limit;
val *= PAGE_SIZE;
break;
case RES_USAGE:
- if (!cg_proto)
+ if (memcg == root_mem_cgroup)
val = atomic_long_read(&tcp_memory_allocated);
else
- val = page_counter_read(&cg_proto->memory_allocated);
+ val = page_counter_read(&memcg->skmem);
val *= PAGE_SIZE;
break;
case RES_FAILCNT:
- if (!cg_proto)
- return 0;
- val = cg_proto->memory_allocated.failcnt;
+ val = memcg->skmem.failcnt;
break;
case RES_MAX_USAGE:
- if (!cg_proto)
- return 0;
- val = cg_proto->memory_allocated.watermark;
+ if (memcg == root_mem_cgroup)
+ val = 0;
+ else
+ val = memcg->skmem.watermark;
val *= PAGE_SIZE;
break;
default:
@@ -178,20 +83,14 @@ static u64 tcp_cgroup_read(struct cgroup_subsys_state *css, struct cftype *cft)
static ssize_t tcp_cgroup_reset(struct kernfs_open_file *of,
char *buf, size_t nbytes, loff_t off)
{
- struct mem_cgroup *memcg;
- struct cg_proto *cg_proto;
-
- memcg = mem_cgroup_from_css(of_css(of));
- cg_proto = tcp_prot.proto_cgroup(memcg);
- if (!cg_proto)
- return nbytes;
+ struct mem_cgroup *memcg = mem_cgroup_from_css(of_css(of));
switch (of_cft(of)->private) {
case RES_MAX_USAGE:
- page_counter_reset_watermark(&cg_proto->memory_allocated);
+ page_counter_reset_watermark(&memcg->skmem);
break;
case RES_FAILCNT:
- cg_proto->memory_allocated.failcnt = 0;
+ memcg->skmem.failcnt = 0;
break;
}
diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
index 19adedb..b496fc9 100644
--- a/net/ipv4/tcp_output.c
+++ b/net/ipv4/tcp_output.c
@@ -2819,13 +2819,15 @@ begin_fwd:
*/
void sk_forced_mem_schedule(struct sock *sk, int size)
{
- int amt, status;
+ int amt;
if (size <= sk->sk_forward_alloc)
return;
amt = sk_mem_pages(size);
sk->sk_forward_alloc += amt * SK_MEM_QUANTUM;
- sk_memory_allocated_add(sk, amt, &status);
+ sk_memory_allocated_add(sk, amt);
+ if (mem_cgroup_do_sockets() && sk->sk_memcg)
+ mem_cgroup_charge_skmem(sk->sk_memcg, amt);
}
/* Send a FIN. The caller locks the socket for us.
diff --git a/net/ipv6/tcp_ipv6.c b/net/ipv6/tcp_ipv6.c
index f495d18..cf19e65 100644
--- a/net/ipv6/tcp_ipv6.c
+++ b/net/ipv6/tcp_ipv6.c
@@ -1862,9 +1862,6 @@ struct proto tcpv6_prot = {
.compat_setsockopt = compat_tcp_setsockopt,
.compat_getsockopt = compat_tcp_getsockopt,
#endif
-#ifdef CONFIG_MEMCG_KMEM
- .proto_cgroup = tcp_proto_cgroup,
-#endif
.clear_sk = tcp_v6_clear_sk,
};
--
2.6.1
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@virtuozzo.com> |
|---|---|
| Date | 2015-10-22 20:50 +0200 |
| Subject | Re: [PATCH 3/8] net: consolidate memcg socket buffer tracking and accounting |
| Message-ID | <qmt61-2wa-1@gated-at.bofh.it> |
| In reply to | #1253476 |
On Thu, Oct 22, 2015 at 12:21:31AM -0400, Johannes Weiner wrote: > The tcp memory controller has extensive provisions for future memory > accounting interfaces that won't materialize after all. Cut the code > base down to what's actually used, now and in the likely future. > > - There won't be any different protocol counters in the future, so a > direct sock->sk_memcg linkage is enough. This eliminates a lot of > callback maze and boilerplate code, and restores most of the socket > allocation code to pre-tcp_memcontrol state. > > - There won't be a tcp control soft limit, so integrating the memcg In fact, the code is ready for the "soft" limit (I mean min, pressure, max tuple), it just lacks a knob. > code into the global skmem limiting scheme complicates things > unnecessarily. Replace all that with simple and clear charge and > uncharge calls--hidden behind a jump label--to account skb memory. > > - The previous jump label code was an elaborate state machine that > tracked the number of cgroups with an active socket limit in order > to enable the skmem tracking and accounting code only when actively > necessary. But this is overengineered: it was meant to protect the > people who never use this feature in the first place. Simply enable > the branches once when the first limit is set until the next reboot. > ... > @@ -1136,9 +1090,6 @@ static inline bool sk_under_memory_pressure(const struct sock *sk) > if (!sk->sk_prot->memory_pressure) > return false; > > - if (mem_cgroup_sockets_enabled && sk->sk_cgrp) > - return !!sk->sk_cgrp->memory_pressure; > - AFAIU, now we won't shrink the window on hitting the limit, i.e. this patch subtly changes the behavior of the existing knobs, potentially breaking them. Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2015-10-22 21:20 +0200 |
| Subject | Re: [PATCH 3/8] net: consolidate memcg socket buffer tracking and accounting |
| Message-ID | <qmtz5-3kA-39@gated-at.bofh.it> |
| In reply to | #1254076 |
On Thu, Oct 22, 2015 at 09:46:12PM +0300, Vladimir Davydov wrote: > On Thu, Oct 22, 2015 at 12:21:31AM -0400, Johannes Weiner wrote: > > The tcp memory controller has extensive provisions for future memory > > accounting interfaces that won't materialize after all. Cut the code > > base down to what's actually used, now and in the likely future. > > > > - There won't be any different protocol counters in the future, so a > > direct sock->sk_memcg linkage is enough. This eliminates a lot of > > callback maze and boilerplate code, and restores most of the socket > > allocation code to pre-tcp_memcontrol state. > > > > - There won't be a tcp control soft limit, so integrating the memcg > > In fact, the code is ready for the "soft" limit (I mean min, pressure, > max tuple), it just lacks a knob. Yeah, but that's not going to materialize if the entire interface for dedicated tcp throttling is considered obsolete. > > @@ -1136,9 +1090,6 @@ static inline bool sk_under_memory_pressure(const struct sock *sk) > > if (!sk->sk_prot->memory_pressure) > > return false; > > > > - if (mem_cgroup_sockets_enabled && sk->sk_cgrp) > > - return !!sk->sk_cgrp->memory_pressure; > > - > > AFAIU, now we won't shrink the window on hitting the limit, i.e. this > patch subtly changes the behavior of the existing knobs, potentially > breaking them. Hm, but there is no grace period in which something meaningful could happen with the window shrinking, is there? Any buffer allocation is still going to fail hard. I don't see how this would change anything in practice. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@virtuozzo.com> |
|---|---|
| Date | 2015-10-23 15:50 +0200 |
| Subject | Re: [PATCH 3/8] net: consolidate memcg socket buffer tracking and accounting |
| Message-ID | <qmKTg-32t-13@gated-at.bofh.it> |
| In reply to | #1254126 |
On Thu, Oct 22, 2015 at 03:09:43PM -0400, Johannes Weiner wrote: > On Thu, Oct 22, 2015 at 09:46:12PM +0300, Vladimir Davydov wrote: > > On Thu, Oct 22, 2015 at 12:21:31AM -0400, Johannes Weiner wrote: > > > The tcp memory controller has extensive provisions for future memory > > > accounting interfaces that won't materialize after all. Cut the code > > > base down to what's actually used, now and in the likely future. > > > > > > - There won't be any different protocol counters in the future, so a > > > direct sock->sk_memcg linkage is enough. This eliminates a lot of > > > callback maze and boilerplate code, and restores most of the socket > > > allocation code to pre-tcp_memcontrol state. > > > > > > - There won't be a tcp control soft limit, so integrating the memcg > > > > In fact, the code is ready for the "soft" limit (I mean min, pressure, > > max tuple), it just lacks a knob. > > Yeah, but that's not going to materialize if the entire interface for > dedicated tcp throttling is considered obsolete. May be, it shouldn't be. My current understanding is that per memcg tcp window control is necessary, because: - We need to be able to protect a containerized workload from its growing network buffers. Using vmpressure notifications for that does not look reassuring to me. - We need a way to limit network buffers of a particular container, otherwise it can fill the system-wide window throttling other containers, which is unfair. > > > > @@ -1136,9 +1090,6 @@ static inline bool sk_under_memory_pressure(const struct sock *sk) > > > if (!sk->sk_prot->memory_pressure) > > > return false; > > > > > > - if (mem_cgroup_sockets_enabled && sk->sk_cgrp) > > > - return !!sk->sk_cgrp->memory_pressure; > > > - > > > > AFAIU, now we won't shrink the window on hitting the limit, i.e. this > > patch subtly changes the behavior of the existing knobs, potentially > > breaking them. > > Hm, but there is no grace period in which something meaningful could > happen with the window shrinking, is there? Any buffer allocation is > still going to fail hard. AFAIU when we hit the limit, we not only throttle the socket which allocates, but also try to release space reserved by other sockets. After your patch we won't. This looks unfair to me. Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2015-10-23 14:40 +0200 |
| Subject | Re: [PATCH 3/8] net: consolidate memcg socket buffer tracking and accounting |
| Message-ID | <qmJNx-1va-45@gated-at.bofh.it> |
| In reply to | #1253476 |
On Thu 22-10-15 00:21:31, Johannes Weiner wrote:
> The tcp memory controller has extensive provisions for future memory
> accounting interfaces that won't materialize after all. Cut the code
> base down to what's actually used, now and in the likely future.
>
> - There won't be any different protocol counters in the future, so a
> direct sock->sk_memcg linkage is enough. This eliminates a lot of
> callback maze and boilerplate code, and restores most of the socket
> allocation code to pre-tcp_memcontrol state.
>
> - There won't be a tcp control soft limit, so integrating the memcg
> code into the global skmem limiting scheme complicates things
> unnecessarily. Replace all that with simple and clear charge and
> uncharge calls--hidden behind a jump label--to account skb memory.
>
> - The previous jump label code was an elaborate state machine that
> tracked the number of cgroups with an active socket limit in order
> to enable the skmem tracking and accounting code only when actively
> necessary. But this is overengineered: it was meant to protect the
> people who never use this feature in the first place. Simply enable
> the branches once when the first limit is set until the next reboot.
>
> Signed-off-by: Johannes Weiner <hannes@cmpxchg.org>
The changelog is certainly attractive. I have looked through the patch
but my knowledge of the networking subsystem and its memory management
is close to zero so I cannot really do a competent review.
Anyway I support any simplification of the tcp kmem accounting. If
networking people are OK with the changes, including reduction of the
functionality as described by Vladimir then no objections from me for
this to be merged.
Thanks!
> ---
> include/linux/memcontrol.h | 64 ++++++++-----------
> include/net/sock.h | 135 +++------------------------------------
> include/net/tcp.h | 3 -
> include/net/tcp_memcontrol.h | 7 ---
> mm/memcontrol.c | 101 +++++++++++++++--------------
> net/core/sock.c | 78 ++++++-----------------
> net/ipv4/sysctl_net_ipv4.c | 1 -
> net/ipv4/tcp.c | 3 +-
> net/ipv4/tcp_ipv4.c | 9 +--
> net/ipv4/tcp_memcontrol.c | 147 +++++++------------------------------------
> net/ipv4/tcp_output.c | 6 +-
> net/ipv6/tcp_ipv6.c | 3 -
> 12 files changed, 136 insertions(+), 421 deletions(-)
> delete mode 100644 include/net/tcp_memcontrol.h
>
> diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
> index 19ff87b..5b72f83 100644
> --- a/include/linux/memcontrol.h
> +++ b/include/linux/memcontrol.h
> @@ -85,34 +85,6 @@ enum mem_cgroup_events_target {
> MEM_CGROUP_NTARGETS,
> };
>
> -/*
> - * Bits in struct cg_proto.flags
> - */
> -enum cg_proto_flags {
> - /* Currently active and new sockets should be assigned to cgroups */
> - MEMCG_SOCK_ACTIVE,
> - /* It was ever activated; we must disarm static keys on destruction */
> - MEMCG_SOCK_ACTIVATED,
> -};
> -
> -struct cg_proto {
> - struct page_counter memory_allocated; /* Current allocated memory. */
> - struct percpu_counter sockets_allocated; /* Current number of sockets. */
> - int memory_pressure;
> - long sysctl_mem[3];
> - unsigned long flags;
> - /*
> - * memcg field is used to find which memcg we belong directly
> - * Each memcg struct can hold more than one cg_proto, so container_of
> - * won't really cut.
> - *
> - * The elegant solution would be having an inverse function to
> - * proto_cgroup in struct proto, but that means polluting the structure
> - * for everybody, instead of just for memcg users.
> - */
> - struct mem_cgroup *memcg;
> -};
> -
> #ifdef CONFIG_MEMCG
> struct mem_cgroup_stat_cpu {
> long count[MEM_CGROUP_STAT_NSTATS];
> @@ -185,8 +157,15 @@ struct mem_cgroup {
>
> /* Accounted resources */
> struct page_counter memory;
> +
> + /*
> + * Legacy non-resource counters. In unified hierarchy, all
> + * memory is accounted and limited through memcg->memory.
> + * Consumer breakdown happens in the statistics.
> + */
> struct page_counter memsw;
> struct page_counter kmem;
> + struct page_counter skmem;
>
> /* Normal memory consumption range */
> unsigned long low;
> @@ -246,9 +225,6 @@ struct mem_cgroup {
> */
> struct mem_cgroup_stat_cpu __percpu *stat;
>
> -#if defined(CONFIG_MEMCG_KMEM) && defined(CONFIG_INET)
> - struct cg_proto tcp_mem;
> -#endif
> #if defined(CONFIG_MEMCG_KMEM)
> /* Index in the kmem_cache->memcg_params.memcg_caches array */
> int kmemcg_id;
> @@ -676,12 +652,6 @@ void mem_cgroup_count_vm_event(struct mm_struct *mm, enum vm_event_item idx)
> }
> #endif /* CONFIG_MEMCG */
>
> -enum {
> - UNDER_LIMIT,
> - SOFT_LIMIT,
> - OVER_LIMIT,
> -};
> -
> #ifdef CONFIG_CGROUP_WRITEBACK
>
> struct list_head *mem_cgroup_cgwb_list(struct mem_cgroup *memcg);
> @@ -707,15 +677,35 @@ static inline void mem_cgroup_wb_stats(struct bdi_writeback *wb,
>
> struct sock;
> #if defined(CONFIG_INET) && defined(CONFIG_MEMCG_KMEM)
> +extern struct static_key_false mem_cgroup_sockets;
> +static inline bool mem_cgroup_do_sockets(void)
> +{
> + return static_branch_unlikely(&mem_cgroup_sockets);
> +}
> void sock_update_memcg(struct sock *sk);
> void sock_release_memcg(struct sock *sk);
> +bool mem_cgroup_charge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages);
> +void mem_cgroup_uncharge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages);
> #else
> +static inline bool mem_cgroup_do_sockets(void)
> +{
> + return false;
> +}
> static inline void sock_update_memcg(struct sock *sk)
> {
> }
> static inline void sock_release_memcg(struct sock *sk)
> {
> }
> +static inline bool mem_cgroup_charge_skmem(struct mem_cgroup *memcg,
> + unsigned int nr_pages)
> +{
> + return true;
> +}
> +static inline void mem_cgroup_uncharge_skmem(struct mem_cgroup *memcg,
> + unsigned int nr_pages)
> +{
> +}
> #endif /* CONFIG_INET && CONFIG_MEMCG_KMEM */
>
> #ifdef CONFIG_MEMCG_KMEM
> diff --git a/include/net/sock.h b/include/net/sock.h
> index 59a7196..67795fc 100644
> --- a/include/net/sock.h
> +++ b/include/net/sock.h
> @@ -69,22 +69,6 @@
> #include <net/tcp_states.h>
> #include <linux/net_tstamp.h>
>
> -struct cgroup;
> -struct cgroup_subsys;
> -#ifdef CONFIG_NET
> -int mem_cgroup_sockets_init(struct mem_cgroup *memcg, struct cgroup_subsys *ss);
> -void mem_cgroup_sockets_destroy(struct mem_cgroup *memcg);
> -#else
> -static inline
> -int mem_cgroup_sockets_init(struct mem_cgroup *memcg, struct cgroup_subsys *ss)
> -{
> - return 0;
> -}
> -static inline
> -void mem_cgroup_sockets_destroy(struct mem_cgroup *memcg)
> -{
> -}
> -#endif
> /*
> * This structure really needs to be cleaned up.
> * Most of it is for TCP, and not used by any of
> @@ -243,7 +227,6 @@ struct sock_common {
> /* public: */
> };
>
> -struct cg_proto;
> /**
> * struct sock - network layer representation of sockets
> * @__sk_common: shared layout with inet_timewait_sock
> @@ -310,7 +293,7 @@ struct cg_proto;
> * @sk_security: used by security modules
> * @sk_mark: generic packet mark
> * @sk_classid: this socket's cgroup classid
> - * @sk_cgrp: this socket's cgroup-specific proto data
> + * @sk_memcg: this socket's memcg association
> * @sk_write_pending: a write to stream socket waits to start
> * @sk_state_change: callback to indicate change in the state of the sock
> * @sk_data_ready: callback to indicate there is data to be processed
> @@ -447,7 +430,7 @@ struct sock {
> #ifdef CONFIG_CGROUP_NET_CLASSID
> u32 sk_classid;
> #endif
> - struct cg_proto *sk_cgrp;
> + struct mem_cgroup *sk_memcg;
> void (*sk_state_change)(struct sock *sk);
> void (*sk_data_ready)(struct sock *sk);
> void (*sk_write_space)(struct sock *sk);
> @@ -1051,18 +1034,6 @@ struct proto {
> #ifdef SOCK_REFCNT_DEBUG
> atomic_t socks;
> #endif
> -#ifdef CONFIG_MEMCG_KMEM
> - /*
> - * cgroup specific init/deinit functions. Called once for all
> - * protocols that implement it, from cgroups populate function.
> - * This function has to setup any files the protocol want to
> - * appear in the kmem cgroup filesystem.
> - */
> - int (*init_cgroup)(struct mem_cgroup *memcg,
> - struct cgroup_subsys *ss);
> - void (*destroy_cgroup)(struct mem_cgroup *memcg);
> - struct cg_proto *(*proto_cgroup)(struct mem_cgroup *memcg);
> -#endif
> };
>
> int proto_register(struct proto *prot, int alloc_slab);
> @@ -1093,23 +1064,6 @@ static inline void sk_refcnt_debug_release(const struct sock *sk)
> #define sk_refcnt_debug_release(sk) do { } while (0)
> #endif /* SOCK_REFCNT_DEBUG */
>
> -#if defined(CONFIG_MEMCG_KMEM) && defined(CONFIG_NET)
> -extern struct static_key memcg_socket_limit_enabled;
> -static inline struct cg_proto *parent_cg_proto(struct proto *proto,
> - struct cg_proto *cg_proto)
> -{
> - return proto->proto_cgroup(parent_mem_cgroup(cg_proto->memcg));
> -}
> -#define mem_cgroup_sockets_enabled static_key_false(&memcg_socket_limit_enabled)
> -#else
> -#define mem_cgroup_sockets_enabled 0
> -static inline struct cg_proto *parent_cg_proto(struct proto *proto,
> - struct cg_proto *cg_proto)
> -{
> - return NULL;
> -}
> -#endif
> -
> static inline bool sk_stream_memory_free(const struct sock *sk)
> {
> if (sk->sk_wmem_queued >= sk->sk_sndbuf)
> @@ -1136,9 +1090,6 @@ static inline bool sk_under_memory_pressure(const struct sock *sk)
> if (!sk->sk_prot->memory_pressure)
> return false;
>
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
> - return !!sk->sk_cgrp->memory_pressure;
> -
> return !!*sk->sk_prot->memory_pressure;
> }
>
> @@ -1146,61 +1097,19 @@ static inline void sk_leave_memory_pressure(struct sock *sk)
> {
> int *memory_pressure = sk->sk_prot->memory_pressure;
>
> - if (!memory_pressure)
> - return;
> -
> - if (*memory_pressure)
> + if (memory_pressure && *memory_pressure)
> *memory_pressure = 0;
> -
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
> - struct cg_proto *cg_proto = sk->sk_cgrp;
> - struct proto *prot = sk->sk_prot;
> -
> - for (; cg_proto; cg_proto = parent_cg_proto(prot, cg_proto))
> - cg_proto->memory_pressure = 0;
> - }
> -
> }
>
> static inline void sk_enter_memory_pressure(struct sock *sk)
> {
> - if (!sk->sk_prot->enter_memory_pressure)
> - return;
> -
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
> - struct cg_proto *cg_proto = sk->sk_cgrp;
> - struct proto *prot = sk->sk_prot;
> -
> - for (; cg_proto; cg_proto = parent_cg_proto(prot, cg_proto))
> - cg_proto->memory_pressure = 1;
> - }
> -
> - sk->sk_prot->enter_memory_pressure(sk);
> + if (sk->sk_prot->enter_memory_pressure)
> + sk->sk_prot->enter_memory_pressure(sk);
> }
>
> static inline long sk_prot_mem_limits(const struct sock *sk, int index)
> {
> - long *prot = sk->sk_prot->sysctl_mem;
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
> - prot = sk->sk_cgrp->sysctl_mem;
> - return prot[index];
> -}
> -
> -static inline void memcg_memory_allocated_add(struct cg_proto *prot,
> - unsigned long amt,
> - int *parent_status)
> -{
> - page_counter_charge(&prot->memory_allocated, amt);
> -
> - if (page_counter_read(&prot->memory_allocated) >
> - prot->memory_allocated.limit)
> - *parent_status = OVER_LIMIT;
> -}
> -
> -static inline void memcg_memory_allocated_sub(struct cg_proto *prot,
> - unsigned long amt)
> -{
> - page_counter_uncharge(&prot->memory_allocated, amt);
> + return sk->sk_prot->sysctl_mem[index];
> }
>
> static inline long
> @@ -1208,24 +1117,14 @@ sk_memory_allocated(const struct sock *sk)
> {
> struct proto *prot = sk->sk_prot;
>
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
> - return page_counter_read(&sk->sk_cgrp->memory_allocated);
> -
> return atomic_long_read(prot->memory_allocated);
> }
>
> static inline long
> -sk_memory_allocated_add(struct sock *sk, int amt, int *parent_status)
> +sk_memory_allocated_add(struct sock *sk, int amt)
> {
> struct proto *prot = sk->sk_prot;
>
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
> - memcg_memory_allocated_add(sk->sk_cgrp, amt, parent_status);
> - /* update the root cgroup regardless */
> - atomic_long_add_return(amt, prot->memory_allocated);
> - return page_counter_read(&sk->sk_cgrp->memory_allocated);
> - }
> -
> return atomic_long_add_return(amt, prot->memory_allocated);
> }
>
> @@ -1234,9 +1133,6 @@ sk_memory_allocated_sub(struct sock *sk, int amt)
> {
> struct proto *prot = sk->sk_prot;
>
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
> - memcg_memory_allocated_sub(sk->sk_cgrp, amt);
> -
> atomic_long_sub(amt, prot->memory_allocated);
> }
>
> @@ -1244,13 +1140,6 @@ static inline void sk_sockets_allocated_dec(struct sock *sk)
> {
> struct proto *prot = sk->sk_prot;
>
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
> - struct cg_proto *cg_proto = sk->sk_cgrp;
> -
> - for (; cg_proto; cg_proto = parent_cg_proto(prot, cg_proto))
> - percpu_counter_dec(&cg_proto->sockets_allocated);
> - }
> -
> percpu_counter_dec(prot->sockets_allocated);
> }
>
> @@ -1258,13 +1147,6 @@ static inline void sk_sockets_allocated_inc(struct sock *sk)
> {
> struct proto *prot = sk->sk_prot;
>
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
> - struct cg_proto *cg_proto = sk->sk_cgrp;
> -
> - for (; cg_proto; cg_proto = parent_cg_proto(prot, cg_proto))
> - percpu_counter_inc(&cg_proto->sockets_allocated);
> - }
> -
> percpu_counter_inc(prot->sockets_allocated);
> }
>
> @@ -1273,9 +1155,6 @@ sk_sockets_allocated_read_positive(struct sock *sk)
> {
> struct proto *prot = sk->sk_prot;
>
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
> - return percpu_counter_read_positive(&sk->sk_cgrp->sockets_allocated);
> -
> return percpu_counter_read_positive(prot->sockets_allocated);
> }
>
> diff --git a/include/net/tcp.h b/include/net/tcp.h
> index eed94fc..77b6c7e 100644
> --- a/include/net/tcp.h
> +++ b/include/net/tcp.h
> @@ -291,9 +291,6 @@ extern int tcp_memory_pressure;
> /* optimized version of sk_under_memory_pressure() for TCP sockets */
> static inline bool tcp_under_memory_pressure(const struct sock *sk)
> {
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
> - return !!sk->sk_cgrp->memory_pressure;
> -
> return tcp_memory_pressure;
> }
> /*
> diff --git a/include/net/tcp_memcontrol.h b/include/net/tcp_memcontrol.h
> deleted file mode 100644
> index 05b94d9..0000000
> --- a/include/net/tcp_memcontrol.h
> +++ /dev/null
> @@ -1,7 +0,0 @@
> -#ifndef _TCP_MEMCG_H
> -#define _TCP_MEMCG_H
> -
> -struct cg_proto *tcp_proto_cgroup(struct mem_cgroup *memcg);
> -int tcp_init_cgroup(struct mem_cgroup *memcg, struct cgroup_subsys *ss);
> -void tcp_destroy_cgroup(struct mem_cgroup *memcg);
> -#endif /* _TCP_MEMCG_H */
> diff --git a/mm/memcontrol.c b/mm/memcontrol.c
> index e54f434..c41e6d7 100644
> --- a/mm/memcontrol.c
> +++ b/mm/memcontrol.c
> @@ -66,7 +66,6 @@
> #include "internal.h"
> #include <net/sock.h>
> #include <net/ip.h>
> -#include <net/tcp_memcontrol.h>
> #include "slab.h"
>
> #include <asm/uaccess.h>
> @@ -291,58 +290,68 @@ static inline struct mem_cgroup *mem_cgroup_from_id(unsigned short id)
> /* Writing them here to avoid exposing memcg's inner layout */
> #if defined(CONFIG_INET) && defined(CONFIG_MEMCG_KMEM)
>
> +DEFINE_STATIC_KEY_FALSE(mem_cgroup_sockets);
> +
> void sock_update_memcg(struct sock *sk)
> {
> - if (mem_cgroup_sockets_enabled) {
> - struct mem_cgroup *memcg;
> - struct cg_proto *cg_proto;
> -
> - BUG_ON(!sk->sk_prot->proto_cgroup);
> -
> - /* Socket cloning can throw us here with sk_cgrp already
> - * filled. It won't however, necessarily happen from
> - * process context. So the test for root memcg given
> - * the current task's memcg won't help us in this case.
> - *
> - * Respecting the original socket's memcg is a better
> - * decision in this case.
> - */
> - if (sk->sk_cgrp) {
> - BUG_ON(mem_cgroup_is_root(sk->sk_cgrp->memcg));
> - css_get(&sk->sk_cgrp->memcg->css);
> - return;
> - }
> -
> - rcu_read_lock();
> - memcg = mem_cgroup_from_task(current);
> - cg_proto = sk->sk_prot->proto_cgroup(memcg);
> - if (cg_proto && test_bit(MEMCG_SOCK_ACTIVE, &cg_proto->flags) &&
> - css_tryget_online(&memcg->css)) {
> - sk->sk_cgrp = cg_proto;
> - }
> - rcu_read_unlock();
> + struct mem_cgroup *memcg;
> + /*
> + * Socket cloning can throw us here with sk_cgrp already
> + * filled. It won't however, necessarily happen from
> + * process context. So the test for root memcg given
> + * the current task's memcg won't help us in this case.
> + *
> + * Respecting the original socket's memcg is a better
> + * decision in this case.
> + */
> + if (sk->sk_memcg) {
> + BUG_ON(mem_cgroup_is_root(sk->sk_memcg));
> + css_get(&sk->sk_memcg->css);
> + return;
> }
> +
> + rcu_read_lock();
> + memcg = mem_cgroup_from_task(current);
> + if (css_tryget_online(&memcg->css))
> + sk->sk_memcg = memcg;
> + rcu_read_unlock();
> }
> EXPORT_SYMBOL(sock_update_memcg);
>
> void sock_release_memcg(struct sock *sk)
> {
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp) {
> - struct mem_cgroup *memcg;
> - WARN_ON(!sk->sk_cgrp->memcg);
> - memcg = sk->sk_cgrp->memcg;
> - css_put(&sk->sk_cgrp->memcg->css);
> - }
> + if (sk->sk_memcg)
> + css_put(&sk->sk_memcg->css);
> }
>
> -struct cg_proto *tcp_proto_cgroup(struct mem_cgroup *memcg)
> +/**
> + * mem_cgroup_charge_skmem - charge socket memory
> + * @memcg: memcg to charge
> + * @nr_pages: number of pages to charge
> + *
> + * Charges @nr_pages to @memcg. Returns %true if the charge fit within
> + * the memcg's configured limit, %false if the charge had to be forced.
> + */
> +bool mem_cgroup_charge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages)
> {
> - if (!memcg || mem_cgroup_is_root(memcg))
> - return NULL;
> + struct page_counter *counter;
> +
> + if (page_counter_try_charge(&memcg->skmem, nr_pages, &counter))
> + return true;
>
> - return &memcg->tcp_mem;
> + page_counter_charge(&memcg->skmem, nr_pages);
> + return false;
> +}
> +
> +/**
> + * mem_cgroup_uncharge_skmem - uncharge socket memory
> + * @memcg: memcg to uncharge
> + * @nr_pages: number of pages to uncharge
> + */
> +void mem_cgroup_uncharge_skmem(struct mem_cgroup *memcg, unsigned int nr_pages)
> +{
> + page_counter_uncharge(&memcg->skmem, nr_pages);
> }
> -EXPORT_SYMBOL(tcp_proto_cgroup);
>
> #endif
>
> @@ -3592,13 +3601,7 @@ static int mem_cgroup_oom_control_write(struct cgroup_subsys_state *css,
> #ifdef CONFIG_MEMCG_KMEM
> static int memcg_init_kmem(struct mem_cgroup *memcg, struct cgroup_subsys *ss)
> {
> - int ret;
> -
> - ret = memcg_propagate_kmem(memcg);
> - if (ret)
> - return ret;
> -
> - return mem_cgroup_sockets_init(memcg, ss);
> + return memcg_propagate_kmem(memcg);
> }
>
> static void memcg_deactivate_kmem(struct mem_cgroup *memcg)
> @@ -3654,7 +3657,6 @@ static void memcg_destroy_kmem(struct mem_cgroup *memcg)
> static_key_slow_dec(&memcg_kmem_enabled_key);
> WARN_ON(page_counter_read(&memcg->kmem));
> }
> - mem_cgroup_sockets_destroy(memcg);
> }
> #else
> static int memcg_init_kmem(struct mem_cgroup *memcg, struct cgroup_subsys *ss)
> @@ -4218,6 +4220,7 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *parent_css)
> memcg->soft_limit = PAGE_COUNTER_MAX;
> page_counter_init(&memcg->memsw, NULL);
> page_counter_init(&memcg->kmem, NULL);
> + page_counter_init(&memcg->skmem, NULL);
> }
>
> memcg->last_scanned_node = MAX_NUMNODES;
> @@ -4266,6 +4269,7 @@ mem_cgroup_css_online(struct cgroup_subsys_state *css)
> memcg->soft_limit = PAGE_COUNTER_MAX;
> page_counter_init(&memcg->memsw, &parent->memsw);
> page_counter_init(&memcg->kmem, &parent->kmem);
> + page_counter_init(&memcg->skmem, &parent->skmem);
>
> /*
> * No need to take a reference to the parent because cgroup
> @@ -4277,6 +4281,7 @@ mem_cgroup_css_online(struct cgroup_subsys_state *css)
> memcg->soft_limit = PAGE_COUNTER_MAX;
> page_counter_init(&memcg->memsw, NULL);
> page_counter_init(&memcg->kmem, NULL);
> + page_counter_init(&memcg->skmem, NULL);
> /*
> * Deeper hierachy with use_hierarchy == false doesn't make
> * much sense so let cgroup subsystem know about this
> diff --git a/net/core/sock.c b/net/core/sock.c
> index 0fafd27..0debff5 100644
> --- a/net/core/sock.c
> +++ b/net/core/sock.c
> @@ -194,44 +194,6 @@ bool sk_net_capable(const struct sock *sk, int cap)
> }
> EXPORT_SYMBOL(sk_net_capable);
>
> -
> -#ifdef CONFIG_MEMCG_KMEM
> -int mem_cgroup_sockets_init(struct mem_cgroup *memcg, struct cgroup_subsys *ss)
> -{
> - struct proto *proto;
> - int ret = 0;
> -
> - mutex_lock(&proto_list_mutex);
> - list_for_each_entry(proto, &proto_list, node) {
> - if (proto->init_cgroup) {
> - ret = proto->init_cgroup(memcg, ss);
> - if (ret)
> - goto out;
> - }
> - }
> -
> - mutex_unlock(&proto_list_mutex);
> - return ret;
> -out:
> - list_for_each_entry_continue_reverse(proto, &proto_list, node)
> - if (proto->destroy_cgroup)
> - proto->destroy_cgroup(memcg);
> - mutex_unlock(&proto_list_mutex);
> - return ret;
> -}
> -
> -void mem_cgroup_sockets_destroy(struct mem_cgroup *memcg)
> -{
> - struct proto *proto;
> -
> - mutex_lock(&proto_list_mutex);
> - list_for_each_entry_reverse(proto, &proto_list, node)
> - if (proto->destroy_cgroup)
> - proto->destroy_cgroup(memcg);
> - mutex_unlock(&proto_list_mutex);
> -}
> -#endif
> -
> /*
> * Each address family might have different locking rules, so we have
> * one slock key per address family:
> @@ -239,11 +201,6 @@ void mem_cgroup_sockets_destroy(struct mem_cgroup *memcg)
> static struct lock_class_key af_family_keys[AF_MAX];
> static struct lock_class_key af_family_slock_keys[AF_MAX];
>
> -#if defined(CONFIG_MEMCG_KMEM)
> -struct static_key memcg_socket_limit_enabled;
> -EXPORT_SYMBOL(memcg_socket_limit_enabled);
> -#endif
> -
> /*
> * Make lock validator output more readable. (we pre-construct these
> * strings build-time, so that runtime initialization of socket
> @@ -1476,12 +1433,6 @@ void sk_free(struct sock *sk)
> }
> EXPORT_SYMBOL(sk_free);
>
> -static void sk_update_clone(const struct sock *sk, struct sock *newsk)
> -{
> - if (mem_cgroup_sockets_enabled && sk->sk_cgrp)
> - sock_update_memcg(newsk);
> -}
> -
> /**
> * sk_clone_lock - clone a socket, and lock its clone
> * @sk: the socket to clone
> @@ -1577,7 +1528,8 @@ struct sock *sk_clone_lock(const struct sock *sk, const gfp_t priority)
> sk_set_socket(newsk, NULL);
> newsk->sk_wq = NULL;
>
> - sk_update_clone(sk, newsk);
> + if (mem_cgroup_do_sockets())
> + sock_update_memcg(newsk);
>
> if (newsk->sk_prot->sockets_allocated)
> sk_sockets_allocated_inc(newsk);
> @@ -2036,27 +1988,27 @@ int __sk_mem_schedule(struct sock *sk, int size, int kind)
> struct proto *prot = sk->sk_prot;
> int amt = sk_mem_pages(size);
> long allocated;
> - int parent_status = UNDER_LIMIT;
>
> sk->sk_forward_alloc += amt * SK_MEM_QUANTUM;
>
> - allocated = sk_memory_allocated_add(sk, amt, &parent_status);
> + allocated = sk_memory_allocated_add(sk, amt);
> +
> + if (mem_cgroup_do_sockets() && sk->sk_memcg &&
> + !mem_cgroup_charge_skmem(sk->sk_memcg, amt))
> + goto suppress_allocation;
>
> /* Under limit. */
> - if (parent_status == UNDER_LIMIT &&
> - allocated <= sk_prot_mem_limits(sk, 0)) {
> + if (allocated <= sk_prot_mem_limits(sk, 0)) {
> sk_leave_memory_pressure(sk);
> return 1;
> }
>
> - /* Under pressure. (we or our parents) */
> - if ((parent_status > SOFT_LIMIT) ||
> - allocated > sk_prot_mem_limits(sk, 1))
> + /* Under pressure. */
> + if (allocated > sk_prot_mem_limits(sk, 1))
> sk_enter_memory_pressure(sk);
>
> - /* Over hard limit (we or our parents) */
> - if ((parent_status == OVER_LIMIT) ||
> - (allocated > sk_prot_mem_limits(sk, 2)))
> + /* Over hard limit. */
> + if (allocated > sk_prot_mem_limits(sk, 2))
> goto suppress_allocation;
>
> /* guarantee minimum buffer size under pressure */
> @@ -2105,6 +2057,9 @@ suppress_allocation:
>
> sk_memory_allocated_sub(sk, amt);
>
> + if (mem_cgroup_do_sockets() && sk->sk_memcg)
> + mem_cgroup_uncharge_skmem(sk->sk_memcg, amt);
> +
> return 0;
> }
> EXPORT_SYMBOL(__sk_mem_schedule);
> @@ -2120,6 +2075,9 @@ void __sk_mem_reclaim(struct sock *sk, int amount)
> sk_memory_allocated_sub(sk, amount);
> sk->sk_forward_alloc -= amount << SK_MEM_QUANTUM_SHIFT;
>
> + if (mem_cgroup_do_sockets() && sk->sk_memcg)
> + mem_cgroup_uncharge_skmem(sk->sk_memcg, amount);
> +
> if (sk_under_memory_pressure(sk) &&
> (sk_memory_allocated(sk) < sk_prot_mem_limits(sk, 0)))
> sk_leave_memory_pressure(sk);
> diff --git a/net/ipv4/sysctl_net_ipv4.c b/net/ipv4/sysctl_net_ipv4.c
> index 894da3a..1f00819 100644
> --- a/net/ipv4/sysctl_net_ipv4.c
> +++ b/net/ipv4/sysctl_net_ipv4.c
> @@ -24,7 +24,6 @@
> #include <net/cipso_ipv4.h>
> #include <net/inet_frag.h>
> #include <net/ping.h>
> -#include <net/tcp_memcontrol.h>
>
> static int zero;
> static int one = 1;
> diff --git a/net/ipv4/tcp.c b/net/ipv4/tcp.c
> index ac1bdbb..ec931c0 100644
> --- a/net/ipv4/tcp.c
> +++ b/net/ipv4/tcp.c
> @@ -421,7 +421,8 @@ void tcp_init_sock(struct sock *sk)
> sk->sk_rcvbuf = sysctl_tcp_rmem[1];
>
> local_bh_disable();
> - sock_update_memcg(sk);
> + if (mem_cgroup_do_sockets())
> + sock_update_memcg(sk);
> sk_sockets_allocated_inc(sk);
> local_bh_enable();
> }
> diff --git a/net/ipv4/tcp_ipv4.c b/net/ipv4/tcp_ipv4.c
> index 30dd45c..bb5f4f2 100644
> --- a/net/ipv4/tcp_ipv4.c
> +++ b/net/ipv4/tcp_ipv4.c
> @@ -73,7 +73,6 @@
> #include <net/timewait_sock.h>
> #include <net/xfrm.h>
> #include <net/secure_seq.h>
> -#include <net/tcp_memcontrol.h>
> #include <net/busy_poll.h>
>
> #include <linux/inet.h>
> @@ -1808,7 +1807,8 @@ void tcp_v4_destroy_sock(struct sock *sk)
> tcp_saved_syn_free(tp);
>
> sk_sockets_allocated_dec(sk);
> - sock_release_memcg(sk);
> + if (mem_cgroup_do_sockets())
> + sock_release_memcg(sk);
> }
> EXPORT_SYMBOL(tcp_v4_destroy_sock);
>
> @@ -2330,11 +2330,6 @@ struct proto tcp_prot = {
> .compat_setsockopt = compat_tcp_setsockopt,
> .compat_getsockopt = compat_tcp_getsockopt,
> #endif
> -#ifdef CONFIG_MEMCG_KMEM
> - .init_cgroup = tcp_init_cgroup,
> - .destroy_cgroup = tcp_destroy_cgroup,
> - .proto_cgroup = tcp_proto_cgroup,
> -#endif
> };
> EXPORT_SYMBOL(tcp_prot);
>
> diff --git a/net/ipv4/tcp_memcontrol.c b/net/ipv4/tcp_memcontrol.c
> index 2379c1b..09a37eb 100644
> --- a/net/ipv4/tcp_memcontrol.c
> +++ b/net/ipv4/tcp_memcontrol.c
> @@ -1,107 +1,10 @@
> -#include <net/tcp.h>
> -#include <net/tcp_memcontrol.h>
> -#include <net/sock.h>
> -#include <net/ip.h>
> -#include <linux/nsproxy.h>
> +#include <linux/page_counter.h>
> #include <linux/memcontrol.h>
> +#include <linux/cgroup.h>
> #include <linux/module.h>
> -
> -int tcp_init_cgroup(struct mem_cgroup *memcg, struct cgroup_subsys *ss)
> -{
> - /*
> - * The root cgroup does not use page_counters, but rather,
> - * rely on the data already collected by the network
> - * subsystem
> - */
> - struct mem_cgroup *parent = parent_mem_cgroup(memcg);
> - struct page_counter *counter_parent = NULL;
> - struct cg_proto *cg_proto, *parent_cg;
> -
> - cg_proto = tcp_prot.proto_cgroup(memcg);
> - if (!cg_proto)
> - return 0;
> -
> - cg_proto->sysctl_mem[0] = sysctl_tcp_mem[0];
> - cg_proto->sysctl_mem[1] = sysctl_tcp_mem[1];
> - cg_proto->sysctl_mem[2] = sysctl_tcp_mem[2];
> - cg_proto->memory_pressure = 0;
> - cg_proto->memcg = memcg;
> -
> - parent_cg = tcp_prot.proto_cgroup(parent);
> - if (parent_cg)
> - counter_parent = &parent_cg->memory_allocated;
> -
> - page_counter_init(&cg_proto->memory_allocated, counter_parent);
> - percpu_counter_init(&cg_proto->sockets_allocated, 0, GFP_KERNEL);
> -
> - return 0;
> -}
> -EXPORT_SYMBOL(tcp_init_cgroup);
> -
> -void tcp_destroy_cgroup(struct mem_cgroup *memcg)
> -{
> - struct cg_proto *cg_proto;
> -
> - cg_proto = tcp_prot.proto_cgroup(memcg);
> - if (!cg_proto)
> - return;
> -
> - percpu_counter_destroy(&cg_proto->sockets_allocated);
> -
> - if (test_bit(MEMCG_SOCK_ACTIVATED, &cg_proto->flags))
> - static_key_slow_dec(&memcg_socket_limit_enabled);
> -
> -}
> -EXPORT_SYMBOL(tcp_destroy_cgroup);
> -
> -static int tcp_update_limit(struct mem_cgroup *memcg, unsigned long nr_pages)
> -{
> - struct cg_proto *cg_proto;
> - int i;
> - int ret;
> -
> - cg_proto = tcp_prot.proto_cgroup(memcg);
> - if (!cg_proto)
> - return -EINVAL;
> -
> - ret = page_counter_limit(&cg_proto->memory_allocated, nr_pages);
> - if (ret)
> - return ret;
> -
> - for (i = 0; i < 3; i++)
> - cg_proto->sysctl_mem[i] = min_t(long, nr_pages,
> - sysctl_tcp_mem[i]);
> -
> - if (nr_pages == PAGE_COUNTER_MAX)
> - clear_bit(MEMCG_SOCK_ACTIVE, &cg_proto->flags);
> - else {
> - /*
> - * The active bit needs to be written after the static_key
> - * update. This is what guarantees that the socket activation
> - * function is the last one to run. See sock_update_memcg() for
> - * details, and note that we don't mark any socket as belonging
> - * to this memcg until that flag is up.
> - *
> - * We need to do this, because static_keys will span multiple
> - * sites, but we can't control their order. If we mark a socket
> - * as accounted, but the accounting functions are not patched in
> - * yet, we'll lose accounting.
> - *
> - * We never race with the readers in sock_update_memcg(),
> - * because when this value change, the code to process it is not
> - * patched in yet.
> - *
> - * The activated bit is used to guarantee that no two writers
> - * will do the update in the same memcg. Without that, we can't
> - * properly shutdown the static key.
> - */
> - if (!test_and_set_bit(MEMCG_SOCK_ACTIVATED, &cg_proto->flags))
> - static_key_slow_inc(&memcg_socket_limit_enabled);
> - set_bit(MEMCG_SOCK_ACTIVE, &cg_proto->flags);
> - }
> -
> - return 0;
> -}
> +#include <linux/kernfs.h>
> +#include <linux/mutex.h>
> +#include <net/tcp.h>
>
> enum {
> RES_USAGE,
> @@ -124,11 +27,17 @@ static ssize_t tcp_cgroup_write(struct kernfs_open_file *of,
> switch (of_cft(of)->private) {
> case RES_LIMIT:
> /* see memcontrol.c */
> + if (memcg == root_mem_cgroup) {
> + ret = -EINVAL;
> + break;
> + }
> ret = page_counter_memparse(buf, "-1", &nr_pages);
> if (ret)
> break;
> mutex_lock(&tcp_limit_mutex);
> - ret = tcp_update_limit(memcg, nr_pages);
> + ret = page_counter_limit(&memcg->skmem, nr_pages);
> + if (!ret)
> + static_branch_enable(&mem_cgroup_sockets);
> mutex_unlock(&tcp_limit_mutex);
> break;
> default:
> @@ -141,32 +50,28 @@ static ssize_t tcp_cgroup_write(struct kernfs_open_file *of,
> static u64 tcp_cgroup_read(struct cgroup_subsys_state *css, struct cftype *cft)
> {
> struct mem_cgroup *memcg = mem_cgroup_from_css(css);
> - struct cg_proto *cg_proto = tcp_prot.proto_cgroup(memcg);
> u64 val;
>
> switch (cft->private) {
> case RES_LIMIT:
> - if (!cg_proto)
> - return PAGE_COUNTER_MAX;
> - val = cg_proto->memory_allocated.limit;
> + val = memcg->skmem.limit;
> val *= PAGE_SIZE;
> break;
> case RES_USAGE:
> - if (!cg_proto)
> + if (memcg == root_mem_cgroup)
> val = atomic_long_read(&tcp_memory_allocated);
> else
> - val = page_counter_read(&cg_proto->memory_allocated);
> + val = page_counter_read(&memcg->skmem);
> val *= PAGE_SIZE;
> break;
> case RES_FAILCNT:
> - if (!cg_proto)
> - return 0;
> - val = cg_proto->memory_allocated.failcnt;
> + val = memcg->skmem.failcnt;
> break;
> case RES_MAX_USAGE:
> - if (!cg_proto)
> - return 0;
> - val = cg_proto->memory_allocated.watermark;
> + if (memcg == root_mem_cgroup)
> + val = 0;
> + else
> + val = memcg->skmem.watermark;
> val *= PAGE_SIZE;
> break;
> default:
> @@ -178,20 +83,14 @@ static u64 tcp_cgroup_read(struct cgroup_subsys_state *css, struct cftype *cft)
> static ssize_t tcp_cgroup_reset(struct kernfs_open_file *of,
> char *buf, size_t nbytes, loff_t off)
> {
> - struct mem_cgroup *memcg;
> - struct cg_proto *cg_proto;
> -
> - memcg = mem_cgroup_from_css(of_css(of));
> - cg_proto = tcp_prot.proto_cgroup(memcg);
> - if (!cg_proto)
> - return nbytes;
> + struct mem_cgroup *memcg = mem_cgroup_from_css(of_css(of));
>
> switch (of_cft(of)->private) {
> case RES_MAX_USAGE:
> - page_counter_reset_watermark(&cg_proto->memory_allocated);
> + page_counter_reset_watermark(&memcg->skmem);
> break;
> case RES_FAILCNT:
> - cg_proto->memory_allocated.failcnt = 0;
> + memcg->skmem.failcnt = 0;
> break;
> }
>
> diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c
> index 19adedb..b496fc9 100644
> --- a/net/ipv4/tcp_output.c
> +++ b/net/ipv4/tcp_output.c
> @@ -2819,13 +2819,15 @@ begin_fwd:
> */
> void sk_forced_mem_schedule(struct sock *sk, int size)
> {
> - int amt, status;
> + int amt;
>
> if (size <= sk->sk_forward_alloc)
> return;
> amt = sk_mem_pages(size);
> sk->sk_forward_alloc += amt * SK_MEM_QUANTUM;
> - sk_memory_allocated_add(sk, amt, &status);
> + sk_memory_allocated_add(sk, amt);
> + if (mem_cgroup_do_sockets() && sk->sk_memcg)
> + mem_cgroup_charge_skmem(sk->sk_memcg, amt);
> }
>
> /* Send a FIN. The caller locks the socket for us.
> diff --git a/net/ipv6/tcp_ipv6.c b/net/ipv6/tcp_ipv6.c
> index f495d18..cf19e65 100644
> --- a/net/ipv6/tcp_ipv6.c
> +++ b/net/ipv6/tcp_ipv6.c
> @@ -1862,9 +1862,6 @@ struct proto tcpv6_prot = {
> .compat_setsockopt = compat_tcp_setsockopt,
> .compat_getsockopt = compat_tcp_getsockopt,
> #endif
> -#ifdef CONFIG_MEMCG_KMEM
> - .proto_cgroup = tcp_proto_cgroup,
> -#endif
> .clear_sk = tcp_v6_clear_sk,
> };
>
> --
> 2.6.1
--
Michal Hocko
SUSE Labs
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@virtuozzo.com> |
|---|---|
| Date | 2015-10-22 20:50 +0200 |
| Subject | Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy |
| Message-ID | <qmt62-2wa-19@gated-at.bofh.it> |
| In reply to | #1253466 |
Hi Johannes, On Thu, Oct 22, 2015 at 12:21:28AM -0400, Johannes Weiner wrote: ... > Patch #5 adds accounting and tracking of socket memory to the unified > hierarchy memory controller, as described above. It uses the existing > per-cpu charge caches and triggers high limit reclaim asynchroneously. > > Patch #8 uses the vmpressure extension to equalize pressure between > the pages tracked natively by the VM and socket buffer pages. As the > pool is shared, it makes sense that while natively tracked pages are > under duress the network transmit windows are also not increased. First of all, I've no experience in networking, so I'm likely to be mistaken. Nevertheless I beg to disagree that this patch set is a step in the right direction. Here goes why. I admit that your idea to get rid of explicit tcp window control knobs and size it dynamically basing on memory pressure instead does sound tempting, but I don't think it'd always work. The problem is that in contrast to, say, dcache, we can't shrink tcp buffers AFAIU, we can only stop growing them. Now suppose a system hasn't experienced memory pressure for a while. If we don't have explicit tcp window limit, tcp buffers on such a system might have eaten almost all available memory (because of network load/problems). If a user workload that needs a significant amount of memory is started suddenly then, the network code will receive a notification and surely stop growing buffers, but all those buffers accumulated won't disappear instantly. As a result, the workload might be unable to find enough free memory and have no choice but invoke OOM killer. This looks unexpected from the user POV. That said, I think we do need per memcg tcp window control similar to what we have system-wide. In other words, Glauber's work makes sense to me. You might want to point me at my RFC patch where I proposed to revert it (https://lkml.org/lkml/2014/9/12/401). Well, I've changed my mind since then. Now I think I was mistaken, luckily I was stopped. However, I may be mistaken again :-) Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2015-10-26 18:30 +0100 |
| Subject | Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy |
| Message-ID | <qnTKO-2CZ-11@gated-at.bofh.it> |
| In reply to | #1254082 |
On Thu, Oct 22, 2015 at 09:45:10PM +0300, Vladimir Davydov wrote: > Hi Johannes, > > On Thu, Oct 22, 2015 at 12:21:28AM -0400, Johannes Weiner wrote: > ... > > Patch #5 adds accounting and tracking of socket memory to the unified > > hierarchy memory controller, as described above. It uses the existing > > per-cpu charge caches and triggers high limit reclaim asynchroneously. > > > > Patch #8 uses the vmpressure extension to equalize pressure between > > the pages tracked natively by the VM and socket buffer pages. As the > > pool is shared, it makes sense that while natively tracked pages are > > under duress the network transmit windows are also not increased. > > First of all, I've no experience in networking, so I'm likely to be > mistaken. Nevertheless I beg to disagree that this patch set is a step > in the right direction. Here goes why. > > I admit that your idea to get rid of explicit tcp window control knobs > and size it dynamically basing on memory pressure instead does sound > tempting, but I don't think it'd always work. The problem is that in > contrast to, say, dcache, we can't shrink tcp buffers AFAIU, we can only > stop growing them. Now suppose a system hasn't experienced memory > pressure for a while. If we don't have explicit tcp window limit, tcp > buffers on such a system might have eaten almost all available memory > (because of network load/problems). If a user workload that needs a > significant amount of memory is started suddenly then, the network code > will receive a notification and surely stop growing buffers, but all > those buffers accumulated won't disappear instantly. As a result, the > workload might be unable to find enough free memory and have no choice > but invoke OOM killer. This looks unexpected from the user POV. I'm not getting rid of those knobs, I'm just reusing the old socket accounting infrastructure in an attempt to make the memory accounting feature useful to more people in cgroups v2 (unified hierarchy). We can always come back to think about per-cgroup tcp window limits in the unified hierarchy, my patches don't get in the way of this. I'm not removing the knobs in cgroups v1 and I'm not preventing them in v2. But regardless of tcp window control, we need to account socket memory in the main memory accounting pool where pressure is shared (to the best of our abilities) between all accounted memory consumers. From an interface standpoint alone, I don't think it's reasonable to ask users per default to limit different consumers on a case by case basis. I certainly have no problem with finetuning for scenarios you describe above, but with memory.current, memory.high, memory.max we are providing a generic interface to account and contain memory consumption of workloads. This has to include all major memory consumers to make semantical sense. But also, there are people right now for whom the socket buffers cause system OOM, but the existing memcg's hard tcp window limitq that exists absolutely wrecks network performance for them. It's not usable the way it is. It'd be much better to have the socket buffers exert pressure on the shared pool, and then propagate the overall pressure back to individual consumers with reclaim, shrinkers, vmpressure etc. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@virtuozzo.com> |
|---|---|
| Date | 2015-10-27 09:50 +0100 |
| Subject | Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy |
| Message-ID | <qo878-30x-11@gated-at.bofh.it> |
| In reply to | #1256180 |
On Mon, Oct 26, 2015 at 01:22:16PM -0400, Johannes Weiner wrote: > On Thu, Oct 22, 2015 at 09:45:10PM +0300, Vladimir Davydov wrote: > > Hi Johannes, > > > > On Thu, Oct 22, 2015 at 12:21:28AM -0400, Johannes Weiner wrote: > > ... > > > Patch #5 adds accounting and tracking of socket memory to the unified > > > hierarchy memory controller, as described above. It uses the existing > > > per-cpu charge caches and triggers high limit reclaim asynchroneously. > > > > > > Patch #8 uses the vmpressure extension to equalize pressure between > > > the pages tracked natively by the VM and socket buffer pages. As the > > > pool is shared, it makes sense that while natively tracked pages are > > > under duress the network transmit windows are also not increased. > > > > First of all, I've no experience in networking, so I'm likely to be > > mistaken. Nevertheless I beg to disagree that this patch set is a step > > in the right direction. Here goes why. > > > > I admit that your idea to get rid of explicit tcp window control knobs > > and size it dynamically basing on memory pressure instead does sound > > tempting, but I don't think it'd always work. The problem is that in > > contrast to, say, dcache, we can't shrink tcp buffers AFAIU, we can only > > stop growing them. Now suppose a system hasn't experienced memory > > pressure for a while. If we don't have explicit tcp window limit, tcp > > buffers on such a system might have eaten almost all available memory > > (because of network load/problems). If a user workload that needs a > > significant amount of memory is started suddenly then, the network code > > will receive a notification and surely stop growing buffers, but all > > those buffers accumulated won't disappear instantly. As a result, the > > workload might be unable to find enough free memory and have no choice > > but invoke OOM killer. This looks unexpected from the user POV. > > I'm not getting rid of those knobs, I'm just reusing the old socket > accounting infrastructure in an attempt to make the memory accounting > feature useful to more people in cgroups v2 (unified hierarchy). > My understanding is that in the meantime you effectively break the existing per memcg tcp window control logic. > We can always come back to think about per-cgroup tcp window limits in > the unified hierarchy, my patches don't get in the way of this. I'm > not removing the knobs in cgroups v1 and I'm not preventing them in v2. > > But regardless of tcp window control, we need to account socket memory > in the main memory accounting pool where pressure is shared (to the > best of our abilities) between all accounted memory consumers. > No objections to this point. However, I really don't like the idea to charge tcp window size to memory.current instead of charging individual pages consumed by the workload for storing socket buffers, because it is inconsistent with what we have now. Can't we charge individual skb pages as we do in case of other kmem allocations? > From an interface standpoint alone, I don't think it's reasonable to > ask users per default to limit different consumers on a case by case > basis. I certainly have no problem with finetuning for scenarios you > describe above, but with memory.current, memory.high, memory.max we > are providing a generic interface to account and contain memory > consumption of workloads. This has to include all major memory > consumers to make semantical sense. We can propose a reasonable default as we do in the global case. > > But also, there are people right now for whom the socket buffers cause > system OOM, but the existing memcg's hard tcp window limitq that > exists absolutely wrecks network performance for them. It's not usable > the way it is. It'd be much better to have the socket buffers exert > pressure on the shared pool, and then propagate the overall pressure > back to individual consumers with reclaim, shrinkers, vmpressure etc. > This might or might not work. I'm not an expert to judge. But if you do this only for memcg leaving the global case as it is, networking people won't budge IMO. So could you please start such a major rework from the global case? Could you please try to deprecate the tcp window limits not only in the legacy memcg hierarchy, but also system-wide in order to attract attention of networking experts? Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2015-10-27 17:10 +0100 |
| Subject | Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy |
| Message-ID | <qoeYX-7jh-57@gated-at.bofh.it> |
| In reply to | #1256577 |
On Tue, Oct 27, 2015 at 11:43:21AM +0300, Vladimir Davydov wrote: > On Mon, Oct 26, 2015 at 01:22:16PM -0400, Johannes Weiner wrote: > > I'm not getting rid of those knobs, I'm just reusing the old socket > > accounting infrastructure in an attempt to make the memory accounting > > feature useful to more people in cgroups v2 (unified hierarchy). > > My understanding is that in the meantime you effectively break the > existing per memcg tcp window control logic. That's not my intention, this stuff has to keep working. I'm assuming you mean the changes to sk_enter_memory_pressure() when hitting the charge limit; let me address this in the other subthread. > > We can always come back to think about per-cgroup tcp window limits in > > the unified hierarchy, my patches don't get in the way of this. I'm > > not removing the knobs in cgroups v1 and I'm not preventing them in v2. > > > > But regardless of tcp window control, we need to account socket memory > > in the main memory accounting pool where pressure is shared (to the > > best of our abilities) between all accounted memory consumers. > > > > No objections to this point. However, I really don't like the idea to > charge tcp window size to memory.current instead of charging individual > pages consumed by the workload for storing socket buffers, because it is > inconsistent with what we have now. Can't we charge individual skb pages > as we do in case of other kmem allocations? Absolutely, both work for me. I chose that route because it's where the networking code already tracks and accounts memory consumed, so it seemed like a better site to hook into. But I understand your concerns. We want to track this stuff as close to the memory allocators as possible. > > But also, there are people right now for whom the socket buffers cause > > system OOM, but the existing memcg's hard tcp window limitq that > > exists absolutely wrecks network performance for them. It's not usable > > the way it is. It'd be much better to have the socket buffers exert > > pressure on the shared pool, and then propagate the overall pressure > > back to individual consumers with reclaim, shrinkers, vmpressure etc. > > This might or might not work. I'm not an expert to judge. But if you do > this only for memcg leaving the global case as it is, networking people > won't budge IMO. So could you please start such a major rework from the > global case? Could you please try to deprecate the tcp window limits not > only in the legacy memcg hierarchy, but also system-wide in order to > attract attention of networking experts? I'm definitely interested in addressing this globally as well. The idea behind this was to use the memcg part as a testbed. cgroup2 is going to be new and people are prepared for hiccups when migrating their applications to it; and they can roll back to cgroup1 and tcp window limits at any time should they run into problems in production. So this seemed like a good way to prove a new mechanism before rolling it out to every single Linux setup, rather than switch everybody over after the limited scope testing I can do as a developer on my own. Keep in mind that my patches are not committing anything in terms of interface, so we retain all the freedom to fix and tune the way this is implemented, including the freedom to re-add tcp window limits in case the pressure balancing is not a comprehensive solution. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@virtuozzo.com> |
|---|---|
| Date | 2015-10-28 09:30 +0100 |
| Subject | Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy |
| Message-ID | <qouhl-af-37@gated-at.bofh.it> |
| In reply to | #1256890 |
On Tue, Oct 27, 2015 at 09:01:08AM -0700, Johannes Weiner wrote: ... > > > But regardless of tcp window control, we need to account socket memory > > > in the main memory accounting pool where pressure is shared (to the > > > best of our abilities) between all accounted memory consumers. > > > > > > > No objections to this point. However, I really don't like the idea to > > charge tcp window size to memory.current instead of charging individual > > pages consumed by the workload for storing socket buffers, because it is > > inconsistent with what we have now. Can't we charge individual skb pages > > as we do in case of other kmem allocations? > > Absolutely, both work for me. I chose that route because it's where > the networking code already tracks and accounts memory consumed, so it > seemed like a better site to hook into. > > But I understand your concerns. We want to track this stuff as close > to the memory allocators as possible. Exactly. > > > > But also, there are people right now for whom the socket buffers cause > > > system OOM, but the existing memcg's hard tcp window limitq that > > > exists absolutely wrecks network performance for them. It's not usable > > > the way it is. It'd be much better to have the socket buffers exert > > > pressure on the shared pool, and then propagate the overall pressure > > > back to individual consumers with reclaim, shrinkers, vmpressure etc. > > > > This might or might not work. I'm not an expert to judge. But if you do > > this only for memcg leaving the global case as it is, networking people > > won't budge IMO. So could you please start such a major rework from the > > global case? Could you please try to deprecate the tcp window limits not > > only in the legacy memcg hierarchy, but also system-wide in order to > > attract attention of networking experts? > > I'm definitely interested in addressing this globally as well. > > The idea behind this was to use the memcg part as a testbed. cgroup2 > is going to be new and people are prepared for hiccups when migrating > their applications to it; and they can roll back to cgroup1 and tcp > window limits at any time should they run into problems in production. Then you'd better not touch existing tcp limits at all, because they just work, and the logic behind them is very close to that of global tcp limits. I don't think one can simplify it somehow. Moreover, frankly I still have my reservations about this vmpressure propagation to skb you're proposing. It might work, but I doubt it will allow us to throw away explicit tcp limit, as I explained previously. So, even with your approach I think we can still need per memcg tcp limit *unless* you get rid of global tcp limit somehow. > > So this seemed like a good way to prove a new mechanism before rolling > it out to every single Linux setup, rather than switch everybody over > after the limited scope testing I can do as a developer on my own. > > Keep in mind that my patches are not committing anything in terms of > interface, so we retain all the freedom to fix and tune the way this > is implemented, including the freedom to re-add tcp window limits in > case the pressure balancing is not a comprehensive solution. > I really dislike this kind of proof. It looks like you're trying to push something you think is right covertly, w/o having a proper discussion with networking people and then say that it just works and hence should be done globally, but what if it won't? Revert it? We already have a lot of dubious stuff in memcg that should be reverted, so let's please try to avoid this kind of mistakes in future. Note, I say "w/o having a proper discussion with networking people", because I don't think they will really care *unless* you change the global logic, simply because most of them aren't very interested in memcg AFAICS. That effectively means you loose a chance to listen to networking experts, who could point you at design flaws and propose an improvement right away. Let's please not miss such an opportunity. You said that you'd seen this problem happen w/o cgroups, so you have a use case that might need fixing at the global level. IMO it shouldn't be difficult to prepare an RFC patch for the global case first and see what people think about it. Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2015-10-28 20:00 +0100 |
| Subject | Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy |
| Message-ID | <qoE6Z-6pV-9@gated-at.bofh.it> |
| In reply to | #1257787 |
On Wed, Oct 28, 2015 at 11:20:03AM +0300, Vladimir Davydov wrote: > Then you'd better not touch existing tcp limits at all, because they > just work, and the logic behind them is very close to that of global tcp > limits. I don't think one can simplify it somehow. Uhm, no, there is a crapload of boilerplate code and complication that seems entirely unnecessary. The only thing missing from my patch seems to be the part where it enters memory pressure state when the limit is hit. I'm adding this for completeness, but I doubt it even matters. > Moreover, frankly I still have my reservations about this vmpressure > propagation to skb you're proposing. It might work, but I doubt it > will allow us to throw away explicit tcp limit, as I explained > previously. So, even with your approach I think we can still need > per memcg tcp limit *unless* you get rid of global tcp limit > somehow. Having the hard limit as a failsafe (or a minimum for other consumers) is one thing, and certainly something I'm open to for cgroupv2, should we have problems with load startup up after a socket memory landgrab. That being said, if the VM is struggling to reclaim pages, or is even swapping, it makes perfect sense to let the socket memory scheduler know it shouldn't continue to increase its footprint until the VM recovers. Regardless of any hard limitations/minimum guarantees. This is what my patch does and it seems pretty straight-forward to me. I don't really understand why this is so controversial. The *next* step would be to figure out whether we can actually *reclaim* memory in the network subsystem--shrink windows and steal buffers back--and that might even be an avenue to replace tcp window limits. But it's not necessary for *this* patch series to be useful. > > So this seemed like a good way to prove a new mechanism before rolling > > it out to every single Linux setup, rather than switch everybody over > > after the limited scope testing I can do as a developer on my own. > > > > Keep in mind that my patches are not committing anything in terms of > > interface, so we retain all the freedom to fix and tune the way this > > is implemented, including the freedom to re-add tcp window limits in > > case the pressure balancing is not a comprehensive solution. > > I really dislike this kind of proof. It looks like you're trying to > push something you think is right covertly, w/o having a proper > discussion with networking people and then say that it just works > and hence should be done globally, but what if it won't? Revert it? > We already have a lot of dubious stuff in memcg that should be > reverted, so let's please try to avoid this kind of mistakes in > future. Note, I say "w/o having a proper discussion with networking > people", because I don't think they will really care *unless* you > change the global logic, simply because most of them aren't very > interested in memcg AFAICS. Come on, Dave is the first To and netdev is CC'd. They might not care about memcg, but "pushing things covertly" is a bit of a stretch. > That effectively means you loose a chance to listen to networking > experts, who could point you at design flaws and propose an improvement > right away. Let's please not miss such an opportunity. You said that > you'd seen this problem happen w/o cgroups, so you have a use case that > might need fixing at the global level. IMO it shouldn't be difficult to > prepare an RFC patch for the global case first and see what people think > about it. No, the problem we are running into is when network memory is not tracked per cgroup. The lack of containment means that the socket memory consumption of individual cgroups can trigger system OOM. We tried using the per-memcg tcp limits, and that prevents the OOMs for sure, but it's horrendous for network performance. There is no "stop growing" phase, it just keeps going full throttle until it hits the wall hard. Now, we could probably try to replicate the global knobs and add a per-memcg soft limit. But you know better than anyone else how hard it is to estimate the overall workingset size of a workload, and the margins on containerized loads are razor-thin. Performance is much more sensitive to input errors, and often times parameters must be adjusted continuously during the runtime of a workload. It'd be disasterous to rely on yet more static, error-prone user input here. What all this means to me is that fixing it on the cgroup level has higher priority. But it also means that once we figured it out under such a high-pressure environment, it's much easier to apply to the global case and potentially replace the soft limit there. This seems like a better approach to me than starting globally, only to realize that the solution is not workable for cgroups and we need yet something else. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vladimir Davydov <vdavydov@virtuozzo.com> |
|---|---|
| Date | 2015-10-29 10:30 +0100 |
| Subject | Re: [PATCH 0/8] mm: memcontrol: account socket memory in unified hierarchy |
| Message-ID | <qoRGW-6NC-29@gated-at.bofh.it> |
| In reply to | #1258373 |
On Wed, Oct 28, 2015 at 11:58:10AM -0700, Johannes Weiner wrote: > On Wed, Oct 28, 2015 at 11:20:03AM +0300, Vladimir Davydov wrote: > > Then you'd better not touch existing tcp limits at all, because they > > just work, and the logic behind them is very close to that of global tcp > > limits. I don't think one can simplify it somehow. > > Uhm, no, there is a crapload of boilerplate code and complication that > seems entirely unnecessary. The only thing missing from my patch seems > to be the part where it enters memory pressure state when the limit is > hit. I'm adding this for completeness, but I doubt it even matters. > > > Moreover, frankly I still have my reservations about this vmpressure > > propagation to skb you're proposing. It might work, but I doubt it > > will allow us to throw away explicit tcp limit, as I explained > > previously. So, even with your approach I think we can still need > > per memcg tcp limit *unless* you get rid of global tcp limit > > somehow. > > Having the hard limit as a failsafe (or a minimum for other consumers) > is one thing, and certainly something I'm open to for cgroupv2, should > we have problems with load startup up after a socket memory landgrab. > > That being said, if the VM is struggling to reclaim pages, or is even > swapping, it makes perfect sense to let the socket memory scheduler > know it shouldn't continue to increase its footprint until the VM > recovers. Regardless of any hard limitations/minimum guarantees. > > This is what my patch does and it seems pretty straight-forward to > me. I don't really understand why this is so controversial. I'm not arguing that the idea behind this patch set is necessarily bad. Quite the contrary, it does look interesting to me. I'm just saying that IMO it can't replace hard/soft limits. It probably could if it was possible to shrink buffers, but I don't think it's feasible, even theoretically. That's why I propose not to change the behavior of the existing per memcg tcp limit at all. And frankly I don't get why you are so keen on simplifying it. You say it's a "crapload of boilerplate code". Well, I don't see how it is - it just replicates global knobs and I don't see how it could be done in a better way. The code is hidden behind jump labels, so the overhead is zero if it isn't used. If you really dislike this code, we can isolate it under a separate config option. But all right, I don't rule out the possibility that the code could be simplified. If you do that w/o breaking it, that'll be OK to me, but I don't see why it should be related to this particular patch set. > > The *next* step would be to figure out whether we can actually > *reclaim* memory in the network subsystem--shrink windows and steal > buffers back--and that might even be an avenue to replace tcp window > limits. But it's not necessary for *this* patch series to be useful. Again, I don't think we can *reclaim* network memory, but you're right. > > > > So this seemed like a good way to prove a new mechanism before rolling > > > it out to every single Linux setup, rather than switch everybody over > > > after the limited scope testing I can do as a developer on my own. > > > > > > Keep in mind that my patches are not committing anything in terms of > > > interface, so we retain all the freedom to fix and tune the way this > > > is implemented, including the freedom to re-add tcp window limits in > > > case the pressure balancing is not a comprehensive solution. > > > > I really dislike this kind of proof. It looks like you're trying to > > push something you think is right covertly, w/o having a proper > > discussion with networking people and then say that it just works > > and hence should be done globally, but what if it won't? Revert it? > > We already have a lot of dubious stuff in memcg that should be > > reverted, so let's please try to avoid this kind of mistakes in > > future. Note, I say "w/o having a proper discussion with networking > > people", because I don't think they will really care *unless* you > > change the global logic, simply because most of them aren't very > > interested in memcg AFAICS. > > Come on, Dave is the first To and netdev is CC'd. They might not care > about memcg, but "pushing things covertly" is a bit of a stretch. Sorry if it sounded rude to you. I just look back at my experience patching slab internals to make kmem accountable, and AFAICS Christoph didn't really care about *what* I was doing, he only cared about the global case - if there was no performance degradation when kmemcg was disabled, he was usually fine with it, even if from the memcg pov it was a crap. Anyway, I can't force you to patch the global case first or simultaneously with the memcg case, so let's just hope I'm a bit too overcautious. > > > That effectively means you loose a chance to listen to networking > > experts, who could point you at design flaws and propose an improvement > > right away. Let's please not miss such an opportunity. You said that > > you'd seen this problem happen w/o cgroups, so you have a use case that > > might need fixing at the global level. IMO it shouldn't be difficult to > > prepare an RFC patch for the global case first and see what people think > > about it. > > No, the problem we are running into is when network memory is not > tracked per cgroup. The lack of containment means that the socket > memory consumption of individual cgroups can trigger system OOM. > > We tried using the per-memcg tcp limits, and that prevents the OOMs > for sure, but it's horrendous for network performance. There is no > "stop growing" phase, it just keeps going full throttle until it hits > the wall hard. > > Now, we could probably try to replicate the global knobs and add a > per-memcg soft limit. But you know better than anyone else how hard it > is to estimate the overall workingset size of a workload, and the > margins on containerized loads are razor-thin. Performance is much > more sensitive to input errors, and often times parameters must be > adjusted continuously during the runtime of a workload. It'd be > disasterous to rely on yet more static, error-prone user input here. Yeah, but the dynamic approach proposed in your patch set doesn't guarantee we won't hit OOM in memcg due to overgrown buffers. It just reduces this possibility. Of course, memcg OOM is far not as disastrous as the global one, but still it usually means the workload breakage. The static approach is error-prone for sure, but it has existed for years and worked satisfactory AFAIK. > > What all this means to me is that fixing it on the cgroup level has > higher priority. But it also means that once we figured it out under > such a high-pressure environment, it's much easier to apply to the > global case and potentially replace the soft limit there. > > This seems like a better approach to me than starting globally, only > to realize that the solution is not workable for cgroups and we need > yet something else. > Are we in rush? I think if you try your approach at the global level and fail, it's still good, because it will probably give us all a better understanding of the problem. If you successfully fix the global case, but then realize that it doesn't fit memcg, it's even better, because you actually fixed a problem. If you patch both global and memcg cases, it's perfect. But of course, that's my understanding and I may be mistaken. Let's hope you're right. Thanks, Vladimir -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web