Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1288457 > unrolled thread

[PATCH 0/7] Add swap accounting to cgroup2

Started byVladimir Davydov <vdavydov@virtuozzo.com>
First post2015-12-10 12:40 +0100
Last post2015-12-12 17:20 +0100
Articles 17 on this page of 37 — 5 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH 0/7] Add swap accounting to cgroup2 Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-10 12:40 +0100
    [PATCH 4/7] swap.h: move memcg related stuff to the end of the file Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-10 12:40 +0100
      Re: [PATCH 4/7] swap.h: move memcg related stuff to the end of the  file Johannes Weiner <hannes@cmpxchg.org> - 2015-12-11 20:30 +0100
    [PATCH 3/7] mm: memcontrol: replace mem_cgroup_lruvec_online with mem_cgroup_online Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-10 12:40 +0100
      Re: [PATCH 3/7] mm: memcontrol: replace mem_cgroup_lruvec_online  with mem_cgroup_online Johannes Weiner <hannes@cmpxchg.org> - 2015-12-11 20:30 +0100
    [PATCH 2/7] mm: vmscan: pass memcg to get_scan_count() Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-10 12:40 +0100
      Re: [PATCH 2/7] mm: vmscan: pass memcg to get_scan_count() Johannes Weiner <hannes@cmpxchg.org> - 2015-12-11 20:30 +0100
    [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-10 12:40 +0100
      Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Johannes Weiner <hannes@cmpxchg.org> - 2015-12-10 17:10 +0100
        Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-10 18:10 +0100
      Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Kamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com> - 2015-12-11 04:00 +0100
        Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-11 08:50 +0100
      Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Michal Hocko <mhocko@kernel.org> - 2015-12-14 16:40 +0100
        Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Johannes Weiner <hannes@cmpxchg.org> - 2015-12-14 16:50 +0100
        Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-14 20:50 +0100
          Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 One Thousand Gnomes <gnomes@lxorguk.ukuu.org.uk> - 2015-12-14 21:00 +0100
          Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Kamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com> - 2015-12-15 04:30 +0100
            Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-15 12:10 +0100
              Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Kamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com> - 2015-12-16 03:50 +0100
            Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Johannes Weiner <hannes@cmpxchg.org> - 2015-12-15 16:00 +0100
              Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Kamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com> - 2015-12-16 04:20 +0100
                Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Johannes Weiner <hannes@cmpxchg.org> - 2015-12-16 12:10 +0100
                  Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Kamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com> - 2015-12-17 03:50 +0100
                    Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Johannes Weiner <hannes@cmpxchg.org> - 2015-12-17 04:40 +0100
                      Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Kamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com> - 2015-12-17 05:40 +0100
          Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Michal Hocko <mhocko@kernel.org> - 2015-12-15 18:30 +0100
            Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Johannes Weiner <hannes@cmpxchg.org> - 2015-12-15 21:30 +0100
            Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Kamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com> - 2015-12-16 05:00 +0100
        Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Kamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com> - 2015-12-15 04:30 +0100
          Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-15 09:40 +0100
            Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2 Kamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com> - 2015-12-15 10:40 +0100
    [PATCH 7/7] Documentation: cgroup: add memory.swap.{current,max} description Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-10 12:50 +0100
      Re: [PATCH 7/7] Documentation: cgroup: add memory.swap.{current,max}  description Johannes Weiner <hannes@cmpxchg.org> - 2015-12-11 20:50 +0100
        Re: [PATCH 7/7] Documentation: cgroup: add memory.swap.{current,max}  description Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-12 17:20 +0100
    [PATCH 6/7] mm: free swap cache aggressively if memcg swap is full Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-10 12:50 +0100
      Re: [PATCH 6/7] mm: free swap cache aggressively if memcg swap is  full Johannes Weiner <hannes@cmpxchg.org> - 2015-12-11 20:40 +0100
        Re: [PATCH 6/7] mm: free swap cache aggressively if memcg swap is  full Vladimir Davydov <vdavydov@virtuozzo.com> - 2015-12-12 17:20 +0100

Page 2 of 2 — ← Prev page 1 [2]


#1292719 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromKamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com>
Date2015-12-16 04:20 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qGaNc-4Jx-11@gated-at.bofh.it>
In reply to#1292237
On 2015/12/15 23:50, Johannes Weiner wrote:
> On Tue, Dec 15, 2015 at 12:22:41PM +0900, Kamezawa Hiroyuki wrote:
>> On 2015/12/15 4:42, Vladimir Davydov wrote:
>>> Anyway, if you don't trust a container you'd better set the hard memory
>>> limit so that it can't hurt others no matter what it runs and how it
>>> tweaks its sub-tree knobs.
>>
>> Limiting swap can easily cause "OOM-Killer even while there are available swap"
>> with easy mistake. Can't you add "swap excess" switch to sysctl to allow global
>> memory reclaim can ignore swap limitation ?
>
> That never worked with a combined memory+swap limit, either. How could
> it? The parent might swap you out under pressure, but simply touching
> a few of your anon pages causes them to get swapped back in, thrashing
> with whatever the parent was trying to do. Your ability to swap it out
> is simply no protection against a group touching its pages.
>
> Allowing the parent to exceed swap with separate counters makes even
> less sense, because every page swapped out frees up a page of memory
> that the child can reuse. For every swap page that exceeds the limit,
> the child gets a free memory page! The child doesn't even have to
> cause swapin, it can just steal whatever the parent tried to free up,
> and meanwhile its combined memory & swap footprint explodes.
>
Sure.

> The answer is and always should have been: don't overcommit untrusted
> cgroups. Think of swap as a resource you distribute, not as breathing
> room for the parents to rely on. Because it can't and could never.
>
ok, don't overcommmit.

> And the new separate swap counter makes this explicit.
>
Hmm, my requests are
  - set the same capabilities as mlock() to set swap.limit=0
  - swap-full notification via vmpressure or something mechanism.
  - OOM-Killer's available memory calculation may be corrupted, please check.
  - force swap-in at reducing swap.limit

Thanks,
-Kame


--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1292899 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromJohannes Weiner <hannes@cmpxchg.org>
Date2015-12-16 12:10 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qGi82-Za-9@gated-at.bofh.it>
In reply to#1292719
On Wed, Dec 16, 2015 at 12:18:30PM +0900, Kamezawa Hiroyuki wrote:
> Hmm, my requests are
>  - set the same capabilities as mlock() to set swap.limit=0

Setting swap.max is already privileged operation.

>  - swap-full notification via vmpressure or something mechanism.

Why?

>  - OOM-Killer's available memory calculation may be corrupted, please check.

Vladimir updated mem_cgroup_get_limit().

>  - force swap-in at reducing swap.limit

Why?
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1293563 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromKamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com>
Date2015-12-17 03:50 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qGwNI-1K8-7@gated-at.bofh.it>
In reply to#1292899
On 2015/12/16 20:09, Johannes Weiner wrote:
> On Wed, Dec 16, 2015 at 12:18:30PM +0900, Kamezawa Hiroyuki wrote:
>> Hmm, my requests are
>>   - set the same capabilities as mlock() to set swap.limit=0
>
> Setting swap.max is already privileged operation.
>
Sure.

>>   - swap-full notification via vmpressure or something mechanism.
>
> Why?
>

I think it's a sign of unhealthy condition, starting file cache drop rate to rise.
But I forgot that there are resource threshold notifier already. Does the notifier work
for swap.usage ?

>>   - OOM-Killer's available memory calculation may be corrupted, please check.
>
> Vladimir updated mem_cgroup_get_limit().
>
I'll check it.

>>   - force swap-in at reducing swap.limit
>
> Why?
>
If full, swap.limit cannot be reduced even if there are available memory in a cgroup.
Another cgroup cannot make use of the swap resource while it's occupied by other cgroup.
The job scheduler should have a chance to fix the situation.

Thanks,
-Kame

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1293595 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromJohannes Weiner <hannes@cmpxchg.org>
Date2015-12-17 04:40 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qGxA6-2li-7@gated-at.bofh.it>
In reply to#1293563
On Thu, Dec 17, 2015 at 11:46:27AM +0900, Kamezawa Hiroyuki wrote:
> On 2015/12/16 20:09, Johannes Weiner wrote:
> >On Wed, Dec 16, 2015 at 12:18:30PM +0900, Kamezawa Hiroyuki wrote:
> >>  - swap-full notification via vmpressure or something mechanism.
> >
> >Why?
> >
> 
> I think it's a sign of unhealthy condition, starting file cache drop rate to rise.
> But I forgot that there are resource threshold notifier already. Does the notifier work
> for swap.usage ?

That will be reflected in vmpressure or other distress mechanisms. I'm
not convinced "ran out of swap space" needs special casing in any way.

> >>  - force swap-in at reducing swap.limit
> >
> >Why?
> >
> If full, swap.limit cannot be reduced even if there are available memory in a cgroup.
> Another cgroup cannot make use of the swap resource while it's occupied by other cgroup.
> The job scheduler should have a chance to fix the situation.

I don't see why swap space allowance would need to be as dynamically
adjustable as the memory allowance. There is usually no need to be as
tight with swap space as with memory, and the performance penalty of
swapping, even with flash drives, is high enough that swap space acts
as an overflow vessel rather than be part of the regularly backing of
the anonymous/shmem working set. It really is NOT obvious that swap
space would need to be adjusted on the fly, and that it's important
that reducing the limit will be reflected in consumption right away.

We shouldn't be adding hundreds of lines of likely terrible heuristics
code* on speculation that somebody MIGHT find this useful in real life.
We should wait until we are presented with a real usecase that applies
to a whole class of users, and then see what the true requirements are.

* If a group has 200M swapped out and the swap limit is reduced by 10M
below the current consumption, which pages would you swap in? There is
no LRU list for swap space.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1293606 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromKamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com>
Date2015-12-17 05:40 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qGyw9-2VN-15@gated-at.bofh.it>
In reply to#1293595
On 2015/12/17 12:32, Johannes Weiner wrote:
> On Thu, Dec 17, 2015 at 11:46:27AM +0900, Kamezawa Hiroyuki wrote:
>> On 2015/12/16 20:09, Johannes Weiner wrote:
>>> On Wed, Dec 16, 2015 at 12:18:30PM +0900, Kamezawa Hiroyuki wrote:
>>>>   - swap-full notification via vmpressure or something mechanism.
>>>
>>> Why?
>>>
>>
>> I think it's a sign of unhealthy condition, starting file cache drop rate to rise.
>> But I forgot that there are resource threshold notifier already. Does the notifier work
>> for swap.usage ?
>
> That will be reflected in vmpressure or other distress mechanisms. I'm
> not convinced "ran out of swap space" needs special casing in any way.
>
Most users checks swap space shortage as "system alarm" in enterprise systems.
At least, our customers checks swap-full.

>>>>   - force swap-in at reducing swap.limit
>>>
>>> Why?
>>>
>> If full, swap.limit cannot be reduced even if there are available memory in a cgroup.
>> Another cgroup cannot make use of the swap resource while it's occupied by other cgroup.
>> The job scheduler should have a chance to fix the situation.
>
> I don't see why swap space allowance would need to be as dynamically
> adjustable as the memory allowance. There is usually no need to be as
> tight with swap space as with memory, and the performance penalty of
> swapping, even with flash drives, is high enough that swap space acts
> as an overflow vessel rather than be part of the regularly backing of
> the anonymous/shmem working set. It really is NOT obvious that swap
> space would need to be adjusted on the fly, and that it's important
> that reducing the limit will be reflected in consumption right away.
>

With my OS support experience, some customers consider swap-space as a resource.


> We shouldn't be adding hundreds of lines of likely terrible heuristics
> code* on speculation that somebody MIGHT find this useful in real life.
> We should wait until we are presented with a real usecase that applies
> to a whole class of users, and then see what the true requirements are.
>
ok, we should wait.  I'm just guessing (japanese) HPC people will want the
feature for their job control. I hear many programs relies on swap.

> * If a group has 200M swapped out and the swap limit is reduced by 10M
> below the current consumption, which pages would you swap in? There is
> no LRU list for swap space.
>
If a rotation can happen when a swap-in-by-real-pagefault, random swap-in
at reducing swap.limit will work enough.

Thanks,
-Kame



--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1292383 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromMichal Hocko <mhocko@kernel.org>
Date2015-12-15 18:30 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qG1Ae-7eq-17@gated-at.bofh.it>
In reply to#1291511
On Mon 14-12-15 22:42:58, Vladimir Davydov wrote:
> On Mon, Dec 14, 2015 at 04:30:37PM +0100, Michal Hocko wrote:
> > On Thu 10-12-15 14:39:14, Vladimir Davydov wrote:
> > > In the legacy hierarchy we charge memsw, which is dubious, because:
> > > 
> > >  - memsw.limit must be >= memory.limit, so it is impossible to limit
> > >    swap usage less than memory usage. Taking into account the fact that
> > >    the primary limiting mechanism in the unified hierarchy is
> > >    memory.high while memory.limit is either left unset or set to a very
> > >    large value, moving memsw.limit knob to the unified hierarchy would
> > >    effectively make it impossible to limit swap usage according to the
> > >    user preference.
> > > 
> > >  - memsw.usage != memory.usage + swap.usage, because a page occupying
> > >    both swap entry and a swap cache page is charged only once to memsw
> > >    counter. As a result, it is possible to effectively eat up to
> > >    memory.limit of memory pages *and* memsw.limit of swap entries, which
> > >    looks unexpected.
> > > 
> > > That said, we should provide a different swap limiting mechanism for
> > > cgroup2.
> > > This patch adds mem_cgroup->swap counter, which charges the actual
> > > number of swap entries used by a cgroup. It is only charged in the
> > > unified hierarchy, while the legacy hierarchy memsw logic is left
> > > intact.
> > 
> > I agree that the previous semantic was awkward. The problem I can see
> > with this approach is that once the swap limit is reached the anon
> > memory pressure might spill over to other and unrelated memcgs during
> > the global memory pressure. I guess this is what Kame referred to as
> > anon would become mlocked basically. This would be even more of an issue
> > with resource delegation to sub-hierarchies because nobody will prevent
> > setting the swap amount to a small value and use that as an anon memory
> > protection.
> 
> AFAICS such anon memory protection has a side-effect: real-life
> workloads need page cache to run smoothly (at least for mapping
> executables). Disabling swapping would switch pressure to page caches,
> resulting in performance degradation. So, I don't think per memcg swap
> limit can be abused to boost your workload on an overcommitted system.

Well, you can trash on the page cache which could slow down the workload
but the executable pages get an additional protection so this might be
not sufficient and still trigger a massive disruption on the global level.

> If you mean malicious users, well, they already have plenty ways to eat
> all available memory up to the hard limit by creating unreclaimable
> kernel objects.
> 
> Anyway, if you don't trust a container you'd better set the hard memory
> limit so that it can't hurt others no matter what it runs and how it
> tweaks its sub-tree knobs.

I completely agree that malicious/untrusted users absolutely have to
be capped by the hard limit. Then the separate swap limit would work
for sure. But I am less convinced about usefulness of the rigid (to
the global memory pressure) swap limit without the hard limit. All the
memory that could have been swapped out will make a memory pressure to
the rest of the system without being punished for it too much. Memcg
is allowed to grow over the high limit (in the current implementation)
without any way to shrink back in other words.

My understanding was that the primary use case for the swap limit is to
handle potential (not only malicious but also unexpectedly misbehaving
application) anon memory consumption runaways more gracefully without
the massive disruption on the global level. I simply didn't see swap
space partitioning as important enough because an alternative to swap
usage is to consume primary memory which is a more precious resource
IMO. Swap storage is really cheap and runtime expandable resource which
is not the case for the primary memory in general. Maybe there are other
use cases I am not aware of, though. Do you want to guarantee the swap
availability?

Just to make it clear. I am not against the new way of the swap
accounting. It is much more clear then the previous one. I am just
worried it allows for an easy misconfiguration and we do not have any
measures to help the global system healthiness. I am OK with the patch
if we document the risk for now. I still think we will end up doing some
heuristic to throttle for a large unreclaimable high limit excess in the
future but I agree this shouldn't be the prerequisite.

> ...
> > My question now is. Is the knob usable/useful even without additional
> > heuristics? Do we want to protect swap space so rigidly that a swap
> > limited memcg can cause bigger problems than without the swap limit
> > globally?
> 
> Hmm, I don't see why problems might get bigger with per memcg swap limit
> than w/o it.

because the reclaim fairness between different memcg hierarchies will be
severely affected.

> W/o swap limit, a memcg can eat all swap space on the host
> and disable swapping for everyone, not just for itself alone.

this is true of course but this is not very much different from pushing
everybody else to the swap while eating the unreclaimable anonymous
memory and eventually hit the OOM killer.

-- 
Michal Hocko
SUSE Labs
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1292519 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromJohannes Weiner <hannes@cmpxchg.org>
Date2015-12-15 21:30 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qG4op-Bj-5@gated-at.bofh.it>
In reply to#1292383
On Tue, Dec 15, 2015 at 06:21:28PM +0100, Michal Hocko wrote:
> > AFAICS such anon memory protection has a side-effect: real-life
> > workloads need page cache to run smoothly (at least for mapping
> > executables). Disabling swapping would switch pressure to page caches,
> > resulting in performance degradation. So, I don't think per memcg swap
> > limit can be abused to boost your workload on an overcommitted system.
> 
> Well, you can trash on the page cache which could slow down the workload
> but the executable pages get an additional protection so this might be
> not sufficient and still trigger a massive disruption on the global level.

No, this is a real consequence. If you fill your available memory with
mostly unreclaimable memory and your executables start thrashing you
might not make forward progress for hours. We don't have a swap token
for page cache.

> Just to make it clear. I am not against the new way of the swap
> accounting. It is much more clear then the previous one. I am just
> worried it allows for an easy misconfiguration and we do not have any
> measures to help the global system healthiness. I am OK with the patch
> if we document the risk for now. I still think we will end up doing some
> heuristic to throttle for a large unreclaimable high limit excess in the
> future but I agree this shouldn't be the prerequisite.

It's unclear to me how the previous memory+swap counters did anything
tangible for global system health with malicious/buggy workloads. If
anything, the previous model seems to encourage blatant overcommit of
workloads on the flawed assumption that global pressure could always
claw back memory, including anonymous pages of untrusted workloads,
which does not actually work in practice. So I'm not sure what new
risk you are referring to here.

As far as the high limit goes, its job is to contain cache growth and
throttle applications during somewhat higher-than-expected consumption
peaks; not to contain "large unreclaimable high limit excess" from
buggy or malicious applications, that's what the hard limit is for.

All in all, it seems to me we should leave this discussion to actual
problems arising in the real world. There is a lot of unfocussed
speculation in this thread about things that might go wrong, without
much thought put into whether these scenarios are even meaningful or
real or whether they are new problems that come with the swap limit.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1292748 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromKamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com>
Date2015-12-16 05:00 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qGbpV-4Ya-17@gated-at.bofh.it>
In reply to#1292383
On 2015/12/16 2:21, Michal Hocko wrote:
  
> I completely agree that malicious/untrusted users absolutely have to
> be capped by the hard limit. Then the separate swap limit would work
> for sure. But I am less convinced about usefulness of the rigid (to
> the global memory pressure) swap limit without the hard limit. All the
> memory that could have been swapped out will make a memory pressure to
> the rest of the system without being punished for it too much. Memcg
> is allowed to grow over the high limit (in the current implementation)
> without any way to shrink back in other words.
>
> My understanding was that the primary use case for the swap limit is to
> handle potential (not only malicious but also unexpectedly misbehaving
> application) anon memory consumption runaways more gracefully without
> the massive disruption on the global level. I simply didn't see swap
> space partitioning as important enough because an alternative to swap
> usage is to consume primary memory which is a more precious resource
> IMO. Swap storage is really cheap and runtime expandable resource which
> is not the case for the primary memory in general. Maybe there are other
> use cases I am not aware of, though. Do you want to guarantee the swap
> availability?
>

At the first implementation, NEC guy explained their use case in HPC area.
At that time, there was no swap support.

Considering 2 workloads partitioned into group A, B. total swap was 100GB.
   A: memory.limit = 40G
   B: memory.limit = 40G

Job scheduler runs applications in A and B in turn. Apps in A stops while Apps in B running.

If App-A requires 120GB of anonymous memory, it uses 80GB of swap. So, App-B can use only
20GB of swap. This can cause trouble if App-B needs 100GB of anonymous memory.
They need some knob to control amount of swap per cgroup.

The point is, at least for their customer, the swap is "resource", which should be under
control. With their use case, memory usage and swap usage has the same meaning. So,
mem+swap limit doesn't cause trouble.

Thanks,
-Kame




--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1291811 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromKamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com>
Date2015-12-15 04:30 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qFOtj-6Va-1@gated-at.bofh.it>
In reply to#1291285
On 2015/12/15 0:30, Michal Hocko wrote:
> On Thu 10-12-15 14:39:14, Vladimir Davydov wrote:
>> In the legacy hierarchy we charge memsw, which is dubious, because:
>>
>>   - memsw.limit must be >= memory.limit, so it is impossible to limit
>>     swap usage less than memory usage. Taking into account the fact that
>>     the primary limiting mechanism in the unified hierarchy is
>>     memory.high while memory.limit is either left unset or set to a very
>>     large value, moving memsw.limit knob to the unified hierarchy would
>>     effectively make it impossible to limit swap usage according to the
>>     user preference.
>>
>>   - memsw.usage != memory.usage + swap.usage, because a page occupying
>>     both swap entry and a swap cache page is charged only once to memsw
>>     counter. As a result, it is possible to effectively eat up to
>>     memory.limit of memory pages *and* memsw.limit of swap entries, which
>>     looks unexpected.
>>
>> That said, we should provide a different swap limiting mechanism for
>> cgroup2.
>> This patch adds mem_cgroup->swap counter, which charges the actual
>> number of swap entries used by a cgroup. It is only charged in the
>> unified hierarchy, while the legacy hierarchy memsw logic is left
>> intact.
>
> I agree that the previous semantic was awkward. The problem I can see
> with this approach is that once the swap limit is reached the anon
> memory pressure might spill over to other and unrelated memcgs during
> the global memory pressure. I guess this is what Kame referred to as
> anon would become mlocked basically. This would be even more of an issue
> with resource delegation to sub-hierarchies because nobody will prevent
> setting the swap amount to a small value and use that as an anon memory
> protection.
>
> I guess this was the reason why this approach hasn't been chosen before

Yes. At that age, "never break global VM" was the policy. And "mlock" can be
used for attacking system.

> but I think we can come up with a way to stop the run away consumption
> even when the swap is accounted separately. All of them are quite nasty
> but let me try.
>
> We could allow charges to fail even for the high limit if the excess is
> way above the amount of reclaimable memory in the given memcg/hierarchy.
> A runaway load would be stopped before it can cause a considerable
> damage outside of its hierarchy this way even when the swap limit
> is configured small.
> Now that goes against the high limit semantic which should only throttle
> the consumer and shouldn't cause any functional failures but maybe this
> is acceptable for the overall system stability. An alternative would
> be to throttle in the high limit reclaim context proportionally to
> the excess. This is normally done by the reclaim itself but with no
> reclaimable memory this wouldn't work that way.
>
This seems hard to use for users who want to control resource precisely
even if stability is good.

> Another option would be to ignore the swap limit during the global
> reclaim. This wouldn't stop the runaway loads but they would at least
> see their fair share of the reclaim. The swap excess could be then used
> as a "handicap" for a more aggressive throttling during high limit reclaim
> or to trigger hard limit sooner.
>
This seems to work. But users need to understand swap-limit can be exceeded.

> Or we could teach the global OOM killer to select abusive anon memory
> users with restricted swap. That would require to iterate through all
> memcgs and checks whether their anon consumption is in a large excess to
> their swap limit and fallback to the memcg OOM victim selection if that
> is the case. This adds more complexity to the OOM killer path so I am
> not sure this is generally acceptable, though.
>

I think this is not acceptable.

> My question now is. Is the knob usable/useful even without additional
> heuristics? Do we want to protect swap space so rigidly that a swap
> limited memcg can cause bigger problems than without the swap limit
> globally?
>

swap requires some limit. If not, an application can eat up all swap
and it will not be never freed until the application access it or
swapoff runs.

Thanks,
-Kame

>> The swap usage can be monitored using new memory.swap.current file and
>> limited using memory.swap.max.
>>
>> Signed-off-by: Vladimir Davydov <vdavydov@virtuozzo.com>
>> ---
>>   include/linux/memcontrol.h |   1 +
>>   include/linux/swap.h       |   5 ++
>>   mm/memcontrol.c            | 123 +++++++++++++++++++++++++++++++++++++++++----
>>   mm/shmem.c                 |   4 ++
>>   mm/swap_state.c            |   5 ++
>>   5 files changed, 129 insertions(+), 9 deletions(-)
>
> [...]
>


--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1291933 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromVladimir Davydov <vdavydov@virtuozzo.com>
Date2015-12-15 09:40 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qFTjk-1xA-7@gated-at.bofh.it>
In reply to#1291811
On Tue, Dec 15, 2015 at 12:12:40PM +0900, Kamezawa Hiroyuki wrote:
> On 2015/12/15 0:30, Michal Hocko wrote:
> >On Thu 10-12-15 14:39:14, Vladimir Davydov wrote:
> >>In the legacy hierarchy we charge memsw, which is dubious, because:
> >>
> >>  - memsw.limit must be >= memory.limit, so it is impossible to limit
> >>    swap usage less than memory usage. Taking into account the fact that
> >>    the primary limiting mechanism in the unified hierarchy is
> >>    memory.high while memory.limit is either left unset or set to a very
> >>    large value, moving memsw.limit knob to the unified hierarchy would
> >>    effectively make it impossible to limit swap usage according to the
> >>    user preference.
> >>
> >>  - memsw.usage != memory.usage + swap.usage, because a page occupying
> >>    both swap entry and a swap cache page is charged only once to memsw
> >>    counter. As a result, it is possible to effectively eat up to
> >>    memory.limit of memory pages *and* memsw.limit of swap entries, which
> >>    looks unexpected.
> >>
> >>That said, we should provide a different swap limiting mechanism for
> >>cgroup2.
> >>This patch adds mem_cgroup->swap counter, which charges the actual
> >>number of swap entries used by a cgroup. It is only charged in the
> >>unified hierarchy, while the legacy hierarchy memsw logic is left
> >>intact.
> >
> >I agree that the previous semantic was awkward. The problem I can see
> >with this approach is that once the swap limit is reached the anon
> >memory pressure might spill over to other and unrelated memcgs during
> >the global memory pressure. I guess this is what Kame referred to as
> >anon would become mlocked basically. This would be even more of an issue
> >with resource delegation to sub-hierarchies because nobody will prevent
> >setting the swap amount to a small value and use that as an anon memory
> >protection.
> >
> >I guess this was the reason why this approach hasn't been chosen before
> 
> Yes. At that age, "never break global VM" was the policy. And "mlock" can be
> used for attacking system.

If we are talking about "attacking system" from inside a container,
there are much easier and disruptive ways, e.g. running a fork-bomb or
creating pipes - such memory can't be reclaimed and global OOM killer
won't help.

Thanks,
Vladimir
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1291976 — Re: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2

FromKamezawa Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com>
Date2015-12-15 10:40 +0100
SubjectRe: [PATCH 1/7] mm: memcontrol: charge swap to cgroup2
Message-ID<qFUfo-2at-7@gated-at.bofh.it>
In reply to#1291933
On 2015/12/15 17:30, Vladimir Davydov wrote:
> On Tue, Dec 15, 2015 at 12:12:40PM +0900, Kamezawa Hiroyuki wrote:
>> On 2015/12/15 0:30, Michal Hocko wrote:
>>> On Thu 10-12-15 14:39:14, Vladimir Davydov wrote:
>>>> In the legacy hierarchy we charge memsw, which is dubious, because:
>>>>
>>>>   - memsw.limit must be >= memory.limit, so it is impossible to limit
>>>>     swap usage less than memory usage. Taking into account the fact that
>>>>     the primary limiting mechanism in the unified hierarchy is
>>>>     memory.high while memory.limit is either left unset or set to a very
>>>>     large value, moving memsw.limit knob to the unified hierarchy would
>>>>     effectively make it impossible to limit swap usage according to the
>>>>     user preference.
>>>>
>>>>   - memsw.usage != memory.usage + swap.usage, because a page occupying
>>>>     both swap entry and a swap cache page is charged only once to memsw
>>>>     counter. As a result, it is possible to effectively eat up to
>>>>     memory.limit of memory pages *and* memsw.limit of swap entries, which
>>>>     looks unexpected.
>>>>
>>>> That said, we should provide a different swap limiting mechanism for
>>>> cgroup2.
>>>> This patch adds mem_cgroup->swap counter, which charges the actual
>>>> number of swap entries used by a cgroup. It is only charged in the
>>>> unified hierarchy, while the legacy hierarchy memsw logic is left
>>>> intact.
>>>
>>> I agree that the previous semantic was awkward. The problem I can see
>>> with this approach is that once the swap limit is reached the anon
>>> memory pressure might spill over to other and unrelated memcgs during
>>> the global memory pressure. I guess this is what Kame referred to as
>>> anon would become mlocked basically. This would be even more of an issue
>>> with resource delegation to sub-hierarchies because nobody will prevent
>>> setting the swap amount to a small value and use that as an anon memory
>>> protection.
>>>
>>> I guess this was the reason why this approach hasn't been chosen before
>>
>> Yes. At that age, "never break global VM" was the policy. And "mlock" can be
>> used for attacking system.
>
> If we are talking about "attacking system" from inside a container,
> there are much easier and disruptive ways, e.g. running a fork-bomb or
> creating pipes - such memory can't be reclaimed and global OOM killer
> won't help.

You're right. We just wanted to avoid affecting global memory reclaim by
each cgroup settings.

Thanks,
-Kame



--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1288466 — [PATCH 7/7] Documentation: cgroup: add memory.swap.{current,max} description

FromVladimir Davydov <vdavydov@virtuozzo.com>
Date2015-12-10 12:50 +0100
Subject[PATCH 7/7] Documentation: cgroup: add memory.swap.{current,max} description
Message-ID<qE7Tr-5EP-1@gated-at.bofh.it>
In reply to#1288457
Signed-off-by: Vladimir Davydov <vdavydov@virtuozzo.com>
---
 Documentation/cgroup.txt | 16 ++++++++++++++++
 1 file changed, 16 insertions(+)

diff --git a/Documentation/cgroup.txt b/Documentation/cgroup.txt
index 31d1f7bf12a1..21c6c013c339 100644
--- a/Documentation/cgroup.txt
+++ b/Documentation/cgroup.txt
@@ -819,6 +819,22 @@ PAGE_SIZE multiple when read back.
 		the cgroup.  This may not exactly match the number of
 		processes killed but should generally be close.
 
+  memory.swap.current
+
+	A read-only single value file which exists on non-root
+	cgroups.
+
+	The total amount of swap currently being used by the cgroup
+	and its descendants.
+
+  memory.swap.max
+
+	A read-write single value file which exists on non-root
+	cgroups.  The default is "max".
+
+	Swap usage hard limit.  If a cgroup's swap usage reaches this
+	limit, anonymous meomry of the cgroup will not be swapped out.
+
 
 5-2-2. General Usage
 
-- 
2.1.4

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289865 — Re: [PATCH 7/7] Documentation: cgroup: add memory.swap.{current,max} description

FromJohannes Weiner <hannes@cmpxchg.org>
Date2015-12-11 20:50 +0100
SubjectRe: [PATCH 7/7] Documentation: cgroup: add memory.swap.{current,max} description
Message-ID<qEBRv-mG-33@gated-at.bofh.it>
In reply to#1288466
On Thu, Dec 10, 2015 at 02:39:20PM +0300, Vladimir Davydov wrote:
> Signed-off-by: Vladimir Davydov <vdavydov@virtuozzo.com>

Acked-by: Johannes Weiner <hannes@cmpxchg.org>

Can we include a blurb for R-5-1 of cgroups.txt as well to explain why
cgroup2 has a new swap interface? I had already written something up
in the past, pasted below, feel free to use it if you like. Otherwise,
you had pretty good reasons in your swap controller changelog as well.

---

- The combined memory+swap accounting and limiting is replaced by real
  control over swap space.

  memory.swap.current

       The amount memory of this subtree that has been swapped to
       disk.

  memory.swap.max

       The maximum amount of memory this subtree is allowed to swap
       to disk.

  The main argument for a combined memory+swap facility in the
  original cgroup design was that global or parental pressure would
  always be able to swap all anonymous memory of a child group,
  regardless of the child's own (possibly untrusted) configuration.
  However, untrusted groups can sabotage swapping by other means--such
  as referencing its anonymous memory in a tight loop--and an admin
  can not assume full swappability when overcommitting untrusted jobs.

  For trusted jobs, on the other hand, a combined counter is not an
  intuitive userspace interface, and it flies in the face of the idea
  that cgroup controllers should account and limit specific physical
  resources. Swap space is a resource like all others in the system,
  and that's why unified hierarchy allows distributing it separately.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1290219 — Re: [PATCH 7/7] Documentation: cgroup: add memory.swap.{current,max} description

FromVladimir Davydov <vdavydov@virtuozzo.com>
Date2015-12-12 17:20 +0100
SubjectRe: [PATCH 7/7] Documentation: cgroup: add memory.swap.{current,max} description
Message-ID<qEV3P-4oQ-1@gated-at.bofh.it>
In reply to#1289865
On Fri, Dec 11, 2015 at 02:42:54PM -0500, Johannes Weiner wrote:
> On Thu, Dec 10, 2015 at 02:39:20PM +0300, Vladimir Davydov wrote:
> > Signed-off-by: Vladimir Davydov <vdavydov@virtuozzo.com>
> 
> Acked-by: Johannes Weiner <hannes@cmpxchg.org>
> 
> Can we include a blurb for R-5-1 of cgroups.txt as well to explain why
> cgroup2 has a new swap interface? I had already written something up
> in the past, pasted below, feel free to use it if you like. Otherwise,
> you had pretty good reasons in your swap controller changelog as well.

Will do in v2.

Thanks!
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1288468 — [PATCH 6/7] mm: free swap cache aggressively if memcg swap is full

FromVladimir Davydov <vdavydov@virtuozzo.com>
Date2015-12-10 12:50 +0100
Subject[PATCH 6/7] mm: free swap cache aggressively if memcg swap is full
Message-ID<qE7Tr-5EP-5@gated-at.bofh.it>
In reply to#1288457
Swap cache pages are freed aggressively if swap is nearly full (>50%
currently), because otherwise we are likely to stop scanning anonymous
when we near the swap limit even if there is plenty of freeable swap
cache pages. We should follow the same trend in case of memory cgroup,
which has its own swap limit.

Signed-off-by: Vladimir Davydov <vdavydov@virtuozzo.com>
---
 include/linux/swap.h |  6 ++++++
 mm/memcontrol.c      | 23 +++++++++++++++++++++++
 mm/memory.c          |  3 ++-
 mm/swapfile.c        |  2 +-
 mm/vmscan.c          |  2 +-
 5 files changed, 33 insertions(+), 3 deletions(-)

diff --git a/include/linux/swap.h b/include/linux/swap.h
index e3344d8ca2e9..1d708860be97 100644
--- a/include/linux/swap.h
+++ b/include/linux/swap.h
@@ -552,6 +552,7 @@ extern void mem_cgroup_swapout(struct page *page, swp_entry_t entry);
 extern int mem_cgroup_charge_swap(struct page *page, swp_entry_t entry);
 extern void mem_cgroup_uncharge_swap(swp_entry_t entry);
 extern long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg);
+extern bool mem_cgroup_swap_full(struct page *page);
 #else
 static inline void mem_cgroup_swapout(struct page *page, swp_entry_t entry)
 {
@@ -570,6 +571,11 @@ static inline long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
 {
 	return get_nr_swap_pages();
 }
+
+static inline bool mem_cgroup_swap_full(struct page *page)
+{
+	return vm_swap_full();
+}
 #endif
 
 #endif /* __KERNEL__*/
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 2ee823d62f80..e5bd43340cd8 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -5839,6 +5839,29 @@ long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
 	return nr_swap_pages;
 }
 
+bool mem_cgroup_swap_full(struct page *page)
+{
+	struct mem_cgroup *memcg;
+
+	VM_BUG_ON_PAGE(!PageLocked(page), page);
+
+	if (vm_swap_full())
+		return true;
+	if (!do_swap_account || !PageSwapCache(page))
+		return false;
+
+	memcg = page->mem_cgroup;
+	if (!memcg)
+		return false;
+
+	for (; memcg != root_mem_cgroup; memcg = parent_mem_cgroup(memcg)) {
+		if (page_counter_read(&memcg->swap) * 2 >=
+				READ_ONCE(memcg->swap.limit))
+			return true;
+	}
+	return false;
+}
+
 /* for remember boot option*/
 #ifdef CONFIG_MEMCG_SWAP_ENABLED
 static int really_do_swap_account __initdata = 1;
diff --git a/mm/memory.c b/mm/memory.c
index 3b115dcaa26e..2bd6a78c142b 100644
--- a/mm/memory.c
+++ b/mm/memory.c
@@ -2563,7 +2563,8 @@ int do_swap_page(struct mm_struct *mm, struct vm_area_struct *vma,
 	}
 
 	swap_free(entry);
-	if (vm_swap_full() || (vma->vm_flags & VM_LOCKED) || PageMlocked(page))
+	if (mem_cgroup_swap_full(page) ||
+	    (vma->vm_flags & VM_LOCKED) || PageMlocked(page))
 		try_to_free_swap(page);
 	unlock_page(page);
 	if (page != swapcache) {
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 7073faecb38f..c0aba04f7a59 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -1011,7 +1011,7 @@ int free_swap_and_cache(swp_entry_t entry)
 		 * Also recheck PageSwapCache now page is locked (above).
 		 */
 		if (PageSwapCache(page) && !PageWriteback(page) &&
-				(!page_mapped(page) || vm_swap_full())) {
+		    (!page_mapped(page) || mem_cgroup_swap_full(page))) {
 			delete_from_swap_cache(page);
 			SetPageDirty(page);
 		}
diff --git a/mm/vmscan.c b/mm/vmscan.c
index ab52d865d922..1cd88e9b0383 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -1206,7 +1206,7 @@ cull_mlocked:
 
 activate_locked:
 		/* Not a candidate for swapping, so reclaim swap space. */
-		if (PageSwapCache(page) && vm_swap_full())
+		if (PageSwapCache(page) && mem_cgroup_swap_full(page))
 			try_to_free_swap(page);
 		VM_BUG_ON_PAGE(PageActive(page), page);
 		SetPageActive(page);
-- 
2.1.4

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1289856 — Re: [PATCH 6/7] mm: free swap cache aggressively if memcg swap is full

FromJohannes Weiner <hannes@cmpxchg.org>
Date2015-12-11 20:40 +0100
SubjectRe: [PATCH 6/7] mm: free swap cache aggressively if memcg swap is full
Message-ID<qEBHQ-ir-21@gated-at.bofh.it>
In reply to#1288468
On Thu, Dec 10, 2015 at 02:39:19PM +0300, Vladimir Davydov wrote:
> Swap cache pages are freed aggressively if swap is nearly full (>50%
> currently), because otherwise we are likely to stop scanning anonymous
> when we near the swap limit even if there is plenty of freeable swap
> cache pages. We should follow the same trend in case of memory cgroup,
> which has its own swap limit.
> 
> Signed-off-by: Vladimir Davydov <vdavydov@virtuozzo.com>

Acked-by: Johannes Weiner <hannes@cmpxchg.org>

One note:

> @@ -5839,6 +5839,29 @@ long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
>  	return nr_swap_pages;
>  }
>  
> +bool mem_cgroup_swap_full(struct page *page)
> +{
> +	struct mem_cgroup *memcg;
> +
> +	VM_BUG_ON_PAGE(!PageLocked(page), page);
> +
> +	if (vm_swap_full())
> +		return true;
> +	if (!do_swap_account || !PageSwapCache(page))
> +		return false;

The callers establish PageSwapCache() under the page lock, which makes
sense since they only inquire about the swap state when deciding what
to do with a swapcache page at hand. So this check seems unnecessary.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1290222 — Re: [PATCH 6/7] mm: free swap cache aggressively if memcg swap is full

FromVladimir Davydov <vdavydov@virtuozzo.com>
Date2015-12-12 17:20 +0100
SubjectRe: [PATCH 6/7] mm: free swap cache aggressively if memcg swap is full
Message-ID<qEV3P-4oQ-15@gated-at.bofh.it>
In reply to#1289856
On Fri, Dec 11, 2015 at 02:33:58PM -0500, Johannes Weiner wrote:
> On Thu, Dec 10, 2015 at 02:39:19PM +0300, Vladimir Davydov wrote:
> > Swap cache pages are freed aggressively if swap is nearly full (>50%
> > currently), because otherwise we are likely to stop scanning anonymous
> > when we near the swap limit even if there is plenty of freeable swap
> > cache pages. We should follow the same trend in case of memory cgroup,
> > which has its own swap limit.
> > 
> > Signed-off-by: Vladimir Davydov <vdavydov@virtuozzo.com>
> 
> Acked-by: Johannes Weiner <hannes@cmpxchg.org>
> 
> One note:
> 
> > @@ -5839,6 +5839,29 @@ long mem_cgroup_get_nr_swap_pages(struct mem_cgroup *memcg)
> >  	return nr_swap_pages;
> >  }
> >  
> > +bool mem_cgroup_swap_full(struct page *page)
> > +{
> > +	struct mem_cgroup *memcg;
> > +
> > +	VM_BUG_ON_PAGE(!PageLocked(page), page);
> > +
> > +	if (vm_swap_full())
> > +		return true;
> > +	if (!do_swap_account || !PageSwapCache(page))
> > +		return false;
> 
> The callers establish PageSwapCache() under the page lock, which makes
> sense since they only inquire about the swap state when deciding what
> to do with a swapcache page at hand. So this check seems unnecessary.

Yeah, you're right, we don't need it here. Will remove it in v2.

Besides, I think I should have inserted cgroup_subsys_on_dflt check in
this function so that it wouldn't check memcg->swap limit in case the
legacy hierarchy is used. Will do.

Thanks,
Vladimir
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web