Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1700315 > unrolled thread
| Started by | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| First post | 2017-07-31 20:50 +0200 |
| Last post | 2017-08-01 14:30 +0200 |
| Articles | 6 — 3 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH 3/3] mm/sched: memdelay: memory health interface for systems and workloads Johannes Weiner <hannes@cmpxchg.org> - 2017-07-31 20:50 +0200
Re: [PATCH 3/3] mm/sched: memdelay: memory health interface for systems and workloads Mike Galbraith <efault@gmx.de> - 2017-07-31 22:00 +0200
Re: [PATCH 3/3] mm/sched: memdelay: memory health interface for systems and workloads Johannes Weiner <hannes@cmpxchg.org> - 2017-07-31 22:40 +0200
Re: [PATCH 3/3] mm/sched: memdelay: memory health interface for systems and workloads Mike Galbraith <efault@gmx.de> - 2017-08-01 04:30 +0200
Re: [PATCH 3/3] mm/sched: memdelay: memory health interface for systems and workloads Peter Zijlstra <peterz@infradead.org> - 2017-08-01 10:00 +0200
Re: [PATCH 3/3] mm/sched: memdelay: memory health interface for systems and workloads Johannes Weiner <hannes@cmpxchg.org> - 2017-08-01 14:30 +0200
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2017-07-31 20:50 +0200 |
| Subject | Re: [PATCH 3/3] mm/sched: memdelay: memory health interface for systems and workloads |
| Message-ID | <u9nVo-1jR-19@gated-at.bofh.it> |
On Mon, Jul 31, 2017 at 10:31:11AM +0200, Peter Zijlstra wrote:
> On Sun, Jul 30, 2017 at 11:28:13AM -0400, Johannes Weiner wrote:
> > On Sat, Jul 29, 2017 at 11:10:55AM +0200, Peter Zijlstra wrote:
> > > On Thu, Jul 27, 2017 at 11:30:10AM -0400, Johannes Weiner wrote:
> > > > +static void domain_cpu_update(struct memdelay_domain *md, int cpu,
> > > > + int old, int new)
> > > > +{
> > > > + enum memdelay_domain_state state;
> > > > + struct memdelay_domain_cpu *mdc;
> > > > + unsigned long now, delta;
> > > > + unsigned long flags;
> > > > +
> > > > + mdc = per_cpu_ptr(md->mdcs, cpu);
> > > > + spin_lock_irqsave(&mdc->lock, flags);
> > >
> > > Afaict this is inside scheduler locks, this cannot be a spinlock. Also,
> > > do we really want to add more atomics there?
> >
> > I think we should be able to get away without an additional lock and
> > rely on the rq lock instead. schedule, enqueue, dequeue already hold
> > it, memdelay_enter/leave could be added. I need to think about what to
> > do with try_to_wake_up in order to get the cpu move accounting inside
> > the locked section of ttwu_queue(), but that should be doable too.
>
> So could you start by describing what actual statistics we need? Because
> as is the scheduler already does a gazillion stats and why can't re
> repurpose some of those?
If that's possible, that would be great of course.
We want to be able to tell how many tasks in a domain (the system or a
memory cgroup) are inside a memdelay section as opposed to how many
are in a "productive" state such as runnable or iowait. Then derive
from that whether the domain as a whole is unproductive (all non-idle
tasks memdelayed), or partially unproductive (some delayed, but CPUs
are productive or there are iowait tasks). Then derive the percentages
of walltime the domain spends partially or fully unproductive.
For that we need per-domain counters for
1) nr of tasks in memdelay sections
2) nr of iowait or runnable/queued tasks that are NOT inside
memdelay sections
The memdelay and runnable counts need to be per-cpu as well. (The idea
is this: if you have one CPU and some tasks are delayed while others
are runnable, you're 100% partially productive, as the CPU is fully
used. But if you have two CPUs, and the tasks on one CPU are all
runnable while the tasks on the others are all delayed, the domain is
50% of the time fully unproductive (and not 100% partially productive)
as half the available CPU time is being squandered by delays).
On the system-level, we already count runnable/queued per cpu through
rq->nr_running.
However, we need to distinguish between productive runnables and tasks
that are in runnable while in a memdelay section (doing reclaim). The
current counters don't do that.
Lastly, and somewhat obscurely, the presence of runnable tasks means
that usually the domain is at least partially productive. But if the
CPU is used by a task in a memdelay section (direct reclaim), the
domain is fully unproductive (unless there are iowait tasks in the
domain, since they make "progress" without CPU). So we need to track
task_current() && task_memdelayed() per-domain per-cpu as well.
Now, thinking only about the system-level, we could split
rq->nr_running into a sets of delayed and non-delayed counters
(present them as sum in all current read sides).
Adding an rq counter for tasks inside memdelay sections should be
straight-forward as well (except for maybe the migration cost of that
state between CPUs in ttwu that Mike pointed out).
That leaves the question of how to track these numbers per cgroup at
an acceptable cost. The idea for a tree of cgroups is that walltime
impact of delays at each level is reported for all tasks at or below
that level. E.g. a leave group aggregates the state of its own tasks,
the root/system aggregates the state of all tasks in the system; hence
the propagation of the task state counters up the hierarchy.
[toc] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-07-31 22:00 +0200 |
| Message-ID | <u9p17-1XO-5@gated-at.bofh.it> |
| In reply to | #1700315 |
On Mon, 2017-07-31 at 14:41 -0400, Johannes Weiner wrote: > > Adding an rq counter for tasks inside memdelay sections should be > straight-forward as well (except for maybe the migration cost of that > state between CPUs in ttwu that Mike pointed out). What I pointed out should be easily eliminated (zero use case). > That leaves the question of how to track these numbers per cgroup at > an acceptable cost. The idea for a tree of cgroups is that walltime > impact of delays at each level is reported for all tasks at or below > that level. E.g. a leave group aggregates the state of its own tasks, > the root/system aggregates the state of all tasks in the system; hence > the propagation of the task state counters up the hierarchy. The crux of the biscuit is where exactly the investment return lies. Gathering of these numbers ain't gonna be free, no matter how hard you try, and you're plugging into paths where every cycle added is made of userspace hide. -Mike
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2017-07-31 22:40 +0200 |
| Message-ID | <u9pDP-2pW-3@gated-at.bofh.it> |
| In reply to | #1700358 |
On Mon, Jul 31, 2017 at 09:49:39PM +0200, Mike Galbraith wrote: > On Mon, 2017-07-31 at 14:41 -0400, Johannes Weiner wrote: > > > > Adding an rq counter for tasks inside memdelay sections should be > > straight-forward as well (except for maybe the migration cost of that > > state between CPUs in ttwu that Mike pointed out). > > What I pointed out should be easily eliminated (zero use case). How so? > > That leaves the question of how to track these numbers per cgroup at > > an acceptable cost. The idea for a tree of cgroups is that walltime > > impact of delays at each level is reported for all tasks at or below > > that level. E.g. a leave group aggregates the state of its own tasks, > > the root/system aggregates the state of all tasks in the system; hence > > the propagation of the task state counters up the hierarchy. > > The crux of the biscuit is where exactly the investment return lies. > Gathering of these numbers ain't gonna be free, no matter how hard you > try, and you're plugging into paths where every cycle added is made of > userspace hide. Right. But how to implement it sanely and optimize for cycles, and whether we want to default-enable this interface are two separate conversations. It makes sense to me to first make the implementation as lightweight on cycles and maintainability as possible, and then worry about the cost / benefit defaults of the shipped Linux kernel afterwards. That goes for the purely informative userspace interface, anyway. The easily-provoked thrashing livelock I have described in the email to Andrew is a different matter. If the OOM killer requires hooking up to this metric to fix it, it won't be optional. But the OOM code isn't part of this series yet, so again a conversation best had later, IMO. PS: I'm stealing the "made of userspace hide" thing.
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <efault@gmx.de> |
|---|---|
| Date | 2017-08-01 04:30 +0200 |
| Message-ID | <u9v6x-5RK-1@gated-at.bofh.it> |
| In reply to | #1700379 |
On Mon, 2017-07-31 at 16:38 -0400, Johannes Weiner wrote: > On Mon, Jul 31, 2017 at 09:49:39PM +0200, Mike Galbraith wrote: > > On Mon, 2017-07-31 at 14:41 -0400, Johannes Weiner wrote: > > > > > > Adding an rq counter for tasks inside memdelay sections should be > > > straight-forward as well (except for maybe the migration cost of that > > > state between CPUs in ttwu that Mike pointed out). > > > > What I pointed out should be easily eliminated (zero use case). > > How so? I was thinking along the lines of schedstat_enabled(). > > > That leaves the question of how to track these numbers per cgroup at > > > an acceptable cost. The idea for a tree of cgroups is that walltime > > > impact of delays at each level is reported for all tasks at or below > > > that level. E.g. a leave group aggregates the state of its own tasks, > > > the root/system aggregates the state of all tasks in the system; hence > > > the propagation of the task state counters up the hierarchy. > > > > The crux of the biscuit is where exactly the investment return lies. > > Gathering of these numbers ain't gonna be free, no matter how hard you > > try, and you're plugging into paths where every cycle added is made of > > userspace hide. > > Right. But how to implement it sanely and optimize for cycles, and > whether we want to default-enable this interface are two separate > conversations. > > It makes sense to me to first make the implementation as lightweight > on cycles and maintainability as possible, and then worry about the > cost / benefit defaults of the shipped Linux kernel afterwards. > > That goes for the purely informative userspace interface, anyway. The > easily-provoked thrashing livelock I have described in the email to > Andrew is a different matter. If the OOM killer requires hooking up to > this metric to fix it, it won't be optional. But the OOM code isn't > part of this series yet, so again a conversation best had later, IMO. If that "the many must pay a toll to save the few" conversation ever happens, just recall me registering my boo/hiss in advance. I don't have to feel guilty about not liking the idea of making donations to feed the poor starving proggies ;-) -Mike
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-08-01 10:00 +0200 |
| Message-ID | <u9AfU-pF-3@gated-at.bofh.it> |
| In reply to | #1700315 |
On Mon, Jul 31, 2017 at 02:41:42PM -0400, Johannes Weiner wrote: > On Mon, Jul 31, 2017 at 10:31:11AM +0200, Peter Zijlstra wrote: > > So could you start by describing what actual statistics we need? Because > > as is the scheduler already does a gazillion stats and why can't re > > repurpose some of those? > > If that's possible, that would be great of course. > > We want to be able to tell how many tasks in a domain (the system or a > memory cgroup) are inside a memdelay section as opposed to how many And you haven't even defined wth a memdelay section is yet.. > are in a "productive" state such as runnable or iowait. Then derive > from that whether the domain as a whole is unproductive (all non-idle > tasks memdelayed), or partially unproductive (some delayed, but CPUs > are productive or there are iowait tasks). Then derive the percentages > of walltime the domain spends partially or fully unproductive. > > For that we need per-domain counters for > > 1) nr of tasks in memdelay sections > 2) nr of iowait or runnable/queued tasks that are NOT inside > memdelay sections And I still have no clue..
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2017-08-01 14:30 +0200 |
| Message-ID | <u9Etc-3XT-17@gated-at.bofh.it> |
| In reply to | #1700710 |
On Tue, Aug 01, 2017 at 09:57:28AM +0200, Peter Zijlstra wrote: > On Mon, Jul 31, 2017 at 02:41:42PM -0400, Johannes Weiner wrote: > > On Mon, Jul 31, 2017 at 10:31:11AM +0200, Peter Zijlstra wrote: > > > > So could you start by describing what actual statistics we need? Because > > > as is the scheduler already does a gazillion stats and why can't re > > > repurpose some of those? > > > > If that's possible, that would be great of course. > > > > We want to be able to tell how many tasks in a domain (the system or a > > memory cgroup) are inside a memdelay section as opposed to how many > > And you haven't even defined wth a memdelay section is yet.. It's what a task is in after it calls memdelay_enter() and before it calls memdelay_leave(). Tasks mark themselves to be in a memory section when they know to perform work that is necessary due to a lack of memory, such as waiting for a refault or a direct reclaim invocation. From the patch: +/** + * memdelay_enter - mark the beginning of a memory delay section + * @flags: flags to handle nested memdelay sections + * + * Marks the calling task as being delayed due to a lack of memory, + * such as waiting for a workingset refault or performing reclaim. + */ +/** + * memdelay_leave - mark the end of a memory delay section + * @flags: flags to handle nested memdelay sections + * + * Marks the calling task as no longer delayed due to memory. + */ where a reclaim callsite looks like this (decluttered): memdelay_enter() nr_reclaimed = do_try_to_free_pages() memdelay_leave() That's what defines the "unproductive due to lack of memory" state of a task. Time spent in that state weighed against time spent while the task is productive - runnable or in iowait while not in a memdelay section - gives the memory health of the task. And the system and cgroup states/health can be derived from task states as described: > > are in a "productive" state such as runnable or iowait. Then derive > > from that whether the domain as a whole is unproductive (all non-idle > > tasks memdelayed), or partially unproductive (some delayed, but CPUs > > are productive or there are iowait tasks). Then derive the percentages > > of walltime the domain spends partially or fully unproductive. > > > > For that we need per-domain counters for > > > > 1) nr of tasks in memdelay sections > > 2) nr of iowait or runnable/queued tasks that are NOT inside > > memdelay sections
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web