Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1239987 > unrolled thread

CFS scheduler unfairly prefers pinned tasks

Started bypaul.szabo@sydney.edu.au
First post2015-10-06 00:00 +0200
Last post2015-10-10 10:00 +0200
Articles 20 on this page of 46 — 5 participants

Back to article view | Back to linux.kernel


Contents

  CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-06 00:00 +0200
    Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-06 04:50 +0200
      Re: CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-06 12:10 +0200
        Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-06 14:20 +0200
          Re: CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-06 22:50 +0200
            Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-07 03:30 +0200
      Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-08 10:30 +0200
        Re: CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-08 13:00 +0200
          Re: CFS scheduler unfairly prefers pinned tasks Peter Zijlstra <peterz@infradead.org> - 2015-10-08 13:30 +0200
            [patch] sched: disable task group re-weighting on the desktop Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-10 15:30 +0200
              Re: [patch] sched: disable task group re-weighting on the desktop Peter Zijlstra <peterz@infradead.org> - 2015-10-10 19:10 +0200
                Re: [patch] sched: disable task group re-weighting on the desktop Peter Zijlstra <peterz@infradead.org> - 2015-10-10 19:20 +0200
                Re: [patch] sched: disable task group re-weighting on the desktop Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-11 04:30 +0200
                  4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-11 19:50 +0200
                    Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-12 09:30 +0200
                      Re: 4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-12 09:50 +0200
                        Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-12 10:10 +0200
                          Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-12 10:50 +0200
                            Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-12 11:20 +0200
                              Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-12 12:10 +0200
                                Re: 4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-12 12:30 +0200
                                  Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-13 05:50 +0200
                                    Re: 4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-13 06:10 +0200
                                      Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-13 06:40 +0200
                                    Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-13 10:10 +0200
                                      Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-13 10:20 +0200
                                        Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-13 10:30 +0200
                                      Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-13 10:30 +0200
                                Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-12 13:50 +0200
                                  Re: 4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-13 04:30 +0200
                                  Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-13 05:30 +0200
                                    Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-13 10:10 +0200
                          Re: 4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-12 10:50 +0200
              Re: [patch] sched: disable task group re-weighting on the desktop paul.szabo@sydney.edu.au - 2015-10-10 22:20 +0200
                Re: [patch] sched: disable task group re-weighting on the desktop Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-11 04:40 +0200
                  Re: [patch] sched: disable task group re-weighting on the desktop paul.szabo@sydney.edu.au - 2015-10-11 11:30 +0200
                    Re: [patch] sched: disable task group re-weighting on the desktop Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-11 15:00 +0200
              Re: [patch] sched: disable task group re-weighting on the desktop paul.szabo@sydney.edu.au - 2015-10-11 21:50 +0200
                Re: [patch] sched: disable task group re-weighting on the desktop Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-12 04:00 +0200
          Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-08 16:30 +0200
            Re: CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-09 00:00 +0200
              Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-09 04:00 +0200
              Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-09 04:50 +0200
                Re: CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-11 11:50 +0200
        Re: CFS scheduler unfairly prefers pinned tasks Wanpeng Li <wanpeng.li@hotmail.com> - 2015-10-10 06:10 +0200
          Re: CFS scheduler unfairly prefers pinned tasks Wanpeng Li <wanpeng.li@hotmail.com> - 2015-10-10 10:00 +0200

Page 1 of 3  [1] 2 3  Next page →


#1239987 — CFS scheduler unfairly prefers pinned tasks

Frompaul.szabo@sydney.edu.au
Date2015-10-06 00:00 +0200
SubjectCFS scheduler unfairly prefers pinned tasks
Message-ID<qglXB-37e-27@gated-at.bofh.it>
The Linux CFS scheduler prefers pinned tasks and unfairly
gives more CPU time to tasks that have set CPU affinity.
This effect is observed with or without CGROUP controls.

To demonstrate: on an otherwise idle machine, as some user
run several processes pinned to each CPU, one for each CPU
(as many as CPUs present in the system) e.g. for a quad-core
non-HyperThreaded machine:

  taskset -c 0 perl -e 'while(1){1}' &
  taskset -c 1 perl -e 'while(1){1}' &
  taskset -c 2 perl -e 'while(1){1}' &
  taskset -c 3 perl -e 'while(1){1}' &

and (as that same or some other user) run some without
pinning:

  perl -e 'while(1){1}' &
  perl -e 'while(1){1}' &

and use e.g.   top   to observe that the pinned processes get
more CPU time than "fair".

Fairness is obtained when either:
 - there are as many un-pinned processes as CPUs; or
 - with CGROUP controls and the two kinds of processes run by
   different users, when there is just one un-pinned process; or
 - if the pinning is turned off for these processes (or they
   are started without).

Any insight is welcome!

---

I would appreciate replies direct to me as I am not subscribed to the
linux-kernel mailing list (but will try to watch the archives).

This bug is also reported to Debian, please see
  http://bugs.debian.org/800945

I use Debian with the 3.16 kernel, have not yet tried 4.* kernels.


Thanks, Paul

Paul Szabo   psz@maths.usyd.edu.au   http://www.maths.usyd.edu.au/u/psz/
School of Mathematics and Statistics   University of Sydney    Australia
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1240125

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2015-10-06 04:50 +0200
Message-ID<qgqud-1kk-1@gated-at.bofh.it>
In reply to#1239987
On Tue, 2015-10-06 at 08:48 +1100, paul.szabo@sydney.edu.au wrote:
> The Linux CFS scheduler prefers pinned tasks and unfairly
> gives more CPU time to tasks that have set CPU affinity.
> This effect is observed with or without CGROUP controls.
> 
> To demonstrate: on an otherwise idle machine, as some user
> run several processes pinned to each CPU, one for each CPU
> (as many as CPUs present in the system) e.g. for a quad-core
> non-HyperThreaded machine:
> 
>   taskset -c 0 perl -e 'while(1){1}' &
>   taskset -c 1 perl -e 'while(1){1}' &
>   taskset -c 2 perl -e 'while(1){1}' &
>   taskset -c 3 perl -e 'while(1){1}' &
> 
> and (as that same or some other user) run some without
> pinning:
> 
>   perl -e 'while(1){1}' &
>   perl -e 'while(1){1}' &
> 
> and use e.g.   top   to observe that the pinned processes get
> more CPU time than "fair".
> 
> Fairness is obtained when either:
>  - there are as many un-pinned processes as CPUs; or
>  - with CGROUP controls and the two kinds of processes run by
>    different users, when there is just one un-pinned process; or
>  - if the pinning is turned off for these processes (or they
>    are started without).
> 
> Any insight is welcome!

If they can all migrate, load balancing can move any of them to try to
fix the permanent imbalance, so they'll all bounce about sharing a CPU
with some other hog, and it all kinda sorta works out.

When most are pinned, to make it work out long term you'd have to be
short term unfair, walking the unpinned minority around the box in a
carefully orchestrated dance... and have omniscient powers that assure
that none of the tasks you're trying to equalize is gonna do something
rude like leave, sleep, fork or whatever, and muck up the grand plan.

	-Mike

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1240336

Frompaul.szabo@sydney.edu.au
Date2015-10-06 12:10 +0200
Message-ID<qgxm2-30l-21@gated-at.bofh.it>
In reply to#1240125
Dear Mike,

>> .. CFS ... unfairly gives more CPU time to [pinned] tasks ...
>
> If they can all migrate, load balancing can move any of them to try to
> fix the permanent imbalance, so they'll all bounce about sharing a CPU
> with some other hog, and it all kinda sorta works out.
>
> When most are pinned, to make it work out long term you'd have to be
> short term unfair, walking the unpinned minority around the box in a
> carefully orchestrated dance... and have omniscient powers that assure
> that none of the tasks you're trying to equalize is gonna do something
> rude like leave, sleep, fork or whatever, and muck up the grand plan.

Could not your argument be turned around: for a pinned task it is harder
to find an idle CPU, so they should get less time?

But really... those pinned tasks do not hog the CPU forever. Whatever
kicks them off: could not that be done just a little earlier?

And further... the CFS is meant to be fair, using things like vruntime
to preempt, and throttling. Why are those pinned tasks not preempted or
throttled?

Thanks, Paul

Paul Szabo   psz@maths.usyd.edu.au   http://www.maths.usyd.edu.au/u/psz/
School of Mathematics and Statistics   University of Sydney    Australia
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1240399

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2015-10-06 14:20 +0200
Message-ID<qgznP-5Sf-11@gated-at.bofh.it>
In reply to#1240336
On Tue, 2015-10-06 at 21:06 +1100, paul.szabo@sydney.edu.au wrote:

> And further... the CFS is meant to be fair, using things like vruntime
> to preempt, and throttling. Why are those pinned tasks not preempted or
> throttled?

Imagine you own a 8192 CPU box for a moment, all CPUs having one pinned
task, plus one extra unpinned task, and ponder what would have to happen
in order to meet your utilization expectation.  <time passes>  Right.

What you're seeing is not a bug.  No task can occupy more than one CPU
at a time, making space reservation on multiple CPUs a very bad idea.

	-Mike

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1240983

Frompaul.szabo@sydney.edu.au
Date2015-10-06 22:50 +0200
Message-ID<qgHlo-p0-7@gated-at.bofh.it>
In reply to#1240399
Dear Mike,

>> ... the CFS is meant to be fair, using things like vruntime
>> to preempt, and throttling. Why are those pinned tasks not preempted or
>> throttled?
>
> Imagine you own a 8192 CPU box for a moment, all CPUs having one pinned
> task, plus one extra unpinned task, and ponder what would have to happen
> in order to meet your utilization expectation. ...

Sorry but the kernel contradicts. As per my original report, things are
"fair" in the case of:
 - with CGROUP controls and the two kinds of processes run by
   different users, when there is just one un-pinned process
and that is so on my quad-core i5-3470 baby or my 32-core 4*E5-4627v2
server (and everywhere that I tested). The kernel is smart and gets it
right for one un-pinned process: why not for two?

Now re-testing further (on some machines with CGROUP): on the i5-3470
things are fair still with one un-pinned (become un-fair with two), on
the 4*E5-4627v2 are fair still with 4 un-pinned (become un-fair with 5).
Does this suggest that the kernel does things right within each physical
CPU, but breaks across several (or exact contrary)? Maybe not: on a
2*E5530 machine, things are fair with just one un-pinned and un-fair
with 2 already.

> What you're seeing is not a bug.  No task can occupy more than one CPU
> at a time, making space reservation on multiple CPUs a very bad idea.

I agree that pinning may be bad... should not the kernel penalize the
badly pinned processes?

Cheers, Paul

Paul Szabo   psz@maths.usyd.edu.au   http://www.maths.usyd.edu.au/u/psz/
School of Mathematics and Statistics   University of Sydney    Australia
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1241118

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2015-10-07 03:30 +0200
Message-ID<qgLIl-6JQ-5@gated-at.bofh.it>
In reply to#1240983
On Wed, 2015-10-07 at 07:44 +1100, paul.szabo@sydney.edu.au wrote:

> I agree that pinning may be bad... should not the kernel penalize the
> badly pinned processes?

I didn't say pinning is bad, I said was what you're seeing is not a bug.

	-Mike

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1242048

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2015-10-08 10:30 +0200
Message-ID<qheKm-6pG-17@gated-at.bofh.it>
In reply to#1240125
On Tue, 2015-10-06 at 04:45 +0200, Mike Galbraith wrote:
> On Tue, 2015-10-06 at 08:48 +1100, paul.szabo@sydney.edu.au wrote:
> > The Linux CFS scheduler prefers pinned tasks and unfairly
> > gives more CPU time to tasks that have set CPU affinity.
> > This effect is observed with or without CGROUP controls.
> > 
> > To demonstrate: on an otherwise idle machine, as some user
> > run several processes pinned to each CPU, one for each CPU
> > (as many as CPUs present in the system) e.g. for a quad-core
> > non-HyperThreaded machine:
> > 
> >   taskset -c 0 perl -e 'while(1){1}' &
> >   taskset -c 1 perl -e 'while(1){1}' &
> >   taskset -c 2 perl -e 'while(1){1}' &
> >   taskset -c 3 perl -e 'while(1){1}' &
> > 
> > and (as that same or some other user) run some without
> > pinning:
> > 
> >   perl -e 'while(1){1}' &
> >   perl -e 'while(1){1}' &
> > 
> > and use e.g.   top   to observe that the pinned processes get
> > more CPU time than "fair".

I see a fairness issue with pinned tasks and group scheduling, but one
opposite to your complaint.
 
Two task groups, one with 8 hogs (oink), one with 1 (pert), all are pinned.
  PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ P COMMAND
 3269 root      20   0    4060    724    648 R 100.0 0.004   1:00.02 1 oink
 3270 root      20   0    4060    652    576 R 100.0 0.004   0:59.84 2 oink
 3271 root      20   0    4060    692    616 R 100.0 0.004   0:59.95 3 oink
 3274 root      20   0    4060    608    532 R 100.0 0.004   1:00.01 6 oink
 3273 root      20   0    4060    728    652 R 99.90 0.005   0:59.98 5 oink
 3272 root      20   0    4060    644    568 R 99.51 0.004   0:59.80 4 oink
 3268 root      20   0    4060    612    536 R 99.41 0.004   0:59.67 0 oink
 3279 root      20   0    8312    804    708 R 88.83 0.005   0:53.06 7 pert
 3275 root      20   0    4060    656    580 R 11.07 0.004   0:06.98 7 oink
.
That group share math would make a huge compute group with progress
checkpoints sharing an SGI monster with one other hog amusing to watch.
  
	-Mike

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1242192

Frompaul.szabo@sydney.edu.au
Date2015-10-08 13:00 +0200
Message-ID<qhh5w-1dD-11@gated-at.bofh.it>
In reply to#1242048
Dear Mike,

> I see a fairness issue ... but one opposite to your complaint.

Why is that opposite? I think it would be fair for the one pert process
to get 100% CPU, the many oink processes can get everything else. That
one oink is lowly 10% (when others are 100%) is of no consequence.

What happens when you un-pin pert: does it get 100%? What if you run two
perts? Have you reproduced my observations?

---

Good to see that you agree on the fairness issue... it MUST be fixed!
CFS might be wrong or wasteful, but never unfair.

Cheers, Paul

Paul Szabo   psz@maths.usyd.edu.au   http://www.maths.usyd.edu.au/u/psz/
School of Mathematics and Statistics   University of Sydney    Australia
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1242220

FromPeter Zijlstra <peterz@infradead.org>
Date2015-10-08 13:30 +0200
Message-ID<qhhyy-21q-35@gated-at.bofh.it>
In reply to#1242192
On Thu, Oct 08, 2015 at 09:54:21PM +1100, paul.szabo@sydney.edu.au wrote:
> Good to see that you agree on the fairness issue... it MUST be fixed!
> CFS might be wrong or wasteful, but never unfair.

I've not yet had time to look at the case at hand, but there are wat is
called 'infeasible weight' scenarios for which it is impossible to be
fair.

Also, CFS must remain a practical scheduler, which places bounds on the
amount of weird cases we can deal with.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1243896 — [patch] sched: disable task group re-weighting on the desktop

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2015-10-10 15:30 +0200
Subject[patch] sched: disable task group re-weighting on the desktop
Message-ID<qi2nM-29y-1@gated-at.bofh.it>
In reply to#1242220
On Thu, 2015-10-08 at 13:19 +0200, Peter Zijlstra wrote:
> On Thu, Oct 08, 2015 at 09:54:21PM +1100, paul.szabo@sydney.edu.au wrote:
> > Good to see that you agree on the fairness issue... it MUST be fixed!
> > CFS might be wrong or wasteful, but never unfair.
> 
> I've not yet had time to look at the case at hand, but there are wat is
> called 'infeasible weight' scenarios for which it is impossible to be
> fair.

And sometimes, group wide fairness ain't all that wonderful anyway.

> Also, CFS must remain a practical scheduler, which places bounds on the
> amount of weird cases we can deal with.

Yup, and on a practical note...

master, 1 group of 8 (oink) vs 8 groups of 1 (pert)
  PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ P COMMAND                                                                                                                                 
 5618 root      20   0    8312    840    744 R 90.46 0.005   1:40.48 0 pert                                                                                                                                    
 5630 root      20   0    8312    720    624 R 90.46 0.004   1:38.40 4 pert                                                                                                                                    
 5615 root      20   0    8312    768    672 R 89.48 0.005   1:39.25 6 pert                                                                                                                                    
 5621 root      20   0    8312    792    696 R 89.34 0.005   1:38.49 2 pert                                                                                                                                    
 5627 root      20   0    8312    760    664 R 89.06 0.005   1:36.53 5 pert                                                                                                                                    
 5645 root      20   0    8312    804    708 R 89.06 0.005   1:34.69 1 pert                                                                                                                                    
 5624 root      20   0    8312    716    620 R 88.64 0.004   1:38.45 7 pert                                                                                                                                    
 5612 root      20   0    8312    716    620 R 83.03 0.004   1:40.11 3 pert                                                                                                                                    
 5633 root      20   0    8312    792    696 R 10.94 0.005   0:11.59 4 oink                                                                                                                                    
 5635 root      20   0    8312    804    708 R 10.80 0.005   0:11.74 2 oink                                                                                                                                    
 5637 root      20   0    8312    796    700 R 10.80 0.005   0:11.34 5 oink                                                                                                                                    
 5639 root      20   0    8312    836    740 R 10.80 0.005   0:11.71 2 oink                                                                                                                                    
 5634 root      20   0    8312    840    744 R 10.66 0.005   0:11.36 7 oink                                                                                                                                    
 5636 root      20   0    8312    756    660 R 10.66 0.005   0:11.68 1 oink                                                                                                                                    
 5640 root      20   0    8312    752    656 R 10.10 0.005   0:11.41 7 oink                                                                                                                                    
 5638 root      20   0    8312    804    708 R 9.818 0.005   0:11.99 7 oink

Avg 98.2s per group vs 92.8s for the 8 task group.  Not _perfect_, but ok.

Before reading further, now would be a good time for readers to chant the
"perfect is the enemy of good" mantra, pretending my not so scientific
measurements had actually shown perfect group wide distribution.  You're
gonna see good, and it doesn't resemble perfect.. which is good ;-)

master+, 1 group of 8 (oink) vs 8 groups of 1 (pert)
  PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ P COMMAND                                                                                                                                 
19269 root      20   0    8312    716    620 R 77.25 0.004   1:39.43 2 pert                                                                                                                                    
19263 root      20   0    8312    752    656 R 76.65 0.005   1:43.70 7 pert                                                                                                                                    
19257 root      20   0    8312    760    664 R 72.85 0.005   1:37.08 5 pert                                                                                                                                    
19260 root      20   0    8312    804    704 R 71.86 0.005   1:40.42 1 pert                                                                                                                                    
19273 root      20   0    8312    748    652 R 71.26 0.005   1:41.98 6 pert                                                                                                                                    
19266 root      20   0    8312    752    656 R 67.47 0.005   1:41.69 4 pert                                                                                                                                    
19254 root      20   0    8312    744    648 R 61.28 0.005   1:42.88 4 pert                                                                                                                                    
19277 root      20   0    8312    836    740 R 56.29 0.005   0:46.16 5 oink                                                                                                                                    
19281 root      20   0    8312    768    672 R 55.89 0.005   0:42.05 0 oink                                                                                                                                    
19283 root      20   0    8312    840    744 R 44.91 0.005   0:53.05 3 oink                                                                                                                                    
19282 root      20   0    8312    800    704 R 30.74 0.005   0:41.70 3 oink                                                                                                                                    
19284 root      20   0    8312    724    628 R 28.14 0.004   0:42.08 3 oink                                                                                                                                    
19278 root      20   0    8312    752    656 R 25.15 0.005   0:42.26 3 oink                                                                                                                                    
19280 root      20   0    8312    756    660 R 24.35 0.005   0:40.39 3 oink                                                                                                                                    
19279 root      20   0    8312    836    740 R 23.95 0.005   0:45.71 3 oink

Avg 101.6s per pert group vs 353.4s for the 8 task oink group.  Not remotely
fair total group utilization wise.  Ah, but now onward to interactivity...

master, 8 groups of 1 (pert) vs desktop (mplayer BigBuckBunny-DivXPlusHD.mkv)
  PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ P COMMAND                                                                                                                                 
 4068 root      20   0    8312    724    628 R 99.64 0.004   1:04.32 6 pert                                                                                                                                    
 4065 root      20   0    8312    744    648 R 99.45 0.005   1:04.92 5 pert                                                                                                                                    
 4071 root      20   0    8312    748    652 R 99.27 0.005   1:03.12 7 pert                                                                                                                                    
 4077 root      20   0    8312    840    744 R 98.72 0.005   1:01.46 3 pert                                                                                                                                    
 4074 root      20   0    8312    796    700 R 98.18 0.005   1:03.38 1 pert                                                                                                                                    
 4079 root      20   0    8312    720    624 R 97.99 0.004   1:01.45 4 pert                                                                                                                                    
 4062 root      20   0    8312    836    740 R 96.72 0.005   1:03.44 0 pert                                                                                                                                    
 4059 root      20   0    8312    720    624 R 94.16 0.004   1:04.92 2 pert                                                                                                                                    
 4082 root      20   0 1094400 154324  33592 S 4.197 0.954   0:02.69 0 mplayer                                                                                                                                 
 1029 root      20   0  465332 151540  40816 R 3.285 0.937   0:24.59 2 Xorg                                                                                                                                    
 1773 root      20   0  662592  73308  42012 S 2.007 0.453   0:12.84 5 konsole                                                                                                                                 
  771 root      20   0   11416   1964   1824 S 0.730 0.012   0:10.45 0 rngd                                                                                                                                    
 1722 root      20   0 2866772  65224  51152 S 0.365 0.403   0:03.44 2 kwin                                                                                                                                    
 1769 root      20   0  711684  54212  38020 S 0.182 0.335   0:00.39 1 kmix

That is NOT good.  Mplayer and friends need more than that.  Interactivity
is _horrible_, and buck is an unwatchable mess (no biggy, I know every frame).

master+, 8 groups of 1 (pert) vs desktop (mplayer BigBuckBunny-DivXPlusHD.mkv)
  PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ P COMMAND                                                                                                                                 
 4346 root      20   0    8312    756    660 R 99.20 0.005   0:59.89 5 pert                                                                                                                                    
 4349 root      20   0    8312    748    652 R 98.80 0.005   1:00.77 6 pert                                                                                                                                    
 4343 root      20   0    8312    720    624 R 94.81 0.004   1:02.11 2 pert                                                                                                                                    
 4331 root      20   0    8312    724    628 R 91.22 0.004   1:01.16 3 pert                                                                                                                                    
 4340 root      20   0    8312    720    624 R 91.22 0.004   1:01.06 7 pert                                                                                                                                    
 4328 root      20   0    8312    836    740 R 90.42 0.005   1:00.07 4 pert                                                                                                                                    
 4334 root      20   0    8312    756    660 R 87.82 0.005   0:59.84 1 pert                                                                                                                                    
 4337 root      20   0    8312    824    728 R 76.85 0.005   0:52.20 0 pert                                                                                                                                    
 4352 root      20   0 1058812 123876  33388 S 29.34 0.766   0:25.01 3 mplayer                                                                                                                                 
 1029 root      20   0  471168 156748  40316 R 22.36 0.969   0:42.23 3 Xorg                                                                                                                                    
 1773 root      20   0  663080  74176  42012 S 4.192 0.459   0:17.98 1 konsole                                                                                                                                 
  771 root      20   0   11416   1964   1824 R 1.198 0.012   0:13.45 0 rngd                                                                                                                                    
 1722 root      20   0 2866880  65340  51152 R 0.599 0.404   0:04.87 3 kwin                                                                                                                                    
 1788 root       9 -11  516744  11932   8536 S 0.599 0.074   0:01.01 0 pulseaudio                                                                                                                              
 1733 root      20   0 3369480 141564  71776 S 0.200 0.875   0:05.51 1 plasma-desktop                                                                                                                          

That's good.  Interactivity is fine, I can't even tell pert groups exist by
watching buck kick squirrel butt for the 10387th time.  With master, and one 8
hog group vs desktop/mplayer, I can see the hog group interfere with mplayer.
Add another hog group, mplayer lurches quite badly.  I can feel even one group
while using mouse wheel to scroll through mail.  With master+, I see/feel none
of that unpleasantness.

Conclusion: task group re-weighting is the mortal enemy of a good desktop.

sched: disable task group re-weighting on the desktop

Task group wide utilization based weight may work well for servers, but it
is horrible on the desktop.  8 groups of 1 hog demoloshes interactivity, 1
group of 8 hogs has noticable impact, 2 such groups is very very noticable.

Turn it off if autogroup is enabled, and add a feature to let people set the
definition of fair to what serves them best.  For the desktop, fixed group
weight wins hands down, no contest....

Signed-off-by: Mike Galbraith <umgwanakikbuit@gmail.com>
---
 kernel/sched/fair.c     |   10 ++++++----
 kernel/sched/features.h |   14 ++++++++++++++
 2 files changed, 20 insertions(+), 4 deletions(-)

--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -2372,6 +2372,8 @@ static long calc_cfs_shares(struct cfs_r
 {
 	long tg_weight, load, shares;
 
+	if (!sched_feat(SMP_FAIR_GROUPS))
+		return tg->shares;
 	tg_weight = calc_tg_weight(tg, cfs_rq);
 	load = cfs_rq_load_avg(cfs_rq);
 
@@ -2420,10 +2422,10 @@ static void update_cfs_shares(struct cfs
 	se = tg->se[cpu_of(rq_of(cfs_rq))];
 	if (!se || throttled_hierarchy(cfs_rq))
 		return;
-#ifndef CONFIG_SMP
-	if (likely(se->load.weight == tg->shares))
-		return;
-#endif
+	if (!IS_ENABLED(CONFIG_SMP) || !sched_feat(SMP_FAIR_GROUPS)) {
+		if (likely(se->load.weight == tg->shares))
+			return;
+	}
 	shares = calc_cfs_shares(cfs_rq, tg);
 
 	reweight_entity(cfs_rq_of(se), se, shares);
--- a/kernel/sched/features.h
+++ b/kernel/sched/features.h
@@ -88,3 +88,17 @@ SCHED_FEAT(LB_MIN, false)
  */
 SCHED_FEAT(NUMA,	true)
 #endif
+
+#if defined(CONFIG_SMP) && defined(CONFIG_FAIR_GROUP_SCHED)
+/*
+ * With SMP_FAIR_GROUPS set, activity group wide determines share for
+ * all froup members.  This does very bad things to interactivity when
+ * a desktop box is heavily loaded.  Default to off when autogroup is
+ * enabled, and let all users set it to what works best for them.
+ */
+#ifndef CONFIG_SCHED_AUTOGROUP
+SCHED_FEAT(SMP_FAIR_GROUPS, true)
+#else
+SCHED_FEAT(SMP_FAIR_GROUPS, false)
+#endif
+#endif


--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1243957 — Re: [patch] sched: disable task group re-weighting on the desktop

FromPeter Zijlstra <peterz@infradead.org>
Date2015-10-10 19:10 +0200
SubjectRe: [patch] sched: disable task group re-weighting on the desktop
Message-ID<qi5OH-7ei-17@gated-at.bofh.it>
In reply to#1243896
On Sat, Oct 10, 2015 at 03:22:49PM +0200, Mike Galbraith wrote:
> Ah, but now onward to interactivity...
>
> master, 8 groups of 1 (pert) vs desktop (mplayer BigBuckBunny-DivXPlusHD.mkv)
>   PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ P COMMAND
>  4068 root      20   0    8312    724    628 R 99.64 0.004   1:04.32 6 pert
>  4065 root      20   0    8312    744    648 R 99.45 0.005   1:04.92 5 pert
>  4071 root      20   0    8312    748    652 R 99.27 0.005   1:03.12 7 pert
>  4077 root      20   0    8312    840    744 R 98.72 0.005   1:01.46 3 pert
>  4074 root      20   0    8312    796    700 R 98.18 0.005   1:03.38 1 pert
>  4079 root      20   0    8312    720    624 R 97.99 0.004   1:01.45 4 pert
>  4062 root      20   0    8312    836    740 R 96.72 0.005   1:03.44 0 pert
>  4059 root      20   0    8312    720    624 R 94.16 0.004   1:04.92 2 pert
>  4082 root      20   0 1094400 154324  33592 S 4.197 0.954   0:02.69 0 mplayer
>  1029 root      20   0  465332 151540  40816 R 3.285 0.937   0:24.59 2 Xorg
>  1773 root      20   0  662592  73308  42012 S 2.007 0.453   0:12.84 5 konsole
>   771 root      20   0   11416   1964   1824 S 0.730 0.012   0:10.45 0 rngd
>  1722 root      20   0 2866772  65224  51152 S 0.365 0.403   0:03.44 2 kwin
>  1769 root      20   0  711684  54212  38020 S 0.182 0.335   0:00.39 1 kmix
>
> That is NOT good.  Mplayer and friends need more than that.  Interactivity
> is _horrible_, and buck is an unwatchable mess (no biggy, I know every frame).
>
> master+, 8 groups of 1 (pert) vs desktop (mplayer BigBuckBunny-DivXPlusHD.mkv)
>   PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ P COMMAND
>  4346 root      20   0    8312    756    660 R 99.20 0.005   0:59.89 5 pert
>  4349 root      20   0    8312    748    652 R 98.80 0.005   1:00.77 6 pert
>  4343 root      20   0    8312    720    624 R 94.81 0.004   1:02.11 2 pert
>  4331 root      20   0    8312    724    628 R 91.22 0.004   1:01.16 3 pert
>  4340 root      20   0    8312    720    624 R 91.22 0.004   1:01.06 7 pert
>  4328 root      20   0    8312    836    740 R 90.42 0.005   1:00.07 4 pert
>  4334 root      20   0    8312    756    660 R 87.82 0.005   0:59.84 1 pert
>  4337 root      20   0    8312    824    728 R 76.85 0.005   0:52.20 0 pert
>  4352 root      20   0 1058812 123876  33388 S 29.34 0.766   0:25.01 3 mplayer
>  1029 root      20   0  471168 156748  40316 R 22.36 0.969   0:42.23 3 Xorg
>  1773 root      20   0  663080  74176  42012 S 4.192 0.459   0:17.98 1 konsole
>   771 root      20   0   11416   1964   1824 R 1.198 0.012   0:13.45 0 rngd
>  1722 root      20   0 2866880  65340  51152 R 0.599 0.404   0:04.87 3 kwin
>  1788 root       9 -11  516744  11932   8536 S 0.599 0.074   0:01.01 0 pulseaudio
>  1733 root      20   0 3369480 141564  71776 S 0.200 0.875   0:05.51 1 plasma-desktop
>
> That's good.  Interactivity is fine...

But the patch is most horrible.. :/ It completely destroys everything
group scheduling is supposed to be.

What are these oink/pert things? Both spinners just with amusing names
to distinguish them?

Is the interactivity the same (horrible) at fe32d3cd5e8e (ie, before the
load tracking rewrite from Yuyang)?
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1243959 — Re: [patch] sched: disable task group re-weighting on the desktop

FromPeter Zijlstra <peterz@infradead.org>
Date2015-10-10 19:20 +0200
SubjectRe: [patch] sched: disable task group re-weighting on the desktop
Message-ID<qi5Yl-7pK-5@gated-at.bofh.it>
In reply to#1243957
On Sat, Oct 10, 2015 at 07:01:42PM +0200, Peter Zijlstra wrote:
> On Sat, Oct 10, 2015 at 03:22:49PM +0200, Mike Galbraith wrote:
> > Ah, but now onward to interactivity...
> >
> > master, 8 groups of 1 (pert) vs desktop (mplayer BigBuckBunny-DivXPlusHD.mkv)
> >   PID USER      PR  NI    VIRT    RES    SHR S  %CPU  %MEM     TIME+ P COMMAND
> >  4068 root      20   0    8312    724    628 R 99.64 0.004   1:04.32 6 pert
> >  4065 root      20   0    8312    744    648 R 99.45 0.005   1:04.92 5 pert
> >  4071 root      20   0    8312    748    652 R 99.27 0.005   1:03.12 7 pert
> >  4077 root      20   0    8312    840    744 R 98.72 0.005   1:01.46 3 pert
> >  4074 root      20   0    8312    796    700 R 98.18 0.005   1:03.38 1 pert
> >  4079 root      20   0    8312    720    624 R 97.99 0.004   1:01.45 4 pert
> >  4062 root      20   0    8312    836    740 R 96.72 0.005   1:03.44 0 pert
> >  4059 root      20   0    8312    720    624 R 94.16 0.004   1:04.92 2 pert
> >  4082 root      20   0 1094400 154324  33592 S 4.197 0.954   0:02.69 0 mplayer
> >  1029 root      20   0  465332 151540  40816 R 3.285 0.937   0:24.59 2 Xorg
> >  1773 root      20   0  662592  73308  42012 S 2.007 0.453   0:12.84 5 konsole
> >   771 root      20   0   11416   1964   1824 S 0.730 0.012   0:10.45 0 rngd
> >  1722 root      20   0 2866772  65224  51152 S 0.365 0.403   0:03.44 2 kwin
> >  1769 root      20   0  711684  54212  38020 S 0.182 0.335   0:00.39 1 kmix

Ah wait, so you have 8 groups of cycle soakers vs 1 group of desktop?
That means your desktop will get 1/9 th of the total time and that is
almost so:

 2.69+24.59+12.84+10.45+3.44+.39 = 54.4

vs

 (64.32+64.92+63.12+61.46+63.38+61.45+63.44+64.92)/8 = 63.37625

Which isn't too far off.

This really appears to be a case where you get what you ask for.

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1244048 — Re: [patch] sched: disable task group re-weighting on the desktop

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2015-10-11 04:30 +0200
SubjectRe: [patch] sched: disable task group re-weighting on the desktop
Message-ID<qieyB-30m-5@gated-at.bofh.it>
In reply to#1243957
On Sat, 2015-10-10 at 19:01 +0200, Peter Zijlstra wrote:

> But the patch is most horrible.. :/ It completely destroys everything
> group scheduling is supposed to be.

Yeah, and it works great but...

> What are these oink/pert things? Both spinners just with amusing names
> to distinguish them?

(yeah).

> Is the interactivity the same (horrible) at fe32d3cd5e8e (ie, before the
> load tracking rewrite from Yuyang)?

...you're right.  fe32d3cd5e8e isn't as good as master with a big dent
in its skull, but it is far from the ugly beast I clubbed to death.

	-Mike

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1244175 — 4.3 group scheduling regression

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2015-10-11 19:50 +0200
Subject4.3 group scheduling regression
Message-ID<qisUV-71l-1@gated-at.bofh.it>
In reply to#1244048
(change subject, CCs)

On Sun, 2015-10-11 at 04:25 +0200, Mike Galbraith wrote:

> > Is the interactivity the same (horrible) at fe32d3cd5e8e (ie, before the
> > load tracking rewrite from Yuyang)?

It is the rewrite, 9d89c257dfb9c51a532d69397f6eed75e5168c35.

Watching 8 single hog groups vs 1 tbench group, master vs 4.2.3, I saw
no big hairy difference, just as 1 group of 8 hogs vs 8 groups of 1.

8 single hog groups vs the less hungry mplayer otoh is quite different

100 second scripted recordings:
(note: "testo" is kde konsole acting as task group launch vehicle)

master
 -----------------------------------------------------------------------------------------------------------------
  Task                  |   Runtime ms  | Switches | Average delay ms | Maximum delay ms | Maximum delay at       |
 -----------------------------------------------------------------------------------------------------------------
  oink:(8)              | 787637.964 ms |    16242 | avg:    0.557 ms | max:   68.993 ms | max at:    239.126118 s
  mplayer:(25)          |   5477.234 ms |     8504 | avg:   16.395 ms | max: 2100.233 ms | max at:    282.850734 s
  Xorg:997              |   1773.218 ms |     4680 | avg:    4.857 ms | max: 1640.194 ms | max at:    285.660210 s
  konsole:1789          |    649.323 ms |     1261 | avg:    6.747 ms | max:  156.282 ms | max at:    265.548523 s
  testo:(9)             |    454.046 ms |     2867 | avg:    5.961 ms | max:  276.371 ms | max at:    245.511282 s
  plasma-desktop:1753   |    223.251 ms |     1582 | avg:    4.220 ms | max:  299.354 ms | max at:    337.242542 s
  kwin:1745             |    156.746 ms |     2879 | avg:    2.398 ms | max:  355.765 ms | max at:    337.242490 s
  pulseaudio:1797       |     60.268 ms |     2573 | avg:    0.695 ms | max:   36.069 ms | max at:    292.318120 s
  threaded-ml:3477      |     47.076 ms |     3878 | avg:    7.083 ms | max: 1898.940 ms | max at:    254.919367 s
  perf:3437             |     28.525 ms |        4 | avg:  129.042 ms | max:  498.816 ms | max at:    336.102154 s

4.2.3
 -----------------------------------------------------------------------------------------------------------------
  Task                  |   Runtime ms  | Switches | Average delay ms | Maximum delay ms | Maximum delay at       |
 -----------------------------------------------------------------------------------------------------------------
  oink:(8)              | 741307.292 ms |    42325 | avg:    1.276 ms | max:   23.598 ms | max at:    192.459790 s
  mplayer:(25)          |  35296.804 ms |    35423 | avg:    1.715 ms | max:   71.972 ms | max at:    128.737783 s
  Xorg:929              |  13257.917 ms |    21583 | avg:    0.091 ms | max:   27.983 ms | max at:    102.272376 s
  testo:(9)             |   2315.080 ms |    13213 | avg:    0.133 ms | max:    6.632 ms | max at:    201.422570 s
  konsole:1747          |    938.939 ms |     1458 | avg:    0.096 ms | max:   15.006 ms | max at:    102.260294 s
  kwin:1703             |    815.384 ms |    17376 | avg:    0.464 ms | max:    9.311 ms | max at:    119.026179 s
  pulseaudio:1762       |    396.168 ms |    14338 | avg:    0.020 ms | max:    6.514 ms | max at:    115.928179 s
  threaded-ml:3477      |    310.132 ms |    23966 | avg:    0.428 ms | max:   27.974 ms | max at:    134.100588 s
  plasma-desktop:1711   |    239.232 ms |     1577 | avg:    0.048 ms | max:    7.072 ms | max at:    102.060279 s
  perf:3434             |     65.705 ms |        2 | avg:    0.054 ms | max:    0.105 ms | max at:    102.011221 s

master, mplayer solo reference
 -----------------------------------------------------------------------------------------------------------------
  Task                  |   Runtime ms  | Switches | Average delay ms | Maximum delay ms | Maximum delay at       |
 -----------------------------------------------------------------------------------------------------------------
  mplayer:(25)          |  32171.732 ms |    18416 | avg:    0.012 ms | max:    4.405 ms | max at:   4911.226038 s
  Xorg:948              |  14271.286 ms |    17396 | avg:    0.016 ms | max:    0.082 ms | max at:   4911.243020 s
  testo:4121            |   3594.784 ms |    11607 | avg:    0.015 ms | max:    0.078 ms | max at:   4981.705240 s
  kwin:1650             |   1209.387 ms |    17562 | avg:    0.012 ms | max:    1.612 ms | max at:   4911.245523 s
  konsole:1728          |    967.914 ms |     1498 | avg:    0.007 ms | max:    0.048 ms | max at:   4997.903759 s
  pulseaudio:1750       |    684.342 ms |    14460 | avg:    0.013 ms | max:    0.552 ms | max at:   4957.743502 s
  threaded-ml:4153      |    641.893 ms |    15748 | avg:    0.016 ms | max:    2.201 ms | max at:   4923.928810 s
  plasma-desktop:1658   |    150.068 ms |      569 | avg:    0.011 ms | max:    0.390 ms | max at:   4911.258650 s
  perf:4126             |     43.854 ms |        3 | avg:    0.022 ms | max:    0.051 ms | max at:   4959.327694 s

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1244432 — Re: 4.3 group scheduling regression

FromPeter Zijlstra <peterz@infradead.org>
Date2015-10-12 09:30 +0200
SubjectRe: 4.3 group scheduling regression
Message-ID<qiFIu-pP-7@gated-at.bofh.it>
In reply to#1244175
On Sun, Oct 11, 2015 at 07:42:01PM +0200, Mike Galbraith wrote:
> (change subject, CCs)
> 
> On Sun, 2015-10-11 at 04:25 +0200, Mike Galbraith wrote:
> 
> > > Is the interactivity the same (horrible) at fe32d3cd5e8e (ie, before the
> > > load tracking rewrite from Yuyang)?
> 
> It is the rewrite, 9d89c257dfb9c51a532d69397f6eed75e5168c35.

Just to be sure, so 9d89c257dfb9^1 is good, while 9d89c257dfb9 is bad?

And *groan*, _just_ the thing I need on a monday morning ;-)
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1244450 — Re: 4.3 group scheduling regression

FromMike Galbraith <umgwanakikbuti@gmail.com>
Date2015-10-12 09:50 +0200
SubjectRe: 4.3 group scheduling regression
Message-ID<qiG1P-MT-7@gated-at.bofh.it>
In reply to#1244432
On Mon, 2015-10-12 at 09:23 +0200, Peter Zijlstra wrote:
> On Sun, Oct 11, 2015 at 07:42:01PM +0200, Mike Galbraith wrote:
> > (change subject, CCs)
> > 
> > On Sun, 2015-10-11 at 04:25 +0200, Mike Galbraith wrote:
> > 
> > > > Is the interactivity the same (horrible) at fe32d3cd5e8e (ie, before the
> > > > load tracking rewrite from Yuyang)?
> > 
> > It is the rewrite, 9d89c257dfb9c51a532d69397f6eed75e5168c35.
> 
> Just to be sure, so 9d89c257dfb9^1 is good, while 9d89c257dfb9 is bad?

Yeah, I went ahead and bisected.
 
> And *groan*, _just_ the thing I need on a monday morning ;-)

Sorry 'bout that.

It's odd to me that things look pretty much the same good/bad tree with
hogs vs hogs or hogs vs tbench (with top anyway, just adding up times).
Seems Xorg+mplayer more or less playing cross group ping-pong must be
the BadThing trigger.

	-Mike


--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1244468 — Re: 4.3 group scheduling regression

FromPeter Zijlstra <peterz@infradead.org>
Date2015-10-12 10:10 +0200
SubjectRe: 4.3 group scheduling regression
Message-ID<qiGld-1p0-15@gated-at.bofh.it>
In reply to#1244450
On Mon, Oct 12, 2015 at 09:44:57AM +0200, Mike Galbraith wrote:

> It's odd to me that things look pretty much the same good/bad tree with
> hogs vs hogs or hogs vs tbench (with top anyway, just adding up times).
> Seems Xorg+mplayer more or less playing cross group ping-pong must be
> the BadThing trigger.

Ohh, wait, Xorg and mplayer are _not_ in the same group? I was assuming
you had your entire user session in 1 (auto) group and was competing
against 8 manual cgroups.

So how exactly are things configured?
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1244489 — Re: 4.3 group scheduling regression

FromYuyang Du <yuyang.du@intel.com>
Date2015-10-12 10:50 +0200
SubjectRe: 4.3 group scheduling regression
Message-ID<qiGXT-28r-5@gated-at.bofh.it>
In reply to#1244468
Good morning, Peter.

On Mon, Oct 12, 2015 at 10:04:07AM +0200, Peter Zijlstra wrote:
> On Mon, Oct 12, 2015 at 09:44:57AM +0200, Mike Galbraith wrote:
> 
> > It's odd to me that things look pretty much the same good/bad tree with
> > hogs vs hogs or hogs vs tbench (with top anyway, just adding up times).
> > Seems Xorg+mplayer more or less playing cross group ping-pong must be
> > the BadThing trigger.
>
> Ohh, wait, Xorg and mplayer are _not_ in the same group? I was assuming
> you had your entire user session in 1 (auto) group and was competing
> against 8 manual cgroups.
> 
> So how exactly are things configured?
 
Hmm... my impression is the naughty boy mplayer (+Xorg) isn't favored, due 
to the per CPU group entity share distribution. Let me dig more.

Sorry.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1244516 — Re: 4.3 group scheduling regression

FromPeter Zijlstra <peterz@infradead.org>
Date2015-10-12 11:20 +0200
SubjectRe: 4.3 group scheduling regression
Message-ID<qiHqW-2Xd-19@gated-at.bofh.it>
In reply to#1244489
On Mon, Oct 12, 2015 at 08:53:51AM +0800, Yuyang Du wrote:
> Good morning, Peter.
> 
> On Mon, Oct 12, 2015 at 10:04:07AM +0200, Peter Zijlstra wrote:
> > On Mon, Oct 12, 2015 at 09:44:57AM +0200, Mike Galbraith wrote:
> > 
> > > It's odd to me that things look pretty much the same good/bad tree with
> > > hogs vs hogs or hogs vs tbench (with top anyway, just adding up times).
> > > Seems Xorg+mplayer more or less playing cross group ping-pong must be
> > > the BadThing trigger.
> >
> > Ohh, wait, Xorg and mplayer are _not_ in the same group? I was assuming
> > you had your entire user session in 1 (auto) group and was competing
> > against 8 manual cgroups.
> > 
> > So how exactly are things configured?
>  
> Hmm... my impression is the naughty boy mplayer (+Xorg) isn't favored, due 
> to the per CPU group entity share distribution. Let me dig more.

So in the old code we had 'magic' to deal with the case where a cgroup
was consuming less than 1 cpu's worth of runtime. For example, a single
task running in the group.

In that scenario it might be possible that the group entity weight:

	se->weight = (tg->shares * cfs_rq->weight) / tg->weight;

Strongly deviates from the tg->shares; you want the single task reflect
the full group shares to the next level; due to the whole distributed
approximation stuff.

I see you've deleted all that code; see the former
__update_group_entity_contrib().

It could be that we need to bring that back. But let me think a little
bit more on this.. I'm having a hard time waking :/
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1244559 — Re: 4.3 group scheduling regression

FromYuyang Du <yuyang.du@intel.com>
Date2015-10-12 12:10 +0200
SubjectRe: 4.3 group scheduling regression
Message-ID<qiIdl-488-13@gated-at.bofh.it>
In reply to#1244516
On Mon, Oct 12, 2015 at 11:12:06AM +0200, Peter Zijlstra wrote:
> On Mon, Oct 12, 2015 at 08:53:51AM +0800, Yuyang Du wrote:
> > Good morning, Peter.
> > 
> > On Mon, Oct 12, 2015 at 10:04:07AM +0200, Peter Zijlstra wrote:
> > > On Mon, Oct 12, 2015 at 09:44:57AM +0200, Mike Galbraith wrote:
> > > 
> > > > It's odd to me that things look pretty much the same good/bad tree with
> > > > hogs vs hogs or hogs vs tbench (with top anyway, just adding up times).
> > > > Seems Xorg+mplayer more or less playing cross group ping-pong must be
> > > > the BadThing trigger.
> > >
> > > Ohh, wait, Xorg and mplayer are _not_ in the same group? I was assuming
> > > you had your entire user session in 1 (auto) group and was competing
> > > against 8 manual cgroups.
> > > 
> > > So how exactly are things configured?
> >  
> > Hmm... my impression is the naughty boy mplayer (+Xorg) isn't favored, due 
> > to the per CPU group entity share distribution. Let me dig more.
> 
> So in the old code we had 'magic' to deal with the case where a cgroup
> was consuming less than 1 cpu's worth of runtime. For example, a single
> task running in the group.
> 
> In that scenario it might be possible that the group entity weight:
> 
> 	se->weight = (tg->shares * cfs_rq->weight) / tg->weight;
> 
> Strongly deviates from the tg->shares; you want the single task reflect
> the full group shares to the next level; due to the whole distributed
> approximation stuff.

Yeah, I thought so.
 
> I see you've deleted all that code; see the former
> __update_group_entity_contrib().
 
Probably not there, it actually was an icky way to adjust things.

> It could be that we need to bring that back. But let me think a little
> bit more on this.. I'm having a hard time waking :/

I am guessing it is in calc_tg_weight(), and naughty boys do make them more
favored, what a reality...

Mike, beg you test the following?

--

diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 4df37a4..b184da0 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -2370,7 +2370,7 @@ static inline long calc_tg_weight(struct task_group *tg, struct cfs_rq *cfs_rq)
 	 */
 	tg_weight = atomic_long_read(&tg->load_avg);
 	tg_weight -= cfs_rq->tg_load_avg_contrib;
-	tg_weight += cfs_rq_load_avg(cfs_rq);
+	tg_weight += cfs_rq->load.weight;
 
 	return tg_weight;
 }
@@ -2380,7 +2380,7 @@ static long calc_cfs_shares(struct cfs_rq *cfs_rq, struct task_group *tg)
 	long tg_weight, load, shares;
 
 	tg_weight = calc_tg_weight(tg, cfs_rq);
-	load = cfs_rq_load_avg(cfs_rq);
+	load = cfs_rq->load.weight;
 
 	shares = (tg->shares * load);
 	if (tg_weight)
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | linux.kernel


csiph-web