Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1240125 > unrolled thread
| Started by | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| First post | 2015-10-06 04:50 +0200 |
| Last post | 2015-10-10 10:00 +0200 |
| Articles | 20 on this page of 45 — 5 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-06 04:50 +0200
Re: CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-06 12:10 +0200
Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-06 14:20 +0200
Re: CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-06 22:50 +0200
Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-07 03:30 +0200
Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-08 10:30 +0200
Re: CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-08 13:00 +0200
Re: CFS scheduler unfairly prefers pinned tasks Peter Zijlstra <peterz@infradead.org> - 2015-10-08 13:30 +0200
[patch] sched: disable task group re-weighting on the desktop Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-10 15:30 +0200
Re: [patch] sched: disable task group re-weighting on the desktop Peter Zijlstra <peterz@infradead.org> - 2015-10-10 19:10 +0200
Re: [patch] sched: disable task group re-weighting on the desktop Peter Zijlstra <peterz@infradead.org> - 2015-10-10 19:20 +0200
Re: [patch] sched: disable task group re-weighting on the desktop Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-11 04:30 +0200
4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-11 19:50 +0200
Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-12 09:30 +0200
Re: 4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-12 09:50 +0200
Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-12 10:10 +0200
Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-12 10:50 +0200
Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-12 11:20 +0200
Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-12 12:10 +0200
Re: 4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-12 12:30 +0200
Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-13 05:50 +0200
Re: 4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-13 06:10 +0200
Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-13 06:40 +0200
Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-13 10:10 +0200
Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-13 10:20 +0200
Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-13 10:30 +0200
Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-13 10:30 +0200
Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-12 13:50 +0200
Re: 4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-13 04:30 +0200
Re: 4.3 group scheduling regression Yuyang Du <yuyang.du@intel.com> - 2015-10-13 05:30 +0200
Re: 4.3 group scheduling regression Peter Zijlstra <peterz@infradead.org> - 2015-10-13 10:10 +0200
Re: 4.3 group scheduling regression Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-12 10:50 +0200
Re: [patch] sched: disable task group re-weighting on the desktop paul.szabo@sydney.edu.au - 2015-10-10 22:20 +0200
Re: [patch] sched: disable task group re-weighting on the desktop Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-11 04:40 +0200
Re: [patch] sched: disable task group re-weighting on the desktop paul.szabo@sydney.edu.au - 2015-10-11 11:30 +0200
Re: [patch] sched: disable task group re-weighting on the desktop Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-11 15:00 +0200
Re: [patch] sched: disable task group re-weighting on the desktop paul.szabo@sydney.edu.au - 2015-10-11 21:50 +0200
Re: [patch] sched: disable task group re-weighting on the desktop Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-12 04:00 +0200
Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-08 16:30 +0200
Re: CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-09 00:00 +0200
Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-09 04:00 +0200
Re: CFS scheduler unfairly prefers pinned tasks Mike Galbraith <umgwanakikbuti@gmail.com> - 2015-10-09 04:50 +0200
Re: CFS scheduler unfairly prefers pinned tasks paul.szabo@sydney.edu.au - 2015-10-11 11:50 +0200
Re: CFS scheduler unfairly prefers pinned tasks Wanpeng Li <wanpeng.li@hotmail.com> - 2015-10-10 06:10 +0200
Re: CFS scheduler unfairly prefers pinned tasks Wanpeng Li <wanpeng.li@hotmail.com> - 2015-10-10 10:00 +0200
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-10-13 05:50 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qiYL8-2Uy-13@gated-at.bofh.it> |
| In reply to | #1244575 |
On Mon, Oct 12, 2015 at 12:23:31PM +0200, Mike Galbraith wrote:
> On Mon, 2015-10-12 at 10:12 +0800, Yuyang Du wrote:
>
> > I am guessing it is in calc_tg_weight(), and naughty boys do make them more
> > favored, what a reality...
> >
> > Mike, beg you test the following?
>
> Wow, that was quick. Dinky patch made it all better.
>
> -----------------------------------------------------------------------------------------------------------------
> Task | Runtime ms | Switches | Average delay ms | Maximum delay ms | Maximum delay at |
> -----------------------------------------------------------------------------------------------------------------
> oink:(8) | 739056.970 ms | 27270 | avg: 2.043 ms | max: 29.105 ms | max at: 339.988310 s
> mplayer:(25) | 36448.997 ms | 44670 | avg: 1.886 ms | max: 72.808 ms | max at: 302.153121 s
> Xorg:988 | 13334.908 ms | 22210 | avg: 0.081 ms | max: 25.005 ms | max at: 269.068666 s
> testo:(9) | 2558.540 ms | 13703 | avg: 0.124 ms | max: 6.412 ms | max at: 279.235272 s
> konsole:1781 | 1084.316 ms | 1457 | avg: 0.006 ms | max: 1.039 ms | max at: 268.863379 s
> kwin:1734 | 879.645 ms | 17855 | avg: 0.458 ms | max: 15.788 ms | max at: 268.854992 s
> pulseaudio:1808 | 356.334 ms | 15023 | avg: 0.028 ms | max: 6.134 ms | max at: 324.479766 s
> threaded-ml:3483 | 292.782 ms | 25769 | avg: 0.364 ms | max: 40.387 ms | max at: 294.550515 s
> plasma-desktop:1745 | 265.055 ms | 1470 | avg: 0.102 ms | max: 21.886 ms | max at: 267.724902 s
> perf:3439 | 61.677 ms | 2 | avg: 0.117 ms | max: 0.232 ms | max at: 367.043889 s
Phew...
I think maybe the real disease is the tg->load_avg is not updated in time.
I.e., it is after migrate, the source cfs_rq does not decrease its contribution
to the parent's tg->load_avg fast enough.
--
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 4df37a4..3dba883 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -2686,12 +2686,13 @@ static inline u64 cfs_rq_clock_task(struct cfs_rq *cfs_rq);
static inline int update_cfs_rq_load_avg(u64 now, struct cfs_rq *cfs_rq)
{
struct sched_avg *sa = &cfs_rq->avg;
- int decayed;
+ int decayed, updated = 0;
if (atomic_long_read(&cfs_rq->removed_load_avg)) {
long r = atomic_long_xchg(&cfs_rq->removed_load_avg, 0);
sa->load_avg = max_t(long, sa->load_avg - r, 0);
sa->load_sum = max_t(s64, sa->load_sum - r * LOAD_AVG_MAX, 0);
+ updated = 1;
}
if (atomic_long_read(&cfs_rq->removed_util_avg)) {
@@ -2708,7 +2709,7 @@ static inline int update_cfs_rq_load_avg(u64 now, struct cfs_rq *cfs_rq)
cfs_rq->load_last_update_time_copy = sa->last_update_time;
#endif
- return decayed;
+ return decayed | updated;
}
/* Update task and its cfs_rq load average */
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2015-10-13 06:10 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qiZ4t-3w3-9@gated-at.bofh.it> |
| In reply to | #1245323 |
On Tue, 2015-10-13 at 03:55 +0800, Yuyang Du wrote:
> On Mon, Oct 12, 2015 at 12:23:31PM +0200, Mike Galbraith wrote:
> > On Mon, 2015-10-12 at 10:12 +0800, Yuyang Du wrote:
> >
> > > I am guessing it is in calc_tg_weight(), and naughty boys do make them more
> > > favored, what a reality...
> > >
> > > Mike, beg you test the following?
> >
> > Wow, that was quick. Dinky patch made it all better.
> >
> > -----------------------------------------------------------------------------------------------------------------
> > Task | Runtime ms | Switches | Average delay ms | Maximum delay ms | Maximum delay at |
> > -----------------------------------------------------------------------------------------------------------------
> > oink:(8) | 739056.970 ms | 27270 | avg: 2.043 ms | max: 29.105 ms | max at: 339.988310 s
> > mplayer:(25) | 36448.997 ms | 44670 | avg: 1.886 ms | max: 72.808 ms | max at: 302.153121 s
> > Xorg:988 | 13334.908 ms | 22210 | avg: 0.081 ms | max: 25.005 ms | max at: 269.068666 s
> > testo:(9) | 2558.540 ms | 13703 | avg: 0.124 ms | max: 6.412 ms | max at: 279.235272 s
> > konsole:1781 | 1084.316 ms | 1457 | avg: 0.006 ms | max: 1.039 ms | max at: 268.863379 s
> > kwin:1734 | 879.645 ms | 17855 | avg: 0.458 ms | max: 15.788 ms | max at: 268.854992 s
> > pulseaudio:1808 | 356.334 ms | 15023 | avg: 0.028 ms | max: 6.134 ms | max at: 324.479766 s
> > threaded-ml:3483 | 292.782 ms | 25769 | avg: 0.364 ms | max: 40.387 ms | max at: 294.550515 s
> > plasma-desktop:1745 | 265.055 ms | 1470 | avg: 0.102 ms | max: 21.886 ms | max at: 267.724902 s
> > perf:3439 | 61.677 ms | 2 | avg: 0.117 ms | max: 0.232 ms | max at: 367.043889 s
>
> Phew...
>
> I think maybe the real disease is the tg->load_avg is not updated in time.
> I.e., it is after migrate, the source cfs_rq does not decrease its contribution
> to the parent's tg->load_avg fast enough.
It sounded like you wanted me to run the below alone. If so, it's a nogo.
-----------------------------------------------------------------------------------------------------------------
Task | Runtime ms | Switches | Average delay ms | Maximum delay ms | Maximum delay at |
-----------------------------------------------------------------------------------------------------------------
oink:(8) | 787001.236 ms | 21641 | avg: 0.377 ms | max: 21.991 ms | max at: 51.504005 s
mplayer:(25) | 4256.224 ms | 7264 | avg: 19.698 ms | max: 2087.489 ms | max at: 115.294922 s
Xorg:1011 | 1507.958 ms | 4081 | avg: 8.349 ms | max: 1652.200 ms | max at: 126.908021 s
konsole:1752 | 697.806 ms | 1186 | avg: 5.749 ms | max: 160.189 ms | max at: 53.037952 s
testo:(9) | 438.164 ms | 2551 | avg: 6.616 ms | max: 215.527 ms | max at: 117.302455 s
plasma-desktop:1716 | 280.418 ms | 1624 | avg: 3.701 ms | max: 574.806 ms | max at: 53.582261 s
kwin:1708 | 144.986 ms | 2422 | avg: 3.301 ms | max: 315.707 ms | max at: 116.555721 s
> --
>
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 4df37a4..3dba883 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -2686,12 +2686,13 @@ static inline u64 cfs_rq_clock_task(struct cfs_rq *cfs_rq);
> static inline int update_cfs_rq_load_avg(u64 now, struct cfs_rq *cfs_rq)
> {
> struct sched_avg *sa = &cfs_rq->avg;
> - int decayed;
> + int decayed, updated = 0;
>
> if (atomic_long_read(&cfs_rq->removed_load_avg)) {
> long r = atomic_long_xchg(&cfs_rq->removed_load_avg, 0);
> sa->load_avg = max_t(long, sa->load_avg - r, 0);
> sa->load_sum = max_t(s64, sa->load_sum - r * LOAD_AVG_MAX, 0);
> + updated = 1;
> }
>
> if (atomic_long_read(&cfs_rq->removed_util_avg)) {
> @@ -2708,7 +2709,7 @@ static inline int update_cfs_rq_load_avg(u64 now, struct cfs_rq *cfs_rq)
> cfs_rq->load_last_update_time_copy = sa->last_update_time;
> #endif
>
> - return decayed;
> + return decayed | updated;
> }
>
> /* Update task and its cfs_rq load average */
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-10-13 06:40 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qiZxw-46G-7@gated-at.bofh.it> |
| In reply to | #1245328 |
On Tue, Oct 13, 2015 at 06:08:34AM +0200, Mike Galbraith wrote:
> It sounded like you wanted me to run the below alone. If so, it's a nogo.
Yes, thanks.
Then it is the sad fact that after migrate and removed_load_avg is added
in migrate_task_rq_fair(), we don't get a chance to update the tg so fast
that at the destination the mplayer is weighted to the group's share.
> -----------------------------------------------------------------------------------------------------------------
> Task | Runtime ms | Switches | Average delay ms | Maximum delay ms | Maximum delay at |
> -----------------------------------------------------------------------------------------------------------------
> oink:(8) | 787001.236 ms | 21641 | avg: 0.377 ms | max: 21.991 ms | max at: 51.504005 s
> mplayer:(25) | 4256.224 ms | 7264 | avg: 19.698 ms | max: 2087.489 ms | max at: 115.294922 s
> Xorg:1011 | 1507.958 ms | 4081 | avg: 8.349 ms | max: 1652.200 ms | max at: 126.908021 s
> konsole:1752 | 697.806 ms | 1186 | avg: 5.749 ms | max: 160.189 ms | max at: 53.037952 s
> testo:(9) | 438.164 ms | 2551 | avg: 6.616 ms | max: 215.527 ms | max at: 117.302455 s
> plasma-desktop:1716 | 280.418 ms | 1624 | avg: 3.701 ms | max: 574.806 ms | max at: 53.582261 s
> kwin:1708 | 144.986 ms | 2422 | avg: 3.301 ms | max: 315.707 ms | max at: 116.555721 s
>
> > --
> >
> > diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> > index 4df37a4..3dba883 100644
> > --- a/kernel/sched/fair.c
> > +++ b/kernel/sched/fair.c
> > @@ -2686,12 +2686,13 @@ static inline u64 cfs_rq_clock_task(struct cfs_rq *cfs_rq);
> > static inline int update_cfs_rq_load_avg(u64 now, struct cfs_rq *cfs_rq)
> > {
> > struct sched_avg *sa = &cfs_rq->avg;
> > - int decayed;
> > + int decayed, updated = 0;
> >
> > if (atomic_long_read(&cfs_rq->removed_load_avg)) {
> > long r = atomic_long_xchg(&cfs_rq->removed_load_avg, 0);
> > sa->load_avg = max_t(long, sa->load_avg - r, 0);
> > sa->load_sum = max_t(s64, sa->load_sum - r * LOAD_AVG_MAX, 0);
> > + updated = 1;
> > }
> >
> > if (atomic_long_read(&cfs_rq->removed_util_avg)) {
> > @@ -2708,7 +2709,7 @@ static inline int update_cfs_rq_load_avg(u64 now, struct cfs_rq *cfs_rq)
> > cfs_rq->load_last_update_time_copy = sa->last_update_time;
> > #endif
> >
> > - return decayed;
> > + return decayed | updated;
A typo: decayed || updated, but shouldn't make any difference.
> > }
> >
> > /* Update task and its cfs_rq load average */
>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-10-13 10:10 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qj2OL-A2-37@gated-at.bofh.it> |
| In reply to | #1245323 |
On Tue, Oct 13, 2015 at 03:55:17AM +0800, Yuyang Du wrote: > I think maybe the real disease is the tg->load_avg is not updated in time. > I.e., it is after migrate, the source cfs_rq does not decrease its contribution > to the parent's tg->load_avg fast enough. No, using the load_avg for shares calculation seems wrong; that would mean we'd first have to ramp up the avg before you react. You want to react quickly to actual load changes, esp. going up. We use the avg to guess the global group load, since that's the best compromise we have, but locally it doesn't make sense to use the avg if we have the actual values. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-10-13 10:20 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qj2Yq-Lw-17@gated-at.bofh.it> |
| In reply to | #1245438 |
On Tue, Oct 13, 2015 at 10:06:48AM +0200, Peter Zijlstra wrote: > On Tue, Oct 13, 2015 at 03:55:17AM +0800, Yuyang Du wrote: > > > I think maybe the real disease is the tg->load_avg is not updated in time. > > I.e., it is after migrate, the source cfs_rq does not decrease its contribution > > to the parent's tg->load_avg fast enough. > > No, using the load_avg for shares calculation seems wrong; that would > mean we'd first have to ramp up the avg before you react. > > You want to react quickly to actual load changes, esp. going up. > > We use the avg to guess the global group load, since that's the best > compromise we have, but locally it doesn't make sense to use the avg if > we have the actual values. That is, can you send the original patch with a Changelog etc.. so that I can press 'A' :-) -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-10-13 10:30 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qj386-Y7-13@gated-at.bofh.it> |
| In reply to | #1245444 |
On Tue, Oct 13, 2015 at 10:10:23AM +0200, Peter Zijlstra wrote: > On Tue, Oct 13, 2015 at 10:06:48AM +0200, Peter Zijlstra wrote: > > On Tue, Oct 13, 2015 at 03:55:17AM +0800, Yuyang Du wrote: > > > > > I think maybe the real disease is the tg->load_avg is not updated in time. > > > I.e., it is after migrate, the source cfs_rq does not decrease its contribution > > > to the parent's tg->load_avg fast enough. > > > > No, using the load_avg for shares calculation seems wrong; that would > > mean we'd first have to ramp up the avg before you react. > > > > You want to react quickly to actual load changes, esp. going up. > > > > We use the avg to guess the global group load, since that's the best > > compromise we have, but locally it doesn't make sense to use the avg if > > we have the actual values. > > That is, can you send the original patch with a Changelog etc.. so that > I can press 'A' :-) Sure, in minutes, :) -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-10-13 10:30 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qj386-Y7-21@gated-at.bofh.it> |
| In reply to | #1245438 |
On Tue, Oct 13, 2015 at 10:06:48AM +0200, Peter Zijlstra wrote: > On Tue, Oct 13, 2015 at 03:55:17AM +0800, Yuyang Du wrote: > > > I think maybe the real disease is the tg->load_avg is not updated in time. > > I.e., it is after migrate, the source cfs_rq does not decrease its contribution > > to the parent's tg->load_avg fast enough. > > No, using the load_avg for shares calculation seems wrong; that would > mean we'd first have to ramp up the avg before you react. > > You want to react quickly to actual load changes, esp. going up. > > We use the avg to guess the global group load, since that's the best > compromise we have, but locally it doesn't make sense to use the avg if > we have the actual values. In Mike's case, since the mplayer group has only one active task, after the task migrates, the source cfs_rq should have zero contrib to the tg, so at the destination, the group entity should have the entire tg's share. It is just the zeroing can be that fast we need. But yes, in a general case, the load_avg (that has the blocked load) is likely to lag behind. Using the actual load.weight to accelerate the process makes sense. It is especially helpful to the less hungry tasks. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-10-12 13:50 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qiJM6-6go-13@gated-at.bofh.it> |
| In reply to | #1244559 |
On Mon, Oct 12, 2015 at 10:12:31AM +0800, Yuyang Du wrote:
> On Mon, Oct 12, 2015 at 11:12:06AM +0200, Peter Zijlstra wrote:
> > So in the old code we had 'magic' to deal with the case where a cgroup
> > was consuming less than 1 cpu's worth of runtime. For example, a single
> > task running in the group.
> >
> > In that scenario it might be possible that the group entity weight:
> >
> > se->weight = (tg->shares * cfs_rq->weight) / tg->weight;
> >
> > Strongly deviates from the tg->shares; you want the single task reflect
> > the full group shares to the next level; due to the whole distributed
> > approximation stuff.
>
> Yeah, I thought so.
>
> > I see you've deleted all that code; see the former
> > __update_group_entity_contrib().
>
> Probably not there, it actually was an icky way to adjust things.
Yeah, no argument there.
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 4df37a4..b184da0 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -2370,7 +2370,7 @@ static inline long calc_tg_weight(struct task_group *tg, struct cfs_rq *cfs_rq)
> */
> tg_weight = atomic_long_read(&tg->load_avg);
> tg_weight -= cfs_rq->tg_load_avg_contrib;
> - tg_weight += cfs_rq_load_avg(cfs_rq);
> + tg_weight += cfs_rq->load.weight;
>
> return tg_weight;
> }
> @@ -2380,7 +2380,7 @@ static long calc_cfs_shares(struct cfs_rq *cfs_rq, struct task_group *tg)
> long tg_weight, load, shares;
>
> tg_weight = calc_tg_weight(tg, cfs_rq);
> - load = cfs_rq_load_avg(cfs_rq);
> + load = cfs_rq->load.weight;
>
> shares = (tg->shares * load);
> if (tg_weight)
Aah, yes very much so. I completely overlooked that :-(
When calculating shares we very much want the current load, not the load
average.
Also, should we do the below? At this point se->on_rq is still 0 so
reweight_entity() will not update (dequeue/enqueue) the accounting, but
we'll have just accounted the 'old' load.weight.
Doing it this way around we'll first update the weight and then account
it, which seems more accurate.
---
diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
index 700eb548315f..d2efef565aed 100644
--- a/kernel/sched/fair.c
+++ b/kernel/sched/fair.c
@@ -3009,8 +3009,8 @@ enqueue_entity(struct cfs_rq *cfs_rq, struct sched_entity *se, int flags)
*/
update_curr(cfs_rq);
enqueue_entity_load_avg(cfs_rq, se);
- account_entity_enqueue(cfs_rq, se);
update_cfs_shares(cfs_rq);
+ account_entity_enqueue(cfs_rq, se);
if (flags & ENQUEUE_WAKEUP) {
place_entity(cfs_rq, se, 0);
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2015-10-13 04:30 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qiXvH-192-7@gated-at.bofh.it> |
| In reply to | #1244619 |
On Mon, 2015-10-12 at 13:47 +0200, Peter Zijlstra wrote: > Also, should we do the below? Ew. Box said "Either you quilt pop/burn, or I boot windows." ;-) -Mike -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Yuyang Du <yuyang.du@intel.com> |
|---|---|
| Date | 2015-10-13 05:30 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qiYrL-2vS-3@gated-at.bofh.it> |
| In reply to | #1244619 |
On Mon, Oct 12, 2015 at 01:47:23PM +0200, Peter Zijlstra wrote:
>
> Also, should we do the below? At this point se->on_rq is still 0 so
> reweight_entity() will not update (dequeue/enqueue) the accounting, but
> we'll have just accounted the 'old' load.weight.
>
> Doing it this way around we'll first update the weight and then account
> it, which seems more accurate.
I think the original looks ok.
The account_entity_enqueue() adds child entity's load.weight to parent's load:
update_load_add(&cfs_rq->load, se->load.weight)
Then recalculate the shares.
Then reweight_entity() resets the parent entity's load.weight.
> ---
> diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c
> index 700eb548315f..d2efef565aed 100644
> --- a/kernel/sched/fair.c
> +++ b/kernel/sched/fair.c
> @@ -3009,8 +3009,8 @@ enqueue_entity(struct cfs_rq *cfs_rq, struct sched_entity *se, int flags)
> */
> update_curr(cfs_rq);
> enqueue_entity_load_avg(cfs_rq, se);
> - account_entity_enqueue(cfs_rq, se);
> update_cfs_shares(cfs_rq);
> + account_entity_enqueue(cfs_rq, se);
>
> if (flags & ENQUEUE_WAKEUP) {
> place_entity(cfs_rq, se, 0);
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-10-13 10:10 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qj2OL-A2-39@gated-at.bofh.it> |
| In reply to | #1245314 |
On Tue, Oct 13, 2015 at 03:32:47AM +0800, Yuyang Du wrote: > On Mon, Oct 12, 2015 at 01:47:23PM +0200, Peter Zijlstra wrote: > > > > Also, should we do the below? At this point se->on_rq is still 0 so > > reweight_entity() will not update (dequeue/enqueue) the accounting, but > > we'll have just accounted the 'old' load.weight. > > > > Doing it this way around we'll first update the weight and then account > > it, which seems more accurate. > > I think the original looks ok. > > The account_entity_enqueue() adds child entity's load.weight to parent's load: > > update_load_add(&cfs_rq->load, se->load.weight) > > Then recalculate the shares. > > Then reweight_entity() resets the parent entity's load.weight. Yes, some days I should just not be allowed near a keyboard :) -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2015-10-12 10:50 +0200 |
| Subject | Re: 4.3 group scheduling regression |
| Message-ID | <qiGXT-28r-13@gated-at.bofh.it> |
| In reply to | #1244468 |
On Mon, 2015-10-12 at 10:04 +0200, Peter Zijlstra wrote: > On Mon, Oct 12, 2015 at 09:44:57AM +0200, Mike Galbraith wrote: > > > It's odd to me that things look pretty much the same good/bad tree with > > hogs vs hogs or hogs vs tbench (with top anyway, just adding up times). > > Seems Xorg+mplayer more or less playing cross group ping-pong must be > > the BadThing trigger. > > Ohh, wait, Xorg and mplayer are _not_ in the same group? I was assuming > you had your entire user session in 1 (auto) group and was competing > against 8 manual cgroups. > > So how exactly are things configured? I turned autogroup on as to not have to muck about creating groups, so Xorg is in its per session group, and each konsole instance in its. I launched groups via testo (aka konsole) -e <content> in a little script to turn it loose at once to run for 100 seconds and kill itself, but that's not necessary 'course. Start 1 hog in 8 konsole tabs, and mplayer in the 9th, ickiness follows. -Mike -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | paul.szabo@sydney.edu.au |
|---|---|
| Date | 2015-10-10 22:20 +0200 |
| Subject | Re: [patch] sched: disable task group re-weighting on the desktop |
| Message-ID | <qi8My-33u-7@gated-at.bofh.it> |
| In reply to | #1243896 |
Dear Mike, You CCed me on this patch. Is that because you expect this to solve "my" problem also? You had some measurements of many oinks vs many perts or vs "desktop", but not many oinks vs 1 or 2 perts as per my "complaint". You also changed the subject line, so maybe this is all un-related. Thanks, Paul Paul Szabo psz@maths.usyd.edu.au http://www.maths.usyd.edu.au/u/psz/ School of Mathematics and Statistics University of Sydney Australia -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2015-10-11 04:40 +0200 |
| Subject | Re: [patch] sched: disable task group re-weighting on the desktop |
| Message-ID | <qieIh-3cX-3@gated-at.bofh.it> |
| In reply to | #1244005 |
On Sun, 2015-10-11 at 07:14 +1100, paul.szabo@sydney.edu.au wrote: > Dear Mike, > > You CCed me on this patch. Is that because you expect this to solve "my" > problem also? You had some measurements of many oinks vs many perts or > vs "desktop", but not many oinks vs 1 or 2 perts as per my "complaint". > You also changed the subject line, so maybe this is all un-related. I haven't seen the problem you reported. I did stumble upon a problem, but turns out that is only present in master, so yes, un-related. -Mike -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | paul.szabo@sydney.edu.au |
|---|---|
| Date | 2015-10-11 11:30 +0200 |
| Subject | Re: [patch] sched: disable task group re-weighting on the desktop |
| Message-ID | <qil74-49b-13@gated-at.bofh.it> |
| In reply to | #1244050 |
Dear Mike, > ... so yes, un-related. Thanks for clarifying. > I haven't seen the problem you reported. ... You mean you chose not to reproduce: you persisted in pinning your perts, whereas the problem was stated with un-pinned perts (and pinned oinks). But that is OK... others did reproduce, and anyway I believe I have now fixed my problem. (Solution in that "other" email thread.) Cheers, Paul Paul Szabo psz@maths.usyd.edu.au http://www.maths.usyd.edu.au/u/psz/ School of Mathematics and Statistics University of Sydney Australia -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2015-10-11 15:00 +0200 |
| Subject | Re: [patch] sched: disable task group re-weighting on the desktop |
| Message-ID | <qiooi-nn-9@gated-at.bofh.it> |
| In reply to | #1244084 |
On Sun, 2015-10-11 at 20:25 +1100, paul.szabo@sydney.edu.au wrote: > Dear Mike, > > > ... so yes, un-related. > > Thanks for clarifying. > > > I haven't seen the problem you reported. ... > > You mean you chose not to reproduce: you persisted in pinning your > perts.. There was hard data to the contrary in your mailbox as you wrote that. -Mike -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | paul.szabo@sydney.edu.au |
|---|---|
| Date | 2015-10-11 21:50 +0200 |
| Subject | Re: [patch] sched: disable task group re-weighting on the desktop |
| Message-ID | <qiuN4-1iF-13@gated-at.bofh.it> |
| In reply to | #1243896 |
Dear Mike, Did you check whether setting min_- and max_interval e.g. as per https://lkml.org/lkml/2015/10/11/34 would help with your issue (instead of your "horrible gs destroying" patch)? Cheers, Paul Paul Szabo psz@maths.usyd.edu.au http://www.maths.usyd.edu.au/u/psz/ School of Mathematics and Statistics University of Sydney Australia -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2015-10-12 04:00 +0200 |
| Subject | Re: [patch] sched: disable task group re-weighting on the desktop |
| Message-ID | <qiAz7-1bO-3@gated-at.bofh.it> |
| In reply to | #1244200 |
On Mon, 2015-10-12 at 06:46 +1100, paul.szabo@sydney.edu.au wrote: > Dear Mike, > > Did you check whether setting min_- and max_interval e.g. as per > https://lkml.org/lkml/2015/10/11/34 > would help with your issue (instead of your "horrible gs destroying" > patch)? I spent a lot of MY time looking into YOUR problem, only to be accused of actively avoiding reproduction thereof, and now you toss another cute little dart my way. Looking into your problem wasn't a complete waste of my time, as it led me to something that actually looks interesting. Thanks for that, and goodbye. *PLONK* -Mike -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Mike Galbraith <umgwanakikbuti@gmail.com> |
|---|---|
| Date | 2015-10-08 16:30 +0200 |
| Message-ID | <qhkmL-67a-41@gated-at.bofh.it> |
| In reply to | #1242192 |
On Thu, 2015-10-08 at 21:54 +1100, paul.szabo@sydney.edu.au wrote: > Dear Mike, > > > I see a fairness issue ... but one opposite to your complaint. > > Why is that opposite? I think it would be fair for the one pert process > to get 100% CPU, the many oink processes can get everything else. That > one oink is lowly 10% (when others are 100%) is of no consequence. Well, not exactly opposite, only opposite in that the one pert task also receives MORE than it's fair share when unpinned. Two 100$ hogs sharing one CPU should each get 50% of that CPU. The fact that the oink group contains 8 tasks vs 1 for the pert group should be irrelevant, but what that last oinker is getting is 1/9 of a CPU, and there just happen to be 9 runnable tasks total, 1 in group pert, and 8 in group oink. IFF that ratio were to prove to be a constant, AND the oink group were a massively parallel and synchronized compute job on a huge box, that entire compute job would not be slowed down by the factor 2 that a fair distribution would do to it, on say a 1000 core box, it'd be.. utterly dead, because you'd put it out of your misery. vogelweide:~/:[0]# cgexec -g cpu:foo bash vogelweide:~/:[0]# for i in `seq 0 63`; do taskset -c $i cpuhog& done [1] 8025 [2] 8026 ... vogelweide:~/:[130]# cgexec -g cpu:bar bash vogelweide:~/:[130]# taskset -c 63 pert 10 (report every 10 seconds) 2260.91 MHZ CPU perturbation threshold 0.024 usecs. pert/s: 255 >2070.76us: 38 min: 0.05 max:4065.46 avg: 93.83 sum/s: 23946us overhead: 2.39% pert/s: 255 >2070.32us: 37 min: 1.32 max:4039.94 avg: 92.82 sum/s: 23744us overhead: 2.37% pert/s: 253 >2069.85us: 38 min: 0.05 max:4036.44 avg: 94.89 sum/s: 24054us overhead: 2.41% Hm, that's a kinda odd looking number from my 64 core box, but whatever, it's far from fair according to my definition thereof. Poor little oink plus all other cycles not spent in pert's tight loop add up ~24ms/s. > Good to see that you agree on the fairness issue... it MUST be fixed! > CFS might be wrong or wasteful, but never unfair. Weeell, we've disagreed on pretty much everything we've talked about so far, but I can well imagine that what I see in the share update business _could_ be part of your massive compute job woes. -Mike -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | paul.szabo@sydney.edu.au |
|---|---|
| Date | 2015-10-09 00:00 +0200 |
| Message-ID | <qhrog-7CK-43@gated-at.bofh.it> |
| In reply to | #1242419 |
Dear Mike, >>> I see a fairness issue ... but one opposite to your complaint. >> Why is that opposite? ... > > Well, not exactly opposite, only opposite in that the one pert task also > receives MORE than it's fair share when unpinned. Two 100$ hogs sharing > one CPU should each get 50% of that CPU. ... But you are using CGROUPs, grouping all oinks into one group, and the one pert into another: requesting each group to get same total CPU. Since pert has one process only, the most he can get is 100% (not 400%), and it is quite OK for the oinks together to get 700%. > IFF ... massively parallel and synchronized ... You would be making the assumption that you had the machine to yourself: might be the wrong thing to assume. >> Good to see that you agree ... > Weeell, we've disagreed on pretty much everything ... Sorry I disagree: we do agree on the essence. :-) Cheers, Paul Paul Szabo psz@maths.usyd.edu.au http://www.maths.usyd.edu.au/u/psz/ School of Mathematics and Statistics University of Sydney Australia -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web