Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1168570 > unrolled thread

Re: [PATCH v8 2/4] sched: Rewrite runnable load and utilization average tracking

Started byYuyang Du <yuyang.du@intel.com>
First post2015-06-19 09:00 +0200
Last post2015-06-22 08:40 +0200
Articles 4 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH v8 2/4] sched: Rewrite runnable load and utilization  average tracking Yuyang Du <yuyang.du@intel.com> - 2015-06-19 09:00 +0200
    Re: [PATCH v8 2/4] sched: Rewrite runnable load and utilization  average tracking Boqun Feng <boqun.feng@gmail.com> - 2015-06-19 10:00 +0200
      Re: [PATCH v8 2/4] sched: Rewrite runnable load and utilization  average tracking Boqun Feng <boqun.feng@gmail.com> - 2015-06-19 14:30 +0200
        Re: [PATCH v8 2/4] sched: Rewrite runnable load and utilization  average tracking Yuyang Du <yuyang.du@intel.com> - 2015-06-22 08:40 +0200

#1168570 — Re: [PATCH v8 2/4] sched: Rewrite runnable load and utilization average tracking

FromYuyang Du <yuyang.du@intel.com>
Date2015-06-19 09:00 +0200
SubjectRe: [PATCH v8 2/4] sched: Rewrite runnable load and utilization average tracking
Message-ID<pCYro-8ea-13@gated-at.bofh.it>
On Fri, Jun 19, 2015 at 02:00:38PM +0800, Boqun Feng wrote:
> However, update_cfs_rq_load_avg() only updates cfs_rq->avg, the change
> won't be contributed or aggregated to cfs_rq's parent in the
> for_each_leaf_cfs_rq loop, therefore that's actually not a bottom-up
> update.
> 
> To fix this, I think we can add a update_cfs_shares(cfs_rq) after
> update_cfs_rq_load_avg(). Like:
> 
>  	for_each_leaf_cfs_rq(rq, cfs_rq) {
> -		/*
> -		 * Note: We may want to consider periodically releasing
> -		 * rq->lock about these updates so that creating many task
> -		 * groups does not result in continually extending hold time.
> -		 */
> -		__update_blocked_averages_cpu(cfs_rq->tg, rq->cpu);
> +		/* throttled entities do not contribute to load */
> +		if (throttled_hierarchy(cfs_rq))
> +			continue;
> +
> +		update_cfs_rq_load_avg(cfs_rq_clock_task(cfs_rq), cfs_rq);
> +		update_cfs_share(cfs_rq);
>  	}
> 
> However, I think update_cfs_share isn't cheap, because it may do a
> bottom-up update once called. So how about just update the root cfs_rq?
> Like:
> 
> -	/*
> -	 * Iterates the task_group tree in a bottom up fashion, see
> -	 * list_add_leaf_cfs_rq() for details.
> -	 */
> -	for_each_leaf_cfs_rq(rq, cfs_rq) {
> -		/*
> -		 * Note: We may want to consider periodically releasing
> -		 * rq->lock about these updates so that creating many task
> -		 * groups does not result in continually extending hold time.
> -		 */
> -		__update_blocked_averages_cpu(cfs_rq->tg, rq->cpu);
> -	}
> +	update_cfs_rq_load_avg(rq_clock_task(rq), rq->cfs_rq);

Hi Boqun,

Did I get you right:

This rewrite patch does not NEED to aggregate entity's load to cfs_rq,
but rather directly update the cfs_rq's load (both runnable and blocked),
so there is NO NEED to iterate all of the cfs_rqs.

So simply updating the top cfs_rq is already equivalent to the stock.

It is better if we iterate the cfs_rq to update the actually weight
(update_cfs_share), because the weight may have already changed, which
would in turn change the load. But update_cfs_share is not cheap.

Right?

Thanks,
Yuyang
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1168592

FromBoqun Feng <boqun.feng@gmail.com>
Date2015-06-19 10:00 +0200
Message-ID<pCZns-17v-13@gated-at.bofh.it>
In reply to#1168570

[Multipart message — attachments visible in raw view] — view raw

Hi Yuyang,

On Fri, Jun 19, 2015 at 07:05:54AM +0800, Yuyang Du wrote:
> On Fri, Jun 19, 2015 at 02:00:38PM +0800, Boqun Feng wrote:
> > However, update_cfs_rq_load_avg() only updates cfs_rq->avg, the change
> > won't be contributed or aggregated to cfs_rq's parent in the
> > for_each_leaf_cfs_rq loop, therefore that's actually not a bottom-up
> > update.
> > 
> > To fix this, I think we can add a update_cfs_shares(cfs_rq) after
> > update_cfs_rq_load_avg(). Like:
> > 
> >  	for_each_leaf_cfs_rq(rq, cfs_rq) {
> > -		/*
> > -		 * Note: We may want to consider periodically releasing
> > -		 * rq->lock about these updates so that creating many task
> > -		 * groups does not result in continually extending hold time.
> > -		 */
> > -		__update_blocked_averages_cpu(cfs_rq->tg, rq->cpu);
> > +		/* throttled entities do not contribute to load */
> > +		if (throttled_hierarchy(cfs_rq))
> > +			continue;
> > +
> > +		update_cfs_rq_load_avg(cfs_rq_clock_task(cfs_rq), cfs_rq);
> > +		update_cfs_share(cfs_rq);
> >  	}
> > 
> > However, I think update_cfs_share isn't cheap, because it may do a
> > bottom-up update once called. So how about just update the root cfs_rq?
> > Like:
> > 
> > -	/*
> > -	 * Iterates the task_group tree in a bottom up fashion, see
> > -	 * list_add_leaf_cfs_rq() for details.
> > -	 */
> > -	for_each_leaf_cfs_rq(rq, cfs_rq) {
> > -		/*
> > -		 * Note: We may want to consider periodically releasing
> > -		 * rq->lock about these updates so that creating many task
> > -		 * groups does not result in continually extending hold time.
> > -		 */
> > -		__update_blocked_averages_cpu(cfs_rq->tg, rq->cpu);
> > -	}
> > +	update_cfs_rq_load_avg(rq_clock_task(rq), rq->cfs_rq);
> 
> Hi Boqun,
> 
> Did I get you right:
> 
> This rewrite patch does not NEED to aggregate entity's load to cfs_rq,
> but rather directly update the cfs_rq's load (both runnable and blocked),
> so there is NO NEED to iterate all of the cfs_rqs.

Actually, I'm not sure whether we NEED to aggregate or NOT.

> 
> So simply updating the top cfs_rq is already equivalent to the stock.
> 

The stock does have a bottom up update, so simply updating the top
cfs_rq is not equivalent to it. Simply updateing the top cfs_rq is
equivalent to the rewrite patch, because the rewrite patch lacks of the
aggregation.

> It is better if we iterate the cfs_rq to update the actually weight
> (update_cfs_share), because the weight may have already changed, which
> would in turn change the load. But update_cfs_share is not cheap.
> 
> Right?

You get me right for most part ;-)

My points are:

1. We *may not* need to aggregate entity's load to cfs_rq in
update_blocked_averages(), simply updating the top cfs_rq may be just
fine, but I'm not sure, so scheduler experts' insights are needed here.

2. Whether we need to aggregate or not, the update_blocked_averages() in
the rewrite patch could be improved. If we need to aggregate, we have to
add something like update_cfs_shares(). If we don't need, we can just
replace the loop with one update_cfs_rq_load_avg() on root cfs_rq.

I think we'd better to figure out the "may not" part in point 1 first to
get a reasonable implemenation of update_blocked_averages().

Is that clear now?

Thanks and Best Regards,
Boqun

[toc] | [prev] | [next] | [standalone]


#1168737

FromBoqun Feng <boqun.feng@gmail.com>
Date2015-06-19 14:30 +0200
Message-ID<pD3AK-7hF-15@gated-at.bofh.it>
In reply to#1168592

[Multipart message — attachments visible in raw view] — view raw

Hi Yuyang,

On Fri, Jun 19, 2015 at 11:11:16AM +0800, Yuyang Du wrote:
> On Fri, Jun 19, 2015 at 03:57:24PM +0800, Boqun Feng wrote:
> > > 
> > > This rewrite patch does not NEED to aggregate entity's load to cfs_rq,
> > > but rather directly update the cfs_rq's load (both runnable and blocked),
> > > so there is NO NEED to iterate all of the cfs_rqs.
> > 
> > Actually, I'm not sure whether we NEED to aggregate or NOT.
> > 
> > > 
> > > So simply updating the top cfs_rq is already equivalent to the stock.
> > > 
> 
> Ok. By aggregate, the rewrite patch does not need it, because the cfs_rq's
> load is calculated at once with all its runnable and blocked tasks counted,
> assuming the all children's weights are up-to-date, of course. Please refer
> to the changelog to get an idea.
> 
> > 
> > The stock does have a bottom up update, so simply updating the top
> > cfs_rq is not equivalent to it. Simply updateing the top cfs_rq is
> > equivalent to the rewrite patch, because the rewrite patch lacks of the
> > aggregation.
> 
> It is not the rewrite patch "lacks" aggregation, it is needless. The stock
> has to do a bottom-up update and aggregate, because 1) it updates the
> load at an entity granularity, 2) the blocked load is separate.

Yep, you are right, the aggregation is not necessary.

Let me see if I understand you, in the rewrite, when we
update_cfs_rq_load_avg() we need neither to aggregate child's load_avg,
nor to update cfs_rq->load.weight. Because:

1) For the load before cfs_rq->last_update_time, it's already in the
->load_avg, and decay will do the job.
2) For the load from cfs_rq->last_update_time to now, we calculate
with cfs_rq->load.weight, and the weight should be weight at
->last_update_time rather than now.

Right?

> 
> > > It is better if we iterate the cfs_rq to update the actually weight
> > > (update_cfs_share), because the weight may have already changed, which
> > > would in turn change the load. But update_cfs_share is not cheap.
> > > 
> > > Right?
> > 
> > You get me right for most part ;-)
> > 
> > My points are:
> > 
> > 1. We *may not* need to aggregate entity's load to cfs_rq in
> > update_blocked_averages(), simply updating the top cfs_rq may be just
> > fine, but I'm not sure, so scheduler experts' insights are needed here.
>  
> Then I don't need to say anything about this.
> 
> > 2. Whether we need to aggregate or not, the update_blocked_averages() in
> > the rewrite patch could be improved. If we need to aggregate, we have to
> > add something like update_cfs_shares(). If we don't need, we can just
> > replace the loop with one update_cfs_rq_load_avg() on root cfs_rq.
>  
> If update_cfs_shares() is done here, it is good, but probably not necessary
> though. However, we do need to update_tg_load_avg() here, because if cfs_rq's

We may have another problem even we udpate_tg_load_avg(), because after
the loop, for each cfs_rq, ->load.weight is not up-to-date, right? So
next time before we update_cfs_rq_load_avg(), we need to guarantee that
the cfs_rq->load.weight is already updated, right? And IMO, we don't
have that guarantee yet, do we?

> load change, the parent tg's load_avg should change too. I will upload a next
> version soon.
> 
> In addition, an update to the stress + dbench test case:
> 
> I have a Core i7, not a Xeon Nehalem, and I have a patch that may not impact
> the result. Then, the dbench runs at very low CPU utilization ~1%. Boqun said
> this may result from cgroup control, the dbench I/O is low.
> 
> Anyway, I can't reproduce the results, the CPU0's util is 92+%, and other CPUs
> have ~100% util.

Thank you for looking into that problem, and I will test with your new
version of patch ;-)

Thanks,
Boqun

> 
> Thanks,
> Yuyang

[toc] | [prev] | [next] | [standalone]


#1169722

FromYuyang Du <yuyang.du@intel.com>
Date2015-06-22 08:40 +0200
Message-ID<pE3yF-4bg-5@gated-at.bofh.it>
In reply to#1168737
On Fri, Jun 19, 2015 at 08:22:07PM +0800, Boqun Feng wrote:
> > It is not the rewrite patch "lacks" aggregation, it is needless. The stock
> > has to do a bottom-up update and aggregate, because 1) it updates the
> > load at an entity granularity, 2) the blocked load is separate.
> 
> Yep, you are right, the aggregation is not necessary.
> 
> Let me see if I understand you, in the rewrite, when we
> update_cfs_rq_load_avg() we need neither to aggregate child's load_avg,
> nor to update cfs_rq->load.weight. Because:
> 
> 1) For the load before cfs_rq->last_update_time, it's already in the
> ->load_avg, and decay will do the job.
> 2) For the load from cfs_rq->last_update_time to now, we calculate
> with cfs_rq->load.weight, and the weight should be weight at
> ->last_update_time rather than now.
> 
> Right?
 
Yes.

> > If update_cfs_shares() is done here, it is good, but probably not necessary
> > though. However, we do need to update_tg_load_avg() here, because if cfs_rq's
> 
> We may have another problem even we udpate_tg_load_avg(), because after
> the loop, for each cfs_rq, ->load.weight is not up-to-date, right? So
> next time before we update_cfs_rq_load_avg(), we need to guarantee that
> the cfs_rq->load.weight is already updated, right? And IMO, we don't
> have that guarantee yet, do we?

If we update weight, we must update load_avg. But if we update load_avg, we may need
to update weight. Yes, your comment here is valid, but we already update the shares
as needed in the cases when they are "active", update_blocked_averages() is
largely for inactive group entities, so we should be fine here.
 
> > load change, the parent tg's load_avg should change too. I will upload a next
> > version soon.
> > 
> > In addition, an update to the stress + dbench test case:
> > 
> > I have a Core i7, not a Xeon Nehalem, and I have a patch that may not impact
> > the result. Then, the dbench runs at very low CPU utilization ~1%. Boqun said
> > this may result from cgroup control, the dbench I/O is low.
> > 
> > Anyway, I can't reproduce the results, the CPU0's util is 92+%, and other CPUs
> > have ~100% util.
> 
> Thank you for looking into that problem, and I will test with your new
> version of patch ;-)
 
That would be good. I played the dbench "as is", and its output looks pretty fine.

Thanks,
Yuyang
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web