Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1452908 > unrolled thread

[RFC][PATCH 0/7] cpufreq / sched: cpufreq_update_util() flags and iowait boosting

Started by"Rafael J. Wysocki" <rjw@rjwysocki.net>
First post2016-08-01 01:50 +0200
Last post2016-08-08 15:00 +0200
Articles 20 on this page of 40 — 7 participants

Back to article view | Back to linux.kernel


Contents

  [RFC][PATCH 0/7] cpufreq / sched: cpufreq_update_util() flags and iowait boosting "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 01:50 +0200
    [RFC][PATCH 4/7] cpufreq / sched: Add flags argument to cpufreq_update_util() "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 01:50 +0200
      Re: [RFC][PATCH 4/7] cpufreq / sched: Add flags argument to  cpufreq_update_util() Dominik Brodowski <linux@dominikbrodowski.net> - 2016-08-01 10:00 +0200
        Re: [RFC][PATCH 4/7] cpufreq / sched: Add flags argument to cpufreq_update_util() "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 17:10 +0200
          Re: [RFC][PATCH 4/7] cpufreq / sched: Add flags argument to  cpufreq_update_util() Steve Muckle <steve.muckle@linaro.org> - 2016-08-01 22:30 +0200
            Re: [RFC][PATCH 4/7] cpufreq / sched: Add flags argument to cpufreq_update_util() "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-02 01:50 +0200
              Re: [RFC][PATCH 4/7] cpufreq / sched: Add flags argument to  cpufreq_update_util() Steve Muckle <steve.muckle@linaro.org> - 2016-08-02 04:10 +0200
    [RFC][PATCH 3/7] cpufreq / sched: Check cpu_of(rq) in cpufreq_update_util() "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 01:50 +0200
      Re: [RFC][PATCH 3/7] cpufreq / sched: Check cpu_of(rq) in  cpufreq_update_util() Dominik Brodowski <linux@dominikbrodowski.net> - 2016-08-01 10:00 +0200
        Re: [RFC][PATCH 3/7] cpufreq / sched: Check cpu_of(rq) in cpufreq_update_util() "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 17:10 +0200
        Re: [RFC][PATCH 3/7] cpufreq / sched: Check cpu_of(rq) in  cpufreq_update_util() Steve Muckle <steve.muckle@linaro.org> - 2016-08-01 21:50 +0200
          Re: [RFC][PATCH 3/7] cpufreq / sched: Check cpu_of(rq) in cpufreq_update_util() "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-02 01:50 +0200
    [RFC][PATCH 1/7] cpufreq / sched: Make schedutil access utilization data directly "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 01:50 +0200
      Re: [RFC][PATCH 1/7] cpufreq / sched: Make schedutil access  utilization data directly Steve Muckle <steve.muckle@linaro.org> - 2016-08-01 21:40 +0200
        Re: [RFC][PATCH 1/7] cpufreq / sched: Make schedutil access utilization data directly "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-02 02:10 +0200
          Re: [RFC][PATCH 1/7] cpufreq / sched: Make schedutil access  utilization data directly Juri Lelli <juri.lelli@arm.com> - 2016-08-02 12:40 +0200
            Re: [RFC][PATCH 1/7] cpufreq / sched: Make schedutil access  utilization data directly Steve Muckle <steve.muckle@linaro.org> - 2016-08-02 16:40 +0200
              Re: [RFC][PATCH 1/7] cpufreq / sched: Make schedutil access  utilization data directly Juri Lelli <juri.lelli@arm.com> - 2016-08-02 17:00 +0200
          Re: [RFC][PATCH 1/7] cpufreq / sched: Make schedutil access  utilization data directly Peter Zijlstra <peterz@infradead.org> - 2016-08-08 12:40 +0200
    [RFC][PATCH 2/7] cpufreq / sched: Drop cpufreq_trigger_update() "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 01:50 +0200
    [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 01:50 +0200
      Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait  condition Steve Muckle <steve.muckle@linaro.org> - 2016-08-02 03:30 +0200
        Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition "Rafael J. Wysocki" <rafael@kernel.org> - 2016-08-02 03:50 +0200
          Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait  condition Steve Muckle <steve.muckle@linaro.org> - 2016-08-03 00:30 +0200
            Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition "Rafael J. Wysocki" <rafael@kernel.org> - 2016-08-03 00:50 +0200
              Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait  condition Steve Muckle <steve.muckle@linaro.org> - 2016-08-04 04:30 +0200
                Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-04 23:20 +0200
                  Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait  condition Steve Muckle <steve.muckle@linaro.org> - 2016-08-05 00:10 +0200
                    Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-06 23:20 +0200
    [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 01:50 +0200
      RE: [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core "Doug Smythies" <dsmythies@telus.net> - 2016-08-04 09:00 +0200
        Re: [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-06 23:50 +0200
          RE: [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core "Doug Smythies" <dsmythies@telus.net> - 2016-08-09 19:20 +0200
    [RFC][PATCH 6/7] cpufreq: schedutil: Add iowait boosting "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 01:50 +0200
      Re: [RFC][PATCH 6/7] cpufreq: schedutil: Add iowait boosting Steve Muckle <steve.muckle@linaro.org> - 2016-08-02 03:40 +0200
        Re: [RFC][PATCH 6/7] cpufreq: schedutil: Add iowait boosting "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-03 01:00 +0200
    RE: [RFC][PATCH 0/7] cpufreq / sched: cpufreq_update_util() flags and iowait boosting "Doug Smythies" <dsmythies@telus.net> - 2016-08-01 17:30 +0200
      Re: [RFC][PATCH 0/7] cpufreq / sched: cpufreq_update_util() flags and iowait boosting "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-01 18:30 +0200
        Re: [RFC][PATCH 0/7] cpufreq / sched: cpufreq_update_util() flags  and iowait boosting Peter Zijlstra <peterz@infradead.org> - 2016-08-08 13:10 +0200
          Re: [RFC][PATCH 0/7] cpufreq / sched: cpufreq_update_util() flags and iowait boosting "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-08-08 15:00 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1452914 — [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-08-01 01:50 +0200
Subject[RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition
Message-ID<s18Ex-6NS-17@gated-at.bofh.it>
In reply to#1452908
From: Rafael J. Wysocki <rafael.j.wysocki@intel.com>

Testing indicates that it is possible to improve performace
significantly without increasing energy consumption too much by
teaching cpufreq governors to bump up the CPU performance level if
the in_iowait flag is set for the task in enqueue_task_fair().

For this purpose, define a new cpufreq_update_util() flag
UUF_IO and modify enqueue_task_fair() to pass that flag to
cpufreq_update_util() in the in_iowait case.  That generally
requires cpufreq_update_util() to be called directly from there,
because update_load_avg() is not likely to be invoked in that
case.

Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
---
 include/linux/sched.h |    1 +
 kernel/sched/fair.c   |    8 ++++++++
 2 files changed, 9 insertions(+)

Index: linux-pm/kernel/sched/fair.c
===================================================================
--- linux-pm.orig/kernel/sched/fair.c
+++ linux-pm/kernel/sched/fair.c
@@ -4459,6 +4459,14 @@ enqueue_task_fair(struct rq *rq, struct
 	struct cfs_rq *cfs_rq;
 	struct sched_entity *se = &p->se;
 
+	/*
+	 * If in_iowait is set, it is likely that the loops below will not
+	 * trigger any cpufreq utilization updates, so do it here explicitly
+	 * with the IO flag passed.
+	 */
+	if (p->in_iowait)
+		cpufreq_update_util(rq, UUF_IO);
+
 	for_each_sched_entity(se) {
 		if (se->on_rq)
 			break;
Index: linux-pm/include/linux/sched.h
===================================================================
--- linux-pm.orig/include/linux/sched.h
+++ linux-pm/include/linux/sched.h
@@ -3376,6 +3376,7 @@ static inline unsigned long rlimit_max(u
 }
 
 #define UUF_RT	0x01
+#define UUF_IO	0x02
 
 #ifdef CONFIG_CPU_FREQ
 struct update_util_data {

[toc] | [prev] | [next] | [standalone]


#1453529 — Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition

FromSteve Muckle <steve.muckle@linaro.org>
Date2016-08-02 03:30 +0200
SubjectRe: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition
Message-ID<s1wGW-5Vm-15@gated-at.bofh.it>
In reply to#1452914
On Mon, Aug 01, 2016 at 01:37:23AM +0200, Rafael J. Wysocki wrote:
...
> For this purpose, define a new cpufreq_update_util() flag
> UUF_IO and modify enqueue_task_fair() to pass that flag to
> cpufreq_update_util() in the in_iowait case.  That generally
> requires cpufreq_update_util() to be called directly from there,
> because update_load_avg() is not likely to be invoked in that
> case.

I didn't follow why the cpufreq hook won't likely be called if
in_iowait is set? AFAICS update_load_avg() gets called in the second loop
and calls update_cfs_rq_load_avg (triggers the hook).

> 
> Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
> ---
>  include/linux/sched.h |    1 +
>  kernel/sched/fair.c   |    8 ++++++++
>  2 files changed, 9 insertions(+)
> 
> Index: linux-pm/kernel/sched/fair.c
> ===================================================================
> --- linux-pm.orig/kernel/sched/fair.c
> +++ linux-pm/kernel/sched/fair.c
> @@ -4459,6 +4459,14 @@ enqueue_task_fair(struct rq *rq, struct
>  	struct cfs_rq *cfs_rq;
>  	struct sched_entity *se = &p->se;
>  
> +	/*
> +	 * If in_iowait is set, it is likely that the loops below will not
> +	 * trigger any cpufreq utilization updates, so do it here explicitly
> +	 * with the IO flag passed.
> +	 */
> +	if (p->in_iowait)
> +		cpufreq_update_util(rq, UUF_IO);
> +
>  	for_each_sched_entity(se) {
>  		if (se->on_rq)
>  			break;
> Index: linux-pm/include/linux/sched.h
> ===================================================================
> --- linux-pm.orig/include/linux/sched.h
> +++ linux-pm/include/linux/sched.h
> @@ -3376,6 +3376,7 @@ static inline unsigned long rlimit_max(u
>  }
>  
>  #define UUF_RT	0x01
> +#define UUF_IO	0x02
>  
>  #ifdef CONFIG_CPU_FREQ
>  struct update_util_data {

thanks,
Steve

[toc] | [prev] | [next] | [standalone]


#1453536 — Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-08-02 03:50 +0200
SubjectRe: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition
Message-ID<s1x0e-628-3@gated-at.bofh.it>
In reply to#1453529
On Tue, Aug 2, 2016 at 3:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
> On Mon, Aug 01, 2016 at 01:37:23AM +0200, Rafael J. Wysocki wrote:
> ...
>> For this purpose, define a new cpufreq_update_util() flag
>> UUF_IO and modify enqueue_task_fair() to pass that flag to
>> cpufreq_update_util() in the in_iowait case.  That generally
>> requires cpufreq_update_util() to be called directly from there,
>> because update_load_avg() is not likely to be invoked in that
>> case.
>
> I didn't follow why the cpufreq hook won't likely be called if
> in_iowait is set? AFAICS update_load_avg() gets called in the second loop
> and calls update_cfs_rq_load_avg (triggers the hook).

In practice it turns out that in the majority of cases when in_iowait
is set the second loop will not run.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1455509 — Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition

FromSteve Muckle <steve.muckle@linaro.org>
Date2016-08-03 00:30 +0200
SubjectRe: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition
Message-ID<s1Qmd-2er-13@gated-at.bofh.it>
In reply to#1453536
On Tue, Aug 02, 2016 at 03:37:02AM +0200, Rafael J. Wysocki wrote:
> On Tue, Aug 2, 2016 at 3:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
> > On Mon, Aug 01, 2016 at 01:37:23AM +0200, Rafael J. Wysocki wrote:
> > ...
> >> For this purpose, define a new cpufreq_update_util() flag
> >> UUF_IO and modify enqueue_task_fair() to pass that flag to
> >> cpufreq_update_util() in the in_iowait case.  That generally
> >> requires cpufreq_update_util() to be called directly from there,
> >> because update_load_avg() is not likely to be invoked in that
> >> case.
> >
> > I didn't follow why the cpufreq hook won't likely be called if
> > in_iowait is set? AFAICS update_load_avg() gets called in the second loop
> > and calls update_cfs_rq_load_avg (triggers the hook).
> 
> In practice it turns out that in the majority of cases when in_iowait
> is set the second loop will not run.

My understanding of enqueue_task_fair() is that the first loop walks up
the portion of the sched_entity hierarchy that needs to be enqueued, and
the second loop updates the rest of the hierarchy that was already
enqueued.

Even if the se corresponding to the root cfs_rq needs to be enqueued
(meaning the whole hierarchy is traversed in the first loop and the
second loop does nothing), enqueue_entity() on the root cfs_rq should
result in the cpufreq hook being called, via enqueue_entity() ->
enqueue_entity_load_avg() -> update_cfs_rq_load_avg().

I'll keep looking to see if I've misunderstood something in here.

thanks,
Steve

[toc] | [prev] | [next] | [standalone]


#1455515 — Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2016-08-03 00:50 +0200
SubjectRe: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition
Message-ID<s1QFz-2lk-15@gated-at.bofh.it>
In reply to#1455509
On Wed, Aug 3, 2016 at 12:02 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
> On Tue, Aug 02, 2016 at 03:37:02AM +0200, Rafael J. Wysocki wrote:
>> On Tue, Aug 2, 2016 at 3:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
>> > On Mon, Aug 01, 2016 at 01:37:23AM +0200, Rafael J. Wysocki wrote:
>> > ...
>> >> For this purpose, define a new cpufreq_update_util() flag
>> >> UUF_IO and modify enqueue_task_fair() to pass that flag to
>> >> cpufreq_update_util() in the in_iowait case.  That generally
>> >> requires cpufreq_update_util() to be called directly from there,
>> >> because update_load_avg() is not likely to be invoked in that
>> >> case.
>> >
>> > I didn't follow why the cpufreq hook won't likely be called if
>> > in_iowait is set? AFAICS update_load_avg() gets called in the second loop
>> > and calls update_cfs_rq_load_avg (triggers the hook).
>>
>> In practice it turns out that in the majority of cases when in_iowait
>> is set the second loop will not run.
>
> My understanding of enqueue_task_fair() is that the first loop walks up
> the portion of the sched_entity hierarchy that needs to be enqueued, and
> the second loop updates the rest of the hierarchy that was already
> enqueued.
>
> Even if the se corresponding to the root cfs_rq needs to be enqueued
> (meaning the whole hierarchy is traversed in the first loop and the
> second loop does nothing), enqueue_entity() on the root cfs_rq should
> result in the cpufreq hook being called, via enqueue_entity() ->
> enqueue_entity_load_avg() -> update_cfs_rq_load_avg().

But then it's rather difficult to pass the IO flag to this one, isn't it?

Essentially, the problem is to pass "IO" to cpufreq_update_util() when
p->in_iowait is set.

If you can find a clever way to do it without adding an extra call
site, that's fine by me, but in any case the extra
cpufreq_update_util() invocation should not be too expensive.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1456119 — Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition

FromSteve Muckle <steve.muckle@linaro.org>
Date2016-08-04 04:30 +0200
SubjectRe: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition
Message-ID<s2gA1-2HH-5@gated-at.bofh.it>
In reply to#1455515
On Wed, Aug 03, 2016 at 12:38:20AM +0200, Rafael J. Wysocki wrote:
> On Wed, Aug 3, 2016 at 12:02 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
> > On Tue, Aug 02, 2016 at 03:37:02AM +0200, Rafael J. Wysocki wrote:
> >> On Tue, Aug 2, 2016 at 3:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
> >> > On Mon, Aug 01, 2016 at 01:37:23AM +0200, Rafael J. Wysocki wrote:
> >> > ...
> >> >> For this purpose, define a new cpufreq_update_util() flag
> >> >> UUF_IO and modify enqueue_task_fair() to pass that flag to
> >> >> cpufreq_update_util() in the in_iowait case.  That generally
> >> >> requires cpufreq_update_util() to be called directly from there,
> >> >> because update_load_avg() is not likely to be invoked in that
> >> >> case.
> >> >
> >> > I didn't follow why the cpufreq hook won't likely be called if
> >> > in_iowait is set? AFAICS update_load_avg() gets called in the second loop
> >> > and calls update_cfs_rq_load_avg (triggers the hook).
> >>
> >> In practice it turns out that in the majority of cases when in_iowait
> >> is set the second loop will not run.
> >
> > My understanding of enqueue_task_fair() is that the first loop walks up
> > the portion of the sched_entity hierarchy that needs to be enqueued, and
> > the second loop updates the rest of the hierarchy that was already
> > enqueued.
> >
> > Even if the se corresponding to the root cfs_rq needs to be enqueued
> > (meaning the whole hierarchy is traversed in the first loop and the
> > second loop does nothing), enqueue_entity() on the root cfs_rq should
> > result in the cpufreq hook being called, via enqueue_entity() ->
> > enqueue_entity_load_avg() -> update_cfs_rq_load_avg().
> 
> But then it's rather difficult to pass the IO flag to this one, isn't it?
> 
> Essentially, the problem is to pass "IO" to cpufreq_update_util() when
> p->in_iowait is set.
> 
> If you can find a clever way to do it without adding an extra call
> site, that's fine by me, but in any case the extra
> cpufreq_update_util() invocation should not be too expensive.

I was under the impression that function pointer calls were more
expensive, and in the shared policy case there is a nontrivial amount of
code that is run in schedutil (including taking a spinlock) before we'd
see sugov_should_update_freq() return false and bail.

Agreed that getting knowledge of p->in_iowait down to the existing hook
is not easy. I spent some time fiddling with that. It seemed doable but
somewhat gross due to the required flag passing and modifications
to enqueue_entity, update_load_avg, etc. If it is decided that it is worth
pursuing I can keep working on it and post a draft.

But I also wonder if the hooks are in the best location.  They are
currently deep in the PELT code. This may make sense from a theoretical
standpoint, calling them whenever a root cfs_rq utilization changes, but
it also makes the hooks difficult to correlate (for policy purposes such
as this iowait change) with higher level logical events like a task
wakeup. Or load balance where we probably want to call the hook just
once after a load balance is complete.

This is also an issue for the remote wakeup case where I currently have
another invocation of the hook in check_preempt_curr(), so I can know if
preemption was triggered and skip a remote schedutil update in that case
to avoid a duplicate IPI.

It seems to me worth evaluating if a higher level set of hook locations
could be used. One possibility is higher up in CFS:
- enqueue_task_fair, dequeue_task_fair
- scheduler_tick
- active_load_balance_cpu_stop, load_balance

Though this wouldn't solve my issue with check_preempt_curr. That would
probably require going further up the stack to try_to_wake_up() etc. Not
yet sure what the other hook locations would be at that level.

thanks,
Steve

[toc] | [prev] | [next] | [standalone]


#1456735 — Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-08-04 23:20 +0200
SubjectRe: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition
Message-ID<s2ydA-6EW-33@gated-at.bofh.it>
In reply to#1456119
On Wednesday, August 03, 2016 07:24:18 PM Steve Muckle wrote:
> On Wed, Aug 03, 2016 at 12:38:20AM +0200, Rafael J. Wysocki wrote:
> > On Wed, Aug 3, 2016 at 12:02 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
> > > On Tue, Aug 02, 2016 at 03:37:02AM +0200, Rafael J. Wysocki wrote:
> > >> On Tue, Aug 2, 2016 at 3:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
> > >> > On Mon, Aug 01, 2016 at 01:37:23AM +0200, Rafael J. Wysocki wrote:
> > >> > ...
> > >> >> For this purpose, define a new cpufreq_update_util() flag
> > >> >> UUF_IO and modify enqueue_task_fair() to pass that flag to
> > >> >> cpufreq_update_util() in the in_iowait case.  That generally
> > >> >> requires cpufreq_update_util() to be called directly from there,
> > >> >> because update_load_avg() is not likely to be invoked in that
> > >> >> case.
> > >> >
> > >> > I didn't follow why the cpufreq hook won't likely be called if
> > >> > in_iowait is set? AFAICS update_load_avg() gets called in the second loop
> > >> > and calls update_cfs_rq_load_avg (triggers the hook).
> > >>
> > >> In practice it turns out that in the majority of cases when in_iowait
> > >> is set the second loop will not run.
> > >
> > > My understanding of enqueue_task_fair() is that the first loop walks up
> > > the portion of the sched_entity hierarchy that needs to be enqueued, and
> > > the second loop updates the rest of the hierarchy that was already
> > > enqueued.
> > >
> > > Even if the se corresponding to the root cfs_rq needs to be enqueued
> > > (meaning the whole hierarchy is traversed in the first loop and the
> > > second loop does nothing), enqueue_entity() on the root cfs_rq should
> > > result in the cpufreq hook being called, via enqueue_entity() ->
> > > enqueue_entity_load_avg() -> update_cfs_rq_load_avg().
> > 
> > But then it's rather difficult to pass the IO flag to this one, isn't it?
> > 
> > Essentially, the problem is to pass "IO" to cpufreq_update_util() when
> > p->in_iowait is set.
> > 
> > If you can find a clever way to do it without adding an extra call
> > site, that's fine by me, but in any case the extra
> > cpufreq_update_util() invocation should not be too expensive.
> 
> I was under the impression that function pointer calls were more
> expensive, and in the shared policy case there is a nontrivial amount of
> code that is run in schedutil (including taking a spinlock) before we'd
> see sugov_should_update_freq() return false and bail.

That's correct in principle, but we only do that if p->in_iowait is set,
which is somewhat special anyway and doesn't happen every time for sure.

So while there is overhead theoretically, I'm not even sure if it is measurable.

> Agreed that getting knowledge of p->in_iowait down to the existing hook
> is not easy. I spent some time fiddling with that. It seemed doable but
> somewhat gross due to the required flag passing and modifications
> to enqueue_entity, update_load_avg, etc. If it is decided that it is worth
> pursuing I can keep working on it and post a draft.

Well, that's a Peter's call. :-)

> But I also wonder if the hooks are in the best location.  They are
> currently deep in the PELT code. This may make sense from a theoretical
> standpoint, calling them whenever a root cfs_rq utilization changes, but
> it also makes the hooks difficult to correlate (for policy purposes such
> as this iowait change) with higher level logical events like a task
> wakeup. Or load balance where we probably want to call the hook just
> once after a load balance is complete.

I generally agree.  We still need to ensure that the hools will be invoked
frequently enough, though, even if HZ is 100.

> This is also an issue for the remote wakeup case where I currently have
> another invocation of the hook in check_preempt_curr(), so I can know if
> preemption was triggered and skip a remote schedutil update in that case
> to avoid a duplicate IPI.
> 
> It seems to me worth evaluating if a higher level set of hook locations
> could be used. One possibility is higher up in CFS:
> - enqueue_task_fair, dequeue_task_fair
> - scheduler_tick
> - active_load_balance_cpu_stop, load_balance

Agreed, that's worth checking.

> Though this wouldn't solve my issue with check_preempt_curr. That would
> probably require going further up the stack to try_to_wake_up() etc. Not
> yet sure what the other hook locations would be at that level.

That's probably too far away from the root cfs_rq utilization changes IMO.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1456752 — Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition

FromSteve Muckle <steve.muckle@linaro.org>
Date2016-08-05 00:10 +0200
SubjectRe: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition
Message-ID<s2yZX-7nK-17@gated-at.bofh.it>
In reply to#1456735
On Thu, Aug 04, 2016 at 11:19:00PM +0200, Rafael J. Wysocki wrote:
> On Wednesday, August 03, 2016 07:24:18 PM Steve Muckle wrote:
> > On Wed, Aug 03, 2016 at 12:38:20AM +0200, Rafael J. Wysocki wrote:
> > > On Wed, Aug 3, 2016 at 12:02 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
> > > > On Tue, Aug 02, 2016 at 03:37:02AM +0200, Rafael J. Wysocki wrote:
> > > >> On Tue, Aug 2, 2016 at 3:22 AM, Steve Muckle <steve.muckle@linaro.org> wrote:
> > > >> > On Mon, Aug 01, 2016 at 01:37:23AM +0200, Rafael J. Wysocki wrote:
> > > >> > ...
> > > >> >> For this purpose, define a new cpufreq_update_util() flag
> > > >> >> UUF_IO and modify enqueue_task_fair() to pass that flag to
> > > >> >> cpufreq_update_util() in the in_iowait case.  That generally
> > > >> >> requires cpufreq_update_util() to be called directly from there,
> > > >> >> because update_load_avg() is not likely to be invoked in that
> > > >> >> case.
> > > >> >
> > > >> > I didn't follow why the cpufreq hook won't likely be called if
> > > >> > in_iowait is set? AFAICS update_load_avg() gets called in the second loop
> > > >> > and calls update_cfs_rq_load_avg (triggers the hook).
> > > >>
> > > >> In practice it turns out that in the majority of cases when in_iowait
> > > >> is set the second loop will not run.
> > > >
> > > > My understanding of enqueue_task_fair() is that the first loop walks up
> > > > the portion of the sched_entity hierarchy that needs to be enqueued, and
> > > > the second loop updates the rest of the hierarchy that was already
> > > > enqueued.
> > > >
> > > > Even if the se corresponding to the root cfs_rq needs to be enqueued
> > > > (meaning the whole hierarchy is traversed in the first loop and the
> > > > second loop does nothing), enqueue_entity() on the root cfs_rq should
> > > > result in the cpufreq hook being called, via enqueue_entity() ->
> > > > enqueue_entity_load_avg() -> update_cfs_rq_load_avg().
> > > 
> > > But then it's rather difficult to pass the IO flag to this one, isn't it?
> > > 
> > > Essentially, the problem is to pass "IO" to cpufreq_update_util() when
> > > p->in_iowait is set.
> > > 
> > > If you can find a clever way to do it without adding an extra call
> > > site, that's fine by me, but in any case the extra
> > > cpufreq_update_util() invocation should not be too expensive.
> > 
> > I was under the impression that function pointer calls were more
> > expensive, and in the shared policy case there is a nontrivial amount of
> > code that is run in schedutil (including taking a spinlock) before we'd
> > see sugov_should_update_freq() return false and bail.
> 
> That's correct in principle, but we only do that if p->in_iowait is set,
> which is somewhat special anyway and doesn't happen every time for sure.
> 
> So while there is overhead theoretically, I'm not even sure if it is measurable.

Ok my worry was if there were IO-heavy workloads that would
hammer this path, but I don't know of any specifically or how often this
path can be taken.

> 
> > Agreed that getting knowledge of p->in_iowait down to the existing hook
> > is not easy. I spent some time fiddling with that. It seemed doable but
> > somewhat gross due to the required flag passing and modifications
> > to enqueue_entity, update_load_avg, etc. If it is decided that it is worth
> > pursuing I can keep working on it and post a draft.
> 
> Well, that's a Peter's call. :-)
> 
> > But I also wonder if the hooks are in the best location.  They are
> > currently deep in the PELT code. This may make sense from a theoretical
> > standpoint, calling them whenever a root cfs_rq utilization changes, but
> > it also makes the hooks difficult to correlate (for policy purposes such
> > as this iowait change) with higher level logical events like a task
> > wakeup. Or load balance where we probably want to call the hook just
> > once after a load balance is complete.
> 
> I generally agree.  We still need to ensure that the hools will be invoked
> frequently enough, though, even if HZ is 100.
> 
> > This is also an issue for the remote wakeup case where I currently have
> > another invocation of the hook in check_preempt_curr(), so I can know if
> > preemption was triggered and skip a remote schedutil update in that case
> > to avoid a duplicate IPI.
> > 
> > It seems to me worth evaluating if a higher level set of hook locations
> > could be used. One possibility is higher up in CFS:
> > - enqueue_task_fair, dequeue_task_fair
> > - scheduler_tick
> > - active_load_balance_cpu_stop, load_balance
> 
> Agreed, that's worth checking.
> 
> > Though this wouldn't solve my issue with check_preempt_curr. That would
> > probably require going further up the stack to try_to_wake_up() etc. Not
> > yet sure what the other hook locations would be at that level.
> 
> That's probably too far away from the root cfs_rq utilization changes IMO.

Is your concern that the rate of hook calls would be decreased?

thanks,
Steve

[toc] | [prev] | [next] | [standalone]


#1457366 — Re: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-08-06 23:20 +0200
SubjectRe: [RFC][PATCH 5/7] cpufreq / sched: UUF_IO flag to indicate iowait condition
Message-ID<s3haG-2Dy-15@gated-at.bofh.it>
In reply to#1456752
On Thursday, August 04, 2016 03:09:08 PM Steve Muckle wrote:
> On Thu, Aug 04, 2016 at 11:19:00PM +0200, Rafael J. Wysocki wrote:

[cut]
 
> > > This is also an issue for the remote wakeup case where I currently have
> > > another invocation of the hook in check_preempt_curr(), so I can know if
> > > preemption was triggered and skip a remote schedutil update in that case
> > > to avoid a duplicate IPI.
> > > 
> > > It seems to me worth evaluating if a higher level set of hook locations
> > > could be used. One possibility is higher up in CFS:
> > > - enqueue_task_fair, dequeue_task_fair
> > > - scheduler_tick
> > > - active_load_balance_cpu_stop, load_balance
> > 
> > Agreed, that's worth checking.
> > 
> > > Though this wouldn't solve my issue with check_preempt_curr. That would
> > > probably require going further up the stack to try_to_wake_up() etc. Not
> > > yet sure what the other hook locations would be at that level.
> > 
> > That's probably too far away from the root cfs_rq utilization changes IMO.
> 
> Is your concern that the rate of hook calls would be decreased?

It might be decreased, but also we might end up using utilization values
that wouldn't reflect the current situation (eg. if the hook is called before
update_load_avg(), the util value used by the governor may not be adequate).

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1452915 — [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-08-01 01:50 +0200
Subject[RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core
Message-ID<s18Ex-6NS-11@gated-at.bofh.it>
In reply to#1452908
From: Rafael J. Wysocki <rafael.j.wysocki@intel.com>

The PID-base P-state selection algorithm used by intel_pstate for
Core processors is based on very weak foundations.  Namely, its
decisions are mostly based on the values of the APERF and MPERF
feedback registers and it only estimates the actual utilization to
check if it is not extremely low (in order to avoid getting stuck
in the highest P-state in that case).

Since it generally causes the CPU P-state to ramp up quickly, it
leads to satisfactory performance, but the metric used by it is only
really valid when the CPU changes P-states by itself (ie. in the turbo
range) and if the P-state value set by the driver is treated by the
CPU as the upper limit on turbo P-states selected by it.

As a result, the only case when P-states are reduced by that
algorithm is when the CPU has just come out of idle, but in that
particular case it would have been better to bump up the P-state
instead.  That causes some benchmarks to behave erratically and
attempts to improve the situation lead to excessive energy
consumption, because they make the CPU stay in very high P-states
almost all the time.

Consequently, the only viable way to fix that is to replace the
erroneous algorithm entirely with a better one.

To that end, notice that setting the P-state proportional to the
actual CPU utilization (measured with the help of MPERF and TSC)
generally leads to reasonable behavior, but it does not reflect
the "performance boosting" nature of the current P-state
selection algorithm.  It may be made more similar to that
algorithm, though, by adding iowait boosting to it.

Specifically, if the P-state is bumped up to the maximum after
receiving the UUF_IO flag via cpufreq_update_util(), it will
allow tasks that were previously waiting on I/O to get the full
capacity of the CPU when they are ready to process data again and
that should lead to the desired performance increase overall
without sacrificing too much energy.

For this reason, use the above approach for Core processors in
intel_pstate.

Original-by: Srinivas Pandruvada <srinivas.pandruvada@linux.intel.com>
Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
---
 drivers/cpufreq/intel_pstate.c |   43 +++++++++++++++++++++++++++++++++++++++--
 include/linux/sched.h          |    3 ++
 kernel/sched/sched.h           |    3 --
 3 files changed, 44 insertions(+), 5 deletions(-)

Index: linux-pm/drivers/cpufreq/intel_pstate.c
===================================================================
--- linux-pm.orig/drivers/cpufreq/intel_pstate.c
+++ linux-pm/drivers/cpufreq/intel_pstate.c
@@ -181,6 +181,8 @@ struct _pid {
  * @cpu:		CPU number for this instance data
  * @update_util:	CPUFreq utility callback information
  * @update_util_set:	CPUFreq utility callback is set
+ * @iowait_boost:	iowait-related boost fraction
+ * @last_update:	Time of the last update.
  * @pstate:		Stores P state limits for this CPU
  * @vid:		Stores VID limits for this CPU
  * @pid:		Stores PID parameters for this CPU
@@ -206,6 +208,7 @@ struct cpudata {
 	struct vid_data vid;
 	struct _pid pid;
 
+	u64	last_update;
 	u64	last_sample_time;
 	u64	prev_aperf;
 	u64	prev_mperf;
@@ -216,6 +219,7 @@ struct cpudata {
 	struct acpi_processor_performance acpi_perf_data;
 	bool valid_pss_table;
 #endif
+	unsigned int iowait_boost;
 };
 
 static struct cpudata **all_cpu_data;
@@ -229,6 +233,7 @@ static struct cpudata **all_cpu_data;
  * @p_gain_pct:		PID proportional gain
  * @i_gain_pct:		PID integral gain
  * @d_gain_pct:		PID derivative gain
+ * @boost_iowait:	Whether or not to use iowait boosting.
  *
  * Stores per CPU model static PID configuration data.
  */
@@ -240,6 +245,7 @@ struct pstate_adjust_policy {
 	int p_gain_pct;
 	int d_gain_pct;
 	int i_gain_pct;
+	bool boost_iowait;
 };
 
 /**
@@ -277,6 +283,7 @@ struct cpu_defaults {
 	struct pstate_funcs funcs;
 };
 
+static inline int32_t get_target_pstate_default(struct cpudata *cpu);
 static inline int32_t get_target_pstate_use_performance(struct cpudata *cpu);
 static inline int32_t get_target_pstate_use_cpu_load(struct cpudata *cpu);
 
@@ -1017,6 +1024,7 @@ static struct cpu_defaults core_params =
 		.p_gain_pct = 20,
 		.d_gain_pct = 0,
 		.i_gain_pct = 0,
+		.boost_iowait = true,
 	},
 	.funcs = {
 		.get_max = core_get_max_pstate,
@@ -1025,7 +1033,7 @@ static struct cpu_defaults core_params =
 		.get_turbo = core_get_turbo_pstate,
 		.get_scaling = core_get_scaling,
 		.get_val = core_get_val,
-		.get_target_pstate = get_target_pstate_use_performance,
+		.get_target_pstate = get_target_pstate_default,
 	},
 };
 
@@ -1290,6 +1298,24 @@ static inline int32_t get_target_pstate_
 	return cpu->pstate.current_pstate - pid_calc(&cpu->pid, perf_scaled);
 }
 
+static inline int32_t get_target_pstate_default(struct cpudata *cpu)
+{
+	struct sample *sample = &cpu->sample;
+	int32_t busy_frac;
+	int pstate;
+
+	busy_frac = div_fp(sample->mperf, sample->tsc);
+	sample->busy_scaled = busy_frac * 100;
+
+	if (busy_frac < cpu->iowait_boost)
+		busy_frac = cpu->iowait_boost;
+
+	cpu->iowait_boost >>= 1;
+
+	pstate = cpu->pstate.turbo_pstate;
+	return fp_toint((pstate + (pstate >> 2)) * busy_frac);
+}
+
 static inline void intel_pstate_update_pstate(struct cpudata *cpu, int pstate)
 {
 	int max_perf, min_perf;
@@ -1332,8 +1358,21 @@ static void intel_pstate_update_util(str
 				     unsigned int flags)
 {
 	struct cpudata *cpu = container_of(data, struct cpudata, update_util);
-	u64 delta_ns = time - cpu->sample.time;
+	u64 delta_ns;
+
+	if (pid_params.boost_iowait) {
+		if (flags & UUF_IO) {
+			cpu->iowait_boost = int_tofp(1);
+		} else if (cpu->iowait_boost) {
+			/* Clear iowait_boost if the CPU may have been idle. */
+			delta_ns = time - cpu->last_update;
+			if (delta_ns > TICK_NSEC)
+				cpu->iowait_boost = 0;
+		}
+		cpu->last_update = time;
+	}
 
+	delta_ns = time - cpu->sample.time;
 	if ((s64)delta_ns >= pid_params.sample_rate_ns) {
 		bool sample_taken = intel_pstate_sample(cpu, time);
 

[toc] | [prev] | [next] | [standalone]


#1456194 — RE: [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core

From"Doug Smythies" <dsmythies@telus.net>
Date2016-08-04 09:00 +0200
SubjectRE: [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core
Message-ID<s2kNk-5m7-5@gated-at.bofh.it>
In reply to#1452915
On 2016.08.03 21:19 Doug Smythies wrote:

Re-sending without the previously attached graph.

Hi Rafael,

Hope this feedback and test results help.

On 2016.07.31 16:49 Rafael J. Wysocki wrote:

> The PID-base P-state selection algorithm used by intel_pstate for
> Core processors is based on very weak foundations.

Agree, very very much.

...[cut]...

> Consequently, the only viable way to fix that is to replace the
> erroneous algorithm entirely with a better one.

Agree, very very much.

> To that end, notice that setting the P-state proportional to the
> actual CPU utilization (measured with the help of MPERF and TSC)
> generally leads to reasonable behavior, but it does not reflect
> the "performance boosting" nature of the current P-state
> selection algorithm. It may be made more similar to that
> algorithm, though, by adding iowait boosting to it.

Which is good and does help a lot for the IO case, but issues
remain for the compute case.

...[cut]...

> +static inline int32_t get_target_pstate_default(struct cpudata *cpu)
> +{
> +	struct sample *sample = &cpu->sample;
> +	int32_t busy_frac;
> +	int pstate;
> +
> +	busy_frac = div_fp(sample->mperf, sample->tsc);
> +	sample->busy_scaled = busy_frac * 100;
> +
> +	if (busy_frac < cpu->iowait_boost)
> +		busy_frac = cpu->iowait_boost;
> +
> +	cpu->iowait_boost >>= 1;
> +
> +	pstate = cpu->pstate.turbo_pstate;
> +	return fp_toint((pstate + (pstate >> 2)) * busy_frac);
> +}
> +

The response curve is not normalized on the lower end to the minimum
pstate for the processor, meaning the overall response will vary
between processors as a function of minimum pstate.

The clamping at maximum pstate at about 80% load seems at little high
to me. Work I have done in various attempts to bring back the use of actual load
has always ended up achieving maximum pstate before 80% load for best results.
Even the get_target_pstate_cpu_load people reach the max pstate faster, and they
are more about energy than performance.
What was the criteria for the decision here? Are test results available for review
and/or duplication by others?


Several tests were done with this patch set.
The patch set would not apply to kernel 4.7, but did apply fine to a 4.7+ kernel
(I did as of 7a66ecf) from a few days ago.

Test 1: Phoronix ffmpeg test (less time is better):
Reason: Because it suffers from rotating amongst CPUs in an odd way, challenging for CPU frequency scaling drivers.
This test tends to be an indicator of potential troubles with some games.
Criteria: (Dirk Brandewie): Must match or better acpi_cpufreq - ondemand.
With patch set: 15.8 Seconds average and 24.51 package watts.
Without patch set: 11.61 Seconds average and 27.59 watts.
Conclusion: Significant reduction in performance with proposed patch set.

Tests 2, 3, 4: Phoronix apache, kernel compile, and postmark tests.
Conclusion: All were similar with and without the patch set, with perhaps a slight
improvement in power consumption for the postmark test with the patch set.

Test 5: Random reads within a largish (50 gigabytes) file.
Reason: Because it was a test I used to use with other include or not include IOWAIT work.
Conclusion: no difference with and without the patch set, likely due to domination by 
long seek times (the file is on a normal disk, not an SSD).

Test 6: Sequential read of a largish (50 gigabytes) file.
Reason: Because it was a test I used to use with other include or not include IOWAIT work.
With patch set: 288.38 Seconds; 177.544 MB/Sec; 6.83 Watts.
Without patch set: 292.38 Seconds; 174.99 MB/Sec; 7.08 Watts.
Conclusion: Better performance and better power with the patch set.

Test 7: Compile the kernel 9 times.
Reason: Just because it was a very revealing test during the
"intel_pstate: Increase hold-off time before busyness is scaled"
discussion / thread(s).
Conclusion: no difference with and without the patch set.

Test 8: pipe-test between cores.
Reason: Just because it was so useful during the
"cross core scheduling frequency drop bisected to 0c313cb20732"
discussion / thread(s).
With patch set: 73.166 Sec; 3.6576 usec/loop; 2278.53 Joules.
Without Patch set: 74.444 Sec; 3.7205 usec/loop; 2338.79 Joules.
Conclusion: Slightly better performance and better energy with the patch set.

Test 9: Dougs_specpower simulator (20% load):
Time is fixed, less energy is better.
Reason: During the long
"[intel-pstate driver regression] processor frequency very high even if in idle"
and subsequent https://bugzilla.kernel.org/show_bug.cgi?id=115771
discussion / thread(s), some sort of test was needed to try to mimic what Srinivas
was getting on his fancy SpecPower test platform. So far at least, this test does that.
Only the 20% load case was created, because that was the biggest problem case back then.
With patch set: 4 tests at an average of 7197 Joules per test, relatively high CPU frequencies.
Without the patch set: 4 tests at an average of 5956 Joules per test, relatively low CPU frequencies.
Conclusion: 21% energy regression with the patch set.
Note: Newer processors might do better than my older i7-2600K.

Test 10: measure the frequency response curve, fixed work packet method,
75 hertz work / sleep frequency (all CPU, no IOWAIT):
Reason: To compare to some older data and observe overall.
png graph attached - might get stripped from the distribution lists.
Conclusions: Tends to oscillate, suggesting some sort of damping is needed.
However, any filtering tends to increase the step function load rise time
(see test 11 below, I think there is some wiggle room here).
See also graph which has: with and without patch set; performance mode (for reference);
Philippe Longepe's cpu_load method also with setpoint 40 (for reference); one of my previous
attempts at a load related patch set from quite sometime ago (for reference).

Test 11: Look at the step function load response. From no load to 100% on one CPU (CPU load only, no IO).
While there is a graph, it is not attached:
Conclusion: The step function response is greatly improved (virtually one sample time max).
It would probably be O.K. to slow it down a little with a filter so as to reduce the
tendency to oscillate under periodic load conditions (to a point, at least. A low enough frequency will
always oscillate) (see the graph for test10).
 
Questions:
Is there a migration plan? i.e. will there be an attempt to merge the current cpu_load method
and this method into one method? Then possibly the PID controller could be eliminated.

... Doug

[toc] | [prev] | [next] | [standalone]


#1457377 — Re: [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-08-06 23:50 +0200
SubjectRe: [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core
Message-ID<s3hDI-2QM-15@gated-at.bofh.it>
In reply to#1456194
On Wednesday, August 03, 2016 11:53:23 PM Doug Smythies wrote:
> On 2016.08.03 21:19 Doug Smythies wrote:
> 
> Re-sending without the previously attached graph.
> 
> Hi Rafael,
> 
> Hope this feedback and test results help.
> 
> On 2016.07.31 16:49 Rafael J. Wysocki wrote:
> 
> > The PID-base P-state selection algorithm used by intel_pstate for
> > Core processors is based on very weak foundations.
> 
> Agree, very very much.
> 
> ...[cut]...
> 
> > Consequently, the only viable way to fix that is to replace the
> > erroneous algorithm entirely with a better one.
> 
> Agree, very very much.
> 
> > To that end, notice that setting the P-state proportional to the
> > actual CPU utilization (measured with the help of MPERF and TSC)
> > generally leads to reasonable behavior, but it does not reflect
> > the "performance boosting" nature of the current P-state
> > selection algorithm. It may be made more similar to that
> > algorithm, though, by adding iowait boosting to it.
> 
> Which is good and does help a lot for the IO case, but issues
> remain for the compute case.
> 
> ...[cut]...
> 
> > +static inline int32_t get_target_pstate_default(struct cpudata *cpu)
> > +{
> > +	struct sample *sample = &cpu->sample;
> > +	int32_t busy_frac;
> > +	int pstate;
> > +
> > +	busy_frac = div_fp(sample->mperf, sample->tsc);
> > +	sample->busy_scaled = busy_frac * 100;
> > +
> > +	if (busy_frac < cpu->iowait_boost)
> > +		busy_frac = cpu->iowait_boost;
> > +
> > +	cpu->iowait_boost >>= 1;
> > +
> > +	pstate = cpu->pstate.turbo_pstate;
> > +	return fp_toint((pstate + (pstate >> 2)) * busy_frac);
> > +}
> > +
> 
> The response curve is not normalized on the lower end to the minimum
> pstate for the processor, meaning the overall response will vary
> between processors as a function of minimum pstate.

But that's OK IMO.

Mapping busy_frac = 0 to the minimum P-state would over-provision workloads
with small values of busy_frac.

> The clamping at maximum pstate at about 80% load seems at little high
> to me. Work I have done in various attempts to bring back the use of actual load
> has always ended up achieving maximum pstate before 80% load for best results.
> Even the get_target_pstate_cpu_load people reach the max pstate faster, and they
> are more about energy than performance.
> What was the criteria for the decision here? Are test results available for review
> and/or duplication by others?

This follows the coefficient used by the schedutil governor, but then the
metric is different, so quite possibly a different value may work better here.

We'll test other values before applying this for sure. :-)

> 
> Several tests were done with this patch set.
> The patch set would not apply to kernel 4.7, but did apply fine to a 4.7+ kernel
> (I did as of 7a66ecf) from a few days ago.
> 
> Test 1: Phoronix ffmpeg test (less time is better):
> Reason: Because it suffers from rotating amongst CPUs in an odd way, challenging for CPU frequency scaling drivers.
> This test tends to be an indicator of potential troubles with some games.
> Criteria: (Dirk Brandewie): Must match or better acpi_cpufreq - ondemand.
> With patch set: 15.8 Seconds average and 24.51 package watts.
> Without patch set: 11.61 Seconds average and 27.59 watts.
> Conclusion: Significant reduction in performance with proposed patch set.
> 
> Tests 2, 3, 4: Phoronix apache, kernel compile, and postmark tests.
> Conclusion: All were similar with and without the patch set, with perhaps a slight
> improvement in power consumption for the postmark test with the patch set.
> 
> Test 5: Random reads within a largish (50 gigabytes) file.
> Reason: Because it was a test I used to use with other include or not include IOWAIT work.
> Conclusion: no difference with and without the patch set, likely due to domination by 
> long seek times (the file is on a normal disk, not an SSD).
> 
> Test 6: Sequential read of a largish (50 gigabytes) file.
> Reason: Because it was a test I used to use with other include or not include IOWAIT work.
> With patch set: 288.38 Seconds; 177.544 MB/Sec; 6.83 Watts.
> Without patch set: 292.38 Seconds; 174.99 MB/Sec; 7.08 Watts.
> Conclusion: Better performance and better power with the patch set.
> 
> Test 7: Compile the kernel 9 times.
> Reason: Just because it was a very revealing test during the
> "intel_pstate: Increase hold-off time before busyness is scaled"
> discussion / thread(s).
> Conclusion: no difference with and without the patch set.
> 
> Test 8: pipe-test between cores.
> Reason: Just because it was so useful during the
> "cross core scheduling frequency drop bisected to 0c313cb20732"
> discussion / thread(s).
> With patch set: 73.166 Sec; 3.6576 usec/loop; 2278.53 Joules.
> Without Patch set: 74.444 Sec; 3.7205 usec/loop; 2338.79 Joules.
> Conclusion: Slightly better performance and better energy with the patch set.
> 
> Test 9: Dougs_specpower simulator (20% load):
> Time is fixed, less energy is better.
> Reason: During the long
> "[intel-pstate driver regression] processor frequency very high even if in idle"
> and subsequent https://bugzilla.kernel.org/show_bug.cgi?id=115771
> discussion / thread(s), some sort of test was needed to try to mimic what Srinivas
> was getting on his fancy SpecPower test platform. So far at least, this test does that.
> Only the 20% load case was created, because that was the biggest problem case back then.
> With patch set: 4 tests at an average of 7197 Joules per test, relatively high CPU frequencies.
> Without the patch set: 4 tests at an average of 5956 Joules per test, relatively low CPU frequencies.
> Conclusion: 21% energy regression with the patch set.
> Note: Newer processors might do better than my older i7-2600K.
> 
> Test 10: measure the frequency response curve, fixed work packet method,
> 75 hertz work / sleep frequency (all CPU, no IOWAIT):
> Reason: To compare to some older data and observe overall.
> png graph attached - might get stripped from the distribution lists.
> Conclusions: Tends to oscillate, suggesting some sort of damping is needed.
> However, any filtering tends to increase the step function load rise time
> (see test 11 below, I think there is some wiggle room here).
> See also graph which has: with and without patch set; performance mode (for reference);
> Philippe Longepe's cpu_load method also with setpoint 40 (for reference); one of my previous
> attempts at a load related patch set from quite sometime ago (for reference).
> 
> Test 11: Look at the step function load response. From no load to 100% on one CPU (CPU load only, no IO).
> While there is a graph, it is not attached:
> Conclusion: The step function response is greatly improved (virtually one sample time max).
> It would probably be O.K. to slow it down a little with a filter so as to reduce the
> tendency to oscillate under periodic load conditions (to a point, at least. A low enough frequency will
> always oscillate) (see the graph for test10).

All of the above is useful information, thanks for taking the time to do that
work!

> Questions:
> Is there a migration plan?

Not yet.  We have quite a lot of testing to do first.

> i.e. will there be an attempt to merge the current cpu_load method
> and this method into one method?

Quite possibly if the results are good enough.

> Then possibly the PID controller could be eliminated.

Right.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1459007 — RE: [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core

From"Doug Smythies" <dsmythies@telus.net>
Date2016-08-09 19:20 +0200
SubjectRE: [RFC][PATCH 7/7] cpufreq: intel_pstate: Change P-state selection algorithm for Core
Message-ID<s4iR4-24N-25@gated-at.bofh.it>
In reply to#1457377
On 2016.08.05 17:02 Rafael J. Wysocki wrote:
> On 2016.08.03 21:19 Doug Smythies wrote:
>> On 2016.07.31 16:49 Rafael J. Wysocki wrote:
>>
>> 
>>> +static inline int32_t get_target_pstate_default(struct cpudata *cpu)
>>> +{
>>> +	struct sample *sample = &cpu->sample;
>>> +	int32_t busy_frac;
>>> +	int pstate;
>>> +
>>> +	busy_frac = div_fp(sample->mperf, sample->tsc);
>>> +	sample->busy_scaled = busy_frac * 100;
>>> +
>>> +	if (busy_frac < cpu->iowait_boost)
>>> +		busy_frac = cpu->iowait_boost;
>>> +
>>> +	cpu->iowait_boost >>= 1;
>>> +
>>> +	pstate = cpu->pstate.turbo_pstate;
>>> +	return fp_toint((pstate + (pstate >> 2)) * busy_frac);
>>> +}
>>> +
>> 
>> The response curve is not normalized on the lower end to the minimum
>> pstate for the processor, meaning the overall response will vary
>> between processors as a function of minimum pstate.

> But that's OK IMO.
>
> Mapping busy_frac = 0 to the minimum P-state would over-provision workloads
> with small values of busy_frac.

Agreed, mapping busy_frac = 0 to the minimum Pstate would be a bad thing to do.

However, that is not what I meant. I meant that the mapping of busy-frac = N to
the minimum pstate for the processor should give the same "N" (within granularity),
regardless of the processor.

Example, my processor, i7-2600K: max pstate = 38; min pstate = 16.
Load before going to pstate, 17:  17 = (38 + 38/4) * load
Load = N = 35.8 %

Example, something like, i5-3337U (I think, I don't actually have one):
max pstate = 27; min pstate = 8.
Load before going to pstate, 9: 9 = (27 + 27/4) * load
Load =  N = 26.7 %

It was a couple of years ago, so I should re-do the sensitivity
analysis/testing, but I concluded that the performance / energy tradeoff
was somewhat sensitive to "N".
 
I am suggesting that the response curve, or transfer function,
needs to be normalized, for any processor, to:

  Max pstate |              __________
             |             /
             |            /
             |           /
             |          /
             |         /
   Min pstate| _______/
             |__________________________
              |       |     |          |
              0%      N     M         100%
                      CPU load

Currently M ~= 80%

One time (not re-based since kernel 4.3) I did have a proposed solution [1],
but it was expensive in terms of extra multiplies and divides.

[1]: http://marc.info/?l=linux-pm&m=142881187323474&w=2

>> The clamping at maximum pstate at about 80% load seems at little high
>> to me. Work I have done in various attempts to bring back the use of actual load
>> has always ended up achieving maximum pstate before 80% load for best results.
>> Even the get_target_pstate_cpu_load people reach the max pstate faster, and they
>> are more about energy than performance.
>> What was the criteria for the decision here? Are test results available for review
>> and/or duplication by others?

> This follows the coefficient used by the schedutil governor, but then the
> metric is different, so quite possibly a different value may work better here.
> 
> We'll test other values before applying this for sure. :-)

I am now testing this change to the code (for M ~= 67%; N ~= 30% (my CPU); N ~= 22% (i5-3337U)):

diff --git a/drivers/cpufreq/intel_pstate.c b/drivers/cpufreq/intel_pstate.c
index 8b2bdb7..909d441 100644
--- a/drivers/cpufreq/intel_pstate.c
+++ b/drivers/cpufreq/intel_pstate.c
@@ -1313,7 +1313,7 @@ static inline int32_t get_target_pstate_default(struct cpudata *cpu)
        cpu->iowait_boost >>= 1;

        pstate = cpu->pstate.turbo_pstate;
-       return fp_toint((pstate + (pstate >> 2)) * busy_frac);
+       return fp_toint((pstate + (pstate >> 1)) * busy_frac);
 }

 static inline void intel_pstate_update_pstate(struct cpudata *cpu, int pstate)

> 
> Several tests were done with this patch set.
...[cut]...

>> Questions:
>> Is there a migration plan?
>
> Not yet.  We have quite a lot of testing to do first.
>
>> i.e. will there be an attempt to merge the current cpu_load method
>> and this method into one method?
>
> Quite possibly if the results are good enough.
>
>> Then possibly the PID controller could be eliminated.
>
> Right.

I think this change is important, and I'll help with it as best as I can.

... Doug

A related CPU frequency Vs. Load graph will be sent to Rafael and Srinivas off-list.

[toc] | [prev] | [next] | [standalone]


#1452916 — [RFC][PATCH 6/7] cpufreq: schedutil: Add iowait boosting

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-08-01 01:50 +0200
Subject[RFC][PATCH 6/7] cpufreq: schedutil: Add iowait boosting
Message-ID<s18Ex-6NS-19@gated-at.bofh.it>
In reply to#1452908
From: Rafael J. Wysocki <rafael.j.wysocki@intel.com>

Modify the schedutil cpufreq governor to boost the CPU frequency
if the UUF_IO flag is passed to it via cpufreq_update_util().

If that happens, the frequency is set to the maximum during
the first update after receiving the UUF_IO flag and then the
boost is reduced by half during each following update.

Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
---
 kernel/sched/cpufreq_schedutil.c |   61 ++++++++++++++++++++++++++++++++++-----
 1 file changed, 54 insertions(+), 7 deletions(-)

Index: linux-pm/kernel/sched/cpufreq_schedutil.c
===================================================================
--- linux-pm.orig/kernel/sched/cpufreq_schedutil.c
+++ linux-pm/kernel/sched/cpufreq_schedutil.c
@@ -48,11 +48,13 @@ struct sugov_cpu {
 	struct sugov_policy *sg_policy;
 
 	unsigned int cached_raw_freq;
+	unsigned long iowait_boost;
+	unsigned long iowait_boost_max;
+	u64 last_update;
 
 	/* The fields below are only needed when sharing a policy. */
 	unsigned long util;
 	unsigned long max;
-	u64 last_update;
 	unsigned int flags;
 };
 
@@ -172,22 +174,58 @@ static void sugov_get_util(unsigned long
 	}
 }
 
+static void sugov_set_iowait_boost(struct sugov_cpu *sg_cpu, u64 time,
+				   unsigned int flags)
+{
+	if (flags & UUF_IO) {
+		sg_cpu->iowait_boost = sg_cpu->iowait_boost_max;
+	} else if (sg_cpu->iowait_boost) {
+		s64 delta_ns = time - sg_cpu->last_update;
+
+		/* Clear iowait_boost if the CPU apprears to have been idle. */
+		if (delta_ns > TICK_NSEC)
+			sg_cpu->iowait_boost = 0;
+	}
+}
+
+static void sugov_iowait_boost(struct sugov_cpu *sg_cpu, unsigned long *util,
+			       unsigned long *max)
+{
+	unsigned long boost_util = sg_cpu->iowait_boost;
+	unsigned long boost_max = sg_cpu->iowait_boost_max;
+
+	if (!boost_util)
+		return;
+
+	if (*util * boost_max < *max * boost_util) {
+		*util = boost_util;
+		*max = boost_max;
+	}
+	sg_cpu->iowait_boost >>= 1;
+}
+
 static void sugov_update_single(struct update_util_data *hook, u64 time,
 				unsigned int flags)
 {
 	struct sugov_cpu *sg_cpu = container_of(hook, struct sugov_cpu, update_util);
 	struct sugov_policy *sg_policy = sg_cpu->sg_policy;
-	struct cpufreq_policy *policy = sg_policy->policy;
 	unsigned long util, max;
 	unsigned int next_f;
 
+	sugov_set_iowait_boost(sg_cpu, time, flags);
+	sg_cpu->last_update = time;
+
 	if (!sugov_should_update_freq(sg_policy, time))
 		return;
 
 	sugov_get_util(&util, &max, flags);
 
-	next_f = flags & UUF_RT ? policy->cpuinfo.max_freq :
-				  get_next_freq(sg_cpu, util, max);
+	if (flags & UUF_RT) {
+		next_f = sg_policy->policy->cpuinfo.max_freq;
+	} else {
+		sugov_iowait_boost(sg_cpu, &util, &max);
+		next_f = get_next_freq(sg_cpu, util, max);
+	}
 	sugov_update_commit(sg_policy, time, next_f);
 }
 
@@ -204,6 +242,8 @@ static unsigned int sugov_next_freq_shar
 	if (flags & UUF_RT)
 		return max_f;
 
+	sugov_iowait_boost(sg_cpu, &util, &max);
+
 	for_each_cpu(j, policy->cpus) {
 		struct sugov_cpu *j_sg_cpu;
 		unsigned long j_util, j_max;
@@ -218,12 +258,13 @@ static unsigned int sugov_next_freq_shar
 		 * frequency update and the time elapsed between the last update
 		 * of the CPU utilization and the last frequency update is long
 		 * enough, don't take the CPU into account as it probably is
-		 * idle now.
+		 * idle now (and clear iowait_boost for it).
 		 */
 		delta_ns = last_freq_update_time - j_sg_cpu->last_update;
-		if (delta_ns > TICK_NSEC)
+		if (delta_ns > TICK_NSEC) {
+			j_sg_cpu->iowait_boost = 0;
 			continue;
-
+		}
 		if (j_sg_cpu->flags & UUF_RT)
 			return max_f;
 
@@ -233,6 +274,8 @@ static unsigned int sugov_next_freq_shar
 			util = j_util;
 			max = j_max;
 		}
+
+		sugov_iowait_boost(j_sg_cpu, &util, &max);
 	}
 
 	return get_next_freq(sg_cpu, util, max);
@@ -253,6 +296,8 @@ static void sugov_update_shared(struct u
 	sg_cpu->util = util;
 	sg_cpu->max = max;
 	sg_cpu->flags = flags;
+
+	sugov_set_iowait_boost(sg_cpu, time, flags);
 	sg_cpu->last_update = time;
 
 	if (sugov_should_update_freq(sg_policy, time)) {
@@ -485,6 +530,8 @@ static int sugov_start(struct cpufreq_po
 			sg_cpu->flags = UUF_RT;
 			sg_cpu->last_update = 0;
 			sg_cpu->cached_raw_freq = 0;
+			sg_cpu->iowait_boost = 0;
+			sg_cpu->iowait_boost_max = policy->cpuinfo.max_freq;
 			cpufreq_add_update_util_hook(cpu, &sg_cpu->update_util,
 						     sugov_update_shared);
 		} else {

[toc] | [prev] | [next] | [standalone]


#1453533 — Re: [RFC][PATCH 6/7] cpufreq: schedutil: Add iowait boosting

FromSteve Muckle <steve.muckle@linaro.org>
Date2016-08-02 03:40 +0200
SubjectRe: [RFC][PATCH 6/7] cpufreq: schedutil: Add iowait boosting
Message-ID<s1wQx-5YM-9@gated-at.bofh.it>
In reply to#1452916
On Mon, Aug 01, 2016 at 01:37:59AM +0200, Rafael J. Wysocki wrote:
> From: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
> 
> Modify the schedutil cpufreq governor to boost the CPU frequency
> if the UUF_IO flag is passed to it via cpufreq_update_util().
> 
> If that happens, the frequency is set to the maximum during
> the first update after receiving the UUF_IO flag and then the
> boost is reduced by half during each following update.

Were these changes to schedutil part of the positive test results
mentioned in patch 5? Or are those just from intel pstate?

I was nervous about the effect of this on power and tested a couple low
power usecases. The platform is the Hikey 96board (8 core ARM A53,
single CPUfreq domain) running AOSP Android and schedutil backported to
kernel 4.4. These tests run mp3 and mpeg4 playback for a little while,
recording total energy consumption during the test along with frequency
residency.

As the results below show I did not measure an appreciable effect - if
anything things may be slightly better with the patches.

The hardcoding of a non-tunable boosting scheme makes me nervous but
perhaps it could be revisited if some platform or configuration shows
a noticeable regression?

Testcase	Energy	/----- CPU frequency residency -----\
		(J)	208000	432000	729000	960000	1200000
mp3-before-1	26.822	47.27%	24.79%	16.23%	5.20%	6.52%
mp3-before-2	26.817	41.70%	28.75%	17.62%	5.17%	6.75%
mp3-before-3	26.65	42.48%	28.65%	17.25%	5.07%	6.55%
mp3-after-1	26.667	42.51%	27.38%	18.00%	5.40%	6.71%
mp3-after-2	26.777	48.37%	24.15%	15.68%	4.55%	7.25%
mp3-after-3	26.806	41.93%	27.71%	18.35%	4.78%	7.35%

mpeg4-before-1	26.024	18.41%	60.09%	13.16%	0.49%	7.85%
mpeg4-before-2	25.147	20.47%	64.80%	8.44%	1.37%	4.91%
mpeg4-before-3	25.007	19.18%	66.08%	10.01%	0.59%	4.22%
mpeg4-after-1	25.598	19.77%	61.33%	11.63%	0.79%	6.48%
mpeg4-after-2	25.18	22.31%	62.78%	8.83%	1.18%	4.90%
mpeg4-after-3	25.162	21.59%	64.88%	8.29%	0.49%	4.71%

thanks,
Steve

[toc] | [prev] | [next] | [standalone]


#1455522 — Re: [RFC][PATCH 6/7] cpufreq: schedutil: Add iowait boosting

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-08-03 01:00 +0200
SubjectRe: [RFC][PATCH 6/7] cpufreq: schedutil: Add iowait boosting
Message-ID<s1QPf-2oO-1@gated-at.bofh.it>
In reply to#1453533
On Monday, August 01, 2016 06:35:31 PM Steve Muckle wrote:
> On Mon, Aug 01, 2016 at 01:37:59AM +0200, Rafael J. Wysocki wrote:
> > From: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
> > 
> > Modify the schedutil cpufreq governor to boost the CPU frequency
> > if the UUF_IO flag is passed to it via cpufreq_update_util().
> > 
> > If that happens, the frequency is set to the maximum during
> > the first update after receiving the UUF_IO flag and then the
> > boost is reduced by half during each following update.
> 
> Were these changes to schedutil part of the positive test results
> mentioned in patch 5? Or are those just from intel pstate?
> 
> I was nervous about the effect of this on power and tested a couple low
> power usecases. The platform is the Hikey 96board (8 core ARM A53,
> single CPUfreq domain) running AOSP Android and schedutil backported to
> kernel 4.4. These tests run mp3 and mpeg4 playback for a little while,
> recording total energy consumption during the test along with frequency
> residency.
> 
> As the results below show I did not measure an appreciable effect - if
> anything things may be slightly better with the patches.
> 
> The hardcoding of a non-tunable boosting scheme makes me nervous but
> perhaps it could be revisited if some platform or configuration shows
> a noticeable regression?

That would be my approach. :-)

I'm not a big fan of tunables in general, as there are only a few people
who actually set them to anything different from the default and then they
get a lot of focus (even though they are after super-corner cases sometimes).

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1453271

From"Doug Smythies" <dsmythies@telus.net>
Date2016-08-01 17:30 +0200
Message-ID<s1nke-834-9@gated-at.bofh.it>
In reply to#1452908
On 2016.07.31 16:32 Rafael J. Wysocki wrote:

> Hi,
>
> Admittedly, this hasn't been tested yet, so no promises and you have been
> warned.  It builds, though (on x86-64 at least).

It would not build for me until I changed the kernel configuration file like so:

$ scripts/diffconfig .config .config_rjw_b_2
 CPU_FREQ_GOV_SCHEDUTIL m -> y

(the above is actually backwards, it is "y" that works.)

Otherwise I got:

Kernel: arch/x86/boot/bzImage is ready  (#86)
  MODPOST 3318 modules
ERROR: "dl_bw_cpus" [kernel/sched/cpufreq_schedutil.ko] undefined!
ERROR: "dl_bw_of" [kernel/sched/cpufreq_schedutil.ko] undefined!
ERROR: "runqueues" [kernel/sched/cpufreq_schedutil.ko] undefined!
scripts/Makefile.modpost:91: recipe for target '__modpost' failed
make[4]: *** [__modpost] Error 1
Makefile:1186: recipe for target 'modules' failed

... Doug

[toc] | [prev] | [next] | [standalone]


#1453302

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-08-01 18:30 +0200
Message-ID<s1ogi-eO-17@gated-at.bofh.it>
In reply to#1453271
On Monday, August 01, 2016 08:26:54 AM Doug Smythies wrote:
> On 2016.07.31 16:32 Rafael J. Wysocki wrote:
> 
> > Hi,
> >
> > Admittedly, this hasn't been tested yet, so no promises and you have been
> > warned.  It builds, though (on x86-64 at least).
> 
> It would not build for me until I changed the kernel configuration file like so:
> 
> $ scripts/diffconfig .config .config_rjw_b_2
>  CPU_FREQ_GOV_SCHEDUTIL m -> y
> 
> (the above is actually backwards, it is "y" that works.)
> 
> Otherwise I got:
> 
> Kernel: arch/x86/boot/bzImage is ready  (#86)
>   MODPOST 3318 modules
> ERROR: "dl_bw_cpus" [kernel/sched/cpufreq_schedutil.ko] undefined!
> ERROR: "dl_bw_of" [kernel/sched/cpufreq_schedutil.ko] undefined!
> ERROR: "runqueues" [kernel/sched/cpufreq_schedutil.ko] undefined!
> scripts/Makefile.modpost:91: recipe for target '__modpost' failed
> make[4]: *** [__modpost] Error 1
> Makefile:1186: recipe for target 'modules' failed

There are some export_symbol declarations missing, I'll fix that up.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1457704 — Re: [RFC][PATCH 0/7] cpufreq / sched: cpufreq_update_util() flags and iowait boosting

FromPeter Zijlstra <peterz@infradead.org>
Date2016-08-08 13:10 +0200
SubjectRe: [RFC][PATCH 0/7] cpufreq / sched: cpufreq_update_util() flags and iowait boosting
Message-ID<s3QBs-A4-21@gated-at.bofh.it>
In reply to#1453302
On Mon, Aug 01, 2016 at 06:30:27PM +0200, Rafael J. Wysocki wrote:
> >  CPU_FREQ_GOV_SCHEDUTIL m -> y

> > ERROR: "runqueues" [kernel/sched/cpufreq_schedutil.ko] undefined!

> There are some export_symbol declarations missing, I'll fix that up.

Do you really need that thing to be a module? I would raelly like to not
export runqueues. People have had 'crazy' ideas in the past.

[toc] | [prev] | [next] | [standalone]


#1457764

From"Rafael J. Wysocki" <rjw@rjwysocki.net>
Date2016-08-08 15:00 +0200
Message-ID<s3SjT-1vg-15@gated-at.bofh.it>
In reply to#1457704
On Monday, August 08, 2016 01:08:22 PM Peter Zijlstra wrote:
> On Mon, Aug 01, 2016 at 06:30:27PM +0200, Rafael J. Wysocki wrote:
> > >  CPU_FREQ_GOV_SCHEDUTIL m -> y
> 
> > > ERROR: "runqueues" [kernel/sched/cpufreq_schedutil.ko] undefined!
> 
> > There are some export_symbol declarations missing, I'll fix that up.
> 
> Do you really need that thing to be a module? I would raelly like to not
> export runqueues. People have had 'crazy' ideas in the past.

It doesn't have to be a module, but it is today.

Well, I guess accessing scheduler data directly is a good enough resone for
making it non-modular. :-)

Thanks,
Rafael

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web