Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1683892 > unrolled thread

[RFC PATCH v1 00/11] Create fast idle path for short idle periods

Started byAubrey Li <aubrey.li@intel.com>
First post2017-07-10 03:50 +0200
Last post2017-07-11 11:10 +0200
Articles 20 on this page of 84 — 10 participants

Back to article view | Back to linux.kernel


Contents

  [RFC PATCH v1 00/11] Create fast idle path for short idle periods Aubrey Li <aubrey.li@intel.com> - 2017-07-10 03:50 +0200
    [RFC PATCH v1 04/11] sched/idle: make the fast idle path for short idle periods Aubrey Li <aubrey.li@intel.com> - 2017-07-10 03:50 +0200
      Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for  short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-11 15:00 +0200
        Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for  short idle periods Frederic Weisbecker <fweisbec@gmail.com> - 2017-07-11 18:40 +0200
          Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for  short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-11 20:20 +0200
            Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for  short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-12 05:30 +0200
              Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for  short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 07:10 +0200
                Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for  short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-12 07:30 +0200
          Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for  short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 14:30 +0200
    [RFC PATCH v1 05/11] cpuidle: update idle statistics before cpuidle governor Aubrey Li <aubrey.li@intel.com> - 2017-07-10 03:50 +0200
    [RFC PATCH v1 08/11] cpuidle: menu: remove reduplicative implementation Aubrey Li <aubrey.li@intel.com> - 2017-07-10 04:00 +0200
    [RFC PATCH v1 07/11] cpuidle: make idle residency update more generic Aubrey Li <aubrey.li@intel.com> - 2017-07-10 04:00 +0200
    [RFC PATCH v1 03/11] cpuidle: introduce cpuidle governor for idle prediction Aubrey Li <aubrey.li@intel.com> - 2017-07-10 04:00 +0200
      Re: [RFC PATCH v1 03/11] cpuidle: introduce cpuidle governor for  idle prediction Peter Zijlstra <peterz@infradead.org> - 2017-07-12 14:20 +0200
    [RFC PATCH v1 09/11] cpuidle: menu: feed cpuidle prediction to menu governor Aubrey Li <aubrey.li@intel.com> - 2017-07-10 04:00 +0200
    Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-10 10:50 +0200
      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Wanpeng Li <kernellwp@gmail.com> - 2017-07-10 11:40 +0200
      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-10 16:00 +0200
      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-10 16:50 +0200
        Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-10 18:50 +0200
          Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-10 19:30 +0200
            Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-11 06:50 +0200
              Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-11 11:50 +0200
                Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Frederic Weisbecker <fweisbec@gmail.com> - 2017-07-11 18:10 +0200
                  Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-11 18:40 +0200
                    Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-11 20:10 +0200
                      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 14:00 +0200
                        Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 18:00 +0200
                          Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 19:50 +0200
                            Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 21:00 +0200
                              Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 21:10 +0200
                      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 14:30 +0200
                        Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 18:00 +0200
                          Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 19:20 +0200
                            Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 20:00 +0200
                              Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 21:00 +0200
                            Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 20:50 +0200
                              Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-13 10:40 +0200
                    Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-12 06:20 +0200
                      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 10:40 +0200
                        Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-12 23:40 +0200
                          Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-13 10:40 +0200
                            Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-13 16:50 +0200
                              Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-13 17:00 +0200
                                Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-13 17:20 +0200
                                  Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-13 20:30 +0200
                                    Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-14 06:00 +0200
                                      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-14 17:40 +0200
                                        Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-14 18:00 +0200
                                          Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-14 18:10 +0200
                                            Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-17 11:30 +0200
                                              Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-17 15:50 +0200
                                        Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-14 18:00 +0200
                                          Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-14 18:10 +0200
                                            Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-14 18:30 +0200
                                              Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-17 21:30 +0200
                                                Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Arjan van de Ven <arjan@linux.intel.com> - 2017-07-17 21:30 +0200
                                                  Re: [RFC PATCH v1 00/11] Create fast idle path for short idle  periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-17 21:50 +0200
                                                    Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Arjan van de Ven <arjan@linux.intel.com> - 2017-07-17 22:00 +0200
                                                      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle  periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-17 22:10 +0200
                                                Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Arjan van de Ven <arjan@linux.intel.com> - 2017-07-17 21:50 +0200
                                                  Re: [RFC PATCH v1 00/11] Create fast idle path for short idle  periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-17 22:00 +0200
                                                    Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Arjan van de Ven <arjan@linux.intel.com> - 2017-07-17 22:00 +0200
                                                  Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-18 05:30 +0200
                                                Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-18 05:20 +0200
                                                  Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-18 06:50 +0200
                                                    Re: [RFC PATCH v1 00/11] Create fast idle path for short idle  periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-18 08:50 +0200
                                                      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-18 09:00 +0200
                                                        Re: [RFC PATCH v1 00/11] Create fast idle path for short idle  periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-18 09:20 +0200
                                                        Re: [RFC PATCH v1 00/11] Create fast idle path for short idle  periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-18 09:30 +0200
                                                      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-18 09:00 +0200
                                        Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Arjan van de Ven <arjan@linux.intel.com> - 2017-07-14 18:00 +0200
                                Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-13 17:30 +0200
                                  Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-14 05:50 +0200
                                    Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-14 06:10 +0200
                                      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-17 15:30 +0200
                                        Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-17 16:00 +0200
                                          Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-17 16:10 +0200
                      Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 14:20 +0200
                  Re: [RFC PATCH v1 00/11] Create fast idle path for short idle  periods Christoph Lameter <cl@linux.com> - 2017-07-11 20:00 +0200
                    Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-12 04:10 +0200
                  Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-12 04:40 +0200
                  Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 20:20 +0200
            Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-11 11:10 +0200

Page 3 of 5 — ← Prev page 1 2 [3] 4 5  Next page →


#1686050

FromAndi Kleen <ak@linux.intel.com>
Date2017-07-12 23:40 +0200
Message-ID<u2xwt-56I-13@gated-at.bofh.it>
In reply to#1685631
On Wed, Jul 12, 2017 at 10:34:10AM +0200, Peter Zijlstra wrote:
> On Wed, Jul 12, 2017 at 12:15:08PM +0800, Li, Aubrey wrote:
> > Okay, the difference is that Mike's patch uses a very simple algorithm to make the decision.
> 
> No, the difference is that we don't end up with duplication of a metric
> ton of code.

What do you mean? There isn't much duplication from the fast path
in Aubrey's patch kit.

It just moves some code around from the cpuidle governor to be generic,
that accounts for the bulk of the changes. It's just moving however,
not adding.

> It uses the normal idle path, it just makes the NOHZ enter fail.

Which is only a small part of the problem.

-Andi

[toc] | [prev] | [next] | [standalone]


#1686381

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-13 10:40 +0200
Message-ID<u2HPb-3f1-5@gated-at.bofh.it>
In reply to#1686050
On Wed, Jul 12, 2017 at 02:32:40PM -0700, Andi Kleen wrote:
> On Wed, Jul 12, 2017 at 10:34:10AM +0200, Peter Zijlstra wrote:
> > On Wed, Jul 12, 2017 at 12:15:08PM +0800, Li, Aubrey wrote:
> > > Okay, the difference is that Mike's patch uses a very simple algorithm to make the decision.
> > 
> > No, the difference is that we don't end up with duplication of a metric
> > ton of code.
> 
> What do you mean? There isn't much duplication from the fast path
> in Aubrey's patch kit.

A whole second idle path is one too many. We're not going to have
duplicate idle paths.

> It just moves some code around from the cpuidle governor to be generic,
> that accounts for the bulk of the changes. It's just moving however,
> not adding.

It wasn't at first glance evident it was a pure move because he does it
over a bunch of patches. Also, that code will not be moved to the
generic code, people are working on alternatives and very much rely on
this being a governor thing.

> > It uses the normal idle path, it just makes the NOHZ enter fail.
> 
> Which is only a small part of the problem.

Given the data so far provided it was by far the biggest problem. If you
want more things changed, you really have to give more data.

[toc] | [prev] | [next] | [standalone]


#1686594

From"Li, Aubrey" <aubrey.li@linux.intel.com>
Date2017-07-13 16:50 +0200
Message-ID<u2NBg-6SY-19@gated-at.bofh.it>
In reply to#1686381
On 2017/7/13 16:36, Peter Zijlstra wrote:
> On Wed, Jul 12, 2017 at 02:32:40PM -0700, Andi Kleen wrote:
> 
>>> It uses the normal idle path, it just makes the NOHZ enter fail.
>>
>> Which is only a small part of the problem.
> 
> Given the data so far provided it was by far the biggest problem. If you
> want more things changed, you really have to give more data.
> 

I have a data between arch_cpu_idle_enter and arch_cpu_idle_exit, this already
excluded HW sleep.

- totally from arch_cpu_idle_enter entry to arch_cpu_idle_exit return costs
  9122ns - 15318ns.
---- In this period(arch idle), rcu_idle_enter costs 1985ns - 2262ns, rcu_idle_exit
     costs 1813ns - 3507ns

Besides RCU, the period includes c-state selection on X86, a few timestamp updates
and a few computations in menu governor. Also, deep HW-cstate latency can be up
to 100+ microseconds, even if the system is very busy, CPU still has chance to enter
deep cstate, which I guess some outburst workloads are not happy with it.

That's my major concern without a fast idle path.

Thanks,
-Aubrey

[toc] | [prev] | [next] | [standalone]


#1686601

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-13 17:00 +0200
Message-ID<u2NKW-6Wd-17@gated-at.bofh.it>
In reply to#1686594
On Thu, Jul 13, 2017 at 10:48:55PM +0800, Li, Aubrey wrote:

> - totally from arch_cpu_idle_enter entry to arch_cpu_idle_exit return costs
>   9122ns - 15318ns.
> ---- In this period(arch idle), rcu_idle_enter costs 1985ns - 2262ns, rcu_idle_exit
>      costs 1813ns - 3507ns
> 
> Besides RCU,

So Paul wants more details on where RCU hurts so we can try to fix.

> the period includes c-state selection on X86, a few timestamp updates
> and a few computations in menu governor. Also, deep HW-cstate latency can be up
> to 100+ microseconds, even if the system is very busy, CPU still has chance to enter
> deep cstate, which I guess some outburst workloads are not happy with it.
> 
> That's my major concern without a fast idle path.

Fixing C-state selection by creating an alternative idle path sounds so
very wrong.

[toc] | [prev] | [next] | [standalone]


#1686620

From"Li, Aubrey" <aubrey.li@linux.intel.com>
Date2017-07-13 17:20 +0200
Message-ID<u2O4j-7hE-27@gated-at.bofh.it>
In reply to#1686601
On 2017/7/13 22:53, Peter Zijlstra wrote:
> On Thu, Jul 13, 2017 at 10:48:55PM +0800, Li, Aubrey wrote:
> 
>> - totally from arch_cpu_idle_enter entry to arch_cpu_idle_exit return costs
>>   9122ns - 15318ns.
>> ---- In this period(arch idle), rcu_idle_enter costs 1985ns - 2262ns, rcu_idle_exit
>>      costs 1813ns - 3507ns
>>
>> Besides RCU,
> 
> So Paul wants more details on where RCU hurts so we can try to fix.
>
If we can call RCU idle enter/exit after tick is really stopped, instead of
call it every idle, I think it's fine. Then we can skip stopping tick if we need
fast idle.
 
>> the period includes c-state selection on X86, a few timestamp updates
>> and a few computations in menu governor. Also, deep HW-cstate latency can be up
>> to 100+ microseconds, even if the system is very busy, CPU still has chance to enter
>> deep cstate, which I guess some outburst workloads are not happy with it.
>>
>> That's my major concern without a fast idle path.
> 
> Fixing C-state selection by creating an alternative idle path sounds so
> very wrong.

This only happens on the arch which has multiple hardware idle cstates, like
Intel's processor. As long as we want to support multiple cstates, we have to
make a selection(with cost of timestamp update and computation). That's fine
in the normal idle path, but if we want a fast idle switch, we can make a 
tradeoff to use a low-latency one directly, that's why I proposed a fast idle
path, so that we don't need to mix fast idle condition judgement in both idle
entry and idle exit path.

Thanks,
-Aubrey

[toc] | [prev] | [next] | [standalone]


#1686834

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-13 20:30 +0200
Message-ID<u2R29-FW-1@gated-at.bofh.it>
In reply to#1686620
On Thu, Jul 13, 2017 at 11:13:28PM +0800, Li, Aubrey wrote:
> On 2017/7/13 22:53, Peter Zijlstra wrote:

> > Fixing C-state selection by creating an alternative idle path sounds so
> > very wrong.
> 
> This only happens on the arch which has multiple hardware idle cstates, like
> Intel's processor. As long as we want to support multiple cstates, we have to
> make a selection(with cost of timestamp update and computation). That's fine
> in the normal idle path, but if we want a fast idle switch, we can make a 
> tradeoff to use a low-latency one directly, that's why I proposed a fast idle
> path, so that we don't need to mix fast idle condition judgement in both idle
> entry and idle exit path.

That doesn't make sense. If you can decide to pick a shallow C state in
any way, you can fix the general selection too.

[toc] | [prev] | [next] | [standalone]


#1687054

From"Li, Aubrey" <aubrey.li@linux.intel.com>
Date2017-07-14 06:00 +0200
Message-ID<u2ZVL-69C-7@gated-at.bofh.it>
In reply to#1686834
On 2017/7/14 2:28, Peter Zijlstra wrote:
> On Thu, Jul 13, 2017 at 11:13:28PM +0800, Li, Aubrey wrote:
>> On 2017/7/13 22:53, Peter Zijlstra wrote:
> 
>>> Fixing C-state selection by creating an alternative idle path sounds so
>>> very wrong.
>>
>> This only happens on the arch which has multiple hardware idle cstates, like
>> Intel's processor. As long as we want to support multiple cstates, we have to
>> make a selection(with cost of timestamp update and computation). That's fine
>> in the normal idle path, but if we want a fast idle switch, we can make a 
>> tradeoff to use a low-latency one directly, that's why I proposed a fast idle
>> path, so that we don't need to mix fast idle condition judgement in both idle
>> entry and idle exit path.
> 
> That doesn't make sense. If you can decide to pick a shallow C state in
> any way, you can fix the general selection too.
> 

Okay, maybe something like the following make sense? Give a hint to
cpuidle_idle_call() to indicate a fast idle.

--------------------------------------------------------
diff --git a/kernel/sched/idle.c b/kernel/sched/idle.c
index ef63adc..3165e99 100644
--- a/kernel/sched/idle.c
+++ b/kernel/sched/idle.c
@@ -152,7 +152,7 @@ static void cpuidle_idle_call(void)
	 */
	rcu_idle_enter();
 
-	if (cpuidle_not_available(drv, dev)) {
+	if (cpuidle_not_available(drv, dev) || this_is_a_fast_idle) {
		default_idle_call();
		goto exit_idle;
	}

[toc] | [prev] | [next] | [standalone]


#1687517

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-14 17:40 +0200
Message-ID<u3aRb-5mz-3@gated-at.bofh.it>
In reply to#1687054
On Fri, Jul 14, 2017 at 11:56:33AM +0800, Li, Aubrey wrote:
> On 2017/7/14 2:28, Peter Zijlstra wrote:
> > On Thu, Jul 13, 2017 at 11:13:28PM +0800, Li, Aubrey wrote:
> >> On 2017/7/13 22:53, Peter Zijlstra wrote:
> > 
> >>> Fixing C-state selection by creating an alternative idle path sounds so
> >>> very wrong.
> >>
> >> This only happens on the arch which has multiple hardware idle cstates, like
> >> Intel's processor. As long as we want to support multiple cstates, we have to
> >> make a selection(with cost of timestamp update and computation). That's fine
> >> in the normal idle path, but if we want a fast idle switch, we can make a 
> >> tradeoff to use a low-latency one directly, that's why I proposed a fast idle
> >> path, so that we don't need to mix fast idle condition judgement in both idle
> >> entry and idle exit path.
> > 
> > That doesn't make sense. If you can decide to pick a shallow C state in
> > any way, you can fix the general selection too.
> > 
> 
> Okay, maybe something like the following make sense? Give a hint to
> cpuidle_idle_call() to indicate a fast idle.
> 
> --------------------------------------------------------
> diff --git a/kernel/sched/idle.c b/kernel/sched/idle.c
> index ef63adc..3165e99 100644
> --- a/kernel/sched/idle.c
> +++ b/kernel/sched/idle.c
> @@ -152,7 +152,7 @@ static void cpuidle_idle_call(void)
> 	 */
> 	rcu_idle_enter();
>  
> -	if (cpuidle_not_available(drv, dev)) {
> +	if (cpuidle_not_available(drv, dev) || this_is_a_fast_idle) {
> 		default_idle_call();
> 		goto exit_idle;
> 	}

No, that's wrong. We want to fix the normal C state selection process to
pick the right C state.

The fast-idle criteria could cut off a whole bunch of available C
states. We need to understand why our current C state pick is wrong and
amend the algorithm to do better. Not just bolt something on the side.

[toc] | [prev] | [next] | [standalone]


#1687523

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-14 18:00 +0200
Message-ID<u3bax-5vu-1@gated-at.bofh.it>
In reply to#1687517
On Fri, Jul 14, 2017 at 08:52:28AM -0700, Arjan van de Ven wrote:
> On 7/14/2017 8:38 AM, Peter Zijlstra wrote:
> > No, that's wrong. We want to fix the normal C state selection process to
> > pick the right C state.
> > 
> > The fast-idle criteria could cut off a whole bunch of available C
> > states. We need to understand why our current C state pick is wrong and
> > amend the algorithm to do better. Not just bolt something on the side.
> 
> I can see a fast path through selection if you know the upper bound of any
> selection is just 1 state.
> 
> But also, how much of this is about "C1 be fast" versus "selecting C1 is slow"

I got the impression its about we need to select C1 for longer. But the
fact that the patches don't in fact answer any of these questions,
they're wrong in principle ;-)

[toc] | [prev] | [next] | [standalone]


#1687543

FromAndi Kleen <ak@linux.intel.com>
Date2017-07-14 18:10 +0200
Message-ID<u3bkf-5Oq-41@gated-at.bofh.it>
In reply to#1687523
On Fri, Jul 14, 2017 at 05:58:53PM +0200, Peter Zijlstra wrote:
> On Fri, Jul 14, 2017 at 08:52:28AM -0700, Arjan van de Ven wrote:
> > On 7/14/2017 8:38 AM, Peter Zijlstra wrote:
> > > No, that's wrong. We want to fix the normal C state selection process to
> > > pick the right C state.
> > > 
> > > The fast-idle criteria could cut off a whole bunch of available C
> > > states. We need to understand why our current C state pick is wrong and
> > > amend the algorithm to do better. Not just bolt something on the side.
> > 
> > I can see a fast path through selection if you know the upper bound of any
> > selection is just 1 state.

fast idle doesn't have an upper bound.

If the prediction exceeds the fast idle threshold any C state can be used.

It's just another state (fast C1), but right now it has an own threshold
which may be different from standard C1.

> > 
> > But also, how much of this is about "C1 be fast" versus "selecting C1 is slow"
> 
> I got the impression its about we need to select C1 for longer. But the
> fact that the patches don't in fact answer any of these questions,
> they're wrong in principle ;-)

For that workload. Tuning idle thresholds is a complex trade off.

-Andi

[toc] | [prev] | [next] | [standalone]


#1688821

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-17 11:30 +0200
Message-ID<u4avN-3bA-27@gated-at.bofh.it>
In reply to#1687543
On Fri, Jul 14, 2017 at 09:03:14AM -0700, Andi Kleen wrote:
> fast idle doesn't have an upper bound.
> 
> If the prediction exceeds the fast idle threshold any C state can be used.
> 
> It's just another state (fast C1), but right now it has an own threshold
> which may be different from standard C1.

Given it uses the same estimate we end up with:

select_c_state(idle_est)
{
	if (idle_est < fast_threshold)
		return C1;

	if (idle_est < C1_threshold)
		return C1;
	if (idle_est < C2_threshold)
		return C2;
	/* ... */

	return C6
}

Now, unless you're mister Turnbull, C2 will never get selected when
fast_threshold > C2_threshold.

Which is wrong. If you want to effectively scale the selection of C1,
why not also change the C2 and further criteria.

[toc] | [prev] | [next] | [standalone]


#1689058

From"Li, Aubrey" <aubrey.li@linux.intel.com>
Date2017-07-17 15:50 +0200
Message-ID<u4ezo-5Kj-29@gated-at.bofh.it>
In reply to#1688821
On 2017/7/17 17:21, Peter Zijlstra wrote:
> On Fri, Jul 14, 2017 at 09:03:14AM -0700, Andi Kleen wrote:
>> fast idle doesn't have an upper bound.
>>
>> If the prediction exceeds the fast idle threshold any C state can be used.
>>
>> It's just another state (fast C1), but right now it has an own threshold
>> which may be different from standard C1.
> 
> Given it uses the same estimate we end up with:
> 
> 
> Now, unless you're mister Turnbull, C2 will never get selected when
> fast_threshold > C2_threshold.
> 
> Which is wrong. If you want to effectively scale the selection of C1,
> why not also change the C2 and further criteria.
> 
That's not our intention, I think. As long as the predicted coming idle
period > threshold, we'll enter normal idle path, you can select any supported
c-states, tick can also be turned off for power saving. Any deferrable stuff
can be invoke as well because we'll sleep long enough.

Thanks,
-Aubrey

[toc] | [prev] | [next] | [standalone]


#1687527

FromAndi Kleen <ak@linux.intel.com>
Date2017-07-14 18:00 +0200
Message-ID<u3bay-5vu-17@gated-at.bofh.it>
In reply to#1687517
> > -	if (cpuidle_not_available(drv, dev)) {
> > +	if (cpuidle_not_available(drv, dev) || this_is_a_fast_idle) {
> > 		default_idle_call();
> > 		goto exit_idle;
> > 	}
> 
> No, that's wrong. We want to fix the normal C state selection process to
> pick the right C state.
> 
> The fast-idle criteria could cut off a whole bunch of available C
> states. We need to understand why our current C state pick is wrong and
> amend the algorithm to do better. Not just bolt something on the side.

Fast idle uses the same predictor as the current C state governor.

The only difference is that it uses a different threshold for C1.
Likely that's the cause. If it was using the same threshold the
decision would be the same.

The thresholds are coming either from the tables in intel idle,
or from ACPI (let's assume the first)

That means either: the intel idle C1 threshold on the system Aubrey
tested on is too high, or the fast idle threshold is too low.

But that would be only true for the workload he tested.
It may well be that it's not that great for another.

The numbers in the standard intel_idle are reasonably tested
with many workloads, so I guess it would be safer to pick that one.
Unless someone wants to revisit these tables.

-Andi

[toc] | [prev] | [next] | [standalone]


#1687535

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-14 18:10 +0200
Message-ID<u3bke-5Oq-9@gated-at.bofh.it>
In reply to#1687527
On Fri, Jul 14, 2017 at 08:53:56AM -0700, Andi Kleen wrote:
> > > -	if (cpuidle_not_available(drv, dev)) {
> > > +	if (cpuidle_not_available(drv, dev) || this_is_a_fast_idle) {
> > > 		default_idle_call();
> > > 		goto exit_idle;
> > > 	}
> > 
> > No, that's wrong. We want to fix the normal C state selection process to
> > pick the right C state.
> > 
> > The fast-idle criteria could cut off a whole bunch of available C
> > states. We need to understand why our current C state pick is wrong and
> > amend the algorithm to do better. Not just bolt something on the side.
> 
> Fast idle uses the same predictor as the current C state governor.
> 
> The only difference is that it uses a different threshold for C1.
> Likely that's the cause. If it was using the same threshold the
> decision would be the same.

Right, so its selecting C1 for longer. That in turn means we could now
never select C2; because the fast-idle threshold is longer than our C2
time.

Which I feel is wrong; because if we're getting C1 wrong, what says
we're then getting the rest right.

> The thresholds are coming either from the tables in intel idle,
> or from ACPI (let's assume the first)
> 
> That means either: the intel idle C1 threshold on the system Aubrey
> tested on is too high, or the fast idle threshold is too low.

Or our predictor is doing it wrong. It could be its over-estimating idle
duration. For example, suppose we have an idle distribution of:

	40% < C1
	60% > C2

And we end up selecting C2. Even though in many of our sleeps we really
wanted C1.

And as said; Daniel has been working on a better predictor -- now he's
probably not used it on the network workload you're looking at, so that
might be something to consider.

[toc] | [prev] | [next] | [standalone]


#1687550

FromAndi Kleen <ak@linux.intel.com>
Date2017-07-14 18:30 +0200
Message-ID<u3bDA-5W4-11@gated-at.bofh.it>
In reply to#1687535
> And as said; Daniel has been working on a better predictor -- now he's
> probably not used it on the network workload you're looking at, so that
> might be something to consider.

Deriving a better idle predictor is a bit orthogonal to fast idle.

It would be a good idea to set the fast idle threshold to be the
same as intel_idle would use on that platform and rerun the benchmarks 

Then the Cx pattern should be mostly identical, and fast idle can be
evaluated on its own benefits.

-Andi

[toc] | [prev] | [next] | [standalone]


#1689384

FromPeter Zijlstra <peterz@infradead.org>
Date2017-07-17 21:30 +0200
Message-ID<u4jSq-Gx-31@gated-at.bofh.it>
In reply to#1687550
On Fri, Jul 14, 2017 at 09:26:19AM -0700, Andi Kleen wrote:
> > And as said; Daniel has been working on a better predictor -- now he's
> > probably not used it on the network workload you're looking at, so that
> > might be something to consider.
> 
> Deriving a better idle predictor is a bit orthogonal to fast idle.

No. If you want a different C state selected we need to fix the current
C state selector. We're not going to tinker.

And the predictor is probably the most fundamental part of the whole C
state selection logic.

Now I think the problem is that the current predictor goes for an
average idle duration. This means that we, on average, get it wrong 50%
of the time. For performance that's bad.

If you want to improve the worst case, we need to consider a cumulative
distribution function, and staying with the Gaussian assumption already
present, that would mean using:

	 1              x - mu
CDF(x) = - [ 1 + erf(-------------) ]
	 2           sigma sqrt(2)

Where, per the normal convention mu is the average and sigma^2 the
variance. See also:

  https://en.wikipedia.org/wiki/Normal_distribution

We then solve CDF(x) = n% to find the x for which we get it wrong n% of
the time (IIRC something like: 'mu - 2sigma' ends up being 5% or so).

This conceptually gets us better exit latency for the cases where we got
it wrong before, and practically pushes down the estimate which gets us
C1 longer.

Of course, this all assumes a Gaussian distribution to begin with, if we
get bimodal (or worse) distributions we can still get it wrong. To fix
that, we'd need to do something better than what we currently have.

[toc] | [prev] | [next] | [standalone]


#1689386

FromArjan van de Ven <arjan@linux.intel.com>
Date2017-07-17 21:30 +0200
Message-ID<u4jSr-Gx-37@gated-at.bofh.it>
In reply to#1689384
On 7/17/2017 12:23 PM, Peter Zijlstra wrote:
> Now I think the problem is that the current predictor goes for an
> average idle duration. This means that we, on average, get it wrong 50%
> of the time. For performance that's bad.

that's not really what it does; it looks at next tick
and then discounts that based on history;
(with different discounts for different order of magnitude)

[toc] | [prev] | [next] | [standalone]


#1689402 — Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods

FromThomas Gleixner <tglx@linutronix.de>
Date2017-07-17 21:50 +0200
SubjectRe: [RFC PATCH v1 00/11] Create fast idle path for short idle periods
Message-ID<u4kbM-Pt-27@gated-at.bofh.it>
In reply to#1689386
On Mon, 17 Jul 2017, Arjan van de Ven wrote:
> On 7/17/2017 12:23 PM, Peter Zijlstra wrote:
> > Now I think the problem is that the current predictor goes for an
> > average idle duration. This means that we, on average, get it wrong 50%
> > of the time. For performance that's bad.
> 
> that's not really what it does; it looks at next tick
> and then discounts that based on history;
> (with different discounts for different order of magnitude)

next tick is the worst thing to look at for interrupt heavy workloads as
the next tick (as computed by the nohz code) can be far away, while the I/O
interrupts come in at a high frequency.

That's where Daniel Lezcanos work of predicting interrupts comes in and
that's the right solution to the problem. The core infrastructure has been
merged, just the idle/cpufreq users are not there yet. All you need to do
is to select CONFIG_IRQ_TIMINGS and use the statistics generated there.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1689421

FromArjan van de Ven <arjan@linux.intel.com>
Date2017-07-17 22:00 +0200
Message-ID<u4kls-SQ-27@gated-at.bofh.it>
In reply to#1689402
On 7/17/2017 12:46 PM, Thomas Gleixner wrote:
> On Mon, 17 Jul 2017, Arjan van de Ven wrote:
>> On 7/17/2017 12:23 PM, Peter Zijlstra wrote:
>>> Now I think the problem is that the current predictor goes for an
>>> average idle duration. This means that we, on average, get it wrong 50%
>>> of the time. For performance that's bad.
>>
>> that's not really what it does; it looks at next tick
>> and then discounts that based on history;
>> (with different discounts for different order of magnitude)
>
> next tick is the worst thing to look at for interrupt heavy workloads as

well it was better than what was there before (without discount and without detecting
repeated patterns)

> the next tick (as computed by the nohz code) can be far away, while the I/O
> interrupts come in at a high frequency.
>
> That's where Daniel Lezcanos work of predicting interrupts comes in and
> that's the right solution to the problem. The core infrastructure has been
> merged, just the idle/cpufreq users are not there yet. All you need to do
> is to select CONFIG_IRQ_TIMINGS and use the statistics generated there.
>

yes ;-)

also note that the predictor does not need to perfect, on most systems C states are
an order of magnitude apart in terms of power/performance/latency so if you get the general
order of magnitude right the predictor is doing its job.

(this is not universally true, but physics of power gating/etc tend to drive to this conclusion;
the cost of implementing an extra state very close to another state means that the HW folks are unlikely
to do the less power saving state of the two to save their cost and testing effort)

[toc] | [prev] | [next] | [standalone]


#1689424 — Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods

FromThomas Gleixner <tglx@linutronix.de>
Date2017-07-17 22:10 +0200
SubjectRe: [RFC PATCH v1 00/11] Create fast idle path for short idle periods
Message-ID<u4kv7-1bl-11@gated-at.bofh.it>
In reply to#1689421
On Mon, 17 Jul 2017, Arjan van de Ven wrote:
> On 7/17/2017 12:46 PM, Thomas Gleixner wrote:
> > That's where Daniel Lezcanos work of predicting interrupts comes in and
> > that's the right solution to the problem. The core infrastructure has been
> > merged, just the idle/cpufreq users are not there yet. All you need to do
> > is to select CONFIG_IRQ_TIMINGS and use the statistics generated there.
> > 
> yes ;-)

:)

> also note that the predictor does not need to perfect, on most systems C
> states are an order of magnitude apart in terms of
> power/performance/latency so if you get the general order of magnitude
> right the predictor is doing its job.

So it would be interesting just to enable the irq timings stuff and compare
the outcome as a first step. That should be reasonably simple to implement
and would give us also some information of how that code behaves on larger
systems.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


Page 3 of 5 — ← Prev page 1 2 [3] 4 5  Next page →

Back to top | Article view | linux.kernel


csiph-web