Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1683892 > unrolled thread
| Started by | Aubrey Li <aubrey.li@intel.com> |
|---|---|
| First post | 2017-07-10 03:50 +0200 |
| Last post | 2017-07-11 11:10 +0200 |
| Articles | 20 on this page of 84 — 10 participants |
Back to article view | Back to linux.kernel
[RFC PATCH v1 00/11] Create fast idle path for short idle periods Aubrey Li <aubrey.li@intel.com> - 2017-07-10 03:50 +0200
[RFC PATCH v1 04/11] sched/idle: make the fast idle path for short idle periods Aubrey Li <aubrey.li@intel.com> - 2017-07-10 03:50 +0200
Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-11 15:00 +0200
Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for short idle periods Frederic Weisbecker <fweisbec@gmail.com> - 2017-07-11 18:40 +0200
Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-11 20:20 +0200
Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-12 05:30 +0200
Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 07:10 +0200
Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-12 07:30 +0200
Re: [RFC PATCH v1 04/11] sched/idle: make the fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 14:30 +0200
[RFC PATCH v1 05/11] cpuidle: update idle statistics before cpuidle governor Aubrey Li <aubrey.li@intel.com> - 2017-07-10 03:50 +0200
[RFC PATCH v1 08/11] cpuidle: menu: remove reduplicative implementation Aubrey Li <aubrey.li@intel.com> - 2017-07-10 04:00 +0200
[RFC PATCH v1 07/11] cpuidle: make idle residency update more generic Aubrey Li <aubrey.li@intel.com> - 2017-07-10 04:00 +0200
[RFC PATCH v1 03/11] cpuidle: introduce cpuidle governor for idle prediction Aubrey Li <aubrey.li@intel.com> - 2017-07-10 04:00 +0200
Re: [RFC PATCH v1 03/11] cpuidle: introduce cpuidle governor for idle prediction Peter Zijlstra <peterz@infradead.org> - 2017-07-12 14:20 +0200
[RFC PATCH v1 09/11] cpuidle: menu: feed cpuidle prediction to menu governor Aubrey Li <aubrey.li@intel.com> - 2017-07-10 04:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-10 10:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Wanpeng Li <kernellwp@gmail.com> - 2017-07-10 11:40 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-10 16:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-10 16:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-10 18:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-10 19:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-11 06:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-11 11:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Frederic Weisbecker <fweisbec@gmail.com> - 2017-07-11 18:10 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-11 18:40 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-11 20:10 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 14:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 18:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 19:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 21:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 21:10 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 14:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 18:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 19:20 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 20:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 21:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-12 20:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-13 10:40 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-12 06:20 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 10:40 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-12 23:40 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-13 10:40 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-13 16:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-13 17:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-13 17:20 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-13 20:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-14 06:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-14 17:40 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-14 18:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-14 18:10 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-17 11:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-17 15:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-14 18:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-14 18:10 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-14 18:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-17 21:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Arjan van de Ven <arjan@linux.intel.com> - 2017-07-17 21:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-17 21:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Arjan van de Ven <arjan@linux.intel.com> - 2017-07-17 22:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-17 22:10 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Arjan van de Ven <arjan@linux.intel.com> - 2017-07-17 21:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-17 22:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Arjan van de Ven <arjan@linux.intel.com> - 2017-07-17 22:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-18 05:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-18 05:20 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-18 06:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-18 08:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Andi Kleen <ak@linux.intel.com> - 2017-07-18 09:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-18 09:20 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Thomas Gleixner <tglx@linutronix.de> - 2017-07-18 09:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-18 09:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Arjan van de Ven <arjan@linux.intel.com> - 2017-07-14 18:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-13 17:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-14 05:50 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-07-14 06:10 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-17 15:30 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-17 16:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-17 16:10 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 14:20 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Christoph Lameter <cl@linux.com> - 2017-07-11 20:00 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-12 04:10 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods "Li, Aubrey" <aubrey.li@linux.intel.com> - 2017-07-12 04:40 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-12 20:20 +0200
Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods Peter Zijlstra <peterz@infradead.org> - 2017-07-11 11:10 +0200
Page 4 of 5 — ← Prev page 1 2 3 [4] 5 Next page →
| From | Arjan van de Ven <arjan@linux.intel.com> |
|---|---|
| Date | 2017-07-17 21:50 +0200 |
| Message-ID | <u4kbN-Pt-51@gated-at.bofh.it> |
| In reply to | #1689384 |
On 7/17/2017 12:23 PM, Peter Zijlstra wrote: > Of course, this all assumes a Gaussian distribution to begin with, if we > get bimodal (or worse) distributions we can still get it wrong. To fix > that, we'd need to do something better than what we currently have. > fwiw some time ago I made a chart for predicted vs actual so you can sort of judge the distribution of things visually http://git.fenrus.org/tmp/linux2.png
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-07-17 22:00 +0200 |
| Subject | Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods |
| Message-ID | <u4klr-SQ-5@gated-at.bofh.it> |
| In reply to | #1689411 |
On Mon, 17 Jul 2017, Arjan van de Ven wrote: > On 7/17/2017 12:23 PM, Peter Zijlstra wrote: > > Of course, this all assumes a Gaussian distribution to begin with, if we > > get bimodal (or worse) distributions we can still get it wrong. To fix > > that, we'd need to do something better than what we currently have. > > > > fwiw some time ago I made a chart for predicted vs actual so you can sort > of judge the distribution of things visually Predicted by what? Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Arjan van de Ven <arjan@linux.intel.com> |
|---|---|
| Date | 2017-07-17 22:00 +0200 |
| Message-ID | <u4kls-SQ-31@gated-at.bofh.it> |
| In reply to | #1689414 |
On 7/17/2017 12:53 PM, Thomas Gleixner wrote: > On Mon, 17 Jul 2017, Arjan van de Ven wrote: >> On 7/17/2017 12:23 PM, Peter Zijlstra wrote: >>> Of course, this all assumes a Gaussian distribution to begin with, if we >>> get bimodal (or worse) distributions we can still get it wrong. To fix >>> that, we'd need to do something better than what we currently have. >>> >> >> fwiw some time ago I made a chart for predicted vs actual so you can sort >> of judge the distribution of things visually > > Predicted by what? this chart was with the current linux predictor http://git.fenrus.org/tmp/timer.png is what you get if you JUST use the next timer ;-) (which way back linux was doing)
[toc] | [prev] | [next] | [standalone]
| From | "Li, Aubrey" <aubrey.li@linux.intel.com> |
|---|---|
| Date | 2017-07-18 05:30 +0200 |
| Message-ID | <u4rmV-5vl-15@gated-at.bofh.it> |
| In reply to | #1689411 |
On 2017/7/18 3:48, Arjan van de Ven wrote: > On 7/17/2017 12:23 PM, Peter Zijlstra wrote: >> Of course, this all assumes a Gaussian distribution to begin with, if we >> get bimodal (or worse) distributions we can still get it wrong. To fix >> that, we'd need to do something better than what we currently have. >> > > fwiw some time ago I made a chart for predicted vs actual so you can sort > of judge the distribution of things visually > > http://git.fenrus.org/tmp/linux2.png > > This does not look like a Gaussian, does it? I mean abs(expected - actual). Thanks, -Aubrey
[toc] | [prev] | [next] | [standalone]
| From | "Li, Aubrey" <aubrey.li@linux.intel.com> |
|---|---|
| Date | 2017-07-18 05:20 +0200 |
| Message-ID | <u4rdf-5rV-1@gated-at.bofh.it> |
| In reply to | #1689384 |
On 2017/7/18 3:23, Peter Zijlstra wrote: > On Fri, Jul 14, 2017 at 09:26:19AM -0700, Andi Kleen wrote: >>> And as said; Daniel has been working on a better predictor -- now he's >>> probably not used it on the network workload you're looking at, so that >>> might be something to consider. >> >> Deriving a better idle predictor is a bit orthogonal to fast idle. > > No. If you want a different C state selected we need to fix the current > C state selector. We're not going to tinker. > > And the predictor is probably the most fundamental part of the whole C > state selection logic. > > Now I think the problem is that the current predictor goes for an > average idle duration. This means that we, on average, get it wrong 50% > of the time. For performance that's bad. > > If you want to improve the worst case, we need to consider a cumulative > distribution function, and staying with the Gaussian assumption already > present, that would mean using: > > 1 x - mu > CDF(x) = - [ 1 + erf(-------------) ] > 2 sigma sqrt(2) > > Where, per the normal convention mu is the average and sigma^2 the > variance. See also: > > https://en.wikipedia.org/wiki/Normal_distribution > > We then solve CDF(x) = n% to find the x for which we get it wrong n% of > the time (IIRC something like: 'mu - 2sigma' ends up being 5% or so). > > This conceptually gets us better exit latency for the cases where we got > it wrong before, and practically pushes down the estimate which gets us > C1 longer. > > Of course, this all assumes a Gaussian distribution to begin with, if we > get bimodal (or worse) distributions we can still get it wrong. To fix > that, we'd need to do something better than what we currently have. > Maybe you are talking about applying some machine learning algorithm online to fit a multivariate normal distribution, :) Well, back to the problem, when the scheduler picks up idle thread, it does not look at the history, nor make the prediction. So it's possible it has to switch back a task ASAP when it's going into idle(very common under some workloads). That is, (idle_entry + idle_exit) > idle. If the system has multiple hardware idle states, then: (idle_entry + idle_exit + HW_entry + HW_exit) > HW_sleep So we eventually want the idle path lighter than what we currently have. A complex predictor may have high accuracy, but the cost could be high as well. We need a tradeoff here IMHO. I'll check Daniel's work to understand how/if it's better than menu governor. Thanks, -Aubrey
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <ak@linux.intel.com> |
|---|---|
| Date | 2017-07-18 06:50 +0200 |
| Message-ID | <u4sCl-6f0-3@gated-at.bofh.it> |
| In reply to | #1689628 |
> We need a tradeoff here IMHO. I'll check Daniel's work to understand how/if > it's better than menu governor. I still would like to see how the fast path without the C1 heuristic works. Fast pathing is a different concept from a better predictor. IMHO we need both, but the first is likely lower hanging fruit. -Andi
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-07-18 08:50 +0200 |
| Subject | Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods |
| Message-ID | <u4uut-7ol-11@gated-at.bofh.it> |
| In reply to | #1689690 |
On Mon, 17 Jul 2017, Andi Kleen wrote: > > We need a tradeoff here IMHO. I'll check Daniel's work to understand how/if > > it's better than menu governor. > > I still would like to see how the fast path without the C1 heuristic works. > > Fast pathing is a different concept from a better predictor. IMHO we need > both, but the first is likely lower hanging fruit. Hacking something on the side is always the lower hanging fruit as it avoids solving the hard problems. As Peter said already, that's not going to happen unless there is a real technical reason why the general path cannot be fixed. So far there is no proof for that. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <ak@linux.intel.com> |
|---|---|
| Date | 2017-07-18 09:00 +0200 |
| Message-ID | <u4uEa-7rF-13@gated-at.bofh.it> |
| In reply to | #1689778 |
On Tue, Jul 18, 2017 at 08:43:53AM +0200, Thomas Gleixner wrote: > On Mon, 17 Jul 2017, Andi Kleen wrote: > > > > We need a tradeoff here IMHO. I'll check Daniel's work to understand how/if > > > it's better than menu governor. > > > > I still would like to see how the fast path without the C1 heuristic works. > > > > Fast pathing is a different concept from a better predictor. IMHO we need > > both, but the first is likely lower hanging fruit. > > Hacking something on the side is always the lower hanging fruit as it > avoids solving the hard problems. As Peter said already, that's not going > to happen unless there is a real technical reason why the general path > cannot be fixed. So far there is no proof for that. You didn't look at Aubrey's data? There are some unavoidable slow operations in the current path -- e.g. reprograming the timer for NOHZ. But we don't need that for really short idle periods, because as you pointed out they never get woken up by the tick. Similar for other things like RCU. I don't see how you can avoid that other than without a fast path mechanism. Clearly these operations are eventually needed, just not all the time for short sleeps. Now in theory you could have lots of little fast paths in all the individual operations that check this individually, but I don't see how that is better than a single simple fast path. -Andi
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-07-18 09:20 +0200 |
| Subject | Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods |
| Message-ID | <u4uXw-7P3-13@gated-at.bofh.it> |
| In reply to | #1689787 |
On Mon, 17 Jul 2017, Andi Kleen wrote: > On Tue, Jul 18, 2017 at 08:43:53AM +0200, Thomas Gleixner wrote: > > On Mon, 17 Jul 2017, Andi Kleen wrote: > > > > > > We need a tradeoff here IMHO. I'll check Daniel's work to understand how/if > > > > it's better than menu governor. > > > > > > I still would like to see how the fast path without the C1 heuristic works. > > > > > > Fast pathing is a different concept from a better predictor. IMHO we need > > > both, but the first is likely lower hanging fruit. > > > > Hacking something on the side is always the lower hanging fruit as it > > avoids solving the hard problems. As Peter said already, that's not going > > to happen unless there is a real technical reason why the general path > > cannot be fixed. So far there is no proof for that. > > You didn't look at Aubrey's data? > > There are some unavoidable slow operations in the current path -- e.g. > reprograming the timer for NOHZ. But we don't need that for really > short idle periods, because as you pointed out they never get woken > up by the tick. > > Similar for other things like RCU. > > I don't see how you can avoid that other than without a fast path mechanism. You can very well avoid it by taking the irq timings or whatever other information into account for the NOHZ decision. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-07-18 09:30 +0200 |
| Subject | Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods |
| Message-ID | <u4v7c-7Sa-19@gated-at.bofh.it> |
| In reply to | #1689787 |
On Mon, 17 Jul 2017, Andi Kleen wrote: > On Tue, Jul 18, 2017 at 08:43:53AM +0200, Thomas Gleixner wrote: > > On Mon, 17 Jul 2017, Andi Kleen wrote: > > > > > > We need a tradeoff here IMHO. I'll check Daniel's work to understand how/if > > > > it's better than menu governor. > > > > > > I still would like to see how the fast path without the C1 heuristic works. > > > > > > Fast pathing is a different concept from a better predictor. IMHO we need > > > both, but the first is likely lower hanging fruit. > > > > Hacking something on the side is always the lower hanging fruit as it > > avoids solving the hard problems. As Peter said already, that's not going > > to happen unless there is a real technical reason why the general path > > cannot be fixed. So far there is no proof for that. > > You didn't look at Aubrey's data? I did, but that data is no proof that it is unfixable. It's just data describing the current situation, not more not less. > There are some unavoidable slow operations in the current path -- e.g. That's the whole point: current path, IOW current implementation. This implementation is not set in stone and we rather fix it than just creating a side channel and leave everything else as is. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | "Li, Aubrey" <aubrey.li@linux.intel.com> |
|---|---|
| Date | 2017-07-18 09:00 +0200 |
| Message-ID | <u4uEa-7rF-15@gated-at.bofh.it> |
| In reply to | #1689778 |
On 2017/7/18 14:43, Thomas Gleixner wrote: > On Mon, 17 Jul 2017, Andi Kleen wrote: > >>> We need a tradeoff here IMHO. I'll check Daniel's work to understand how/if >>> it's better than menu governor. >> >> I still would like to see how the fast path without the C1 heuristic works. >> >> Fast pathing is a different concept from a better predictor. IMHO we need >> both, but the first is likely lower hanging fruit. > > Hacking something on the side is always the lower hanging fruit as it > avoids solving the hard problems. As Peter said already, that's not going > to happen unless there is a real technical reason why the general path > cannot be fixed. So far there is no proof for that. > Let me try to make a summary, please correct me if I was wrong. 1) for quiet_vmstat, we are agreed to move to another place where tick is really stopped. 2) for rcu idle enter/exit, I measured the details which Paul provided, and the result matches with what I have measured before, nothing notable found. But it still makes more sense if we can make rcu idle enter/exit hooked with tick off. (it's possible other workloads behave differently) 3) for tick nohz idle, we want to skip if the coming idle is short. If we can skip the tick nohz idle, we then skip all the items depending on it. But, there are two hard points: 3.1) how to compute the period of the coming idle. My current proposal is to use two factors in the current idle menu governor. There are two possible options from Peter and Thomas, the one is to use scheduler idle estimate, which is task activity based, the other is to use the statistics generated from irq timings work. 3.2) how to determine if the idle is short or long. My current proposal is to use a tunable value via /sys, while Peter prefers an auto-adjust mechanism. I didn't get the details of an auto-adjust mechanism yet 4) for idle loop, my proposal introduces a simple one to use default idle routine directly, while Peter and Thomas suggest we fix c-state selection in the existing idle path. Thanks, -Aubrey
[toc] | [prev] | [next] | [standalone]
| From | Arjan van de Ven <arjan@linux.intel.com> |
|---|---|
| Date | 2017-07-14 18:00 +0200 |
| Message-ID | <u3bax-5vu-3@gated-at.bofh.it> |
| In reply to | #1687517 |
On 7/14/2017 8:38 AM, Peter Zijlstra wrote: > No, that's wrong. We want to fix the normal C state selection process to > pick the right C state. > > The fast-idle criteria could cut off a whole bunch of available C > states. We need to understand why our current C state pick is wrong and > amend the algorithm to do better. Not just bolt something on the side. I can see a fast path through selection if you know the upper bound of any selection is just 1 state. But also, how much of this is about "C1 be fast" versus "selecting C1 is slow" a lot of the patches in the thread seem to be about making a lighter/faster C1, which is reasonable (you can even argue we might end up with 2 C1s, one fast one full feature)
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-13 17:30 +0200 |
| Message-ID | <u2OdX-7kO-7@gated-at.bofh.it> |
| In reply to | #1686601 |
On Thu, Jul 13, 2017 at 04:53:11PM +0200, Peter Zijlstra wrote: > On Thu, Jul 13, 2017 at 10:48:55PM +0800, Li, Aubrey wrote: > > > - totally from arch_cpu_idle_enter entry to arch_cpu_idle_exit return costs > > 9122ns - 15318ns. > > ---- In this period(arch idle), rcu_idle_enter costs 1985ns - 2262ns, rcu_idle_exit > > costs 1813ns - 3507ns > > > > Besides RCU, > > So Paul wants more details on where RCU hurts so we can try to fix. More specifically: rcu_needs_cpu(), rcu_prepare_for_idle(), rcu_cleanup_after_idle(), rcu_eqs_enter(), rcu_eqs_enter_common(), rcu_dynticks_eqs_enter(), do_nocb_deferred_wakeup(), rcu_dynticks_task_enter(), rcu_eqs_exit(), rcu_eqs_exit_common(), rcu_dynticks_task_exit(), rcu_dynticks_eqs_exit(). The first three (rcu_needs_cpu(), rcu_prepare_for_idle(), and rcu_cleanup_after_idle()) should not be significant unless you have CONFIG_RCU_FAST_NO_HZ=y. If you do, it would be interesting to learn how often invoke_rcu_core() is invoked from rcu_prepare_for_idle() and rcu_cleanup_after_idle(), as this can raise softirq. Also rcu_accelerate_cbs() and rcu_try_advance_all_cbs(). Knowing which of these is causing the most trouble might help me reduce the overhead in the current idle path. Also, how big is this system? If you can say, about what is the cost of a cache miss to some other CPU's cache? Thanx, Paul > > the period includes c-state selection on X86, a few timestamp updates > > and a few computations in menu governor. Also, deep HW-cstate latency can be up > > to 100+ microseconds, even if the system is very busy, CPU still has chance to enter > > deep cstate, which I guess some outburst workloads are not happy with it. > > > > That's my major concern without a fast idle path. > > Fixing C-state selection by creating an alternative idle path sounds so > very wrong. >
[toc] | [prev] | [next] | [standalone]
| From | "Li, Aubrey" <aubrey.li@linux.intel.com> |
|---|---|
| Date | 2017-07-14 05:50 +0200 |
| Message-ID | <u2ZM5-66v-1@gated-at.bofh.it> |
| In reply to | #1686628 |
On 2017/7/13 23:20, Paul E. McKenney wrote: > On Thu, Jul 13, 2017 at 04:53:11PM +0200, Peter Zijlstra wrote: >> On Thu, Jul 13, 2017 at 10:48:55PM +0800, Li, Aubrey wrote: >> >>> - totally from arch_cpu_idle_enter entry to arch_cpu_idle_exit return costs >>> 9122ns - 15318ns. >>> ---- In this period(arch idle), rcu_idle_enter costs 1985ns - 2262ns, rcu_idle_exit >>> costs 1813ns - 3507ns >>> >>> Besides RCU, >> >> So Paul wants more details on where RCU hurts so we can try to fix. > > More specifically: rcu_needs_cpu(), rcu_prepare_for_idle(), > rcu_cleanup_after_idle(), rcu_eqs_enter(), rcu_eqs_enter_common(), > rcu_dynticks_eqs_enter(), do_nocb_deferred_wakeup(), > rcu_dynticks_task_enter(), rcu_eqs_exit(), rcu_eqs_exit_common(), > rcu_dynticks_task_exit(), rcu_dynticks_eqs_exit(). > > The first three (rcu_needs_cpu(), rcu_prepare_for_idle(), and > rcu_cleanup_after_idle()) should not be significant unless you have > CONFIG_RCU_FAST_NO_HZ=y. If you do, it would be interesting to learn > how often invoke_rcu_core() is invoked from rcu_prepare_for_idle() > and rcu_cleanup_after_idle(), as this can raise softirq. Also > rcu_accelerate_cbs() and rcu_try_advance_all_cbs(). > > Knowing which of these is causing the most trouble might help me > reduce the overhead in the current idle path. > I don't have details of these functions, I can measure if you want. Do you have preferred workload for the measurement? > Also, how big is this system? If you can say, about what is the cost > of a cache miss to some other CPU's cache? > The system has two NUMA nodes. nproc returns 104. local memory access is ~100 ns and remote memory access is ~200ns, reported by mgen. Does this address your question? Thanks, -Aubrey
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2017-07-14 06:10 +0200 |
| Message-ID | <u305s-6t9-5@gated-at.bofh.it> |
| In reply to | #1687051 |
On Fri, Jul 14, 2017 at 11:47:32AM +0800, Li, Aubrey wrote: > On 2017/7/13 23:20, Paul E. McKenney wrote: > > On Thu, Jul 13, 2017 at 04:53:11PM +0200, Peter Zijlstra wrote: > >> On Thu, Jul 13, 2017 at 10:48:55PM +0800, Li, Aubrey wrote: > >> > >>> - totally from arch_cpu_idle_enter entry to arch_cpu_idle_exit return costs > >>> 9122ns - 15318ns. > >>> ---- In this period(arch idle), rcu_idle_enter costs 1985ns - 2262ns, rcu_idle_exit > >>> costs 1813ns - 3507ns > >>> > >>> Besides RCU, > >> > >> So Paul wants more details on where RCU hurts so we can try to fix. > > > > More specifically: rcu_needs_cpu(), rcu_prepare_for_idle(), > > rcu_cleanup_after_idle(), rcu_eqs_enter(), rcu_eqs_enter_common(), > > rcu_dynticks_eqs_enter(), do_nocb_deferred_wakeup(), > > rcu_dynticks_task_enter(), rcu_eqs_exit(), rcu_eqs_exit_common(), > > rcu_dynticks_task_exit(), rcu_dynticks_eqs_exit(). > > > > The first three (rcu_needs_cpu(), rcu_prepare_for_idle(), and > > rcu_cleanup_after_idle()) should not be significant unless you have > > CONFIG_RCU_FAST_NO_HZ=y. If you do, it would be interesting to learn > > how often invoke_rcu_core() is invoked from rcu_prepare_for_idle() > > and rcu_cleanup_after_idle(), as this can raise softirq. Also > > rcu_accelerate_cbs() and rcu_try_advance_all_cbs(). > > > > Knowing which of these is causing the most trouble might help me > > reduce the overhead in the current idle path. > > > I don't have details of these functions, I can measure if you want. > Do you have preferred workload for the measurement? I do not have a specific workload in mind. Could you please choose one with very frequent transitions to and from idle? > > Also, how big is this system? If you can say, about what is the cost > > of a cache miss to some other CPU's cache? > > > The system has two NUMA nodes. nproc returns 104. local memory access is > ~100 ns and remote memory access is ~200ns, reported by mgen. Does this > address your question? Very much so, thank you! This will allow me to correctly interpret time spent in the above functions. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | "Li, Aubrey" <aubrey.li@linux.intel.com> |
|---|---|
| Date | 2017-07-17 15:30 +0200 |
| Message-ID | <u4eg1-5DQ-1@gated-at.bofh.it> |
| In reply to | #1687058 |
On 2017/7/14 12:05, Paul E. McKenney wrote: > > More specifically: rcu_needs_cpu(), rcu_prepare_for_idle(), > rcu_cleanup_after_idle(), rcu_eqs_enter(), rcu_eqs_enter_common(), > rcu_dynticks_eqs_enter(), do_nocb_deferred_wakeup(), > rcu_dynticks_task_enter(), rcu_eqs_exit(), rcu_eqs_exit_common(), > rcu_dynticks_task_exit(), rcu_dynticks_eqs_exit(). > > The first three (rcu_needs_cpu(), rcu_prepare_for_idle(), and > rcu_cleanup_after_idle()) should not be significant unless you have > CONFIG_RCU_FAST_NO_HZ=y. If you do, it would be interesting to learn > how often invoke_rcu_core() is invoked from rcu_prepare_for_idle() > and rcu_cleanup_after_idle(), as this can raise softirq. Also > rcu_accelerate_cbs() and rcu_try_advance_all_cbs(). > > Knowing which of these is causing the most trouble might help me > reduce the overhead in the current idle path. I measured two cases, nothing notable found. The one is CONFIG_NO_HZ_IDLE=y, so the following function is just empty. rcu_prepare_for_idle(): NULL rcu_cleanup_after_idle(): NULL do_nocb_deferred_wakeup(): NULL rcu_dynticks_task_enter(): NULL rcu_dynticks_task_exit(): NULL And the following functions are traced separately, for each function I traced 3 times by intel_PT, for each time the sampling period is 1-second. num means the times the function is invoked in 1-second. (min, max, avg) is the function duration, the unit is nano-second. rcu_needs_cpu(): 1) num: 6110 min: 3 max: 564 avg: 17.0 2) num: 16535 min: 3 max: 683 avg: 18.0 3) num: 8815 min: 3 max: 394 avg: 20.0 rcu_eqs_enter(): 1) num: 7956 min: 17 max: 656 avg: 32.0 2) num: 9170 min: 17 max: 1075 avg: 35.0 3) num: 8604 min: 17 max: 859 avg: 29.0 rcu_eqs_enter_common(): 1) num: 14676 min: 15 max: 620 avg: 28.0 2) num: 11180 min: 15 max: 795 avg: 30.0 3) num: 11484 min: 15 max: 725 avg: 29.0 rcu_dynticks_eqs_enter(): 1) num: 11035 min: 10 max: 580 avg: 17.0 2) num: 15518 min: 10 max: 456 avg: 16.0 3) num: 15320 min: 10 max: 454 avg: 19.0 rcu_eqs_exit(): 1) num: 11080 min: 14 max: 893 avg: 23.0 2) num: 13526 min: 14 max: 640 avg: 23.0 3) num: 12534 min: 14 max: 630 avg: 22.0 rcu_eqs_exit_common(): 1) num: 18002 min: 12 max: 553 avg: 17.0 2) num: 10570 min: 11 max: 485 avg: 17.0 3) num: 13628 min: 11 max: 567 avg: 16.0 rcu_dynticks_eqs_exit(): 1) num: 11195 min: 11 max: 436 avg: 16.0 2) num: 11808 min: 10 max: 506 avg: 16.0 3) num: 8132 min: 10 max: 546 avg: 15.0 ============================================================================== The other case is CONFIG_NO_HZ_FULL, I also enabled the required config to make all the functions not empty. rcu_needs_cpu(): 1) num: 8530 min: 5 max: 770 avg: 13.0 2) num: 9965 min: 5 max: 518 avg: 12.0 3) num: 12503 min: 5 max: 755 avg: 16.0 rcu_prepare_for_idle(): 1) num: 11662 min: 5 max: 684 avg: 9.0 2) num: 15294 min: 5 max: 676 avg: 9.0 3) num: 14332 min: 5 max: 524 avg: 9.0 rcu_cleanup_after_idle(): 1) num: 13584 min: 4 max: 657 avg: 6.0 2) num: 9102 min: 4 max: 529 avg: 5.0 3) num: 10648 min: 4 max: 471 avg: 6.0 rcu_eqs_enter(): 1) num: 14222 min: 26 max: 745 avg: 54.0 2) num: 12502 min: 26 max: 650 avg: 53.0 3) num: 11834 min: 26 max: 863 avg: 52.0 rcu_eqs_enter_common(): 1) num: 16792 min: 24 max: 973 avg: 43.0 2) num: 19755 min: 24 max: 898 avg: 45.0 3) num: 8167 min: 24 max: 722 avg: 42.0 rcu_dynticks_eqs_enter(): 1) num: 11605 min: 10 max: 532 avg: 14.0 2) num: 10438 min: 9 max: 554 avg: 14.0 3) num: 19816 min: 9 max: 701 avg: 14.0 do_nocb_deferred_wakeup(): 1) num: 15348 min: 1 max: 459 avg: 3.0 2) num: 12822 min: 1 max: 564 avg: 4.0 3) num: 8272 min: 0 max: 458 avg: 3.0 rcu_dynticks_task_enter(): 1) num: 6358 min: 1 max: 268 avg: 1.0 2) num: 11128 min: 1 max: 360 avg: 1.0 3) num: 20516 min: 1 max: 214 avg: 1.0 rcu_eqs_exit(): 1) num: 16117 min: 20 max: 782 avg: 43.0 2) num: 11042 min: 20 max: 775 avg: 47.0 3) num: 16499 min: 20 max: 752 avg: 46.0 rcu_eqs_exit_common(): 1) num: 12584 min: 17 max: 703 avg: 28.0 2) num: 17412 min: 17 max: 759 avg: 28.0 3) num: 16733 min: 17 max: 798 avg: 29.0 rcu_dynticks_task_exit(): 1) num: 11730 min: 1 max: 528 avg: 4.0 2) num: 18840 min: 1 max: 581 avg: 5.0 3) num: 9815 min: 1 max: 381 avg: 4.0 rcu_dynticks_eqs_exit(): 1) num: 10902 min: 9 max: 557 avg: 13.0 2) num: 19474 min: 9 max: 563 avg: 13.0 3) num: 11865 min: 9 max: 672 avg: 12.0 Please let me know if there is some data not reasonable, I can revisit again. Thanks, -Aubrey
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-17 16:00 +0200 |
| Message-ID | <u4eJ5-5NI-21@gated-at.bofh.it> |
| In reply to | #1689032 |
On Mon, Jul 17, 2017 at 09:24:51PM +0800, Li, Aubrey wrote:
> On 2017/7/14 12:05, Paul E. McKenney wrote:
> >
> > More specifically: rcu_needs_cpu(), rcu_prepare_for_idle(),
> > rcu_cleanup_after_idle(), rcu_eqs_enter(), rcu_eqs_enter_common(),
> > rcu_dynticks_eqs_enter(), do_nocb_deferred_wakeup(),
> > rcu_dynticks_task_enter(), rcu_eqs_exit(), rcu_eqs_exit_common(),
> > rcu_dynticks_task_exit(), rcu_dynticks_eqs_exit().
> >
> > The first three (rcu_needs_cpu(), rcu_prepare_for_idle(), and
> > rcu_cleanup_after_idle()) should not be significant unless you have
> > CONFIG_RCU_FAST_NO_HZ=y. If you do, it would be interesting to learn
> > how often invoke_rcu_core() is invoked from rcu_prepare_for_idle()
> > and rcu_cleanup_after_idle(), as this can raise softirq. Also
> > rcu_accelerate_cbs() and rcu_try_advance_all_cbs().
> >
> > Knowing which of these is causing the most trouble might help me
> > reduce the overhead in the current idle path.
>
> I measured two cases, nothing notable found.
So skipping rcu_idle_{enter,exit}() is not in fact needed at all?
[toc] | [prev] | [next] | [standalone]
| From | "Li, Aubrey" <aubrey.li@linux.intel.com> |
|---|---|
| Date | 2017-07-17 16:10 +0200 |
| Message-ID | <u4eSJ-65Z-1@gated-at.bofh.it> |
| In reply to | #1689075 |
On 2017/7/17 21:58, Peter Zijlstra wrote:
> On Mon, Jul 17, 2017 at 09:24:51PM +0800, Li, Aubrey wrote:
>> On 2017/7/14 12:05, Paul E. McKenney wrote:
>>>
>>> More specifically: rcu_needs_cpu(), rcu_prepare_for_idle(),
>>> rcu_cleanup_after_idle(), rcu_eqs_enter(), rcu_eqs_enter_common(),
>>> rcu_dynticks_eqs_enter(), do_nocb_deferred_wakeup(),
>>> rcu_dynticks_task_enter(), rcu_eqs_exit(), rcu_eqs_exit_common(),
>>> rcu_dynticks_task_exit(), rcu_dynticks_eqs_exit().
>>>
>>> The first three (rcu_needs_cpu(), rcu_prepare_for_idle(), and
>>> rcu_cleanup_after_idle()) should not be significant unless you have
>>> CONFIG_RCU_FAST_NO_HZ=y. If you do, it would be interesting to learn
>>> how often invoke_rcu_core() is invoked from rcu_prepare_for_idle()
>>> and rcu_cleanup_after_idle(), as this can raise softirq. Also
>>> rcu_accelerate_cbs() and rcu_try_advance_all_cbs().
>>>
>>> Knowing which of these is causing the most trouble might help me
>>> reduce the overhead in the current idle path.
>>
>> I measured two cases, nothing notable found.
>
> So skipping rcu_idle_{enter,exit}() is not in fact needed at all?
>
I think it would make more sense if we still put them under the case
where tick is really stopped.
Thanks,
-Aubrey
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-07-12 14:20 +0200 |
| Message-ID | <u2oMy-88t-17@gated-at.bofh.it> |
| In reply to | #1685510 |
On Wed, Jul 12, 2017 at 12:15:08PM +0800, Li, Aubrey wrote: > While my proposal is trying to leverage the prediction functionality > of the existing idle menu governor, which works very well for a long > time. Oh, so you've missed the emails where people say its shit? ;-) Look for the emails of Daniel Lezcano who has been working on better estimating the IRQ periodicity. Some of that has recently been merged but it hasn't got any users yet afaik. But as said before, I'm not convinced the actual idle time is the right measure for NOHZ, because many interrupts do not in fact re-enable the tick.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2017-07-11 20:00 +0200 |
| Subject | Re: [RFC PATCH v1 00/11] Create fast idle path for short idle periods |
| Message-ID | <u27C2-5m0-21@gated-at.bofh.it> |
| In reply to | #1685174 |
On Tue, 11 Jul 2017, Frederic Weisbecker wrote:
> > --- a/kernel/time/tick-sched.c
> > +++ b/kernel/time/tick-sched.c
> > @@ -787,6 +787,7 @@ static ktime_t tick_nohz_stop_sched_tick(struct tick_sched *ts,
> > if (!ts->tick_stopped) {
> > calc_load_nohz_start();
> > cpu_load_update_nohz_start();
> > + quiet_vmstat();
>
> This patch seems to make sense. Christoph?
Ok makes sense to me too. Was never entirely sure where the proper place
would be to call it.
[toc] | [prev] | [next] | [standalone]
Page 4 of 5 — ← Prev page 1 2 3 [4] 5 Next page →
Back to top | Article view | linux.kernel
csiph-web