Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1668143 > unrolled thread
| Started by | Len Brown <lenb@kernel.org> |
|---|---|
| First post | 2017-06-17 05:10 +0200 |
| Last post | 2017-06-19 14:50 +0200 |
| Articles | 3 — 2 participants |
Back to article view | Back to linux.kernel
[GIT PULL] x86,cpufreq: unify APERF/MPERF computation Len Brown <lenb@kernel.org> - 2017-06-17 05:10 +0200
[PATCH 3/4] intel_pstate: delete scheduler hook in HWP mode Len Brown <lenb@kernel.org> - 2017-06-17 05:10 +0200
Re: [GIT PULL] x86,cpufreq: unify APERF/MPERF computation "Rafael J. Wysocki" <rafael@kernel.org> - 2017-06-19 14:50 +0200
| From | Len Brown <lenb@kernel.org> |
|---|---|
| Date | 2017-06-17 05:10 +0200 |
| Subject | [GIT PULL] x86,cpufreq: unify APERF/MPERF computation |
| Message-ID | <tTchz-11g-3@gated-at.bofh.it> |
In-Reply-To:
Hi Rafael,
This patch series has 3 goals:
1. Make "cpu MHz" in /proc/cpuinfo supportable.
2. Make /sys/.../cpufreq/scaling_cur_freq meaningful
and consistent on modern x86 systems.
3. Use 1. and 2. to remove scheduler and cpufreq overhead
There are 3 main changes since this series was proposed
about a year ago:
This update responds to distro feedback to make /proc/cpuinfo
"cpu MHz" constant. Originally, we had proposed making it return
the same dynamic value as cpufreq sysfs.
Some community members suggested that sysfs MHz values should
be meaninful, even down to 10ms intervals. So this has been
changed, versus the original proposal to not re-compute
at intervals shorter than 100ms.
(For those who really care about observing frequency, the
recommendation remains to use turbostat(8) or equivalent utility,
which can reliably measure concurrent intervals of arbitrary length)
The intel_pstate sampling mechanism has changed.
Originally this series removed an intel_pstate timer in HWP mode.
Now it removes the analogous scheduler call-back.
Most recently, in response to posting this patch on the list
about 10-days ago, the patch to remove frequency calculation
from inside intel_pstate was dropped, in order to maintain compatibility
with tracing scripts. Also, the order of the last two patches
has been exchanged.
Please let me know if you see any issues with this series.
thanks!
Len Brown, Intel Open Source Technology Center
The following changes since commit 3c2993b8c6143d8a5793746a54eba8f86f95240f:
Linux 4.12-rc4 (2017-06-04 16:47:43 -0700)
are available in the git repository at:
git://git.kernel.org/pub/scm/linux/kernel/git/lenb/linux.git x86
for you to fetch changes up to d020eed98440faa4a529c621f881aa9fda296956:
intel_pstate: skip scheduler hook when in "performance" mode. (2017-06-16 19:11:13 -0700)
----------------------------------------------------------------
Len Brown (4):
x86: do not use cpufreq_quick_get() for /proc/cpuinfo "cpu MHz"
x86: use common aperfmperf_khz_on_cpu() to calculate KHz using APERF/MPERF
intel_pstate: delete scheduler hook in HWP mode
intel_pstate: skip scheduler hook when in "performance" mode.
arch/x86/kernel/cpu/Makefile | 1 +
arch/x86/kernel/cpu/aperfmperf.c | 82 ++++++++++++++++++++++++++++++++++++++++
arch/x86/kernel/cpu/proc.c | 10 +----
drivers/cpufreq/cpufreq.c | 7 +++-
drivers/cpufreq/intel_pstate.c | 18 +++------
include/linux/cpufreq.h | 13 +++++++
6 files changed, 109 insertions(+), 22 deletions(-)
create mode 100644 arch/x86/kernel/cpu/aperfmperf.c
[toc] | [next] | [standalone]
| From | Len Brown <lenb@kernel.org> |
|---|---|
| Date | 2017-06-17 05:10 +0200 |
| Subject | [PATCH 3/4] intel_pstate: delete scheduler hook in HWP mode |
| Message-ID | <tTchA-11g-13@gated-at.bofh.it> |
| In reply to | #1668143 |
From: Len Brown <len.brown@intel.com>
The cpufreq/scaling_cur_freq sysfs attribute is now provided by
shared x86 cpufreq code on modern x86 systems, including
all systems supported by the intel_pstate driver.
In HWP mode, maintaining that value was the sole purpose of
the scheduler hook, intel_pstate_update_util_hwp(),
so it can now be removed.
Signed-off-by: Len Brown <len.brown@intel.com>
---
drivers/cpufreq/intel_pstate.c | 14 +++-----------
1 file changed, 3 insertions(+), 11 deletions(-)
diff --git a/drivers/cpufreq/intel_pstate.c b/drivers/cpufreq/intel_pstate.c
index b7de5bd..4ec5668 100644
--- a/drivers/cpufreq/intel_pstate.c
+++ b/drivers/cpufreq/intel_pstate.c
@@ -1732,16 +1732,6 @@ static void intel_pstate_adjust_pstate(struct cpudata *cpu, int target_pstate)
fp_toint(cpu->iowait_boost * 100));
}
-static void intel_pstate_update_util_hwp(struct update_util_data *data,
- u64 time, unsigned int flags)
-{
- struct cpudata *cpu = container_of(data, struct cpudata, update_util);
- u64 delta_ns = time - cpu->sample.time;
-
- if ((s64)delta_ns >= INTEL_PSTATE_HWP_SAMPLING_INTERVAL)
- intel_pstate_sample(cpu, time);
-}
-
static void intel_pstate_update_util_pid(struct update_util_data *data,
u64 time, unsigned int flags)
{
@@ -1933,6 +1923,9 @@ static void intel_pstate_set_update_util_hook(unsigned int cpu_num)
{
struct cpudata *cpu = all_cpu_data[cpu_num];
+ if (hwp_active)
+ return;
+
if (cpu->update_util_set)
return;
@@ -2557,7 +2550,6 @@ static int __init intel_pstate_init(void)
} else {
hwp_active++;
intel_pstate.attr = hwp_cpufreq_attrs;
- pstate_funcs.update_util = intel_pstate_update_util_hwp;
goto hwp_cpu_matched;
}
} else {
--
2.7.4
[toc] | [prev] | [next] | [standalone]
| From | "Rafael J. Wysocki" <rafael@kernel.org> |
|---|---|
| Date | 2017-06-19 14:50 +0200 |
| Message-ID | <tU4hY-2HA-21@gated-at.bofh.it> |
| In reply to | #1668143 |
On Sat, Jun 17, 2017 at 5:03 AM, Len Brown <lenb@kernel.org> wrote: > > In-Reply-To: > > Hi Rafael, > > This patch series has 3 goals: > > 1. Make "cpu MHz" in /proc/cpuinfo supportable. > > 2. Make /sys/.../cpufreq/scaling_cur_freq meaningful > and consistent on modern x86 systems. > > 3. Use 1. and 2. to remove scheduler and cpufreq overhead > > There are 3 main changes since this series was proposed > about a year ago: > > This update responds to distro feedback to make /proc/cpuinfo > "cpu MHz" constant. Originally, we had proposed making it return > the same dynamic value as cpufreq sysfs. > > Some community members suggested that sysfs MHz values should > be meaninful, even down to 10ms intervals. So this has been > changed, versus the original proposal to not re-compute > at intervals shorter than 100ms. > > (For those who really care about observing frequency, the > recommendation remains to use turbostat(8) or equivalent utility, > which can reliably measure concurrent intervals of arbitrary length) > > The intel_pstate sampling mechanism has changed. > Originally this series removed an intel_pstate timer in HWP mode. > Now it removes the analogous scheduler call-back. > > Most recently, in response to posting this patch on the list > about 10-days ago, the patch to remove frequency calculation > from inside intel_pstate was dropped, in order to maintain compatibility > with tracing scripts. Also, the order of the last two patches > has been exchanged. > > Please let me know if you see any issues with this series. > > thanks! > Len Brown, Intel Open Source Technology Center > > The following changes since commit 3c2993b8c6143d8a5793746a54eba8f86f95240f: > > Linux 4.12-rc4 (2017-06-04 16:47:43 -0700) > > are available in the git repository at: > > git://git.kernel.org/pub/scm/linux/kernel/git/lenb/linux.git x86 > > for you to fetch changes up to d020eed98440faa4a529c621f881aa9fda296956: > > intel_pstate: skip scheduler hook when in "performance" mode. (2017-06-16 19:11:13 -0700) > > ---------------------------------------------------------------- > Len Brown (4): > x86: do not use cpufreq_quick_get() for /proc/cpuinfo "cpu MHz" > x86: use common aperfmperf_khz_on_cpu() to calculate KHz using APERF/MPERF > intel_pstate: delete scheduler hook in HWP mode > intel_pstate: skip scheduler hook when in "performance" mode. I'd like to hear from the x86 maintainers about the first two patches. At least I'd like to know that there are no objections there. Thanks, Rafael
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web