Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1390741 > unrolled thread
| Started by | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| First post | 2016-04-29 06:50 +0200 |
| Last post | 2016-04-29 23:20 +0200 |
| Articles | 10 on this page of 30 — 3 participants |
Back to article view | Back to linux.kernel
[PATCH 00/32] 2nd Iteration of Cache QoS Monitoring support. David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 06:50 +0200
[PATCH 01/32] perf/x86/intel/cqm: temporarily remove MBM from CQM and cleanup David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 06:50 +0200
Re: [PATCH 01/32] perf/x86/intel/cqm: temporarily remove MBM from CQM and cleanup Vikas Shivappa <vikas.shivappa@linux.intel.com> - 2016-04-29 22:30 +0200
[PATCH 24/32] perf/x86/intel/cqm: use PERF_INACTIVE_*_READ_* flags in CQM David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 06:50 +0200
[PATCH 25/32] sched: introduce the finish_arch_pre_lock_switch() scheduler hook David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 06:50 +0200
Re: [PATCH 25/32] sched: introduce the finish_arch_pre_lock_switch() scheduler hook Peter Zijlstra <peterz@infradead.org> - 2016-04-29 11:00 +0200
Re: [PATCH 25/32] sched: introduce the finish_arch_pre_lock_switch() scheduler hook David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 20:50 +0200
Re: [PATCH 25/32] sched: introduce the finish_arch_pre_lock_switch() scheduler hook Vikas Shivappa <vikas.shivappa@linux.intel.com> - 2016-04-29 22:30 +0200
Re: [PATCH 25/32] sched: introduce the finish_arch_pre_lock_switch() scheduler hook David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 23:00 +0200
[PATCH 14/32] perf/x86/intel/cqm: add preallocation of anodes David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 06:50 +0200
[PATCH 32/32] perf/stat: revamp error handling for snapshot and per_pkg events David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 06:50 +0200
[PATCH 09/32] perf/x86/intel/cqm: add per-package RMIDs, data and locks David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 06:50 +0200
Re: [PATCH 09/32] perf/x86/intel/cqm: add per-package RMIDs, data and locks Vikas Shivappa <vikas.shivappa@linux.intel.com> - 2016-04-29 23:00 +0200
[PATCH 29/32] perf,perf/x86,perf/powerpc,perf/arm,perf/*: add int error return to pmu::read David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 06:50 +0200
[PATCH 06/32] x86/intel,cqm: add CONFIG_INTEL_RDT configuration flag and refactor PQR David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 06:50 +0200
[PATCH 15/32] perf/core: add hooks to expose architecture specific features in perf_cgroup David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
[PATCH 18/32] perf/x86/intel/cqm: use pmu::event_terminate David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
[PATCH 16/32] perf/x86/intel/cqm: add cgroup support David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
[PATCH 17/32] perf/core: adding pmu::event_terminate David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
[PATCH 05/32] perf/core: remove unused pmu->count David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
[PATCH 04/32] perf/x86/intel/cqm: make read of RMIDs per package (Temporal) David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
[PATCH 11/32] perf/x86/intel/cqm: (I)state and limbo prmids David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
[PATCH 12/32] perf/x86/intel/cqm: add per-package RMID rotation David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
[PATCH 08/32] perf/x86/intel/cqm: prepare for next patches David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
Re: [PATCH 08/32] perf/x86/intel/cqm: prepare for next patches Peter Zijlstra <peterz@infradead.org> - 2016-04-29 11:20 +0200
[PATCH 07/32] perf/x86/intel/cqm: separate CQM PMU's attributes from x86 PMU David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
[PATCH 19/32] perf/core: introduce PMU event flag PERF_CGROUP_NO_RECURSION David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
[PATCH 10/32] perf/x86/intel/cqm: basic RMID hierarchy with per package rmids David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 07:00 +0200
Re: [PATCH 00/32] 2nd Iteration of Cache QoS Monitoring support. Vikas Shivappa <vikas.shivappa@linux.intel.com> - 2016-04-29 23:10 +0200
Re: [PATCH 00/32] 2nd Iteration of Cache QoS Monitoring support. David Carrillo-Cisneros <davidcc@google.com> - 2016-04-29 23:20 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2016-04-29 07:00 +0200 |
| Subject | [PATCH 04/32] perf/x86/intel/cqm: make read of RMIDs per package (Temporal) |
| Message-ID | <rt8H1-8vJ-33@gated-at.bofh.it> |
| In reply to | #1390741 |
The previous version of Intel's CQM introduced pmu::count as a replacement
for reading CQM events. This was done to avoid using an IPI to read the
CQM occupancy event when reading events attached to a thread.
Using pmu->count in place of pmu->read is inconsistent with the usage by
other PMUs and introduces several problems such as:
1) pmu::read for thread events returns bogus values when called from
interrupts disabled contexts.
2) perf_event_count(), behavior depends on whether interruptions are
enabled or not.
3) perf_event_count() will always read a fresh value from the PMU, which
is inconsistent with the behavior of other events.
4) perf_event_count() will perform slow MSR read and writes and IPIs.
This patches removes pmu::count from CQM and makes pmu::read always
read from the local socket (package). Future patches will add a mechanism
to add the event count from other packages.
This patch also removes the unused field rmid_usecnt from intel_pqr_state.
Reviewed-by: Stephane Eranian <eranian@google.com>
Signed-off-by: David Carrillo-Cisneros <davidcc@google.com>
---
arch/x86/events/intel/cqm.c | 125 ++++++--------------------------------------
1 file changed, 16 insertions(+), 109 deletions(-)
diff --git a/arch/x86/events/intel/cqm.c b/arch/x86/events/intel/cqm.c
index 3c1e247..afd60dd 100644
--- a/arch/x86/events/intel/cqm.c
+++ b/arch/x86/events/intel/cqm.c
@@ -20,7 +20,6 @@ static unsigned int cqm_l3_scale; /* supposedly cacheline size */
* struct intel_pqr_state - State cache for the PQR MSR
* @rmid: The cached Resource Monitoring ID
* @closid: The cached Class Of Service ID
- * @rmid_usecnt: The usage counter for rmid
*
* The upper 32 bits of MSR_IA32_PQR_ASSOC contain closid and the
* lower 10 bits rmid. The update to MSR_IA32_PQR_ASSOC always
@@ -32,7 +31,6 @@ static unsigned int cqm_l3_scale; /* supposedly cacheline size */
struct intel_pqr_state {
u32 rmid;
u32 closid;
- int rmid_usecnt;
};
/*
@@ -44,6 +42,19 @@ struct intel_pqr_state {
static DEFINE_PER_CPU(struct intel_pqr_state, pqr_state);
/*
+ * Updates caller cpu's cache.
+ */
+static inline void __update_pqr_rmid(u32 rmid)
+{
+ struct intel_pqr_state *state = this_cpu_ptr(&pqr_state);
+
+ if (state->rmid == rmid)
+ return;
+ state->rmid = rmid;
+ wrmsr(MSR_IA32_PQR_ASSOC, rmid, state->closid);
+}
+
+/*
* Protects cache_cgroups and cqm_rmid_free_lru and cqm_rmid_limbo_lru.
* Also protects event->hw.cqm_rmid
*
@@ -309,7 +320,7 @@ struct rmid_read {
atomic64_t value;
};
-static void __intel_cqm_event_count(void *info);
+static void intel_cqm_event_read(struct perf_event *event);
/*
* If we fail to assign a new RMID for intel_cqm_rotation_rmid because
@@ -376,12 +387,6 @@ static void intel_cqm_event_read(struct perf_event *event)
u32 rmid;
u64 val;
- /*
- * Task events are handled by intel_cqm_event_count().
- */
- if (event->cpu == -1)
- return;
-
raw_spin_lock_irqsave(&cache_lock, flags);
rmid = event->hw.cqm_rmid;
@@ -401,123 +406,28 @@ out:
raw_spin_unlock_irqrestore(&cache_lock, flags);
}
-static void __intel_cqm_event_count(void *info)
-{
- struct rmid_read *rr = info;
- u64 val;
-
- val = __rmid_read(rr->rmid);
-
- if (val & (RMID_VAL_ERROR | RMID_VAL_UNAVAIL))
- return;
-
- atomic64_add(val, &rr->value);
-}
-
static inline bool cqm_group_leader(struct perf_event *event)
{
return !list_empty(&event->hw.cqm_groups_entry);
}
-static u64 intel_cqm_event_count(struct perf_event *event)
-{
- unsigned long flags;
- struct rmid_read rr = {
- .value = ATOMIC64_INIT(0),
- };
-
- /*
- * We only need to worry about task events. System-wide events
- * are handled like usual, i.e. entirely with
- * intel_cqm_event_read().
- */
- if (event->cpu != -1)
- return __perf_event_count(event);
-
- /*
- * Only the group leader gets to report values. This stops us
- * reporting duplicate values to userspace, and gives us a clear
- * rule for which task gets to report the values.
- *
- * Note that it is impossible to attribute these values to
- * specific packages - we forfeit that ability when we create
- * task events.
- */
- if (!cqm_group_leader(event))
- return 0;
-
- /*
- * Getting up-to-date values requires an SMP IPI which is not
- * possible if we're being called in interrupt context. Return
- * the cached values instead.
- */
- if (unlikely(in_interrupt()))
- goto out;
-
- /*
- * Notice that we don't perform the reading of an RMID
- * atomically, because we can't hold a spin lock across the
- * IPIs.
- *
- * Speculatively perform the read, since @event might be
- * assigned a different (possibly invalid) RMID while we're
- * busying performing the IPI calls. It's therefore necessary to
- * check @event's RMID afterwards, and if it has changed,
- * discard the result of the read.
- */
- rr.rmid = ACCESS_ONCE(event->hw.cqm_rmid);
-
- if (!__rmid_valid(rr.rmid))
- goto out;
-
- on_each_cpu_mask(&cqm_cpumask, __intel_cqm_event_count, &rr, 1);
-
- raw_spin_lock_irqsave(&cache_lock, flags);
- if (event->hw.cqm_rmid == rr.rmid)
- local64_set(&event->count, atomic64_read(&rr.value));
- raw_spin_unlock_irqrestore(&cache_lock, flags);
-out:
- return __perf_event_count(event);
-}
-
static void intel_cqm_event_start(struct perf_event *event, int mode)
{
- struct intel_pqr_state *state = this_cpu_ptr(&pqr_state);
- u32 rmid = event->hw.cqm_rmid;
-
if (!(event->hw.cqm_state & PERF_HES_STOPPED))
return;
event->hw.cqm_state &= ~PERF_HES_STOPPED;
-
- if (state->rmid_usecnt++) {
- if (!WARN_ON_ONCE(state->rmid != rmid))
- return;
- } else {
- WARN_ON_ONCE(state->rmid);
- }
-
- state->rmid = rmid;
- wrmsr(MSR_IA32_PQR_ASSOC, rmid, state->closid);
+ __update_pqr_rmid(event->hw.cqm_rmid);
}
static void intel_cqm_event_stop(struct perf_event *event, int mode)
{
- struct intel_pqr_state *state = this_cpu_ptr(&pqr_state);
-
if (event->hw.cqm_state & PERF_HES_STOPPED)
return;
event->hw.cqm_state |= PERF_HES_STOPPED;
-
intel_cqm_event_read(event);
-
- if (!--state->rmid_usecnt) {
- state->rmid = 0;
- wrmsr(MSR_IA32_PQR_ASSOC, 0, state->closid);
- } else {
- WARN_ON_ONCE(!state->rmid);
- }
+ __update_pqr_rmid(0);
}
static int intel_cqm_event_add(struct perf_event *event, int mode)
@@ -534,7 +444,6 @@ static int intel_cqm_event_add(struct perf_event *event, int mode)
intel_cqm_event_start(event, mode);
raw_spin_unlock_irqrestore(&cache_lock, flags);
-
return 0;
}
@@ -720,7 +629,6 @@ static struct pmu intel_cqm_pmu = {
.start = intel_cqm_event_start,
.stop = intel_cqm_event_stop,
.read = intel_cqm_event_read,
- .count = intel_cqm_event_count,
};
static inline void cqm_pick_event_reader(int cpu)
@@ -743,7 +651,6 @@ static void intel_cqm_cpu_starting(unsigned int cpu)
state->rmid = 0;
state->closid = 0;
- state->rmid_usecnt = 0;
WARN_ON(c->x86_cache_max_rmid != cqm_max_rmid);
WARN_ON(c->x86_cache_occ_scale != cqm_l3_scale);
--
2.8.0.rc3.226.g39d4020
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2016-04-29 07:00 +0200 |
| Subject | [PATCH 11/32] perf/x86/intel/cqm: (I)state and limbo prmids |
| Message-ID | <rt8H1-8vJ-37@gated-at.bofh.it> |
| In reply to | #1390741 |
CQM defines a dirty threshold that is the minimum number of dirty
cache lines that a prmid can hold before being eligible to be reused.
This threshold is zero unless there exist significant contention of prmids
(more on this on the patch that introduces rotation of RMIDs).
A limbo prmid is a prmid that is no longer utilized by any pmonr, yet, its
occupancy exceeds the dirty threshold. This is a consequence of the
hardware design that do not provide a mechanism to flush cache lines
associated with a RMID.
If no pmonr schedules a limbo prmid, it's expected that it's occupancy
will eventually drop below the dirty threshold. Nevertheless, the cache
lines tagged to a limbo prmid still hold valid occupancy for the previous
owner of the prmid. This creates a difference in the way the occupancy of
pmonr is read depending on whether it has hold a prmid recently or not.
This patch introduces the (I)state mentioned in previous changelog.
The (I)state is a superstate conformed by two substates:
- (IL)state: (I)state with limbo prmid, this pmonr held a prmid in
(A)state before its transition to (I)state.
- (IN)state: (I)state without limbo prmid, this pmonr did not held a
prmid recently.
A pmonr in (IL)state keeps the reference to its former prmid in the field
limbo_prmid, this occupancy is counted towards the occupancy
of the ancestors of the pmonr, reducing the error caused by stealing
of prmids during RMID rotation.
In future patches (rotation logic), the occupancy of limbo_prmids is
polled periodically and (IL)state pmonrs with limbo prmids that had become
clean will transition to (IN)state.
Reviewed-by: Stephane Eranian <eranian@google.com>
Signed-off-by: David Carrillo-Cisneros <davidcc@google.com>
---
arch/x86/events/intel/cqm.c | 203 ++++++++++++++++++++++++++++++++++++++++++--
arch/x86/events/intel/cqm.h | 88 +++++++++++++++++--
2 files changed, 277 insertions(+), 14 deletions(-)
diff --git a/arch/x86/events/intel/cqm.c b/arch/x86/events/intel/cqm.c
index 65551bb..caf7152 100644
--- a/arch/x86/events/intel/cqm.c
+++ b/arch/x86/events/intel/cqm.c
@@ -39,16 +39,34 @@ struct monr *monr_hrchy_root;
struct pkg_data *cqm_pkgs_data[PQR_MAX_NR_PKGS];
+static inline bool __pmonr__in_istate(struct pmonr *pmonr)
+{
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+ return pmonr->ancestor_pmonr;
+}
+
+static inline bool __pmonr__in_ilstate(struct pmonr *pmonr)
+{
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+ return __pmonr__in_istate(pmonr) && pmonr->limbo_prmid;
+}
+
+static inline bool __pmonr__in_instate(struct pmonr *pmonr)
+{
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+ return __pmonr__in_istate(pmonr) && !__pmonr__in_ilstate(pmonr);
+}
+
static inline bool __pmonr__in_astate(struct pmonr *pmonr)
{
lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
- return pmonr->prmid;
+ return pmonr->prmid && !pmonr->ancestor_pmonr;
}
static inline bool __pmonr__in_ustate(struct pmonr *pmonr)
{
lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
- return !pmonr->prmid;
+ return !pmonr->prmid && !pmonr->ancestor_pmonr;
}
static inline bool monr__is_root(struct monr *monr)
@@ -210,9 +228,12 @@ static int pkg_data_init_cpu(int cpu)
INIT_LIST_HEAD(&pkg_data->free_prmids_pool);
INIT_LIST_HEAD(&pkg_data->active_prmids_pool);
+ INIT_LIST_HEAD(&pkg_data->pmonr_limbo_prmids_pool);
INIT_LIST_HEAD(&pkg_data->nopmonr_limbo_prmids_pool);
INIT_LIST_HEAD(&pkg_data->astate_pmonrs_lru);
+ INIT_LIST_HEAD(&pkg_data->istate_pmonrs_lru);
+ INIT_LIST_HEAD(&pkg_data->ilstate_pmonrs_lru);
mutex_init(&pkg_data->pkg_data_mutex);
raw_spin_lock_init(&pkg_data->pkg_data_lock);
@@ -261,7 +282,15 @@ static struct pmonr *pmonr_alloc(int cpu)
if (!pmonr)
return ERR_PTR(-ENOMEM);
+ pmonr->ancestor_pmonr = NULL;
+
+ /*
+ * Since (A)state and (I)state have union in members,
+ * initialize one of them only.
+ */
+ INIT_LIST_HEAD(&pmonr->pmonr_deps_head);
pmonr->prmid = NULL;
+ INIT_LIST_HEAD(&pmonr->limbo_rotation_entry);
pmonr->monr = NULL;
INIT_LIST_HEAD(&pmonr->rotation_entry);
@@ -327,6 +356,44 @@ __pmonr__finish_to_astate(struct pmonr *pmonr, struct prmid *prmid)
atomic64_set(&pmonr->prmid_summary_atomic, summary.value);
}
+/*
+ * Transition to (A)state from (IN)state, given a valid prmid.
+ * Cannot fail. Updates ancestor dependants to use this pmonr as new ancestor.
+ */
+static inline void
+__pmonr__instate_to_astate(struct pmonr *pmonr, struct prmid *prmid)
+{
+ struct pmonr *pos, *tmp, *ancestor;
+ union prmid_summary old_summary, summary;
+
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+
+ /* If in (I) state, cannot have limbo_prmid, otherwise prmid
+ * in function's argument is superfluous.
+ */
+ WARN_ON_ONCE(pmonr->limbo_prmid);
+
+ /* Do not depend on ancestor_pmonr anymore. Make it (A)state. */
+ ancestor = pmonr->ancestor_pmonr;
+ list_del_init(&pmonr->pmonr_deps_entry);
+ pmonr->ancestor_pmonr = NULL;
+ __pmonr__finish_to_astate(pmonr, prmid);
+
+ /* Update ex ancestor's dependants that are pmonr descendants. */
+ list_for_each_entry_safe(pos, tmp, &ancestor->pmonr_deps_head,
+ pmonr_deps_entry) {
+ if (!__monr_hrchy_is_ancestor(monr_hrchy_root,
+ pmonr->monr, pos->monr))
+ continue;
+ list_move_tail(&pos->pmonr_deps_entry, &pmonr->pmonr_deps_head);
+ pos->ancestor_pmonr = pmonr;
+ old_summary.value = atomic64_read(&pos->prmid_summary_atomic);
+ summary.sched_rmid = prmid->rmid;
+ summary.read_rmid = old_summary.read_rmid;
+ atomic64_set(&pos->prmid_summary_atomic, summary.value);
+ }
+}
+
static inline void
__pmonr__ustate_to_astate(struct pmonr *pmonr, struct prmid *prmid)
{
@@ -334,9 +401,59 @@ __pmonr__ustate_to_astate(struct pmonr *pmonr, struct prmid *prmid)
__pmonr__finish_to_astate(pmonr, prmid);
}
+/*
+ * Find lowest active ancestor.
+ * Always successful since monr_hrchy_root is always in (A)state.
+ */
+static struct monr *
+__monr_hrchy__find_laa(struct monr *monr, u16 pkg_id)
+{
+ lockdep_assert_held(&cqm_pkgs_data[pkg_id]->pkg_data_lock);
+
+ while ((monr = monr->parent)) {
+ if (__pmonr__in_astate(monr->pmonrs[pkg_id]))
+ return monr;
+ }
+ /* Should have hitted monr_hrchy_root */
+ WARN_ON_ONCE(true);
+ return NULL;
+}
+
+/*
+ * __pmnor__move_dependants: Move dependants from one ancestor to another.
+ * @old: Old ancestor.
+ * @new: New ancestor.
+ *
+ * To be called on valid pmonrs. @new must be ancestor of @old.
+ */
+static inline void
+__pmonr__move_dependants(struct pmonr *old, struct pmonr *new)
+{
+ struct pmonr *dep;
+ union prmid_summary old_summary, summary;
+
+ WARN_ON_ONCE(old->pkg_id != new->pkg_id);
+ lockdep_assert_held(&__pkg_data(old, pkg_data_lock));
+
+ /* Update this pmonr dependencies to use new ancestor. */
+ list_for_each_entry(dep, &old->pmonr_deps_head, pmonr_deps_entry) {
+ /* Set next summary for dependent pmonrs. */
+ dep->ancestor_pmonr = new;
+
+ old_summary.value = atomic64_read(&dep->prmid_summary_atomic);
+ summary.sched_rmid = new->prmid->rmid;
+ summary.read_rmid = old_summary.read_rmid;
+ atomic64_set(&dep->prmid_summary_atomic, summary.value);
+ }
+ list_splice_tail_init(&old->pmonr_deps_head,
+ &new->pmonr_deps_head);
+}
+
static inline void
__pmonr__to_ustate(struct pmonr *pmonr)
{
+ struct pmonr *ancestor;
+ u16 pkg_id = pmonr->pkg_id;
union prmid_summary summary;
lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
@@ -350,9 +467,27 @@ __pmonr__to_ustate(struct pmonr *pmonr)
if (__pmonr__in_astate(pmonr)) {
WARN_ON_ONCE(!pmonr->prmid);
+ ancestor = __monr_hrchy__find_laa(
+ pmonr->monr, pkg_id)->pmonrs[pkg_id];
+ WARN_ON_ONCE(!ancestor);
+ __pmonr__move_dependants(pmonr, ancestor);
list_move_tail(&pmonr->prmid->pool_entry,
&__pkg_data(pmonr, nopmonr_limbo_prmids_pool));
pmonr->prmid = NULL;
+ } else if (__pmonr__in_istate(pmonr)) {
+ list_del_init(&pmonr->pmonr_deps_entry);
+ /* limbo_prmid is already in limbo pool */
+ if (__pmonr__in_ilstate(pmonr)) {
+ WARN_ON(!pmonr->limbo_prmid);
+ list_move_tail(
+ &pmonr->limbo_prmid->pool_entry,
+ &__pkg_data(pmonr, nopmonr_limbo_prmids_pool));
+
+ pmonr->limbo_prmid = NULL;
+ list_del_init(&pmonr->limbo_rotation_entry);
+ } else {
+ }
+ pmonr->ancestor_pmonr = NULL;
} else {
WARN_ON_ONCE(true);
return;
@@ -367,6 +502,62 @@ __pmonr__to_ustate(struct pmonr *pmonr)
WARN_ON_ONCE(!__pmonr__in_ustate(pmonr));
}
+static inline void __pmonr__set_istate_summary(struct pmonr *pmonr)
+{
+ union prmid_summary summary;
+
+ summary.sched_rmid = pmonr->ancestor_pmonr->prmid->rmid;
+ summary.read_rmid =
+ pmonr->limbo_prmid ? pmonr->limbo_prmid->rmid : INVALID_RMID;
+ atomic64_set(
+ &pmonr->prmid_summary_atomic, summary.value);
+}
+
+/*
+ * Transition to (I)state from no (I)state..
+ * Finds a valid ancestor transversing monr_hrchy. Cannot fail.
+ */
+static inline void
+__pmonr__to_istate(struct pmonr *pmonr)
+{
+ struct pmonr *ancestor;
+ u16 pkg_id = pmonr->pkg_id;
+
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+
+ if (!(__pmonr__in_ustate(pmonr) || __pmonr__in_astate(pmonr))) {
+ /* Invalid initial state. */
+ WARN_ON_ONCE(true);
+ return;
+ }
+
+ ancestor = __monr_hrchy__find_laa(pmonr->monr, pkg_id)->pmonrs[pkg_id];
+ WARN_ON_ONCE(!ancestor);
+
+ if (__pmonr__in_astate(pmonr)) {
+ /* Active pmonr->prmid becomes limbo in transition to (I)state.
+ * Note that pmonr->prmid and pmonr->limbo_prmid are in an
+ * union, so no need to copy.
+ */
+ __pmonr__move_dependants(pmonr, ancestor);
+ list_move_tail(&pmonr->limbo_prmid->pool_entry,
+ &__pkg_data(pmonr, pmonr_limbo_prmids_pool));
+ }
+
+ pmonr->ancestor_pmonr = ancestor;
+ list_add_tail(&pmonr->pmonr_deps_entry, &ancestor->pmonr_deps_head);
+
+ list_move_tail(
+ &pmonr->rotation_entry, &__pkg_data(pmonr, istate_pmonrs_lru));
+
+ if (pmonr->limbo_prmid)
+ list_move_tail(&pmonr->limbo_rotation_entry,
+ &__pkg_data(pmonr, ilstate_pmonrs_lru));
+
+ __pmonr__set_istate_summary(pmonr);
+
+}
+
static int intel_cqm_setup_pkg_prmid_pools(u16 pkg_id)
{
int r;
@@ -538,11 +729,11 @@ monr_hrchy_get_next_prmid_summary(struct pmonr *pmonr)
*/
WARN_ON_ONCE(!__pmonr__in_ustate(pmonr));
- if (!list_empty(&__pkg_data(pmonr, free_prmids_pool))) {
+ if (list_empty(&__pkg_data(pmonr, free_prmids_pool))) {
/* Failed to obtain an valid rmid in this package for this
- * monr. In next patches it will transition to (I)state.
- * For now, stay in (U)state (do nothing)..
+ * monr. Use an inherited one.
*/
+ __pmonr__to_istate(pmonr);
} else {
/* Transition to (A)state using free prmid. */
__pmonr__ustate_to_astate(
@@ -796,7 +987,7 @@ static int intel_cqm_event_add(struct perf_event *event, int mode)
__intel_cqm_event_start(event, summary);
/* (I)state pmonrs cannot report occupancy for themselves. */
- return 0;
+ return prmid_summary__is_istate(summary) ? -1 : 0;
}
static void intel_cqm_event_destroy(struct perf_event *event)
diff --git a/arch/x86/events/intel/cqm.h b/arch/x86/events/intel/cqm.h
index 81092f2..22635bc 100644
--- a/arch/x86/events/intel/cqm.h
+++ b/arch/x86/events/intel/cqm.h
@@ -60,11 +60,11 @@ static inline int cqm_prmid_update(struct prmid *prmid);
* The combination of values in sched_rmid and read_rmid indicate the state of
* the associated pmonr (see pmonr comments) as follows:
* pmonr state
- * | (A)state (U)state
+ * | (A)state (IN)state (IL)state (U)state
* ----------------------------------------------------------------------------
- * sched_rmid | pmonr.prmid INVALID_RMID
- * read_rmid | pmonr.prmid INVALID_RMID
- * (or 0)
+ * sched_rmid | pmonr.prmid ancestor.prmid ancestor.prmid INVALID_RMID
+ * read_rmid | pmonr.prmid INVALID_RMID pmonr.limbo_prmid INVALID_RMID
+ * (or 0)
*
* The combination sched_rmid == INVALID_RMID and read_rmid == 0 for (U)state
* denotes that the flag MONR_MON_ACTIVE is set in the monr associated with
@@ -88,6 +88,13 @@ inline bool prmid_summary__is_ustate(union prmid_summary summ)
return summ.sched_rmid == INVALID_RMID;
}
+/* A pmonr in (I)state (either (IN)state or (IL)state. */
+inline bool prmid_summary__is_istate(union prmid_summary summ)
+{
+ return summ.sched_rmid != INVALID_RMID &&
+ summ.sched_rmid != summ.read_rmid;
+}
+
inline bool prmid_summary__is_mon_active(union prmid_summary summ)
{
/* If not in (U)state, then MONR_MON_ACTIVE must be set. */
@@ -98,9 +105,26 @@ inline bool prmid_summary__is_mon_active(union prmid_summary summ)
struct monr;
/* struct pmonr: Node of per-package hierarchy of MONitored Resources.
+ * @ancestor_pmonr: lowest active pmonr whose monr is ancestor of
+ * this pmonr's monr.
+ * @pmonr_deps_head: List of pmonrs without prmid that use
+ * this pmonr's prmid -when in (A)state-.
* @prmid: The prmid of this pmonr -when in (A)state-.
- * @rotation_entry: List entry to attach to astate_pmonrs_lru
- * in pkg_data.
+ * @pmonr_deps_entry: Entry into ancestor's @pmonr_deps_head
+ * -when inheriting, (I)state-.
+ * @limbo_prmid: A prmid previously used by this pmonr and that
+ * has not been reused yet and therefore contain
+ * occupancy that should be counted towards this
+ * pmonr's occupancy.
+ * The limbo_prmid can be reused in the same pmonr
+ * in the next transition to (A) state, even if
+ * the occupancy of @limbo_prmid is not below the
+ * dirty threshold, reducing the need of free
+ * prmids.
+ * @limbo_rotation_entry: List entry to attach to ilstate_pmonrs_lru when
+ * this pmonr is in (IL)state.
+ * @rotation_entry: List entry to attach to either astate_pmonrs_lru
+ * or ilstate_pmonrs_lru in pkg_data.
* @monr: The monr that contains this pmonr.
* @pkg_id: Auxiliar variable with pkg id for this pmonr.
* @prmid_summary_atomic: Atomic accesor to store a union prmid_summary
@@ -112,6 +136,15 @@ struct monr;
* pmonr, the pmonr utilizes the rmid of its ancestor.
* A pmonr is always in one of the following states:
* - (A)ctive: Has @prmid assigned, @ancestor_pmonr must be NULL.
+ * - (I)nherited: The prmid used is "Inherited" from @ancestor_pmonr.
+ * @ancestor_pmonr must be set. @prmid is unused. This is
+ * a super-state composed of two substates:
+ *
+ * - (IL)state: A pmonr in (I)state that has a valid limbo_prmid.
+ * - (IN)state: A pmonr in (I)state with NO valid limbo_prmid.
+ *
+ * When the distintion between the two substates is
+ * no relevant, the pmonr is simply in the (I)state.
* - (U)nused: No @ancestor_pmonr and no @prmid, hence no available
* prmid and no inhering one either. Not in rotation list.
* This state is unschedulable and a prmid
@@ -122,13 +155,41 @@ struct monr;
* The state transitions are:
* (U) : The initial state. Starts there after allocation.
* (U) -> (A): If on first sched (or initialization) pmonr receives a prmid.
+ * (U) -> (I): If on first sched (or initialization) pmonr cannot find a free
+ * prmid and resort to use its ancestor's.
+ * (A) -> (I): On stealing of prmid from pmonr (by rotation logic only).
* (A) -> (U): On destruction of monr.
+ * (I) -> (A): On receiving a free prmid or on reuse of its @limbo_prmid (by
+ * rotation logic only).
+ * (I) -> (U): On destruction of pmonr.
+ *
+ * Note that the (I) -> (A) transition makes monitoring available, but can
+ * introduce error due to cache lines allocated before the transition. Such
+ * error is likely to decrease over time.
+ * When entering an (I) state, the reported count of the event is unavaiable.
*
- * Each pmonr is contained by a monr.
+ * Each pmonr is contained by a monr. Each monr forms a system-wide hierarchy
+ * that is used by the pmrs to find ancestors and dependants. The per-package
+ * hierarchy spanned by the pmrs follows the monr hierarchy except by
+ * collapsing the nodes in (I)state into a super-node that contains an (A)state
+ * pmonr and all of its dependants (pmonr in pmonr_deps_head).
*/
struct pmonr {
- struct prmid *prmid;
+ /* If set, pmonr is in (I)state. */
+ struct pmonr *ancestor_pmonr;
+
+ union{
+ struct { /* (A)state variables. */
+ struct list_head pmonr_deps_head;
+ struct prmid *prmid;
+ };
+ struct { /* (I)state variables. */
+ struct list_head pmonr_deps_entry;
+ struct prmid *limbo_prmid;
+ struct list_head limbo_rotation_entry;
+ };
+ };
struct monr *monr;
struct list_head rotation_entry;
@@ -146,10 +207,17 @@ struct pmonr {
* XXX: Make it an array of prmids.
* @free_prmid_pool: Free prmids.
* @active_prmid_pool: prmids associated with a (A)state pmonr.
+ * @pmonr_limbo_prmid_pool: limbo prmids referenced by the limbo_prmid of a
+ * pmonr in (I)state.
* @nopmonr_limbo_prmid_pool: prmids in limbo state that are not referenced
* by a pmonr.
* @astate_pmonrs_lru: pmonrs in (A)state. LRU in increasing order of
* pmonr.last_enter_astate.
+ * @istate_pmonrs_lru: pmors In (I)state with no limbo_prmid. LRU in
+ * increasing order of pmonr.last_enter_istate.
+ * @ilsate_pmonrs_lru: pmonrs in (IL)state, these pmonrs have a valid
+ * limbo_prmid. It's a subset of istate_pmonrs_lru.
+ * Sorted increasingly by pmonr.last_enter_istate.
* @pkg_data_mutex: Hold for stability when modifying pmonrs
* hierarchy.
* @pkg_data_lock: Hold to protect variables that may be accessed
@@ -171,9 +239,13 @@ struct pkg_data {
/* Can be modified during task switch with (U)state -> (A)state. */
struct list_head active_prmids_pool;
/* Only modified during rotation logic and deletion. */
+ struct list_head pmonr_limbo_prmids_pool;
struct list_head nopmonr_limbo_prmids_pool;
struct list_head astate_pmonrs_lru;
+ /* Superset of ilstate_pmonrs_lru. */
+ struct list_head istate_pmonrs_lru;
+ struct list_head ilstate_pmonrs_lru;
struct mutex pkg_data_mutex;
raw_spinlock_t pkg_data_lock;
--
2.8.0.rc3.226.g39d4020
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2016-04-29 07:00 +0200 |
| Subject | [PATCH 12/32] perf/x86/intel/cqm: add per-package RMID rotation |
| Message-ID | <rt8H1-8vJ-27@gated-at.bofh.it> |
| In reply to | #1390741 |
This version of RMID rotation improves over original one by:
1. Being per-package. No need for IPIs to test for occupancy.
2. Since the monr hierarchy removed the potential conflicts between
events, the new RMID rotation logic does not need to check and
resolve conflicts.
3. No need to mantain an unused RMID as rotation_rmid, effectively
freeing one RMID per package.
4. Guarantee that monitored events and cgroups with a valid RMID keep
the RMID for an user configurable time: __cqm_min_mon_slice ms.
Previously, it was likely to receive a RMID in one execution of the
rotation logic just to have it removed in the next. That was
specially problematic in the presence of events conflict
(ie. cgroup events and thread events in a descendant cgroup).
5. Do not increase the dirty threshold unless strictly necessary to make
progress. Previous version simultaneously stole RMIDs and increased
the dirty threshold (the maximum number of cache lines with spurious
occupancy associated with a "clean" RMID). This version makes sure
that increasing the dirty threshold is the only way to make progress
in the RMID rotation (the case when too many RMID in limbo do not
drop occupancy despite having spent enough time in limbo) before
increasing the threshold.
This change reduces spurious occupancy as a source of error.
6. Do not steal RMIDs unnecesarily. Thanks to a more detailed
bookeeping, this patch guarantees that the number of RMIDs in limbo
do not exceed the number of RMIDs needed by pmonrs currently waiting
for an RMID.
7. Reutilize dirty limbo RMIDs when appropriate. In this new version, a
stolen RMID remains referenced by its former pmonr owner until it is
reutilized by another pmonr or it is moved from limbo into the pool
of free RMIDs.
These RMIDs that are referenced and in limbo are not written into the
MSR_IA32_PQR_ASSOC msr, therefore, they have the chance to drop
occupancy as any other limbo RMID. If the pmonr with a limbo RMID is
to be activated, then it reuses its former RMID even if its still
dirty. The occupancy attributed to that RMID is part of the pmonr
occupancy and therefore reusing the RMID even when dirty decreases
the error of the read.
This feature decreases the negative impact of RMIDs that do not drop
occupancy in the efficiency of the rotation logic.
For an user perspective, the behavior of the new rotation logic is
controlled by SLO type parameters:
__cqm_min_mon_slice : Minimum time a monr is to be monitored
before being eligible by rotation logic to loss any of its RMIDs.
__cqm_max_wait_mon : Maximum time a monr can be deactivated
before forcing rotation logic to be more aggresive (stealing more
RMIDs per iteration).
__cqm_min_progress_rate: Minimum number of pmonrs that must be
activated per second to consider that rotation logic's progress
is acceptable.
Since the minimum progress rate is a SLO, the magnitude of the rotation
period (the rtimer_interval_ms) do not control the speed of RMID rotation,
it only controls the frequency at which rotation logic is executed.
Reviewed-by: Stephane Eranian <eranian@google.com>
Signed-off-by: David Carrillo-Cisneros <davidcc@google.com>
---
arch/x86/events/intel/cqm.c | 727 ++++++++++++++++++++++++++++++++++++++++++++
arch/x86/events/intel/cqm.h | 59 +++-
2 files changed, 784 insertions(+), 2 deletions(-)
diff --git a/arch/x86/events/intel/cqm.c b/arch/x86/events/intel/cqm.c
index caf7152..31f0fd6 100644
--- a/arch/x86/events/intel/cqm.c
+++ b/arch/x86/events/intel/cqm.c
@@ -235,9 +235,14 @@ static int pkg_data_init_cpu(int cpu)
INIT_LIST_HEAD(&pkg_data->istate_pmonrs_lru);
INIT_LIST_HEAD(&pkg_data->ilstate_pmonrs_lru);
+ pkg_data->nr_instate_pmonrs = 0;
+ pkg_data->nr_ilstate_pmonrs = 0;
+
mutex_init(&pkg_data->pkg_data_mutex);
raw_spin_lock_init(&pkg_data->pkg_data_lock);
+ INIT_DELAYED_WORK(
+ &pkg_data->rotation_work, intel_cqm_rmid_rotation_work);
/* XXX: Chose randomly*/
pkg_data->rotation_cpu = cpu;
@@ -295,6 +300,10 @@ static struct pmonr *pmonr_alloc(int cpu)
pmonr->monr = NULL;
INIT_LIST_HEAD(&pmonr->rotation_entry);
+ pmonr->last_enter_istate = 0;
+ pmonr->last_enter_astate = 0;
+ pmonr->nr_enter_istate = 0;
+
pmonr->pkg_id = topology_physical_package_id(cpu);
summary.sched_rmid = INVALID_RMID;
summary.read_rmid = INVALID_RMID;
@@ -346,6 +355,8 @@ __pmonr__finish_to_astate(struct pmonr *pmonr, struct prmid *prmid)
pmonr->prmid = prmid;
+ pmonr->last_enter_astate = jiffies;
+
list_move_tail(
&prmid->pool_entry, &__pkg_data(pmonr, active_prmids_pool));
list_move_tail(
@@ -373,6 +384,8 @@ __pmonr__instate_to_astate(struct pmonr *pmonr, struct prmid *prmid)
*/
WARN_ON_ONCE(pmonr->limbo_prmid);
+ __pkg_data(pmonr, nr_instate_pmonrs)--;
+
/* Do not depend on ancestor_pmonr anymore. Make it (A)state. */
ancestor = pmonr->ancestor_pmonr;
list_del_init(&pmonr->pmonr_deps_entry);
@@ -394,6 +407,28 @@ __pmonr__instate_to_astate(struct pmonr *pmonr, struct prmid *prmid)
}
}
+/*
+ * Transition from (IL)state to (A)state.
+ */
+static inline void
+__pmonr__ilstate_to_astate(struct pmonr *pmonr)
+{
+ struct prmid *prmid;
+
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+ WARN_ON_ONCE(!pmonr->limbo_prmid);
+
+ prmid = pmonr->limbo_prmid;
+ pmonr->limbo_prmid = NULL;
+ list_del_init(&pmonr->limbo_rotation_entry);
+
+ __pkg_data(pmonr, nr_ilstate_pmonrs)--;
+ __pkg_data(pmonr, nr_instate_pmonrs)++;
+ list_del_init(&prmid->pool_entry);
+
+ __pmonr__instate_to_astate(pmonr, prmid);
+}
+
static inline void
__pmonr__ustate_to_astate(struct pmonr *pmonr, struct prmid *prmid)
{
@@ -485,7 +520,9 @@ __pmonr__to_ustate(struct pmonr *pmonr)
pmonr->limbo_prmid = NULL;
list_del_init(&pmonr->limbo_rotation_entry);
+ __pkg_data(pmonr, nr_ilstate_pmonrs)--;
} else {
+ __pkg_data(pmonr, nr_instate_pmonrs)--;
}
pmonr->ancestor_pmonr = NULL;
} else {
@@ -542,6 +579,9 @@ __pmonr__to_istate(struct pmonr *pmonr)
__pmonr__move_dependants(pmonr, ancestor);
list_move_tail(&pmonr->limbo_prmid->pool_entry,
&__pkg_data(pmonr, pmonr_limbo_prmids_pool));
+ __pkg_data(pmonr, nr_ilstate_pmonrs)++;
+ } else {
+ __pkg_data(pmonr, nr_instate_pmonrs)++;
}
pmonr->ancestor_pmonr = ancestor;
@@ -554,10 +594,51 @@ __pmonr__to_istate(struct pmonr *pmonr)
list_move_tail(&pmonr->limbo_rotation_entry,
&__pkg_data(pmonr, ilstate_pmonrs_lru));
+ pmonr->last_enter_istate = jiffies;
+ pmonr->nr_enter_istate++;
+
__pmonr__set_istate_summary(pmonr);
}
+static inline void
+__pmonr__ilstate_to_instate(struct pmonr *pmonr)
+{
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+
+ list_move_tail(&pmonr->limbo_prmid->pool_entry,
+ &__pkg_data(pmonr, free_prmids_pool));
+ pmonr->limbo_prmid = NULL;
+
+ __pkg_data(pmonr, nr_ilstate_pmonrs)--;
+ __pkg_data(pmonr, nr_instate_pmonrs)++;
+
+ list_del_init(&pmonr->limbo_rotation_entry);
+ __pmonr__set_istate_summary(pmonr);
+}
+
+/* Count all limbo prmids, including the ones still attached to pmonrs.
+ * Maximum number of prmids is fixed by hw and generally small.
+ */
+static int count_limbo_prmids(struct pkg_data *pkg_data)
+{
+ unsigned int c = 0;
+ struct prmid *prmid;
+
+ lockdep_assert_held(&pkg_data->pkg_data_mutex);
+
+ list_for_each_entry(
+ prmid, &pkg_data->pmonr_limbo_prmids_pool, pool_entry) {
+ c++;
+ }
+ list_for_each_entry(
+ prmid, &pkg_data->nopmonr_limbo_prmids_pool, pool_entry) {
+ c++;
+ }
+
+ return c;
+}
+
static int intel_cqm_setup_pkg_prmid_pools(u16 pkg_id)
{
int r;
@@ -871,8 +952,652 @@ static bool __match_event(struct perf_event *a, struct perf_event *b)
return false;
}
+/*
+ * Try to reuse limbo prmid's for pmonrs at the front of ilstate_pmonrs_lru.
+ */
+static int __try_reuse_ilstate_pmonrs(struct pkg_data *pkg_data)
+{
+ int reused = 0;
+ struct pmonr *pmonr;
+
+ lockdep_assert_held(&pkg_data->pkg_data_mutex);
+ lockdep_assert_held(&pkg_data->pkg_data_lock);
+
+ while ((pmonr = list_first_entry_or_null(
+ &pkg_data->istate_pmonrs_lru, struct pmonr, rotation_entry))) {
+
+ if (__pmonr__in_instate(pmonr))
+ break;
+ __pmonr__ilstate_to_astate(pmonr);
+ reused++;
+ }
+ return reused;
+}
+
+static int try_reuse_ilstate_pmonrs(struct pkg_data *pkg_data)
+{
+ int reused;
+ unsigned long flags;
+#ifdef CONFIG_LOCKDEP
+ u16 pkg_id = topology_physical_package_id(smp_processor_id());
+#endif
+
+ lockdep_assert_held(&pkg_data->pkg_data_mutex);
+
+ raw_spin_lock_irqsave_nested(&pkg_data->pkg_data_lock, flags, pkg_id);
+ reused = __try_reuse_ilstate_pmonrs(pkg_data);
+ raw_spin_unlock_irqrestore(&pkg_data->pkg_data_lock, flags);
+ return reused;
+}
+
+
+/*
+ * A monr is only readable when all it's used pmonrs have a RMID.
+ * Therefore, the time a monr entered (A)state is the maximum of the
+ * last_enter_astate times for all (A)state pmonrs if no pmonr is in (I)state.
+ * A monr with any pmonr in (I)state has no entered (A)state.
+ * Returns monr_enter_astate time if available, otherwise min_inh_pkg is
+ * set to the smallest pkg_id where the monr's pmnor is in (I)state and
+ * the return value is undefined.
+ */
+static unsigned long
+__monr__last_enter_astate(struct monr *monr, int *min_inh_pkg)
+{
+ struct pkg_data *pkg_data;
+ u16 pkg_id;
+ unsigned long flags, astate_time = 0;
+
+ *min_inh_pkg = -1;
+ cqm_pkg_id_for_each_online(pkg_id) {
+ struct pmonr *pmonr;
+
+ if (min_inh_pkg >= 0)
+ break;
+
+ raw_spin_lock_irqsave_nested(
+ &pkg_data->pkg_data_lock, flags, pkg_id);
+
+ pmonr = monr->pmonrs[pkg_id];
+ if (__pmonr__in_istate(pmonr) && min_inh_pkg < 0)
+ *min_inh_pkg = pkg_id;
+ else if (__pmonr__in_astate(pmonr) &&
+ astate_time < pmonr->last_enter_astate)
+ astate_time = pmonr->last_enter_astate;
+
+ raw_spin_unlock_irqrestore(&pkg_data->pkg_data_lock, flags);
+ }
+ return astate_time;
+}
+
+/*
+ * Steal as many rmids as possible.
+ * Transition pmonrs that have stayed at least __cqm_min_mon_slice in
+ * (A)state to (I)state.
+ */
+static inline int
+__try_steal_active_pmonrs(
+ struct pkg_data *pkg_data, unsigned int max_to_steal)
+{
+ struct pmonr *pmonr, *tmp;
+ int nr_stolen = 0, min_inh_pkg;
+ u16 pkg_id = topology_physical_package_id(smp_processor_id());
+ unsigned long flags, monr_astate_end_time, now = jiffies;
+ struct list_head *alist = &pkg_data->astate_pmonrs_lru;
+
+ lockdep_assert_held(&pkg_data->pkg_data_mutex);
+
+ /* pmonrs don't leave astate outside of rotation logic.
+ * The pkg mutex protects against the pmonr leaving
+ * astate_pmonrs_lru. The raw_spin_lock protects these list
+ * operations from list insertions at tail coming from the
+ * sched logic ( (U)state -> (A)state )
+ */
+ raw_spin_lock_irqsave_nested(&pkg_data->pkg_data_lock, flags, pkg_id);
+
+ pmonr = list_first_entry(alist, struct pmonr, rotation_entry);
+ WARN_ON_ONCE(pmonr != monr_hrchy_root->pmonrs[pkg_id]);
+ WARN_ON_ONCE(pmonr->pkg_id != pkg_id);
+
+ list_for_each_entry_safe_continue(pmonr, tmp, alist, rotation_entry) {
+ bool steal_rmid = false;
+
+ WARN_ON_ONCE(!__pmonr__in_astate(pmonr));
+ WARN_ON_ONCE(pmonr->pkg_id != pkg_id);
+
+ raw_spin_unlock_irqrestore(&pkg_data->pkg_data_lock, flags);
+
+ monr_astate_end_time =
+ __monr__last_enter_astate(pmonr->monr, &min_inh_pkg) +
+ __cqm_min_mon_slice;
+
+ /* pmonr in this pkg is supposed to be in (A)state. */
+ WARN_ON_ONCE(min_inh_pkg == pkg_id);
+
+ /* Steal a pmonr if:
+ * 1) Any pmonr in a pkg with pkg_id < local pkg_id is
+ * in (I)state.
+ * 2) It's monr has been active for enough time.
+ * Note that since the min_inh_pkg for a monr cannot decrease
+ * while the monr is not active, then the monr eventually will
+ * become active again despite the stealing of pmonrs in pkgs
+ * with id larger than min_inh_pkg.
+ */
+ if (min_inh_pkg >= 0 && min_inh_pkg < pkg_id)
+ steal_rmid = true;
+ if (min_inh_pkg < 0 && monr_astate_end_time <= now)
+ steal_rmid = true;
+
+ raw_spin_lock_irqsave_nested(
+ &pkg_data->pkg_data_lock, flags, pkg_id);
+ if (!steal_rmid)
+ continue;
+
+ __pmonr__to_istate(pmonr);
+ nr_stolen++;
+ if (nr_stolen == max_to_steal)
+ break;
+ }
+
+ raw_spin_unlock_irqrestore(&pkg_data->pkg_data_lock, flags);
+
+ return nr_stolen;
+}
+
+/* It will remove the prmid from the list its attached, if used. */
+static inline int __try_use_free_prmid(struct pkg_data *pkg_data,
+ struct prmid *prmid, bool *succeed)
+{
+ struct pmonr *pmonr;
+ int nr_activated = 0;
+
+ lockdep_assert_held(&pkg_data->pkg_data_mutex);
+ lockdep_assert_held(&pkg_data->pkg_data_lock);
+
+ *succeed = false;
+ nr_activated += __try_reuse_ilstate_pmonrs(pkg_data);
+ pmonr = list_first_entry_or_null(&pkg_data->istate_pmonrs_lru,
+ struct pmonr, rotation_entry);
+ if (!pmonr)
+ return nr_activated;
+ WARN_ON_ONCE(__pmonr__in_ilstate(pmonr));
+ WARN_ON_ONCE(!__pmonr__in_instate(pmonr));
+
+ /* the state transition function will move the prmid to
+ * the active lru list.
+ */
+ __pmonr__instate_to_astate(pmonr, prmid);
+ nr_activated++;
+ *succeed = true;
+ return nr_activated;
+}
+
+static inline int __try_use_free_prmids(struct pkg_data *pkg_data)
+{
+ struct prmid *prmid, *tmp_prmid;
+ unsigned long flags;
+ int nr_activated = 0;
+ bool succeed;
+#ifdef CONFIG_DEBUG_SPINLOCK
+ u16 pkg_id = topology_physical_package_id(smp_processor_id());
+#endif
+
+ lockdep_assert_held(&pkg_data->pkg_data_mutex);
+ /* Lock protects free_prmids_pool, istate_pmonrs_lru and
+ * the monr hrchy.
+ */
+ raw_spin_lock_irqsave_nested(&pkg_data->pkg_data_lock, flags, pkg_id);
+
+ list_for_each_entry_safe(prmid, tmp_prmid,
+ &pkg_data->free_prmids_pool, pool_entry) {
+
+ /* Removes the free prmid if used. */
+ nr_activated += __try_use_free_prmid(pkg_data,
+ prmid, &succeed);
+ }
+
+ nr_activated += __try_reuse_ilstate_pmonrs(pkg_data);
+ raw_spin_unlock_irqrestore(&pkg_data->pkg_data_lock, flags);
+
+ return nr_activated;
+}
+
+/* Update prmid's of pmonrs in ilstate. To mantain fairness of rotation
+ * logic, Try to activate (IN)state pmonrs with recovered prmids when
+ * possible rather than simply adding them to free rmids list. This prevents,
+ * ustate pmonrs (pmonrs that haven't wait in istate_pmonrs_lru) to obtain
+ * the newly available RMIDs before those waiting in queue.
+ */
+static inline int
+__try_free_ilstate_prmids(struct pkg_data *pkg_data,
+ unsigned int cqm_threshold,
+ unsigned int *min_occupancy_dirty)
+{
+ struct pmonr *pmonr, *tmp_pmonr, *istate_pmonr;
+ struct prmid *prmid;
+ unsigned long flags;
+ u64 val;
+ bool succeed;
+ int ret, nr_activated = 0;
+#ifdef CONFIG_LOCKDEP
+ u16 pkg_id = topology_physical_package_id(smp_processor_id());
+#endif
+
+ lockdep_assert_held(&pkg_data->pkg_data_mutex);
+
+ WARN_ON_ONCE(try_reuse_ilstate_pmonrs(pkg_data));
+
+ /* No need to acquire pkg lock to iterate over ilstate_pmonrs_lru
+ * since only rotation logic modifies it.
+ */
+ list_for_each_entry_safe(
+ pmonr, tmp_pmonr,
+ &pkg_data->ilstate_pmonrs_lru, limbo_rotation_entry) {
+
+ if (WARN_ON_ONCE(list_empty(&pkg_data->istate_pmonrs_lru)))
+ return nr_activated;
+
+ istate_pmonr = list_first_entry(&pkg_data->istate_pmonrs_lru,
+ struct pmonr, rotation_entry);
+
+ if (pmonr == istate_pmonr) {
+ raw_spin_lock_irqsave_nested(
+ &pkg_data->pkg_data_lock, flags, pkg_id);
+
+ nr_activated++;
+ __pmonr__ilstate_to_astate(pmonr);
+
+ raw_spin_unlock_irqrestore(
+ &pkg_data->pkg_data_lock, flags);
+ continue;
+ }
+
+ ret = __cqm_prmid_update(pmonr->limbo_prmid,
+ __rmid_min_update_time);
+ if (WARN_ON_ONCE(ret < 0))
+ continue;
+
+ val = atomic64_read(&pmonr->limbo_prmid->last_read_value);
+ if (val > cqm_threshold) {
+ if (val < *min_occupancy_dirty)
+ *min_occupancy_dirty = val;
+ continue;
+ }
+
+ raw_spin_lock_irqsave_nested(
+ &pkg_data->pkg_data_lock, flags, pkg_id);
+
+ prmid = pmonr->limbo_prmid;
+
+ /* moves the prmid to free_prmids_pool. */
+ __pmonr__ilstate_to_instate(pmonr);
+
+ /* Do not affect ilstate_pmonrs_lru.
+ * If succeeds, prmid will end in active_prmids_pool,
+ * otherwise, stays in free_prmids_pool where the
+ * ilstate_to_instate transition left it.
+ */
+ nr_activated += __try_use_free_prmid(pkg_data,
+ prmid, &succeed);
+
+ raw_spin_unlock_irqrestore(&pkg_data->pkg_data_lock, flags);
+ }
+ return nr_activated;
+}
+
+/* Update limbo prmid's no associated to a pmonr. To mantain fairness of
+ * rotation logic, Try to activate (IN)state pmonrs with recovered prmids when
+ * possible rather than simply adding them to free rmids list. This prevents,
+ * ustate pmonrs (pmonrs that haven't wait in istate_pmonrs_lru) to obtain
+ * the newly available RMIDs before those waiting in queue.
+ */
+static inline int
+__try_free_limbo_prmids(struct pkg_data *pkg_data,
+ unsigned int cqm_threshold,
+ unsigned int *min_occupancy_dirty)
+{
+ struct prmid *prmid, *tmp_prmid;
+ unsigned long flags;
+ bool succeed;
+ int ret, nr_activated = 0;
+
+#ifdef CONFIG_LOCKDEP
+ u16 pkg_id = topology_physical_package_id(smp_processor_id());
+#endif
+ u64 val;
+
+ lockdep_assert_held(&pkg_data->pkg_data_mutex);
+
+ list_for_each_entry_safe(
+ prmid, tmp_prmid,
+ &pkg_data->nopmonr_limbo_prmids_pool, pool_entry) {
+
+ /* If min update time is good enough for user, it is good
+ * enough for rotation.
+ */
+ ret = __cqm_prmid_update(prmid, __rmid_min_update_time);
+ if (WARN_ON_ONCE(ret < 0))
+ continue;
+
+ val = atomic64_read(&prmid->last_read_value);
+ if (val > cqm_threshold) {
+ if (val < *min_occupancy_dirty)
+ *min_occupancy_dirty = val;
+ continue;
+ }
+ raw_spin_lock_irqsave_nested(
+ &pkg_data->pkg_data_lock, flags, pkg_id);
+
+ nr_activated = __try_use_free_prmid(pkg_data, prmid, &succeed);
+ if (!succeed)
+ list_move_tail(&prmid->pool_entry,
+ &pkg_data->free_prmids_pool);
+
+ raw_spin_unlock_irqrestore(&pkg_data->pkg_data_lock, flags);
+ }
+ return nr_activated;
+}
+
+/*
+ * Activate (I)state pmonrs.
+ *
+ * @min_occupancy_dirty: pointer to store the minimum occupancy of any
+ * dirty prmid.
+ *
+ * Try to activate as many pmonrs as possible before utilizing limbo prmids
+ * pointed by ilstate pmonrs in order to minimize the number of dirty rmids
+ * that move to other pmonr when cqm_threshold > 0.
+ */
+static int __try_activate_istate_pmonrs(
+ struct pkg_data *pkg_data, unsigned int cqm_threshold,
+ unsigned int *min_occupancy_dirty)
+{
+ int nr_activated = 0;
+
+ lockdep_assert_held(&pkg_data->pkg_data_mutex);
+
+ /* Start reusing limbo prmids no pointed by any ilstate pmonr. */
+ nr_activated += __try_free_limbo_prmids(pkg_data, cqm_threshold,
+ min_occupancy_dirty);
+
+ /* Try to use newly available free prmids */
+ nr_activated += __try_use_free_prmids(pkg_data);
+
+ /* Continue reusing limbo prmids pointed by a ilstate pmonr. */
+ nr_activated += __try_free_ilstate_prmids(pkg_data, cqm_threshold,
+ min_occupancy_dirty);
+ /* Try to use newly available free prmids */
+ nr_activated += __try_use_free_prmids(pkg_data);
+
+ WARN_ON_ONCE(try_reuse_ilstate_pmonrs(pkg_data));
+ return nr_activated;
+}
+
+/* Number of pmonrs that have been in (I)state for at least min_wait_jiffies.
+ * XXX: Use rcu to access to istate_pmonrs_lru.
+ */
+static int
+count_istate_pmonrs(struct pkg_data *pkg_data,
+ unsigned int min_wait_jiffies, bool exclude_limbo)
+{
+ unsigned long flags;
+ unsigned int c = 0;
+ struct pmonr *pmonr;
+#ifdef CONFIG_DEBUG_SPINLOCK
+ u16 pkg_id = topology_physical_package_id(smp_processor_id());
+#endif
+
+ lockdep_assert_held(&pkg_data->pkg_data_mutex);
+
+ raw_spin_lock_irqsave_nested(&pkg_data->pkg_data_lock, flags, pkg_id);
+ list_for_each_entry(
+ pmonr, &pkg_data->istate_pmonrs_lru, rotation_entry) {
+
+ if (jiffies - pmonr->last_enter_istate < min_wait_jiffies)
+ break;
+
+ WARN_ON_ONCE(!__pmonr__in_istate(pmonr));
+ if (exclude_limbo && __pmonr__in_ilstate(pmonr))
+ continue;
+ c++;
+ }
+ raw_spin_unlock_irqrestore(&pkg_data->pkg_data_lock, flags);
+
+ return c;
+}
+
+static inline int
+read_nr_instate_pmonrs(struct pkg_data *pkg_data, u16 pkg_id) {
+ unsigned long flags;
+ int n;
+
+ raw_spin_lock_irqsave_nested(&pkg_data->pkg_data_lock, flags, pkg_id);
+ n = READ_ONCE(cqm_pkgs_data[pkg_id]->nr_instate_pmonrs);
+ raw_spin_unlock_irqrestore(&pkg_data->pkg_data_lock, flags);
+ WARN_ON_ONCE(n < 0);
+ return n;
+}
+
+/*
+ * Rotate RMIDs among rpgks.
+ *
+ * For reads to be meaningful valid rmids had to be programmed for
+ * enough time to capture enough instances of cache allocation/retirement
+ * to yield useful occupancy values. The approach to handle that problem
+ * is to guarantee that every pmonr will spend at least T time in (A)state
+ * when such transition has occurred and hope that T is long enough.
+ *
+ * The hardware retains occupancy for 'old' tags, even after changing rmid
+ * for a task/cgroup. To workaround this problem, we keep retired rmids
+ * as limbo in each pmonr and use their occupancy. Also we prefer reusing
+ * such limbo rmids rather than free ones since their residual occupancy
+ * is valid occupancy for the task/cgroup.
+ *
+ * Rotation works by taking away an RMID from a group (the old RMID),
+ * and assigning the free RMID to another group (the new RMID). We must
+ * then wait for the old RMID to not be used (no cachelines tagged).
+ * This ensure that all cachelines are tagged with 'active' RMIDs. At
+ * this point we can start reading values for the new RMID and treat the
+ * old RMID as the free RMID for the next rotation.
+ */
+static void
+__intel_cqm_rmid_rotate(struct pkg_data *pkg_data,
+ unsigned int nr_max_limbo,
+ unsigned int nr_min_activated)
+{
+ int nr_instate, nr_to_steal, nr_stolen, nr_slo_violated;
+ int limbo_cushion = 0;
+ unsigned int cqm_threshold = 0, min_occupancy_dirty;
+ u16 pkg_id = topology_physical_package_id(smp_processor_id());
+
+ /*
+ * To avoid locking the process, keep track of pmonrs that
+ * are activated during this execution of rotaton logic, so
+ * we don't have to rely on the state of the pmonrs lists
+ * to estimate progress, that can be modified during
+ * creation and destruction of events and cgroups.
+ */
+ int nr_activated = 0;
+
+ mutex_lock_nested(&pkg_data->pkg_data_mutex, pkg_id);
+
+ /*
+ * Since ilstates are created only during stealing or destroying pmonrs,
+ * but destroy requires pkg_data_mutex, then it is only necessary to
+ * try to reuse ilstate once per call. Furthermore, new ilstates during
+ * iteration in rotation logic is an error.
+ */
+ nr_activated += try_reuse_ilstate_pmonrs(pkg_data);
+
+again:
+ nr_stolen = 0;
+ min_occupancy_dirty = UINT_MAX;
+ /*
+ * Three types of actions are taken in rotation logic:
+ * 1) Try to activate pmonrs using limbo RMIDs.
+ * 2) Steal more RMIDs. Ideally the number of RMIDs in limbo equals
+ * the number of pmonrs in (I)state plus the limbo_cushion aimed to
+ * compensate for limbo RMIDs that do no drop occupancy fast enough.
+ * The actual number stolen is constrained
+ * prevent having more than nr_max_limbo RMIDs in limbo.
+ * 3) Increase cqm_threshold so even RMIDs with residual occupancy
+ * are utilized to activate (I)state primds. Doing so increases the
+ * error in the reported value in a way undetectable to the user, so
+ * it is left as a last resource.
+ */
+
+ /* Verify all available ilimbo where activated where they
+ * were supposed to.
+ */
+ WARN_ON_ONCE(try_reuse_ilstate_pmonrs(pkg_data) > 0);
+
+ /* Activate all pmonrs that we can by recycling rmids in limbo */
+ nr_activated += __try_activate_istate_pmonrs(
+ pkg_data, cqm_threshold, &min_occupancy_dirty);
+
+ /* Count nr of pmonrs that are inherited and do not have limbo_prmid */
+ nr_instate = read_nr_instate_pmonrs(pkg_data, pkg_id);
+ WARN_ON_ONCE(nr_instate < 0);
+ /*
+ * If no pmonr needs rmid, then it's time to let go. pmonrs in ilimbo
+ * are not counted since the limbo_prmid can be reused, once its time
+ * to activate them.
+ */
+ if (nr_instate == 0)
+ goto exit;
+
+ WARN_ON_ONCE(!list_empty(&pkg_data->free_prmids_pool));
+ WARN_ON_ONCE(try_reuse_ilstate_pmonrs(pkg_data) > 0);
+
+ /* There are still pmonrs waiting for RMID, check if the SLO about
+ * _cqm_max_wait_mon has been violated. If so, use a more
+ * aggresive version of RMID stealing and reutilization.
+ */
+ nr_slo_violated = count_istate_pmonrs(
+ pkg_data, msecs_to_jiffies(__cqm_max_wait_mon), false);
+
+ /* First measure against SLO violation is to increase number of stolen
+ * RMIDs beyond the number of pmonrs waiting for RMID. The magnitud of
+ * the limbo_cushion is proportional to nr_slo_violated (but
+ * arbitarily weighthed).
+ */
+ if (nr_slo_violated)
+ limbo_cushion = (nr_slo_violated + 1) / 2;
+
+ /*
+ * Need more free rmids. Steal RMIDs from active pmonrs and place them
+ * into limbo lru. Steal enough to have high chances that eventually
+ * occupancy of enough RMIDs in limbo will drop enough to be reused
+ * (the limbo_cushion).
+ */
+ nr_to_steal = min(nr_instate + limbo_cushion,
+ max(0, (int)nr_max_limbo -
+ count_limbo_prmids(pkg_data)));
+
+ if (nr_to_steal)
+ nr_stolen = __try_steal_active_pmonrs(pkg_data, nr_to_steal);
+
+ /* Already stole as many as possible, finish if no SLO violations. */
+ if (!nr_slo_violated)
+ goto exit;
+
+ /*
+ * There are SLO violations due to recycling RMIDs not progressing
+ * fast enough. Possible (non-exclusive) causal factors are:
+ * 1) Too many RMIDs in limbo do not drop occupancy despite having
+ * spent a "reasonable" time in limbo lru.
+ * 2) RMIDs in limbo have not been for long enough to have drop
+ * occupancy, but they will within "reasonable" time.
+ *
+ * If (2) only, it is ok to wait, since eventually the rmids
+ * will rotate. If (1), there is a danger of being stuck, in that case
+ * the dirty threshold, cqm_threshold, must be increased.
+ * The notion of "reasonable" time is ambiguous since the more SLOs
+ * violations, the more urgent it is to rotate. For now just try
+ * to guarantee any progress is made (activate at least one prmid
+ * with SLO violated).
+ */
+
+ /* Using the minimum observed occupancy in dirty rmids guarantees to
+ * to recover at least one rmid per iteration. Check if constrainst
+ * would allow to use such threshold, otherwise makes no sense to
+ * retry.
+ */
+ if (nr_activated < nr_min_activated && min_occupancy_dirty <=
+ READ_ONCE(__intel_cqm_max_threshold) / cqm_l3_scale) {
+
+ cqm_threshold = min_occupancy_dirty;
+ goto again;
+ }
+exit:
+ mutex_unlock(&pkg_data->pkg_data_mutex);
+}
+
static struct pmu intel_cqm_pmu;
+/* Rotation only needs to be run when there is any pmonr in (I)state. */
+static bool intel_cqm_need_rotation(u16 pkg_id)
+{
+
+ struct pkg_data *pkg_data;
+ bool need_rot;
+
+ pkg_data = cqm_pkgs_data[pkg_id];
+
+ mutex_lock_nested(&pkg_data->pkg_data_mutex, pkg_id);
+ /* Rotation is needed if prmids in limbo need to be recycled or if
+ * there are pmonrs in (I)state.
+ */
+ need_rot = !list_empty(&pkg_data->nopmonr_limbo_prmids_pool) ||
+ !list_empty(&pkg_data->istate_pmonrs_lru);
+
+ mutex_unlock(&pkg_data->pkg_data_mutex);
+ return need_rot;
+}
+
+/*
+ * Schedule rotation in one package.
+ */
+static void __intel_cqm_schedule_rotation_for_pkg(u16 pkg_id)
+{
+ struct pkg_data *pkg_data;
+ unsigned long delay;
+
+ delay = msecs_to_jiffies(intel_cqm_pmu.hrtimer_interval_ms);
+ pkg_data = cqm_pkgs_data[pkg_id];
+ schedule_delayed_work_on(
+ pkg_data->rotation_cpu, &pkg_data->rotation_work, delay);
+}
+
+/*
+ * Schedule rotation and rmid's timed update in all packages.
+ * Reescheduling will stop when no longer needed.
+ */
+static void intel_cqm_schedule_work_all_pkgs(void)
+{
+ int pkg_id;
+
+ cqm_pkg_id_for_each_online(pkg_id)
+ __intel_cqm_schedule_rotation_for_pkg(pkg_id);
+}
+
+static void intel_cqm_rmid_rotation_work(struct work_struct *work)
+{
+ struct pkg_data *pkg_data = container_of(
+ to_delayed_work(work), struct pkg_data, rotation_work);
+ /* Allow max 25% of RMIDs to be in limbo. */
+ unsigned int max_limbo_rmids = max(1u, (pkg_data->max_rmid + 1) / 4);
+ unsigned int min_activated = max(1u, (intel_cqm_pmu.hrtimer_interval_ms
+ * __cqm_min_progress_rate) / 1000);
+ u16 pkg_id = topology_physical_package_id(pkg_data->rotation_cpu);
+
+ WARN_ON_ONCE(pkg_data != cqm_pkgs_data[pkg_id]);
+
+ __intel_cqm_rmid_rotate(pkg_data, max_limbo_rmids, min_activated);
+
+ if (intel_cqm_need_rotation(pkg_id))
+ __intel_cqm_schedule_rotation_for_pkg(pkg_id);
+}
+
/*
* Find a group and setup RMID.
*
@@ -1099,6 +1824,8 @@ static int intel_cqm_event_init(struct perf_event *event)
mutex_unlock(&cqm_mutex);
+ intel_cqm_schedule_work_all_pkgs();
+
return 0;
}
diff --git a/arch/x86/events/intel/cqm.h b/arch/x86/events/intel/cqm.h
index 22635bc..b0e1698 100644
--- a/arch/x86/events/intel/cqm.h
+++ b/arch/x86/events/intel/cqm.h
@@ -123,9 +123,16 @@ struct monr;
* prmids.
* @limbo_rotation_entry: List entry to attach to ilstate_pmonrs_lru when
* this pmonr is in (IL)state.
- * @rotation_entry: List entry to attach to either astate_pmonrs_lru
- * or ilstate_pmonrs_lru in pkg_data.
+ * @last_enter_istate: Time last enter (I)state.
+ * @last_enter_astate: Time last enter (A)state. Used in rotation logic
+ * to guarantee that each pmonr gets a minimum
+ * time in (A)state.
+ * @rotation_entry: List entry to attach to pmonr rotation lists in
+ * pkg_data.
* @monr: The monr that contains this pmonr.
+ * @nr_enter_istate: Track number of times entered (I)state. Useful
+ * signal to diagnose excessive contention for
+ * rmids in this package.
* @pkg_id: Auxiliar variable with pkg id for this pmonr.
* @prmid_summary_atomic: Atomic accesor to store a union prmid_summary
* that represent the state of this pmonr.
@@ -194,6 +201,10 @@ struct pmonr {
struct monr *monr;
struct list_head rotation_entry;
+ unsigned long last_enter_istate;
+ unsigned long last_enter_astate;
+ unsigned int nr_enter_istate;
+
u16 pkg_id;
/* all writers are sync'ed by package's lock. */
@@ -218,6 +229,7 @@ struct pmonr {
* @ilsate_pmonrs_lru: pmonrs in (IL)state, these pmonrs have a valid
* limbo_prmid. It's a subset of istate_pmonrs_lru.
* Sorted increasingly by pmonr.last_enter_istate.
+ * @nr_inherited_pmonrs nr of pmonrs in any of the (I)state substates.
* @pkg_data_mutex: Hold for stability when modifying pmonrs
* hierarchy.
* @pkg_data_lock: Hold to protect variables that may be accessed
@@ -226,6 +238,7 @@ struct pmonr {
* hierarchy.
* @rotation_cpu: CPU to run @rotation_work on, it must be in the
* package associated to this instance of pkg_data.
+ * @rotation_work: Task that performs rotation of prmids.
*/
struct pkg_data {
u32 max_rmid;
@@ -247,9 +260,13 @@ struct pkg_data {
struct list_head istate_pmonrs_lru;
struct list_head ilstate_pmonrs_lru;
+ int nr_instate_pmonrs;
+ int nr_ilstate_pmonrs;
+
struct mutex pkg_data_mutex;
raw_spinlock_t pkg_data_lock;
+ struct delayed_work rotation_work;
int rotation_cpu;
};
@@ -410,6 +427,44 @@ static inline int monr_hrchy_count_held_raw_spin_locks(void)
#define CQM_DEFAULT_ROTATION_PERIOD 1200 /* ms */
/*
+ * Rotation function.
+ * Rotation logic runs per-package. In each package, if free rmids are needed,
+ * it will steal prmids from the pmonr that has been the longest time in
+ * (A)state.
+ * The hardware provides to way to signal that a rmid will be reused, therefore,
+ * before reusing a rmid that has been stolen, the rmid should stay for some
+ * in a "limbo" state where is not associated to any thread, hoping that the
+ * cache lines allocated for this rmid will eventually be replaced.
+ */
+static void intel_cqm_rmid_rotation_work(struct work_struct *work);
+
+/*
+ * Service Level Objectives (SLO) for the rotation logic.
+ *
+ * @__cqm_min_duration_mon_slice: Minimum duration of a monitored slice.
+ * @__cqm_max_wait_monitor: Maximum time that a pmonr can pass waiting for an
+ * RMID without rotation logic making any progress. Once elapsed for any
+ * prmid, the reusing threshold (__intel_cqm_max_threshold) can be increased,
+ * potentially increasing the speed at which RMIDs are reused, but potentially
+ * introducing measurement error.
+ */
+#define CQM_DEFAULT_MIN_MON_SLICE 2000 /* ms */
+static unsigned int __cqm_min_mon_slice = CQM_DEFAULT_MIN_MON_SLICE;
+
+#define CQM_DEFAULT_MAX_WAIT_MON 20000 /* ms */
+static unsigned int __cqm_max_wait_mon = CQM_DEFAULT_MAX_WAIT_MON;
+
+#define CQM_DEFAULT_MIN_PROGRESS_RATE 1 /* activated pmonrs per second */
+static unsigned int __cqm_min_progress_rate = CQM_DEFAULT_MIN_PROGRESS_RATE;
+/*
+ * If we fail to assign any RMID for intel_cqm_rotation because cachelines are
+ * still tagged with RMIDs in limbo even after having stolen enough rmids (a
+ * maximum number of rmids in limbo at any time), then we increment the dirty
+ * threshold to allow at least one RMID to be recycled. This mitigates the
+ * problem caused when cachelines tagged with a RMID are not evicted but
+ * it introduces error in the occupancy reads but allows the rotation of rmids
+ * to proceed.
+ *
* __intel_cqm_max_threshold provides an upper bound on the threshold,
* and is measured in bytes because it's exposed to userland.
* It's units are bytes must be scaled by cqm_l3_scale to obtain cache lines.
--
2.8.0.rc3.226.g39d4020
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2016-04-29 07:00 +0200 |
| Subject | [PATCH 08/32] perf/x86/intel/cqm: prepare for next patches |
| Message-ID | <rt8H1-8vJ-35@gated-at.bofh.it> |
| In reply to | #1390741 |
Move code around, delete unnecesary code and do some renaming in
in order to increase readibility of next patches. Create cqm.h file.
Reviewed-by: Stephane Eranian <eranian@google.com>
Signed-off-by: David Carrillo-Cisneros <davidcc@google.com>
---
arch/x86/events/intel/cqm.c | 170 +++++++++++++++-----------------------------
arch/x86/events/intel/cqm.h | 42 +++++++++++
include/linux/perf_event.h | 8 +--
3 files changed, 103 insertions(+), 117 deletions(-)
create mode 100644 arch/x86/events/intel/cqm.h
diff --git a/arch/x86/events/intel/cqm.c b/arch/x86/events/intel/cqm.c
index d5eac8f..f678014 100644
--- a/arch/x86/events/intel/cqm.c
+++ b/arch/x86/events/intel/cqm.c
@@ -4,10 +4,9 @@
* Based very, very heavily on work by Peter Zijlstra.
*/
-#include <linux/perf_event.h>
#include <linux/slab.h>
#include <asm/cpu_device_id.h>
-#include <asm/pqr_common.h>
+#include "cqm.h"
#include "../perf_event.h"
#define MSR_IA32_QM_CTR 0x0c8e
@@ -16,13 +15,26 @@
static u32 cqm_max_rmid = -1;
static unsigned int cqm_l3_scale; /* supposedly cacheline size */
+#define RMID_VAL_ERROR (1ULL << 63)
+#define RMID_VAL_UNAVAIL (1ULL << 62)
+
+#define QOS_L3_OCCUP_EVENT_ID (1 << 0)
+
+#define QOS_EVENT_MASK QOS_L3_OCCUP_EVENT_ID
+
+#define CQM_EVENT_ATTR_STR(_name, v, str) \
+static struct perf_pmu_events_attr event_attr_##v = { \
+ .attr = __ATTR(_name, 0444, perf_event_sysfs_show, NULL), \
+ .id = 0, \
+ .event_str = str, \
+}
+
/*
* Updates caller cpu's cache.
*/
static inline void __update_pqr_rmid(u32 rmid)
{
struct intel_pqr_state *state = this_cpu_ptr(&pqr_state);
-
if (state->rmid == rmid)
return;
state->rmid = rmid;
@@ -30,37 +42,18 @@ static inline void __update_pqr_rmid(u32 rmid)
}
/*
- * Protects cache_cgroups and cqm_rmid_free_lru and cqm_rmid_limbo_lru.
- * Also protects event->hw.cqm_rmid
- *
- * Hold either for stability, both for modification of ->hw.cqm_rmid.
- */
-static DEFINE_MUTEX(cache_mutex);
-static DEFINE_RAW_SPINLOCK(cache_lock);
-
-#define CQM_EVENT_ATTR_STR(_name, v, str) \
-static struct perf_pmu_events_attr event_attr_##v = { \
- .attr = __ATTR(_name, 0444, perf_event_sysfs_show, NULL), \
- .id = 0, \
- .event_str = str, \
-}
-
-/*
* Groups of events that have the same target(s), one RMID per group.
+ * Protected by cqm_mutex.
*/
static LIST_HEAD(cache_groups);
+static DEFINE_MUTEX(cqm_mutex);
+static DEFINE_RAW_SPINLOCK(cache_lock);
/*
* Mask of CPUs for reading CQM values. We only need one per-socket.
*/
static cpumask_t cqm_cpumask;
-#define RMID_VAL_ERROR (1ULL << 63)
-#define RMID_VAL_UNAVAIL (1ULL << 62)
-
-#define QOS_L3_OCCUP_EVENT_ID (1 << 0)
-
-#define QOS_EVENT_MASK QOS_L3_OCCUP_EVENT_ID
/*
* This is central to the rotation algorithm in __intel_cqm_rmid_rotate().
@@ -71,8 +64,6 @@ static cpumask_t cqm_cpumask;
*/
static u32 intel_cqm_rotation_rmid;
-#define INVALID_RMID (-1)
-
/*
* Is @rmid valid for programming the hardware?
*
@@ -140,7 +131,7 @@ struct cqm_rmid_entry {
* rotation worker moves RMIDs from the limbo list to the free list once
* the occupancy value drops below __intel_cqm_threshold.
*
- * Both lists are protected by cache_mutex.
+ * Both lists are protected by cqm_mutex.
*/
static LIST_HEAD(cqm_rmid_free_lru);
static LIST_HEAD(cqm_rmid_limbo_lru);
@@ -172,13 +163,13 @@ static inline struct cqm_rmid_entry *__rmid_entry(u32 rmid)
/*
* Returns < 0 on fail.
*
- * We expect to be called with cache_mutex held.
+ * We expect to be called with cqm_mutex held.
*/
static u32 __get_rmid(void)
{
struct cqm_rmid_entry *entry;
- lockdep_assert_held(&cache_mutex);
+ lockdep_assert_held(&cqm_mutex);
if (list_empty(&cqm_rmid_free_lru))
return INVALID_RMID;
@@ -193,7 +184,7 @@ static void __put_rmid(u32 rmid)
{
struct cqm_rmid_entry *entry;
- lockdep_assert_held(&cache_mutex);
+ lockdep_assert_held(&cqm_mutex);
WARN_ON(!__rmid_valid(rmid));
entry = __rmid_entry(rmid);
@@ -237,9 +228,9 @@ static int intel_cqm_setup_rmid_cache(void)
entry = __rmid_entry(0);
list_del(&entry->list);
- mutex_lock(&cache_mutex);
+ mutex_lock(&cqm_mutex);
intel_cqm_rotation_rmid = __get_rmid();
- mutex_unlock(&cache_mutex);
+ mutex_unlock(&cqm_mutex);
return 0;
fail:
@@ -250,6 +241,7 @@ fail:
return -ENOMEM;
}
+
/*
* Determine if @a and @b measure the same set of tasks.
*
@@ -287,49 +279,11 @@ static bool __match_event(struct perf_event *a, struct perf_event *b)
return false;
}
-#ifdef CONFIG_CGROUP_PERF
-static inline struct perf_cgroup *event_to_cgroup(struct perf_event *event)
-{
- if (event->attach_state & PERF_ATTACH_TASK)
- return perf_cgroup_from_task(event->hw.target, event->ctx);
-
- return event->cgrp;
-}
-#endif
-
struct rmid_read {
u32 rmid;
atomic64_t value;
};
-static void intel_cqm_event_read(struct perf_event *event);
-
-/*
- * If we fail to assign a new RMID for intel_cqm_rotation_rmid because
- * cachelines are still tagged with RMIDs in limbo, we progressively
- * increment the threshold until we find an RMID in limbo with <=
- * __intel_cqm_threshold lines tagged. This is designed to mitigate the
- * problem where cachelines tagged with an RMID are not steadily being
- * evicted.
- *
- * On successful rotations we decrease the threshold back towards zero.
- *
- * __intel_cqm_max_threshold provides an upper bound on the threshold,
- * and is measured in bytes because it's exposed to userland.
- */
-static unsigned int __intel_cqm_threshold;
-static unsigned int __intel_cqm_max_threshold;
-
-/*
- * Initially use this constant for both the limbo queue time and the
- * rotation timer interval, pmu::hrtimer_interval_ms.
- *
- * They don't need to be the same, but the two are related since if you
- * rotate faster than you recycle RMIDs, you may run out of available
- * RMIDs.
- */
-#define RMID_DEFAULT_QUEUE_TIME 250 /* ms */
-
static struct pmu intel_cqm_pmu;
/*
@@ -344,7 +298,7 @@ static void intel_cqm_setup_event(struct perf_event *event,
bool conflict = false;
u32 rmid;
- list_for_each_entry(iter, &cache_groups, hw.cqm_groups_entry) {
+ list_for_each_entry(iter, &cache_groups, hw.cqm_event_groups_entry) {
rmid = iter->hw.cqm_rmid;
if (__match_event(iter, event)) {
@@ -390,24 +344,24 @@ out:
static inline bool cqm_group_leader(struct perf_event *event)
{
- return !list_empty(&event->hw.cqm_groups_entry);
+ return !list_empty(&event->hw.cqm_event_groups_entry);
}
static void intel_cqm_event_start(struct perf_event *event, int mode)
{
- if (!(event->hw.cqm_state & PERF_HES_STOPPED))
+ if (!(event->hw.state & PERF_HES_STOPPED))
return;
- event->hw.cqm_state &= ~PERF_HES_STOPPED;
+ event->hw.state &= ~PERF_HES_STOPPED;
__update_pqr_rmid(event->hw.cqm_rmid);
}
static void intel_cqm_event_stop(struct perf_event *event, int mode)
{
- if (event->hw.cqm_state & PERF_HES_STOPPED)
+ if (event->hw.state & PERF_HES_STOPPED)
return;
- event->hw.cqm_state |= PERF_HES_STOPPED;
+ event->hw.state |= PERF_HES_STOPPED;
intel_cqm_event_read(event);
__update_pqr_rmid(0);
}
@@ -419,7 +373,7 @@ static int intel_cqm_event_add(struct perf_event *event, int mode)
raw_spin_lock_irqsave(&cache_lock, flags);
- event->hw.cqm_state = PERF_HES_STOPPED;
+ event->hw.state = PERF_HES_STOPPED;
rmid = event->hw.cqm_rmid;
if (__rmid_valid(rmid) && (mode & PERF_EF_START))
@@ -433,16 +387,16 @@ static void intel_cqm_event_destroy(struct perf_event *event)
{
struct perf_event *group_other = NULL;
- mutex_lock(&cache_mutex);
+ mutex_lock(&cqm_mutex);
/*
* If there's another event in this group...
*/
- if (!list_empty(&event->hw.cqm_group_entry)) {
- group_other = list_first_entry(&event->hw.cqm_group_entry,
+ if (!list_empty(&event->hw.cqm_event_group_entry)) {
+ group_other = list_first_entry(&event->hw.cqm_event_group_entry,
struct perf_event,
- hw.cqm_group_entry);
- list_del(&event->hw.cqm_group_entry);
+ hw.cqm_event_group_entry);
+ list_del(&event->hw.cqm_event_group_entry);
}
/*
@@ -454,18 +408,18 @@ static void intel_cqm_event_destroy(struct perf_event *event)
* destroy the group and return the RMID.
*/
if (group_other) {
- list_replace(&event->hw.cqm_groups_entry,
- &group_other->hw.cqm_groups_entry);
+ list_replace(&event->hw.cqm_event_groups_entry,
+ &group_other->hw.cqm_event_groups_entry);
} else {
u32 rmid = event->hw.cqm_rmid;
if (__rmid_valid(rmid))
__put_rmid(rmid);
- list_del(&event->hw.cqm_groups_entry);
+ list_del(&event->hw.cqm_event_groups_entry);
}
}
- mutex_unlock(&cache_mutex);
+ mutex_unlock(&cqm_mutex);
}
static int intel_cqm_event_init(struct perf_event *event)
@@ -488,25 +442,26 @@ static int intel_cqm_event_init(struct perf_event *event)
event->attr.sample_period) /* no sampling */
return -EINVAL;
- INIT_LIST_HEAD(&event->hw.cqm_group_entry);
- INIT_LIST_HEAD(&event->hw.cqm_groups_entry);
+ INIT_LIST_HEAD(&event->hw.cqm_event_groups_entry);
+ INIT_LIST_HEAD(&event->hw.cqm_event_group_entry);
event->destroy = intel_cqm_event_destroy;
- mutex_lock(&cache_mutex);
+ mutex_lock(&cqm_mutex);
+
/* Will also set rmid */
intel_cqm_setup_event(event, &group);
if (group) {
- list_add_tail(&event->hw.cqm_group_entry,
- &group->hw.cqm_group_entry);
+ list_add_tail(&event->hw.cqm_event_group_entry,
+ &group->hw.cqm_event_group_entry);
} else {
- list_add_tail(&event->hw.cqm_groups_entry,
- &cache_groups);
+ list_add_tail(&event->hw.cqm_event_groups_entry,
+ &cache_groups);
}
- mutex_unlock(&cache_mutex);
+ mutex_unlock(&cqm_mutex);
return 0;
}
@@ -543,14 +498,14 @@ static struct attribute_group intel_cqm_format_group = {
};
static ssize_t
-max_recycle_threshold_show(struct device *dev, struct device_attribute *attr,
- char *page)
+max_recycle_threshold_show(
+ struct device *dev, struct device_attribute *attr, char *page)
{
ssize_t rv;
- mutex_lock(&cache_mutex);
+ mutex_lock(&cqm_mutex);
rv = snprintf(page, PAGE_SIZE-1, "%u\n", __intel_cqm_max_threshold);
- mutex_unlock(&cache_mutex);
+ mutex_unlock(&cqm_mutex);
return rv;
}
@@ -560,25 +515,16 @@ max_recycle_threshold_store(struct device *dev,
struct device_attribute *attr,
const char *buf, size_t count)
{
- unsigned int bytes, cachelines;
+ unsigned int bytes;
int ret;
ret = kstrtouint(buf, 0, &bytes);
if (ret)
return ret;
- mutex_lock(&cache_mutex);
-
+ mutex_lock(&cqm_mutex);
__intel_cqm_max_threshold = bytes;
- cachelines = bytes / cqm_l3_scale;
-
- /*
- * The new maximum takes effect immediately.
- */
- if (__intel_cqm_threshold > cachelines)
- __intel_cqm_threshold = cachelines;
-
- mutex_unlock(&cache_mutex);
+ mutex_unlock(&cqm_mutex);
return count;
}
@@ -602,7 +548,7 @@ static const struct attribute_group *intel_cqm_attr_groups[] = {
};
static struct pmu intel_cqm_pmu = {
- .hrtimer_interval_ms = RMID_DEFAULT_QUEUE_TIME,
+ .hrtimer_interval_ms = CQM_DEFAULT_ROTATION_PERIOD,
.attr_groups = intel_cqm_attr_groups,
.task_ctx_nr = perf_sw_context,
.event_init = intel_cqm_event_init,
diff --git a/arch/x86/events/intel/cqm.h b/arch/x86/events/intel/cqm.h
new file mode 100644
index 0000000..e25d0a1
--- /dev/null
+++ b/arch/x86/events/intel/cqm.h
@@ -0,0 +1,42 @@
+/*
+ * Intel Cache Quality-of-Service Monitoring (CQM) support.
+ *
+ * A Resource Manager ID (RMID) is a u32 value that, when programmed in a
+ * logical CPU, will allow the LLC cache to associate the changes in occupancy
+ * generated by that cpu (cache lines allocations - deallocations) to the RMID.
+ * If an rmid has been assigned to a thread T long enough for all cache lines
+ * used by T to be allocated, then the occupancy reported by the hardware is
+ * equal to the total cache occupancy for T.
+ *
+ * Groups of threads that are to be monitored together (such as cgroups
+ * or processes) can shared a RMID.
+ *
+ * This driver implements a tree hierarchy of Monitored Resources (monr). Each
+ * monr is a cgroup, a process or a thread that needs one single RMID.
+ */
+
+#include <linux/perf_event.h>
+#include <asm/pqr_common.h>
+
+/*
+ * Minimum time elapsed between reads of occupancy value for an RMID when
+ * transversing the monr hierarchy.
+ */
+#define RMID_DEFAULT_MIN_UPDATE_TIME 20 /* ms */
+
+# define INVALID_RMID (-1)
+
+/*
+ * Time between execution of rotation logic. The frequency of execution does
+ * not affect the rate at which RMIDs are recycled, except by the delay by the
+ * delay updating the prmid's and their pools.
+ * The rotation period is stored in pmu->hrtimer_interval_ms.
+ */
+#define CQM_DEFAULT_ROTATION_PERIOD 1200 /* ms */
+
+/*
+ * __intel_cqm_max_threshold provides an upper bound on the threshold,
+ * and is measured in bytes because it's exposed to userland.
+ * It's units are bytes must be scaled by cqm_l3_scale to obtain cache lines.
+ */
+static unsigned int __intel_cqm_max_threshold;
diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
index 3a847bf..5eb7dea 100644
--- a/include/linux/perf_event.h
+++ b/include/linux/perf_event.h
@@ -120,11 +120,9 @@ struct hw_perf_event {
};
#ifdef CONFIG_INTEL_RDT
struct { /* intel_cqm */
- int cqm_state;
- u32 cqm_rmid;
- struct list_head cqm_events_entry;
- struct list_head cqm_groups_entry;
- struct list_head cqm_group_entry;
+ u32 cqm_rmid;
+ struct list_head cqm_event_group_entry;
+ struct list_head cqm_event_groups_entry;
};
#endif
struct { /* itrace */
--
2.8.0.rc3.226.g39d4020
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-04-29 11:20 +0200 |
| Subject | Re: [PATCH 08/32] perf/x86/intel/cqm: prepare for next patches |
| Message-ID | <rtcKC-3oy-5@gated-at.bofh.it> |
| In reply to | #1390761 |
On Thu, Apr 28, 2016 at 09:43:14PM -0700, David Carrillo-Cisneros wrote: > Move code around, delete unnecesary code and do some renaming in > in order to increase readibility of next patches. Create cqm.h file. *sigh*, this is a royal pain in the backside to review. Please just completely wipe the old driver in patch 1, preserve _nothing_. Then start adding bits back, in gradual coherent pieces. Like that msr write optimization you need new hooks for, that should be a patch doing just that, optimize, it should not introduce new functionality etc.. This piecewise removal of small bits makes it entirely hard to see the complete picture of what is introduced here.
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2016-04-29 07:00 +0200 |
| Subject | [PATCH 07/32] perf/x86/intel/cqm: separate CQM PMU's attributes from x86 PMU |
| Message-ID | <rt8H1-8vJ-39@gated-at.bofh.it> |
| In reply to | #1390741 |
Create a CQM_EVENT_ATTR_STR to use in CQM to remove dependency
on the unrelated x86's PMU EVENT_ATTR_STR.
Reviewed-by: Stephane Eranian <eranian@google.com>
Signed-off-by: David Carrillo-Cisneros <davidcc@google.com>
---
arch/x86/events/intel/cqm.c | 17 ++++++++++++-----
1 file changed, 12 insertions(+), 5 deletions(-)
diff --git a/arch/x86/events/intel/cqm.c b/arch/x86/events/intel/cqm.c
index 8457dd0..d5eac8f 100644
--- a/arch/x86/events/intel/cqm.c
+++ b/arch/x86/events/intel/cqm.c
@@ -38,6 +38,13 @@ static inline void __update_pqr_rmid(u32 rmid)
static DEFINE_MUTEX(cache_mutex);
static DEFINE_RAW_SPINLOCK(cache_lock);
+#define CQM_EVENT_ATTR_STR(_name, v, str) \
+static struct perf_pmu_events_attr event_attr_##v = { \
+ .attr = __ATTR(_name, 0444, perf_event_sysfs_show, NULL), \
+ .id = 0, \
+ .event_str = str, \
+}
+
/*
* Groups of events that have the same target(s), one RMID per group.
*/
@@ -504,11 +511,11 @@ static int intel_cqm_event_init(struct perf_event *event)
return 0;
}
-EVENT_ATTR_STR(llc_occupancy, intel_cqm_llc, "event=0x01");
-EVENT_ATTR_STR(llc_occupancy.per-pkg, intel_cqm_llc_pkg, "1");
-EVENT_ATTR_STR(llc_occupancy.unit, intel_cqm_llc_unit, "Bytes");
-EVENT_ATTR_STR(llc_occupancy.scale, intel_cqm_llc_scale, NULL);
-EVENT_ATTR_STR(llc_occupancy.snapshot, intel_cqm_llc_snapshot, "1");
+CQM_EVENT_ATTR_STR(llc_occupancy, intel_cqm_llc, "event=0x01");
+CQM_EVENT_ATTR_STR(llc_occupancy.per-pkg, intel_cqm_llc_pkg, "1");
+CQM_EVENT_ATTR_STR(llc_occupancy.unit, intel_cqm_llc_unit, "Bytes");
+CQM_EVENT_ATTR_STR(llc_occupancy.scale, intel_cqm_llc_scale, NULL);
+CQM_EVENT_ATTR_STR(llc_occupancy.snapshot, intel_cqm_llc_snapshot, "1");
static struct attribute *intel_cqm_events_attr[] = {
EVENT_PTR(intel_cqm_llc),
--
2.8.0.rc3.226.g39d4020
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2016-04-29 07:00 +0200 |
| Subject | [PATCH 19/32] perf/core: introduce PMU event flag PERF_CGROUP_NO_RECURSION |
| Message-ID | <rt8H1-8vJ-45@gated-at.bofh.it> |
| In reply to | #1390741 |
Some events, such as Intel's CQM llc_occupancy, need small deviations
from the traditional behavior in the generic code in a way that depends
on the event itself (and known by the PMU) and not in a field of
perf_event_attrs.
An example is the recursive scope for cgroups: The generic code handles
cgroup hierarchy for a cgroup C by simultaneously adding to the PMU
the events of all cgroups that are ancestors of C. This approach is
incompatible with the CQM hw that only allows one RMID per virtual core
at a time. CQM's PMU work-arounds this limitation by internally
maintaining the hierarchical dependency between monitored cgroups and
only requires that the generic code adds current cgroup's event to
the PMU.
The introduction of the flag PERF_CGROUP_NO_RECURSION allows the PMU to
signal the generic code to avoid using recursive cgroup scope for
llc_occupancy events, preventing an undesired overwrite of RMIDs.
The PERF_CGROUP_NO_RECURSION, introduced in this patch, is the first flag
of this type, more will be added in this patch series.
To keep things tidy, this patch introduces the flag field pmu_event_flag,
intended to contain all flags that:
- Are not user-configurable event attributes (not suitable for
perf_event_attributes).
- Are known by the PMU during initialization of struct perf_event.
- Signal something to the generic code.
Reviewed-by: Stephane Eranian <eranian@google.com>
Signed-off-by: David Carrillo-Cisneros <davidcc@google.com>
---
include/linux/perf_event.h | 10 ++++++++++
kernel/events/core.c | 3 +++
2 files changed, 13 insertions(+)
diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
index 81e29c6..e4c58b0 100644
--- a/include/linux/perf_event.h
+++ b/include/linux/perf_event.h
@@ -594,9 +594,19 @@ struct perf_event {
#endif
struct list_head sb_list;
+
+ /* Flags to generic code set by PMU. */
+ int pmu_event_flags;
+
#endif /* CONFIG_PERF_EVENTS */
};
+/*
+ * Possible flags for mpu_event_flags.
+ */
+/* Do not enable cgroup events in descendant cgroups. */
+#define PERF_CGROUP_NO_RECURSION (1 << 0)
+
/**
* struct perf_event_context - event context structure
*
diff --git a/kernel/events/core.c b/kernel/events/core.c
index 2a868a6..33961ec 100644
--- a/kernel/events/core.c
+++ b/kernel/events/core.c
@@ -545,6 +545,9 @@ perf_cgroup_match(struct perf_event *event)
if (!cpuctx->cgrp)
return false;
+ if (event->pmu_event_flags & PERF_CGROUP_NO_RECURSION)
+ return cpuctx->cgrp->css.cgroup == event->cgrp->css.cgroup;
+
/*
* Cgroup scoping is recursive. An event enabled for a cgroup is
* also enabled for all its descendant cgroups. If @cpuctx's
--
2.8.0.rc3.226.g39d4020
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2016-04-29 07:00 +0200 |
| Subject | [PATCH 10/32] perf/x86/intel/cqm: basic RMID hierarchy with per package rmids |
| Message-ID | <rt8H1-8vJ-41@gated-at.bofh.it> |
| In reply to | #1390741 |
Cgroups and/or tasks that require to be monitored using a RMID
are abstracted as a MOnitored Resources (monr's). A CQM event points
to a monr to read occupancy (and in the future other attributes) of the
RMIDs associated to the monr.
The monrs form a hierarchy that captures the dependency within the
monitored cgroups and/or tasks/threads. The monr of a cgroup A which
contains another monitored cgroup, B, is an ancestor of B's monr.
Each monr contains one Package MONitored Resource (pmonr) per package.
The monitoring of a monr in a package starts when its corresponding
pmonr receives an RMID for that package (a prmid).
The prmids are lazily assigned to a pmonr the first time a thread
using the monr is scheduled in the package. When a pmonr with a
valid prmid is scheduled, that pmonr's prmid's RMID is written to the
msr MSR_IA32_PQR_ASSOC. If no prmid is available, the prmid of the lowest
ancestor in the monr hierarchy with a valid prmid for that package is
used instead.
A pmonr can be in one of following three states:
- (A)ctive: When it has a prmid available.
- (I)nherited: When no prmid is available. In this state, it "borrows"
the prmid of its lowest ancestor in (A)ctive state during sched in
(writes its ancestor's RMID into hw while any associated thread is
executed). But, since the "borrowed" prmid do not monitor the
occupancy of this monr, the monr cannot report occupancy individually.
- (U)nused: When the monr does not have a prmid yet and have no failed
acquiring one (either because no thread has been scheduled while
monitoring for this pmonr is active or because it has been completed
a transition to (U)state, ie. termination of the associated
event/cgroup).
To avoid synchronization overhead, each prmid contains a prmid_summary.
The union prmid_summary is a concise representation of the prmid state
and its raw RMIDs. Due to its size, the prmid_summary can be read
atomically without a LOCK instruction. Every state transition atomically
updates the prmid_summary. This avoids locking during sched in and out
of threads, except in the cases that a prmid needs to be allocated,
but this only occurs the first time a monr is scheduled in a package.
This patch introduces a first iteration of the monr hierarchy
that maintains two levels: the root monr, at top, and all other monrs
as leaves. The root monr is always (A)ctive.
This patch also implements the essential mechanism of per-package lazy
allocation of RMID.
The (I)state and the transitions from and to it are introduced in the
next patch in this series.
Reviewed-by: Stephane Eranian <eranian@google.com>
Signed-off-by: David Carrillo-Cisneros <davidcc@google.com>
---
arch/x86/events/intel/cqm.c | 633 ++++++++++++++++++++++++++++++++++++--------
arch/x86/events/intel/cqm.h | 149 +++++++++++
include/linux/perf_event.h | 2 +-
3 files changed, 674 insertions(+), 110 deletions(-)
diff --git a/arch/x86/events/intel/cqm.c b/arch/x86/events/intel/cqm.c
index 541e515..65551bb 100644
--- a/arch/x86/events/intel/cqm.c
+++ b/arch/x86/events/intel/cqm.c
@@ -35,28 +35,66 @@ static struct perf_pmu_events_attr event_attr_##v = { \
static LIST_HEAD(cache_groups);
static DEFINE_MUTEX(cqm_mutex);
+struct monr *monr_hrchy_root;
+
struct pkg_data *cqm_pkgs_data[PQR_MAX_NR_PKGS];
-/*
- * Is @rmid valid for programming the hardware?
- *
- * rmid 0 is reserved by the hardware for all non-monitored tasks, which
- * means that we should never come across an rmid with that value.
- * Likewise, an rmid value of -1 is used to indicate "no rmid currently
- * assigned" and is used as part of the rotation code.
- */
-static inline bool __rmid_valid(u32 rmid)
+static inline bool __pmonr__in_astate(struct pmonr *pmonr)
{
- if (!rmid || rmid == INVALID_RMID)
- return false;
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+ return pmonr->prmid;
+}
- return true;
+static inline bool __pmonr__in_ustate(struct pmonr *pmonr)
+{
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+ return !pmonr->prmid;
}
-static u64 __rmid_read(u32 rmid)
+static inline bool monr__is_root(struct monr *monr)
{
- /* XXX: Placeholder, will be removed in next patch. */
- return 0;
+ return monr_hrchy_root == monr;
+}
+
+static inline bool monr__is_mon_active(struct monr *monr)
+{
+ return monr->flags & MONR_MON_ACTIVE;
+}
+
+static inline void __monr__set_summary_read_rmid(struct monr *monr, u32 rmid)
+{
+ int i;
+ struct pmonr *pmonr;
+ union prmid_summary summary;
+
+ monr_hrchy_assert_held_raw_spin_locks();
+
+ cqm_pkg_id_for_each_online(i) {
+ pmonr = monr->pmonrs[i];
+ WARN_ON_ONCE(!__pmonr__in_ustate(pmonr));
+ summary.value = atomic64_read(&pmonr->prmid_summary_atomic);
+ summary.read_rmid = rmid;
+ atomic64_set(&pmonr->prmid_summary_atomic, summary.value);
+ }
+}
+
+static inline void __monr__set_mon_active(struct monr *monr)
+{
+ monr_hrchy_assert_held_raw_spin_locks();
+ __monr__set_summary_read_rmid(monr, 0);
+ monr->flags |= MONR_MON_ACTIVE;
+}
+
+/*
+ * All pmonrs must be in (U)state.
+ * clearing MONR_MON_ACTIVE prevents (U)state prmids from transitioning
+ * to another state.
+ */
+static inline void __monr__clear_mon_active(struct monr *monr)
+{
+ monr_hrchy_assert_held_raw_spin_locks();
+ __monr__set_summary_read_rmid(monr, INVALID_RMID);
+ monr->flags &= ~MONR_MON_ACTIVE;
}
/*
@@ -133,22 +171,6 @@ static inline bool __valid_pkg_id(u16 pkg_id)
return pkg_id < PQR_MAX_NR_PKGS;
}
-/*
- * Returns < 0 on fail.
- *
- * We expect to be called with cache_mutex held.
- */
-static u32 __get_rmid(void)
-{
- /* XXX: Placeholder, will be removed in next patch. */
- return 0;
-}
-
-static void __put_rmid(u32 rmid)
-{
- /* XXX: Placeholder, will be removed in next patch. */
-}
-
/* Init cqm pkg_data for @cpu 's package. */
static int pkg_data_init_cpu(int cpu)
{
@@ -187,6 +209,10 @@ static int pkg_data_init_cpu(int cpu)
}
INIT_LIST_HEAD(&pkg_data->free_prmids_pool);
+ INIT_LIST_HEAD(&pkg_data->active_prmids_pool);
+ INIT_LIST_HEAD(&pkg_data->nopmonr_limbo_prmids_pool);
+
+ INIT_LIST_HEAD(&pkg_data->astate_pmonrs_lru);
mutex_init(&pkg_data->pkg_data_mutex);
raw_spin_lock_init(&pkg_data->pkg_data_lock);
@@ -225,12 +251,129 @@ __prmid_from_rmid(u16 pkg_id, u32 rmid)
return prmid;
}
+static struct pmonr *pmonr_alloc(int cpu)
+{
+ struct pmonr *pmonr;
+ union prmid_summary summary;
+
+ pmonr = kmalloc_node(sizeof(struct pmonr),
+ GFP_KERNEL, cpu_to_node(cpu));
+ if (!pmonr)
+ return ERR_PTR(-ENOMEM);
+
+ pmonr->prmid = NULL;
+
+ pmonr->monr = NULL;
+ INIT_LIST_HEAD(&pmonr->rotation_entry);
+
+ pmonr->pkg_id = topology_physical_package_id(cpu);
+ summary.sched_rmid = INVALID_RMID;
+ summary.read_rmid = INVALID_RMID;
+ atomic64_set(&pmonr->prmid_summary_atomic, summary.value);
+
+ return pmonr;
+}
+
+static void pmonr_dealloc(struct pmonr *pmonr)
+{
+ kfree(pmonr);
+}
+
+/*
+ * @root: Common ancestor.
+ * a bust be distinct to b.
+ * @true if a is ancestor of b.
+ */
+static inline bool
+__monr_hrchy_is_ancestor(struct monr *root,
+ struct monr *a, struct monr *b)
+{
+ WARN_ON_ONCE(!root || !a || !b);
+ WARN_ON_ONCE(a == b);
+
+ if (root == a)
+ return true;
+ if (root == b)
+ return false;
+
+ b = b->parent;
+ /* Break at the root */
+ while (b != root) {
+ WARN_ON_ONCE(!b);
+ if (a == b)
+ return true;
+ b = b->parent;
+ }
+ return false;
+}
+
+/* helper function to finish transition to astate. */
+static inline void
+__pmonr__finish_to_astate(struct pmonr *pmonr, struct prmid *prmid)
+{
+ union prmid_summary summary;
+
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+
+ pmonr->prmid = prmid;
+
+ list_move_tail(
+ &prmid->pool_entry, &__pkg_data(pmonr, active_prmids_pool));
+ list_move_tail(
+ &pmonr->rotation_entry, &__pkg_data(pmonr, astate_pmonrs_lru));
+
+ summary.sched_rmid = pmonr->prmid->rmid;
+ summary.read_rmid = pmonr->prmid->rmid;
+ atomic64_set(&pmonr->prmid_summary_atomic, summary.value);
+}
+
+static inline void
+__pmonr__ustate_to_astate(struct pmonr *pmonr, struct prmid *prmid)
+{
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+ __pmonr__finish_to_astate(pmonr, prmid);
+}
+
+static inline void
+__pmonr__to_ustate(struct pmonr *pmonr)
+{
+ union prmid_summary summary;
+
+ lockdep_assert_held(&__pkg_data(pmonr, pkg_data_lock));
+
+ /* Do not warn on re-enter state for (U)state, to simplify cleanup
+ * of initialized states that were not scheduled.
+ */
+ if (__pmonr__in_ustate(pmonr))
+ return;
+
+ if (__pmonr__in_astate(pmonr)) {
+ WARN_ON_ONCE(!pmonr->prmid);
+
+ list_move_tail(&pmonr->prmid->pool_entry,
+ &__pkg_data(pmonr, nopmonr_limbo_prmids_pool));
+ pmonr->prmid = NULL;
+ } else {
+ WARN_ON_ONCE(true);
+ return;
+ }
+ list_del_init(&pmonr->rotation_entry);
+
+ summary.sched_rmid = INVALID_RMID;
+ summary.read_rmid =
+ monr__is_mon_active(pmonr->monr) ? 0 : INVALID_RMID;
+
+ atomic64_set(&pmonr->prmid_summary_atomic, summary.value);
+ WARN_ON_ONCE(!__pmonr__in_ustate(pmonr));
+}
+
static int intel_cqm_setup_pkg_prmid_pools(u16 pkg_id)
{
int r;
unsigned long flags;
struct prmid *prmid;
struct pkg_data *pkg_data = cqm_pkgs_data[pkg_id];
+ struct pmonr *root_pmonr;
if (!__valid_pkg_id(pkg_id))
return -EINVAL;
@@ -252,12 +395,13 @@ static int intel_cqm_setup_pkg_prmid_pools(u16 pkg_id)
&pkg_data->pkg_data_lock, flags, pkg_id);
pkg_data->prmids_by_rmid[r] = prmid;
+ list_add_tail(&prmid->pool_entry, &pkg_data->free_prmids_pool);
/* RMID 0 is special and makes the root of rmid hierarchy. */
- if (r != 0)
- list_add_tail(&prmid->pool_entry,
- &pkg_data->free_prmids_pool);
-
+ if (r == 0) {
+ root_pmonr = monr_hrchy_root->pmonrs[pkg_id];
+ __pmonr__ustate_to_astate(root_pmonr, prmid);
+ }
raw_spin_unlock_irqrestore(&pkg_data->pkg_data_lock, flags);
}
return 0;
@@ -273,6 +417,232 @@ fail:
}
+/* Alloc monr with all pmonrs in (U)state. */
+static struct monr *monr_alloc(void)
+{
+ int i;
+ struct pmonr *pmonr;
+ struct monr *monr;
+
+ monr = kmalloc(sizeof(struct monr), GFP_KERNEL);
+
+ if (!monr)
+ return ERR_PTR(-ENOMEM);
+
+ monr->flags = 0;
+ monr->parent = NULL;
+ INIT_LIST_HEAD(&monr->children);
+ INIT_LIST_HEAD(&monr->parent_entry);
+ monr->mon_event_group = NULL;
+
+ /* Iterate over all pkgs, even unitialized ones. */
+ for (i = 0; i < PQR_MAX_NR_PKGS; i++) {
+ /* Do not create pmonrs for unitialized packages. */
+ if (!cqm_pkgs_data[i]) {
+ monr->pmonrs[i] = NULL;
+ continue;
+ }
+ /* Rotation cpu is on pmonr's package. */
+ pmonr = pmonr_alloc(cqm_pkgs_data[i]->rotation_cpu);
+ if (IS_ERR(pmonr))
+ goto clean_pmonrs;
+ pmonr->monr = monr;
+ monr->pmonrs[i] = pmonr;
+ }
+ return monr;
+
+clean_pmonrs:
+ while (i--) {
+ if (cqm_pkgs_data[i])
+ kfree(monr->pmonrs[i]);
+ }
+ kfree(monr);
+ return ERR_PTR(PTR_ERR(pmonr));
+}
+
+/* Only can dealloc monrs with all pmonrs in (U)state. */
+static void monr_dealloc(struct monr *monr)
+{
+ int i;
+
+ cqm_pkg_id_for_each_online(i)
+ pmonr_dealloc(monr->pmonrs[i]);
+
+ kfree(monr);
+}
+
+/*
+ * Wrappers for monr manipulation in events.
+ *
+ */
+static inline struct monr *monr_from_event(struct perf_event *event)
+{
+ return (struct monr *) READ_ONCE(event->hw.cqm_monr);
+}
+
+static inline void event_set_monr(struct perf_event *event, struct monr *monr)
+{
+ WRITE_ONCE(event->hw.cqm_monr, monr);
+}
+
+/*
+ * Always finds a rmid_entry to schedule. To be called during scheduler.
+ * A fast path that only uses read_lock for common case when rmid for current
+ * package has been used before.
+ * On failure, verify that monr is active, if it is, try to obtain a free rmid
+ * and set pmonr to (A)state.
+ * On failure, transverse up monr_hrchy until finding one prmid for this
+ * pkg_id and set pmonr to (I)state.
+ * Called during task switch, it will set pmonr's prmid_summary to reflect the
+ * sched and read rmids that reflect pmonr's state.
+ */
+static inline void
+monr_hrchy_get_next_prmid_summary(struct pmonr *pmonr)
+{
+ union prmid_summary summary;
+
+ /*
+ * First, do lock-free fastpath.
+ */
+ summary.value = atomic64_read(&pmonr->prmid_summary_atomic);
+ if (summary.sched_rmid != INVALID_RMID)
+ return;
+
+ if (!prmid_summary__is_mon_active(summary))
+ return;
+
+ /*
+ * Lock-free path failed at first attempt. Now acquire lock and repeat
+ * in case the monr was modified in the mean time.
+ * This time try to obtain free rmid and update pmonr accordingly,
+ * instead of failing fast.
+ */
+ raw_spin_lock_nested(&__pkg_data(pmonr, pkg_data_lock), pmonr->pkg_id);
+
+ summary.value = atomic64_read(&pmonr->prmid_summary_atomic);
+ if (summary.sched_rmid != INVALID_RMID) {
+ raw_spin_unlock(&__pkg_data(pmonr, pkg_data_lock));
+ return;
+ }
+
+ /* Do not try to obtain RMID if monr is not active. */
+ if (!prmid_summary__is_mon_active(summary)) {
+ raw_spin_unlock(&__pkg_data(pmonr, pkg_data_lock));
+ return;
+ }
+
+ /*
+ * Can only fail if it was in (U)state.
+ * Try to obtain a free prmid and go to (A)state, if not possible,
+ * it should go to (I)state.
+ */
+ WARN_ON_ONCE(!__pmonr__in_ustate(pmonr));
+
+ if (!list_empty(&__pkg_data(pmonr, free_prmids_pool))) {
+ /* Failed to obtain an valid rmid in this package for this
+ * monr. In next patches it will transition to (I)state.
+ * For now, stay in (U)state (do nothing)..
+ */
+ } else {
+ /* Transition to (A)state using free prmid. */
+ __pmonr__ustate_to_astate(
+ pmonr,
+ list_first_entry(&__pkg_data(pmonr, free_prmids_pool),
+ struct prmid, pool_entry));
+ }
+ raw_spin_unlock(&__pkg_data(pmonr, pkg_data_lock));
+}
+
+static inline void __assert_monr_is_leaf(struct monr *monr)
+{
+ int i;
+
+ monr_hrchy_assert_held_mutexes();
+ monr_hrchy_assert_held_raw_spin_locks();
+
+ cqm_pkg_id_for_each_online(i)
+ WARN_ON_ONCE(!__pmonr__in_ustate(monr->pmonrs[i]));
+
+ WARN_ON_ONCE(!list_empty(&monr->children));
+}
+
+static inline void
+__monr_hrchy_insert_leaf(struct monr *monr, struct monr *parent)
+{
+ monr_hrchy_assert_held_mutexes();
+ monr_hrchy_assert_held_raw_spin_locks();
+
+ __assert_monr_is_leaf(monr);
+
+ list_add_tail(&monr->parent_entry, &parent->children);
+ monr->parent = parent;
+}
+
+static inline void
+__monr_hrchy_remove_leaf(struct monr *monr)
+{
+ /* Since root cannot be removed, monr must have a parent */
+ WARN_ON_ONCE(!monr->parent);
+
+ monr_hrchy_assert_held_mutexes();
+ monr_hrchy_assert_held_raw_spin_locks();
+
+ __assert_monr_is_leaf(monr);
+
+ list_del_init(&monr->parent_entry);
+ monr->parent = NULL;
+}
+
+static int __monr_hrchy_attach_cpu_event(struct perf_event *event)
+{
+ lockdep_assert_held(&cqm_mutex);
+ WARN_ON_ONCE(monr_from_event(event));
+
+ event_set_monr(event, monr_hrchy_root);
+ return 0;
+}
+
+/* task events are always leaves in the monr_hierarchy */
+static int __monr_hrchy_attach_task_event(struct perf_event *event,
+ struct monr *parent_monr)
+{
+ struct monr *monr;
+ unsigned long flags;
+ int i;
+
+ lockdep_assert_held(&cqm_mutex);
+
+ monr = monr_alloc();
+ if (IS_ERR(monr))
+ return PTR_ERR(monr);
+ event_set_monr(event, monr);
+ monr->mon_event_group = event;
+
+ monr_hrchy_acquire_locks(flags, i);
+ __monr_hrchy_insert_leaf(monr, parent_monr);
+ __monr__set_mon_active(monr);
+ monr_hrchy_release_locks(flags, i);
+
+ return 0;
+}
+
+/*
+ * Find appropriate position in hierarchy and set monr. Create new
+ * monr if necessary.
+ * Locks rmid hrchy.
+ */
+static int monr_hrchy_attach_event(struct perf_event *event)
+{
+ struct monr *monr_parent;
+
+ if (!event->cgrp && !(event->attach_state & PERF_ATTACH_TASK))
+ return __monr_hrchy_attach_cpu_event(event);
+
+ /* Two-levels hierarchy: Root and all event monr underneath it. */
+ monr_parent = monr_hrchy_root;
+ return __monr_hrchy_attach_task_event(event, monr_parent);
+}
+
/*
* Determine if @a and @b measure the same set of tasks.
*
@@ -291,7 +661,7 @@ static bool __match_event(struct perf_event *a, struct perf_event *b)
return false;
#endif
- /* If not task event, we're machine wide */
+ /* If not task event, it's a a cgroup or a non-task cpu event. */
if (!(b->attach_state & PERF_ATTACH_TASK))
return true;
@@ -310,69 +680,51 @@ static bool __match_event(struct perf_event *a, struct perf_event *b)
return false;
}
-struct rmid_read {
- u32 rmid;
- atomic64_t value;
-};
-
static struct pmu intel_cqm_pmu;
/*
* Find a group and setup RMID.
*
- * If we're part of a group, we use the group's RMID.
+ * If we're part of a group, we use the group's monr.
*/
-static void intel_cqm_setup_event(struct perf_event *event,
- struct perf_event **group)
+static int
+intel_cqm_setup_event(struct perf_event *event, struct perf_event **group)
{
struct perf_event *iter;
- bool conflict = false;
- u32 rmid;
+ struct monr *monr;
+ *group = NULL;
- list_for_each_entry(iter, &cache_groups, hw.cqm_event_groups_entry) {
- rmid = iter->hw.cqm_rmid;
+ lockdep_assert_held(&cqm_mutex);
+ list_for_each_entry(iter, &cache_groups, hw.cqm_event_groups_entry) {
+ monr = monr_from_event(iter);
if (__match_event(iter, event)) {
- /* All tasks in a group share an RMID */
- event->hw.cqm_rmid = rmid;
+ /* All tasks in a group share an monr. */
+ event_set_monr(event, monr);
*group = iter;
- return;
+ return 0;
}
}
-
- if (conflict)
- rmid = INVALID_RMID;
- else
- rmid = __get_rmid();
-
- event->hw.cqm_rmid = rmid;
+ /*
+ * Since no match was found, create a new monr and set this
+ * event as head of a new cache group. All events in this cache group
+ * will share the monr.
+ */
+ return monr_hrchy_attach_event(event);
}
+/* Read current package immediately and remote pkg (if any) from cache. */
static void intel_cqm_event_read(struct perf_event *event)
{
- unsigned long flags;
- u32 rmid;
- u64 val;
+ union prmid_summary summary;
+ struct prmid *prmid;
u16 pkg_id = topology_physical_package_id(smp_processor_id());
+ struct pmonr *pmonr = monr_from_event(event)->pmonrs[pkg_id];
- raw_spin_lock_irqsave(&cqm_pkgs_data[pkg_id]->pkg_data_lock, flags);
- rmid = event->hw.cqm_rmid;
-
- if (!__rmid_valid(rmid))
- goto out;
-
- val = __rmid_read(rmid);
-
- /*
- * Ignore this reading on error states and do not update the value.
- */
- if (val & (RMID_VAL_ERROR | RMID_VAL_UNAVAIL))
- goto out;
-
- local64_set(&event->count, val);
-out:
- raw_spin_unlock_irqrestore(
- &cqm_pkgs_data[pkg_id]->pkg_data_lock, flags);
+ summary.value = atomic64_read(&pmonr->prmid_summary_atomic);
+ prmid = __prmid_from_rmid(pkg_id, summary.read_rmid);
+ cqm_prmid_update(prmid);
+ local64_set(&event->count, atomic64_read(&prmid->last_read_value));
}
static inline bool cqm_group_leader(struct perf_event *event)
@@ -380,52 +732,81 @@ static inline bool cqm_group_leader(struct perf_event *event)
return !list_empty(&event->hw.cqm_event_groups_entry);
}
-static void intel_cqm_event_start(struct perf_event *event, int mode)
+static inline void __intel_cqm_event_start(
+ struct perf_event *event, union prmid_summary summary)
{
u16 pkg_id = topology_physical_package_id(smp_processor_id());
if (!(event->hw.state & PERF_HES_STOPPED))
return;
event->hw.state &= ~PERF_HES_STOPPED;
- __update_pqr_prmid(__prmid_from_rmid(pkg_id, event->hw.cqm_rmid));
+ __update_pqr_prmid(__prmid_from_rmid(pkg_id, summary.sched_rmid));
+}
+
+static void intel_cqm_event_start(struct perf_event *event, int mode)
+{
+ union prmid_summary summary;
+ u16 pkg_id = topology_physical_package_id(smp_processor_id());
+ struct pmonr *pmonr = monr_from_event(event)->pmonrs[pkg_id];
+
+ /* Utilize most up to date pmonr summary. */
+ monr_hrchy_get_next_prmid_summary(pmonr);
+ summary.value = atomic64_read(&pmonr->prmid_summary_atomic);
+ __intel_cqm_event_start(event, summary);
}
static void intel_cqm_event_stop(struct perf_event *event, int mode)
{
+ union prmid_summary summary;
u16 pkg_id = topology_physical_package_id(smp_processor_id());
+ struct pmonr *root_pmonr = monr_hrchy_root->pmonrs[pkg_id];
+
if (event->hw.state & PERF_HES_STOPPED)
return;
event->hw.state |= PERF_HES_STOPPED;
- intel_cqm_event_read(event);
- __update_pqr_prmid(__prmid_from_rmid(pkg_id, 0));
+
+ summary.value = atomic64_read(&root_pmonr->prmid_summary_atomic);
+ /* Occupancy of CQM events is obtained at read. No need to read
+ * when event is stopped since read on inactive cpus succeed.
+ */
+ __update_pqr_prmid(__prmid_from_rmid(pkg_id, summary.sched_rmid));
}
static int intel_cqm_event_add(struct perf_event *event, int mode)
{
- unsigned long flags;
- u32 rmid;
+ struct monr *monr;
+ struct pmonr *pmonr;
+ union prmid_summary summary;
u16 pkg_id = topology_physical_package_id(smp_processor_id());
- raw_spin_lock_irqsave(&cqm_pkgs_data[pkg_id]->pkg_data_lock, flags);
+ monr = monr_from_event(event);
+ pmonr = monr->pmonrs[pkg_id];
event->hw.state = PERF_HES_STOPPED;
- rmid = event->hw.cqm_rmid;
- if (__rmid_valid(rmid) && (mode & PERF_EF_START))
- intel_cqm_event_start(event, mode);
+ /* Utilize most up to date pmonr summary. */
+ monr_hrchy_get_next_prmid_summary(pmonr);
+ summary.value = atomic64_read(&pmonr->prmid_summary_atomic);
+
+ if (!prmid_summary__is_mon_active(summary))
+ return -1;
- raw_spin_unlock_irqrestore(
- &cqm_pkgs_data[pkg_id]->pkg_data_lock, flags);
+ if (mode & PERF_EF_START)
+ __intel_cqm_event_start(event, summary);
+
+ /* (I)state pmonrs cannot report occupancy for themselves. */
return 0;
}
static void intel_cqm_event_destroy(struct perf_event *event)
{
struct perf_event *group_other = NULL;
+ struct monr *monr;
+ int i;
+ unsigned long flags;
mutex_lock(&cqm_mutex);
-
/*
* If there's another event in this group...
*/
@@ -435,33 +816,56 @@ static void intel_cqm_event_destroy(struct perf_event *event)
hw.cqm_event_group_entry);
list_del(&event->hw.cqm_event_group_entry);
}
-
/*
* And we're the group leader..
*/
- if (cqm_group_leader(event)) {
- /*
- * If there was a group_other, make that leader, otherwise
- * destroy the group and return the RMID.
- */
- if (group_other) {
- list_replace(&event->hw.cqm_event_groups_entry,
- &group_other->hw.cqm_event_groups_entry);
- } else {
- u32 rmid = event->hw.cqm_rmid;
-
- if (__rmid_valid(rmid))
- __put_rmid(rmid);
- list_del(&event->hw.cqm_event_groups_entry);
- }
+ if (!cqm_group_leader(event))
+ goto exit;
+
+ monr = monr_from_event(event);
+
+ /*
+ * If there was a group_other, make that leader, otherwise
+ * destroy the group and return the RMID.
+ */
+ if (group_other) {
+ /* Update monr reference to group head. */
+ monr->mon_event_group = group_other;
+ list_replace(&event->hw.cqm_event_groups_entry,
+ &group_other->hw.cqm_event_groups_entry);
+ goto exit;
}
+ /*
+ * Event is the only event in cache group.
+ */
+
+ event_set_monr(event, NULL);
+ list_del(&event->hw.cqm_event_groups_entry);
+
+ if (monr__is_root(monr))
+ goto exit;
+
+ /* Transition all pmonrs to (U)state. */
+ monr_hrchy_acquire_locks(flags, i);
+
+ cqm_pkg_id_for_each_online(i)
+ __pmonr__to_ustate(monr->pmonrs[i]);
+
+ __monr__clear_mon_active(monr);
+ monr->mon_event_group = NULL;
+ __monr_hrchy_remove_leaf(monr);
+ monr_hrchy_release_locks(flags, i);
+
+ monr_dealloc(monr);
+exit:
mutex_unlock(&cqm_mutex);
}
static int intel_cqm_event_init(struct perf_event *event)
{
struct perf_event *group = NULL;
+ int ret;
if (event->attr.type != intel_cqm_pmu.type)
return -ENOENT;
@@ -488,7 +892,11 @@ static int intel_cqm_event_init(struct perf_event *event)
/* Will also set rmid */
- intel_cqm_setup_event(event, &group);
+ ret = intel_cqm_setup_event(event, &group);
+ if (ret) {
+ mutex_unlock(&cqm_mutex);
+ return ret;
+ }
if (group) {
list_add_tail(&event->hw.cqm_event_group_entry,
@@ -697,6 +1105,12 @@ static int __init intel_cqm_init(void)
goto error;
}
+ monr_hrchy_root = monr_alloc();
+ if (IS_ERR(monr_hrchy_root)) {
+ ret = PTR_ERR(monr_hrchy_root);
+ goto error;
+ }
+
/* Select the minimum of the maximum rmids to use as limit for
* threshold. XXX: per-package threshold.
*/
@@ -705,6 +1119,7 @@ static int __init intel_cqm_init(void)
min_max_rmid = cqm_pkgs_data[i]->max_rmid;
intel_cqm_setup_pkg_prmid_pools(i);
}
+ monr_hrchy_root->flags |= MONR_MON_ACTIVE;
/*
* A reasonable upper limit on the max threshold is the number
diff --git a/arch/x86/events/intel/cqm.h b/arch/x86/events/intel/cqm.h
index a25d49b..81092f2 100644
--- a/arch/x86/events/intel/cqm.h
+++ b/arch/x86/events/intel/cqm.h
@@ -45,14 +45,111 @@ static unsigned int __rmid_min_update_time = RMID_DEFAULT_MIN_UPDATE_TIME;
static inline int cqm_prmid_update(struct prmid *prmid);
+/*
+ * union prmid_summary: Machine-size summary of a pmonr's prmid state.
+ * @value: One word accesor.
+ * @rmid: rmid for prmid.
+ * @sched_rmid: The rmid to write in the PQR MSR.
+ * @read_rmid: The rmid to read occupancy from.
+ *
+ * The prmid_summarys are read atomically and without the need of LOCK
+ * instructions during event and group scheduling in task context switch.
+ * They are set when a prmid change state and allow lock-free fast paths for
+ * RMID scheduling and RMID read for the common case when prmid does not need
+ * to change state.
+ * The combination of values in sched_rmid and read_rmid indicate the state of
+ * the associated pmonr (see pmonr comments) as follows:
+ * pmonr state
+ * | (A)state (U)state
+ * ----------------------------------------------------------------------------
+ * sched_rmid | pmonr.prmid INVALID_RMID
+ * read_rmid | pmonr.prmid INVALID_RMID
+ * (or 0)
+ *
+ * The combination sched_rmid == INVALID_RMID and read_rmid == 0 for (U)state
+ * denotes that the flag MONR_MON_ACTIVE is set in the monr associated with
+ * the pmonr for this prmid_summary.
+ */
+union prmid_summary {
+ long long value;
+ struct {
+ u32 sched_rmid;
+ u32 read_rmid;
+ };
+};
+
# define INVALID_RMID (-1)
+/* A pmonr in (U)state has no sched_rmid, read_rmid can be 0 or INVALID_RMID
+ * depending on whether monitoring is active or not.
+ */
+inline bool prmid_summary__is_ustate(union prmid_summary summ)
+{
+ return summ.sched_rmid == INVALID_RMID;
+}
+
+inline bool prmid_summary__is_mon_active(union prmid_summary summ)
+{
+ /* If not in (U)state, then MONR_MON_ACTIVE must be set. */
+ return summ.sched_rmid != INVALID_RMID ||
+ summ.read_rmid == 0;
+}
+
+struct monr;
+
+/* struct pmonr: Node of per-package hierarchy of MONitored Resources.
+ * @prmid: The prmid of this pmonr -when in (A)state-.
+ * @rotation_entry: List entry to attach to astate_pmonrs_lru
+ * in pkg_data.
+ * @monr: The monr that contains this pmonr.
+ * @pkg_id: Auxiliar variable with pkg id for this pmonr.
+ * @prmid_summary_atomic: Atomic accesor to store a union prmid_summary
+ * that represent the state of this pmonr.
+ *
+ * A pmonr forms a per-package hierarchy of prmids. Each one represents a
+ * resource to be monitored and can hold a prmid. Due to rmid scarcity,
+ * rmids can be recycled and rotated. When a rmid is not available for this
+ * pmonr, the pmonr utilizes the rmid of its ancestor.
+ * A pmonr is always in one of the following states:
+ * - (A)ctive: Has @prmid assigned, @ancestor_pmonr must be NULL.
+ * - (U)nused: No @ancestor_pmonr and no @prmid, hence no available
+ * prmid and no inhering one either. Not in rotation list.
+ * This state is unschedulable and a prmid
+ * should be found (either o free one or ancestor's) before
+ * scheduling a thread with (U)state pmonr in
+ * a cpu in this package.
+ *
+ * The state transitions are:
+ * (U) : The initial state. Starts there after allocation.
+ * (U) -> (A): If on first sched (or initialization) pmonr receives a prmid.
+ * (A) -> (U): On destruction of monr.
+ *
+ * Each pmonr is contained by a monr.
+ */
+struct pmonr {
+
+ struct prmid *prmid;
+
+ struct monr *monr;
+ struct list_head rotation_entry;
+
+ u16 pkg_id;
+
+ /* all writers are sync'ed by package's lock. */
+ atomic64_t prmid_summary_atomic;
+};
+
/*
* struct pkg_data: Per-package CQM data.
* @max_rmid: Max rmid valid for cpus in this package.
* @prmids_by_rmid: Utility mapping between rmid values and prmids.
* XXX: Make it an array of prmids.
* @free_prmid_pool: Free prmids.
+ * @active_prmid_pool: prmids associated with a (A)state pmonr.
+ * @nopmonr_limbo_prmid_pool: prmids in limbo state that are not referenced
+ * by a pmonr.
+ * @astate_pmonrs_lru: pmonrs in (A)state. LRU in increasing order of
+ * pmonr.last_enter_astate.
* @pkg_data_mutex: Hold for stability when modifying pmonrs
* hierarchy.
* @pkg_data_lock: Hold to protect variables that may be accessed
@@ -71,6 +168,12 @@ struct pkg_data {
* Pools of prmids used in rotation logic.
*/
struct list_head free_prmids_pool;
+ /* Can be modified during task switch with (U)state -> (A)state. */
+ struct list_head active_prmids_pool;
+ /* Only modified during rotation logic and deletion. */
+ struct list_head nopmonr_limbo_prmids_pool;
+
+ struct list_head astate_pmonrs_lru;
struct mutex pkg_data_mutex;
raw_spinlock_t pkg_data_lock;
@@ -78,6 +181,52 @@ struct pkg_data {
int rotation_cpu;
};
+/*
+ * Flags for monr.
+ */
+#define MONR_MON_ACTIVE 0x1
+
+/*
+ * struct monr: MONitored Resource.
+ * @flags: Flags field for monr (XXX: More flags will be added
+ * with MBM).
+ * @mon_event_group: The head of event's group that use this monr, if any.
+ * @parent: Parent in monr hierarchy.
+ * @children: List of children in monr hierarchy.
+ * @parent_entry: Entry in parent's children list.
+ * @pmonrs: Per-package pmonr for this monr.
+ *
+ * Each cgroup or thread that requires a RMID will have a corresponding
+ * monr in the system-wide hierarchy reflecting it's position in the
+ * cgroup/thread hierarchy.
+ * An monr is assigned to every CQM event and/or monitored cgroups when
+ * monitoring is activated and that instance's address do not change during
+ * the lifetime of the event or cgroup.
+ *
+ * On creation, the monr has flags cleared and all its pmonrs in (U)state.
+ * The flag MONR_MON_ACTIVE must be set to enable any transition out of
+ * (U)state to occur.
+ */
+struct monr {
+ u16 flags;
+ /* Back reference pointers */
+ struct perf_event *mon_event_group;
+
+ struct monr *parent;
+ struct list_head children;
+ struct list_head parent_entry;
+ struct pmonr *pmonrs[PQR_MAX_NR_PKGS];
+};
+
+/*
+ * Root for system-wide hierarchy of monr.
+ * A per-package raw_spin_lock protects changes to the per-pkg elements of
+ * the monr hierarchy.
+ * To modify the monr hierarchy, must hold all locks in each package
+ * using packaged-id as nesting parameter.
+ */
+extern struct monr *monr_hrchy_root;
+
extern struct pkg_data *cqm_pkgs_data[PQR_MAX_NR_PKGS];
static inline u16 __cqm_pkgs_data_next_online(u16 pkg_id)
diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h
index 5eb7dea..bf29258 100644
--- a/include/linux/perf_event.h
+++ b/include/linux/perf_event.h
@@ -120,7 +120,7 @@ struct hw_perf_event {
};
#ifdef CONFIG_INTEL_RDT
struct { /* intel_cqm */
- u32 cqm_rmid;
+ void *cqm_monr;
struct list_head cqm_event_group_entry;
struct list_head cqm_event_groups_entry;
};
--
2.8.0.rc3.226.g39d4020
[toc] | [prev] | [next] | [standalone]
| From | Vikas Shivappa <vikas.shivappa@linux.intel.com> |
|---|---|
| Date | 2016-04-29 23:10 +0200 |
| Message-ID | <rtnPH-4gp-3@gated-at.bofh.it> |
| In reply to | #1390741 |
On Thu, 28 Apr 2016, David Carrillo-Cisneros wrote: > This series introduces the next iteration of kernel support for the > Cache QoS Monitoring (CQM) technology available in Intel Xeon processors. Wondering what is the kernel version this compiles on ? Thanks, Vikas > > One of the main limitations of the previous version is the inability > to simultaneously monitor: > 1) cpu event and any other event in that cpu. > 2) cgroup events for cgroups in same descendancy line. > 3) cgroup events and any thread event of a cgroup in the same > descendancy line. > > Another limitation is that monitoring for a cgroup was enabled/disabled by > the existence of a perf event for that cgroup. Since the event > llc_occupancy measures changes in occupancy rather than total occupancy, > in order to read meaningful llc_occupancy values, an event should be > enabled for a long enough period of time. The overhead in context switches > caused by the perf events is undesired in some sensitive scenarios. > > This series of patches addresses the shortcomings mentioned above and, > add some other improvements. The main changes are: > - No more potential conflicts between different events. New > version builds a hierarchy of RMIDs that captures the dependency > between monitored cgroups. llc_occupancy for cgroup is the sum of > llc_occupancies for that cgroup RMID and all other RMIDs in the > cgroups subtree (both monitored cgroups and threads). > > - A cgroup integration that allows to monitor the a cgroup without > creating a perf event, decreasing the context switch overhead. > Monitoring is controlled by a boolean cgroup subsystem attribute > in each perf cgroup, this is: > > echo 1 > cgroup_path/perf_event.cqm_cont_monitoring > > starts CQM monitoring whether or not there is a perf_event > attached to the cgroup. Setting the attribute to 0 makes > monitoring dependent on the existence of a perf_event. > A perf_event is always required in order to read llc_occupancy. > This cgroup integration uses Intel's PQR code and is intended to > be used by upcoming versions of Intel's CAT. > > - A more stable rotation algorithm: New algorithm uses SLOs that > guarantee: > - A minimum of enabled time for monitored cgroups and > threads. > - A maximum time disabled before error is introduced by > reusing dirty RMIDs. > - A minimum rate at which RMIDs recycling must progress. > > - Reduced impact of stealing/rotation of RMIDs: The new algorithm > accounts the residual occupancy held by limbo RMIDs towards the > former owner of the limbo RMID, decreasing the error introduced > by RMID rotation. > It also allows a limbo RMID to be reused by its former owner when > appropriate, decreasing the potential error of reusing dirty RMIDs > and allowing to make progress even if most limbo RMIDs do not > drop occupancy fast enough. > > - Elimination of pmu::count: perf generic's perf_event_count() > perform a quick add of atomic types. The introduction of > pmu::count in the previous CQM series to read occupancy for thread > events changed the behavior of perf_event_count() by performing a > potentially slow IPI and write/read to MSR. It also made pmu::read > to have different behaviors depending on whether the event was a > cpu/cgroup event or a thread. This patches serie removes the custom > pmu::count from CQM and provides a consistent behavior for all > calls of perf_event_read . > > - Added error return for pmu::read: Reads to CQM events may fail > due to stealing of RMIDs, even after successfully adding an event > to a PMU. This patch series expands pmu::read with an int return > value and propagates the error to callers that can fail > (ie. perf_read). > The ability to fail of pmu::read is consistent with the recent > changes that allow perf_event_read to fail for transactional > reading of event groups. > > - Introduces the field pmu_event_flags that contain flags set by > the PMU to signal variations on the default behavior to perf's > generic code. In this series, three flags are introduced: > - PERF_CGROUP_NO_RECURSION : Signals generic code to add > events of the cgroup ancestors of a cgroup. > - PERF_INACTIVE_CPU_READ_PKG: Signals generic coda that > this CPU event can be read in any CPU in its event::cpu's > package, even if the event is not active. > - PERF_INACTIVE_EV_READ_ANY_CPU: Signals generic code that > this event can be read in any CPU in any package in the > system even if the event is not active. > Using the above flags takes advantage of the CQM's hw ability to > read llc_occupancy even when the associated perf event is not > running in a CPU. > > This patch series also updates the perf tool to fix error handling and to > better handle the idiosyncrasies of snapshot and per-pkg events. > > David Carrillo-Cisneros (31): > perf/x86/intel/cqm: temporarily remove MBM from CQM and cleanup > perf/x86/intel/cqm: remove check for conflicting events > perf/x86/intel/cqm: remove all code for rotation of RMIDs > perf/x86/intel/cqm: make read of RMIDs per package (Temporal) > perf/core: remove unused pmu->count > x86/intel,cqm: add CONFIG_INTEL_RDT configuration flag and refactor > PQR > perf/x86/intel/cqm: separate CQM PMU's attributes from x86 PMU > perf/x86/intel/cqm: prepare for next patches > perf/x86/intel/cqm: add per-package RMIDs, data and locks > perf/x86/intel/cqm: basic RMID hierarchy with per package rmids > perf/x86/intel/cqm: (I)state and limbo prmids > perf/x86/intel/cqm: add per-package RMID rotation > perf/x86/intel/cqm: add polled update of RMID's llc_occupancy > perf/x86/intel/cqm: add preallocation of anodes > perf/core: add hooks to expose architecture specific features in > perf_cgroup > perf/x86/intel/cqm: add cgroup support > perf/core: adding pmu::event_terminate > perf/x86/intel/cqm: use pmu::event_terminate > perf/core: introduce PMU event flag PERF_CGROUP_NO_RECURSION > x86/intel/cqm: use PERF_CGROUP_NO_RECURSION in CQM > perf/x86/intel/cqm: handle inherit event and inherit_stat flag > perf/x86/intel/cqm: introduce read_subtree > perf/core: introduce PERF_INACTIVE_*_READ_* flags > perf/x86/intel/cqm: use PERF_INACTIVE_*_READ_* flags in CQM > sched: introduce the finish_arch_pre_lock_switch() scheduler hook > perf/x86/intel/cqm: integrate CQM cgroups with scheduler > perf/core: add perf_event cgroup hooks for subsystem attributes > perf/x86/intel/cqm: add CQM attributes to perf_event cgroup > perf,perf/x86,perf/powerpc,perf/arm,perf/*: add int error return to > pmu::read > perf,perf/x86: add hook perf_event_arch_exec > perf/stat: revamp error handling for snapshot and per_pkg events > > Stephane Eranian (1): > perf/stat: fix bug in handling events in error state > > arch/alpha/kernel/perf_event.c | 3 +- > arch/arc/kernel/perf_event.c | 3 +- > arch/arm64/include/asm/hw_breakpoint.h | 2 +- > arch/arm64/kernel/hw_breakpoint.c | 3 +- > arch/metag/kernel/perf/perf_event.c | 5 +- > arch/mips/kernel/perf_event_mipsxx.c | 3 +- > arch/powerpc/include/asm/hw_breakpoint.h | 2 +- > arch/powerpc/kernel/hw_breakpoint.c | 3 +- > arch/powerpc/perf/core-book3s.c | 11 +- > arch/powerpc/perf/core-fsl-emb.c | 5 +- > arch/powerpc/perf/hv-24x7.c | 5 +- > arch/powerpc/perf/hv-gpci.c | 3 +- > arch/s390/kernel/perf_cpum_cf.c | 5 +- > arch/s390/kernel/perf_cpum_sf.c | 3 +- > arch/sh/include/asm/hw_breakpoint.h | 2 +- > arch/sh/kernel/hw_breakpoint.c | 3 +- > arch/sparc/kernel/perf_event.c | 2 +- > arch/tile/kernel/perf_event.c | 3 +- > arch/x86/Kconfig | 6 + > arch/x86/events/amd/ibs.c | 2 +- > arch/x86/events/amd/iommu.c | 5 +- > arch/x86/events/amd/uncore.c | 3 +- > arch/x86/events/core.c | 3 +- > arch/x86/events/intel/Makefile | 3 +- > arch/x86/events/intel/bts.c | 3 +- > arch/x86/events/intel/cqm.c | 3847 +++++++++++++++++++++--------- > arch/x86/events/intel/cqm.h | 519 ++++ > arch/x86/events/intel/cstate.c | 3 +- > arch/x86/events/intel/pt.c | 3 +- > arch/x86/events/intel/rapl.c | 3 +- > arch/x86/events/intel/uncore.c | 3 +- > arch/x86/events/intel/uncore.h | 2 +- > arch/x86/events/msr.c | 3 +- > arch/x86/include/asm/hw_breakpoint.h | 2 +- > arch/x86/include/asm/perf_event.h | 41 + > arch/x86/include/asm/pqr_common.h | 74 + > arch/x86/include/asm/processor.h | 4 + > arch/x86/kernel/cpu/Makefile | 4 + > arch/x86/kernel/cpu/pqr_common.c | 43 + > arch/x86/kernel/hw_breakpoint.c | 3 +- > arch/x86/kvm/pmu.h | 10 +- > drivers/bus/arm-cci.c | 3 +- > drivers/bus/arm-ccn.c | 3 +- > drivers/perf/arm_pmu.c | 3 +- > include/linux/perf_event.h | 91 +- > kernel/events/core.c | 170 +- > kernel/sched/core.c | 1 + > kernel/sched/sched.h | 3 + > kernel/trace/bpf_trace.c | 5 +- > tools/perf/builtin-stat.c | 43 +- > tools/perf/util/counts.h | 19 + > tools/perf/util/evsel.c | 44 +- > tools/perf/util/evsel.h | 8 +- > tools/perf/util/stat.c | 35 +- > 54 files changed, 3746 insertions(+), 1337 deletions(-) > create mode 100644 arch/x86/events/intel/cqm.h > create mode 100644 arch/x86/include/asm/pqr_common.h > create mode 100644 arch/x86/kernel/cpu/pqr_common.c > > -- > 2.8.0.rc3.226.g39d4020 > >
[toc] | [prev] | [next] | [standalone]
| From | David Carrillo-Cisneros <davidcc@google.com> |
|---|---|
| Date | 2016-04-29 23:20 +0200 |
| Message-ID | <rtnZo-4me-21@gated-at.bofh.it> |
| In reply to | #1391427 |
peterz/queue perf/core On Fri, Apr 29, 2016 at 2:06 PM Vikas Shivappa <vikas.shivappa@linux.intel.com> wrote: > > > > On Thu, 28 Apr 2016, David Carrillo-Cisneros wrote: > > > This series introduces the next iteration of kernel support for the > > Cache QoS Monitoring (CQM) technology available in Intel Xeon processors. > > Wondering what is the kernel version this compiles on ? > > Thanks, > Vikas > > > > > One of the main limitations of the previous version is the inability > > to simultaneously monitor: > > 1) cpu event and any other event in that cpu. > > 2) cgroup events for cgroups in same descendancy line. > > 3) cgroup events and any thread event of a cgroup in the same > > descendancy line. > > > > Another limitation is that monitoring for a cgroup was enabled/disabled by > > the existence of a perf event for that cgroup. Since the event > > llc_occupancy measures changes in occupancy rather than total occupancy, > > in order to read meaningful llc_occupancy values, an event should be > > enabled for a long enough period of time. The overhead in context switches > > caused by the perf events is undesired in some sensitive scenarios. > > > > This series of patches addresses the shortcomings mentioned above and, > > add some other improvements. The main changes are: > > - No more potential conflicts between different events. New > > version builds a hierarchy of RMIDs that captures the dependency > > between monitored cgroups. llc_occupancy for cgroup is the sum of > > llc_occupancies for that cgroup RMID and all other RMIDs in the > > cgroups subtree (both monitored cgroups and threads). > > > > - A cgroup integration that allows to monitor the a cgroup without > > creating a perf event, decreasing the context switch overhead. > > Monitoring is controlled by a boolean cgroup subsystem attribute > > in each perf cgroup, this is: > > > > echo 1 > cgroup_path/perf_event.cqm_cont_monitoring > > > > starts CQM monitoring whether or not there is a perf_event > > attached to the cgroup. Setting the attribute to 0 makes > > monitoring dependent on the existence of a perf_event. > > A perf_event is always required in order to read llc_occupancy. > > This cgroup integration uses Intel's PQR code and is intended to > > be used by upcoming versions of Intel's CAT. > > > > - A more stable rotation algorithm: New algorithm uses SLOs that > > guarantee: > > - A minimum of enabled time for monitored cgroups and > > threads. > > - A maximum time disabled before error is introduced by > > reusing dirty RMIDs. > > - A minimum rate at which RMIDs recycling must progress. > > > > - Reduced impact of stealing/rotation of RMIDs: The new algorithm > > accounts the residual occupancy held by limbo RMIDs towards the > > former owner of the limbo RMID, decreasing the error introduced > > by RMID rotation. > > It also allows a limbo RMID to be reused by its former owner when > > appropriate, decreasing the potential error of reusing dirty RMIDs > > and allowing to make progress even if most limbo RMIDs do not > > drop occupancy fast enough. > > > > - Elimination of pmu::count: perf generic's perf_event_count() > > perform a quick add of atomic types. The introduction of > > pmu::count in the previous CQM series to read occupancy for thread > > events changed the behavior of perf_event_count() by performing a > > potentially slow IPI and write/read to MSR. It also made pmu::read > > to have different behaviors depending on whether the event was a > > cpu/cgroup event or a thread. This patches serie removes the custom > > pmu::count from CQM and provides a consistent behavior for all > > calls of perf_event_read . > > > > - Added error return for pmu::read: Reads to CQM events may fail > > due to stealing of RMIDs, even after successfully adding an event > > to a PMU. This patch series expands pmu::read with an int return > > value and propagates the error to callers that can fail > > (ie. perf_read). > > The ability to fail of pmu::read is consistent with the recent > > changes that allow perf_event_read to fail for transactional > > reading of event groups. > > > > - Introduces the field pmu_event_flags that contain flags set by > > the PMU to signal variations on the default behavior to perf's > > generic code. In this series, three flags are introduced: > > - PERF_CGROUP_NO_RECURSION : Signals generic code to add > > events of the cgroup ancestors of a cgroup. > > - PERF_INACTIVE_CPU_READ_PKG: Signals generic coda that > > this CPU event can be read in any CPU in its event::cpu's > > package, even if the event is not active. > > - PERF_INACTIVE_EV_READ_ANY_CPU: Signals generic code that > > this event can be read in any CPU in any package in the > > system even if the event is not active. > > Using the above flags takes advantage of the CQM's hw ability to > > read llc_occupancy even when the associated perf event is not > > running in a CPU. > > > > This patch series also updates the perf tool to fix error handling and to > > better handle the idiosyncrasies of snapshot and per-pkg events. > > > > David Carrillo-Cisneros (31): > > perf/x86/intel/cqm: temporarily remove MBM from CQM and cleanup > > perf/x86/intel/cqm: remove check for conflicting events > > perf/x86/intel/cqm: remove all code for rotation of RMIDs > > perf/x86/intel/cqm: make read of RMIDs per package (Temporal) > > perf/core: remove unused pmu->count > > x86/intel,cqm: add CONFIG_INTEL_RDT configuration flag and refactor > > PQR > > perf/x86/intel/cqm: separate CQM PMU's attributes from x86 PMU > > perf/x86/intel/cqm: prepare for next patches > > perf/x86/intel/cqm: add per-package RMIDs, data and locks > > perf/x86/intel/cqm: basic RMID hierarchy with per package rmids > > perf/x86/intel/cqm: (I)state and limbo prmids > > perf/x86/intel/cqm: add per-package RMID rotation > > perf/x86/intel/cqm: add polled update of RMID's llc_occupancy > > perf/x86/intel/cqm: add preallocation of anodes > > perf/core: add hooks to expose architecture specific features in > > perf_cgroup > > perf/x86/intel/cqm: add cgroup support > > perf/core: adding pmu::event_terminate > > perf/x86/intel/cqm: use pmu::event_terminate > > perf/core: introduce PMU event flag PERF_CGROUP_NO_RECURSION > > x86/intel/cqm: use PERF_CGROUP_NO_RECURSION in CQM > > perf/x86/intel/cqm: handle inherit event and inherit_stat flag > > perf/x86/intel/cqm: introduce read_subtree > > perf/core: introduce PERF_INACTIVE_*_READ_* flags > > perf/x86/intel/cqm: use PERF_INACTIVE_*_READ_* flags in CQM > > sched: introduce the finish_arch_pre_lock_switch() scheduler hook > > perf/x86/intel/cqm: integrate CQM cgroups with scheduler > > perf/core: add perf_event cgroup hooks for subsystem attributes > > perf/x86/intel/cqm: add CQM attributes to perf_event cgroup > > perf,perf/x86,perf/powerpc,perf/arm,perf/*: add int error return to > > pmu::read > > perf,perf/x86: add hook perf_event_arch_exec > > perf/stat: revamp error handling for snapshot and per_pkg events > > > > Stephane Eranian (1): > > perf/stat: fix bug in handling events in error state > > > > arch/alpha/kernel/perf_event.c | 3 +- > > arch/arc/kernel/perf_event.c | 3 +- > > arch/arm64/include/asm/hw_breakpoint.h | 2 +- > > arch/arm64/kernel/hw_breakpoint.c | 3 +- > > arch/metag/kernel/perf/perf_event.c | 5 +- > > arch/mips/kernel/perf_event_mipsxx.c | 3 +- > > arch/powerpc/include/asm/hw_breakpoint.h | 2 +- > > arch/powerpc/kernel/hw_breakpoint.c | 3 +- > > arch/powerpc/perf/core-book3s.c | 11 +- > > arch/powerpc/perf/core-fsl-emb.c | 5 +- > > arch/powerpc/perf/hv-24x7.c | 5 +- > > arch/powerpc/perf/hv-gpci.c | 3 +- > > arch/s390/kernel/perf_cpum_cf.c | 5 +- > > arch/s390/kernel/perf_cpum_sf.c | 3 +- > > arch/sh/include/asm/hw_breakpoint.h | 2 +- > > arch/sh/kernel/hw_breakpoint.c | 3 +- > > arch/sparc/kernel/perf_event.c | 2 +- > > arch/tile/kernel/perf_event.c | 3 +- > > arch/x86/Kconfig | 6 + > > arch/x86/events/amd/ibs.c | 2 +- > > arch/x86/events/amd/iommu.c | 5 +- > > arch/x86/events/amd/uncore.c | 3 +- > > arch/x86/events/core.c | 3 +- > > arch/x86/events/intel/Makefile | 3 +- > > arch/x86/events/intel/bts.c | 3 +- > > arch/x86/events/intel/cqm.c | 3847 +++++++++++++++++++++--------- > > arch/x86/events/intel/cqm.h | 519 ++++ > > arch/x86/events/intel/cstate.c | 3 +- > > arch/x86/events/intel/pt.c | 3 +- > > arch/x86/events/intel/rapl.c | 3 +- > > arch/x86/events/intel/uncore.c | 3 +- > > arch/x86/events/intel/uncore.h | 2 +- > > arch/x86/events/msr.c | 3 +- > > arch/x86/include/asm/hw_breakpoint.h | 2 +- > > arch/x86/include/asm/perf_event.h | 41 + > > arch/x86/include/asm/pqr_common.h | 74 + > > arch/x86/include/asm/processor.h | 4 + > > arch/x86/kernel/cpu/Makefile | 4 + > > arch/x86/kernel/cpu/pqr_common.c | 43 + > > arch/x86/kernel/hw_breakpoint.c | 3 +- > > arch/x86/kvm/pmu.h | 10 +- > > drivers/bus/arm-cci.c | 3 +- > > drivers/bus/arm-ccn.c | 3 +- > > drivers/perf/arm_pmu.c | 3 +- > > include/linux/perf_event.h | 91 +- > > kernel/events/core.c | 170 +- > > kernel/sched/core.c | 1 + > > kernel/sched/sched.h | 3 + > > kernel/trace/bpf_trace.c | 5 +- > > tools/perf/builtin-stat.c | 43 +- > > tools/perf/util/counts.h | 19 + > > tools/perf/util/evsel.c | 44 +- > > tools/perf/util/evsel.h | 8 +- > > tools/perf/util/stat.c | 35 +- > > 54 files changed, 3746 insertions(+), 1337 deletions(-) > > create mode 100644 arch/x86/events/intel/cqm.h > > create mode 100644 arch/x86/include/asm/pqr_common.h > > create mode 100644 arch/x86/kernel/cpu/pqr_common.c > > > > -- > > 2.8.0.rc3.226.g39d4020 > > > >
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web