Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1469414 > unrolled thread
| Started by | Matt Fleming <matt@codeblueprint.co.uk> |
|---|---|
| First post | 2016-08-24 15:20 +0200 |
| Last post | 2016-08-25 06:40 +0200 |
| Articles | 4 — 3 participants |
Back to article view | Back to linux.kernel
[PATCH v2] perf/x86/amd: Make HW_CACHE_REFERENCES and HW_CACHE_MISSES measure L2 Matt Fleming <matt@codeblueprint.co.uk> - 2016-08-24 15:20 +0200
Re: [PATCH v2] perf/x86/amd: Make HW_CACHE_REFERENCES and HW_CACHE_MISSES measure L2 Borislav Petkov <bp@alien8.de> - 2016-08-24 17:00 +0200
Re: [PATCH v2] perf/x86/amd: Make HW_CACHE_REFERENCES and HW_CACHE_MISSES measure L2 Peter Zijlstra <peterz@infradead.org> - 2016-08-24 20:30 +0200
Re: [PATCH v2] perf/x86/amd: Make HW_CACHE_REFERENCES and HW_CACHE_MISSES measure L2 Borislav Petkov <bp@alien8.de> - 2016-08-25 06:40 +0200
| From | Matt Fleming <matt@codeblueprint.co.uk> |
|---|---|
| Date | 2016-08-24 15:20 +0200 |
| Subject | [PATCH v2] perf/x86/amd: Make HW_CACHE_REFERENCES and HW_CACHE_MISSES measure L2 |
| Message-ID | <s9Gg1-2I8-13@gated-at.bofh.it> |
While the Intel PMU monitors the LLC when perf enables the
HW_CACHE_REFERENCES and HW_CACHE_MISSES events, these events monitor
L1 instruction cache fetches (0x0080) and instruction cache misses
(0x0081) on the AMD PMU.
This is extremely confusing when monitoring the same workload across
Intel and AMD machines, since parameters like,
$ perf stat -e cache-references,cache-misses
measure completely different things.
Instead, make the AMD PMU measure instruction/data cache and TLB fill
requests to the L2 and instruction/data cache and TLB misses in the L2
when HW_CACHE_REFERENCES and HW_CACHE_MISSES are enabled,
respectively. That way the events measure unified caches on both
platforms.
Signed-off-by: Matt Fleming <matt@codeblueprint.co.uk>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Ingo Molnar <mingo@kernel.org>
Cc: Borislav Petkov <bp@alien8.de>
Cc: <stable@vger.kernel.org>
---
Changes in v2:
- Update the KVM AMD PMU code
- Also measure TLB hits/misses in the L2
arch/x86/events/amd/core.c | 4 ++--
arch/x86/kvm/pmu_amd.c | 4 ++--
2 files changed, 4 insertions(+), 4 deletions(-)
diff --git a/arch/x86/events/amd/core.c b/arch/x86/events/amd/core.c
index e07a22bb9308..f5f4b3fbbbc2 100644
--- a/arch/x86/events/amd/core.c
+++ b/arch/x86/events/amd/core.c
@@ -119,8 +119,8 @@ static const u64 amd_perfmon_event_map[PERF_COUNT_HW_MAX] =
{
[PERF_COUNT_HW_CPU_CYCLES] = 0x0076,
[PERF_COUNT_HW_INSTRUCTIONS] = 0x00c0,
- [PERF_COUNT_HW_CACHE_REFERENCES] = 0x0080,
- [PERF_COUNT_HW_CACHE_MISSES] = 0x0081,
+ [PERF_COUNT_HW_CACHE_REFERENCES] = 0x077d,
+ [PERF_COUNT_HW_CACHE_MISSES] = 0x077e,
[PERF_COUNT_HW_BRANCH_INSTRUCTIONS] = 0x00c2,
[PERF_COUNT_HW_BRANCH_MISSES] = 0x00c3,
[PERF_COUNT_HW_STALLED_CYCLES_FRONTEND] = 0x00d0, /* "Decoder empty" event */
diff --git a/arch/x86/kvm/pmu_amd.c b/arch/x86/kvm/pmu_amd.c
index 39b91127ef07..cd944435dfbd 100644
--- a/arch/x86/kvm/pmu_amd.c
+++ b/arch/x86/kvm/pmu_amd.c
@@ -23,8 +23,8 @@
static struct kvm_event_hw_type_mapping amd_event_mapping[] = {
[0] = { 0x76, 0x00, PERF_COUNT_HW_CPU_CYCLES },
[1] = { 0xc0, 0x00, PERF_COUNT_HW_INSTRUCTIONS },
- [2] = { 0x80, 0x00, PERF_COUNT_HW_CACHE_REFERENCES },
- [3] = { 0x81, 0x00, PERF_COUNT_HW_CACHE_MISSES },
+ [2] = { 0x7d, 0x07, PERF_COUNT_HW_CACHE_REFERENCES },
+ [3] = { 0x7e, 0x07, PERF_COUNT_HW_CACHE_MISSES },
[4] = { 0xc2, 0x00, PERF_COUNT_HW_BRANCH_INSTRUCTIONS },
[5] = { 0xc3, 0x00, PERF_COUNT_HW_BRANCH_MISSES },
[6] = { 0xd0, 0x00, PERF_COUNT_HW_STALLED_CYCLES_FRONTEND },
--
2.7.3
[toc] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2016-08-24 17:00 +0200 |
| Subject | Re: [PATCH v2] perf/x86/amd: Make HW_CACHE_REFERENCES and HW_CACHE_MISSES measure L2 |
| Message-ID | <s9HON-3F4-7@gated-at.bofh.it> |
| In reply to | #1469414 |
On Wed, Aug 24, 2016 at 02:12:08PM +0100, Matt Fleming wrote:
> While the Intel PMU monitors the LLC when perf enables the
> HW_CACHE_REFERENCES and HW_CACHE_MISSES events, these events monitor
> L1 instruction cache fetches (0x0080) and instruction cache misses
> (0x0081) on the AMD PMU.
>
> This is extremely confusing when monitoring the same workload across
> Intel and AMD machines, since parameters like,
>
> $ perf stat -e cache-references,cache-misses
>
> measure completely different things.
>
> Instead, make the AMD PMU measure instruction/data cache and TLB fill
> requests to the L2 and instruction/data cache and TLB misses in the L2
> when HW_CACHE_REFERENCES and HW_CACHE_MISSES are enabled,
> respectively. That way the events measure unified caches on both
> platforms.
I'm still not really sure about this: we can't really compare L3 to L2
access patterns - it is almost as comparing apples to oranges. Can we
use the Intel L2 events instead?
I mean, this makes much more sense to me because:
* you *actually* compare the same cache levels
* you have L2 *everywhere* vs L3 (and L4) which are sometimes not present on
thin clients
People who want LLC can enable them with -e additionally...
Hmmm.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-08-24 20:30 +0200 |
| Subject | Re: [PATCH v2] perf/x86/amd: Make HW_CACHE_REFERENCES and HW_CACHE_MISSES measure L2 |
| Message-ID | <s9L61-62m-13@gated-at.bofh.it> |
| In reply to | #1469505 |
On Wed, Aug 24, 2016 at 04:55:14PM +0200, Borislav Petkov wrote: > On Wed, Aug 24, 2016 at 02:12:08PM +0100, Matt Fleming wrote: > > While the Intel PMU monitors the LLC when perf enables the > > HW_CACHE_REFERENCES and HW_CACHE_MISSES events, these events monitor > > L1 instruction cache fetches (0x0080) and instruction cache misses > > (0x0081) on the AMD PMU. > > > > This is extremely confusing when monitoring the same workload across > > Intel and AMD machines, since parameters like, > > > > $ perf stat -e cache-references,cache-misses > > > > measure completely different things. > > > > Instead, make the AMD PMU measure instruction/data cache and TLB fill > > requests to the L2 and instruction/data cache and TLB misses in the L2 > > when HW_CACHE_REFERENCES and HW_CACHE_MISSES are enabled, > > respectively. That way the events measure unified caches on both > > platforms. > > I'm still not really sure about this: we can't really compare L3 to L2 > access patterns - it is almost as comparing apples to oranges. Can we > use the Intel L2 events instead? They're not meant to be comparable between machines. I wouldn't even compare the LLC numbers between two different Intel parts. These events are meant to profile a workload on the machine you run them on. Big cache-miss/ref ratios indicate you loose performance because of the memory subsystem and or data structure layout. And afaict AMD parts, even those that have L3, cannot provide L3 numbers on a per task basis, so these L2 numbers are the best we have.
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2016-08-25 06:40 +0200 |
| Subject | Re: [PATCH v2] perf/x86/amd: Make HW_CACHE_REFERENCES and HW_CACHE_MISSES measure L2 |
| Message-ID | <s9UCl-4wb-17@gated-at.bofh.it> |
| In reply to | #1469630 |
(dropping stable@ from CC)
On Wed, Aug 24, 2016 at 08:27:06PM +0200, Peter Zijlstra wrote:
> They're not meant to be comparable between machines. I wouldn't even
> compare the LLC numbers between two different Intel parts.
>
> These events are meant to profile a workload on the machine you run them
> on. Big cache-miss/ref ratios indicate you loose performance because of
> the memory subsystem and or data structure layout.
Ah ok, then I've misunderstood Matt's justification in the commit message.
FWIW: Acked-by: Borislav Petkov <bp@suse.de>
Thanks.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web