Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1493241 > unrolled thread

Re: [PATCHv4 00/57] perf c2c: Add new tool to analyze cacheline contention on NUMA systems

Started byPeter Zijlstra <peterz@infradead.org>
First post2016-09-29 11:20 +0200
Last post2016-10-01 15:50 +0200
Articles 3 — 3 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCHv4 00/57] perf c2c: Add new tool to analyze cacheline  contention on NUMA systems Peter Zijlstra <peterz@infradead.org> - 2016-09-29 11:20 +0200
    Re: [PATCHv4 00/57] perf c2c: Add new tool to analyze cacheline  contention on NUMA systems Arnaldo Carvalho de Melo <acme@kernel.org> - 2016-09-29 17:00 +0200
    Re: [PATCHv4 00/57] perf c2c: Add new tool to analyze cacheline  contention on NUMA systems Joe Mario <jmario@redhat.com> - 2016-10-01 15:50 +0200

#1493241 — Re: [PATCHv4 00/57] perf c2c: Add new tool to analyze cacheline contention on NUMA systems

FromPeter Zijlstra <peterz@infradead.org>
Date2016-09-29 11:20 +0200
SubjectRe: [PATCHv4 00/57] perf c2c: Add new tool to analyze cacheline contention on NUMA systems
Message-ID<smFFw-28g-33@gated-at.bofh.it>
On Thu, Sep 22, 2016 at 05:36:28PM +0200, Jiri Olsa wrote:
> hi,
> sending new version of c2c patches (v3) originally posted in here:
>   http://lwn.net/Articles/588866/
> 
> I took the old set and reworked it to fit into current upstream code.
> It follows the same logic as original patch and provides (almost) the
> same stdio interface. In addition new TUI interface was added.
> 
> The perf c2c tool provides means for Shared Data C2C/HITM analysis.
> It allows you to track down the cacheline contentions. The tool is
> based on x86's load latency and precise store facility events provided
> by Intel CPUs.
> 
> The tool was tested by Joe Mario and has proven to be useful and found
> some cachelines contentions. Joe also wrote a blog about c2c tool with
> examples located in here:
> 
>   https://joemario.github.io/blog/2016/09/01/c2c-blog/
> 
> v4 changes:
>   - 4 patches already queued
>   - used u32 for c2c_stats instead of int [Stanislav]
>   - fixed NO_SLANG=1 compilation [Kim]
>   - add __hist_entry__snprintf helper [Arnaldo]
> 
> Code is also available in:
>   git://git.kernel.org/pub/scm/linux/kernel/git/jolsa/perf.git
>   perf/c2c_v4
> 
> Testing:
>   $ perf c2c record -a [workload]
>   $ perf c2c report [--stdio]
>   $ man perf-c2c
> 
>   It's most likely you won't generate any remote HITMs on common
>   laptops, so to get results for local HITMs please use:
> 
>   $ perf c2c report -d lcl [--stdio]

I'll just keep repeating; this is not the tool I want :-( I'll not block
this tool, but I also think its far less usable than it should've been.

  https://lkml.kernel.org/r/20151209093402.GM6356@twins.programming.kicks-ass.net

What I want is a tool that maps memop events (any PEBS memops) back to a
'type::member' form and sorts on that. That doesn't rely on the PEBS
'Data Linear Address' field, as that is useless for dynamically
allocated bits. Instead it would use the IP and Dwarf information to
deduce the 'type::member' of the memop.

I want pahole like output, showing me where the hits (green) and misses
(red) are in a structure.

I want to be able to 'perf memops report -EC task_struct' and see the
expanded task_struct (as per 'pahole -EC task_struct') annotated, not a
data address for each task in my workload (which could be 100+ and
entirely useless).

Currently this is somewhat involved, since Dwarf doesn't include type
information for all memops, so we'd have to disassemble and interpret,
which while tedious is possible.

However, afaik, Stephane has been working with their tools team to get
additional DWARF info to make this easier. Stephane, any updates on
that?

[toc] | [next] | [standalone]


#1493523

FromArnaldo Carvalho de Melo <acme@kernel.org>
Date2016-09-29 17:00 +0200
Message-ID<smKYx-5qg-3@gated-at.bofh.it>
In reply to#1493241
Em Thu, Sep 29, 2016 at 11:19:12AM +0200, Peter Zijlstra escreveu:
> On Thu, Sep 22, 2016 at 05:36:28PM +0200, Jiri Olsa wrote:
> > sending new version of c2c patches (v3) originally posted in here:
> >   http://lwn.net/Articles/588866/

> I'll just keep repeating; this is not the tool I want :-( I'll not block
> this tool, but I also think its far less usable than it should've been.

Well, I think its an experimentation with using that info, one that
people have been using and seemingly finding and fixing problems.

Requires more work than the way you describe(d) various times, tho,
indeed. :-\
 
>   https://lkml.kernel.org/r/20151209093402.GM6356@twins.programming.kicks-ass.net
 
> What I want is a tool that maps memop events (any PEBS memops) back to a
> 'type::member' form and sorts on that. That doesn't rely on the PEBS
> 'Data Linear Address' field, as that is useless for dynamically
> allocated bits. Instead it would use the IP and Dwarf information to
> deduce the 'type::member' of the memop.
 
> I want pahole like output, showing me where the hits (green) and misses
> (red) are in a structure.
 
> I want to be able to 'perf memops report -EC task_struct' and see the
> expanded task_struct (as per 'pahole -EC task_struct') annotated, not a
> data address for each task in my workload (which could be 100+ and
> entirely useless).
 
> Currently this is somewhat involved, since Dwarf doesn't include type
> information for all memops, so we'd have to disassemble and interpret,
> which while tedious is possible.
 
> However, afaik, Stephane has been working with their tools team to get
> additional DWARF info to make this easier. Stephane, any updates on
> that?

Yeah, that would be interesting to know, I for one, due to the c2c
effort + this other work Stephane mentioned some time ago, moved working
on such a pahole based tool to the backburner, lots of other patches to
review, test, even proof read to then process all the time :-\

- Arnaldo

[toc] | [prev] | [next] | [standalone]


#1494409

FromJoe Mario <jmario@redhat.com>
Date2016-10-01 15:50 +0200
Message-ID<snsPT-H2-5@gated-at.bofh.it>
In reply to#1493241
On 09/29/2016 05:19 AM, Peter Zijlstra wrote:
  
>
> What I want is a tool that maps memop events (any PEBS memops) back to a
> 'type::member' form and sorts on that. That doesn't rely on the PEBS
> 'Data Linear Address' field, as that is useless for dynamically
> allocated bits. Instead it would use the IP and Dwarf information to
> deduce the 'type::member' of the memop.
>
> I want pahole like output, showing me where the hits (green) and misses
> (red) are in a structure.

I agree that would give valuable insight, but it needs to be
in addition to what this c2c provides today, and not a replacement for.

Ten years ago Robert Hundt created that pahole-style output as a developer option
to the HP-UX compiler.  It used compiler feedback to compute every struct
accessed by the application, with exact counts for all reads and writes to
every struct member.  It even had affinity information to show how often
field members were accessed together in time.

He and I ran it on numerous large applications.  It was awesome, but it
did fall short in a few places that Jiri's c2c patches provide, such as
being able to:

- distinguish where the concurrent cacheline accesses came from (e.g, which
   cores, and which nodes).

- see where the loads got resolved from, (local cache, local memory, remote
   cache, remote memory).

- see if the hot structs were cacheline aligned or not.

- see if more than one hot struct shares a cachline.

- see how costly, via load latencies, the contention is.

- see, among all the accesses to a cachline, which thread or process is
   causing the most harm.

- insight into how many other threads/processes are contending for a
   cacheline (and who they are).

The above info has been critical to understanding how best to tackle the
contention uncovered for all those who have used the "perf c2c" prototype.

So yes, the pahole-style addition would be a plus and it would make it easier
to map it back to the struct, but make sure to preserve what the current
"perfc2c" provides that the pahole-style output will not.

Joe

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web