Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1685158 > unrolled thread

[PATCH v3 00/10] x86: ORC unwinder (previously undwarf)

Started byJosh Poimboeuf <jpoimboe@redhat.com>
First post2017-07-11 17:40 +0200
Last post2017-07-13 14:10 +0200
Articles 11 on this page of 31 — 8 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-07-11 17:40 +0200
    Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Ingo Molnar <mingo@kernel.org> - 2017-07-12 10:30 +0200
      Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-07-12 16:50 +0200
        Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Ingo Molnar <mingo@kernel.org> - 2017-07-12 21:30 +0200
          Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-07-14 19:20 +0200
    Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Andres Freund <andres@anarazel.de> - 2017-07-12 23:50 +0200
      Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-07-13 00:40 +0200
        Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Andres Freund <andres@anarazel.de> - 2017-07-13 00:40 +0200
          Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-07-13 00:50 +0200
            Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Andres Freund <andres@anarazel.de> - 2017-07-13 01:00 +0200
        Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Peter Zijlstra <peterz@infradead.org> - 2017-07-13 09:20 +0200
          Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Peter Zijlstra <peterz@infradead.org> - 2017-07-13 11:00 +0200
            Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Ingo Molnar <mingo@kernel.org> - 2017-07-13 11:20 +0200
              Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-07-13 14:20 +0200
                Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-07-13 14:30 +0200
                  Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-07-13 14:40 +0200
                    Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Ingo Molnar <mingo@kernel.org> - 2017-07-14 10:40 +0200
                Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Ingo Molnar <mingo@kernel.org> - 2017-07-14 10:30 +0200
          Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Peter Zijlstra <peterz@infradead.org> - 2017-07-13 11:00 +0200
    Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Andi Kleen <andi@firstfloor.org> - 2017-07-13 00:40 +0200
      Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-07-13 00:50 +0200
        Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Andi Kleen <andi@firstfloor.org> - 2017-07-13 06:30 +0200
          Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-07-13 15:20 +0200
        Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Ingo Molnar <mingo@kernel.org> - 2017-07-13 11:30 +0200
      Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Andy Lutomirski <luto@kernel.org> - 2017-07-13 01:30 +0200
      Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Mike Galbraith <efault@gmx.de> - 2017-07-13 05:10 +0200
        Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Andi Kleen <andi@firstfloor.org> - 2017-07-13 06:20 +0200
          Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Mike Galbraith <efault@gmx.de> - 2017-07-13 06:40 +0200
            Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Andi Kleen <andi@firstfloor.org> - 2017-07-13 06:50 +0200
              Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Mike Galbraith <efault@gmx.de> - 2017-07-13 07:30 +0200
              Re: [PATCH v3 00/10] x86: ORC unwinder (previously undwarf) Jiri Kosina <jikos@kernel.org> - 2017-07-13 14:10 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1686094

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2017-07-13 00:50 +0200
Message-ID<u2yCd-5NI-9@gated-at.bofh.it>
In reply to#1686079
On Wed, Jul 12, 2017 at 03:30:31PM -0700, Andi Kleen wrote:
> Josh Poimboeuf <jpoimboe@redhat.com> writes:
> >
> > The ORC data format does have a few downsides compared to DWARF.  The
> > ORC unwind tables take up ~1MB more memory than DWARF eh_frame tables.
> >
> Can we have an option to just use dwarf instead? For people
> who don't want to waste a MB+ to solve a problem that doesn't
> exist (as proven by many years of opensuse kernel experience)
> 
> As far as I can tell this whole thing has only downsides compared
> to the dwarf unwinder that was earlier proposed. I don't see
> a single advantage.

Improved speed, reliability, maintainability.  Are those not advantages?

-- 
Josh

[toc] | [prev] | [next] | [standalone]


#1686237

FromAndi Kleen <andi@firstfloor.org>
Date2017-07-13 06:30 +0200
Message-ID<u2DVf-PH-9@gated-at.bofh.it>
In reply to#1686094
On Wed, Jul 12, 2017 at 05:47:59PM -0500, Josh Poimboeuf wrote:
> On Wed, Jul 12, 2017 at 03:30:31PM -0700, Andi Kleen wrote:
> > Josh Poimboeuf <jpoimboe@redhat.com> writes:
> > >
> > > The ORC data format does have a few downsides compared to DWARF.  The
> > > ORC unwind tables take up ~1MB more memory than DWARF eh_frame tables.
> > >
> > Can we have an option to just use dwarf instead? For people
> > who don't want to waste a MB+ to solve a problem that doesn't
> > exist (as proven by many years of opensuse kernel experience)
> > 
> > As far as I can tell this whole thing has only downsides compared
> > to the dwarf unwinder that was earlier proposed. I don't see
> > a single advantage.
> 
> Improved speed, reliability, maintainability.  Are those not advantages?

Ok. We'll see how it works out.

The memory overhead is quite bad though. You're basically undoing many
years of efforts to shrink kernel text. I hope this can be still
done better.

-Andi

[toc] | [prev] | [next] | [standalone]


#1686527

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2017-07-13 15:20 +0200
Message-ID<u2Mca-664-17@gated-at.bofh.it>
In reply to#1686237
On Wed, Jul 12, 2017 at 09:29:17PM -0700, Andi Kleen wrote:
> On Wed, Jul 12, 2017 at 05:47:59PM -0500, Josh Poimboeuf wrote:
> > On Wed, Jul 12, 2017 at 03:30:31PM -0700, Andi Kleen wrote:
> > > Josh Poimboeuf <jpoimboe@redhat.com> writes:
> > > >
> > > > The ORC data format does have a few downsides compared to DWARF.  The
> > > > ORC unwind tables take up ~1MB more memory than DWARF eh_frame tables.
> > > >
> > > Can we have an option to just use dwarf instead? For people
> > > who don't want to waste a MB+ to solve a problem that doesn't
> > > exist (as proven by many years of opensuse kernel experience)
> > > 
> > > As far as I can tell this whole thing has only downsides compared
> > > to the dwarf unwinder that was earlier proposed. I don't see
> > > a single advantage.
> > 
> > Improved speed, reliability, maintainability.  Are those not advantages?
> 
> Ok. We'll see how it works out.
> 
> The memory overhead is quite bad though. You're basically undoing many
> years of efforts to shrink kernel text. I hope this can be still
> done better.

If we're talking *text*, this further shrinks text size by 3% because
frame pointers can be disabled.

As far as the data size goes, is anyone *truly* impacted by that extra
1MB or so?  If you're enabling a DWARF/ORC unwinder, you're already
signing up for a few extra megs anyway.

I do have a vague idea about how to reduce the data size, if/when the
size becomes a problem.  Basically there's a *lot* of duplication in the
ORC data:

  $ tools/objtool/objtool orc dump vmlinux | wc -l
  311095

  $ tools/objtool/objtool orc dump vmlinux |cut -d' ' -f2- |sort |uniq |wc -l
  345

So that's over 300,000 6-byte entries, only 345 of which are unique.
There should be a way to compress that.  However, it will probably
require sacrificing some combination of speed and simplicity.

-- 
Josh

[toc] | [prev] | [next] | [standalone]


#1686402

FromIngo Molnar <mingo@kernel.org>
Date2017-07-13 11:30 +0200
Message-ID<u2IBA-3Pg-11@gated-at.bofh.it>
In reply to#1686094
* Josh Poimboeuf <jpoimboe@redhat.com> wrote:

> On Wed, Jul 12, 2017 at 03:30:31PM -0700, Andi Kleen wrote:
> > Josh Poimboeuf <jpoimboe@redhat.com> writes:
> > >
> > > The ORC data format does have a few downsides compared to DWARF.  The
> > > ORC unwind tables take up ~1MB more memory than DWARF eh_frame tables.
> > >
> > Can we have an option to just use dwarf instead? For people
> > who don't want to waste a MB+ to solve a problem that doesn't
> > exist (as proven by many years of opensuse kernel experience)
> > 
> > As far as I can tell this whole thing has only downsides compared
> > to the dwarf unwinder that was earlier proposed. I don't see
> > a single advantage.
> 
> Improved speed, reliability, maintainability.  Are those not advantages?

Exactly, and all these advantages of the ORC debuginfo over DWARF debuginfo are 
enabled by an unwinding optimized data format that the kernel project generates, 
controls and is able to trust inherently.

DWARF generated by external tooling can just never reach that level of trust, 
without insane amounts of formal verification.

Even if ORC was _slower_ its reliability would be reason enough to merge. The fact 
that it's 20-40 times faster than the DWARF unwinder is really just icing on the 
cake.

BTW., as a side note, (and I hope my optimism isn't premature), I believe the ORC 
unwinder is a prime example of where Linus's stubborness resisting poor concepts 
paid off in the long run: had we merged the DWARF unwinder years ago we'd never 
have gained the ORC unwinder. We quite literally had to wait over a decade, but 
good things happened in the end.

Thanks,

	Ingo

[toc] | [prev] | [next] | [standalone]


#1686122

FromAndy Lutomirski <luto@kernel.org>
Date2017-07-13 01:30 +0200
Message-ID<u2zeV-6fP-9@gated-at.bofh.it>
In reply to#1686079
On Wed, Jul 12, 2017 at 3:30 PM, Andi Kleen <andi@firstfloor.org> wrote:
> Josh Poimboeuf <jpoimboe@redhat.com> writes:
>>
>> The ORC data format does have a few downsides compared to DWARF.  The
>> ORC unwind tables take up ~1MB more memory than DWARF eh_frame tables.
>>
> Can we have an option to just use dwarf instead? For people
> who don't want to waste a MB+ to solve a problem that doesn't
> exist (as proven by many years of opensuse kernel experience)
>
> As far as I can tell this whole thing has only downsides compared
> to the dwarf unwinder that was earlier proposed. I don't see
> a single advantage.
>

If someone wanted to write an in-kernel DWARF parser that hooked into
the same machinery that Josh is using and comes with a complete formal
verification package, I might not object, with the caveat that it's
likely to be *much* slower than ORC.  By complete formal verification,
I mean that a user tool running the exact same code, compiled with
strong sanitization, should decode the tables for every single kernel
IP and confirm that (a) the output is sane, (b) the output matches
what objtool says it should do and (c) doesn't crash.

I'm not sure I see the point, though.  I also think that Linus would
object, since I asked him quite recently and he said he'd object.

[toc] | [prev] | [next] | [standalone]


#1686204

FromMike Galbraith <efault@gmx.de>
Date2017-07-13 05:10 +0200
Message-ID<u2CFP-5N-1@gated-at.bofh.it>
In reply to#1686079
On Wed, 2017-07-12 at 15:30 -0700, Andi Kleen wrote:
> Josh Poimboeuf <jpoimboe@redhat.com> writes:
> >
> > The ORC data format does have a few downsides compared to DWARF.  The
> > ORC unwind tables take up ~1MB more memory than DWARF eh_frame tables.
> >
> Can we have an option to just use dwarf instead? For people
> who don't want to waste a MB+ to solve a problem that doesn't
> exist (as proven by many years of opensuse kernel experience)

Sure the dwarf unwinder works well for crashes, but at the price of
demolishing ftrace/perf utility.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1686233

FromAndi Kleen <andi@firstfloor.org>
Date2017-07-13 06:20 +0200
Message-ID<u2DLz-JB-1@gated-at.bofh.it>
In reply to#1686204
On Thu, Jul 13, 2017 at 05:03:00AM +0200, Mike Galbraith wrote:
> On Wed, 2017-07-12 at 15:30 -0700, Andi Kleen wrote:
> > Josh Poimboeuf <jpoimboe@redhat.com> writes:
> > >
> > > The ORC data format does have a few downsides compared to DWARF.  The
> > > ORC unwind tables take up ~1MB more memory than DWARF eh_frame tables.
> > >
> > Can we have an option to just use dwarf instead? For people
> > who don't want to waste a MB+ to solve a problem that doesn't
> > exist (as proven by many years of opensuse kernel experience)
> 
> Sure the dwarf unwinder works well for crashes, but at the price of
> demolishing ftrace/perf utility.

You mean the unwind performance?

That's a valid concern, but neither ORC nor dwarf are likely
to address it. However most usages of ftrace/perf shouldn't be that
depending on unwind performance -- just lower the frequency of your
events. 

The only possible win is if the win from not using FP code is
significant enough. On the x86 side the only modern CPUs that should really
care about this are Atoms.

-Andi

[toc] | [prev] | [next] | [standalone]


#1686241

FromMike Galbraith <efault@gmx.de>
Date2017-07-13 06:40 +0200
Message-ID<u2E4V-Uj-13@gated-at.bofh.it>
In reply to#1686233
On Wed, 2017-07-12 at 21:15 -0700, Andi Kleen wrote:
> On Thu, Jul 13, 2017 at 05:03:00AM +0200, Mike Galbraith wrote:
> > On Wed, 2017-07-12 at 15:30 -0700, Andi Kleen wrote:
> > > Josh Poimboeuf <jpoimboe@redhat.com> writes:
> > > >
> > > > The ORC data format does have a few downsides compared to DWARF.  The
> > > > ORC unwind tables take up ~1MB more memory than DWARF eh_frame tables.
> > > >
> > > Can we have an option to just use dwarf instead? For people
> > > who don't want to waste a MB+ to solve a problem that doesn't
> > > exist (as proven by many years of opensuse kernel experience)
> > 
> > Sure the dwarf unwinder works well for crashes, but at the price of
> > demolishing ftrace/perf utility.
> 
> You mean the unwind performance?

Yeah, it hurts.. massively, has even been known to kill big boxen.

> That's a valid concern, but neither ORC nor dwarf are likely
> to address it. However most usages of ftrace/perf shouldn't be that
> depending on unwind performance -- just lower the frequency of your
> events. 
> 
> The only possible win is if the win from not using FP code is
> significant enough. On the x86 side the only modern CPUs that should really
> care about this are Atoms.

Nope, they all care.  Measure performance delta of fast/light stuff.

Maybe I'm expecting too much good stuff to follow, but don't spoil it
for me, I think I'm looking at a real winner :)

	-Mike

[toc] | [prev] | [next] | [standalone]


#1686242

FromAndi Kleen <andi@firstfloor.org>
Date2017-07-13 06:50 +0200
Message-ID<u2EeB-XA-1@gated-at.bofh.it>
In reply to#1686241
On Thu, Jul 13, 2017 at 06:28:43AM +0200, Mike Galbraith wrote:
> On Wed, 2017-07-12 at 21:15 -0700, Andi Kleen wrote:
> > On Thu, Jul 13, 2017 at 05:03:00AM +0200, Mike Galbraith wrote:
> > > On Wed, 2017-07-12 at 15:30 -0700, Andi Kleen wrote:
> > > > Josh Poimboeuf <jpoimboe@redhat.com> writes:
> > > > >
> > > > > The ORC data format does have a few downsides compared to DWARF.  The
> > > > > ORC unwind tables take up ~1MB more memory than DWARF eh_frame tables.
> > > > >
> > > > Can we have an option to just use dwarf instead? For people
> > > > who don't want to waste a MB+ to solve a problem that doesn't
> > > > exist (as proven by many years of opensuse kernel experience)
> > > 
> > > Sure the dwarf unwinder works well for crashes, but at the price of
> > > demolishing ftrace/perf utility.
> > 
> > You mean the unwind performance?
> 
> Yeah, it hurts.. massively, has even been known to kill big boxen.

Why was that? 

> 
> > That's a valid concern, but neither ORC nor dwarf are likely
> > to address it. However most usages of ftrace/perf shouldn't be that
> > depending on unwind performance -- just lower the frequency of your
> > events. 
> > 
> > The only possible win is if the win from not using FP code is
> > significant enough. On the x86 side the only modern CPUs that should really
> > care about this are Atoms.
> 
> Nope, they all care.  Measure performance delta of fast/light stuff.

Well if your test cares that much about function overhead you may want to try
LTO. It can get rid of a lot of functions by doing cross file
inlining.

https://git.kernel.org/pub/scm/linux/kernel/git/ak/linux-misc.git/log/?h=lto-411-2

> Maybe I'm expecting too much good stuff to follow, but don't spoil it
> for me, I think I'm looking at a real winner :)

It's somewhat surprising. It would be good to under stand why that
happens. Is it icache misses, data cache misses for the stack, or
simply more instructions executed, or worse tail calls?

-Andi

[toc] | [prev] | [next] | [standalone]


#1686250

FromMike Galbraith <efault@gmx.de>
Date2017-07-13 07:30 +0200
Message-ID<u2ERj-1qH-1@gated-at.bofh.it>
In reply to#1686242
On Wed, 2017-07-12 at 21:40 -0700, Andi Kleen wrote:
> On Thu, Jul 13, 2017 at 06:28:43AM +0200, Mike Galbraith wrote:
> > On Wed, 2017-07-12 at 21:15 -0700, Andi Kleen wrote:
> > > On Thu, Jul 13, 2017 at 05:03:00AM +0200, Mike Galbraith wrote:
> > > > On Wed, 2017-07-12 at 15:30 -0700, Andi Kleen wrote:
> > > > > Josh Poimboeuf <jpoimboe@redhat.com> writes:
> > > > > >
> > > > > > The ORC data format does have a few downsides compared to DWARF.  The
> > > > > > ORC unwind tables take up ~1MB more memory than DWARF eh_frame tables.
> > > > > >
> > > > > Can we have an option to just use dwarf instead? For people
> > > > > who don't want to waste a MB+ to solve a problem that doesn't
> > > > > exist (as proven by many years of opensuse kernel experience)
> > > > 
> > > > Sure the dwarf unwinder works well for crashes, but at the price of
> > > > demolishing ftrace/perf utility.
> > > 
> > > You mean the unwind performance?
> > 
> > Yeah, it hurts.. massively, has even been known to kill big boxen.
> 
> Why was that?

Presuming you mean the big box bit, danged if I know, I haven't
personally met that, only the massive overhead.

> > > That's a valid concern, but neither ORC nor dwarf are likely
> > > to address it. However most usages of ftrace/perf shouldn't be that
> > > depending on unwind performance -- just lower the frequency of your
> > > events. 
> > > 
> > > The only possible win is if the win from not using FP code is
> > > significant enough. On the x86 side the only modern CPUs that should really
> > > care about this are Atoms.
> > 
> > Nope, they all care.  Measure performance delta of fast/light stuff.
> 
> Well if your test cares that much about function overhead you may want to try
> LTO. It can get rid of a lot of functions by doing cross file
> inlining.
> 
> https://git.kernel.org/pub/scm/linux/kernel/git/ak/linux-misc.git/log/?h=lto-411-2
> 
> > Maybe I'm expecting too much good stuff to follow, but don't spoil it
> > for me, I think I'm looking at a real winner :)
> 
> It's somewhat surprising. It would be good to under stand why that
> happens. Is it icache misses, data cache misses for the stack, or
> simply more instructions executed, or worse tail calls?

No idea.  It was speculated that it was register loss, but I played
with that, saw nearly zero delta until I stole too many.

	-Mike

[toc] | [prev] | [next] | [standalone]


#1686481

FromJiri Kosina <jikos@kernel.org>
Date2017-07-13 14:10 +0200
Message-ID<u2L6q-5tf-7@gated-at.bofh.it>
In reply to#1686242
On Wed, 12 Jul 2017, Andi Kleen wrote:

> It's somewhat surprising. It would be good to under stand why that 
> happens. Is it icache misses, data cache misses for the stack, or simply 
> more instructions executed, or worse tail calls?

http://lkml.kernel.org/r/20170602104048.jkkzssljsompjdwy@suse.de

-- 
Jiri Kosina
SUSE Labs

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web