Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1461644 > unrolled thread

[PATCH v3 0/7] x86: Rewrite switch_to()

Started byBrian Gerst <brgerst@gmail.com>
First post2016-08-13 18:40 +0200
Last post2016-08-17 23:30 +0200
Articles 3 on this page of 23 — 7 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH v3 0/7] x86: Rewrite switch_to() Brian Gerst <brgerst@gmail.com> - 2016-08-13 18:40 +0200
    [PATCH v3 1/7] x86-32, kgdb: Don't use thread.ip in sleeping_thread_to_gdb_regs() Brian Gerst <brgerst@gmail.com> - 2016-08-13 18:40 +0200
      [tip:x86/asm] sched/x86/32, kgdb: Don't use thread.ip in  sleeping_thread_to_gdb_regs() tip-bot for Brian Gerst <tipbot@zytor.com> - 2016-08-24 15:50 +0200
    [PATCH v3 7/7] Revert "sched: Mark __schedule() stack frame as non-standard" Brian Gerst <brgerst@gmail.com> - 2016-08-13 18:40 +0200
      [tip:x86/asm] sched: Remove __schedule() non-standard frame  annotation tip-bot for Brian Gerst <tipbot@zytor.com> - 2016-08-24 15:50 +0200
    [PATCH v3 2/7] x86-64, kgdb: clear GDB_PS on 64-bit Brian Gerst <brgerst@gmail.com> - 2016-08-13 18:40 +0200
      [tip:x86/asm] sched/x86/64, kgdb: Clear GDB_PS on 64-bit tip-bot for Brian Gerst <tipbot@zytor.com> - 2016-08-24 15:50 +0200
    [PATCH v3 3/7] x86: Add struct inactive_task_frame Brian Gerst <brgerst@gmail.com> - 2016-08-13 18:40 +0200
      [tip:x86/asm] sched/x86: Add 'struct inactive_task_frame' to better  document the sleeping task stack frame tip-bot for Brian Gerst <tipbot@zytor.com> - 2016-08-24 15:20 +0200
    [PATCH v3 4/7] x86: Rewrite switch_to() code Brian Gerst <brgerst@gmail.com> - 2016-08-13 18:50 +0200
      [tip:x86/asm] sched/x86: Rewrite the switch_to() code tip-bot for Brian Gerst <tipbot@zytor.com> - 2016-08-24 15:20 +0200
    [PATCH v3 6/7] x86: Fix thread_saved_pc() Brian Gerst <brgerst@gmail.com> - 2016-08-13 18:50 +0200
      [tip:x86/asm] sched/x86: Fix thread_saved_pc() tip-bot for Brian Gerst <tipbot@zytor.com> - 2016-08-24 15:20 +0200
    Re: [PATCH v3 0/7] x86: Rewrite switch_to() Linus Torvalds <torvalds@linux-foundation.org> - 2016-08-13 19:20 +0200
      Re: [PATCH v3 0/7] x86: Rewrite switch_to() Ingo Molnar <mingo@kernel.org> - 2016-08-14 10:30 +0200
        Re: [PATCH v3 0/7] x86: Rewrite switch_to() Andy Lutomirski <luto@amacapital.net> - 2016-08-14 13:20 +0200
          Re: [PATCH v3 0/7] x86: Rewrite switch_to() Herbert Xu <herbert@gondor.apana.org.au> - 2016-08-17 07:20 +0200
        Re: [PATCH v3 0/7] x86: Rewrite switch_to() Brian Gerst <brgerst@gmail.com> - 2016-08-14 16:20 +0200
          Re: [PATCH v3 0/7] x86: Rewrite switch_to() Ingo Molnar <mingo@kernel.org> - 2016-08-15 07:20 +0200
            Re: [PATCH v3 0/7] x86: Rewrite switch_to() Brian Gerst <brgerst@gmail.com> - 2016-08-15 13:50 +0200
            Re: [PATCH v3 0/7] x86: Rewrite switch_to() Andy Lutomirski <luto@amacapital.net> - 2016-08-17 23:30 +0200
      Re: [PATCH v3 0/7] x86: Rewrite switch_to() Brian Gerst <brgerst@gmail.com> - 2016-08-14 12:30 +0200
    Re: [PATCH v3 0/7] x86: Rewrite switch_to() Josh Poimboeuf <jpoimboe@redhat.com> - 2016-08-17 23:30 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1464812

FromAndy Lutomirski <luto@amacapital.net>
Date2016-08-17 23:30 +0200
Message-ID<s7gzo-6R5-25@gated-at.bofh.it>
In reply to#1462553
On Aug 15, 2016 8:10 AM, "Ingo Molnar" <mingo@kernel.org> wrote:
>
>
> * Brian Gerst <brgerst@gmail.com> wrote:
>
> > > Something like this:
> > >
> > >   taskset 1 perf stat -a -e '{instructions,cycles}' --repeat 10 perf bench sched pipe
> > >
> > > ... will give a very good idea about the general impact of these changes on
> > > context switch overhead.
> >
> > Before:
> >  Performance counter stats for 'system wide' (10 runs):
> >
> >     12,010,932,128      instructions              #    1.03  insn per
> > cycle                                              ( +-  0.31% )
> >     11,691,797,513      cycles
> >                ( +-  0.76% )
> >
> >        3.487329979 seconds time elapsed
> >           ( +-  0.78% )
> >
> > After:
> >  Performance counter stats for 'system wide' (10 runs):
> >
> >     12,097,706,506      instructions              #    1.04  insn per
> > cycle                                              ( +-  0.14% )
> >     11,612,167,742      cycles
> >                ( +-  0.81% )
> >
> >        3.451278789 seconds time elapsed
> >           ( +-  0.82% )
> >
> > The numbers with or without this patch series are roughly the same.
> > There is noticeable variation in the numbers each time I run it, so
> > I'm not sure how good of a benchmark this is.
>
> Weird, I get an order of magnitude lower noise:
>
>  triton:~/tip> taskset 1 perf stat -a -e '{instructions,cycles}' --repeat 10 perf bench sched pipe >/dev/null
>
>  Performance counter stats for 'system wide' (10 runs):
>
>     11,503,026,062      instructions              #    1.23  insn per cycle                                              ( +-  2.64% )
>      9,377,410,613      cycles                                                        ( +-  2.05% )
>
>        1.669425407 seconds time elapsed                                          ( +-  0.12% )
>
> But note that I also had '--sync' for perf stat and did a >/dev/null at the end to
> make sure no terminal output and subsequent Xorg activities interfere. Also, full
> screen terminal.
>
> Maybe try 'taskset 4' as well to put the workload on another CPU, if the first CPU
> is busier than the others?
>
> (Any Hyperthreading on your test system?)
>

I've never investigated for real, but I suspect that cgroups are a big
part of it.  If you do a regular perf recording, I think you'll find
that nearly all of the time is in the scheduler.

[toc] | [prev] | [next] | [standalone]


#1461772

FromBrian Gerst <brgerst@gmail.com>
Date2016-08-14 12:30 +0200
Message-ID<s5YXU-5oj-29@gated-at.bofh.it>
In reply to#1461660
On Sat, Aug 13, 2016 at 1:16 PM, Linus Torvalds
<torvalds@linux-foundation.org> wrote:
> On Sat, Aug 13, 2016 at 9:38 AM, Brian Gerst <brgerst@gmail.com> wrote:
>> This patch set simplifies the switch_to() code, by moving the stack switch
>> code out of line into an asm stub before calling __switch_to().  This ends
>> up being more readable, and using the C calling convention instead of
>> clobbering all registers improves code generation.  It also allows newly
>> forked processes to construct a special stack frame to seamlessly flow
>> to ret_from_fork, instead of using a test and branch, or an unbalanced
>> call/ret.
>
> Do you have performance numbers? Is it noticeable/measurable?

How do I measure it?  The perf documentation isn't easy to understand.

It shouldn't be a significant change.  On a 64-bit defconfig build,
__schedule() shrinks by 103 bytes.  It's hard to analyse what exactly
changes, but it's likely that GCC can allocate registers better
without all the clobbers of the old inline asm version interfering.
The new stub adds just 39 bytes.

--
Brian Gerst

[toc] | [prev] | [next] | [standalone]


#1464811

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2016-08-17 23:30 +0200
Message-ID<s7gzo-6R5-23@gated-at.bofh.it>
In reply to#1461644
On Sat, Aug 13, 2016 at 12:38:15PM -0400, Brian Gerst wrote:
> This patch set simplifies the switch_to() code, by moving the stack switch
> code out of line into an asm stub before calling __switch_to().  This ends
> up being more readable, and using the C calling convention instead of
> clobbering all registers improves code generation.  It also allows newly
> forked processes to construct a special stack frame to seamlessly flow
> to ret_from_fork, instead of using a test and branch, or an unbalanced
> call/ret.
> 
> Changes from v2:
> - Updated comments around kernel threads being uncommon for fork, etc.
> - Removed STACK_FRAME_NON_STANDARD annotation from __schedule() per Josh Poimboeuf
> - A few minor cleanups added

There are a few minor conflicts with my x86 stack dump patch set, but
for the most part they should be orthogonal.

For the series:

  Reviewed-by: Josh Poimboeuf <jpoimboe@redhat.com>

-- 
Josh

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web