Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1216546

Re: [PATCH 0/7] x86 vdso32 cleanups

From Andy Lutomirski <luto@amacapital.net>
Newsgroups linux.kernel
Subject Re: [PATCH 0/7] x86 vdso32 cleanups
Date 2015-09-01 03:40 +0200
Message-ID <q3IIi-eh-27@gated-at.bofh.it> (permalink)
References <q2QeS-5LL-5@gated-at.bofh.it> <q2R1g-6Wa-9@gated-at.bofh.it> <q3ib7-3Zh-9@gated-at.bofh.it> <q3nua-39B-3@gated-at.bofh.it> <q3IIi-eh-29@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


On Mon, Aug 31, 2015 at 6:19 PM, Andy Lutomirski <luto@amacapital.net> wrote:
>
> On Sun, Aug 30, 2015 at 7:52 PM, Andy Lutomirski <luto@amacapital.net> wrote:
>>
>> On Sun, Aug 30, 2015 at 2:18 PM, Brian Gerst <brgerst@gmail.com> wrote:
>> > On Sat, Aug 29, 2015 at 12:10 PM, Andy Lutomirski <luto@amacapital.net> wrote:
>> >> On Sat, Aug 29, 2015 at 8:20 AM, Brian Gerst <brgerst@gmail.com> wrote:
>> >>> This patch set contains several cleanups to the 32-bit VDSO.  The
>> >>> main change is to only build one VDSO image, and select the syscall
>> >>> entry point at runtime.
>> >>
>> >> Oh no, we have dueling patches!
>> >>
>> >> I have a 2/3 finished series that cleans up the AT_SYSINFO mess
>> >> differently, as I outlined earlier.  I've only done the compat and
>> >> common bits (no 32-bit native support quite yet), and it enters
>> >> successfully on Intel using SYSENTER and on (fake) AMD using SYSCALL.
>> >> The SYSRET bit isn't there yet.
>> >>
>> >> Other than some ifdeffery, the final system_call.S looks like this:
>> >>
>> >> https://git.kernel.org/cgit/linux/kernel/git/luto/linux.git/tree/arch/x86/entry/vdso/vdso32/system_call.S?h=x86/entry_compat
>> >>
>> >> The meat is (sorry for whitespace damage):
>> >>
>> >> .text
>> >> .globl __kernel_vsyscall
>> >> .type __kernel_vsyscall,@function
>> >> ALIGN
>> >> __kernel_vsyscall:
>> >> CFI_STARTPROC
>> >> /*
>> >> * Reshuffle regs so that all of any of the entry instructions
>> >> * will preserve enough state.
>> >> */
>> >> pushl %edx
>> >> CFI_ADJUST_CFA_OFFSET 4
>> >> CFI_REL_OFFSET edx, 0
>> >> pushl %ecx
>> >> CFI_ADJUST_CFA_OFFSET 4
>> >> CFI_REL_OFFSET ecx, 0
>> >> movl %esp, %ecx
>> >>
>> >> #ifdef CONFIG_X86_64
>> >> /* If SYSENTER is available, use it. */
>> >> ALTERNATIVE_2 "", "sysenter", X86_FEATURE_SYSENTER32, \
>> >>                  "syscall",  X86_FEATURE_SYSCALL32
>> >> #endif
>> >>
>> >> /* Enter using int $0x80 */
>> >> movl (%esp), %ecx
>> >> int $0x80
>> >> GLOBAL(int80_landing_pad)
>> >>
>> >> /* Restore ECX and EDX in case they were clobbered. */
>> >> popl %ecx
>> >> CFI_RESTORE ecx
>> >> CFI_ADJUST_CFA_OFFSET -4
>> >> popl %edx
>> >> CFI_RESTORE edx
>> >> CFI_ADJUST_CFA_OFFSET -4
>> >> ret
>> >> CFI_ENDPROC
>> >>
>> >> .size __kernel_vsyscall,.-__kernel_vsyscall
>> >> .previous
>> >>
>> >> And that's it.
>> >>
>> >> What do you think?  This comes with massively cleaned up kernel-side
>> >> asm as well as a test case that actually validates the CFI directives.
>> >>
>> >> Certainly, a bunch of your patches make sense regardless, and I'll
>> >> review them and add them to my queue soon.
>> >>
>> >> --Andy
>> >
>> > How does the performance compare to the original?  Looking at the
>> > disassembly, there are two added function calls, and it reloads the
>> > args from the stack instead of just shuffling registers.
>>
>> The replacement is dramatically faster, which means I probably
>> benchmarked it wrong.  I'll try again in a day or two.
>
>
> It's enough slower to be problematic.  I need to figure out how to trace it properly.  (Hmm?  Maybe it's time to learn how to get perf on the host to trace a KVM guest.)
>
> Everything is and was hilariously slow with context tracking on.  That needs to get fixed, and hopefully once this entry stuff is done someone will do the other end of it.
>

I got random errors from perf kvm, but I think I found at least part
of the issue.  The two irqs_disabled() calls in common.c are kind of
expensive.  I should disable them on non-lockdep kernels.

The context tracking hooks are also too expensive, even when disabled.
I should do something to optimize those.  Hello, static keys?  This
doesn't affect syscalls, though.

With context tracking off and the irqs_disabled checks commented out,
we're probably doing well enough.  We can always tweak the C code and
aggressively force inlining if we want a few cycles back.

--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

Back to linux.kernel | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

[PATCH 0/7] x86 vdso32 cleanups Brian Gerst <brgerst@gmail.com> - 2015-08-29 17:30 +0200
  [PATCH 3/7] x86/vdso32: Remove unused vdso-fakesections.c Brian Gerst <brgerst@gmail.com> - 2015-08-29 17:30 +0200
    Re: [PATCH 3/7] x86/vdso32: Remove unused vdso-fakesections.c Andy Lutomirski <luto@amacapital.net> - 2015-08-30 18:50 +0200
  [PATCH 4/7] x86/vdso32: Build single vdso32 image Brian Gerst <brgerst@gmail.com> - 2015-08-29 17:30 +0200
  [PATCH 6/7] x86/vdso32/xen: Move VDSO_NOTE_NONEGSEG_BIT define Brian Gerst <brgerst@gmail.com> - 2015-08-29 17:30 +0200
  [PATCH 5/7] x86/vdso: Merge 32-bit and 64-bit source files Brian Gerst <brgerst@gmail.com> - 2015-08-29 17:30 +0200
  Re: [PATCH 0/7] x86 vdso32 cleanups Andy Lutomirski <luto@amacapital.net> - 2015-08-29 18:20 +0200
    Re: [PATCH 0/7] x86 vdso32 cleanups Brian Gerst <brgerst@gmail.com> - 2015-08-30 23:20 +0200
      Re: [PATCH 0/7] x86 vdso32 cleanups Andy Lutomirski <luto@amacapital.net> - 2015-08-31 05:00 +0200
        Re: [PATCH 0/7] x86 vdso32 cleanups Andy Lutomirski <luto@amacapital.net> - 2015-09-01 03:40 +0200
          Re: [PATCH 0/7] x86 vdso32 cleanups Andy Lutomirski <luto@amacapital.net> - 2015-09-02 00:00 +0200

csiph-web