Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1390507 > unrolled thread
| Started by | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| First post | 2016-04-28 22:50 +0200 |
| Last post | 2016-05-10 13:50 +0200 |
| Articles | 20 on this page of 72 — 11 participants |
Back to article view | Back to linux.kernel
[RFC PATCH v2 00/18] livepatch: hybrid consistency model Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 22:50 +0200
[RFC PATCH v2 14/18] livepatch: remove unnecessary object loaded check Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 22:50 +0200
[RFC PATCH v2 10/18] livepatch/powerpc: add TIF_PATCH_PENDING thread flag Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 22:50 +0200
Re: [RFC PATCH v2 10/18] livepatch/powerpc: add TIF_PATCH_PENDING thread flag Petr Mladek <pmladek@suse.com> - 2016-05-03 11:10 +0200
Re: [RFC PATCH v2 10/18] livepatch/powerpc: add TIF_PATCH_PENDING thread flag Miroslav Benes <mbenes@suse.cz> - 2016-05-03 14:10 +0200
[RFC PATCH v2 16/18] livepatch: store function sizes Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 22:50 +0200
[RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 22:50 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Andy Lutomirski <luto@amacapital.net> - 2016-04-29 20:10 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Andy Lutomirski <luto@amacapital.net> - 2016-04-29 22:20 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-29 22:30 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Andy Lutomirski <luto@amacapital.net> - 2016-04-29 22:40 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-29 23:30 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Andy Lutomirski <luto@amacapital.net> - 2016-04-29 23:40 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Jiri Kosina <jikos@kernel.org> - 2016-04-30 00:20 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-30 01:00 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Andy Lutomirski <luto@amacapital.net> - 2016-04-30 02:20 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-30 00:50 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Andy Lutomirski <luto@amacapital.net> - 2016-04-30 02:10 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-02 16:00 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Andy Lutomirski <luto@amacapital.net> - 2016-05-02 18:00 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-02 19:40 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Andy Lutomirski <luto@amacapital.net> - 2016-05-02 20:20 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Ingo Molnar <mingo@kernel.org> - 2016-05-02 20:40 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-02 21:50 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Jiri Kosina <jikos@kernel.org> - 2016-05-02 22:00 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Jiri Kosina <jikos@kernel.org> - 2016-05-02 22:10 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Andy Lutomirski <luto@amacapital.net> - 2016-05-03 02:50 +0200
RE: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking David Laight <David.Laight@ACULAB.COM> - 2016-05-04 17:20 +0200
Re: [RFC PATCH v2 05/18] sched: add task flag for preempt IRQ tracking Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-29 22:20 +0200
[RFC PATCH v2 09/18] livepatch/x86: add TIF_PATCH_PENDING thread flag Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 22:50 +0200
Re: [RFC PATCH v2 09/18] livepatch/x86: add TIF_PATCH_PENDING thread flag Andy Lutomirski <luto@amacapital.net> - 2016-04-29 20:10 +0200
Re: [RFC PATCH v2 09/18] livepatch/x86: add TIF_PATCH_PENDING thread flag Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-29 22:20 +0200
[RFC PATCH v2 02/18] x86/asm/head: use a common function for starting CPUs Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 22:50 +0200
[RFC PATCH v2 13/18] livepatch: separate enabled and patched states Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 22:50 +0200
Re: [RFC PATCH v2 13/18] livepatch: separate enabled and patched states Petr Mladek <pmladek@suse.com> - 2016-05-03 11:40 +0200
Re: [RFC PATCH v2 13/18] livepatch: separate enabled and patched states Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-03 15:50 +0200
[RFC PATCH v2 06/18] x86: dump_trace() error handling Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 22:50 +0200
Re: [RFC PATCH v2 06/18] x86: dump_trace() error handling Minfei Huang <mnghuan@gmail.com> - 2016-04-29 15:50 +0200
Re: [RFC PATCH v2 06/18] x86: dump_trace() error handling Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-29 16:10 +0200
[RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 23:00 +0200
Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks Brian Gerst <brgerst@gmail.com> - 2016-04-29 20:50 +0200
Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-29 22:30 +0200
Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks Andy Lutomirski <luto@kernel.org> - 2016-04-29 21:40 +0200
Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-29 23:00 +0200
Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks Andy Lutomirski <luto@amacapital.net> - 2016-04-29 23:40 +0200
Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-30 01:30 +0200
Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks Andy Lutomirski <luto@amacapital.net> - 2016-04-30 02:20 +0200
[RFC PATCH v2 04/18] x86: move _stext marker before head code Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 23:00 +0200
[RFC PATCH v2 01/18] x86/asm/head: clean up initial stack variable Josh Poimboeuf <jpoimboe@redhat.com> - 2016-04-28 23:00 +0200
Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Petr Mladek <pmladek@suse.com> - 2016-05-04 10:50 +0200
Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-04 18:00 +0200
Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Miroslav Benes <mbenes@suse.cz> - 2016-05-05 11:50 +0200
Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Petr Mladek <pmladek@suse.com> - 2016-05-05 15:10 +0200
barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Petr Mladek <pmladek@suse.com> - 2016-05-04 14:40 +0200
Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Peter Zijlstra <peterz@infradead.org> - 2016-05-04 16:00 +0200
Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-04 19:00 +0200
Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Petr Mladek <pmladek@suse.com> - 2016-05-04 16:20 +0200
Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-04 19:30 +0200
Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Petr Mladek <pmladek@suse.com> - 2016-05-05 13:30 +0200
Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Miroslav Benes <mbenes@suse.cz> - 2016-05-09 17:50 +0200
Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-04 19:10 +0200
Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Petr Mladek <pmladek@suse.com> - 2016-05-05 12:30 +0200
klp_task_patch: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Petr Mladek <pmladek@suse.com> - 2016-05-04 16:50 +0200
Re: klp_task_patch: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Jiri Kosina <jikos@kernel.org> - 2016-05-04 17:00 +0200
Re: klp_task_patch: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-04 20:00 +0200
Re: klp_task_patch: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Petr Mladek <pmladek@suse.com> - 2016-05-05 14:00 +0200
Re: klp_task_patch: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-06 14:40 +0200
Re: klp_task_patch: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Petr Mladek <pmladek@suse.com> - 2016-05-09 14:30 +0200
Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Petr Mladek <pmladek@suse.com> - 2016-05-06 13:40 +0200
Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Josh Poimboeuf <jpoimboe@redhat.com> - 2016-05-06 14:50 +0200
Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Miroslav Benes <mbenes@suse.cz> - 2016-05-09 11:50 +0200
Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model Miroslav Benes <mbenes@suse.cz> - 2016-05-10 13:50 +0200
Page 3 of 4 — ← Prev page 1 2 [3] 4 Next page →
| From | Brian Gerst <brgerst@gmail.com> |
|---|---|
| Date | 2016-04-29 20:50 +0200 |
| Subject | Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks |
| Message-ID | <rtlEe-2fj-5@gated-at.bofh.it> |
| In reply to | #1390522 |
On Thu, Apr 28, 2016 at 4:44 PM, Josh Poimboeuf <jpoimboe@redhat.com> wrote: > Thanks to all the recent x86 entry code refactoring, most tasks' kernel > stacks start at the same offset right above their saved pt_regs, > regardless of which syscall was used to enter the kernel. That creates > a nice convention which makes it straightforward to identify the > "bottom" of the stack, which can be useful for stack walking code which > needs to verify the stack is sane. > > However there are still a few types of tasks which don't yet follow that > convention: > > 1) CPU idle tasks, aka the "swapper" tasks > > 2) freshly forked TIF_FORK tasks which don't have a stack at all > > Make the idle tasks conform to the new stack bottom convention by > starting their stack at a sizeof(pt_regs) offset from the end of the > stack page. > > Signed-off-by: Josh Poimboeuf <jpoimboe@redhat.com> > --- > arch/x86/kernel/head_64.S | 7 ++++--- > 1 file changed, 4 insertions(+), 3 deletions(-) > > diff --git a/arch/x86/kernel/head_64.S b/arch/x86/kernel/head_64.S > index 6dbd2c0..0b12311 100644 > --- a/arch/x86/kernel/head_64.S > +++ b/arch/x86/kernel/head_64.S > @@ -296,8 +296,9 @@ ENTRY(start_cpu) > * REX.W + FF /5 JMP m16:64 Jump far, absolute indirect, > * address given in m16:64. > */ > - movq initial_code(%rip),%rax > - pushq $0 # fake return address to stop unwinder > + call 1f # put return address on stack for unwinder > +1: xorq %rbp, %rbp # clear frame pointer > + movq initial_code(%rip), %rax > pushq $__KERNEL_CS # set correct cs > pushq %rax # target address in negative space > lretq This chunk looks like it should be a separate patch. -- Brian Gerst
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-04-29 22:30 +0200 |
| Subject | Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks |
| Message-ID | <rtnd0-3HL-19@gated-at.bofh.it> |
| In reply to | #1391327 |
On Fri, Apr 29, 2016 at 02:46:10PM -0400, Brian Gerst wrote: > On Thu, Apr 28, 2016 at 4:44 PM, Josh Poimboeuf <jpoimboe@redhat.com> wrote: > > Thanks to all the recent x86 entry code refactoring, most tasks' kernel > > stacks start at the same offset right above their saved pt_regs, > > regardless of which syscall was used to enter the kernel. That creates > > a nice convention which makes it straightforward to identify the > > "bottom" of the stack, which can be useful for stack walking code which > > needs to verify the stack is sane. > > > > However there are still a few types of tasks which don't yet follow that > > convention: > > > > 1) CPU idle tasks, aka the "swapper" tasks > > > > 2) freshly forked TIF_FORK tasks which don't have a stack at all > > > > Make the idle tasks conform to the new stack bottom convention by > > starting their stack at a sizeof(pt_regs) offset from the end of the > > stack page. > > > > Signed-off-by: Josh Poimboeuf <jpoimboe@redhat.com> > > --- > > arch/x86/kernel/head_64.S | 7 ++++--- > > 1 file changed, 4 insertions(+), 3 deletions(-) > > > > diff --git a/arch/x86/kernel/head_64.S b/arch/x86/kernel/head_64.S > > index 6dbd2c0..0b12311 100644 > > --- a/arch/x86/kernel/head_64.S > > +++ b/arch/x86/kernel/head_64.S > > @@ -296,8 +296,9 @@ ENTRY(start_cpu) > > * REX.W + FF /5 JMP m16:64 Jump far, absolute indirect, > > * address given in m16:64. > > */ > > - movq initial_code(%rip),%rax > > - pushq $0 # fake return address to stop unwinder > > + call 1f # put return address on stack for unwinder > > +1: xorq %rbp, %rbp # clear frame pointer > > + movq initial_code(%rip), %rax > > pushq $__KERNEL_CS # set correct cs > > pushq %rax # target address in negative space > > lretq > > This chunk looks like it should be a separate patch. Agreed, thanks. -- Josh
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@kernel.org> |
|---|---|
| Date | 2016-04-29 21:40 +0200 |
| Subject | Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks |
| Message-ID | <rtmqC-34w-21@gated-at.bofh.it> |
| In reply to | #1390522 |
On Thu, Apr 28, 2016 at 1:44 PM, Josh Poimboeuf <jpoimboe@redhat.com> wrote: > Thanks to all the recent x86 entry code refactoring, most tasks' kernel > stacks start at the same offset right above their saved pt_regs, > regardless of which syscall was used to enter the kernel. That creates > a nice convention which makes it straightforward to identify the > "bottom" of the stack, which can be useful for stack walking code which > needs to verify the stack is sane. > > However there are still a few types of tasks which don't yet follow that > convention: > > 1) CPU idle tasks, aka the "swapper" tasks > > 2) freshly forked TIF_FORK tasks which don't have a stack at all > > Make the idle tasks conform to the new stack bottom convention by > starting their stack at a sizeof(pt_regs) offset from the end of the > stack page. > > Signed-off-by: Josh Poimboeuf <jpoimboe@redhat.com> > --- > arch/x86/kernel/head_64.S | 7 ++++--- > 1 file changed, 4 insertions(+), 3 deletions(-) > > diff --git a/arch/x86/kernel/head_64.S b/arch/x86/kernel/head_64.S > index 6dbd2c0..0b12311 100644 > --- a/arch/x86/kernel/head_64.S > +++ b/arch/x86/kernel/head_64.S > @@ -296,8 +296,9 @@ ENTRY(start_cpu) > * REX.W + FF /5 JMP m16:64 Jump far, absolute indirect, > * address given in m16:64. > */ > - movq initial_code(%rip),%rax > - pushq $0 # fake return address to stop unwinder > + call 1f # put return address on stack for unwinder > +1: xorq %rbp, %rbp # clear frame pointer > + movq initial_code(%rip), %rax > pushq $__KERNEL_CS # set correct cs > pushq %rax # target address in negative space > lretq > @@ -325,7 +326,7 @@ ENDPROC(start_cpu0) > GLOBAL(initial_gs) > .quad INIT_PER_CPU_VAR(irq_stack_union) > GLOBAL(initial_stack) > - .quad init_thread_union+THREAD_SIZE-8 > + .quad init_thread_union + THREAD_SIZE - SIZEOF_PTREGS As long as you're doing this, could you also set orig_ax to -1? I remember running into some oddities resulting from orig_ax containing garbage at some point. --Andy
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-04-29 23:00 +0200 |
| Subject | Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks |
| Message-ID | <rtnG2-3Ts-11@gated-at.bofh.it> |
| In reply to | #1391359 |
On Fri, Apr 29, 2016 at 12:39:16PM -0700, Andy Lutomirski wrote: > On Thu, Apr 28, 2016 at 1:44 PM, Josh Poimboeuf <jpoimboe@redhat.com> wrote: > > Thanks to all the recent x86 entry code refactoring, most tasks' kernel > > stacks start at the same offset right above their saved pt_regs, > > regardless of which syscall was used to enter the kernel. That creates > > a nice convention which makes it straightforward to identify the > > "bottom" of the stack, which can be useful for stack walking code which > > needs to verify the stack is sane. > > > > However there are still a few types of tasks which don't yet follow that > > convention: > > > > 1) CPU idle tasks, aka the "swapper" tasks > > > > 2) freshly forked TIF_FORK tasks which don't have a stack at all > > > > Make the idle tasks conform to the new stack bottom convention by > > starting their stack at a sizeof(pt_regs) offset from the end of the > > stack page. > > > > Signed-off-by: Josh Poimboeuf <jpoimboe@redhat.com> > > --- > > arch/x86/kernel/head_64.S | 7 ++++--- > > 1 file changed, 4 insertions(+), 3 deletions(-) > > > > diff --git a/arch/x86/kernel/head_64.S b/arch/x86/kernel/head_64.S > > index 6dbd2c0..0b12311 100644 > > --- a/arch/x86/kernel/head_64.S > > +++ b/arch/x86/kernel/head_64.S > > @@ -296,8 +296,9 @@ ENTRY(start_cpu) > > * REX.W + FF /5 JMP m16:64 Jump far, absolute indirect, > > * address given in m16:64. > > */ > > - movq initial_code(%rip),%rax > > - pushq $0 # fake return address to stop unwinder > > + call 1f # put return address on stack for unwinder > > +1: xorq %rbp, %rbp # clear frame pointer > > + movq initial_code(%rip), %rax > > pushq $__KERNEL_CS # set correct cs > > pushq %rax # target address in negative space > > lretq > > @@ -325,7 +326,7 @@ ENDPROC(start_cpu0) > > GLOBAL(initial_gs) > > .quad INIT_PER_CPU_VAR(irq_stack_union) > > GLOBAL(initial_stack) > > - .quad init_thread_union+THREAD_SIZE-8 > > + .quad init_thread_union + THREAD_SIZE - SIZEOF_PTREGS > > As long as you're doing this, could you also set orig_ax to -1? I > remember running into some oddities resulting from orig_ax containing > garbage at some point. I assume you mean to initialize the orig_rax value in the pt_regs at the bottom of the stack of the idle task? How could that cause a problem? Since the idle task never returns from a system call, I'd assume that memory never gets accessed? -- Josh
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-04-29 23:40 +0200 |
| Subject | Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks |
| Message-ID | <rtoiK-4yd-15@gated-at.bofh.it> |
| In reply to | #1391425 |
On Fri, Apr 29, 2016 at 1:50 PM, Josh Poimboeuf <jpoimboe@redhat.com> wrote: > On Fri, Apr 29, 2016 at 12:39:16PM -0700, Andy Lutomirski wrote: >> On Thu, Apr 28, 2016 at 1:44 PM, Josh Poimboeuf <jpoimboe@redhat.com> wrote: >> > Thanks to all the recent x86 entry code refactoring, most tasks' kernel >> > stacks start at the same offset right above their saved pt_regs, >> > regardless of which syscall was used to enter the kernel. That creates >> > a nice convention which makes it straightforward to identify the >> > "bottom" of the stack, which can be useful for stack walking code which >> > needs to verify the stack is sane. >> > >> > However there are still a few types of tasks which don't yet follow that >> > convention: >> > >> > 1) CPU idle tasks, aka the "swapper" tasks >> > >> > 2) freshly forked TIF_FORK tasks which don't have a stack at all >> > >> > Make the idle tasks conform to the new stack bottom convention by >> > starting their stack at a sizeof(pt_regs) offset from the end of the >> > stack page. >> > >> > Signed-off-by: Josh Poimboeuf <jpoimboe@redhat.com> >> > --- >> > arch/x86/kernel/head_64.S | 7 ++++--- >> > 1 file changed, 4 insertions(+), 3 deletions(-) >> > >> > diff --git a/arch/x86/kernel/head_64.S b/arch/x86/kernel/head_64.S >> > index 6dbd2c0..0b12311 100644 >> > --- a/arch/x86/kernel/head_64.S >> > +++ b/arch/x86/kernel/head_64.S >> > @@ -296,8 +296,9 @@ ENTRY(start_cpu) >> > * REX.W + FF /5 JMP m16:64 Jump far, absolute indirect, >> > * address given in m16:64. >> > */ >> > - movq initial_code(%rip),%rax >> > - pushq $0 # fake return address to stop unwinder >> > + call 1f # put return address on stack for unwinder >> > +1: xorq %rbp, %rbp # clear frame pointer >> > + movq initial_code(%rip), %rax >> > pushq $__KERNEL_CS # set correct cs >> > pushq %rax # target address in negative space >> > lretq >> > @@ -325,7 +326,7 @@ ENDPROC(start_cpu0) >> > GLOBAL(initial_gs) >> > .quad INIT_PER_CPU_VAR(irq_stack_union) >> > GLOBAL(initial_stack) >> > - .quad init_thread_union+THREAD_SIZE-8 >> > + .quad init_thread_union + THREAD_SIZE - SIZEOF_PTREGS >> >> As long as you're doing this, could you also set orig_ax to -1? I >> remember running into some oddities resulting from orig_ax containing >> garbage at some point. > > I assume you mean to initialize the orig_rax value in the pt_regs at the > bottom of the stack of the idle task? > > How could that cause a problem? Since the idle task never returns from > a system call, I'd assume that memory never gets accessed? > Look at collect_syscall in lib/syscall.c > -- > Josh -- Andy Lutomirski AMA Capital Management, LLC
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-04-30 01:30 +0200 |
| Subject | Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks |
| Message-ID | <rtq1c-6eF-17@gated-at.bofh.it> |
| In reply to | #1391452 |
On Fri, Apr 29, 2016 at 02:38:02PM -0700, Andy Lutomirski wrote: > On Fri, Apr 29, 2016 at 1:50 PM, Josh Poimboeuf <jpoimboe@redhat.com> wrote: > > On Fri, Apr 29, 2016 at 12:39:16PM -0700, Andy Lutomirski wrote: > >> On Thu, Apr 28, 2016 at 1:44 PM, Josh Poimboeuf <jpoimboe@redhat.com> wrote: > >> > Thanks to all the recent x86 entry code refactoring, most tasks' kernel > >> > stacks start at the same offset right above their saved pt_regs, > >> > regardless of which syscall was used to enter the kernel. That creates > >> > a nice convention which makes it straightforward to identify the > >> > "bottom" of the stack, which can be useful for stack walking code which > >> > needs to verify the stack is sane. > >> > > >> > However there are still a few types of tasks which don't yet follow that > >> > convention: > >> > > >> > 1) CPU idle tasks, aka the "swapper" tasks > >> > > >> > 2) freshly forked TIF_FORK tasks which don't have a stack at all > >> > > >> > Make the idle tasks conform to the new stack bottom convention by > >> > starting their stack at a sizeof(pt_regs) offset from the end of the > >> > stack page. > >> > > >> > Signed-off-by: Josh Poimboeuf <jpoimboe@redhat.com> > >> > --- > >> > arch/x86/kernel/head_64.S | 7 ++++--- > >> > 1 file changed, 4 insertions(+), 3 deletions(-) > >> > > >> > diff --git a/arch/x86/kernel/head_64.S b/arch/x86/kernel/head_64.S > >> > index 6dbd2c0..0b12311 100644 > >> > --- a/arch/x86/kernel/head_64.S > >> > +++ b/arch/x86/kernel/head_64.S > >> > @@ -296,8 +296,9 @@ ENTRY(start_cpu) > >> > * REX.W + FF /5 JMP m16:64 Jump far, absolute indirect, > >> > * address given in m16:64. > >> > */ > >> > - movq initial_code(%rip),%rax > >> > - pushq $0 # fake return address to stop unwinder > >> > + call 1f # put return address on stack for unwinder > >> > +1: xorq %rbp, %rbp # clear frame pointer > >> > + movq initial_code(%rip), %rax > >> > pushq $__KERNEL_CS # set correct cs > >> > pushq %rax # target address in negative space > >> > lretq > >> > @@ -325,7 +326,7 @@ ENDPROC(start_cpu0) > >> > GLOBAL(initial_gs) > >> > .quad INIT_PER_CPU_VAR(irq_stack_union) > >> > GLOBAL(initial_stack) > >> > - .quad init_thread_union+THREAD_SIZE-8 > >> > + .quad init_thread_union + THREAD_SIZE - SIZEOF_PTREGS > >> > >> As long as you're doing this, could you also set orig_ax to -1? I > >> remember running into some oddities resulting from orig_ax containing > >> garbage at some point. > > > > I assume you mean to initialize the orig_rax value in the pt_regs at the > > bottom of the stack of the idle task? > > > > How could that cause a problem? Since the idle task never returns from > > a system call, I'd assume that memory never gets accessed? > > > > Look at collect_syscall in lib/syscall.c I don't see how collect_syscall() can be called for the per-cpu idle "swapper" tasks (which is what the above code affects). They don't have pids or /proc entries so you can't do /proc/<pid>/syscall on them. -- Josh
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-04-30 02:20 +0200 |
| Subject | Re: [RFC PATCH v2 03/18] x86/asm/head: standardize the bottom of the stack for idle tasks |
| Message-ID | <rtqNz-6V3-7@gated-at.bofh.it> |
| In reply to | #1391525 |
On Apr 29, 2016 4:27 PM, "Josh Poimboeuf" <jpoimboe@redhat.com> wrote: > > On Fri, Apr 29, 2016 at 02:38:02PM -0700, Andy Lutomirski wrote: > > On Fri, Apr 29, 2016 at 1:50 PM, Josh Poimboeuf <jpoimboe@redhat.com> wrote: > > > On Fri, Apr 29, 2016 at 12:39:16PM -0700, Andy Lutomirski wrote: > > >> On Thu, Apr 28, 2016 at 1:44 PM, Josh Poimboeuf <jpoimboe@redhat.com> wrote: > > >> > Thanks to all the recent x86 entry code refactoring, most tasks' kernel > > >> > stacks start at the same offset right above their saved pt_regs, > > >> > regardless of which syscall was used to enter the kernel. That creates > > >> > a nice convention which makes it straightforward to identify the > > >> > "bottom" of the stack, which can be useful for stack walking code which > > >> > needs to verify the stack is sane. > > >> > > > >> > However there are still a few types of tasks which don't yet follow that > > >> > convention: > > >> > > > >> > 1) CPU idle tasks, aka the "swapper" tasks > > >> > > > >> > 2) freshly forked TIF_FORK tasks which don't have a stack at all > > >> > > > >> > Make the idle tasks conform to the new stack bottom convention by > > >> > starting their stack at a sizeof(pt_regs) offset from the end of the > > >> > stack page. > > >> > > > >> > Signed-off-by: Josh Poimboeuf <jpoimboe@redhat.com> > > >> > --- > > >> > arch/x86/kernel/head_64.S | 7 ++++--- > > >> > 1 file changed, 4 insertions(+), 3 deletions(-) > > >> > > > >> > diff --git a/arch/x86/kernel/head_64.S b/arch/x86/kernel/head_64.S > > >> > index 6dbd2c0..0b12311 100644 > > >> > --- a/arch/x86/kernel/head_64.S > > >> > +++ b/arch/x86/kernel/head_64.S > > >> > @@ -296,8 +296,9 @@ ENTRY(start_cpu) > > >> > * REX.W + FF /5 JMP m16:64 Jump far, absolute indirect, > > >> > * address given in m16:64. > > >> > */ > > >> > - movq initial_code(%rip),%rax > > >> > - pushq $0 # fake return address to stop unwinder > > >> > + call 1f # put return address on stack for unwinder > > >> > +1: xorq %rbp, %rbp # clear frame pointer > > >> > + movq initial_code(%rip), %rax > > >> > pushq $__KERNEL_CS # set correct cs > > >> > pushq %rax # target address in negative space > > >> > lretq > > >> > @@ -325,7 +326,7 @@ ENDPROC(start_cpu0) > > >> > GLOBAL(initial_gs) > > >> > .quad INIT_PER_CPU_VAR(irq_stack_union) > > >> > GLOBAL(initial_stack) > > >> > - .quad init_thread_union+THREAD_SIZE-8 > > >> > + .quad init_thread_union + THREAD_SIZE - SIZEOF_PTREGS > > >> > > >> As long as you're doing this, could you also set orig_ax to -1? I > > >> remember running into some oddities resulting from orig_ax containing > > >> garbage at some point. > > > > > > I assume you mean to initialize the orig_rax value in the pt_regs at the > > > bottom of the stack of the idle task? > > > > > > How could that cause a problem? Since the idle task never returns from > > > a system call, I'd assume that memory never gets accessed? > > > > > > > Look at collect_syscall in lib/syscall.c > > I don't see how collect_syscall() can be called for the per-cpu idle > "swapper" tasks (which is what the above code affects). They don't have > pids or /proc entries so you can't do /proc/<pid>/syscall on them. If so, then never mind. --Andy > > -- > Josh
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-04-28 23:00 +0200 |
| Subject | [RFC PATCH v2 04/18] x86: move _stext marker before head code |
| Message-ID | <rt1cv-21g-17@gated-at.bofh.it> |
| In reply to | #1390507 |
When core_kernel_text() is used to determine whether an address on a
task's stack trace is a kernel text address, it incorrectly returns
false for early text addresses for the head code between the _text and
_stext markers.
Head code is text code too, so mark it as such. This seems to match the
intent of other users of the _stext symbol, and it also seems consistent
with what other architectures are already doing.
Signed-off-by: Josh Poimboeuf <jpoimboe@redhat.com>
---
arch/x86/kernel/vmlinux.lds.S | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/arch/x86/kernel/vmlinux.lds.S b/arch/x86/kernel/vmlinux.lds.S
index 4c941f8..79e15ef 100644
--- a/arch/x86/kernel/vmlinux.lds.S
+++ b/arch/x86/kernel/vmlinux.lds.S
@@ -91,10 +91,10 @@ SECTIONS
/* Text and read-only data */
.text : AT(ADDR(.text) - LOAD_OFFSET) {
_text = .;
+ _stext = .;
/* bootstrapping code */
HEAD_TEXT
. = ALIGN(8);
- _stext = .;
TEXT_TEXT
SCHED_TEXT
LOCK_TEXT
--
2.4.11
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-04-28 23:00 +0200 |
| Subject | [RFC PATCH v2 01/18] x86/asm/head: clean up initial stack variable |
| Message-ID | <rt1cv-21g-19@gated-at.bofh.it> |
| In reply to | #1390507 |
The stack_start variable is similar in usage to initial_code and
initial_gs: they're all stored in head_64.S and they're all updated by
SMP and suspend/resume before starting a CPU.
Rename stack_start to initial_stack to be more consistent with the
others.
Also do a few other related cleanups:
- Remove the unused init_rsp variable declaration.
- Remove the ".word 0" statement after the initial_stack definition
because it has no apparent function.
Signed-off-by: Josh Poimboeuf <jpoimboe@redhat.com>
---
arch/x86/include/asm/realmode.h | 2 +-
arch/x86/include/asm/smp.h | 3 ---
arch/x86/kernel/acpi/sleep.c | 2 +-
arch/x86/kernel/head_32.S | 8 ++++----
arch/x86/kernel/head_64.S | 10 ++++------
arch/x86/kernel/smpboot.c | 2 +-
6 files changed, 11 insertions(+), 16 deletions(-)
diff --git a/arch/x86/include/asm/realmode.h b/arch/x86/include/asm/realmode.h
index 9c6b890..677a671 100644
--- a/arch/x86/include/asm/realmode.h
+++ b/arch/x86/include/asm/realmode.h
@@ -44,9 +44,9 @@ struct trampoline_header {
extern struct real_mode_header *real_mode_header;
extern unsigned char real_mode_blob_end[];
-extern unsigned long init_rsp;
extern unsigned long initial_code;
extern unsigned long initial_gs;
+extern unsigned long initial_stack;
extern unsigned char real_mode_blob[];
extern unsigned char real_mode_relocs[];
diff --git a/arch/x86/include/asm/smp.h b/arch/x86/include/asm/smp.h
index 66b0573..a9ac31b 100644
--- a/arch/x86/include/asm/smp.h
+++ b/arch/x86/include/asm/smp.h
@@ -38,9 +38,6 @@ DECLARE_EARLY_PER_CPU_READ_MOSTLY(u16, x86_bios_cpu_apicid);
DECLARE_EARLY_PER_CPU_READ_MOSTLY(int, x86_cpu_to_logical_apicid);
#endif
-/* Static state in head.S used to set up a CPU */
-extern unsigned long stack_start; /* Initial stack pointer address */
-
struct task_struct;
struct smp_ops {
diff --git a/arch/x86/kernel/acpi/sleep.c b/arch/x86/kernel/acpi/sleep.c
index adb3eaf..4858733 100644
--- a/arch/x86/kernel/acpi/sleep.c
+++ b/arch/x86/kernel/acpi/sleep.c
@@ -99,7 +99,7 @@ int x86_acpi_suspend_lowlevel(void)
saved_magic = 0x12345678;
#else /* CONFIG_64BIT */
#ifdef CONFIG_SMP
- stack_start = (unsigned long)temp_stack + sizeof(temp_stack);
+ initial_stack = (unsigned long)temp_stack + sizeof(temp_stack);
early_gdt_descr.address =
(unsigned long)get_cpu_gdt_table(smp_processor_id());
initial_gs = per_cpu_offset(smp_processor_id());
diff --git a/arch/x86/kernel/head_32.S b/arch/x86/kernel/head_32.S
index 6770865..da840be 100644
--- a/arch/x86/kernel/head_32.S
+++ b/arch/x86/kernel/head_32.S
@@ -94,7 +94,7 @@ RESERVE_BRK(pagetables, INIT_MAP_SIZE)
*/
__HEAD
ENTRY(startup_32)
- movl pa(stack_start),%ecx
+ movl pa(initial_stack),%ecx
/* test KEEP_SEGMENTS flag to see if the bootloader is asking
us to not reload segments */
@@ -286,7 +286,7 @@ num_subarch_entries = (. - subarch_entries) / 4
* start_secondary().
*/
ENTRY(start_cpu0)
- movl stack_start, %ecx
+ movl initial_stack, %ecx
movl %ecx, %esp
jmp *(initial_code)
ENDPROC(start_cpu0)
@@ -307,7 +307,7 @@ ENTRY(startup_32_smp)
movl %eax,%es
movl %eax,%fs
movl %eax,%gs
- movl pa(stack_start),%ecx
+ movl pa(initial_stack),%ecx
movl %eax,%ss
leal -__PAGE_OFFSET(%ecx),%esp
@@ -709,7 +709,7 @@ ENTRY(initial_page_table)
.data
.balign 4
-ENTRY(stack_start)
+ENTRY(initial_stack)
.long init_thread_union+THREAD_SIZE
__INITRODATA
diff --git a/arch/x86/kernel/head_64.S b/arch/x86/kernel/head_64.S
index 5df831e..792c3bb 100644
--- a/arch/x86/kernel/head_64.S
+++ b/arch/x86/kernel/head_64.S
@@ -226,7 +226,7 @@ ENTRY(secondary_startup_64)
movq %rax, %cr0
/* Setup a boot time stack */
- movq stack_start(%rip), %rsp
+ movq initial_stack(%rip), %rsp
/* zero EFLAGS after setting rsp */
pushq $0
@@ -309,7 +309,7 @@ ENTRY(secondary_startup_64)
* start_secondary().
*/
ENTRY(start_cpu0)
- movq stack_start(%rip),%rsp
+ movq initial_stack(%rip),%rsp
movq initial_code(%rip),%rax
pushq $0 # fake return address to stop unwinder
pushq $__KERNEL_CS # set correct cs
@@ -318,17 +318,15 @@ ENTRY(start_cpu0)
ENDPROC(start_cpu0)
#endif
- /* SMP bootup changes these two */
+ /* SMP bootup changes these variables */
__REFDATA
.balign 8
GLOBAL(initial_code)
.quad x86_64_start_kernel
GLOBAL(initial_gs)
.quad INIT_PER_CPU_VAR(irq_stack_union)
-
- GLOBAL(stack_start)
+ GLOBAL(initial_stack)
.quad init_thread_union+THREAD_SIZE-8
- .word 0
__FINITDATA
bad_address:
diff --git a/arch/x86/kernel/smpboot.c b/arch/x86/kernel/smpboot.c
index 1fe4130..503682a 100644
--- a/arch/x86/kernel/smpboot.c
+++ b/arch/x86/kernel/smpboot.c
@@ -950,7 +950,7 @@ static int do_boot_cpu(int apicid, int cpu, struct task_struct *idle)
early_gdt_descr.address = (unsigned long)get_cpu_gdt_table(cpu);
initial_code = (unsigned long)start_secondary;
- stack_start = idle->thread.sp;
+ initial_stack = idle->thread.sp;
/*
* Enable the espfix hack for this CPU
--
2.4.11
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-05-04 10:50 +0200 |
| Subject | Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rv0Fj-DB-1@gated-at.bofh.it> |
| In reply to | #1390507 |
On Thu 2016-04-28 15:44:48, Josh Poimboeuf wrote:
> Change livepatch to use a basic per-task consistency model. This is the
> foundation which will eventually enable us to patch those ~10% of
> security patches which change function or data semantics. This is the
> biggest remaining piece needed to make livepatch more generally useful.
> --- a/kernel/fork.c
> +++ b/kernel/fork.c
> @@ -76,6 +76,7 @@
> #include <linux/compiler.h>
> #include <linux/sysctl.h>
> #include <linux/kcov.h>
> +#include <linux/livepatch.h>
>
> #include <asm/pgtable.h>
> #include <asm/pgalloc.h>
> @@ -1586,6 +1587,8 @@ static struct task_struct *copy_process(unsigned long clone_flags,
> p->parent_exec_id = current->self_exec_id;
> }
>
> + klp_copy_process(p);
I am in doubts here. We copy the state from the parent here. It means
that the new process might still need to be converted. But at the same
point print_context_stack_reliable() returns zero without printing
any stack trace when TIF_FORK flag is set. It means that a freshly
forked task might get be converted immediately. I seems that boot
operations are always done when copy_process() is called. But
they are contradicting each other.
I guess that print_context_stack_reliable() should either return
-EINVAL when TIF_FORK is set. Or it should try to print the
stack of the newly forked task.
Or do I miss something, please?
> +
> spin_lock(¤t->sighand->siglock);
>
> /*
[...]
> diff --git a/kernel/livepatch/transition.c b/kernel/livepatch/transition.c
> new file mode 100644
> index 0000000..92819bb
> --- /dev/null
> +++ b/kernel/livepatch/transition.c
> +/*
> + * This function can be called in the middle of an existing transition to
> + * reverse the direction of the target patch state. This can be done to
> + * effectively cancel an existing enable or disable operation if there are any
> + * tasks which are stuck in the initial patch state.
> + */
> +void klp_reverse_transition(void)
> +{
> + struct klp_patch *patch = klp_transition_patch;
> +
> + klp_target_state = !klp_target_state;
> +
> + /*
> + * Ensure that if another CPU goes through the syscall barrier, sees
> + * the TIF_PATCH_PENDING writes in klp_start_transition(), and calls
> + * klp_patch_task(), it also sees the above write to the target state.
> + * Otherwise it can put the task in the wrong universe.
> + */
> + smp_wmb();
> +
> + klp_start_transition();
> + klp_try_complete_transition();
It is a bit strange that we keep the work scheduled. It might be
better to use
mod_delayed_work(system_wq, &klp_work, 0);
Which triggers more ideas from the nitpicking deparment:
I would move the work definition from core.c to transition.c because
it is closely related to klp_try_complete_transition();
When on it. I would make it more clear that the work is related
to transition. Also I would call queue_delayed_work() directly
instead of adding the klp_schedule_work() wrapper. The delay
might be defined using a constant, e.g.
#define KLP_TRANSITION_DELAY round_jiffies_relative(HZ)
queue_delayed_work(system_wq, &klp_transition_work, KLP_TRANSITION_DELAY);
Finally, the following is always called right after
klp_start_transition(), so I would call it from there.
if (!klp_try_complete_transition())
klp_schedule_work();
> +
> + patch->enabled = !patch->enabled;
> +}
> +
It is really great work! I am checking this patch from left, right, top,
and even bottom and all seems to work well together.
Best Regards,
Petr
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-05-04 18:00 +0200 |
| Subject | Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rv7ns-6NC-11@gated-at.bofh.it> |
| In reply to | #1394089 |
On Wed, May 04, 2016 at 10:42:23AM +0200, Petr Mladek wrote:
> On Thu 2016-04-28 15:44:48, Josh Poimboeuf wrote:
> > Change livepatch to use a basic per-task consistency model. This is the
> > foundation which will eventually enable us to patch those ~10% of
> > security patches which change function or data semantics. This is the
> > biggest remaining piece needed to make livepatch more generally useful.
>
> > --- a/kernel/fork.c
> > +++ b/kernel/fork.c
> > @@ -76,6 +76,7 @@
> > #include <linux/compiler.h>
> > #include <linux/sysctl.h>
> > #include <linux/kcov.h>
> > +#include <linux/livepatch.h>
> >
> > #include <asm/pgtable.h>
> > #include <asm/pgalloc.h>
> > @@ -1586,6 +1587,8 @@ static struct task_struct *copy_process(unsigned long clone_flags,
> > p->parent_exec_id = current->self_exec_id;
> > }
> >
> > + klp_copy_process(p);
>
> I am in doubts here. We copy the state from the parent here. It means
> that the new process might still need to be converted. But at the same
> point print_context_stack_reliable() returns zero without printing
> any stack trace when TIF_FORK flag is set. It means that a freshly
> forked task might get be converted immediately. I seems that boot
> operations are always done when copy_process() is called. But
> they are contradicting each other.
>
> I guess that print_context_stack_reliable() should either return
> -EINVAL when TIF_FORK is set. Or it should try to print the
> stack of the newly forked task.
>
> Or do I miss something, please?
Ok, I admit it's confusing.
A newly forked task doesn't *have* a stack (other than the pt_regs frame
it needs for the return to user space), which is why
print_context_stack_reliable() returns success with an empty array of
addresses.
For a little background, see the second switch_to() macro in
arch/x86/include/asm/switch_to.h. When a newly forked task runs for the
first time, it returns from __switch_to() with no stack. It then jumps
straight to ret_from_fork in entry_64.S, calls a few C functions, and
eventually returns to user space. So, assuming we aren't patching entry
code or the switch_to() macro in __schedule(), it should be safe to
patch the task before it does all that.
With the current code, if an unpatched task gets forked, the child will
also be unpatched. In theory, we could go ahead and patch the child
then. In fact, that's what I did in v1.9.
But in v1.9 discussions it was pointed out that someday maybe the
ret_from_fork stuff will get cleaned up and instead the child stack will
be copied from the parent. In that case the child should inherit its
parent's patched state. So we decided to make it more future-proof by
having the child inherit the parent's patched state.
So, having said all that, I'm really not sure what the best approach is
for print_context_stack_reliable(). Right now I'm thinking I'll change
it back to return -EINVAL for a newly forked task, so it'll be more
future-proof: better to have a false positive than a false negative.
Either way it will probably need to be changed again if the
ret_from_fork code gets cleaned up.
> > +
> > spin_lock(¤t->sighand->siglock);
> >
> > /*
>
> [...]
>
> > diff --git a/kernel/livepatch/transition.c b/kernel/livepatch/transition.c
> > new file mode 100644
> > index 0000000..92819bb
> > --- /dev/null
> > +++ b/kernel/livepatch/transition.c
> > +/*
> > + * This function can be called in the middle of an existing transition to
> > + * reverse the direction of the target patch state. This can be done to
> > + * effectively cancel an existing enable or disable operation if there are any
> > + * tasks which are stuck in the initial patch state.
> > + */
> > +void klp_reverse_transition(void)
> > +{
> > + struct klp_patch *patch = klp_transition_patch;
> > +
> > + klp_target_state = !klp_target_state;
> > +
> > + /*
> > + * Ensure that if another CPU goes through the syscall barrier, sees
> > + * the TIF_PATCH_PENDING writes in klp_start_transition(), and calls
> > + * klp_patch_task(), it also sees the above write to the target state.
> > + * Otherwise it can put the task in the wrong universe.
> > + */
> > + smp_wmb();
> > +
> > + klp_start_transition();
> > + klp_try_complete_transition();
>
> It is a bit strange that we keep the work scheduled. It might be
> better to use
>
> mod_delayed_work(system_wq, &klp_work, 0);
True, I think that would be better.
> Which triggers more ideas from the nitpicking deparment:
>
> I would move the work definition from core.c to transition.c because
> it is closely related to klp_try_complete_transition();
That could be good, but there's a slight problem: klp_work_fn() requires
klp_mutex, which is static to core.c. It's kind of nice to keep the use
of the mutex in core.c only.
> When on it. I would make it more clear that the work is related
> to transition.
How would you recommend doing that? How about:
- rename "klp_work" -> "klp_transition_work"
- rename "klp_work_fn" -> "klp_transition_work_fn"
?
> Also I would call queue_delayed_work() directly
> instead of adding the klp_schedule_work() wrapper. The delay
> might be defined using a constant, e.g.
>
> #define KLP_TRANSITION_DELAY round_jiffies_relative(HZ)
>
> queue_delayed_work(system_wq, &klp_transition_work, KLP_TRANSITION_DELAY);
Sure.
> Finally, the following is always called right after
> klp_start_transition(), so I would call it from there.
>
> if (!klp_try_complete_transition())
> klp_schedule_work();
Except for when it's called by klp_reverse_transition(). And it really
depends on whether we want to allow transition.c to use the mutex. I
don't have a strong opinion either way, I may need to think about it
some more.
> > +
> > + patch->enabled = !patch->enabled;
> > +}
> > +
>
> It is really great work! I am checking this patch from left, right, top,
> and even bottom and all seems to work well together.
Great! Thanks a lot for the thorough review!
--
Josh
[toc] | [prev] | [next] | [standalone]
| From | Miroslav Benes <mbenes@suse.cz> |
|---|---|
| Date | 2016-05-05 11:50 +0200 |
| Subject | Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rvo4X-5Xt-35@gated-at.bofh.it> |
| In reply to | #1394491 |
On Wed, 4 May 2016, Josh Poimboeuf wrote: > On Wed, May 04, 2016 at 10:42:23AM +0200, Petr Mladek wrote: > > On Thu 2016-04-28 15:44:48, Josh Poimboeuf wrote: > > > Change livepatch to use a basic per-task consistency model. This is the > > > foundation which will eventually enable us to patch those ~10% of > > > security patches which change function or data semantics. This is the > > > biggest remaining piece needed to make livepatch more generally useful. > > > > > --- a/kernel/fork.c > > > +++ b/kernel/fork.c > > > @@ -76,6 +76,7 @@ > > > #include <linux/compiler.h> > > > #include <linux/sysctl.h> > > > #include <linux/kcov.h> > > > +#include <linux/livepatch.h> > > > > > > #include <asm/pgtable.h> > > > #include <asm/pgalloc.h> > > > @@ -1586,6 +1587,8 @@ static struct task_struct *copy_process(unsigned long clone_flags, > > > p->parent_exec_id = current->self_exec_id; > > > } > > > > > > + klp_copy_process(p); > > > > I am in doubts here. We copy the state from the parent here. It means > > that the new process might still need to be converted. But at the same > > point print_context_stack_reliable() returns zero without printing > > any stack trace when TIF_FORK flag is set. It means that a freshly > > forked task might get be converted immediately. I seems that boot > > operations are always done when copy_process() is called. But > > they are contradicting each other. > > > > I guess that print_context_stack_reliable() should either return > > -EINVAL when TIF_FORK is set. Or it should try to print the > > stack of the newly forked task. > > > > Or do I miss something, please? > > Ok, I admit it's confusing. > > A newly forked task doesn't *have* a stack (other than the pt_regs frame > it needs for the return to user space), which is why > print_context_stack_reliable() returns success with an empty array of > addresses. > > For a little background, see the second switch_to() macro in > arch/x86/include/asm/switch_to.h. When a newly forked task runs for the > first time, it returns from __switch_to() with no stack. It then jumps > straight to ret_from_fork in entry_64.S, calls a few C functions, and > eventually returns to user space. So, assuming we aren't patching entry > code or the switch_to() macro in __schedule(), it should be safe to > patch the task before it does all that. > > With the current code, if an unpatched task gets forked, the child will > also be unpatched. In theory, we could go ahead and patch the child > then. In fact, that's what I did in v1.9. > > But in v1.9 discussions it was pointed out that someday maybe the > ret_from_fork stuff will get cleaned up and instead the child stack will > be copied from the parent. In that case the child should inherit its > parent's patched state. So we decided to make it more future-proof by > having the child inherit the parent's patched state. > > So, having said all that, I'm really not sure what the best approach is > for print_context_stack_reliable(). Right now I'm thinking I'll change > it back to return -EINVAL for a newly forked task, so it'll be more > future-proof: better to have a false positive than a false negative. > Either way it will probably need to be changed again if the > ret_from_fork code gets cleaned up. I'd be for returning -EINVAL. It is a safe play for now. [...] > > Finally, the following is always called right after > > klp_start_transition(), so I would call it from there. > > > > if (!klp_try_complete_transition()) > > klp_schedule_work(); On the other hand it is quite nice to see the sequence init start try_complete there. Just my 2 cents. > Except for when it's called by klp_reverse_transition(). And it really > depends on whether we want to allow transition.c to use the mutex. I > don't have a strong opinion either way, I may need to think about it > some more. Miroslav
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-05-05 15:10 +0200 |
| Subject | Re: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rvrct-B8-1@gated-at.bofh.it> |
| In reply to | #1394491 |
On Wed 2016-05-04 10:51:21, Josh Poimboeuf wrote:
> On Wed, May 04, 2016 at 10:42:23AM +0200, Petr Mladek wrote:
> > On Thu 2016-04-28 15:44:48, Josh Poimboeuf wrote:
> > > Change livepatch to use a basic per-task consistency model. This is the
> > > foundation which will eventually enable us to patch those ~10% of
> > > security patches which change function or data semantics. This is the
> > > biggest remaining piece needed to make livepatch more generally useful.
> >
> > > --- a/kernel/fork.c
> > > +++ b/kernel/fork.c
> > > @@ -76,6 +76,7 @@
> > > #include <linux/compiler.h>
> > > #include <linux/sysctl.h>
> > > #include <linux/kcov.h>
> > > +#include <linux/livepatch.h>
> > >
> > > #include <asm/pgtable.h>
> > > #include <asm/pgalloc.h>
> > > @@ -1586,6 +1587,8 @@ static struct task_struct *copy_process(unsigned long clone_flags,
> > > p->parent_exec_id = current->self_exec_id;
> > > }
> > >
> > > + klp_copy_process(p);
> >
> > I am in doubts here. We copy the state from the parent here. It means
> > that the new process might still need to be converted. But at the same
> > point print_context_stack_reliable() returns zero without printing
> > any stack trace when TIF_FORK flag is set. It means that a freshly
> > forked task might get be converted immediately. I seems that boot
> > operations are always done when copy_process() is called. But
> > they are contradicting each other.
> >
> > I guess that print_context_stack_reliable() should either return
> > -EINVAL when TIF_FORK is set. Or it should try to print the
> > stack of the newly forked task.
> >
> > Or do I miss something, please?
>
> Ok, I admit it's confusing.
>
> A newly forked task doesn't *have* a stack (other than the pt_regs frame
> it needs for the return to user space), which is why
> print_context_stack_reliable() returns success with an empty array of
> addresses.
>
> For a little background, see the second switch_to() macro in
> arch/x86/include/asm/switch_to.h. When a newly forked task runs for the
> first time, it returns from __switch_to() with no stack. It then jumps
> straight to ret_from_fork in entry_64.S, calls a few C functions, and
> eventually returns to user space. So, assuming we aren't patching entry
> code or the switch_to() macro in __schedule(), it should be safe to
> patch the task before it does all that.
This is great explanation. Thanks for it.
> So, having said all that, I'm really not sure what the best approach is
> for print_context_stack_reliable(). Right now I'm thinking I'll change
> it back to return -EINVAL for a newly forked task, so it'll be more
> future-proof: better to have a false positive than a false negative.
> Either way it will probably need to be changed again if the
> ret_from_fork code gets cleaned up.
I would prefer the -EINVAL. It might safe some hairs when anyone
is working on patching the switch_to stuff. Also it is not that
big loss beacuse most tasks will get migrated on the return to
userspace.
It might help a bit with the newly forked kthreads. But there should
be more safe location where the new kthreads might get migrated,
e.g. right before the main function gets called.
> > > diff --git a/kernel/livepatch/transition.c b/kernel/livepatch/transition.c
> > > new file mode 100644
> > > index 0000000..92819bb
> > > --- /dev/null
> > > +++ b/kernel/livepatch/transition.c
> > > +/*
> > > + * This function can be called in the middle of an existing transition to
> > > + * reverse the direction of the target patch state. This can be done to
> > > + * effectively cancel an existing enable or disable operation if there are any
> > > + * tasks which are stuck in the initial patch state.
> > > + */
> > > +void klp_reverse_transition(void)
> > > +{
> > > + struct klp_patch *patch = klp_transition_patch;
> > > +
> > > + klp_target_state = !klp_target_state;
> > > +
> > > + /*
> > > + * Ensure that if another CPU goes through the syscall barrier, sees
> > > + * the TIF_PATCH_PENDING writes in klp_start_transition(), and calls
> > > + * klp_patch_task(), it also sees the above write to the target state.
> > > + * Otherwise it can put the task in the wrong universe.
> > > + */
> > > + smp_wmb();
> > > +
> > > + klp_start_transition();
> > > + klp_try_complete_transition();
> >
> > It is a bit strange that we keep the work scheduled. It might be
> > better to use
> >
> > mod_delayed_work(system_wq, &klp_work, 0);
>
> True, I think that would be better.
>
> > Which triggers more ideas from the nitpicking deparment:
> >
> > I would move the work definition from core.c to transition.c because
> > it is closely related to klp_try_complete_transition();
>
> That could be good, but there's a slight problem: klp_work_fn() requires
> klp_mutex, which is static to core.c. It's kind of nice to keep the use
> of the mutex in core.c only.
I see and am surprised that we take the lock only in core.c ;-)
I do not have a strong opinion then. Just a small one. The lock guards
also operations from the other .c files. I think that it is only a matter
of time when we will need to access it there. But the work is clearly
transition-related. But it is a real nitpicking. I am sorry for it.
> > When on it. I would make it more clear that the work is related
> > to transition.
>
> How would you recommend doing that? How about:
>
> - rename "klp_work" -> "klp_transition_work"
> - rename "klp_work_fn" -> "klp_transition_work_fn"
Yup, sounds better.
> > Also I would call queue_delayed_work() directly
> > instead of adding the klp_schedule_work() wrapper. The delay
> > might be defined using a constant, e.g.
> >
> > #define KLP_TRANSITION_DELAY round_jiffies_relative(HZ)
> >
> > queue_delayed_work(system_wq, &klp_transition_work, KLP_TRANSITION_DELAY);
>
> Sure.
>
> > Finally, the following is always called right after
> > klp_start_transition(), so I would call it from there.
> >
> > if (!klp_try_complete_transition())
> > klp_schedule_work();
>
> Except for when it's called by klp_reverse_transition(). And it really
> depends on whether we want to allow transition.c to use the mutex. I
> don't have a strong opinion either way, I may need to think about it
> some more.
Ah, I had in mind that it could be replaced by that
mod_delayed_work(system_wq, &klp_transition_work, 0);
So, that we would never call klp_try_complete_transition()
directly. Then it could be the same in all situations. But
it might look strange and be ineffective when really starting
the transition. So, maybe forget about it.
Best Regards,
Petr
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-05-04 14:40 +0200 |
| Subject | barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rv4fV-3Mb-29@gated-at.bofh.it> |
| In reply to | #1390507 |
On Thu 2016-04-28 15:44:48, Josh Poimboeuf wrote:
> Change livepatch to use a basic per-task consistency model. This is the
> foundation which will eventually enable us to patch those ~10% of
> security patches which change function or data semantics. This is the
> biggest remaining piece needed to make livepatch more generally useful.
I spent a lot of time with checking the memory barriers. It seems that
they are basically correct. Let me use my own words to show how
I understand it. I hope that it will help others with review.
> diff --git a/kernel/livepatch/patch.c b/kernel/livepatch/patch.c
> index 782fbb5..b3b8639 100644
> --- a/kernel/livepatch/patch.c
> +++ b/kernel/livepatch/patch.c
> @@ -29,6 +29,7 @@
> #include <linux/bug.h>
> #include <linux/printk.h>
> #include "patch.h"
> +#include "transition.h"
>
> static LIST_HEAD(klp_ops);
>
> @@ -58,11 +59,42 @@ static void notrace klp_ftrace_handler(unsigned long ip,
> ops = container_of(fops, struct klp_ops, fops);
>
> rcu_read_lock();
> +
> func = list_first_or_null_rcu(&ops->func_stack, struct klp_func,
> stack_node);
> - if (WARN_ON_ONCE(!func))
> +
> + if (!func)
> goto unlock;
>
> + /*
> + * See the comment for the 2nd smp_wmb() in klp_init_transition() for
> + * an explanation of why this read barrier is needed.
> + */
> + smp_rmb();
I would prefer to be more explicit, e.g.
/*
* Read the right func->transition when the struct appeared on top of
* func_stack. See klp_init_transition and klp_patch_func().
*/
Note that this barrier is not really needed when the patch is being
disabled, see below.
> +
> + if (unlikely(func->transition)) {
> +
> + /*
> + * See the comment for the 1st smp_wmb() in
> + * klp_init_transition() for an explanation of why this read
> + * barrier is needed.
> + */
> + smp_rmb();
Similar here:
/*
* Read the right initial state when func->transition was
* enabled, see klp_init_transition().
*
* Note that the task must never be migrated to the target
* state when being inside this ftrace handler.
*/
We might want to move the second paragraph on top of the function.
It is a basic and important fact. It actually explains why the first
read barrier is not needed when the patch is being disabled.
There are some more details below. I started to check and comment the
barriers from klp_init_transition().
> + if (current->patch_state == KLP_UNPATCHED) {
> + /*
> + * Use the previously patched version of the function.
> + * If no previous patches exist, use the original
> + * function.
> + */
> + func = list_entry_rcu(func->stack_node.next,
> + struct klp_func, stack_node);
> +
> + if (&func->stack_node == &ops->func_stack)
> + goto unlock;
> + }
> + }
> +
> klp_arch_set_pc(regs, (unsigned long)func->new_func);
> unlock:
> rcu_read_unlock();
> diff --git a/kernel/livepatch/transition.c b/kernel/livepatch/transition.c
> new file mode 100644
> index 0000000..92819bb
> --- /dev/null
> +++ b/kernel/livepatch/transition.c
> +/*
> + * klp_patch_task() - change the patched state of a task
> + * @task: The task to change
> + *
> + * Switches the patched state of the task to the set of functions in the target
> + * patch state.
> + */
> +void klp_patch_task(struct task_struct *task)
> +{
> + clear_tsk_thread_flag(task, TIF_PATCH_PENDING);
> +
> + /*
> + * The corresponding write barriers are in klp_init_transition() and
> + * klp_reverse_transition(). See the comments there for an explanation.
> + */
> + smp_rmb();
I would prefer to be more explicit, e.g.
/*
* Read the correct klp_target_state when TIF_PATCH_PENDING was set
* and this function was called. See klp_init_transition() and
* klp_reverse_transition().
*/
> +
> + task->patch_state = klp_target_state;
> +}
The function name confused me few times when klp_target_state
was KLP_UNPATCHED. I suggest to rename it to klp_update_task()
or klp_transit_task().
> +/*
> + * Initialize the global target patch state and all tasks to the initial patch
> + * state, and initialize all function transition states to true in preparation
> + * for patching or unpatching.
> + */
> +void klp_init_transition(struct klp_patch *patch, int state)
> +{
> + struct task_struct *g, *task;
> + unsigned int cpu;
> + struct klp_object *obj;
> + struct klp_func *func;
> + int initial_state = !state;
> +
> + klp_transition_patch = patch;
> +
> + /*
> + * If the patch can be applied or reverted immediately, skip the
> + * per-task transitions.
> + */
> + if (patch->immediate)
> + return;
> +
> + /*
> + * Initialize all tasks to the initial patch state to prepare them for
> + * switching to the target state.
> + */
> + read_lock(&tasklist_lock);
> + for_each_process_thread(g, task)
> + task->patch_state = initial_state;
> + read_unlock(&tasklist_lock);
> +
> + /*
> + * Ditto for the idle "swapper" tasks.
> + */
> + get_online_cpus();
> + for_each_online_cpu(cpu)
> + idle_task(cpu)->patch_state = initial_state;
> + put_online_cpus();
> +
> + /*
> + * Ensure klp_ftrace_handler() sees the task->patch_state updates
> + * before the func->transition updates. Otherwise it could read an
> + * out-of-date task state and pick the wrong function.
> + */
> + smp_wmb();
This barrier is needed when the patch is being disabled. In this case,
the ftrace handler is already in use and the related struct klp_func
are on top of func_stack. The purpose is well described above.
It is not needed when the patch is being enabled because it is
not visible to the ftrace handler at the moment. The barrier below
is enough.
> + /*
> + * Set the func transition states so klp_ftrace_handler() will know to
> + * switch to the transition logic.
> + *
> + * When patching, the funcs aren't yet in the func_stack and will be
> + * made visible to the ftrace handler shortly by the calls to
> + * klp_patch_object().
> + *
> + * When unpatching, the funcs are already in the func_stack and so are
> + * already visible to the ftrace handler.
> + */
> + klp_for_each_object(patch, obj)
> + klp_for_each_func(obj, func)
> + func->transition = true;
> +
> + /*
> + * Set the global target patch state which tasks will switch to. This
> + * has no effect until the TIF_PATCH_PENDING flags get set later.
> + */
> + klp_target_state = state;
> +
> + /*
> + * For the enable path, ensure klp_ftrace_handler() will see the
> + * func->transition updates before the funcs become visible to the
> + * handler. Otherwise the handler may wrongly pick the new func before
> + * the task switches to the patched state.
By other words, it makes sure that the ftrace handler will see
the updated func->transition before the ftrace handler is registered
and/or before the struct klp_func is listed in func_stack.
> + * For the disable path, the funcs are already visible to the handler.
> + * But we still need to ensure the ftrace handler will see the
> + * func->transition updates before the tasks start switching to the
> + * unpatched state. Otherwise the handler can miss a task patch state
> + * change which would result in it wrongly picking the new function.
If this is true, it would mean that the task might be switched when it
is in the middle of klp_ftrace_handler. It would mean that reading
task->patch_state would be racy against the modification by
klp_patch_task().
Note that before we call klp_patch_task(), the task will stay in the
previous state. We are disabling the patch, so the previous state
is that the patch is enabled. It means that it should always use
the new function before klp_patch_task() is called. It means
that it does not matter if it sees func->transition updated or not.
In both cases, it will use the new function.
Fortunately, task->patch_state might be set to KLP_UNPATCHED
only when it is sleeping or on some other safe location, e.g.
when leaving to userspace. In both cases, the barrier is not needed
here.
By other words, this barrier is not needed to synchronize func_stack
and patch->transition when it is being disabled.
> + * This barrier also ensures that if another CPU goes through the
> + * syscall barrier, sees the TIF_PATCH_PENDING writes in
> + * klp_start_transition(), and calls klp_patch_task(), it also sees the
> + * above write to the target state. Otherwise it can put the task in
> + * the wrong universe.
> + */
By other words, it makes sure that klp_patch_task() will assign the
right patch_state. Where klp_patch_task() could not be called
before we set TIF_PATCH_PENDING in klp_start_transition().
> + smp_wmb();
> +}
> +
> +/*
> + * Start the transition to the specified target patch state so tasks can begin
> + * switching to it.
> + */
> +void klp_start_transition(void)
> +{
> + struct task_struct *g, *task;
> + unsigned int cpu;
> +
> + pr_notice("'%s': %s...\n", klp_transition_patch->mod->name,
> + klp_target_state == KLP_PATCHED ? "patching" : "unpatching");
> +
> + /*
> + * If the patch can be applied or reverted immediately, skip the
> + * per-task transitions.
> + */
> + if (klp_transition_patch->immediate)
> + return;
> +
> + /*
> + * Mark all normal tasks as needing a patch state update. As they pass
> + * through the syscall barrier they'll switch over to the target state
> + * (unless we switch them in klp_try_complete_transition() first).
> + */
> + read_lock(&tasklist_lock);
> + for_each_process_thread(g, task)
> + set_tsk_thread_flag(task, TIF_PATCH_PENDING);
A bad intuition might suggest that we do not need to set this flag
when klp_start_transition() is called from klp_reverse_transition()
and the task already is in the right state.
But I think that we actually must set TIF_PATCH_PENDING even in this
case to avoid a possible race. We do not know if klp_patch_task()
is not running at the moment with the previous klp_target_state().
> + read_unlock(&tasklist_lock);
> +
> + /*
> + * Ditto for the idle "swapper" tasks, though they never cross the
> + * syscall barrier. Instead they switch over in cpu_idle_loop().
> + */
> + get_online_cpus();
> + for_each_online_cpu(cpu)
> + set_tsk_thread_flag(idle_task(cpu), TIF_PATCH_PENDING);
> + put_online_cpus();
> +}
> +
> +/*
> + * This function can be called in the middle of an existing transition to
> + * reverse the direction of the target patch state. This can be done to
> + * effectively cancel an existing enable or disable operation if there are any
> + * tasks which are stuck in the initial patch state.
> + */
> +void klp_reverse_transition(void)
> +{
> + struct klp_patch *patch = klp_transition_patch;
> +
> + klp_target_state = !klp_target_state;
> +
> + /*
> + * Ensure that if another CPU goes through the syscall barrier, sees
> + * the TIF_PATCH_PENDING writes in klp_start_transition(), and calls
> + * klp_patch_task(), it also sees the above write to the target state.
> + * Otherwise it can put the task in the wrong universe.
> + */
> + smp_wmb();
Yup, it is the same reason as for the 2nd barrier in klp_init_transition()
regarding klp_target_state and klp_patch_task() that is triggered by
TIF_PATCH_PENDING.
> +
> + klp_start_transition();
> + klp_try_complete_transition();
> +
> + patch->enabled = !patch->enabled;
> +}
> +
Best Regards,
Petr
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-05-04 16:00 +0200 |
| Subject | Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rv5vk-4VB-15@gated-at.bofh.it> |
| In reply to | #1394246 |
On Wed, May 04, 2016 at 02:39:40PM +0200, Petr Mladek wrote: > > + * This barrier also ensures that if another CPU goes through the > > + * syscall barrier, sees the TIF_PATCH_PENDING writes in > > + * klp_start_transition(), and calls klp_patch_task(), it also sees the > > + * above write to the target state. Otherwise it can put the task in > > + * the wrong universe. > > + */ > > By other words, it makes sure that klp_patch_task() will assign the > right patch_state. Where klp_patch_task() could not be called > before we set TIF_PATCH_PENDING in klp_start_transition(). > > > + smp_wmb(); > > +} So I've not read the patch; but ending a function with an smp_wmb() feels wrong. A wmb orders two stores, and I feel both stores should be well visible in the same function.
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-05-04 19:00 +0200 |
| Subject | Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rv8jx-7F4-29@gated-at.bofh.it> |
| In reply to | #1394313 |
On Wed, May 04, 2016 at 03:53:29PM +0200, Peter Zijlstra wrote: > On Wed, May 04, 2016 at 02:39:40PM +0200, Petr Mladek wrote: > > > + * This barrier also ensures that if another CPU goes through the > > > + * syscall barrier, sees the TIF_PATCH_PENDING writes in > > > + * klp_start_transition(), and calls klp_patch_task(), it also sees the > > > + * above write to the target state. Otherwise it can put the task in > > > + * the wrong universe. (oops, missed a "universe" -> "patch state" rename) > > > + */ > > > > By other words, it makes sure that klp_patch_task() will assign the > > right patch_state. Where klp_patch_task() could not be called > > before we set TIF_PATCH_PENDING in klp_start_transition(). > > > > > + smp_wmb(); > > > +} > > So I've not read the patch; but ending a function with an smp_wmb() > feels wrong. > > A wmb orders two stores, and I feel both stores should be well visible > in the same function. Yeah, I would agree with that. And also, it's probably a red flag that the barrier needs *three* paragraphs to describe the various cases its needed for. However, there are some complications: 1) The stores are in separate functions (which is a generally a good thing as it greatly helps the readability of the code). 2) Which stores are being ordered depends on whether the function is called in the enable path or the disable path. 3) Either way it actually orders *two* separate pairs of stores. Anyway I'm thinking I should move that barrier out of klp_init_transition() and into its callers. The stores will still be in separate functions but at least there will be better visibility of where the stores are occurring, and the comments can be a little more focused. -- Josh
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-05-04 16:20 +0200 |
| Subject | Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rv5OF-5wR-1@gated-at.bofh.it> |
| In reply to | #1394246 |
On Wed 2016-05-04 14:39:40, Petr Mladek wrote: > * > * Note that the task must never be migrated to the target > * state when being inside this ftrace handler. > */ > > We might want to move the second paragraph on top of the function. > It is a basic and important fact. It actually explains why the first > read barrier is not needed when the patch is being disabled. I wrote the statement partly intuitively. I think that it is really somehow important. And I am slightly in doubts if we are on the safe side. First, why is it important that the task->patch_state is not switched when being inside the ftrace handler? If we are inside the handler, we are kind-of inside the called function. And the basic idea of this consistency model is that we must not switch a task when it is inside a patched function. This is normally decided by the stack. The handler is a bit special because it is called right before the function. If it was the only patched function on the stack, it would not matter if we choose the new or old code. Both decisions would be safe for the moment. The fun starts when the function calls another patched function. The other patched function must be called consistently with the first one. If the first function was from the patch, the other must be from the patch as well and vice versa. This is why we must not switch task->patch_state dangerously when being inside the ftrace handler. Now I am not sure if this condition is fulfilled. The ftrace handler is called as the very first instruction of the function. Does not it break the stack validity? Could we sleep inside the ftrace handler? Will the patched function be detected on the stack? Or is my brain already too far in the fantasy world? Best regards, Petr
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-05-04 19:30 +0200 |
| Subject | Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rv8Mz-8iq-31@gated-at.bofh.it> |
| In reply to | #1394323 |
On Wed, May 04, 2016 at 04:12:05PM +0200, Petr Mladek wrote: > On Wed 2016-05-04 14:39:40, Petr Mladek wrote: > > * > > * Note that the task must never be migrated to the target > > * state when being inside this ftrace handler. > > */ > > > > We might want to move the second paragraph on top of the function. > > It is a basic and important fact. It actually explains why the first > > read barrier is not needed when the patch is being disabled. > > I wrote the statement partly intuitively. I think that it is really > somehow important. And I am slightly in doubts if we are on the safe side. > > First, why is it important that the task->patch_state is not switched > when being inside the ftrace handler? > > If we are inside the handler, we are kind-of inside the called > function. And the basic idea of this consistency model is that > we must not switch a task when it is inside a patched function. > This is normally decided by the stack. > > The handler is a bit special because it is called right before the > function. If it was the only patched function on the stack, it would > not matter if we choose the new or old code. Both decisions would > be safe for the moment. > > The fun starts when the function calls another patched function. > The other patched function must be called consistently with > the first one. If the first function was from the patch, > the other must be from the patch as well and vice versa. > > This is why we must not switch task->patch_state dangerously > when being inside the ftrace handler. > > Now I am not sure if this condition is fulfilled. The ftrace handler > is called as the very first instruction of the function. Does not > it break the stack validity? Could we sleep inside the ftrace > handler? Will the patched function be detected on the stack? > > Or is my brain already too far in the fantasy world? I think this isn't a possibility. In today's code base, this can't happen because task patch states are only switched when sleeping or when exiting the kernel. The ftrace handler doesn't sleep directly. If it were preempted, it couldn't be switched there either because we consider preempted stacks to be unreliable. In theory, a DWARF stack trace of a preempted task *could* be reliable. But then the DWARF unwinder should be smart enough to see that the original function called the ftrace handler. Right? So the stack would be reliable, but then livepatch would see the original function on the stack and wouldn't switch the task. Does that make sense? -- Josh
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-05-05 13:30 +0200 |
| Subject | Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rvpDH-7rd-5@gated-at.bofh.it> |
| In reply to | #1394573 |
On Wed 2016-05-04 12:25:17, Josh Poimboeuf wrote: > On Wed, May 04, 2016 at 04:12:05PM +0200, Petr Mladek wrote: > > On Wed 2016-05-04 14:39:40, Petr Mladek wrote: > > > * > > > * Note that the task must never be migrated to the target > > > * state when being inside this ftrace handler. > > > */ > > > > > > We might want to move the second paragraph on top of the function. > > > It is a basic and important fact. It actually explains why the first > > > read barrier is not needed when the patch is being disabled. > > > > I wrote the statement partly intuitively. I think that it is really > > somehow important. And I am slightly in doubts if we are on the safe side. > > > > First, why is it important that the task->patch_state is not switched > > when being inside the ftrace handler? > > > > If we are inside the handler, we are kind-of inside the called > > function. And the basic idea of this consistency model is that > > we must not switch a task when it is inside a patched function. > > This is normally decided by the stack. > > > > The handler is a bit special because it is called right before the > > function. If it was the only patched function on the stack, it would > > not matter if we choose the new or old code. Both decisions would > > be safe for the moment. > > > > The fun starts when the function calls another patched function. > > The other patched function must be called consistently with > > the first one. If the first function was from the patch, > > the other must be from the patch as well and vice versa. > > > > This is why we must not switch task->patch_state dangerously > > when being inside the ftrace handler. > > > > Now I am not sure if this condition is fulfilled. The ftrace handler > > is called as the very first instruction of the function. Does not > > it break the stack validity? Could we sleep inside the ftrace > > handler? Will the patched function be detected on the stack? > > > > Or is my brain already too far in the fantasy world? > > I think this isn't a possibility. > > In today's code base, this can't happen because task patch states are > only switched when sleeping or when exiting the kernel. The ftrace > handler doesn't sleep directly. > > If it were preempted, it couldn't be switched there either because we > consider preempted stacks to be unreliable. This was the missing piece. > In theory, a DWARF stack trace of a preempted task *could* be reliable. > But then the DWARF unwinder should be smart enough to see that the > original function called the ftrace handler. Right? So the stack would > be reliable, but then livepatch would see the original function on the > stack and wouldn't switch the task. > > Does that make sense? Yup. I think that we are on the safe side. Thanks for explanation. Best Regards, Petr
[toc] | [prev] | [next] | [standalone]
| From | Miroslav Benes <mbenes@suse.cz> |
|---|---|
| Date | 2016-05-09 17:50 +0200 |
| Subject | Re: barriers: was: [RFC PATCH v2 17/18] livepatch: change to a per-task consistency model |
| Message-ID | <rwVBw-8t-3@gated-at.bofh.it> |
| In reply to | #1394573 |
On Wed, 4 May 2016, Josh Poimboeuf wrote:
> On Wed, May 04, 2016 at 04:12:05PM +0200, Petr Mladek wrote:
> > On Wed 2016-05-04 14:39:40, Petr Mladek wrote:
> > > *
> > > * Note that the task must never be migrated to the target
> > > * state when being inside this ftrace handler.
> > > */
> > >
> > > We might want to move the second paragraph on top of the function.
> > > It is a basic and important fact. It actually explains why the first
> > > read barrier is not needed when the patch is being disabled.
> >
> > I wrote the statement partly intuitively. I think that it is really
> > somehow important. And I am slightly in doubts if we are on the safe side.
> >
> > First, why is it important that the task->patch_state is not switched
> > when being inside the ftrace handler?
> >
> > If we are inside the handler, we are kind-of inside the called
> > function. And the basic idea of this consistency model is that
> > we must not switch a task when it is inside a patched function.
> > This is normally decided by the stack.
> >
> > The handler is a bit special because it is called right before the
> > function. If it was the only patched function on the stack, it would
> > not matter if we choose the new or old code. Both decisions would
> > be safe for the moment.
> >
> > The fun starts when the function calls another patched function.
> > The other patched function must be called consistently with
> > the first one. If the first function was from the patch,
> > the other must be from the patch as well and vice versa.
> >
> > This is why we must not switch task->patch_state dangerously
> > when being inside the ftrace handler.
> >
> > Now I am not sure if this condition is fulfilled. The ftrace handler
> > is called as the very first instruction of the function. Does not
> > it break the stack validity? Could we sleep inside the ftrace
> > handler? Will the patched function be detected on the stack?
> >
> > Or is my brain already too far in the fantasy world?
>
> I think this isn't a possibility.
>
> In today's code base, this can't happen because task patch states are
> only switched when sleeping or when exiting the kernel. The ftrace
> handler doesn't sleep directly.
>
> If it were preempted, it couldn't be switched there either because we
> consider preempted stacks to be unreliable.
And IIRC ftrace handlers cannot sleep and are called with preemption
disabled as of now. The code is a bit obscure, but see
__ftrace_ops_list_func for example. This is "main" ftrace handler that
calls all the registered ones in case FTRACE_OPS_FL_DYNAMIC is set (which
is always true for handlers coming from modules) and CONFIG_PREEMPT is
on. If it is off and there is only one handler registered for a function
dynamic trampoline is used. See commit 12cce594fa8f ("ftrace/x86: Allow
!CONFIG_PREEMPT dynamic ops to use allocated trampolines"). I think
Steven had a plan to implement dynamic trampolines even for
CONFIG_PREEMPT case but he still hasn't done it. It should use RCU_TASKS
infrastructure.
The reason for all the mess is that ftrace needs to be sure that no task
is in the handler when the handler/trampoline is freed.
So we should be safe for now even from this side.
Miroslav
[toc] | [prev] | [next] | [standalone]
Page 3 of 4 — ← Prev page 1 2 [3] 4 Next page →
Back to top | Article view | linux.kernel
csiph-web