Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1531728 > unrolled thread
| Started by | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| First post | 2016-11-28 23:00 +0100 |
| Last post | 2016-11-30 12:10 +0100 |
| Articles | 20 on this page of 29 — 4 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Josh Poimboeuf <jpoimboe@redhat.com> - 2016-11-28 23:00 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-29 01:50 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Josh Poimboeuf <jpoimboe@redhat.com> - 2016-11-29 07:00 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Peter Zijlstra <peterz@infradead.org> - 2016-11-29 10:20 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-29 15:10 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Josh Poimboeuf <jpoimboe@redhat.com> - 2016-11-29 16:10 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Petr Mladek <pmladek@suse.com> - 2016-11-29 17:20 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-29 19:10 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-29 18:00 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Josh Poimboeuf <jpoimboe@redhat.com> - 2016-11-29 18:20 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-29 18:40 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Petr Mladek <pmladek@suse.com> - 2016-11-30 11:00 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-29 11:30 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Peter Zijlstra <peterz@infradead.org> - 2016-11-29 13:50 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-29 16:20 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Petr Mladek <pmladek@suse.com> - 2016-11-29 17:30 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Peter Zijlstra <peterz@infradead.org> - 2016-11-29 18:20 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-29 20:50 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Peter Zijlstra <peterz@infradead.org> - 2016-11-29 21:00 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-29 21:10 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-29 21:40 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Josh Poimboeuf <jpoimboe@redhat.com> - 2016-11-30 20:20 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-11-30 21:00 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Peter Zijlstra <peterz@infradead.org> - 2016-12-01 07:00 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-12-01 13:40 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Peter Zijlstra <peterz@infradead.org> - 2016-12-01 17:50 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2016-12-01 18:10 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Petr Mladek <pmladek@suse.com> - 2016-11-30 11:10 +0100
Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start Peter Zijlstra <peterz@infradead.org> - 2016-11-30 12:10 +0100
Page 1 of 2 [1] 2 Next page →
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-11-28 23:00 +0100 |
| Subject | Re: perf: fuzzer BUG: KASAN: stack-out-of-bounds in __unwind_start |
| Message-ID | <sIC7T-7Cs-11@gated-at.bofh.it> |
On Thu, Nov 24, 2016 at 12:33:48PM -0500, Vince Weaver wrote: > > This is on a skylake machine, linus git as of yesterday after the various > kasan-related fixes went in. Not sure if there were any that hadn't hit > upstream yet. > > Anyway I can't tell from this one what the actual trigger is. After this > mess the fuzzer process was locked, udev started complaining, and it > eventually died completely after a few hours of repeated messages like > this. > > [38898.373183] INFO: rcu_sched self-detected stall on CPU > [38898.378452] 7-...: (5249 ticks this GP) idle=727/140000000000001/0 softirq=3141908/3141908 fqs=2625 > [38898.381211] INFO: rcu_sched detected stalls on CPUs/tasks: > [38898.381214] 0-...: (1 GPs behind) idle=05f/140000000000001/2 softirq=3285458/3285459 fqs=2625 > [38898.381217] 7-...: (5249 ticks this GP) idle=727/140000000000001/0 softirq=3141908/3141908 fqs=2625 > [38898.381218] (detected by 1, t=5252 jiffies, g=3685053, c=3685052, q=32) > [38898.381244] ================================================================== > [38898.381247] BUG: KASAN: stack-out-of-bounds in __unwind_start+0x1a2/0x1c0 at addr ffff8801e9727c28 > [38898.381248] Read of size 8 by task swapper/1/0 > [38898.381250] page:ffffea0007a5c9c0 count:0 mapcount:0 mapping: (null) index:0x0 > [38898.381251] flags: 0x2ffff8000000000() > [38898.381251] page dumped because: kasan: bad access detected > [38898.381328] Memory state around the buggy address: > [38898.381330] ffff8801e9727b00: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 > [38898.381331] ffff8801e9727b80: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 > [38898.381332] >ffff8801e9727c00: 00 00 00 00 f1 f1 f1 f1 00 00 00 00 f3 f3 f3 f3 > [38898.381333] ^ > [38898.381334] ffff8801e9727c80: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 f1 > [38898.381335] ffff8801e9727d00: f1 f1 f1 00 00 00 f4 f2 f2 f2 f2 00 00 00 00 f3 > [38898.381335] ================================================================== > [38898.510702] (t=5284 jiffies g=3685053 c=3685052 q=32) > > (That's all, the report above repeats but no useful things like a > backtrace are ever printed) After looking at the RCU stall detection code, I think the KASAN error and missing stack dump aren't very surprising. RCU calls the scheduler dump_cpu_task() function, which seems inherently problematic: it tries to dump the stack of a task while it's running on another CPU. There are some issues with that: 1) There's no way to find the starting frame of a currently running task from another CPU. In fact, I'm wondering how dump_cpu_task() ever worked at all? It seems like you'd have to get lucky that the sp/bp registers stored by the last call to schedule() happen to point to a currently valid stack frame. 2) Even if there were a way to find the starting frame, it's racy because the target task could be overwriting the stack while we're reading it. 3) IRQ/exception stack dumps would be missing anyway because the stack dump code only looks at the current CPU's interrupt stacks. Maybe dump_cpu_task() should instead run the stack dump directly from the target CPU, e.g. with trigger_single_cpu_backtrace() or smp_call_function_single()? Paul, Peter, Ingo, any thoughts? -- Josh
[toc] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-29 01:50 +0100 |
| Message-ID | <sIEMq-T5-15@gated-at.bofh.it> |
| In reply to | #1531728 |
On Mon, Nov 28, 2016 at 03:54:11PM -0600, Josh Poimboeuf wrote: > On Thu, Nov 24, 2016 at 12:33:48PM -0500, Vince Weaver wrote: > > > > This is on a skylake machine, linus git as of yesterday after the various > > kasan-related fixes went in. Not sure if there were any that hadn't hit > > upstream yet. > > > > Anyway I can't tell from this one what the actual trigger is. After this > > mess the fuzzer process was locked, udev started complaining, and it > > eventually died completely after a few hours of repeated messages like > > this. > > > > [38898.373183] INFO: rcu_sched self-detected stall on CPU > > [38898.378452] 7-...: (5249 ticks this GP) idle=727/140000000000001/0 softirq=3141908/3141908 fqs=2625 > > [38898.381211] INFO: rcu_sched detected stalls on CPUs/tasks: > > [38898.381214] 0-...: (1 GPs behind) idle=05f/140000000000001/2 softirq=3285458/3285459 fqs=2625 > > [38898.381217] 7-...: (5249 ticks this GP) idle=727/140000000000001/0 softirq=3141908/3141908 fqs=2625 > > [38898.381218] (detected by 1, t=5252 jiffies, g=3685053, c=3685052, q=32) > > [38898.381244] ================================================================== > > [38898.381247] BUG: KASAN: stack-out-of-bounds in __unwind_start+0x1a2/0x1c0 at addr ffff8801e9727c28 > > [38898.381248] Read of size 8 by task swapper/1/0 > > [38898.381250] page:ffffea0007a5c9c0 count:0 mapcount:0 mapping: (null) index:0x0 > > [38898.381251] flags: 0x2ffff8000000000() > > [38898.381251] page dumped because: kasan: bad access detected > > [38898.381328] Memory state around the buggy address: > > [38898.381330] ffff8801e9727b00: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 > > [38898.381331] ffff8801e9727b80: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 > > [38898.381332] >ffff8801e9727c00: 00 00 00 00 f1 f1 f1 f1 00 00 00 00 f3 f3 f3 f3 > > [38898.381333] ^ > > [38898.381334] ffff8801e9727c80: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 f1 > > [38898.381335] ffff8801e9727d00: f1 f1 f1 00 00 00 f4 f2 f2 f2 f2 00 00 00 00 f3 > > [38898.381335] ================================================================== > > [38898.510702] (t=5284 jiffies g=3685053 c=3685052 q=32) > > > > (That's all, the report above repeats but no useful things like a > > backtrace are ever printed) > > After looking at the RCU stall detection code, I think the KASAN error > and missing stack dump aren't very surprising. RCU calls the scheduler > dump_cpu_task() function, which seems inherently problematic: it tries > to dump the stack of a task while it's running on another CPU. > > There are some issues with that: > > 1) There's no way to find the starting frame of a currently running task > from another CPU. > > In fact, I'm wondering how dump_cpu_task() ever worked at all? It > seems like you'd have to get lucky that the sp/bp registers stored by > the last call to schedule() happen to point to a currently valid > stack frame. > > 2) Even if there were a way to find the starting frame, it's racy > because the target task could be overwriting the stack while we're > reading it. > > 3) IRQ/exception stack dumps would be missing anyway because the stack > dump code only looks at the current CPU's interrupt stacks. > > Maybe dump_cpu_task() should instead run the stack dump directly from > the target CPU, e.g. with trigger_single_cpu_backtrace() or > smp_call_function_single()? > > Paul, Peter, Ingo, any thoughts? We used to do that, but the resulting NMIs were problematic on some platforms. Perhaps things have gotten better? Thaxn, Paul
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-11-29 07:00 +0100 |
| Message-ID | <sIJCp-4cE-3@gated-at.bofh.it> |
| In reply to | #1531828 |
On Mon, Nov 28, 2016 at 04:40:21PM -0800, Paul E. McKenney wrote:
> On Mon, Nov 28, 2016 at 03:54:11PM -0600, Josh Poimboeuf wrote:
> > On Thu, Nov 24, 2016 at 12:33:48PM -0500, Vince Weaver wrote:
> > >
> > > This is on a skylake machine, linus git as of yesterday after the various
> > > kasan-related fixes went in. Not sure if there were any that hadn't hit
> > > upstream yet.
> > >
> > > Anyway I can't tell from this one what the actual trigger is. After this
> > > mess the fuzzer process was locked, udev started complaining, and it
> > > eventually died completely after a few hours of repeated messages like
> > > this.
> > >
> > > [38898.373183] INFO: rcu_sched self-detected stall on CPU
> > > [38898.378452] 7-...: (5249 ticks this GP) idle=727/140000000000001/0 softirq=3141908/3141908 fqs=2625
> > > [38898.381211] INFO: rcu_sched detected stalls on CPUs/tasks:
> > > [38898.381214] 0-...: (1 GPs behind) idle=05f/140000000000001/2 softirq=3285458/3285459 fqs=2625
> > > [38898.381217] 7-...: (5249 ticks this GP) idle=727/140000000000001/0 softirq=3141908/3141908 fqs=2625
> > > [38898.381218] (detected by 1, t=5252 jiffies, g=3685053, c=3685052, q=32)
> > > [38898.381244] ==================================================================
> > > [38898.381247] BUG: KASAN: stack-out-of-bounds in __unwind_start+0x1a2/0x1c0 at addr ffff8801e9727c28
> > > [38898.381248] Read of size 8 by task swapper/1/0
> > > [38898.381250] page:ffffea0007a5c9c0 count:0 mapcount:0 mapping: (null) index:0x0
> > > [38898.381251] flags: 0x2ffff8000000000()
> > > [38898.381251] page dumped because: kasan: bad access detected
> > > [38898.381328] Memory state around the buggy address:
> > > [38898.381330] ffff8801e9727b00: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
> > > [38898.381331] ffff8801e9727b80: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
> > > [38898.381332] >ffff8801e9727c00: 00 00 00 00 f1 f1 f1 f1 00 00 00 00 f3 f3 f3 f3
> > > [38898.381333] ^
> > > [38898.381334] ffff8801e9727c80: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 f1
> > > [38898.381335] ffff8801e9727d00: f1 f1 f1 00 00 00 f4 f2 f2 f2 f2 00 00 00 00 f3
> > > [38898.381335] ==================================================================
> > > [38898.510702] (t=5284 jiffies g=3685053 c=3685052 q=32)
> > >
> > > (That's all, the report above repeats but no useful things like a
> > > backtrace are ever printed)
> >
> > After looking at the RCU stall detection code, I think the KASAN error
> > and missing stack dump aren't very surprising. RCU calls the scheduler
> > dump_cpu_task() function, which seems inherently problematic: it tries
> > to dump the stack of a task while it's running on another CPU.
> >
> > There are some issues with that:
> >
> > 1) There's no way to find the starting frame of a currently running task
> > from another CPU.
> >
> > In fact, I'm wondering how dump_cpu_task() ever worked at all? It
> > seems like you'd have to get lucky that the sp/bp registers stored by
> > the last call to schedule() happen to point to a currently valid
> > stack frame.
> >
> > 2) Even if there were a way to find the starting frame, it's racy
> > because the target task could be overwriting the stack while we're
> > reading it.
> >
> > 3) IRQ/exception stack dumps would be missing anyway because the stack
> > dump code only looks at the current CPU's interrupt stacks.
> >
> > Maybe dump_cpu_task() should instead run the stack dump directly from
> > the target CPU, e.g. with trigger_single_cpu_backtrace() or
> > smp_call_function_single()?
> >
> > Paul, Peter, Ingo, any thoughts?
>
> We used to do that, but the resulting NMIs were problematic on some
> platforms. Perhaps things have gotten better?
Did a little digging on git blame and found the following commit (which
seems to be the cause of the KASAN warning and missing stack dump):
bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
I presume this commit is still needed because of the NMI printk deadlock
issues which were discussed at Kernel Summit. I guess those issues need
to be sorted out before the above commit can be reverted.
--
Josh
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-11-29 10:20 +0100 |
| Message-ID | <sIMJY-6yn-23@gated-at.bofh.it> |
| In reply to | #1531916 |
On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> > We used to do that, but the resulting NMIs were problematic on some
> > platforms. Perhaps things have gotten better?
>
> Did a little digging on git blame and found the following commit (which
> seems to be the cause of the KASAN warning and missing stack dump):
>
> bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
>
> I presume this commit is still needed because of the NMI printk deadlock
> issues which were discussed at Kernel Summit. I guess those issues need
> to be sorted out before the above commit can be reverted.
so printk should more or less work from NMI, esp. after:
42a0bb3f7138 ("printk/nmi: generic solution for safe printk in NMI")
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-29 15:10 +0100 |
| Message-ID | <sIRgB-12s-5@gated-at.bofh.it> |
| In reply to | #1532020 |
On Tue, Nov 29, 2016 at 10:16:50AM +0100, Peter Zijlstra wrote:
> On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> > > We used to do that, but the resulting NMIs were problematic on some
> > > platforms. Perhaps things have gotten better?
> >
> > Did a little digging on git blame and found the following commit (which
> > seems to be the cause of the KASAN warning and missing stack dump):
> >
> > bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
> >
> > I presume this commit is still needed because of the NMI printk deadlock
> > issues which were discussed at Kernel Summit. I guess those issues need
> > to be sorted out before the above commit can be reverted.
>
> so printk should more or less work from NMI, esp. after:
>
> 42a0bb3f7138 ("printk/nmi: generic solution for safe printk in NMI")
And of course bc1dce514e9b doesn't revert cleanly, but see hand reversion
below. Also, 42a0bb3f7138's commit log calls out MN10300 and Xtensa as
needing more work. Has that happened?
But I really like the fact that RCU CPU stall warnings dump only those
stacks that are likely to be involved, and the patch below goes back
to dumping everyone. Shouldn't be that hard to fix, though...
Thanx, Paul
------------------------------------------------------------------------
commit e7c9d76ed508fe978c6657e33f4de1b160ee4efe
Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Date: Tue Nov 29 05:49:06 2016 -0800
rcu: Once again use NMI-based stack traces in stall warnings
This commit is for all intents and purposes a revert of bc1dce514e9b
("rcu: Don't use NMIs to dump other CPUs' stacks"). The reason to
suppose that this can now safely be reverted is the presence of
42a0bb3f7138 ("printk/nmi: generic solution for safe printk in NMI"),
which is said to have made NMI-based stack dumps safe.
Not-yet-signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Cc: Petr Mladek <pmladek@suse.com>
Cc: Josh Poimboeuf <jpoimboe@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
index 91a68e4e6671..d73ccd4bed86 100644
--- a/kernel/rcu/tree.c
+++ b/kernel/rcu/tree.c
@@ -1396,7 +1396,10 @@ static void rcu_check_gp_kthread_starvation(struct rcu_state *rsp)
}
/*
- * Dump stacks of all tasks running on stalled CPUs.
+ * Dump stacks of all tasks running on stalled CPUs. First try using
+ * NMIs, but fall back to manual remote stack tracing on architectures
+ * that don't support NMI-based stack dumps. The NMI-triggered stack
+ * traces are more accurate because they are printed by the target CPU.
*/
static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
{
@@ -1404,6 +1407,8 @@ static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
unsigned long flags;
struct rcu_node *rnp;
+ if (trigger_all_cpu_backtrace())
+ return;
rcu_for_each_leaf_node(rsp, rnp) {
raw_spin_lock_irqsave_rcu_node(rnp, flags);
if (rnp->qsmask != 0) {
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-11-29 16:10 +0100 |
| Message-ID | <sIScF-1DN-1@gated-at.bofh.it> |
| In reply to | #1532318 |
On Tue, Nov 29, 2016 at 06:07:34AM -0800, Paul E. McKenney wrote:
> On Tue, Nov 29, 2016 at 10:16:50AM +0100, Peter Zijlstra wrote:
> > On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> > > > We used to do that, but the resulting NMIs were problematic on some
> > > > platforms. Perhaps things have gotten better?
> > >
> > > Did a little digging on git blame and found the following commit (which
> > > seems to be the cause of the KASAN warning and missing stack dump):
> > >
> > > bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
> > >
> > > I presume this commit is still needed because of the NMI printk deadlock
> > > issues which were discussed at Kernel Summit. I guess those issues need
> > > to be sorted out before the above commit can be reverted.
> >
> > so printk should more or less work from NMI, esp. after:
> >
> > 42a0bb3f7138 ("printk/nmi: generic solution for safe printk in NMI")
>
> And of course bc1dce514e9b doesn't revert cleanly, but see hand reversion
> below. Also, 42a0bb3f7138's commit log calls out MN10300 and Xtensa as
> needing more work. Has that happened?
Petr M, any idea?
> But I really like the fact that RCU CPU stall warnings dump only those
> stacks that are likely to be involved, and the patch below goes back
> to dumping everyone. Shouldn't be that hard to fix, though...
There's a new trigger_single_cpu_backtrace() function which can be used
for that.
> ------------------------------------------------------------------------
>
> commit e7c9d76ed508fe978c6657e33f4de1b160ee4efe
> Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> Date: Tue Nov 29 05:49:06 2016 -0800
>
> rcu: Once again use NMI-based stack traces in stall warnings
>
> This commit is for all intents and purposes a revert of bc1dce514e9b
> ("rcu: Don't use NMIs to dump other CPUs' stacks"). The reason to
> suppose that this can now safely be reverted is the presence of
> 42a0bb3f7138 ("printk/nmi: generic solution for safe printk in NMI"),
> which is said to have made NMI-based stack dumps safe.
>
> Not-yet-signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> Cc: Petr Mladek <pmladek@suse.com>
> Cc: Josh Poimboeuf <jpoimboe@redhat.com>
> Cc: Peter Zijlstra <peterz@infradead.org>
>
> diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
> index 91a68e4e6671..d73ccd4bed86 100644
> --- a/kernel/rcu/tree.c
> +++ b/kernel/rcu/tree.c
> @@ -1396,7 +1396,10 @@ static void rcu_check_gp_kthread_starvation(struct rcu_state *rsp)
> }
>
> /*
> - * Dump stacks of all tasks running on stalled CPUs.
> + * Dump stacks of all tasks running on stalled CPUs. First try using
> + * NMIs, but fall back to manual remote stack tracing on architectures
> + * that don't support NMI-based stack dumps. The NMI-triggered stack
> + * traces are more accurate because they are printed by the target CPU.
> */
> static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
> {
> @@ -1404,6 +1407,8 @@ static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
> unsigned long flags;
> struct rcu_node *rnp;
>
> + if (trigger_all_cpu_backtrace())
> + return;
> rcu_for_each_leaf_node(rsp, rnp) {
> raw_spin_lock_irqsave_rcu_node(rnp, flags);
> if (rnp->qsmask != 0) {
>
--
Josh
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-11-29 17:20 +0100 |
| Message-ID | <sITiq-2kT-39@gated-at.bofh.it> |
| In reply to | #1532379 |
On Tue 2016-11-29 09:09:17, Josh Poimboeuf wrote:
> On Tue, Nov 29, 2016 at 06:07:34AM -0800, Paul E. McKenney wrote:
> > On Tue, Nov 29, 2016 at 10:16:50AM +0100, Peter Zijlstra wrote:
> > > On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> > > > > We used to do that, but the resulting NMIs were problematic on some
> > > > > platforms. Perhaps things have gotten better?
> > > >
> > > > Did a little digging on git blame and found the following commit (which
> > > > seems to be the cause of the KASAN warning and missing stack dump):
> > > >
> > > > bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
> > > >
> > > > I presume this commit is still needed because of the NMI printk deadlock
> > > > issues which were discussed at Kernel Summit. I guess those issues need
> > > > to be sorted out before the above commit can be reverted.
> > >
> > > so printk should more or less work from NMI, esp. after:
> > >
> > > 42a0bb3f7138 ("printk/nmi: generic solution for safe printk in NMI")
> >
> > And of course bc1dce514e9b doesn't revert cleanly, but see hand reversion
> > below. Also, 42a0bb3f7138's commit log calls out MN10300 and Xtensa as
> > needing more work. Has that happened?
>
> Petr M, any idea?
These two architectures do not support the safe printk in NMI. But
these architectures also do not implement trigger_all_cpu_backtrace()
and other trigger_*_backtrace() functions. Therefore these functions
return false there.
In fact, only very few architectures implement trigger_*_backtrace().
And only few of them use NMI (x86, arm, tile). I have just double
checked that these all use the safe printk in NMI.
By other words, if trigger_all_cpu_backtrace() or
trigger_single_cpu_backtrace() returns true, it should be NMI safe
and you could use it here.
> > But I really like the fact that RCU CPU stall warnings dump only those
> > stacks that are likely to be involved, and the patch below goes back
> > to dumping everyone. Shouldn't be that hard to fix, though...
>
> There's a new trigger_single_cpu_backtrace() function which can be used
> for that.
There is newly also trigger_cpumask_backtrace(struct cpumask *mask)
where you could select more CPUs using the mask. If this is of any help.
Best Regards,
Petr
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-29 19:10 +0100 |
| Message-ID | <sIV0S-3tJ-5@gated-at.bofh.it> |
| In reply to | #1532475 |
On Tue, Nov 29, 2016 at 05:12:46PM +0100, Petr Mladek wrote:
> On Tue 2016-11-29 09:09:17, Josh Poimboeuf wrote:
> > On Tue, Nov 29, 2016 at 06:07:34AM -0800, Paul E. McKenney wrote:
> > > On Tue, Nov 29, 2016 at 10:16:50AM +0100, Peter Zijlstra wrote:
> > > > On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> > > > > > We used to do that, but the resulting NMIs were problematic on some
> > > > > > platforms. Perhaps things have gotten better?
> > > > >
> > > > > Did a little digging on git blame and found the following commit (which
> > > > > seems to be the cause of the KASAN warning and missing stack dump):
> > > > >
> > > > > bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
> > > > >
> > > > > I presume this commit is still needed because of the NMI printk deadlock
> > > > > issues which were discussed at Kernel Summit. I guess those issues need
> > > > > to be sorted out before the above commit can be reverted.
> > > >
> > > > so printk should more or less work from NMI, esp. after:
> > > >
> > > > 42a0bb3f7138 ("printk/nmi: generic solution for safe printk in NMI")
> > >
> > > And of course bc1dce514e9b doesn't revert cleanly, but see hand reversion
> > > below. Also, 42a0bb3f7138's commit log calls out MN10300 and Xtensa as
> > > needing more work. Has that happened?
> >
> > Petr M, any idea?
>
> These two architectures do not support the safe printk in NMI. But
> these architectures also do not implement trigger_all_cpu_backtrace()
> and other trigger_*_backtrace() functions. Therefore these functions
> return false there.
>
> In fact, only very few architectures implement trigger_*_backtrace().
> And only few of them use NMI (x86, arm, tile). I have just double
> checked that these all use the safe printk in NMI.
>
> By other words, if trigger_all_cpu_backtrace() or
> trigger_single_cpu_backtrace() returns true, it should be NMI safe
> and you could use it here.
Good, I will upgrade my commit to Signed-off-by, then.
> > > But I really like the fact that RCU CPU stall warnings dump only those
> > > stacks that are likely to be involved, and the patch below goes back
> > > to dumping everyone. Shouldn't be that hard to fix, though...
> >
> > There's a new trigger_single_cpu_backtrace() function which can be used
> > for that.
>
> There is newly also trigger_cpumask_backtrace(struct cpumask *mask)
> where you could select more CPUs using the mask. If this is of any help.
In my experience, there is almost never a large number of CPUs stalling
a given RCU grace period. But thank you for letting me know about
trigger_cpumask_backtrace(), as it might be useful in the future.
Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-29 18:00 +0100 |
| Message-ID | <sITV9-2B3-53@gated-at.bofh.it> |
| In reply to | #1532379 |
On Tue, Nov 29, 2016 at 09:09:17AM -0600, Josh Poimboeuf wrote:
> On Tue, Nov 29, 2016 at 06:07:34AM -0800, Paul E. McKenney wrote:
> > On Tue, Nov 29, 2016 at 10:16:50AM +0100, Peter Zijlstra wrote:
> > > On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> > > > > We used to do that, but the resulting NMIs were problematic on some
> > > > > platforms. Perhaps things have gotten better?
> > > >
> > > > Did a little digging on git blame and found the following commit (which
> > > > seems to be the cause of the KASAN warning and missing stack dump):
> > > >
> > > > bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
> > > >
> > > > I presume this commit is still needed because of the NMI printk deadlock
> > > > issues which were discussed at Kernel Summit. I guess those issues need
> > > > to be sorted out before the above commit can be reverted.
> > >
> > > so printk should more or less work from NMI, esp. after:
> > >
> > > 42a0bb3f7138 ("printk/nmi: generic solution for safe printk in NMI")
> >
> > And of course bc1dce514e9b doesn't revert cleanly, but see hand reversion
> > below. Also, 42a0bb3f7138's commit log calls out MN10300 and Xtensa as
> > needing more work. Has that happened?
>
> Petr M, any idea?
My Not-yet-signed-off-by is due to this concern, FWIW.
> > But I really like the fact that RCU CPU stall warnings dump only those
> > stacks that are likely to be involved, and the patch below goes back
> > to dumping everyone. Shouldn't be that hard to fix, though...
>
> There's a new trigger_single_cpu_backtrace() function which can be used
> for that.
Even better, thank you! Killed an hour or so of coding, but I must
confess that it was a mercy killing. ;-)
Much nicer (but completely untested) patch below.
Thanx, Paul
------------------------------------------------------------------------
commit d3515ee46e0cff880170e48a05e8f2791b507758
Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Date: Tue Nov 29 05:49:06 2016 -0800
rcu: Once again use NMI-based stack traces in stall warnings
This commit is for all intents and purposes a revert of bc1dce514e9b
("rcu: Don't use NMIs to dump other CPUs' stacks"). The reason to suppose
that this can now safely be reverted is the presence of 42a0bb3f7138
("printk/nmi: generic solution for safe printk in NMI"), which is said
to have made NMI-based stack dumps safe.
However, this reversion keeps one nice property of bc1dce514e9b
("rcu: Don't use NMIs to dump other CPUs' stacks"), namely that
only those CPUs blocking the grace period are dumped. The new
trigger_single_cpu_backtrace() is used to make this happen, as
suggested by Josh Poimboeuf.
Not-yet-signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Cc: Petr Mladek <pmladek@suse.com>
Cc: Josh Poimboeuf <jpoimboe@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
index 91a68e4e6671..ba0e4825be9d 100644
--- a/kernel/rcu/tree.c
+++ b/kernel/rcu/tree.c
@@ -1396,7 +1396,10 @@ static void rcu_check_gp_kthread_starvation(struct rcu_state *rsp)
}
/*
- * Dump stacks of all tasks running on stalled CPUs.
+ * Dump stacks of all tasks running on stalled CPUs. First try using
+ * NMIs, but fall back to manual remote stack tracing on architectures
+ * that don't support NMI-based stack dumps. The NMI-triggered stack
+ * traces are more accurate because they are printed by the target CPU.
*/
static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
{
@@ -1406,11 +1409,10 @@ static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
rcu_for_each_leaf_node(rsp, rnp) {
raw_spin_lock_irqsave_rcu_node(rnp, flags);
- if (rnp->qsmask != 0) {
- for_each_leaf_node_possible_cpu(rnp, cpu)
- if (rnp->qsmask & leaf_node_cpu_bit(rnp, cpu))
+ for_each_leaf_node_possible_cpu(rnp, cpu)
+ if (rnp->qsmask & leaf_node_cpu_bit(rnp, cpu))
+ if (!trigger_single_cpu_backtrace(cpu))
dump_cpu_task(cpu);
- }
raw_spin_unlock_irqrestore_rcu_node(rnp, flags);
}
}
diff --git a/kernel/rcu/tree.h b/kernel/rcu/tree.h
index 7dcdd59d894c..c0a4bf8f1ed0 100644
--- a/kernel/rcu/tree.h
+++ b/kernel/rcu/tree.h
@@ -691,18 +691,6 @@ static inline void rcu_nocb_q_lengths(struct rcu_data *rdp, long *ql, long *qll)
#endif /* #ifdef CONFIG_RCU_TRACE */
/*
- * Place this after a lock-acquisition primitive to guarantee that
- * an UNLOCK+LOCK pair act as a full barrier. This guarantee applies
- * if the UNLOCK and LOCK are executed by the same CPU or if the
- * UNLOCK and LOCK operate on the same lock variable.
- */
-#ifdef CONFIG_PPC
-#define smp_mb__after_unlock_lock() smp_mb() /* Full ordering for lock. */
-#else /* #ifdef CONFIG_PPC */
-#define smp_mb__after_unlock_lock() do { } while (0)
-#endif /* #else #ifdef CONFIG_PPC */
-
-/*
* Wrappers for the rcu_node::lock acquire and release.
*
* Because the rcu_nodes form a tree, the tree traversal locking will observe
[toc] | [prev] | [next] | [standalone]
| From | Josh Poimboeuf <jpoimboe@redhat.com> |
|---|---|
| Date | 2016-11-29 18:20 +0100 |
| Message-ID | <sIUev-2Xs-77@gated-at.bofh.it> |
| In reply to | #1532530 |
On Tue, Nov 29, 2016 at 08:51:52AM -0800, Paul E. McKenney wrote:
> On Tue, Nov 29, 2016 at 09:09:17AM -0600, Josh Poimboeuf wrote:
> > On Tue, Nov 29, 2016 at 06:07:34AM -0800, Paul E. McKenney wrote:
> > > On Tue, Nov 29, 2016 at 10:16:50AM +0100, Peter Zijlstra wrote:
> > > > On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> > > > > > We used to do that, but the resulting NMIs were problematic on some
> > > > > > platforms. Perhaps things have gotten better?
> > > > >
> > > > > Did a little digging on git blame and found the following commit (which
> > > > > seems to be the cause of the KASAN warning and missing stack dump):
> > > > >
> > > > > bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
> > > > >
> > > > > I presume this commit is still needed because of the NMI printk deadlock
> > > > > issues which were discussed at Kernel Summit. I guess those issues need
> > > > > to be sorted out before the above commit can be reverted.
> > > >
> > > > so printk should more or less work from NMI, esp. after:
> > > >
> > > > 42a0bb3f7138 ("printk/nmi: generic solution for safe printk in NMI")
> > >
> > > And of course bc1dce514e9b doesn't revert cleanly, but see hand reversion
> > > below. Also, 42a0bb3f7138's commit log calls out MN10300 and Xtensa as
> > > needing more work. Has that happened?
> >
> > Petr M, any idea?
>
> My Not-yet-signed-off-by is due to this concern, FWIW.
I think Petr's replies have addressed that now.
> > > But I really like the fact that RCU CPU stall warnings dump only those
> > > stacks that are likely to be involved, and the patch below goes back
> > > to dumping everyone. Shouldn't be that hard to fix, though...
> >
> > There's a new trigger_single_cpu_backtrace() function which can be used
> > for that.
>
> Even better, thank you! Killed an hour or so of coding, but I must
> confess that it was a mercy killing. ;-)
Ha :-)
> Much nicer (but completely untested) patch below.
The kernel/rcu/tree.h changes seem intended for another patch?
Otherwise:
Reviewed-by: Josh Poimboeuf <jpoimboe@redhat.com>
Also I think this will fix the KASAN warnings reported by Vince, so you
might add:
Reported-by: Vince Weaver <vincent.weaver@maine.edu>
>
> Thanx, Paul
>
> ------------------------------------------------------------------------
>
> commit d3515ee46e0cff880170e48a05e8f2791b507758
> Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> Date: Tue Nov 29 05:49:06 2016 -0800
>
> rcu: Once again use NMI-based stack traces in stall warnings
>
> This commit is for all intents and purposes a revert of bc1dce514e9b
> ("rcu: Don't use NMIs to dump other CPUs' stacks"). The reason to suppose
> that this can now safely be reverted is the presence of 42a0bb3f7138
> ("printk/nmi: generic solution for safe printk in NMI"), which is said
> to have made NMI-based stack dumps safe.
>
> However, this reversion keeps one nice property of bc1dce514e9b
> ("rcu: Don't use NMIs to dump other CPUs' stacks"), namely that
> only those CPUs blocking the grace period are dumped. The new
> trigger_single_cpu_backtrace() is used to make this happen, as
> suggested by Josh Poimboeuf.
>
> Not-yet-signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> Cc: Petr Mladek <pmladek@suse.com>
> Cc: Josh Poimboeuf <jpoimboe@redhat.com>
> Cc: Peter Zijlstra <peterz@infradead.org>
>
> diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
> index 91a68e4e6671..ba0e4825be9d 100644
> --- a/kernel/rcu/tree.c
> +++ b/kernel/rcu/tree.c
> @@ -1396,7 +1396,10 @@ static void rcu_check_gp_kthread_starvation(struct rcu_state *rsp)
> }
>
> /*
> - * Dump stacks of all tasks running on stalled CPUs.
> + * Dump stacks of all tasks running on stalled CPUs. First try using
> + * NMIs, but fall back to manual remote stack tracing on architectures
> + * that don't support NMI-based stack dumps. The NMI-triggered stack
> + * traces are more accurate because they are printed by the target CPU.
> */
> static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
> {
> @@ -1406,11 +1409,10 @@ static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
>
> rcu_for_each_leaf_node(rsp, rnp) {
> raw_spin_lock_irqsave_rcu_node(rnp, flags);
> - if (rnp->qsmask != 0) {
> - for_each_leaf_node_possible_cpu(rnp, cpu)
> - if (rnp->qsmask & leaf_node_cpu_bit(rnp, cpu))
> + for_each_leaf_node_possible_cpu(rnp, cpu)
> + if (rnp->qsmask & leaf_node_cpu_bit(rnp, cpu))
> + if (!trigger_single_cpu_backtrace(cpu))
> dump_cpu_task(cpu);
> - }
> raw_spin_unlock_irqrestore_rcu_node(rnp, flags);
> }
> }
> diff --git a/kernel/rcu/tree.h b/kernel/rcu/tree.h
> index 7dcdd59d894c..c0a4bf8f1ed0 100644
> --- a/kernel/rcu/tree.h
> +++ b/kernel/rcu/tree.h
> @@ -691,18 +691,6 @@ static inline void rcu_nocb_q_lengths(struct rcu_data *rdp, long *ql, long *qll)
> #endif /* #ifdef CONFIG_RCU_TRACE */
>
> /*
> - * Place this after a lock-acquisition primitive to guarantee that
> - * an UNLOCK+LOCK pair act as a full barrier. This guarantee applies
> - * if the UNLOCK and LOCK are executed by the same CPU or if the
> - * UNLOCK and LOCK operate on the same lock variable.
> - */
> -#ifdef CONFIG_PPC
> -#define smp_mb__after_unlock_lock() smp_mb() /* Full ordering for lock. */
> -#else /* #ifdef CONFIG_PPC */
> -#define smp_mb__after_unlock_lock() do { } while (0)
> -#endif /* #else #ifdef CONFIG_PPC */
> -
> -/*
> * Wrappers for the rcu_node::lock acquire and release.
> *
> * Because the rcu_nodes form a tree, the tree traversal locking will observe
>
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-29 18:40 +0100 |
| Message-ID | <sIUxQ-34p-21@gated-at.bofh.it> |
| In reply to | #1532574 |
On Tue, Nov 29, 2016 at 11:17:25AM -0600, Josh Poimboeuf wrote:
> On Tue, Nov 29, 2016 at 08:51:52AM -0800, Paul E. McKenney wrote:
> > On Tue, Nov 29, 2016 at 09:09:17AM -0600, Josh Poimboeuf wrote:
> > > On Tue, Nov 29, 2016 at 06:07:34AM -0800, Paul E. McKenney wrote:
> > > > On Tue, Nov 29, 2016 at 10:16:50AM +0100, Peter Zijlstra wrote:
> > > > > On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> > > > > > > We used to do that, but the resulting NMIs were problematic on some
> > > > > > > platforms. Perhaps things have gotten better?
> > > > > >
> > > > > > Did a little digging on git blame and found the following commit (which
> > > > > > seems to be the cause of the KASAN warning and missing stack dump):
> > > > > >
> > > > > > bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
> > > > > >
> > > > > > I presume this commit is still needed because of the NMI printk deadlock
> > > > > > issues which were discussed at Kernel Summit. I guess those issues need
> > > > > > to be sorted out before the above commit can be reverted.
> > > > >
> > > > > so printk should more or less work from NMI, esp. after:
> > > > >
> > > > > 42a0bb3f7138 ("printk/nmi: generic solution for safe printk in NMI")
> > > >
> > > > And of course bc1dce514e9b doesn't revert cleanly, but see hand reversion
> > > > below. Also, 42a0bb3f7138's commit log calls out MN10300 and Xtensa as
> > > > needing more work. Has that happened?
> > >
> > > Petr M, any idea?
> >
> > My Not-yet-signed-off-by is due to this concern, FWIW.
>
> I think Petr's replies have addressed that now.
>
> > > > But I really like the fact that RCU CPU stall warnings dump only those
> > > > stacks that are likely to be involved, and the patch below goes back
> > > > to dumping everyone. Shouldn't be that hard to fix, though...
> > >
> > > There's a new trigger_single_cpu_backtrace() function which can be used
> > > for that.
> >
> > Even better, thank you! Killed an hour or so of coding, but I must
> > confess that it was a mercy killing. ;-)
>
> Ha :-)
>
> > Much nicer (but completely untested) patch below.
>
> The kernel/rcu/tree.h changes seem intended for another patch?
Indeed it was, thank you for catching this, fixed.
> Otherwise:
>
> Reviewed-by: Josh Poimboeuf <jpoimboe@redhat.com>
>
> Also I think this will fix the KASAN warnings reported by Vince, so you
> might add:
>
> Reported-by: Vince Weaver <vincent.weaver@maine.edu>
Added both of these, thank you!
Updated (but still untested) commit below.
Thanx, Paul
------------------------------------------------------------------------
commit d3df9bc5fb5d838b049f32a476721eadbc349553
Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Date: Tue Nov 29 05:49:06 2016 -0800
rcu: Once again use NMI-based stack traces in stall warnings
This commit is for all intents and purposes a revert of bc1dce514e9b
("rcu: Don't use NMIs to dump other CPUs' stacks"). The reason to suppose
that this can now safely be reverted is the presence of 42a0bb3f7138
("printk/nmi: generic solution for safe printk in NMI"), which is said
to have made NMI-based stack dumps safe.
However, this reversion keeps one nice property of bc1dce514e9b
("rcu: Don't use NMIs to dump other CPUs' stacks"), namely that
only those CPUs blocking the grace period are dumped. The new
trigger_single_cpu_backtrace() is used to make this happen, as
suggested by Josh Poimboeuf.
Reported-by: Vince Weaver <vincent.weaver@maine.edu>
Not-yet-signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Cc: Petr Mladek <pmladek@suse.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Reviewed-by: Josh Poimboeuf <jpoimboe@redhat.com>
diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
index 91a68e4e6671..ba0e4825be9d 100644
--- a/kernel/rcu/tree.c
+++ b/kernel/rcu/tree.c
@@ -1396,7 +1396,10 @@ static void rcu_check_gp_kthread_starvation(struct rcu_state *rsp)
}
/*
- * Dump stacks of all tasks running on stalled CPUs.
+ * Dump stacks of all tasks running on stalled CPUs. First try using
+ * NMIs, but fall back to manual remote stack tracing on architectures
+ * that don't support NMI-based stack dumps. The NMI-triggered stack
+ * traces are more accurate because they are printed by the target CPU.
*/
static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
{
@@ -1406,11 +1409,10 @@ static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
rcu_for_each_leaf_node(rsp, rnp) {
raw_spin_lock_irqsave_rcu_node(rnp, flags);
- if (rnp->qsmask != 0) {
- for_each_leaf_node_possible_cpu(rnp, cpu)
- if (rnp->qsmask & leaf_node_cpu_bit(rnp, cpu))
+ for_each_leaf_node_possible_cpu(rnp, cpu)
+ if (rnp->qsmask & leaf_node_cpu_bit(rnp, cpu))
+ if (!trigger_single_cpu_backtrace(cpu))
dump_cpu_task(cpu);
- }
raw_spin_unlock_irqrestore_rcu_node(rnp, flags);
}
}
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-11-30 11:00 +0100 |
| Message-ID | <sJ9Qe-4q9-17@gated-at.bofh.it> |
| In reply to | #1532607 |
On Tue 2016-11-29 09:36:00, Paul E. McKenney wrote:
> Updated (but still untested) commit below.
>
>
> Thanx, Paul
>
> ------------------------------------------------------------------------
>
> commit d3df9bc5fb5d838b049f32a476721eadbc349553
> Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> Date: Tue Nov 29 05:49:06 2016 -0800
>
> rcu: Once again use NMI-based stack traces in stall warnings
>
> This commit is for all intents and purposes a revert of bc1dce514e9b
> ("rcu: Don't use NMIs to dump other CPUs' stacks"). The reason to suppose
> that this can now safely be reverted is the presence of 42a0bb3f7138
> ("printk/nmi: generic solution for safe printk in NMI"), which is said
> to have made NMI-based stack dumps safe.
>
> However, this reversion keeps one nice property of bc1dce514e9b
> ("rcu: Don't use NMIs to dump other CPUs' stacks"), namely that
> only those CPUs blocking the grace period are dumped. The new
> trigger_single_cpu_backtrace() is used to make this happen, as
> suggested by Josh Poimboeuf.
>
> Reported-by: Vince Weaver <vincent.weaver@maine.edu>
> Not-yet-signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> Cc: Petr Mladek <pmladek@suse.com>
> Cc: Peter Zijlstra <peterz@infradead.org>
> Reviewed-by: Josh Poimboeuf <jpoimboe@redhat.com>
Looks fine to me.
Reviewed-by: Petr Mladek <pmladek@suse.com>
Best Regards,
Petr
> diff --git a/kernel/rcu/tree.c b/kernel/rcu/tree.c
> index 91a68e4e6671..ba0e4825be9d 100644
> --- a/kernel/rcu/tree.c
> +++ b/kernel/rcu/tree.c
> @@ -1396,7 +1396,10 @@ static void rcu_check_gp_kthread_starvation(struct rcu_state *rsp)
> }
>
> /*
> - * Dump stacks of all tasks running on stalled CPUs.
> + * Dump stacks of all tasks running on stalled CPUs. First try using
> + * NMIs, but fall back to manual remote stack tracing on architectures
> + * that don't support NMI-based stack dumps. The NMI-triggered stack
> + * traces are more accurate because they are printed by the target CPU.
> */
> static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
> {
> @@ -1406,11 +1409,10 @@ static void rcu_dump_cpu_stacks(struct rcu_state *rsp)
>
> rcu_for_each_leaf_node(rsp, rnp) {
> raw_spin_lock_irqsave_rcu_node(rnp, flags);
> - if (rnp->qsmask != 0) {
> - for_each_leaf_node_possible_cpu(rnp, cpu)
> - if (rnp->qsmask & leaf_node_cpu_bit(rnp, cpu))
> + for_each_leaf_node_possible_cpu(rnp, cpu)
> + if (rnp->qsmask & leaf_node_cpu_bit(rnp, cpu))
> + if (!trigger_single_cpu_backtrace(cpu))
> dump_cpu_task(cpu);
> - }
> raw_spin_unlock_irqrestore_rcu_node(rnp, flags);
> }
> }
>
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-29 11:30 +0100 |
| Message-ID | <sINPI-7eO-21@gated-at.bofh.it> |
| In reply to | #1531916 |
On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> On Mon, Nov 28, 2016 at 04:40:21PM -0800, Paul E. McKenney wrote:
> > On Mon, Nov 28, 2016 at 03:54:11PM -0600, Josh Poimboeuf wrote:
> > > On Thu, Nov 24, 2016 at 12:33:48PM -0500, Vince Weaver wrote:
> > > >
> > > > This is on a skylake machine, linus git as of yesterday after the various
> > > > kasan-related fixes went in. Not sure if there were any that hadn't hit
> > > > upstream yet.
> > > >
> > > > Anyway I can't tell from this one what the actual trigger is. After this
> > > > mess the fuzzer process was locked, udev started complaining, and it
> > > > eventually died completely after a few hours of repeated messages like
> > > > this.
> > > >
> > > > [38898.373183] INFO: rcu_sched self-detected stall on CPU
> > > > [38898.378452] 7-...: (5249 ticks this GP) idle=727/140000000000001/0 softirq=3141908/3141908 fqs=2625
> > > > [38898.381211] INFO: rcu_sched detected stalls on CPUs/tasks:
> > > > [38898.381214] 0-...: (1 GPs behind) idle=05f/140000000000001/2 softirq=3285458/3285459 fqs=2625
> > > > [38898.381217] 7-...: (5249 ticks this GP) idle=727/140000000000001/0 softirq=3141908/3141908 fqs=2625
> > > > [38898.381218] (detected by 1, t=5252 jiffies, g=3685053, c=3685052, q=32)
> > > > [38898.381244] ==================================================================
> > > > [38898.381247] BUG: KASAN: stack-out-of-bounds in __unwind_start+0x1a2/0x1c0 at addr ffff8801e9727c28
> > > > [38898.381248] Read of size 8 by task swapper/1/0
> > > > [38898.381250] page:ffffea0007a5c9c0 count:0 mapcount:0 mapping: (null) index:0x0
> > > > [38898.381251] flags: 0x2ffff8000000000()
> > > > [38898.381251] page dumped because: kasan: bad access detected
> > > > [38898.381328] Memory state around the buggy address:
> > > > [38898.381330] ffff8801e9727b00: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
> > > > [38898.381331] ffff8801e9727b80: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
> > > > [38898.381332] >ffff8801e9727c00: 00 00 00 00 f1 f1 f1 f1 00 00 00 00 f3 f3 f3 f3
> > > > [38898.381333] ^
> > > > [38898.381334] ffff8801e9727c80: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 f1
> > > > [38898.381335] ffff8801e9727d00: f1 f1 f1 00 00 00 f4 f2 f2 f2 f2 00 00 00 00 f3
> > > > [38898.381335] ==================================================================
> > > > [38898.510702] (t=5284 jiffies g=3685053 c=3685052 q=32)
> > > >
> > > > (That's all, the report above repeats but no useful things like a
> > > > backtrace are ever printed)
> > >
> > > After looking at the RCU stall detection code, I think the KASAN error
> > > and missing stack dump aren't very surprising. RCU calls the scheduler
> > > dump_cpu_task() function, which seems inherently problematic: it tries
> > > to dump the stack of a task while it's running on another CPU.
> > >
> > > There are some issues with that:
> > >
> > > 1) There's no way to find the starting frame of a currently running task
> > > from another CPU.
> > >
> > > In fact, I'm wondering how dump_cpu_task() ever worked at all? It
> > > seems like you'd have to get lucky that the sp/bp registers stored by
> > > the last call to schedule() happen to point to a currently valid
> > > stack frame.
> > >
> > > 2) Even if there were a way to find the starting frame, it's racy
> > > because the target task could be overwriting the stack while we're
> > > reading it.
> > >
> > > 3) IRQ/exception stack dumps would be missing anyway because the stack
> > > dump code only looks at the current CPU's interrupt stacks.
> > >
> > > Maybe dump_cpu_task() should instead run the stack dump directly from
> > > the target CPU, e.g. with trigger_single_cpu_backtrace() or
> > > smp_call_function_single()?
> > >
> > > Paul, Peter, Ingo, any thoughts?
> >
> > We used to do that, but the resulting NMIs were problematic on some
> > platforms. Perhaps things have gotten better?
>
> Did a little digging on git blame and found the following commit (which
> seems to be the cause of the KASAN warning and missing stack dump):
>
> bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
>
> I presume this commit is still needed because of the NMI printk deadlock
> issues which were discussed at Kernel Summit. I guess those issues need
> to be sorted out before the above commit can be reverted.
Agreed -- it would be very good to revert that commit, but not until
it is safe to do so. ;-)
Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-11-29 13:50 +0100 |
| Message-ID | <sIQ1b-5T-17@gated-at.bofh.it> |
| In reply to | #1531916 |
On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> Did a little digging on git blame and found the following commit (which
> seems to be the cause of the KASAN warning and missing stack dump):
>
> bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
>
> I presume this commit is still needed because of the NMI printk deadlock
> issues which were discussed at Kernel Summit. I guess those issues need
> to be sorted out before the above commit can be reverted.
Also, I most always run with these here patches applied:
https://lkml.kernel.org/r/20161018170830.405990950@infradead.org
People are very busy polishing the turd we call printk, but from where
I'm sitting its terminally and unfixably broken.
I should certainly add a revert of the above commit to the stack of
patches I carry.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-29 16:20 +0100 |
| Message-ID | <sISmm-1He-21@gated-at.bofh.it> |
| In reply to | #1532248 |
On Tue, Nov 29, 2016 at 01:43:23PM +0100, Peter Zijlstra wrote:
> On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
>
> > Did a little digging on git blame and found the following commit (which
> > seems to be the cause of the KASAN warning and missing stack dump):
> >
> > bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
> >
> > I presume this commit is still needed because of the NMI printk deadlock
> > issues which were discussed at Kernel Summit. I guess those issues need
> > to be sorted out before the above commit can be reverted.
>
> Also, I most always run with these here patches applied:
>
> https://lkml.kernel.org/r/20161018170830.405990950@infradead.org
>
> People are very busy polishing the turd we call printk, but from where
> I'm sitting its terminally and unfixably broken.
>
> I should certainly add a revert of the above commit to the stack of
> patches I carry.
This isn't making me feel particularly confident about switching RCU
CPU stall warnings back to NMIs... ;-)
Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-11-29 17:30 +0100 |
| Message-ID | <sITs6-2o0-25@gated-at.bofh.it> |
| In reply to | #1532407 |
On Tue 2016-11-29 07:10:04, Paul E. McKenney wrote:
> On Tue, Nov 29, 2016 at 01:43:23PM +0100, Peter Zijlstra wrote:
> > On Mon, Nov 28, 2016 at 11:52:41PM -0600, Josh Poimboeuf wrote:
> >
> > > Did a little digging on git blame and found the following commit (which
> > > seems to be the cause of the KASAN warning and missing stack dump):
> > >
> > > bc1dce514e9b ("rcu: Don't use NMIs to dump other CPUs' stacks")
> > >
> > > I presume this commit is still needed because of the NMI printk deadlock
> > > issues which were discussed at Kernel Summit. I guess those issues need
> > > to be sorted out before the above commit can be reverted.
> >
> > Also, I most always run with these here patches applied:
> >
> > https://lkml.kernel.org/r/20161018170830.405990950@infradead.org
> >
> > People are very busy polishing the turd we call printk, but from where
> > I'm sitting its terminally and unfixably broken.
I still hope that we could do better :-)
> > I should certainly add a revert of the above commit to the stack of
> > patches I carry.
>
> This isn't making me feel particularly confident about switching RCU
> CPU stall warnings back to NMIs... ;-)
IMHO, trigger_single_cpu_backtrace() is pretty safe at the moment.
It uses per-CPU buffers a lockless way in NMI context. It even makes
sure that the buffers are flushed to the main log buffer and console
once it is back from NMI.
By other words, the deadlocks in NMI context should be gone. The
NMI buffers are flushed using the classic printk(). Therefore
the risk is the same as when you use printk() directly
in rcu_dump_cpu_stacks() now.
Best Regards,
Petr
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-11-29 18:20 +0100 |
| Message-ID | <sIUeu-2Xs-39@gated-at.bofh.it> |
| In reply to | #1532483 |
On Tue, Nov 29, 2016 at 05:29:20PM +0100, Petr Mladek wrote: > > > People are very busy polishing the turd we call printk, but from where > > > I'm sitting its terminally and unfixably broken. > > I still hope that we could do better :-) How? The console drivers are a complete trainwreck, you simply cannot build anything sensible ontop of a trainwreck. And from what I understood from talking to someone (I again forgot who) at LPC, the whole reason people were poking at this is that the block layer (or something thereabouts) prints a gazillion lines of crap when you attach a stupid amount of devices (through FC or other SAN like things). The way we've 'fixed' that in the scheduler (a fairly long time ago) when SGI complained about our printks taking too long (because they had 4096 CPUs), is to simply remove the printks (they're now hidden behind the sched_debug boot param). In any case, as long as printk has a globally serialized 'log', it, per design, will be worse than the console drivers its build upon. And them being shit precludes the entire stack from being useful. It mostly works, most of the time, and that seems to be what Linus wants, since its really the best we can have given the constraints. But for debugging, when you have a UART, it totally blows.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-29 20:50 +0100 |
| Message-ID | <sIWzD-4hH-11@gated-at.bofh.it> |
| In reply to | #1532558 |
On Tue, Nov 29, 2016 at 06:10:38PM +0100, Peter Zijlstra wrote: > On Tue, Nov 29, 2016 at 05:29:20PM +0100, Petr Mladek wrote: > > > > > People are very busy polishing the turd we call printk, but from where > > > > I'm sitting its terminally and unfixably broken. > > > > I still hope that we could do better :-) > > How? The console drivers are a complete trainwreck, you simply cannot > build anything sensible ontop of a trainwreck. > > And from what I understood from talking to someone (I again forgot who) > at LPC, the whole reason people were poking at this is that the block > layer (or something thereabouts) prints a gazillion lines of crap when > you attach a stupid amount of devices (through FC or other SAN like > things). > > The way we've 'fixed' that in the scheduler (a fairly long time ago) > when SGI complained about our printks taking too long (because they had > 4096 CPUs), is to simply remove the printks (they're now hidden behind > the sched_debug boot param). > > > In any case, as long as printk has a globally serialized 'log', it, per > design, will be worse than the console drivers its build upon. And them > being shit precludes the entire stack from being useful. > > It mostly works, most of the time, and that seems to be what Linus > wants, since its really the best we can have given the constraints. But > for debugging, when you have a UART, it totally blows. UART??? They still make those things??? ;-) Thanx, Paul
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-11-29 21:00 +0100 |
| Message-ID | <sIWJj-4l9-25@gated-at.bofh.it> |
| In reply to | #1532716 |
On Tue, Nov 29, 2016 at 11:39:35AM -0800, Paul E. McKenney wrote: > On Tue, Nov 29, 2016 at 06:10:38PM +0100, Peter Zijlstra wrote: > > It mostly works, most of the time, and that seems to be what Linus > > wants, since its really the best we can have given the constraints. But > > for debugging, when you have a UART, it totally blows. > > UART??? They still make those things??? ;-) Yes, most computer like devices actually have them, trouble is, most consumer devices don't have the pins exposed. Luckily most server class hardware still does. And they're absolutely _awesome_ for debugging; getting data out is a matter of trivial MMIO poll loops. Rock solid stuff.
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-11-29 21:10 +0100 |
| Message-ID | <sIWT0-4En-33@gated-at.bofh.it> |
| In reply to | #1532723 |
On Tue, Nov 29, 2016 at 08:52:04PM +0100, Peter Zijlstra wrote: > On Tue, Nov 29, 2016 at 11:39:35AM -0800, Paul E. McKenney wrote: > > On Tue, Nov 29, 2016 at 06:10:38PM +0100, Peter Zijlstra wrote: > > > > It mostly works, most of the time, and that seems to be what Linus > > > wants, since its really the best we can have given the constraints. But > > > for debugging, when you have a UART, it totally blows. > > > > UART??? They still make those things??? ;-) > > Yes, most computer like devices actually have them, trouble is, most > consumer devices don't have the pins exposed. Luckily most server class > hardware still does. > > And they're absolutely _awesome_ for debugging; getting data out is a > matter of trivial MMIO poll loops. Rock solid stuff. They very clearly need to bring the baud rate into the current millenium, many tens of Mbaud at the -very- least. Thanx, Paul
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web