Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1665048 > unrolled thread

Re: CRASH : RCU detected stall

Started byTheodore Ts'o <tytso@mit.edu>
First post2017-06-13 19:40 +0200
Last post2017-06-14 15:00 +0200
Articles 2 — 1 participant

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: CRASH : RCU detected stall Theodore Ts'o <tytso@mit.edu> - 2017-06-13 19:40 +0200
    Re: CRASH : RCU detected stall Theodore Ts'o <tytso@mit.edu> - 2017-06-14 15:00 +0200

#1665048 — Re: CRASH : RCU detected stall

FromTheodore Ts'o <tytso@mit.edu>
Date2017-06-13 19:40 +0200
SubjectRe: CRASH : RCU detected stall
Message-ID<tRXXj-2qr-1@gated-at.bofh.it>
On Tue, Jun 13, 2017 at 07:35:37PM +0430, Ramin Farajpour Cami wrote:
> Hi,
> 
> I've got the following error report while fuzzing the kernel with syzkaller
> version (4.12-rc5)
> 
> https://groups.google.com/forum/#!topic/syzkaller/4e8MkNnRFRQ

Can you reliably reproduce this failure?  If so, please give details.
Unfortunately this sort of RCU self-stall can be caused by any number
of things.

				- Ted

[toc] | [next] | [standalone]


#1665790

FromTheodore Ts'o <tytso@mit.edu>
Date2017-06-14 15:00 +0200
Message-ID<tSg3T-5hv-1@gated-at.bofh.it>
In reply to#1665048
On Wed, Jun 14, 2017 at 10:02:00AM +0430, Ramin Farajpour Cami wrote:
> 
> Unfortunately it's not reproducible. do you have idea about it?

Nope.  Note that this isn't necessarily an ext4 bug.  We have two
complaints about an rcu_sched thread getting staved and an NMI handler
getting taking too long to run:

rcu_sched kthread starved for 22270 jiffies! g2951 c2950 f0x0
INFO: NMI handler (nmi_cpu_backtrace_handler) took too long to run: 2.762 msecs

We also happened be doing writeback on another CPU and the ext4 thread
was doing a slab allocation.

What ultimiately caused the RCU starvation is not at all clear.  Were
we spinning inside the slab allocator, or not?  Getting some
magic-sysrq triggers to see if the PC was always in the slab allocator
would be useful.  And if that's the case, it's not clear what might
have caused us to spinning in the slab allocator.  It could be due to
some slab state getting corrupted by a previous system call, and ext4
was just unlucky enough to do the slab allocation which caused it to
go for a loop.

We just don't have enough information to do any kind of useful
investigation.

					- Ted

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web