Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1633810 > unrolled thread

Re: WARN_ON_ONCE() in process_one_work()?

Started byTejun Heo <tj@kernel.org>
First post2017-05-01 20:50 +0200
Last post2017-05-05 19:20 +0200
Articles 3 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: WARN_ON_ONCE() in process_one_work()? Tejun Heo <tj@kernel.org> - 2017-05-01 20:50 +0200
    Re: WARN_ON_ONCE() in process_one_work()? "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-05-01 21:00 +0200
      Re: WARN_ON_ONCE() in process_one_work()? "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-05-05 19:20 +0200

#1633810 — Re: WARN_ON_ONCE() in process_one_work()?

FromTejun Heo <tj@kernel.org>
Date2017-05-01 20:50 +0200
SubjectRe: WARN_ON_ONCE() in process_one_work()?
Message-ID<tCoyu-4dM-17@gated-at.bofh.it>
Hello, Paul.

On Mon, May 01, 2017 at 11:38:07AM -0700, Paul E. McKenney wrote:
> On Mon, May 01, 2017 at 09:57:47AM -0700, Paul E. McKenney wrote:
> > Hello!
> > 
> > I am hitting this WARN_ON_ONCE() in process_one_work() and am wondering
> > what I did wrong to make this happen:
> 
> Oh, wait...  Rescuer, it says.  Might this be due to the fact that RCU's
> expedited grace periods block within a workqueue handler?  Might this
> in turn run the system out of workqueue kthreads?  If this is the likely
> cause, my approach would be to rework the expected-grace-period workqueue
> handler to return when waiting for the grace period to complete, and to
> replace the current wakeup with a schedule_work() or something similar.

That should be completely fine.  It could just be that the rescuer
path has a bug around CPU hotplug handling.  Can you please confirm
either way on the cpuset usage?

Thanks.

-- 
tejun

[toc] | [next] | [standalone]


#1633813

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-05-01 21:00 +0200
Message-ID<tCoI9-4h9-7@gated-at.bofh.it>
In reply to#1633810
On Mon, May 01, 2017 at 02:44:02PM -0400, Tejun Heo wrote:
> Hello, Paul.
> 
> On Mon, May 01, 2017 at 11:38:07AM -0700, Paul E. McKenney wrote:
> > On Mon, May 01, 2017 at 09:57:47AM -0700, Paul E. McKenney wrote:
> > > Hello!
> > > 
> > > I am hitting this WARN_ON_ONCE() in process_one_work() and am wondering
> > > what I did wrong to make this happen:
> > 
> > Oh, wait...  Rescuer, it says.  Might this be due to the fact that RCU's
> > expedited grace periods block within a workqueue handler?  Might this
> > in turn run the system out of workqueue kthreads?  If this is the likely
> > cause, my approach would be to rework the expected-grace-period workqueue
> > handler to return when waiting for the grace period to complete, and to
> > replace the current wakeup with a schedule_work() or something similar.
> 
> That should be completely fine.  It could just be that the rescuer
> path has a bug around CPU hotplug handling.  Can you please confirm
> either way on the cpuset usage?

I have no explicit cpuset usage or affinity of the workqueue handlers
themselves.

However, this is thus far only happening in CONFIG_NO_HZ_FULL=y runs, in
this case, with the kernel boot parameter nohz_full=2-9 out of 16 CPUs.
IIRC, this sets up a "housekeeping" cpuset that pushes normal tasks away
from the nohz_full CPUs.

I do build with CONFIG_HOTPLUG_CPU=y, and the test does a lot of
hotplugging.  Also, other kthreads (but again, not the workqueue handlers)
do a lot of explicit CPU-affinity manipulation.

								Thanx, Paul

[toc] | [prev] | [next] | [standalone]


#1636505

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-05-05 19:20 +0200
Message-ID<tDP3A-4H2-3@gated-at.bofh.it>
In reply to#1633813
On Mon, May 01, 2017 at 11:58:19AM -0700, Paul E. McKenney wrote:
> On Mon, May 01, 2017 at 02:44:02PM -0400, Tejun Heo wrote:
> > Hello, Paul.
> > 
> > On Mon, May 01, 2017 at 11:38:07AM -0700, Paul E. McKenney wrote:
> > > On Mon, May 01, 2017 at 09:57:47AM -0700, Paul E. McKenney wrote:
> > > > Hello!
> > > > 
> > > > I am hitting this WARN_ON_ONCE() in process_one_work() and am wondering
> > > > what I did wrong to make this happen:
> > > 
> > > Oh, wait...  Rescuer, it says.  Might this be due to the fact that RCU's
> > > expedited grace periods block within a workqueue handler?  Might this
> > > in turn run the system out of workqueue kthreads?  If this is the likely
> > > cause, my approach would be to rework the expected-grace-period workqueue
> > > handler to return when waiting for the grace period to complete, and to
> > > replace the current wakeup with a schedule_work() or something similar.
> > 
> > That should be completely fine.  It could just be that the rescuer
> > path has a bug around CPU hotplug handling.  Can you please confirm
> > either way on the cpuset usage?
> 
> I have no explicit cpuset usage or affinity of the workqueue handlers
> themselves.
> 
> However, this is thus far only happening in CONFIG_NO_HZ_FULL=y runs, in
> this case, with the kernel boot parameter nohz_full=2-9 out of 16 CPUs.
> IIRC, this sets up a "housekeeping" cpuset that pushes normal tasks away
> from the nohz_full CPUs.
> 
> I do build with CONFIG_HOTPLUG_CPU=y, and the test does a lot of
> hotplugging.  Also, other kthreads (but again, not the workqueue handlers)
> do a lot of explicit CPU-affinity manipulation.

Just following up...  I have hit this bug a couple of times over the
past few days.  Anything I can do to help?

							Thanx, Paul

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web