Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1217871 > unrolled thread

Re: Warning in irq_work_queue_on()

Started byFrederic Weisbecker <fweisbec@gmail.com>
First post2015-09-03 00:00 +0200
Last post2015-09-05 22:00 +0200
Articles 6 — 3 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: Warning in irq_work_queue_on() Frederic Weisbecker <fweisbec@gmail.com> - 2015-09-03 00:00 +0200
    Re: Warning in irq_work_queue_on() Peter Zijlstra <peterz@infradead.org> - 2015-09-03 00:30 +0200
      Re: Warning in irq_work_queue_on() Frederic Weisbecker <fweisbec@gmail.com> - 2015-09-03 02:10 +0200
        Re: Warning in irq_work_queue_on() Peter Zijlstra <peterz@infradead.org> - 2015-09-03 10:00 +0200
          Re: Warning in irq_work_queue_on() Frederic Weisbecker <fweisbec@gmail.com> - 2015-09-04 17:20 +0200
            Re: Warning in irq_work_queue_on() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2015-09-05 22:00 +0200

#1217871 — Re: Warning in irq_work_queue_on()

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2015-09-03 00:00 +0200
SubjectRe: Warning in irq_work_queue_on()
Message-ID<q4oeu-JH-1@gated-at.bofh.it>
On Wed, Sep 02, 2015 at 03:44:05PM -0400, Tejun Heo wrote:
> (cc'ing peterz)
> 
> Ooh, this is from irq_work which doesn't have much to do with
> workqueue.  Peter?
> 
> On Mon, Aug 24, 2015 at 05:16:11PM -0700, Paul E. McKenney wrote:
> > Hello, Tejun,
> > 
> > As discussed last week, I am getting an occasional warning out of
> > irq_work_queue_on() WARN_ON_ONCE(cpu_is_offline(cpu)).  The repeat-by
> > seems to be a week or so of rcutorture runs on 16-CPU KVM instances
> > on x86.  So please see below on the off-chance that this is of use.
> > I have also attached a .config file.
> > 
> > Thoughts?
> > 
> > 							Thanx, Paul
> > 
> > ------------------------------------------------------------------------
> > 
> > [  875.702254] ------------[ cut here ]------------
> > [  875.703111] WARNING: CPU: 0 PID: 768 at /home/paulmck/public_git/bisect-linux-rcu/kernel/irq_work.c:69 irq_work_queue_on+0xd4/0x110()
> > [  875.703227] Modules linked in:
> > [  875.703227] CPU: 0 PID: 768 Comm: rcu_torture_rea Tainted: G        W       4.1.0-rc4+ #1
> > [  875.703227] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS Bochs 01/01/2011
> > [  875.703227]  ffffffff81baadd8 ffff88001dc5fce8 ffffffff81895418 00000000000000aa
> > [  875.703227]  0000000000000000 ffff88001dc5fd28 ffffffff810517d5 0000000000015bc0
> > [  875.703227]  0000000000000004 0000000000000004 ffff88001fc8f980 ffff88001fc8d500
> > [  875.703227] Call Trace:
> > [  875.703227]  [<ffffffff81895418>] dump_stack+0x45/0x57
> > [  875.703227]  [<ffffffff810517d5>] warn_slowpath_common+0x85/0xc0
> > [  875.703227]  [<ffffffff810518b5>] warn_slowpath_null+0x15/0x20
> > [  875.703227]  [<ffffffff811119a4>] irq_work_queue_on+0xd4/0x110
> > [  875.703227]  [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50

It happens in nohz full, but I'm not sure the guilty is nohz full.

The problem here is that wake_up_nohz_cpu() selects a CPU that is offline.
But this shouldn't happen. Either it selects a CPU that is in the domain tree,
and I suspect offline CPUs aren't supposed to be there, or it selects the current
CPU. And if the CPU is offlined, it shouldn't be running some kthread...

> > [  875.703227]  [<ffffffff81076384>] wake_up_nohz_cpu+0xb4/0x100
> > [  875.703227]  [<ffffffff810b1196>] internal_add_timer+0x86/0xa0
> > [  875.703227]  [<ffffffff810b30f1>] mod_timer+0xf1/0x1e0
> > [  875.703227]  [<ffffffff810a63a4>] rcu_torture_reader+0x2a4/0x2e0
> > [  875.703227]  [<ffffffff810a63e0>] ? rcu_torture_reader+0x2e0/0x2e0
> > [  875.703227]  [<ffffffff810a6100>] ? rcutorture_trace_dump.part.10+0x20/0x20
> > [  875.703227]  [<ffffffff8106d75d>] kthread+0xcd/0xf0
> > [  875.703227]  [<ffffffff8106d690>] ? kthread_create_on_node+0x180/0x180
> > [  875.703227]  [<ffffffff8189fb92>] ret_from_fork+0x42/0x70
> > [  875.703227]  [<ffffffff8106d690>] ? kthread_create_on_node+0x180/0x180
> > [  875.703227] ---[ end trace 74175128740d0113 ]---
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1217879

FromPeter Zijlstra <peterz@infradead.org>
Date2015-09-03 00:30 +0200
Message-ID<q4oHw-1ws-19@gated-at.bofh.it>
In reply to#1217871
On Wed, Sep 02, 2015 at 11:50:22PM +0200, Frederic Weisbecker wrote:
> > > [  875.703227]  [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50
> 
> It happens in nohz full, but I'm not sure the guilty is nohz full.
> 
> The problem here is that wake_up_nohz_cpu() selects a CPU that is offline.

wake_up_nohz_cpu() doesn't do any such thing. Where does the selection
logic live?

> But this shouldn't happen. Either it selects a CPU that is in the domain tree,
> and I suspect offline CPUs aren't supposed to be there, or it selects the current
> CPU. And if the CPU is offlined, it shouldn't be running some kthread...

Do no assume things like that.. always check with the active mask.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1217954

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2015-09-03 02:10 +0200
Message-ID<q4qgh-3OK-5@gated-at.bofh.it>
In reply to#1217879
On Thu, Sep 03, 2015 at 12:24:27AM +0200, Peter Zijlstra wrote:
> On Wed, Sep 02, 2015 at 11:50:22PM +0200, Frederic Weisbecker wrote:
> > > > [  875.703227]  [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50
> > 
> > It happens in nohz full, but I'm not sure the guilty is nohz full.
> > 
> > The problem here is that wake_up_nohz_cpu() selects a CPU that is offline.
> 
> wake_up_nohz_cpu() doesn't do any such thing. Where does the selection
> logic live?

Err, got confused with get_nohz_timer_target(). But yeah wake_up_nohz_cpu() is
called with a CPU that is chosen by mod_timer() -> get_nohz_timer_target().

> 
> > But this shouldn't happen. Either it selects a CPU that is in the domain tree,
> > and I suspect offline CPUs aren't supposed to be there, or it selects the current
> > CPU. And if the CPU is offlined, it shouldn't be running some kthread...
> 
> Do no assume things like that.. always check with the active mask.

Hmm, so perhaps we need something like this (makes me realize that
the is_housekeeping_cpu() passes the wrong argument, no issue in practice
since nohz full aren't in the domain tree but I still need to fix that along).

diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 0902e4d..2c10a69 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -628,7 +628,7 @@ int get_nohz_timer_target(void)
 
 	rcu_read_lock();
 	for_each_domain(cpu, sd) {
-		for_each_cpu(i, sched_domain_span(sd)) {
+		for_each_cpu_and(i, sched_domain_span(sd), cpu_online_mask) {
 			if (!idle_cpu(i) && is_housekeeping_cpu(cpu)) {
 				cpu = i;
 				goto unlock;

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1218086

FromPeter Zijlstra <peterz@infradead.org>
Date2015-09-03 10:00 +0200
Message-ID<q4xB8-5Ju-19@gated-at.bofh.it>
In reply to#1217954
On Thu, Sep 03, 2015 at 02:03:51AM +0200, Frederic Weisbecker wrote:
> On Thu, Sep 03, 2015 at 12:24:27AM +0200, Peter Zijlstra wrote:
> > On Wed, Sep 02, 2015 at 11:50:22PM +0200, Frederic Weisbecker wrote:
> > > > > [  875.703227]  [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50
> > > 
> > > It happens in nohz full, but I'm not sure the guilty is nohz full.
> > > 
> > > The problem here is that wake_up_nohz_cpu() selects a CPU that is offline.
> > 
> > wake_up_nohz_cpu() doesn't do any such thing. Where does the selection
> > logic live?
> 
> Err, got confused with get_nohz_timer_target(). But yeah wake_up_nohz_cpu() is
> called with a CPU that is chosen by mod_timer() -> get_nohz_timer_target().
> 
> > 
> > > But this shouldn't happen. Either it selects a CPU that is in the domain tree,
> > > and I suspect offline CPUs aren't supposed to be there, or it selects the current
> > > CPU. And if the CPU is offlined, it shouldn't be running some kthread...
> > 
> > Do no assume things like that.. always check with the active mask.
> 
> Hmm, so perhaps we need something like this (makes me realize that
> the is_housekeeping_cpu() passes the wrong argument, no issue in practice
> since nohz full aren't in the domain tree but I still need to fix that along).
> 
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index 0902e4d..2c10a69 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -628,7 +628,7 @@ int get_nohz_timer_target(void)
>  
>  	rcu_read_lock();
>  	for_each_domain(cpu, sd) {
> -		for_each_cpu(i, sched_domain_span(sd)) {
> +		for_each_cpu_and(i, sched_domain_span(sd), cpu_online_mask) {

cpu_active_mask, we clear that when we start killing the cpu. online
only gets cleared once the cpu is actually dead.

>  			if (!idle_cpu(i) && is_housekeeping_cpu(cpu)) {
>  				cpu = i;
>  				goto unlock;
> 
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1219110

FromFrederic Weisbecker <fweisbec@gmail.com>
Date2015-09-04 17:20 +0200
Message-ID<q50Wu-5QK-19@gated-at.bofh.it>
In reply to#1218086
On Thu, Sep 03, 2015 at 09:58:40AM +0200, Peter Zijlstra wrote:
> On Thu, Sep 03, 2015 at 02:03:51AM +0200, Frederic Weisbecker wrote:
> > On Thu, Sep 03, 2015 at 12:24:27AM +0200, Peter Zijlstra wrote:
> > > On Wed, Sep 02, 2015 at 11:50:22PM +0200, Frederic Weisbecker wrote:
> > > > > > [  875.703227]  [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50
> > > > 
> > > > It happens in nohz full, but I'm not sure the guilty is nohz full.
> > > > 
> > > > The problem here is that wake_up_nohz_cpu() selects a CPU that is offline.
> > > 
> > > wake_up_nohz_cpu() doesn't do any such thing. Where does the selection
> > > logic live?
> > 
> > Err, got confused with get_nohz_timer_target(). But yeah wake_up_nohz_cpu() is
> > called with a CPU that is chosen by mod_timer() -> get_nohz_timer_target().
> > 
> > > 
> > > > But this shouldn't happen. Either it selects a CPU that is in the domain tree,
> > > > and I suspect offline CPUs aren't supposed to be there, or it selects the current
> > > > CPU. And if the CPU is offlined, it shouldn't be running some kthread...
> > > 
> > > Do no assume things like that.. always check with the active mask.
> > 
> > Hmm, so perhaps we need something like this (makes me realize that
> > the is_housekeeping_cpu() passes the wrong argument, no issue in practice
> > since nohz full aren't in the domain tree but I still need to fix that along).
> > 
> > diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> > index 0902e4d..2c10a69 100644
> > --- a/kernel/sched/core.c
> > +++ b/kernel/sched/core.c
> > @@ -628,7 +628,7 @@ int get_nohz_timer_target(void)
> >  
> >  	rcu_read_lock();
> >  	for_each_domain(cpu, sd) {
> > -		for_each_cpu(i, sched_domain_span(sd)) {
> > +		for_each_cpu_and(i, sched_domain_span(sd), cpu_online_mask) {
> 
> cpu_active_mask, we clear that when we start killing the cpu. online
> only gets cleared once the cpu is actually dead.

So, after our discussion in IRC, I checked how domains are rebuild on hotplug
ops and it appears that partition_sched_domain() is called on CPU_DOWN_PREPARE
only. The CPU shouldn't be on the domain tree after that.

(Correct me if I'm wrong, I really am not an expert in the domain handling code.
As you said that we can't guarantee that a CPU in the domain tree is in the cpu_online_mask,
I'm likely wrong somewhere).

This is then followed by synchronize_sched(). Which means that after that, the
new version of the CPU domains (with the offlining CPU excluded) is visible
everywhere while the CPU is still in cpu_online_mask.

And finally stop machine runs and the CPU is cleared out of cpu_online_mask.
So I'm probably missing something, otherwise we could find a CPU in the domain
tree that is not in cpu_online_mask.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1219630

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2015-09-05 22:00 +0200
Message-ID<q5rN0-1OZ-7@gated-at.bofh.it>
In reply to#1219110
On Fri, Sep 04, 2015 at 05:11:54PM +0200, Frederic Weisbecker wrote:
> On Thu, Sep 03, 2015 at 09:58:40AM +0200, Peter Zijlstra wrote:
> > On Thu, Sep 03, 2015 at 02:03:51AM +0200, Frederic Weisbecker wrote:
> > > On Thu, Sep 03, 2015 at 12:24:27AM +0200, Peter Zijlstra wrote:
> > > > On Wed, Sep 02, 2015 at 11:50:22PM +0200, Frederic Weisbecker wrote:
> > > > > > > [  875.703227]  [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50
> > > > > 
> > > > > It happens in nohz full, but I'm not sure the guilty is nohz full.
> > > > > 
> > > > > The problem here is that wake_up_nohz_cpu() selects a CPU that is offline.
> > > > 
> > > > wake_up_nohz_cpu() doesn't do any such thing. Where does the selection
> > > > logic live?
> > > 
> > > Err, got confused with get_nohz_timer_target(). But yeah wake_up_nohz_cpu() is
> > > called with a CPU that is chosen by mod_timer() -> get_nohz_timer_target().
> > > 
> > > > 
> > > > > But this shouldn't happen. Either it selects a CPU that is in the domain tree,
> > > > > and I suspect offline CPUs aren't supposed to be there, or it selects the current
> > > > > CPU. And if the CPU is offlined, it shouldn't be running some kthread...
> > > > 
> > > > Do no assume things like that.. always check with the active mask.
> > > 
> > > Hmm, so perhaps we need something like this (makes me realize that
> > > the is_housekeeping_cpu() passes the wrong argument, no issue in practice
> > > since nohz full aren't in the domain tree but I still need to fix that along).
> > > 
> > > diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> > > index 0902e4d..2c10a69 100644
> > > --- a/kernel/sched/core.c
> > > +++ b/kernel/sched/core.c
> > > @@ -628,7 +628,7 @@ int get_nohz_timer_target(void)
> > >  
> > >  	rcu_read_lock();
> > >  	for_each_domain(cpu, sd) {
> > > -		for_each_cpu(i, sched_domain_span(sd)) {
> > > +		for_each_cpu_and(i, sched_domain_span(sd), cpu_online_mask) {
> > 
> > cpu_active_mask, we clear that when we start killing the cpu. online
> > only gets cleared once the cpu is actually dead.
> 
> So, after our discussion in IRC, I checked how domains are rebuild on hotplug
> ops and it appears that partition_sched_domain() is called on CPU_DOWN_PREPARE
> only. The CPU shouldn't be on the domain tree after that.
> 
> (Correct me if I'm wrong, I really am not an expert in the domain handling code.
> As you said that we can't guarantee that a CPU in the domain tree is in the cpu_online_mask,
> I'm likely wrong somewhere).
> 
> This is then followed by synchronize_sched(). Which means that after that, the
> new version of the CPU domains (with the offlining CPU excluded) is visible
> everywhere while the CPU is still in cpu_online_mask.
> 
> And finally stop machine runs and the CPU is cleared out of cpu_online_mask.
> So I'm probably missing something, otherwise we could find a CPU in the domain
> tree that is not in cpu_online_mask.

OK, I have to ask...  Should I be trying Frederic's patch?

At the current failure rate, I will need to be running it for about
a year to give any reasonable conclusion.  :-/

							Thanx, Paul

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web