Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1217871 > unrolled thread
| Started by | Frederic Weisbecker <fweisbec@gmail.com> |
|---|---|
| First post | 2015-09-03 00:00 +0200 |
| Last post | 2015-09-05 22:00 +0200 |
| Articles | 6 — 3 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: Warning in irq_work_queue_on() Frederic Weisbecker <fweisbec@gmail.com> - 2015-09-03 00:00 +0200
Re: Warning in irq_work_queue_on() Peter Zijlstra <peterz@infradead.org> - 2015-09-03 00:30 +0200
Re: Warning in irq_work_queue_on() Frederic Weisbecker <fweisbec@gmail.com> - 2015-09-03 02:10 +0200
Re: Warning in irq_work_queue_on() Peter Zijlstra <peterz@infradead.org> - 2015-09-03 10:00 +0200
Re: Warning in irq_work_queue_on() Frederic Weisbecker <fweisbec@gmail.com> - 2015-09-04 17:20 +0200
Re: Warning in irq_work_queue_on() "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2015-09-05 22:00 +0200
| From | Frederic Weisbecker <fweisbec@gmail.com> |
|---|---|
| Date | 2015-09-03 00:00 +0200 |
| Subject | Re: Warning in irq_work_queue_on() |
| Message-ID | <q4oeu-JH-1@gated-at.bofh.it> |
On Wed, Sep 02, 2015 at 03:44:05PM -0400, Tejun Heo wrote: > (cc'ing peterz) > > Ooh, this is from irq_work which doesn't have much to do with > workqueue. Peter? > > On Mon, Aug 24, 2015 at 05:16:11PM -0700, Paul E. McKenney wrote: > > Hello, Tejun, > > > > As discussed last week, I am getting an occasional warning out of > > irq_work_queue_on() WARN_ON_ONCE(cpu_is_offline(cpu)). The repeat-by > > seems to be a week or so of rcutorture runs on 16-CPU KVM instances > > on x86. So please see below on the off-chance that this is of use. > > I have also attached a .config file. > > > > Thoughts? > > > > Thanx, Paul > > > > ------------------------------------------------------------------------ > > > > [ 875.702254] ------------[ cut here ]------------ > > [ 875.703111] WARNING: CPU: 0 PID: 768 at /home/paulmck/public_git/bisect-linux-rcu/kernel/irq_work.c:69 irq_work_queue_on+0xd4/0x110() > > [ 875.703227] Modules linked in: > > [ 875.703227] CPU: 0 PID: 768 Comm: rcu_torture_rea Tainted: G W 4.1.0-rc4+ #1 > > [ 875.703227] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS Bochs 01/01/2011 > > [ 875.703227] ffffffff81baadd8 ffff88001dc5fce8 ffffffff81895418 00000000000000aa > > [ 875.703227] 0000000000000000 ffff88001dc5fd28 ffffffff810517d5 0000000000015bc0 > > [ 875.703227] 0000000000000004 0000000000000004 ffff88001fc8f980 ffff88001fc8d500 > > [ 875.703227] Call Trace: > > [ 875.703227] [<ffffffff81895418>] dump_stack+0x45/0x57 > > [ 875.703227] [<ffffffff810517d5>] warn_slowpath_common+0x85/0xc0 > > [ 875.703227] [<ffffffff810518b5>] warn_slowpath_null+0x15/0x20 > > [ 875.703227] [<ffffffff811119a4>] irq_work_queue_on+0xd4/0x110 > > [ 875.703227] [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50 It happens in nohz full, but I'm not sure the guilty is nohz full. The problem here is that wake_up_nohz_cpu() selects a CPU that is offline. But this shouldn't happen. Either it selects a CPU that is in the domain tree, and I suspect offline CPUs aren't supposed to be there, or it selects the current CPU. And if the CPU is offlined, it shouldn't be running some kthread... > > [ 875.703227] [<ffffffff81076384>] wake_up_nohz_cpu+0xb4/0x100 > > [ 875.703227] [<ffffffff810b1196>] internal_add_timer+0x86/0xa0 > > [ 875.703227] [<ffffffff810b30f1>] mod_timer+0xf1/0x1e0 > > [ 875.703227] [<ffffffff810a63a4>] rcu_torture_reader+0x2a4/0x2e0 > > [ 875.703227] [<ffffffff810a63e0>] ? rcu_torture_reader+0x2e0/0x2e0 > > [ 875.703227] [<ffffffff810a6100>] ? rcutorture_trace_dump.part.10+0x20/0x20 > > [ 875.703227] [<ffffffff8106d75d>] kthread+0xcd/0xf0 > > [ 875.703227] [<ffffffff8106d690>] ? kthread_create_on_node+0x180/0x180 > > [ 875.703227] [<ffffffff8189fb92>] ret_from_fork+0x42/0x70 > > [ 875.703227] [<ffffffff8106d690>] ? kthread_create_on_node+0x180/0x180 > > [ 875.703227] ---[ end trace 74175128740d0113 ]--- -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-03 00:30 +0200 |
| Message-ID | <q4oHw-1ws-19@gated-at.bofh.it> |
| In reply to | #1217871 |
On Wed, Sep 02, 2015 at 11:50:22PM +0200, Frederic Weisbecker wrote: > > > [ 875.703227] [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50 > > It happens in nohz full, but I'm not sure the guilty is nohz full. > > The problem here is that wake_up_nohz_cpu() selects a CPU that is offline. wake_up_nohz_cpu() doesn't do any such thing. Where does the selection logic live? > But this shouldn't happen. Either it selects a CPU that is in the domain tree, > and I suspect offline CPUs aren't supposed to be there, or it selects the current > CPU. And if the CPU is offlined, it shouldn't be running some kthread... Do no assume things like that.. always check with the active mask. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Frederic Weisbecker <fweisbec@gmail.com> |
|---|---|
| Date | 2015-09-03 02:10 +0200 |
| Message-ID | <q4qgh-3OK-5@gated-at.bofh.it> |
| In reply to | #1217879 |
On Thu, Sep 03, 2015 at 12:24:27AM +0200, Peter Zijlstra wrote:
> On Wed, Sep 02, 2015 at 11:50:22PM +0200, Frederic Weisbecker wrote:
> > > > [ 875.703227] [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50
> >
> > It happens in nohz full, but I'm not sure the guilty is nohz full.
> >
> > The problem here is that wake_up_nohz_cpu() selects a CPU that is offline.
>
> wake_up_nohz_cpu() doesn't do any such thing. Where does the selection
> logic live?
Err, got confused with get_nohz_timer_target(). But yeah wake_up_nohz_cpu() is
called with a CPU that is chosen by mod_timer() -> get_nohz_timer_target().
>
> > But this shouldn't happen. Either it selects a CPU that is in the domain tree,
> > and I suspect offline CPUs aren't supposed to be there, or it selects the current
> > CPU. And if the CPU is offlined, it shouldn't be running some kthread...
>
> Do no assume things like that.. always check with the active mask.
Hmm, so perhaps we need something like this (makes me realize that
the is_housekeeping_cpu() passes the wrong argument, no issue in practice
since nohz full aren't in the domain tree but I still need to fix that along).
diff --git a/kernel/sched/core.c b/kernel/sched/core.c
index 0902e4d..2c10a69 100644
--- a/kernel/sched/core.c
+++ b/kernel/sched/core.c
@@ -628,7 +628,7 @@ int get_nohz_timer_target(void)
rcu_read_lock();
for_each_domain(cpu, sd) {
- for_each_cpu(i, sched_domain_span(sd)) {
+ for_each_cpu_and(i, sched_domain_span(sd), cpu_online_mask) {
if (!idle_cpu(i) && is_housekeeping_cpu(cpu)) {
cpu = i;
goto unlock;
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-03 10:00 +0200 |
| Message-ID | <q4xB8-5Ju-19@gated-at.bofh.it> |
| In reply to | #1217954 |
On Thu, Sep 03, 2015 at 02:03:51AM +0200, Frederic Weisbecker wrote:
> On Thu, Sep 03, 2015 at 12:24:27AM +0200, Peter Zijlstra wrote:
> > On Wed, Sep 02, 2015 at 11:50:22PM +0200, Frederic Weisbecker wrote:
> > > > > [ 875.703227] [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50
> > >
> > > It happens in nohz full, but I'm not sure the guilty is nohz full.
> > >
> > > The problem here is that wake_up_nohz_cpu() selects a CPU that is offline.
> >
> > wake_up_nohz_cpu() doesn't do any such thing. Where does the selection
> > logic live?
>
> Err, got confused with get_nohz_timer_target(). But yeah wake_up_nohz_cpu() is
> called with a CPU that is chosen by mod_timer() -> get_nohz_timer_target().
>
> >
> > > But this shouldn't happen. Either it selects a CPU that is in the domain tree,
> > > and I suspect offline CPUs aren't supposed to be there, or it selects the current
> > > CPU. And if the CPU is offlined, it shouldn't be running some kthread...
> >
> > Do no assume things like that.. always check with the active mask.
>
> Hmm, so perhaps we need something like this (makes me realize that
> the is_housekeeping_cpu() passes the wrong argument, no issue in practice
> since nohz full aren't in the domain tree but I still need to fix that along).
>
> diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> index 0902e4d..2c10a69 100644
> --- a/kernel/sched/core.c
> +++ b/kernel/sched/core.c
> @@ -628,7 +628,7 @@ int get_nohz_timer_target(void)
>
> rcu_read_lock();
> for_each_domain(cpu, sd) {
> - for_each_cpu(i, sched_domain_span(sd)) {
> + for_each_cpu_and(i, sched_domain_span(sd), cpu_online_mask) {
cpu_active_mask, we clear that when we start killing the cpu. online
only gets cleared once the cpu is actually dead.
> if (!idle_cpu(i) && is_housekeeping_cpu(cpu)) {
> cpu = i;
> goto unlock;
>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Frederic Weisbecker <fweisbec@gmail.com> |
|---|---|
| Date | 2015-09-04 17:20 +0200 |
| Message-ID | <q50Wu-5QK-19@gated-at.bofh.it> |
| In reply to | #1218086 |
On Thu, Sep 03, 2015 at 09:58:40AM +0200, Peter Zijlstra wrote:
> On Thu, Sep 03, 2015 at 02:03:51AM +0200, Frederic Weisbecker wrote:
> > On Thu, Sep 03, 2015 at 12:24:27AM +0200, Peter Zijlstra wrote:
> > > On Wed, Sep 02, 2015 at 11:50:22PM +0200, Frederic Weisbecker wrote:
> > > > > > [ 875.703227] [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50
> > > >
> > > > It happens in nohz full, but I'm not sure the guilty is nohz full.
> > > >
> > > > The problem here is that wake_up_nohz_cpu() selects a CPU that is offline.
> > >
> > > wake_up_nohz_cpu() doesn't do any such thing. Where does the selection
> > > logic live?
> >
> > Err, got confused with get_nohz_timer_target(). But yeah wake_up_nohz_cpu() is
> > called with a CPU that is chosen by mod_timer() -> get_nohz_timer_target().
> >
> > >
> > > > But this shouldn't happen. Either it selects a CPU that is in the domain tree,
> > > > and I suspect offline CPUs aren't supposed to be there, or it selects the current
> > > > CPU. And if the CPU is offlined, it shouldn't be running some kthread...
> > >
> > > Do no assume things like that.. always check with the active mask.
> >
> > Hmm, so perhaps we need something like this (makes me realize that
> > the is_housekeeping_cpu() passes the wrong argument, no issue in practice
> > since nohz full aren't in the domain tree but I still need to fix that along).
> >
> > diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> > index 0902e4d..2c10a69 100644
> > --- a/kernel/sched/core.c
> > +++ b/kernel/sched/core.c
> > @@ -628,7 +628,7 @@ int get_nohz_timer_target(void)
> >
> > rcu_read_lock();
> > for_each_domain(cpu, sd) {
> > - for_each_cpu(i, sched_domain_span(sd)) {
> > + for_each_cpu_and(i, sched_domain_span(sd), cpu_online_mask) {
>
> cpu_active_mask, we clear that when we start killing the cpu. online
> only gets cleared once the cpu is actually dead.
So, after our discussion in IRC, I checked how domains are rebuild on hotplug
ops and it appears that partition_sched_domain() is called on CPU_DOWN_PREPARE
only. The CPU shouldn't be on the domain tree after that.
(Correct me if I'm wrong, I really am not an expert in the domain handling code.
As you said that we can't guarantee that a CPU in the domain tree is in the cpu_online_mask,
I'm likely wrong somewhere).
This is then followed by synchronize_sched(). Which means that after that, the
new version of the CPU domains (with the offlining CPU excluded) is visible
everywhere while the CPU is still in cpu_online_mask.
And finally stop machine runs and the CPU is cleared out of cpu_online_mask.
So I'm probably missing something, otherwise we could find a CPU in the domain
tree that is not in cpu_online_mask.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2015-09-05 22:00 +0200 |
| Message-ID | <q5rN0-1OZ-7@gated-at.bofh.it> |
| In reply to | #1219110 |
On Fri, Sep 04, 2015 at 05:11:54PM +0200, Frederic Weisbecker wrote:
> On Thu, Sep 03, 2015 at 09:58:40AM +0200, Peter Zijlstra wrote:
> > On Thu, Sep 03, 2015 at 02:03:51AM +0200, Frederic Weisbecker wrote:
> > > On Thu, Sep 03, 2015 at 12:24:27AM +0200, Peter Zijlstra wrote:
> > > > On Wed, Sep 02, 2015 at 11:50:22PM +0200, Frederic Weisbecker wrote:
> > > > > > > [ 875.703227] [<ffffffff810c2d74>] tick_nohz_full_kick_cpu+0x44/0x50
> > > > >
> > > > > It happens in nohz full, but I'm not sure the guilty is nohz full.
> > > > >
> > > > > The problem here is that wake_up_nohz_cpu() selects a CPU that is offline.
> > > >
> > > > wake_up_nohz_cpu() doesn't do any such thing. Where does the selection
> > > > logic live?
> > >
> > > Err, got confused with get_nohz_timer_target(). But yeah wake_up_nohz_cpu() is
> > > called with a CPU that is chosen by mod_timer() -> get_nohz_timer_target().
> > >
> > > >
> > > > > But this shouldn't happen. Either it selects a CPU that is in the domain tree,
> > > > > and I suspect offline CPUs aren't supposed to be there, or it selects the current
> > > > > CPU. And if the CPU is offlined, it shouldn't be running some kthread...
> > > >
> > > > Do no assume things like that.. always check with the active mask.
> > >
> > > Hmm, so perhaps we need something like this (makes me realize that
> > > the is_housekeeping_cpu() passes the wrong argument, no issue in practice
> > > since nohz full aren't in the domain tree but I still need to fix that along).
> > >
> > > diff --git a/kernel/sched/core.c b/kernel/sched/core.c
> > > index 0902e4d..2c10a69 100644
> > > --- a/kernel/sched/core.c
> > > +++ b/kernel/sched/core.c
> > > @@ -628,7 +628,7 @@ int get_nohz_timer_target(void)
> > >
> > > rcu_read_lock();
> > > for_each_domain(cpu, sd) {
> > > - for_each_cpu(i, sched_domain_span(sd)) {
> > > + for_each_cpu_and(i, sched_domain_span(sd), cpu_online_mask) {
> >
> > cpu_active_mask, we clear that when we start killing the cpu. online
> > only gets cleared once the cpu is actually dead.
>
> So, after our discussion in IRC, I checked how domains are rebuild on hotplug
> ops and it appears that partition_sched_domain() is called on CPU_DOWN_PREPARE
> only. The CPU shouldn't be on the domain tree after that.
>
> (Correct me if I'm wrong, I really am not an expert in the domain handling code.
> As you said that we can't guarantee that a CPU in the domain tree is in the cpu_online_mask,
> I'm likely wrong somewhere).
>
> This is then followed by synchronize_sched(). Which means that after that, the
> new version of the CPU domains (with the offlining CPU excluded) is visible
> everywhere while the CPU is still in cpu_online_mask.
>
> And finally stop machine runs and the CPU is cleared out of cpu_online_mask.
> So I'm probably missing something, otherwise we could find a CPU in the domain
> tree that is not in cpu_online_mask.
OK, I have to ask... Should I be trying Frederic's patch?
At the current failure rate, I will need to be running it for about
a year to give any reasonable conclusion. :-/
Thanx, Paul
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web