Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1394822 > unrolled thread

[PATCH] workqueue: fix rebind bound workers warning

Started byWanpeng Li <kernellwp@gmail.com>
First post2016-05-05 03:50 +0200
Last post2016-05-11 12:30 +0200
Articles 10 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-05 03:50 +0200
    Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-09 09:30 +0200
    Re: [PATCH] workqueue: fix rebind bound workers warning Tejun Heo <tj@kernel.org> - 2016-05-09 19:10 +0200
      Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-10 00:00 +0200
        Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-10 00:20 +0200
        Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-11 01:30 +0200
          Re: [PATCH] workqueue: fix rebind bound workers warning Thomas Gleixner <tglx@linutronix.de> - 2016-05-11 09:40 +0200
            Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-11 10:10 +0200
              Re: [PATCH] workqueue: fix rebind bound workers warning Thomas Gleixner <tglx@linutronix.de> - 2016-05-11 12:10 +0200
                Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-11 12:30 +0200

#1394822 — [PATCH] workqueue: fix rebind bound workers warning

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-05 03:50 +0200
Subject[PATCH] workqueue: fix rebind bound workers warning
Message-ID<rvgAr-6ZQ-5@gated-at.bofh.it>
From: Wanpeng Li <wanpeng.li@hotmail.com>

------------[ cut here ]------------
WARNING: CPU: 0 PID: 16 at kernel/workqueue.c:4559 rebind_workers+0x1c0/0x1d0
Modules linked in:
CPU: 0 PID: 16 Comm: cpuhp/0 Not tainted 4.6.0-rc4+ #31
Hardware name: IBM IBM System x3550 M4 Server -[7914IUW]-/00Y8603, BIOS -[D7E128FUS-1.40]- 07/23/2013
 0000000000000000 ffff881037babb58 ffffffff8139d885 0000000000000010
 0000000000000000 0000000000000000 0000000000000000 ffff881037babba8
 ffffffff8108505d ffff881037ba0000 000011cf3e7d6e60 0000000000000046
Call Trace:
 dump_stack+0x89/0xd4
 __warn+0xfd/0x120
 warn_slowpath_null+0x1d/0x20
 rebind_workers+0x1c0/0x1d0
 workqueue_cpu_up_callback+0xf5/0x1d0
 notifier_call_chain+0x64/0x90
 ? trace_hardirqs_on_caller+0xf2/0x220
 ? notify_prepare+0x80/0x80
 __raw_notifier_call_chain+0xe/0x10
 __cpu_notify+0x35/0x50
 notify_down_prepare+0x5e/0x80
 ? notify_prepare+0x80/0x80
 cpuhp_invoke_callback+0x73/0x330
 ? __schedule+0x33e/0x8a0
 cpuhp_down_callbacks+0x51/0xc0
 cpuhp_thread_fun+0xc1/0xf0
 smpboot_thread_fn+0x159/0x2a0
 ? smpboot_create_threads+0x80/0x80
 kthread+0xef/0x110
 ? wait_for_completion+0xf0/0x120
 ? schedule_tail+0x35/0xf0
 ret_from_fork+0x22/0x50
 ? __init_kthread_worker+0x70/0x70
---[ end trace eb12ae47d2382d8f ]---
notify_down_prepare: attempt to take down CPU 0 failed

This bug can be reproduced by below config w/ nohz_full= all cpus:

CONFIG_BOOTPARAM_HOTPLUG_CPU0=y
CONFIG_DEBUG_HOTPLUG_CPU0=y
CONFIG_NO_HZ_FULL=y

The boot CPU handles housekeeping duty(unbound timers, workqueues, 
timekeeping, ...) on behalf of full dynticks CPUs. It must remain 
online when nohz full is enabled. There is a priority set to every
notifier_blocks:

workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down

So tick_nohz_cpu_down callback failed when down prepare cpu 0, and 
notifier_blocks behind tick_nohz_cpu_down will not be called any 
more, which leads to workers are actually not unbound. Then hotplug 
state machine will fallback to undo and online cpu 0 again. Workers 
will be rebound unconditionally even if they are not unbound and 
trigger the warning in this progress.

This patch fix it by catching !DISASSOCIATED to avoid rebind bound 
workers.

Cc: Tejun Heo <tj@kernel.org>
Cc: Lai Jiangshan <jiangshanlai@gmail.com>
Suggested-by: Lai Jiangshan <jiangshanlai@gmail.com>
Signed-off-by: Wanpeng Li <wanpeng.li@hotmail.com>
---
 kernel/workqueue.c | 6 ++++++
 1 file changed, 6 insertions(+)

diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index 2232ae3..cc18920 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -4525,6 +4525,12 @@ static void rebind_workers(struct worker_pool *pool)
 						  pool->attrs->cpumask) < 0);
 
 	spin_lock_irq(&pool->lock);
+
+	if (!(pool->flags & POOL_DISASSOCIATED)) {
+		spin_unlock_irq(&pool->lock);
+		return;
+	}
+
 	pool->flags &= ~POOL_DISASSOCIATED;
 
 	for_each_pool_worker(worker, pool) {
-- 
1.9.1

[toc] | [next] | [standalone]


#1396704

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-09 09:30 +0200
Message-ID<rwNNH-7J-79@gated-at.bofh.it>
In reply to#1394822
Sorry to quick ping you Tejun, just hope it can catch the upcoming
merge window. :-)
2016-05-05 9:41 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
> From: Wanpeng Li <wanpeng.li@hotmail.com>
>
> ------------[ cut here ]------------
> WARNING: CPU: 0 PID: 16 at kernel/workqueue.c:4559 rebind_workers+0x1c0/0x1d0
> Modules linked in:
> CPU: 0 PID: 16 Comm: cpuhp/0 Not tainted 4.6.0-rc4+ #31
> Hardware name: IBM IBM System x3550 M4 Server -[7914IUW]-/00Y8603, BIOS -[D7E128FUS-1.40]- 07/23/2013
>  0000000000000000 ffff881037babb58 ffffffff8139d885 0000000000000010
>  0000000000000000 0000000000000000 0000000000000000 ffff881037babba8
>  ffffffff8108505d ffff881037ba0000 000011cf3e7d6e60 0000000000000046
> Call Trace:
>  dump_stack+0x89/0xd4
>  __warn+0xfd/0x120
>  warn_slowpath_null+0x1d/0x20
>  rebind_workers+0x1c0/0x1d0
>  workqueue_cpu_up_callback+0xf5/0x1d0
>  notifier_call_chain+0x64/0x90
>  ? trace_hardirqs_on_caller+0xf2/0x220
>  ? notify_prepare+0x80/0x80
>  __raw_notifier_call_chain+0xe/0x10
>  __cpu_notify+0x35/0x50
>  notify_down_prepare+0x5e/0x80
>  ? notify_prepare+0x80/0x80
>  cpuhp_invoke_callback+0x73/0x330
>  ? __schedule+0x33e/0x8a0
>  cpuhp_down_callbacks+0x51/0xc0
>  cpuhp_thread_fun+0xc1/0xf0
>  smpboot_thread_fn+0x159/0x2a0
>  ? smpboot_create_threads+0x80/0x80
>  kthread+0xef/0x110
>  ? wait_for_completion+0xf0/0x120
>  ? schedule_tail+0x35/0xf0
>  ret_from_fork+0x22/0x50
>  ? __init_kthread_worker+0x70/0x70
> ---[ end trace eb12ae47d2382d8f ]---
> notify_down_prepare: attempt to take down CPU 0 failed
>
> This bug can be reproduced by below config w/ nohz_full= all cpus:
>
> CONFIG_BOOTPARAM_HOTPLUG_CPU0=y
> CONFIG_DEBUG_HOTPLUG_CPU0=y
> CONFIG_NO_HZ_FULL=y
>
> The boot CPU handles housekeeping duty(unbound timers, workqueues,
> timekeeping, ...) on behalf of full dynticks CPUs. It must remain
> online when nohz full is enabled. There is a priority set to every
> notifier_blocks:
>
> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down
>
> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and
> notifier_blocks behind tick_nohz_cpu_down will not be called any
> more, which leads to workers are actually not unbound. Then hotplug
> state machine will fallback to undo and online cpu 0 again. Workers
> will be rebound unconditionally even if they are not unbound and
> trigger the warning in this progress.
>
> This patch fix it by catching !DISASSOCIATED to avoid rebind bound
> workers.
>
> Cc: Tejun Heo <tj@kernel.org>
> Cc: Lai Jiangshan <jiangshanlai@gmail.com>
> Suggested-by: Lai Jiangshan <jiangshanlai@gmail.com>
> Signed-off-by: Wanpeng Li <wanpeng.li@hotmail.com>
> ---
>  kernel/workqueue.c | 6 ++++++
>  1 file changed, 6 insertions(+)
>
> diff --git a/kernel/workqueue.c b/kernel/workqueue.c
> index 2232ae3..cc18920 100644
> --- a/kernel/workqueue.c
> +++ b/kernel/workqueue.c
> @@ -4525,6 +4525,12 @@ static void rebind_workers(struct worker_pool *pool)
>                                                   pool->attrs->cpumask) < 0);
>
>         spin_lock_irq(&pool->lock);
> +
> +       if (!(pool->flags & POOL_DISASSOCIATED)) {
> +               spin_unlock_irq(&pool->lock);
> +               return;
> +       }
> +
>         pool->flags &= ~POOL_DISASSOCIATED;
>
>         for_each_pool_worker(worker, pool) {
> --
> 1.9.1
>

[toc] | [prev] | [next] | [standalone]


#1397196

FromTejun Heo <tj@kernel.org>
Date2016-05-09 19:10 +0200
Message-ID<rwWQW-1tE-37@gated-at.bofh.it>
In reply to#1394822
Hello,

On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote:
> The boot CPU handles housekeeping duty(unbound timers, workqueues, 
> timekeeping, ...) on behalf of full dynticks CPUs. It must remain 
> online when nohz full is enabled. There is a priority set to every
> notifier_blocks:
> 
> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down
> 
> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and 
> notifier_blocks behind tick_nohz_cpu_down will not be called any 
> more, which leads to workers are actually not unbound. Then hotplug 
> state machine will fallback to undo and online cpu 0 again. Workers 
> will be rebound unconditionally even if they are not unbound and 
> trigger the warning in this progress.

I'm a bit confused.  Are you saying that the hotplug statemachine may
invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback?

Thanks.

-- 
tejun

[toc] | [prev] | [next] | [standalone]


#1397432

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-10 00:00 +0200
Message-ID<rx1nC-5F3-37@gated-at.bofh.it>
In reply to#1397196
Cc Thomas, the new state machine author,
2016-05-10 1:00 GMT+08:00 Tejun Heo <tj@kernel.org>:
> Hello,
>
> On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote:
>> The boot CPU handles housekeeping duty(unbound timers, workqueues,
>> timekeeping, ...) on behalf of full dynticks CPUs. It must remain
>> online when nohz full is enabled. There is a priority set to every
>> notifier_blocks:
>>
>> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down
>>
>> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and
>> notifier_blocks behind tick_nohz_cpu_down will not be called any
>> more, which leads to workers are actually not unbound. Then hotplug
>> state machine will fallback to undo and online cpu 0 again. Workers
>> will be rebound unconditionally even if they are not unbound and
>> trigger the warning in this progress.
>
> I'm a bit confused.  Are you saying that the hotplug statemachine may
> invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback?

I think so. CPU_DOWN_FAILED is detected in the process of CPU_DOWN_PREPARE

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1397463

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-10 00:20 +0200
Message-ID<rx1GW-6f0-21@gated-at.bofh.it>
In reply to#1397432
Cc Peterz, Frederic,
2016-05-10 5:50 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
> Cc Thomas, the new state machine author,
> 2016-05-10 1:00 GMT+08:00 Tejun Heo <tj@kernel.org>:
>> Hello,
>>
>> On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote:
>>> The boot CPU handles housekeeping duty(unbound timers, workqueues,
>>> timekeeping, ...) on behalf of full dynticks CPUs. It must remain
>>> online when nohz full is enabled. There is a priority set to every
>>> notifier_blocks:
>>>
>>> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down
>>>
>>> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and
>>> notifier_blocks behind tick_nohz_cpu_down will not be called any
>>> more, which leads to workers are actually not unbound. Then hotplug
>>> state machine will fallback to undo and online cpu 0 again. Workers
>>> will be rebound unconditionally even if they are not unbound and
>>> trigger the warning in this progress.
>>
>> I'm a bit confused.  Are you saying that the hotplug statemachine may
>> invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback?
>
> I think so. CPU_DOWN_FAILED is detected in the process of CPU_DOWN_PREPARE
>
> Regards,
> Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1398574

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-11 01:30 +0200
Message-ID<rxpgd-4fX-3@gated-at.bofh.it>
In reply to#1397432
Hi Tejun,
2016-05-10 5:50 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
> Cc Thomas, the new state machine author,
> 2016-05-10 1:00 GMT+08:00 Tejun Heo <tj@kernel.org>:
>> Hello,
>>
>> On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote:
>>> The boot CPU handles housekeeping duty(unbound timers, workqueues,
>>> timekeeping, ...) on behalf of full dynticks CPUs. It must remain
>>> online when nohz full is enabled. There is a priority set to every
>>> notifier_blocks:
>>>
>>> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down
>>>
>>> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and
>>> notifier_blocks behind tick_nohz_cpu_down will not be called any
>>> more, which leads to workers are actually not unbound. Then hotplug
>>> state machine will fallback to undo and online cpu 0 again. Workers
>>> will be rebound unconditionally even if they are not unbound and
>>> trigger the warning in this progress.
>>
>> I'm a bit confused.  Are you saying that the hotplug statemachine may
>> invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback?
>
> I think so. CPU_DOWN_FAILED is detected in the process of CPU_DOWN_PREPARE

So any plan to apply? :-)

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1398743

FromThomas Gleixner <tglx@linutronix.de>
Date2016-05-11 09:40 +0200
Message-ID<rxwUr-3wW-21@gated-at.bofh.it>
In reply to#1398574
On Wed, 11 May 2016, Wanpeng Li wrote:
> Hi Tejun,
> 2016-05-10 5:50 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
> > Cc Thomas, the new state machine author,
> > 2016-05-10 1:00 GMT+08:00 Tejun Heo <tj@kernel.org>:
> >> Hello,
> >>
> >> On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote:
> >>> The boot CPU handles housekeeping duty(unbound timers, workqueues,
> >>> timekeeping, ...) on behalf of full dynticks CPUs. It must remain
> >>> online when nohz full is enabled. There is a priority set to every
> >>> notifier_blocks:
> >>>
> >>> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down
> >>>
> >>> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and
> >>> notifier_blocks behind tick_nohz_cpu_down will not be called any
> >>> more, which leads to workers are actually not unbound. Then hotplug
> >>> state machine will fallback to undo and online cpu 0 again. Workers
> >>> will be rebound unconditionally even if they are not unbound and
> >>> trigger the warning in this progress.
> >>
> >> I'm a bit confused.  Are you saying that the hotplug statemachine may
> >> invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback?
> >
> > I think so. CPU_DOWN_FAILED is detected in the process of CPU_DOWN_PREPARE

Well, no. It's not detected.

If a down prepare callback fails, then DOWN_FAILED is invoked for all
callbacks which have successfully executed DOWN_PREPARE.

But, workqueue has actually two notifiers. One which handles
UP/DOWN_FAILED/ONLINE and one which handles DOWN_PREPARE.

Now look at the priorities of those callbacks:

    CPU_PRI_WORKQUEUE_UP	= 5
    CPU_PRI_WORKQUEUE_DOWN	= -5

So the call order on DOWN_PREPARE is:

   CB 1
   CB ...
   CB workqueue_up() -> Ignores DOWN_PREPARE
   CB ...
   CB X ---> Fails

So we call up to CB X with DOWN_FAILED

   CB 1
   CB ...
   CB workqueue_up() -> Handles DOWN_FAILED
   CB ...
   CB X-1

So the problem is that the workqueue stuff handles DOWN_FAILED in the up
callback, while it should do it in the down callback. Which is not a good idea
either because it wants to be called early on rollback...

Brilliant stuff, isn't it? The hotplug rework will solve this problem because
the callbacks become symetric, but for the existing mess, we need some
workaround in the workqueue code.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1398753

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-11 10:10 +0200
Message-ID<rxxnr-467-9@gated-at.bofh.it>
In reply to#1398743
2016-05-11 15:34 GMT+08:00 Thomas Gleixner <tglx@linutronix.de>:
> On Wed, 11 May 2016, Wanpeng Li wrote:
>> Hi Tejun,
>> 2016-05-10 5:50 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
>> > Cc Thomas, the new state machine author,
>> > 2016-05-10 1:00 GMT+08:00 Tejun Heo <tj@kernel.org>:
>> >> Hello,
>> >>
>> >> On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote:
>> >>> The boot CPU handles housekeeping duty(unbound timers, workqueues,
>> >>> timekeeping, ...) on behalf of full dynticks CPUs. It must remain
>> >>> online when nohz full is enabled. There is a priority set to every
>> >>> notifier_blocks:
>> >>>
>> >>> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down
>> >>>
>> >>> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and
>> >>> notifier_blocks behind tick_nohz_cpu_down will not be called any
>> >>> more, which leads to workers are actually not unbound. Then hotplug
>> >>> state machine will fallback to undo and online cpu 0 again. Workers
>> >>> will be rebound unconditionally even if they are not unbound and
>> >>> trigger the warning in this progress.
>> >>
>> >> I'm a bit confused.  Are you saying that the hotplug statemachine may
>> >> invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback?
>> >
>> > I think so. CPU_DOWN_FAILED is detected in the process of CPU_DOWN_PREPARE
>
> Well, no. It's not detected.
>
> If a down prepare callback fails, then DOWN_FAILED is invoked for all
> callbacks which have successfully executed DOWN_PREPARE.
>
> But, workqueue has actually two notifiers. One which handles
> UP/DOWN_FAILED/ONLINE and one which handles DOWN_PREPARE.
>
> Now look at the priorities of those callbacks:
>
>     CPU_PRI_WORKQUEUE_UP        = 5
>     CPU_PRI_WORKQUEUE_DOWN      = -5
>
> So the call order on DOWN_PREPARE is:
>
>    CB 1
>    CB ...
>    CB workqueue_up() -> Ignores DOWN_PREPARE
>    CB ...
>    CB X ---> Fails
>
> So we call up to CB X with DOWN_FAILED
>
>    CB 1
>    CB ...
>    CB workqueue_up() -> Handles DOWN_FAILED
>    CB ...
>    CB X-1
>
> So the problem is that the workqueue stuff handles DOWN_FAILED in the up
> callback, while it should do it in the down callback. Which is not a good idea
> either because it wants to be called early on rollback...
>
> Brilliant stuff, isn't it? The hotplug rework will solve this problem because
> the callbacks become symetric, but for the existing mess, we need some
> workaround in the workqueue code.

Thanks for your reply, Thomas. :-)
Do you think the current version patch is the right fix/workaround for
the existing mess?

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1398891

FromThomas Gleixner <tglx@linutronix.de>
Date2016-05-11 12:10 +0200
Message-ID<rxzfA-6aG-15@gated-at.bofh.it>
In reply to#1398753
On Wed, 11 May 2016, Wanpeng Li wrote:
> Do you think the current version patch is the right fix/workaround for
> the existing mess?

You might need something stateful related to the hotplug crap, as that
DOWN_FAILED/ONLINE case does a lot of stuff.

Though I leave it to Tejun to decided whether your proposed workaround is good
enough to fix the issue at hand.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1398911

FromWanpeng Li <kernellwp@gmail.com>
Date2016-05-11 12:30 +0200
Message-ID<rxzyX-6m4-25@gated-at.bofh.it>
In reply to#1398891
2016-05-11 18:03 GMT+08:00 Thomas Gleixner <tglx@linutronix.de>:
> On Wed, 11 May 2016, Wanpeng Li wrote:
>> Do you think the current version patch is the right fix/workaround for
>> the existing mess?
>
> You might need something stateful related to the hotplug crap, as that
> DOWN_FAILED/ONLINE case does a lot of stuff.

Yeah, I just mean the current version patch before your finial hotplug
rework. :-)

Regards,
Wanpeng Li

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web