Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1394822 > unrolled thread
| Started by | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| First post | 2016-05-05 03:50 +0200 |
| Last post | 2016-05-11 12:30 +0200 |
| Articles | 10 — 3 participants |
Back to article view | Back to linux.kernel
[PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-05 03:50 +0200
Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-09 09:30 +0200
Re: [PATCH] workqueue: fix rebind bound workers warning Tejun Heo <tj@kernel.org> - 2016-05-09 19:10 +0200
Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-10 00:00 +0200
Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-10 00:20 +0200
Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-11 01:30 +0200
Re: [PATCH] workqueue: fix rebind bound workers warning Thomas Gleixner <tglx@linutronix.de> - 2016-05-11 09:40 +0200
Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-11 10:10 +0200
Re: [PATCH] workqueue: fix rebind bound workers warning Thomas Gleixner <tglx@linutronix.de> - 2016-05-11 12:10 +0200
Re: [PATCH] workqueue: fix rebind bound workers warning Wanpeng Li <kernellwp@gmail.com> - 2016-05-11 12:30 +0200
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-05-05 03:50 +0200 |
| Subject | [PATCH] workqueue: fix rebind bound workers warning |
| Message-ID | <rvgAr-6ZQ-5@gated-at.bofh.it> |
From: Wanpeng Li <wanpeng.li@hotmail.com>
------------[ cut here ]------------
WARNING: CPU: 0 PID: 16 at kernel/workqueue.c:4559 rebind_workers+0x1c0/0x1d0
Modules linked in:
CPU: 0 PID: 16 Comm: cpuhp/0 Not tainted 4.6.0-rc4+ #31
Hardware name: IBM IBM System x3550 M4 Server -[7914IUW]-/00Y8603, BIOS -[D7E128FUS-1.40]- 07/23/2013
0000000000000000 ffff881037babb58 ffffffff8139d885 0000000000000010
0000000000000000 0000000000000000 0000000000000000 ffff881037babba8
ffffffff8108505d ffff881037ba0000 000011cf3e7d6e60 0000000000000046
Call Trace:
dump_stack+0x89/0xd4
__warn+0xfd/0x120
warn_slowpath_null+0x1d/0x20
rebind_workers+0x1c0/0x1d0
workqueue_cpu_up_callback+0xf5/0x1d0
notifier_call_chain+0x64/0x90
? trace_hardirqs_on_caller+0xf2/0x220
? notify_prepare+0x80/0x80
__raw_notifier_call_chain+0xe/0x10
__cpu_notify+0x35/0x50
notify_down_prepare+0x5e/0x80
? notify_prepare+0x80/0x80
cpuhp_invoke_callback+0x73/0x330
? __schedule+0x33e/0x8a0
cpuhp_down_callbacks+0x51/0xc0
cpuhp_thread_fun+0xc1/0xf0
smpboot_thread_fn+0x159/0x2a0
? smpboot_create_threads+0x80/0x80
kthread+0xef/0x110
? wait_for_completion+0xf0/0x120
? schedule_tail+0x35/0xf0
ret_from_fork+0x22/0x50
? __init_kthread_worker+0x70/0x70
---[ end trace eb12ae47d2382d8f ]---
notify_down_prepare: attempt to take down CPU 0 failed
This bug can be reproduced by below config w/ nohz_full= all cpus:
CONFIG_BOOTPARAM_HOTPLUG_CPU0=y
CONFIG_DEBUG_HOTPLUG_CPU0=y
CONFIG_NO_HZ_FULL=y
The boot CPU handles housekeeping duty(unbound timers, workqueues,
timekeeping, ...) on behalf of full dynticks CPUs. It must remain
online when nohz full is enabled. There is a priority set to every
notifier_blocks:
workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down
So tick_nohz_cpu_down callback failed when down prepare cpu 0, and
notifier_blocks behind tick_nohz_cpu_down will not be called any
more, which leads to workers are actually not unbound. Then hotplug
state machine will fallback to undo and online cpu 0 again. Workers
will be rebound unconditionally even if they are not unbound and
trigger the warning in this progress.
This patch fix it by catching !DISASSOCIATED to avoid rebind bound
workers.
Cc: Tejun Heo <tj@kernel.org>
Cc: Lai Jiangshan <jiangshanlai@gmail.com>
Suggested-by: Lai Jiangshan <jiangshanlai@gmail.com>
Signed-off-by: Wanpeng Li <wanpeng.li@hotmail.com>
---
kernel/workqueue.c | 6 ++++++
1 file changed, 6 insertions(+)
diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index 2232ae3..cc18920 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -4525,6 +4525,12 @@ static void rebind_workers(struct worker_pool *pool)
pool->attrs->cpumask) < 0);
spin_lock_irq(&pool->lock);
+
+ if (!(pool->flags & POOL_DISASSOCIATED)) {
+ spin_unlock_irq(&pool->lock);
+ return;
+ }
+
pool->flags &= ~POOL_DISASSOCIATED;
for_each_pool_worker(worker, pool) {
--
1.9.1
[toc] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-05-09 09:30 +0200 |
| Message-ID | <rwNNH-7J-79@gated-at.bofh.it> |
| In reply to | #1394822 |
Sorry to quick ping you Tejun, just hope it can catch the upcoming
merge window. :-)
2016-05-05 9:41 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
> From: Wanpeng Li <wanpeng.li@hotmail.com>
>
> ------------[ cut here ]------------
> WARNING: CPU: 0 PID: 16 at kernel/workqueue.c:4559 rebind_workers+0x1c0/0x1d0
> Modules linked in:
> CPU: 0 PID: 16 Comm: cpuhp/0 Not tainted 4.6.0-rc4+ #31
> Hardware name: IBM IBM System x3550 M4 Server -[7914IUW]-/00Y8603, BIOS -[D7E128FUS-1.40]- 07/23/2013
> 0000000000000000 ffff881037babb58 ffffffff8139d885 0000000000000010
> 0000000000000000 0000000000000000 0000000000000000 ffff881037babba8
> ffffffff8108505d ffff881037ba0000 000011cf3e7d6e60 0000000000000046
> Call Trace:
> dump_stack+0x89/0xd4
> __warn+0xfd/0x120
> warn_slowpath_null+0x1d/0x20
> rebind_workers+0x1c0/0x1d0
> workqueue_cpu_up_callback+0xf5/0x1d0
> notifier_call_chain+0x64/0x90
> ? trace_hardirqs_on_caller+0xf2/0x220
> ? notify_prepare+0x80/0x80
> __raw_notifier_call_chain+0xe/0x10
> __cpu_notify+0x35/0x50
> notify_down_prepare+0x5e/0x80
> ? notify_prepare+0x80/0x80
> cpuhp_invoke_callback+0x73/0x330
> ? __schedule+0x33e/0x8a0
> cpuhp_down_callbacks+0x51/0xc0
> cpuhp_thread_fun+0xc1/0xf0
> smpboot_thread_fn+0x159/0x2a0
> ? smpboot_create_threads+0x80/0x80
> kthread+0xef/0x110
> ? wait_for_completion+0xf0/0x120
> ? schedule_tail+0x35/0xf0
> ret_from_fork+0x22/0x50
> ? __init_kthread_worker+0x70/0x70
> ---[ end trace eb12ae47d2382d8f ]---
> notify_down_prepare: attempt to take down CPU 0 failed
>
> This bug can be reproduced by below config w/ nohz_full= all cpus:
>
> CONFIG_BOOTPARAM_HOTPLUG_CPU0=y
> CONFIG_DEBUG_HOTPLUG_CPU0=y
> CONFIG_NO_HZ_FULL=y
>
> The boot CPU handles housekeeping duty(unbound timers, workqueues,
> timekeeping, ...) on behalf of full dynticks CPUs. It must remain
> online when nohz full is enabled. There is a priority set to every
> notifier_blocks:
>
> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down
>
> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and
> notifier_blocks behind tick_nohz_cpu_down will not be called any
> more, which leads to workers are actually not unbound. Then hotplug
> state machine will fallback to undo and online cpu 0 again. Workers
> will be rebound unconditionally even if they are not unbound and
> trigger the warning in this progress.
>
> This patch fix it by catching !DISASSOCIATED to avoid rebind bound
> workers.
>
> Cc: Tejun Heo <tj@kernel.org>
> Cc: Lai Jiangshan <jiangshanlai@gmail.com>
> Suggested-by: Lai Jiangshan <jiangshanlai@gmail.com>
> Signed-off-by: Wanpeng Li <wanpeng.li@hotmail.com>
> ---
> kernel/workqueue.c | 6 ++++++
> 1 file changed, 6 insertions(+)
>
> diff --git a/kernel/workqueue.c b/kernel/workqueue.c
> index 2232ae3..cc18920 100644
> --- a/kernel/workqueue.c
> +++ b/kernel/workqueue.c
> @@ -4525,6 +4525,12 @@ static void rebind_workers(struct worker_pool *pool)
> pool->attrs->cpumask) < 0);
>
> spin_lock_irq(&pool->lock);
> +
> + if (!(pool->flags & POOL_DISASSOCIATED)) {
> + spin_unlock_irq(&pool->lock);
> + return;
> + }
> +
> pool->flags &= ~POOL_DISASSOCIATED;
>
> for_each_pool_worker(worker, pool) {
> --
> 1.9.1
>
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2016-05-09 19:10 +0200 |
| Message-ID | <rwWQW-1tE-37@gated-at.bofh.it> |
| In reply to | #1394822 |
Hello, On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote: > The boot CPU handles housekeeping duty(unbound timers, workqueues, > timekeeping, ...) on behalf of full dynticks CPUs. It must remain > online when nohz full is enabled. There is a priority set to every > notifier_blocks: > > workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down > > So tick_nohz_cpu_down callback failed when down prepare cpu 0, and > notifier_blocks behind tick_nohz_cpu_down will not be called any > more, which leads to workers are actually not unbound. Then hotplug > state machine will fallback to undo and online cpu 0 again. Workers > will be rebound unconditionally even if they are not unbound and > trigger the warning in this progress. I'm a bit confused. Are you saying that the hotplug statemachine may invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback? Thanks. -- tejun
[toc] | [prev] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-05-10 00:00 +0200 |
| Message-ID | <rx1nC-5F3-37@gated-at.bofh.it> |
| In reply to | #1397196 |
Cc Thomas, the new state machine author, 2016-05-10 1:00 GMT+08:00 Tejun Heo <tj@kernel.org>: > Hello, > > On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote: >> The boot CPU handles housekeeping duty(unbound timers, workqueues, >> timekeeping, ...) on behalf of full dynticks CPUs. It must remain >> online when nohz full is enabled. There is a priority set to every >> notifier_blocks: >> >> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down >> >> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and >> notifier_blocks behind tick_nohz_cpu_down will not be called any >> more, which leads to workers are actually not unbound. Then hotplug >> state machine will fallback to undo and online cpu 0 again. Workers >> will be rebound unconditionally even if they are not unbound and >> trigger the warning in this progress. > > I'm a bit confused. Are you saying that the hotplug statemachine may > invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback? I think so. CPU_DOWN_FAILED is detected in the process of CPU_DOWN_PREPARE Regards, Wanpeng Li
[toc] | [prev] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-05-10 00:20 +0200 |
| Message-ID | <rx1GW-6f0-21@gated-at.bofh.it> |
| In reply to | #1397432 |
Cc Peterz, Frederic, 2016-05-10 5:50 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>: > Cc Thomas, the new state machine author, > 2016-05-10 1:00 GMT+08:00 Tejun Heo <tj@kernel.org>: >> Hello, >> >> On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote: >>> The boot CPU handles housekeeping duty(unbound timers, workqueues, >>> timekeeping, ...) on behalf of full dynticks CPUs. It must remain >>> online when nohz full is enabled. There is a priority set to every >>> notifier_blocks: >>> >>> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down >>> >>> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and >>> notifier_blocks behind tick_nohz_cpu_down will not be called any >>> more, which leads to workers are actually not unbound. Then hotplug >>> state machine will fallback to undo and online cpu 0 again. Workers >>> will be rebound unconditionally even if they are not unbound and >>> trigger the warning in this progress. >> >> I'm a bit confused. Are you saying that the hotplug statemachine may >> invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback? > > I think so. CPU_DOWN_FAILED is detected in the process of CPU_DOWN_PREPARE > > Regards, > Wanpeng Li
[toc] | [prev] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-05-11 01:30 +0200 |
| Message-ID | <rxpgd-4fX-3@gated-at.bofh.it> |
| In reply to | #1397432 |
Hi Tejun, 2016-05-10 5:50 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>: > Cc Thomas, the new state machine author, > 2016-05-10 1:00 GMT+08:00 Tejun Heo <tj@kernel.org>: >> Hello, >> >> On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote: >>> The boot CPU handles housekeeping duty(unbound timers, workqueues, >>> timekeeping, ...) on behalf of full dynticks CPUs. It must remain >>> online when nohz full is enabled. There is a priority set to every >>> notifier_blocks: >>> >>> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down >>> >>> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and >>> notifier_blocks behind tick_nohz_cpu_down will not be called any >>> more, which leads to workers are actually not unbound. Then hotplug >>> state machine will fallback to undo and online cpu 0 again. Workers >>> will be rebound unconditionally even if they are not unbound and >>> trigger the warning in this progress. >> >> I'm a bit confused. Are you saying that the hotplug statemachine may >> invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback? > > I think so. CPU_DOWN_FAILED is detected in the process of CPU_DOWN_PREPARE So any plan to apply? :-) Regards, Wanpeng Li
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-05-11 09:40 +0200 |
| Message-ID | <rxwUr-3wW-21@gated-at.bofh.it> |
| In reply to | #1398574 |
On Wed, 11 May 2016, Wanpeng Li wrote:
> Hi Tejun,
> 2016-05-10 5:50 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>:
> > Cc Thomas, the new state machine author,
> > 2016-05-10 1:00 GMT+08:00 Tejun Heo <tj@kernel.org>:
> >> Hello,
> >>
> >> On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote:
> >>> The boot CPU handles housekeeping duty(unbound timers, workqueues,
> >>> timekeeping, ...) on behalf of full dynticks CPUs. It must remain
> >>> online when nohz full is enabled. There is a priority set to every
> >>> notifier_blocks:
> >>>
> >>> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down
> >>>
> >>> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and
> >>> notifier_blocks behind tick_nohz_cpu_down will not be called any
> >>> more, which leads to workers are actually not unbound. Then hotplug
> >>> state machine will fallback to undo and online cpu 0 again. Workers
> >>> will be rebound unconditionally even if they are not unbound and
> >>> trigger the warning in this progress.
> >>
> >> I'm a bit confused. Are you saying that the hotplug statemachine may
> >> invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback?
> >
> > I think so. CPU_DOWN_FAILED is detected in the process of CPU_DOWN_PREPARE
Well, no. It's not detected.
If a down prepare callback fails, then DOWN_FAILED is invoked for all
callbacks which have successfully executed DOWN_PREPARE.
But, workqueue has actually two notifiers. One which handles
UP/DOWN_FAILED/ONLINE and one which handles DOWN_PREPARE.
Now look at the priorities of those callbacks:
CPU_PRI_WORKQUEUE_UP = 5
CPU_PRI_WORKQUEUE_DOWN = -5
So the call order on DOWN_PREPARE is:
CB 1
CB ...
CB workqueue_up() -> Ignores DOWN_PREPARE
CB ...
CB X ---> Fails
So we call up to CB X with DOWN_FAILED
CB 1
CB ...
CB workqueue_up() -> Handles DOWN_FAILED
CB ...
CB X-1
So the problem is that the workqueue stuff handles DOWN_FAILED in the up
callback, while it should do it in the down callback. Which is not a good idea
either because it wants to be called early on rollback...
Brilliant stuff, isn't it? The hotplug rework will solve this problem because
the callbacks become symetric, but for the existing mess, we need some
workaround in the workqueue code.
Thanks,
tglx
[toc] | [prev] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-05-11 10:10 +0200 |
| Message-ID | <rxxnr-467-9@gated-at.bofh.it> |
| In reply to | #1398743 |
2016-05-11 15:34 GMT+08:00 Thomas Gleixner <tglx@linutronix.de>: > On Wed, 11 May 2016, Wanpeng Li wrote: >> Hi Tejun, >> 2016-05-10 5:50 GMT+08:00 Wanpeng Li <kernellwp@gmail.com>: >> > Cc Thomas, the new state machine author, >> > 2016-05-10 1:00 GMT+08:00 Tejun Heo <tj@kernel.org>: >> >> Hello, >> >> >> >> On Thu, May 05, 2016 at 09:41:31AM +0800, Wanpeng Li wrote: >> >>> The boot CPU handles housekeeping duty(unbound timers, workqueues, >> >>> timekeeping, ...) on behalf of full dynticks CPUs. It must remain >> >>> online when nohz full is enabled. There is a priority set to every >> >>> notifier_blocks: >> >>> >> >>> workqueue_cpu_up > tick_nohz_cpu_down > workqueue_cpu_down >> >>> >> >>> So tick_nohz_cpu_down callback failed when down prepare cpu 0, and >> >>> notifier_blocks behind tick_nohz_cpu_down will not be called any >> >>> more, which leads to workers are actually not unbound. Then hotplug >> >>> state machine will fallback to undo and online cpu 0 again. Workers >> >>> will be rebound unconditionally even if they are not unbound and >> >>> trigger the warning in this progress. >> >> >> >> I'm a bit confused. Are you saying that the hotplug statemachine may >> >> invoke CPU_DOWN_FAILED w/o preceding CPU_DOWN on the same callback? >> > >> > I think so. CPU_DOWN_FAILED is detected in the process of CPU_DOWN_PREPARE > > Well, no. It's not detected. > > If a down prepare callback fails, then DOWN_FAILED is invoked for all > callbacks which have successfully executed DOWN_PREPARE. > > But, workqueue has actually two notifiers. One which handles > UP/DOWN_FAILED/ONLINE and one which handles DOWN_PREPARE. > > Now look at the priorities of those callbacks: > > CPU_PRI_WORKQUEUE_UP = 5 > CPU_PRI_WORKQUEUE_DOWN = -5 > > So the call order on DOWN_PREPARE is: > > CB 1 > CB ... > CB workqueue_up() -> Ignores DOWN_PREPARE > CB ... > CB X ---> Fails > > So we call up to CB X with DOWN_FAILED > > CB 1 > CB ... > CB workqueue_up() -> Handles DOWN_FAILED > CB ... > CB X-1 > > So the problem is that the workqueue stuff handles DOWN_FAILED in the up > callback, while it should do it in the down callback. Which is not a good idea > either because it wants to be called early on rollback... > > Brilliant stuff, isn't it? The hotplug rework will solve this problem because > the callbacks become symetric, but for the existing mess, we need some > workaround in the workqueue code. Thanks for your reply, Thomas. :-) Do you think the current version patch is the right fix/workaround for the existing mess? Regards, Wanpeng Li
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-05-11 12:10 +0200 |
| Message-ID | <rxzfA-6aG-15@gated-at.bofh.it> |
| In reply to | #1398753 |
On Wed, 11 May 2016, Wanpeng Li wrote: > Do you think the current version patch is the right fix/workaround for > the existing mess? You might need something stateful related to the hotplug crap, as that DOWN_FAILED/ONLINE case does a lot of stuff. Though I leave it to Tejun to decided whether your proposed workaround is good enough to fix the issue at hand. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2016-05-11 12:30 +0200 |
| Message-ID | <rxzyX-6m4-25@gated-at.bofh.it> |
| In reply to | #1398891 |
2016-05-11 18:03 GMT+08:00 Thomas Gleixner <tglx@linutronix.de>: > On Wed, 11 May 2016, Wanpeng Li wrote: >> Do you think the current version patch is the right fix/workaround for >> the existing mess? > > You might need something stateful related to the hotplug crap, as that > DOWN_FAILED/ONLINE case does a lot of stuff. Yeah, I just mean the current version patch before your finial hotplug rework. :-) Regards, Wanpeng Li
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web