Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1575081 > unrolled thread
| Started by | Dmitry Vyukov <dvyukov@google.com> |
|---|---|
| First post | 2017-02-06 20:20 +0100 |
| Last post | 2017-02-07 23:40 +0100 |
| Articles | 17 on this page of 57 — 9 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: mm: deadlock between get_online_cpus/pcpu_alloc Dmitry Vyukov <dvyukov@google.com> - 2017-02-06 20:20 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-06 23:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 09:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Vlastimil Babka <vbabka@suse.cz> - 2017-02-07 10:30 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-07 10:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 11:00 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-07 11:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-07 12:20 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-07 10:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Vlastimil Babka <vbabka@suse.cz> - 2017-02-07 11:00 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 11:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-07 11:30 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 11:40 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-07 12:40 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 12:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Vlastimil Babka <vbabka@suse.cz> - 2017-02-07 13:00 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 13:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 13:40 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Vlastimil Babka <vbabka@suse.cz> - 2017-02-07 13:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 13:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Vlastimil Babka <vbabka@suse.cz> - 2017-02-07 15:00 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-07 15:00 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 15:20 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 16:40 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-07 17:30 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 17:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Christoph Lameter <cl@linux.com> - 2017-02-07 18:00 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Thomas Gleixner <tglx@linutronix.de> - 2017-02-07 23:30 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-08 08:40 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Thomas Gleixner <tglx@linutronix.de> - 2017-02-08 13:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-08 13:30 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-08 13:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-08 15:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Peter Zijlstra <peterz@infradead.org> - 2017-02-08 18:00 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Thomas Gleixner <tglx@linutronix.de> - 2017-02-08 15:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Christoph Lameter <cl@linux.com> - 2017-02-08 16:30 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Christoph Lameter <cl@linux.com> - 2017-02-08 17:30 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Thomas Gleixner <tglx@linutronix.de> - 2017-02-08 19:40 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Christoph Lameter <cl@linux.com> - 2017-02-09 04:20 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Thomas Gleixner <tglx@linutronix.de> - 2017-02-09 12:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Christoph Lameter <cl@linux.com> - 2017-02-09 15:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Thomas Gleixner <tglx@linutronix.de> - 2017-02-09 16:40 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Christoph Lameter <cl@linux.com> - 2017-02-09 16:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Thomas Gleixner <tglx@linutronix.de> - 2017-02-09 17:20 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Christoph Lameter <cl@linux.com> - 2017-02-09 18:30 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Thomas Gleixner <tglx@linutronix.de> - 2017-02-09 18:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-09 20:20 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Christoph Lameter <cl@linux.com> - 2017-02-10 19:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-08 18:40 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Christoph Lameter <cl@linux.com> - 2017-02-08 16:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Tejun Heo <tj@kernel.org> - 2017-02-07 18:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 21:40 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Mel Gorman <mgorman@techsingularity.net> - 2017-02-07 14:10 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 14:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp> - 2017-02-07 12:30 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Michal Hocko <mhocko@kernel.org> - 2017-02-07 09:50 +0100
Re: mm: deadlock between get_online_cpus/pcpu_alloc Thomas Gleixner <tglx@linutronix.de> - 2017-02-07 23:40 +0100
Page 3 of 3 — ← Prev page 1 2 [3]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2017-02-09 15:10 +0100 |
| Message-ID | <t8XA7-79s-47@gated-at.bofh.it> |
| In reply to | #1577496 |
On Thu, 9 Feb 2017, Thomas Gleixner wrote: > And how does that solve the problem at hand? Not at all: > > CPU 0 CPU 1 > > for_each_online_cpu(cpu) > ==> cpu = 1 > stop_machine() > set_cpu_online(1, false) > queue_work(cpu1) > > Thanks, Well thats not how I remember stop_machine does work. Doesnt it stop the processing on all cpus otherwise its not a real usable stop. The stop_machine would need to ensure that all cpus cease processing before proceeding.
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-02-09 16:40 +0100 |
| Message-ID | <t8YZd-7Tl-31@gated-at.bofh.it> |
| In reply to | #1577625 |
On Thu, 9 Feb 2017, Christoph Lameter wrote: > On Thu, 9 Feb 2017, Thomas Gleixner wrote: > > > And how does that solve the problem at hand? Not at all: > > > > CPU 0 CPU 1 > > > > for_each_online_cpu(cpu) > > ==> cpu = 1 > > stop_machine() > > set_cpu_online(1, false) > > queue_work(cpu1) > > > > Thanks, > > Well thats not how I remember stop_machine does work. Doesnt it stop the > processing on all cpus otherwise its not a real usable stop. > > The stop_machine would need to ensure that all cpus cease processing > before proceeding. Ok. I try again: CPU 0 CPU 1 for_each_online_cpu(cpu) ==> cpu = 1 stop_machine() Stops processing on all CPUs by preempting the current execution and forcing them into a high priority busy loop with interrupts disabled. context_switch() stomper_thread() busyloop() set_cpu_online(1, false) stop_machine end() release busy looping CPUs context_switch Resumes operation at the preemption point. cpu is still 1 queue_work(cpu == 1) It does exactly what you describe. It stops processing on all other cpus until release, but that does not invalidate any data on those cpus. It's been that way forever. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2017-02-09 16:50 +0100 |
| Message-ID | <t8Z8S-7WR-15@gated-at.bofh.it> |
| In reply to | #1577705 |
On Thu, 9 Feb 2017, Thomas Gleixner wrote: > > The stop_machine would need to ensure that all cpus cease processing > > before proceeding. > > Ok. I try again: > > CPU 0 CPU 1 > for_each_online_cpu(cpu) > ==> cpu = 1 > stop_machine() > > Stops processing on all CPUs by preempting the current execution and > forcing them into a high priority busy loop with interrupts disabled. Exactly that means we are outside of the sections marked with get_online_cpous(). > It does exactly what you describe. It stops processing on all other cpus > until release, but that does not invalidate any data on those cpus. Why would it need to invalidate any data? The change of the cpu masks would need to be done when the machine is stopped. This sounds exactly like what we need and much of it is already there. Lets get rid of get_online_cpus() etc.
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-02-09 17:20 +0100 |
| Message-ID | <t8ZBT-8mq-7@gated-at.bofh.it> |
| In reply to | #1577719 |
On Thu, 9 Feb 2017, Christoph Lameter wrote: > On Thu, 9 Feb 2017, Thomas Gleixner wrote: > > > > The stop_machine would need to ensure that all cpus cease processing > > > before proceeding. > > > > Ok. I try again: > > > > CPU 0 CPU 1 > > for_each_online_cpu(cpu) > > ==> cpu = 1 > > stop_machine() > > > > Stops processing on all CPUs by preempting the current execution and > > forcing them into a high priority busy loop with interrupts disabled. > > Exactly that means we are outside of the sections marked with > get_online_cpous(). > > > It does exactly what you describe. It stops processing on all other cpus > > until release, but that does not invalidate any data on those cpus. > > Why would it need to invalidate any data? The change of the cpu masks > would need to be done when the machine is stopped. This sounds exactly > like what we need and much of it is already there. You are just not getting it, really. The problem is that this for_each_online_cpu() is racy against a concurrent hot unplug and therefor can queue stuff for a not longer online cpu. That's what the mm folks tried to avoid by preventing a CPU hotplug operation before entering that loop. > Lets get rid of get_online_cpus() etc. And that solves what? Can you please start to understand the scope of the whole hotplug machinery including the requirements for get_online_cpus() before you waste everybodys time with your uninformed and halfbaken proposals? Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2017-02-09 18:30 +0100 |
| Message-ID | <t90HE-yG-17@gated-at.bofh.it> |
| In reply to | #1577737 |
On Thu, 9 Feb 2017, Thomas Gleixner wrote: > You are just not getting it, really. > > The problem is that this for_each_online_cpu() is racy against a concurrent > hot unplug and therefor can queue stuff for a not longer online cpu. That's > what the mm folks tried to avoid by preventing a CPU hotplug operation > before entering that loop. With a stop machine action it is NOT racy because the machine goes into a special kernel state that guarantees that key operating system structures are not touched. See mm/page_alloc.c's use of that characteristic to build zonelists. Thus it cannot be executing for_each_online_cpu and related tasks (unless one does not disable preempt .... but that is a given if a spinlock has been taken).. > > Lets get rid of get_online_cpus() etc. > > And that solves what? It gets rid of future issues with serialization in paths were we need to lock and still do for_each_online_cpu(). > Can you please start to understand the scope of the whole hotplug machinery > including the requirements for get_online_cpus() before you waste > everybodys time with your uninformed and halfbaken proposals? Its an obvious solution to the issues that have arisen multiple times with get_online_cpus() within the slab allocators. The hotplug machinery should make things as easy as possible for other people and having these get_online_cpus() everywhere does complicate things.
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-02-09 18:50 +0100 |
| Message-ID | <t910Z-Ft-1@gated-at.bofh.it> |
| In reply to | #1577812 |
On Thu, 9 Feb 2017, Christoph Lameter wrote: > On Thu, 9 Feb 2017, Thomas Gleixner wrote: > > > You are just not getting it, really. > > > > The problem is that this for_each_online_cpu() is racy against a concurrent > > hot unplug and therefor can queue stuff for a not longer online cpu. That's > > what the mm folks tried to avoid by preventing a CPU hotplug operation > > before entering that loop. > > With a stop machine action it is NOT racy because the machine goes into a > special kernel state that guarantees that key operating system structures > are not touched. See mm/page_alloc.c's use of that characteristic to build > zonelists. Thus it cannot be executing for_each_online_cpu and related > tasks (unless one does not disable preempt .... but that is a given if a > spinlock has been taken).. drain_all_pages() is called from preemptible context. So what are you talking about again? > > > Lets get rid of get_online_cpus() etc. > > > > And that solves what? > > It gets rid of future issues with serialization in paths were we need to > lock and still do for_each_online_cpu(). There are code pathes which might sleep inside the loop so get_online_cpus() is the only way to serialize against hotplug. Just because the only tool you know is stop machine it does not make everything an atomic context where it can be applied. > > Can you please start to understand the scope of the whole hotplug machinery > > including the requirements for get_online_cpus() before you waste > > everybodys time with your uninformed and halfbaken proposals? > > Its an obvious solution to the issues that have arisen multiple times with > get_online_cpus() within the slab allocators. The hotplug machinery should > make things as easy as possible for other people and having these > get_online_cpus() everywhere does complicate things. It's no obvious solution to everything. It's context dependend and people have to think hard how to solve their problem within the context they are dealing with. Your 'get rid of get_online_cpus()' mantra does make all of this magically simple. Relying on the fact, that the CPU online bit is cleared via stomp_machine(), which is a horrible operation in various aspects, only applies to a very small subset of problems. You can repeat your mantra another thousand times and that will not make the way larger set of problems magically disappear. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-02-09 20:20 +0100 |
| Message-ID | <t92q5-1Gf-7@gated-at.bofh.it> |
| In reply to | #1577812 |
On Thu 09-02-17 11:22:49, Cristopher Lameter wrote: > On Thu, 9 Feb 2017, Thomas Gleixner wrote: > > > You are just not getting it, really. > > > > The problem is that this for_each_online_cpu() is racy against a concurrent > > hot unplug and therefor can queue stuff for a not longer online cpu. That's > > what the mm folks tried to avoid by preventing a CPU hotplug operation > > before entering that loop. > > With a stop machine action it is NOT racy because the machine goes into a > special kernel state that guarantees that key operating system structures > are not touched. See mm/page_alloc.c's use of that characteristic to build > zonelists. Thus it cannot be executing for_each_online_cpu and related > tasks (unless one does not disable preempt .... but that is a given if a > spinlock has been taken).. Christoph, you are completely ignoring the reality and the code. There is no need for stop_machine nor it is helping anything. As the matter of fact there is a synchronization with the cpu hotplug needed if you want to make a per-cpu specific operations. get_online_cpus is the most straightforward and heavy weight way to do this synchronization but not the only one. As the patch [1] describes we do not really need get_online_cpus in drain_all_pages because we can do _better_. But this is not in any way a generic thing applicable to other code paths. If you disagree then you are free to post patches but hand waving you are doing here is just wasting everybody's time. So please cut it here unless you have specific proposals to improve the current situation. Thanks! [1] http://lkml.kernel.org/r/20170207201950.20482-1-mhocko@kernel.org -- Michal Hocko SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2017-02-10 19:10 +0100 |
| Message-ID | <t9nNT-6VN-17@gated-at.bofh.it> |
| In reply to | #1577889 |
On Thu, 9 Feb 2017, Michal Hocko wrote: > Christoph, you are completely ignoring the reality and the code. There > is no need for stop_machine nor it is helping anything. As the matter > of fact there is a synchronization with the cpu hotplug needed if you > want to make a per-cpu specific operations. get_online_cpus is the > most straightforward and heavy weight way to do this synchronization > but not the only one. As the patch [1] describes we do not really need > get_online_cpus in drain_all_pages because we can do _better_. But > this is not in any way a generic thing applicable to other code paths. > > If you disagree then you are free to post patches but hand waving you > are doing here is just wasting everybody's time. So please cut it here > unless you have specific proposals to improve the current situation. I am fine with the particular solution here for this particular problem. My problem is the general way of having to synchronize via get_online_cpus() because of cpu hotplug operations.
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-02-08 18:40 +0100 |
| Message-ID | <t8Di1-2Aw-13@gated-at.bofh.it> |
| In reply to | #1576646 |
On Wed 08-02-17 09:11:06, Cristopher Lameter wrote: > On Wed, 8 Feb 2017, Michal Hocko wrote: > > > > Huch? stop_machine() is horrible and heavy weight. Don't go there, there > > > must be simpler solutions than that. > > > > Absolutely agreed. We are in the page allocator path so using the > > stop_machine* is just ridiculous. And, in fact, there is a much simpler > > solution [1] > > That is nonsense. stop_machine would be used when adding removing a > processor. There would be no need to synchronize when looping over active > cpus anymore. get_online_cpus() etc would be removed from the hot > path since the cpu masks are guaranteed to be stable. I have no idea what you are trying to say and how this is related to the deadlock we are discussing here. We certainly do not need to add stop_machine the problem. And yeah, dropping get_online_cpus was possible after considering all fallouts. -- Michal Hocko SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2017-02-08 16:10 +0100 |
| Message-ID | <t8C2C-1TA-11@gated-at.bofh.it> |
| In reply to | #1576122 |
On Tue, 7 Feb 2017, Thomas Gleixner wrote: > > Yep. Hotplug events are pretty significant. Using stop_machine_XXXX() etc > > would be advisable and that would avoid the taking of locks and get rid of all the > > ocmplexity, reduce the code size and make the overall system much more > > reliable. > > Huch? stop_machine() is horrible and heavy weight. Don't go there, there > must be simpler solutions than that. Inserting or removing hardware is a heavy process. This would help quite a bit with these issues for loops over active cpus.
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2017-02-07 18:10 +0100 |
| Message-ID | <t8hrc-5Gz-15@gated-at.bofh.it> |
| In reply to | #1575807 |
Hello,
Sorry about the delay.
On Tue, Feb 07, 2017 at 04:34:59PM +0100, Michal Hocko wrote:
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index c3358d4f7932..b6411816787a 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -2343,7 +2343,16 @@ void drain_local_pages(struct zone *zone)
>
> static void drain_local_pages_wq(struct work_struct *work)
> {
> + /*
> + * drain_all_pages doesn't use proper cpu hotplug protection so
> + * we can race with cpu offline when the WQ can move this from
> + * a cpu pinned worker to an unbound one. We can operate on a different
> + * cpu which is allright but we also have to make sure to not move to
> + * a different one.
> + */
> + preempt_disable();
> drain_local_pages(NULL);
> + preempt_enable();
> }
>
> /*
> @@ -2379,12 +2388,6 @@ void drain_all_pages(struct zone *zone)
> }
>
> /*
> - * As this can be called from reclaim context, do not reenter reclaim.
> - * An allocation failure can be handled, it's simply slower
> - */
> - get_online_cpus();
> -
> - /*
> * We don't care about racing with CPU hotplug event
> * as offline notification will cause the notified
> * cpu to drain that CPU pcps and on_each_cpu_mask
> @@ -2423,7 +2426,6 @@ void drain_all_pages(struct zone *zone)
> for_each_cpu(cpu, &cpus_with_pcps)
> flush_work(per_cpu_ptr(&pcpu_drain, cpu));
>
> - put_online_cpus();
> mutex_unlock(&pcpu_drain_mutex);
I think this would work; however, a more canonical way would be
something along the line of...
drain_all_pages()
{
...
spin_lock();
for_each_possible_cpu() {
if (this cpu should get drained) {
queue_work_on(this cpu's work);
}
}
spin_unlock();
...
}
offline_hook()
{
spin_lock();
this cpu should get drained = false;
spin_unlock();
queue_work_on(this cpu's work);
flush_work(this cpu's work);
}
I think what workqueue should do is automatically flush in-flight CPU
work items on CPU offline and erroring out on queue_work_on() on
offline CPUs. And we now actually can do that because we have lifted
the guarantee that queue_work() is local CPU affine some releases ago.
I'll look into it soonish.
For the time being, either approach should be fine. The more
canonical one might be a bit less surprising but the
preempt_disable/enable() change is short and sweet and completely fine
for the case at hand.
Please feel free to add
Acked-by: Tejun Heo <tj@kernel.org>
Thanks.
--
tejun
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-02-07 21:40 +0100 |
| Message-ID | <t8kIq-7B3-29@gated-at.bofh.it> |
| In reply to | #1575893 |
On Tue 07-02-17 12:03:19, Tejun Heo wrote:
> Hello,
>
> Sorry about the delay.
>
> On Tue, Feb 07, 2017 at 04:34:59PM +0100, Michal Hocko wrote:
> > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > index c3358d4f7932..b6411816787a 100644
> > --- a/mm/page_alloc.c
> > +++ b/mm/page_alloc.c
> > @@ -2343,7 +2343,16 @@ void drain_local_pages(struct zone *zone)
> >
> > static void drain_local_pages_wq(struct work_struct *work)
> > {
> > + /*
> > + * drain_all_pages doesn't use proper cpu hotplug protection so
> > + * we can race with cpu offline when the WQ can move this from
> > + * a cpu pinned worker to an unbound one. We can operate on a different
> > + * cpu which is allright but we also have to make sure to not move to
> > + * a different one.
> > + */
> > + preempt_disable();
> > drain_local_pages(NULL);
> > + preempt_enable();
> > }
> >
> > /*
> > @@ -2379,12 +2388,6 @@ void drain_all_pages(struct zone *zone)
> > }
> >
> > /*
> > - * As this can be called from reclaim context, do not reenter reclaim.
> > - * An allocation failure can be handled, it's simply slower
> > - */
> > - get_online_cpus();
> > -
> > - /*
> > * We don't care about racing with CPU hotplug event
> > * as offline notification will cause the notified
> > * cpu to drain that CPU pcps and on_each_cpu_mask
> > @@ -2423,7 +2426,6 @@ void drain_all_pages(struct zone *zone)
> > for_each_cpu(cpu, &cpus_with_pcps)
> > flush_work(per_cpu_ptr(&pcpu_drain, cpu));
> >
> > - put_online_cpus();
> > mutex_unlock(&pcpu_drain_mutex);
>
> I think this would work; however, a more canonical way would be
> something along the line of...
>
> drain_all_pages()
> {
> ...
> spin_lock();
> for_each_possible_cpu() {
> if (this cpu should get drained) {
> queue_work_on(this cpu's work);
> }
> }
> spin_unlock();
> ...
> }
>
> offline_hook()
> {
> spin_lock();
> this cpu should get drained = false;
> spin_unlock();
> queue_work_on(this cpu's work);
> flush_work(this cpu's work);
> }
I see
> I think what workqueue should do is automatically flush in-flight CPU
> work items on CPU offline and erroring out on queue_work_on() on
> offline CPUs. And we now actually can do that because we have lifted
> the guarantee that queue_work() is local CPU affine some releases ago.
> I'll look into it soonish.
>
> For the time being, either approach should be fine. The more
> canonical one might be a bit less surprising but the
> preempt_disable/enable() change is short and sweet and completely fine
> for the case at hand.
Thanks for double checking!
> Please feel free to add
>
> Acked-by: Tejun Heo <tj@kernel.org>
Thanks!
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2017-02-07 14:10 +0100 |
| Message-ID | <t8dGZ-3hn-119@gated-at.bofh.it> |
| In reply to | #1575591 |
On Tue, Feb 07, 2017 at 12:43:27PM +0100, Michal Hocko wrote:
> > Right. The unbind operation can set a mask that is any allowable CPU and
> > the final process_work is not done in a context that prevents
> > preemption.
> >
> > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > index 3b93879990fd..7af165d308c4 100644
> > --- a/mm/page_alloc.c
> > +++ b/mm/page_alloc.c
> > @@ -2342,7 +2342,14 @@ void drain_local_pages(struct zone *zone)
> >
> > static void drain_local_pages_wq(struct work_struct *work)
> > {
> > + /*
> > + * Ordinarily a drain operation is bound to a CPU but may be unbound
> > + * after a CPU hotplug operation so it's necessary to disable
> > + * preemption for the drain to stabilise the CPU ID.
> > + */
> > + preempt_disable();
> > drain_local_pages(NULL);
> > + preempt_enable_no_resched();
> > }
> >
> > /*
> [...]
> > @@ -6711,7 +6714,16 @@ static int page_alloc_cpu_dead(unsigned int cpu)
> > {
> >
> > lru_add_drain_cpu(cpu);
> > +
> > + /*
> > + * A per-cpu drain via a workqueue from drain_all_pages can be
> > + * rescheduled onto an unrelated CPU. That allows the hotplug
> > + * operation and the drain to potentially race on the same
> > + * CPU. Serialise hotplug versus drain using pcpu_drain_mutex
> > + */
> > + mutex_lock(&pcpu_drain_mutex);
> > drain_pages(cpu);
> > + mutex_unlock(&pcpu_drain_mutex);
>
> You cannot put sleepable lock inside the preempt disbaled section...
> We can make it a spinlock right?
>
The CPU down callback can hold a mutex and at least he SLUB callback
already does so. That gives
page_alloc_cpu_dead
mutex_lock
drain_pages
mutex_unlock
drain_all_pages
mutex_lock
queue workqueue
drain_local_pages_wq
preempt_disable
drain_local_pages
drain_pages
preempt_enable
flush queues
mutex_unlock
I must be blind or maybe it's rushing between multiple concerns but which
sleepable lock is of concern?
--
Mel Gorman
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-02-07 14:50 +0100 |
| Message-ID | <t8ejE-3wo-27@gated-at.bofh.it> |
| In reply to | #1575713 |
On Tue 07-02-17 13:03:50, Mel Gorman wrote:
> On Tue, Feb 07, 2017 at 12:43:27PM +0100, Michal Hocko wrote:
> > > Right. The unbind operation can set a mask that is any allowable CPU and
> > > the final process_work is not done in a context that prevents
> > > preemption.
> > >
> > > diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> > > index 3b93879990fd..7af165d308c4 100644
> > > --- a/mm/page_alloc.c
> > > +++ b/mm/page_alloc.c
> > > @@ -2342,7 +2342,14 @@ void drain_local_pages(struct zone *zone)
> > >
> > > static void drain_local_pages_wq(struct work_struct *work)
> > > {
> > > + /*
> > > + * Ordinarily a drain operation is bound to a CPU but may be unbound
> > > + * after a CPU hotplug operation so it's necessary to disable
> > > + * preemption for the drain to stabilise the CPU ID.
> > > + */
> > > + preempt_disable();
> > > drain_local_pages(NULL);
> > > + preempt_enable_no_resched();
> > > }
> > >
> > > /*
> > [...]
> > > @@ -6711,7 +6714,16 @@ static int page_alloc_cpu_dead(unsigned int cpu)
> > > {
> > >
> > > lru_add_drain_cpu(cpu);
> > > +
> > > + /*
> > > + * A per-cpu drain via a workqueue from drain_all_pages can be
> > > + * rescheduled onto an unrelated CPU. That allows the hotplug
> > > + * operation and the drain to potentially race on the same
> > > + * CPU. Serialise hotplug versus drain using pcpu_drain_mutex
> > > + */
> > > + mutex_lock(&pcpu_drain_mutex);
> > > drain_pages(cpu);
> > > + mutex_unlock(&pcpu_drain_mutex);
> >
> > You cannot put sleepable lock inside the preempt disbaled section...
> > We can make it a spinlock right?
> >
>
> The CPU down callback can hold a mutex and at least he SLUB callback
> already does so. That gives
>
> page_alloc_cpu_dead
> mutex_lock
> drain_pages
> mutex_unlock
>
> drain_all_pages
> mutex_lock
> queue workqueue
> drain_local_pages_wq
> preempt_disable
> drain_local_pages
> drain_pages
> preempt_enable
> flush queues
> mutex_unlock
>
> I must be blind or maybe it's rushing between multiple concerns but which
> sleepable lock is of concern?
I thought the cpu hotplug callback was non-preemptible. This is not the
case as mentioned in other reply. The pcpu_drain_mutex in the hotplug
callback is alright. Sorry about the confusion! I am still wondering
whether the lock is really needed. See the other reply.
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp> |
|---|---|
| Date | 2017-02-07 12:30 +0100 |
| Message-ID | <t8c8a-2bv-1@gated-at.bofh.it> |
| In reply to | #1575189 |
On 2017/02/07 7:05, Mel Gorman wrote:
> On Mon, Feb 06, 2017 at 08:13:35PM +0100, Dmitry Vyukov wrote:
>> On Mon, Jan 30, 2017 at 4:48 PM, Dmitry Vyukov <dvyukov@google.com> wrote:
>>> On Sun, Jan 29, 2017 at 6:22 PM, Vlastimil Babka <vbabka@suse.cz> wrote:
>>>> On 29.1.2017 13:44, Dmitry Vyukov wrote:
>>>>> Hello,
>>>>>
>>>>> I've got the following deadlock report while running syzkaller fuzzer
>>>>> on f37208bc3c9c2f811460ef264909dfbc7f605a60:
>>>>>
>>>>> [ INFO: possible circular locking dependency detected ]
>>>>> 4.10.0-rc5-next-20170125 #1 Not tainted
>>>>> -------------------------------------------------------
>>>>> syz-executor3/14255 is trying to acquire lock:
>>>>> (cpu_hotplug.dep_map){++++++}, at: [<ffffffff814271c7>]
>>>>> get_online_cpus+0x37/0x90 kernel/cpu.c:239
>>>>>
>>>>> but task is already holding lock:
>>>>> (pcpu_alloc_mutex){+.+.+.}, at: [<ffffffff81937fee>]
>>>>> pcpu_alloc+0xbfe/0x1290 mm/percpu.c:897
>>>>>
>>>>> which lock already depends on the new lock.
>>>>
>>>> I suspect the dependency comes from recent changes in drain_all_pages(). They
>>>> were later redone (for other reasons, but nice to have another validation) in
>>>> the mmots patch [1], which AFAICS is not yet in mmotm and thus linux-next. Could
>>>> you try if it helps?
>>>
>>> It happened only once on linux-next, so I can't verify the fix. But I
>>> will watch out for other occurrences.
>>
>> Unfortunately it does not seem to help.
>
> I'm a little stuck on how to best handle this. get_online_cpus() can
> halt forever if the hotplug operation is holding the mutex when calling
> pcpu_alloc. One option would be to add a try_get_online_cpus() helper which
> trylocks the mutex. However, given that drain is so unlikely to actually
> make that make a difference when racing against parallel allocations,
> I think this should be acceptable.
>
> Any objections?
Why below change on top of current linux.git (8b1b41ee74f9) insufficient?
I think it can eliminate IPIs a lot without introducing lockdep warnings.
----------
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index f3e0c69..ae6e7aa 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -2354,6 +2354,7 @@ void drain_local_pages(struct zone *zone)
*/
void drain_all_pages(struct zone *zone)
{
+ static DEFINE_MUTEX(lock);
int cpu;
/*
@@ -2362,6 +2363,7 @@ void drain_all_pages(struct zone *zone)
*/
static cpumask_t cpus_with_pcps;
+ mutex_lock(&lock);
/*
* We don't care about racing with CPU hotplug event
* as offline notification will cause the notified
@@ -2394,6 +2396,7 @@ void drain_all_pages(struct zone *zone)
}
on_each_cpu_mask(&cpus_with_pcps, (smp_call_func_t) drain_local_pages,
zone, 1);
+ mutex_unlock(&lock);
}
#ifdef CONFIG_HIBERNATION
----------
By the way, drain_all_pages() is a sleepable context, isn't it?
I don't get get soft lockup using current linux.git (8b1b41ee74f9).
But I trivially get soft lockup if I try above change after reverting
"mm, page_alloc: use static global work_struct for draining per-cpu pages" and
"mm, page_alloc: drain per-cpu pages from workqueue context" on linux-next-20170207 .
List corruption also came in with linux-next-20170207 .
----------
[ 32.672890] ip6_tables: (C) 2000-2006 Netfilter Core Team
[ 33.109860] Ebtables v2.0 registered
[ 33.410293] nf_conntrack version 0.5.0 (16384 buckets, 65536 max)
[ 33.935512] IPv6: ADDRCONF(NETDEV_UP): eno16777728: link is not ready
[ 33.937478] e1000: eno16777728 NIC Link is Up 1000 Mbps Full Duplex, Flow Control: None
[ 33.939777] IPv6: ADDRCONF(NETDEV_CHANGE): eno16777728: link becomes ready
[ 34.194000] Netfilter messages via NETLINK v0.30.
[ 34.258828] ip_set: protocol 6
[ 38.518375] nf_conntrack: default automatic helper assignment has been turned off for security reasons and CT-based firewall rule not found. Use the iptables CT target to attach helpers instead.
[ 43.911609] cp (5167) used greatest stack depth: 10488 bytes left
[ 48.174125] cp (5860) used greatest stack depth: 10224 bytes left
[ 101.075280] ------------[ cut here ]------------
[ 101.076743] WARNING: CPU: 0 PID: 9766 at lib/list_debug.c:25 __list_add_valid+0x46/0xa0
[ 101.079095] list_add corruption. next->prev should be prev (ffffea00013dfce0), but was ffff88006d3e1d40. (next=ffff88006d3e1d40).
[ 101.081863] Modules linked in: nf_conntrack_netbios_ns nf_conntrack_broadcast ip6t_rpfilter ipt_REJECT nf_reject_ipv4 ip6t_REJECT nf_reject_ipv6 xt_conntrack ip_set nfnetlink ebtable_nat ebtable_broute bridge stp llc ip6table_nat nf_conntrack_ipv6 nf_defrag_ipv6 nf_nat_ipv6 ip6table_mangle ip6table_raw iptable_nat nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 nf_nat nf_conntrack iptable_mangle iptable_raw ebtable_filter ebtables ip6table_filter ip6_tables iptable_filter coretemp crct10dif_pclmul crc32_pclmul ghash_clmulni_intel aesni_intel crypto_simd cryptd ppdev glue_helper vmw_balloon pcspkr sg parport_pc parport i2c_piix4 shpchp vmw_vsock_vmci_transport vsock vmw_vmci ip_tables xfs libcrc32c sd_mod sr_mod cdrom ata_generic pata_acpi crc32c_intel vmwgfx serio_raw drm_kms_helper syscopyarea sysfillrect
[ 101.098346] sysimgblt fb_sys_fops ttm mptspi scsi_transport_spi ata_piix mptscsih ahci drm libahci mptbase libata e1000 i2c_core
[ 101.101279] CPU: 0 PID: 9766 Comm: oom-write Tainted: G W 4.10.0-rc7-next-20170207+ #500
[ 101.103519] Hardware name: VMware, Inc. VMware Virtual Platform/440BX Desktop Reference Platform, BIOS 6.00 07/02/2015
[ 101.106031] Call Trace:
[ 101.107069] dump_stack+0x85/0xc9
[ 101.108239] __warn+0xd1/0xf0
[ 101.109566] warn_slowpath_fmt+0x5f/0x80
[ 101.110843] __list_add_valid+0x46/0xa0
[ 101.112187] free_hot_cold_page+0x205/0x460
[ 101.113593] free_hot_cold_page_list+0x3c/0x1c0
[ 101.115046] shrink_page_list+0x4dd/0xd10
[ 101.116388] shrink_inactive_list+0x1c5/0x660
[ 101.117796] shrink_node_memcg+0x535/0x7f0
[ 101.119158] ? mem_cgroup_iter+0x1e0/0x720
[ 101.120873] shrink_node+0xe1/0x310
[ 101.122250] do_try_to_free_pages+0xe1/0x300
[ 101.123611] try_to_free_pages+0x131/0x3f0
[ 101.124998] __alloc_pages_slowpath+0x479/0xe32
[ 101.126421] __alloc_pages_nodemask+0x382/0x3d0
[ 101.128053] alloc_pages_vma+0xae/0x2f0
[ 101.129353] do_anonymous_page+0x111/0x5d0
[ 101.130725] __handle_mm_fault+0xbc9/0xeb0
[ 101.132127] ? sched_clock+0x9/0x10
[ 101.133389] ? sched_clock_cpu+0x11/0xc0
[ 101.134764] handle_mm_fault+0x16b/0x390
[ 101.136147] ? handle_mm_fault+0x49/0x390
[ 101.137545] __do_page_fault+0x24a/0x530
[ 101.138999] do_page_fault+0x30/0x80
[ 101.140375] page_fault+0x28/0x30
[ 101.141685] RIP: 0033:0x4006a0
[ 101.142880] RSP: 002b:00007ffe3a4acc10 EFLAGS: 00010206
[ 101.144638] RAX: 0000000037f8e000 RBX: 0000000080000000 RCX: 00007fc5a1fda650
[ 101.146653] RDX: 0000000000000000 RSI: 00007ffe3a4aca30 RDI: 00007ffe3a4aca30
[ 101.148768] RBP: 00007fc4a2111010 R08: 00007ffe3a4acb40 R09: 00007ffe3a4ac980
[ 101.150893] R10: 0000000000000008 R11: 0000000000000246 R12: 0000000000000007
[ 101.152937] R13: 00007fc4a2111010 R14: 0000000000000000 R15: 0000000000000000
[ 101.155240] ---[ end trace 5d8b63572ab78be3 ]---
[ 128.945470] NMI watchdog: BUG: soft lockup - CPU#2 stuck for 23s! [oom-write:9766]
[ 128.947878] Modules linked in: nf_conntrack_netbios_ns nf_conntrack_broadcast ip6t_rpfilter ipt_REJECT nf_reject_ipv4 ip6t_REJECT nf_reject_ipv6 xt_conntrack ip_set nfnetlink ebtable_nat ebtable_broute bridge stp llc ip6table_nat nf_conntrack_ipv6 nf_defrag_ipv6 nf_nat_ipv6 ip6table_mangle ip6table_raw iptable_nat nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 nf_nat nf_conntrack iptable_mangle iptable_raw ebtable_filter ebtables ip6table_filter ip6_tables iptable_filter coretemp crct10dif_pclmul crc32_pclmul ghash_clmulni_intel aesni_intel crypto_simd cryptd ppdev glue_helper vmw_balloon pcspkr sg parport_pc parport i2c_piix4 shpchp vmw_vsock_vmci_transport vsock vmw_vmci ip_tables xfs libcrc32c sd_mod sr_mod cdrom ata_generic pata_acpi crc32c_intel vmwgfx serio_raw drm_kms_helper syscopyarea sysfillrect
[ 128.966961] sysimgblt fb_sys_fops ttm mptspi scsi_transport_spi ata_piix mptscsih ahci drm libahci mptbase libata e1000 i2c_core
[ 128.970135] irq event stamp: 2491700
[ 128.971666] hardirqs last enabled at (2491699): [<ffffffff817e5d70>] restore_regs_and_iret+0x0/0x1d
[ 128.974560] hardirqs last disabled at (2491700): [<ffffffff817e7198>] apic_timer_interrupt+0x98/0xb0
[ 128.977148] softirqs last enabled at (2445236): [<ffffffff817eab39>] __do_softirq+0x349/0x52d
[ 128.979639] softirqs last disabled at (2445215): [<ffffffff810a99e5>] irq_exit+0xf5/0x110
[ 128.982019] CPU: 2 PID: 9766 Comm: oom-write Tainted: G W 4.10.0-rc7-next-20170207+ #500
[ 128.984598] Hardware name: VMware, Inc. VMware Virtual Platform/440BX Desktop Reference Platform, BIOS 6.00 07/02/2015
[ 128.987490] task: ffff880067ed8040 task.stack: ffffc9001117c000
[ 128.989464] RIP: 0010:smp_call_function_many+0x25c/0x320
[ 128.991305] RSP: 0000:ffffc9001117fad0 EFLAGS: 00000202 ORIG_RAX: ffffffffffffff10
[ 128.993609] RAX: 0000000000000000 RBX: ffff88006d7dd640 RCX: 0000000000000001
[ 128.995823] RDX: 0000000000000001 RSI: ffff88006d3e3398 RDI: ffff88006c528dc8
[ 128.998058] RBP: ffffc9001117fb18 R08: 0000000000000009 R09: 0000000000000000
[ 129.000284] R10: 0000000000000001 R11: ffff88005486d4a8 R12: 0000000000000000
[ 129.002521] R13: ffffffff811fe6e0 R14: 0000000000000000 R15: 0000000000000080
[ 129.004788] FS: 00007fc5a24ec740(0000) GS:ffff88006d600000(0000) knlGS:0000000000000000
[ 129.007235] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 129.009211] CR2: 00007fc4da0c5010 CR3: 0000000044314000 CR4: 00000000001406e0
[ 129.011641] Call Trace:
[ 129.012993] ? page_alloc_cpu_dead+0x30/0x30
[ 129.014701] on_each_cpu_mask+0x30/0xb0
[ 129.016318] drain_all_pages+0x113/0x170
[ 129.017949] __alloc_pages_slowpath+0x520/0xe32
[ 129.019710] __alloc_pages_nodemask+0x382/0x3d0
[ 129.021472] alloc_pages_vma+0xae/0x2f0
[ 129.023080] do_anonymous_page+0x111/0x5d0
[ 129.024760] __handle_mm_fault+0xbc9/0xeb0
[ 129.026449] ? sched_clock+0x9/0x10
[ 129.028042] ? sched_clock_cpu+0x11/0xc0
[ 129.029729] handle_mm_fault+0x16b/0x390
[ 129.031418] ? handle_mm_fault+0x49/0x390
[ 129.033077] __do_page_fault+0x24a/0x530
[ 129.034721] do_page_fault+0x30/0x80
[ 129.036279] page_fault+0x28/0x30
[ 129.038160] RIP: 0033:0x4006a0
[ 129.039600] RSP: 002b:00007ffe3a4acc10 EFLAGS: 00010206
[ 129.041488] RAX: 0000000037fb4000 RBX: 0000000080000000 RCX: 00007fc5a1fda650
[ 129.043758] RDX: 0000000000000000 RSI: 00007ffe3a4aca30 RDI: 00007ffe3a4aca30
[ 129.046037] RBP: 00007fc4a2111010 R08: 00007ffe3a4acb40 R09: 00007ffe3a4ac980
[ 129.048324] R10: 0000000000000008 R11: 0000000000000246 R12: 0000000000000007
[ 129.050587] R13: 00007fc4a2111010 R14: 0000000000000000 R15: 0000000000000000
[ 129.052829] Code: 7b 3e 2b 00 3b 05 69 b6 d5 00 41 89 c4 0f 8d 3f fe ff ff 48 63 d0 48 8b 33 48 03 34 d5 60 c4 ab 81 8b 56 18 83 e2 01 74 0a f3 90 <8b> 4e 18 83 e1 01 75 f6 83 f8 ff 48 8b 7b 08 74 2a 48 63 35 30
[ 156.945366] NMI watchdog: BUG: soft lockup - CPU#2 stuck for 23s! [oom-write:9766]
[ 156.947797] Modules linked in: nf_conntrack_netbios_ns nf_conntrack_broadcast ip6t_rpfilter ipt_REJECT nf_reject_ipv4 ip6t_REJECT nf_reject_ipv6 xt_conntrack ip_set nfnetlink ebtable_nat ebtable_broute bridge stp llc ip6table_nat nf_conntrack_ipv6 nf_defrag_ipv6 nf_nat_ipv6 ip6table_mangle ip6table_raw iptable_nat nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 nf_nat nf_conntrack iptable_mangle iptable_raw ebtable_filter ebtables ip6table_filter ip6_tables iptable_filter coretemp crct10dif_pclmul crc32_pclmul ghash_clmulni_intel aesni_intel crypto_simd cryptd ppdev glue_helper vmw_balloon pcspkr sg parport_pc parport i2c_piix4 shpchp vmw_vsock_vmci_transport vsock vmw_vmci ip_tables xfs libcrc32c sd_mod sr_mod cdrom ata_generic pata_acpi crc32c_intel vmwgfx serio_raw drm_kms_helper syscopyarea sysfillrect
[ 156.968021] sysimgblt fb_sys_fops ttm mptspi scsi_transport_spi ata_piix mptscsih ahci drm libahci mptbase libata e1000 i2c_core
[ 156.971269] irq event stamp: 2547400
[ 156.972865] hardirqs last enabled at (2547399): [<ffffffff817e5d70>] restore_regs_and_iret+0x0/0x1d
[ 156.975579] hardirqs last disabled at (2547400): [<ffffffff817e7198>] apic_timer_interrupt+0x98/0xb0
[ 156.978287] softirqs last enabled at (2445236): [<ffffffff817eab39>] __do_softirq+0x349/0x52d
[ 156.980863] softirqs last disabled at (2445215): [<ffffffff810a99e5>] irq_exit+0xf5/0x110
[ 156.983393] CPU: 2 PID: 9766 Comm: oom-write Tainted: G W L 4.10.0-rc7-next-20170207+ #500
[ 156.986111] Hardware name: VMware, Inc. VMware Virtual Platform/440BX Desktop Reference Platform, BIOS 6.00 07/02/2015
[ 156.989117] task: ffff880067ed8040 task.stack: ffffc9001117c000
[ 156.991179] RIP: 0010:smp_call_function_many+0x25c/0x320
[ 156.993113] RSP: 0000:ffffc9001117fad0 EFLAGS: 00000202 ORIG_RAX: ffffffffffffff10
[ 156.995501] RAX: 0000000000000000 RBX: ffff88006d7dd640 RCX: 0000000000000001
[ 156.997803] RDX: 0000000000000001 RSI: ffff88006d3e3398 RDI: ffff88006c528dc8
[ 157.000111] RBP: ffffc9001117fb18 R08: 0000000000000009 R09: 0000000000000000
[ 157.002402] R10: 0000000000000001 R11: ffff88005486d4a8 R12: 0000000000000000
[ 157.004700] R13: ffffffff811fe6e0 R14: 0000000000000000 R15: 0000000000000080
[ 157.006980] FS: 00007fc5a24ec740(0000) GS:ffff88006d600000(0000) knlGS:0000000000000000
[ 157.009477] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 157.011494] CR2: 00007fc4da0c5010 CR3: 0000000044314000 CR4: 00000000001406e0
[ 157.013844] Call Trace:
[ 157.015231] ? page_alloc_cpu_dead+0x30/0x30
[ 157.016975] on_each_cpu_mask+0x30/0xb0
[ 157.018723] drain_all_pages+0x113/0x170
[ 157.020410] __alloc_pages_slowpath+0x520/0xe32
[ 157.022219] __alloc_pages_nodemask+0x382/0x3d0
[ 157.024036] alloc_pages_vma+0xae/0x2f0
[ 157.025704] do_anonymous_page+0x111/0x5d0
[ 157.027430] __handle_mm_fault+0xbc9/0xeb0
[ 157.029143] ? sched_clock+0x9/0x10
[ 157.030739] ? sched_clock_cpu+0x11/0xc0
[ 157.032416] handle_mm_fault+0x16b/0x390
[ 157.034100] ? handle_mm_fault+0x49/0x390
[ 157.035843] __do_page_fault+0x24a/0x530
[ 157.037537] do_page_fault+0x30/0x80
[ 157.039160] page_fault+0x28/0x30
[ 157.040792] RIP: 0033:0x4006a0
[ 157.042290] RSP: 002b:00007ffe3a4acc10 EFLAGS: 00010206
[ 157.044231] RAX: 0000000037fb4000 RBX: 0000000080000000 RCX: 00007fc5a1fda650
[ 157.046635] RDX: 0000000000000000 RSI: 00007ffe3a4aca30 RDI: 00007ffe3a4aca30
[ 157.048941] RBP: 00007fc4a2111010 R08: 00007ffe3a4acb40 R09: 00007ffe3a4ac980
[ 157.051324] R10: 0000000000000008 R11: 0000000000000246 R12: 0000000000000007
[ 157.053577] R13: 00007fc4a2111010 R14: 0000000000000000 R15: 0000000000000000
[ 157.055822] Code: 7b 3e 2b 00 3b 05 69 b6 d5 00 41 89 c4 0f 8d 3f fe ff ff 48 63 d0 48 8b 33 48 03 34 d5 60 c4 ab 81 8b 56 18 83 e2 01 74 0a f3 90 <8b> 4e 18 83 e1 01 75 f6 83 f8 ff 48 8b 7b 08 74 2a 48 63 35 30
[ 171.241423] BUG: spinlock lockup suspected on CPU#3, swapper/3/0
[ 171.243575] lock: 0xffff88007ffddd00, .magic: dead4ead, .owner: kworker/0:3/9777, .owner_cpu: 0
[ 171.246136] CPU: 3 PID: 0 Comm: swapper/3 Tainted: G W L 4.10.0-rc7-next-20170207+ #500
[ 171.248706] Hardware name: VMware, Inc. VMware Virtual Platform/440BX Desktop Reference Platform, BIOS 6.00 07/02/2015
[ 171.251617] Call Trace:
[ 171.252854] <IRQ>
[ 171.253989] dump_stack+0x85/0xc9
[ 171.255376] spin_dump+0x90/0x95
[ 171.256741] do_raw_spin_lock+0x9a/0x130
[ 171.258246] _raw_spin_lock_irqsave+0x75/0x90
[ 171.259822] ? free_pcppages_bulk+0x37/0x910
[ 171.261359] free_pcppages_bulk+0x37/0x910
[ 171.262843] ? sched_clock_cpu+0x11/0xc0
[ 171.264289] ? sched_clock_tick+0x2d/0x80
[ 171.265747] drain_pages_zone+0x82/0x90
[ 171.267154] ? page_alloc_cpu_dead+0x30/0x30
[ 171.268644] drain_pages+0x3f/0x60
[ 171.269955] drain_local_pages+0x25/0x30
[ 171.271366] flush_smp_call_function_queue+0x7b/0x170
[ 171.273014] generic_smp_call_function_single_interrupt+0x13/0x30
[ 171.274878] smp_call_function_interrupt+0x27/0x40
[ 171.276460] call_function_interrupt+0x9d/0xb0
[ 171.277963] RIP: 0010:native_safe_halt+0x6/0x10
[ 171.279467] RSP: 0018:ffffc900003a3e70 EFLAGS: 00000206 ORIG_RAX: ffffffffffffff03
[ 171.281610] RAX: ffff88006c5c0040 RBX: 0000000000000000 RCX: 0000000000000000
[ 171.283662] RDX: ffff88006c5c0040 RSI: 0000000000000001 RDI: ffff88006c5c0040
[ 171.285727] RBP: ffffc900003a3e70 R08: 0000000000000000 R09: 0000000000000000
[ 171.287778] R10: 0000000000000001 R11: 0000000000000001 R12: 0000000000000003
[ 171.289831] R13: ffff88006c5c0040 R14: ffff88006c5c0040 R15: 0000000000000000
[ 171.291882] </IRQ>
[ 171.292934] default_idle+0x23/0x1d0
[ 171.294287] arch_cpu_idle+0xf/0x20
[ 171.295618] default_idle_call+0x23/0x40
[ 171.297040] do_idle+0x162/0x230
[ 171.298308] cpu_startup_entry+0x71/0x80
[ 171.299732] start_secondary+0x17f/0x1f0
[ 171.301146] start_cpu+0x14/0x14
[ 171.302432] NMI backtrace for cpu 3
[ 171.303764] CPU: 3 PID: 0 Comm: swapper/3 Tainted: G W L 4.10.0-rc7-next-20170207+ #500
[ 171.306218] Hardware name: VMware, Inc. VMware Virtual Platform/440BX Desktop Reference Platform, BIOS 6.00 07/02/2015
[ 171.309065] Call Trace:
[ 171.310500] <IRQ>
[ 171.311569] dump_stack+0x85/0xc9
[ 171.312946] nmi_cpu_backtrace+0xc0/0xe0
[ 171.314396] ? irq_force_complete_move+0x170/0x170
[ 171.316012] nmi_trigger_cpumask_backtrace+0x12a/0x188
[ 171.317679] arch_trigger_cpumask_backtrace+0x19/0x20
[ 171.319294] do_raw_spin_lock+0xa8/0x130
[ 171.320686] _raw_spin_lock_irqsave+0x75/0x90
[ 171.322148] ? free_pcppages_bulk+0x37/0x910
[ 171.323585] free_pcppages_bulk+0x37/0x910
[ 171.325037] ? sched_clock_cpu+0x11/0xc0
[ 171.326615] ? sched_clock_tick+0x2d/0x80
[ 171.327974] drain_pages_zone+0x82/0x90
[ 171.329294] ? page_alloc_cpu_dead+0x30/0x30
[ 171.330707] drain_pages+0x3f/0x60
[ 171.331938] drain_local_pages+0x25/0x30
[ 171.333283] flush_smp_call_function_queue+0x7b/0x170
[ 171.334855] generic_smp_call_function_single_interrupt+0x13/0x30
[ 171.336643] smp_call_function_interrupt+0x27/0x40
[ 171.338456] call_function_interrupt+0x9d/0xb0
[ 171.339911] RIP: 0010:native_safe_halt+0x6/0x10
[ 171.341522] RSP: 0018:ffffc900003a3e70 EFLAGS: 00000206 ORIG_RAX: ffffffffffffff03
[ 171.343860] RAX: ffff88006c5c0040 RBX: 0000000000000000 RCX: 0000000000000000
[ 171.346001] RDX: ffff88006c5c0040 RSI: 0000000000000001 RDI: ffff88006c5c0040
[ 171.348134] RBP: ffffc900003a3e70 R08: 0000000000000000 R09: 0000000000000000
[ 171.350270] R10: 0000000000000001 R11: 0000000000000001 R12: 0000000000000003
[ 171.352384] R13: ffff88006c5c0040 R14: ffff88006c5c0040 R15: 0000000000000000
[ 171.354503] </IRQ>
[ 171.355616] default_idle+0x23/0x1d0
[ 171.357055] arch_cpu_idle+0xf/0x20
[ 171.361577] default_idle_call+0x23/0x40
[ 171.364329] do_idle+0x162/0x230
[ 171.365799] cpu_startup_entry+0x71/0x80
[ 171.367172] start_secondary+0x17f/0x1f0
[ 171.368530] start_cpu+0x14/0x14
[ 171.369748] Sending NMI from CPU 3 to CPUs 0-2:
[ 171.371367] NMI backtrace for cpu 2
[ 171.371368] CPU: 2 PID: 9766 Comm: oom-write Tainted: G W L 4.10.0-rc7-next-20170207+ #500
[ 171.371368] Hardware name: VMware, Inc. VMware Virtual Platform/440BX Desktop Reference Platform, BIOS 6.00 07/02/2015
[ 171.371369] task: ffff880067ed8040 task.stack: ffffc9001117c000
[ 171.371369] RIP: 0010:smp_call_function_many+0x25c/0x320
[ 171.371370] RSP: 0000:ffffc9001117fad0 EFLAGS: 00000202
[ 171.371371] RAX: 0000000000000000 RBX: ffff88006d7dd640 RCX: 0000000000000001
[ 171.371372] RDX: 0000000000000001 RSI: ffff88006d3e3398 RDI: ffff88006c528dc8
[ 171.371372] RBP: ffffc9001117fb18 R08: 0000000000000009 R09: 0000000000000000
[ 171.371372] R10: 0000000000000001 R11: ffff88005486d4a8 R12: 0000000000000000
[ 171.371373] R13: ffffffff811fe6e0 R14: 0000000000000000 R15: 0000000000000080
[ 171.371373] FS: 00007fc5a24ec740(0000) GS:ffff88006d600000(0000) knlGS:0000000000000000
[ 171.371374] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 171.371374] CR2: 00007fc4da0c5010 CR3: 0000000044314000 CR4: 00000000001406e0
[ 171.371375] Call Trace:
[ 171.371375] ? page_alloc_cpu_dead+0x30/0x30
[ 171.371375] on_each_cpu_mask+0x30/0xb0
[ 171.371376] drain_all_pages+0x113/0x170
[ 171.371376] __alloc_pages_slowpath+0x520/0xe32
[ 171.371377] __alloc_pages_nodemask+0x382/0x3d0
[ 171.371377] alloc_pages_vma+0xae/0x2f0
[ 171.371377] do_anonymous_page+0x111/0x5d0
[ 171.371378] __handle_mm_fault+0xbc9/0xeb0
[ 171.371378] ? sched_clock+0x9/0x10
[ 171.371379] ? sched_clock_cpu+0x11/0xc0
[ 171.371379] handle_mm_fault+0x16b/0x390
[ 171.371379] ? handle_mm_fault+0x49/0x390
[ 171.371380] __do_page_fault+0x24a/0x530
[ 171.371380] do_page_fault+0x30/0x80
[ 171.371381] page_fault+0x28/0x30
[ 171.371381] RIP: 0033:0x4006a0
[ 171.371381] RSP: 002b:00007ffe3a4acc10 EFLAGS: 00010206
[ 171.371382] RAX: 0000000037fb4000 RBX: 0000000080000000 RCX: 00007fc5a1fda650
[ 171.371383] RDX: 0000000000000000 RSI: 00007ffe3a4aca30 RDI: 00007ffe3a4aca30
[ 171.371383] RBP: 00007fc4a2111010 R08: 00007ffe3a4acb40 R09: 00007ffe3a4ac980
[ 171.371384] R10: 0000000000000008 R11: 0000000000000246 R12: 0000000000000007
[ 171.371384] R13: 00007fc4a2111010 R14: 0000000000000000 R15: 0000000000000000
[ 171.371385] Code: 7b 3e 2b 00 3b 05 69 b6 d5 00 41 89 c4 0f 8d 3f fe ff ff 48 63 d0 48 8b 33 48 03 34 d5 60 c4 ab 81 8b 56 18 83 e2 01 74 0a f3 90 <8b> 4e 18 83 e1 01 75 f6 83 f8 ff 48 8b 7b 08 74 2a 48 63 35 30
[ 171.372317] NMI backtrace for cpu 0
[ 171.372317] CPU: 0 PID: 9777 Comm: kworker/0:3 Tainted: G W L 4.10.0-rc7-next-20170207+ #500
[ 171.372318] Hardware name: VMware, Inc. VMware Virtual Platform/440BX Desktop Reference Platform, BIOS 6.00 07/02/2015
[ 171.372318] Workqueue: events vmw_fb_dirty_flush [vmwgfx]
[ 171.372319] task: ffff880064082540 task.stack: ffffc90013b54000
[ 171.372320] RIP: 0010:free_pcppages_bulk+0xbb/0x910
[ 171.372320] RSP: 0000:ffff88006d203eb0 EFLAGS: 00000002
[ 171.372321] RAX: 0000000000000000 RBX: 0000000000000000 RCX: 0000000000000010
[ 171.372322] RDX: 0000000000000000 RSI: 00000000357cc759 RDI: ffff88006d3e1d50
[ 171.372322] RBP: ffff88006d203f38 R08: ffff88006d3e1d20 R09: 0000000000000002
[ 171.372323] R10: 0000000000000000 R11: 000000000004f7f0 R12: ffff88007ffdd8f8
[ 171.372323] R13: ffffea00013dfc20 R14: ffffea00013dfc00 R15: ffff88007ffdd740
[ 171.372324] FS: 0000000000000000(0000) GS:ffff88006d200000(0000) knlGS:0000000000000000
[ 171.372324] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 171.372325] CR2: 00007fc4da09f010 CR3: 0000000044314000 CR4: 00000000001406f0
[ 171.372325] Call Trace:
[ 171.372325] <IRQ>
[ 171.372326] ? trace_hardirqs_off+0xd/0x10
[ 171.372326] drain_pages_zone+0x82/0x90
[ 171.372327] ? page_alloc_cpu_dead+0x30/0x30
[ 171.372327] drain_pages+0x3f/0x60
[ 171.372327] drain_local_pages+0x25/0x30
[ 171.372328] flush_smp_call_function_queue+0x7b/0x170
[ 171.372328] generic_smp_call_function_single_interrupt+0x13/0x30
[ 171.372329] smp_call_function_interrupt+0x27/0x40
[ 171.372329] call_function_interrupt+0x9d/0xb0
[ 171.372330] RIP: 0010:memcpy_orig+0x19/0x110
[ 171.372330] RSP: 0000:ffffc90013b57d98 EFLAGS: 00000202 ORIG_RAX: ffffffffffffff03
[ 171.372331] RAX: ffffc90003f1c400 RBX: ffff880063b54a88 RCX: 0000000000000004
[ 171.372331] RDX: 00000000000010e0 RSI: ffffc900037d3900 RDI: ffffc90003f1c6e0
[ 171.372332] RBP: ffffc90013b57e08 R08: 0000000000000000 R09: 0000000000000000
[ 171.372332] R10: 0000000000000000 R11: 0000000000000000 R12: 0000000000001400
[ 171.372333] R13: 000000000000011e R14: ffff880063b55168 R15: ffffc900037d3620
[ 171.372333] </IRQ>
[ 171.372333] ? vmw_fb_dirty_flush+0x1ef/0x2b0 [vmwgfx]
[ 171.372334] process_one_work+0x22b/0x760
[ 171.372334] ? process_one_work+0x194/0x760
[ 171.372335] worker_thread+0x137/0x4b0
[ 171.372335] kthread+0x10f/0x150
[ 171.372335] ? process_one_work+0x760/0x760
[ 171.372336] ? kthread_create_on_node+0x70/0x70
[ 171.372336] ? do_syscall_64+0x6c/0x200
[ 171.372337] ret_from_fork+0x31/0x40
[ 171.372337] Code: 00 00 4d 89 cf 8b 45 98 8b 75 9c 31 d2 4c 8b 85 78 ff ff ff 83 c0 01 83 c6 01 83 f8 03 0f 44 c2 48 63 c8 48 83 c1 01 48 c1 e1 04 <4c> 01 c1 48 8b 39 48 39 f9 74 de 89 45 98 83 fe 03 89 f0 0f 44
[ 171.372357] NMI backtrace for cpu 1
[ 171.372357] CPU: 1 PID: 47 Comm: khugepaged Tainted: G W L 4.10.0-rc7-next-20170207+ #500
[ 171.372358] Hardware name: VMware, Inc. VMware Virtual Platform/440BX Desktop Reference Platform, BIOS 6.00 07/02/2015
[ 171.372358] task: ffff88006bd22540 task.stack: ffffc90000890000
[ 171.372359] RIP: 0010:delay_tsc+0x48/0x70
[ 171.372359] RSP: 0018:ffffc90000893520 EFLAGS: 00000006
[ 171.372360] RAX: 000000004e0ee6be RBX: ffff88007ffddd00 RCX: 000000774e0ee6a6
[ 171.372360] RDX: 0000000000000077 RSI: 0000000000000001 RDI: 0000000000000001
[ 171.372361] RBP: ffffc90000893520 R08: 0000000000000000 R09: 0000000000000000
[ 171.372361] R10: 0000000000000000 R11: 0000000000000000 R12: 00000000a6822110
[ 171.372362] R13: 0000000083093e7f R14: 0000000000000001 R15: ffffea000137a020
[ 171.372362] FS: 0000000000000000(0000) GS:ffff88006d400000(0000) knlGS:0000000000000000
[ 171.372363] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 171.372363] CR2: 00007ff5aeb82000 CR3: 0000000068093000 CR4: 00000000001406e0
[ 171.372363] Call Trace:
[ 171.372364] __delay+0xf/0x20
[ 171.372364] do_raw_spin_lock+0x86/0x130
[ 171.372365] _raw_spin_lock_irqsave+0x75/0x90
[ 171.372365] ? free_pcppages_bulk+0x37/0x910
[ 171.372365] free_pcppages_bulk+0x37/0x910
[ 171.372366] ? __kernel_map_pages+0x87/0x120
[ 171.372366] free_hot_cold_page+0x373/0x460
[ 171.372367] free_hot_cold_page_list+0x3c/0x1c0
[ 171.372367] shrink_page_list+0x4dd/0xd10
[ 171.372368] shrink_inactive_list+0x1c5/0x660
[ 171.372368] shrink_node_memcg+0x535/0x7f0
[ 171.372368] ? mem_cgroup_iter+0x14d/0x720
[ 171.372369] shrink_node+0xe1/0x310
[ 171.372369] do_try_to_free_pages+0xe1/0x300
[ 171.372370] try_to_free_pages+0x131/0x3f0
[ 171.372370] __alloc_pages_slowpath+0x479/0xe32
[ 171.372370] __alloc_pages_nodemask+0x382/0x3d0
[ 171.372371] khugepaged_alloc_page+0x6d/0xd0
[ 171.372371] collapse_huge_page+0x81/0x1240
[ 171.372372] ? sched_clock_cpu+0x11/0xc0
[ 171.372372] ? khugepaged_scan_mm_slot+0xc26/0x1000
[ 171.372373] khugepaged_scan_mm_slot+0xc49/0x1000
[ 171.372373] ? sched_clock_cpu+0x11/0xc0
[ 171.372373] ? finish_wait+0x75/0x90
[ 171.372374] khugepaged+0x327/0x5e0
[ 171.372374] ? remove_wait_queue+0x60/0x60
[ 171.372375] kthread+0x10f/0x150
[ 171.372375] ? khugepaged_scan_mm_slot+0x1000/0x1000
[ 171.372375] ? kthread_create_on_node+0x70/0x70
[ 171.372376] ret_from_fork+0x31/0x40
[ 171.372376] Code: 89 d1 48 c1 e1 20 48 09 c1 eb 1b 65 ff 0d c9 0a c1 7e f3 90 65 ff 05 c0 0a c1 7e 65 8b 05 51 d7 c0 7e 39 c6 75 20 0f ae e8 0f 31 <48> c1 e2 20 48 09 c2 48 89 d0 48 29 c8 48 39 f8 72 ce 65 ff 0d
----------
I also got soft lockup without any change using linux-next-20170202.
Something is wrong with calling multiple CPUs? List corruption?
----------
[ 80.556598] ip6_tables: (C) 2000-2006 Netfilter Core Team
[ 84.020374] IPv6: ADDRCONF(NETDEV_UP): eno16777728: link is not ready
[ 84.024423] e1000: eno16777728 NIC Link is Up 1000 Mbps Full Duplex, Flow Control: None
[ 84.030731] IPv6: ADDRCONF(NETDEV_CHANGE): eno16777728: link becomes ready
[ 84.212613] Ebtables v2.0 registered
[ 86.841246] nf_conntrack version 0.5.0 (16384 buckets, 65536 max)
[ 90.594709] Netfilter messages via NETLINK v0.30.
[ 90.756119] ip_set: protocol 6
[ 161.309259] NMI watchdog: BUG: soft lockup - CPU#3 stuck for 23s! [ip6tables-resto:4210]
[ 162.329982] Modules linked in: nf_conntrack_netbios_ns nf_conntrack_broadcast ip6t_rpfilter ipt_REJECT nf_reject_ipv4 ip6t_REJECT nf_reject_ipv6 xt_conntrack ip_set nfnetlink ebtable_nat ebtable_broute bridge stp llc ip6table_nat nf_conntrack_ipv6 nf_defrag_ipv6 nf_nat_ipv6 ip6table_mangle ip6table_raw iptable_nat nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 nf_nat nf_conntrack iptable_mangle iptable_raw ebtable_filter ebtables ip6table_filter ip6_tables iptable_filter coretemp crct10dif_pclmul crc32_pclmul ghash_clmulni_intel aesni_intel crypto_simd cryptd ppdev glue_helper vmw_balloon pcspkr sg parport_pc parport i2c_piix4 shpchp vmw_vsock_vmci_transport vsock vmw_vmci ip_tables xfs libcrc32c sd_mod sr_mod cdrom ata_generic pata_acpi crc32c_intel serio_raw vmwgfx drm_kms_helper syscopyarea sysfillrect
[ 162.486639] sysimgblt fb_sys_fops ttm e1000 mptspi ahci scsi_transport_spi drm libahci ata_piix mptscsih i2c_core mptbase libata
[ 162.486651] irq event stamp: 306010
[ 162.486656] hardirqs last enabled at (306009): [<ffffffff817e4970>] restore_regs_and_iret+0x0/0x1d
[ 162.486658] hardirqs last disabled at (306010): [<ffffffff817e5d98>] apic_timer_interrupt+0x98/0xb0
[ 162.486661] softirqs last enabled at (306008): [<ffffffff817e9739>] __do_softirq+0x349/0x52d
[ 162.486664] softirqs last disabled at (306001): [<ffffffff810a98c5>] irq_exit+0xf5/0x110
[ 162.486666] CPU: 3 PID: 4210 Comm: ip6tables-resto Not tainted 4.10.0-rc6-next-20170202 #498
[ 162.486667] Hardware name: VMware, Inc. VMware Virtual Platform/440BX Desktop Reference Platform, BIOS 6.00 07/02/2015
[ 162.486668] task: ffff880062eb4a40 task.stack: ffffc900054e8000
[ 162.486671] RIP: 0010:smp_call_function_many+0x25c/0x320
[ 162.486672] RSP: 0018:ffffc900054ebc98 EFLAGS: 00000202 ORIG_RAX: ffffffffffffff10
[ 162.486674] RAX: 0000000000000000 RBX: ffff88006d9dd680 RCX: 0000000000000001
[ 162.486675] RDX: 0000000000000001 RSI: ffff88006d3e3600 RDI: ffff88006c52ad68
[ 162.486675] RBP: ffffc900054ebce0 R08: 0000000000000007 R09: 0000000000000000
[ 162.486676] R10: 0000000000000001 R11: ffff880067ecd768 R12: 0000000000000000
[ 162.486677] R13: ffffffff81080790 R14: ffffc900054ebd18 R15: 0000000000000080
[ 162.486678] FS: 00007f37a4b86740(0000) GS:ffff88006d800000(0000) knlGS:0000000000000000
[ 162.486679] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 162.486680] CR2: 00000000008ed017 CR3: 0000000053f70000 CR4: 00000000001406e0
[ 162.486719] Call Trace:
[ 162.486725] ? x86_configure_nx+0x50/0x50
[ 162.486727] on_each_cpu+0x3b/0xa0
[ 162.486730] flush_tlb_kernel_range+0x79/0x80
[ 162.486734] remove_vm_area+0xb1/0xc0
[ 162.486737] __vunmap+0x2e/0x110
[ 162.486739] vfree+0x2e/0x70
[ 162.486744] do_ip6t_get_ctl+0x2de/0x370 [ip6_tables]
[ 162.486751] nf_getsockopt+0x49/0x70
[ 162.486755] ipv6_getsockopt+0xd3/0x130
[ 162.486758] rawv6_getsockopt+0x2c/0x70
[ 162.486761] sock_common_getsockopt+0x14/0x20
[ 162.486763] SyS_getsockopt+0x77/0xe0
[ 162.486767] do_syscall_64+0x6c/0x200
[ 162.486770] entry_SYSCALL64_slow_path+0x25/0x25
[ 162.486771] RIP: 0033:0x7f37a3d9151a
[ 162.486772] RSP: 002b:00007ffeecccd9e8 EFLAGS: 00000202 ORIG_RAX: 0000000000000037
[ 162.486774] RAX: ffffffffffffffda RBX: 0000000000000003 RCX: 00007f37a3d9151a
[ 162.486774] RDX: 0000000000000041 RSI: 0000000000000029 RDI: 0000000000000003
[ 162.486775] RBP: 00000000008e90c0 R08: 00007ffeecccda30 R09: feff7164736b6865
[ 162.486776] R10: 00000000008e90c0 R11: 0000000000000202 R12: 00007ffeecccdf60
[ 162.486777] R13: 00000000008e9010 R14: 00007ffeecccda40 R15: 0000000000000000
[ 162.486783] Code: bb 39 2b 00 3b 05 a9 40 c7 00 41 89 c4 0f 8d 3f fe ff ff 48 63 d0 48 8b 33 48 03 34 d5 60 c4 ab 81 8b 56 18 83 e2 01 74 0a f3 90 <8b> 4e 18 83 e1 01 75 f6 83 f8 ff 48 8b 7b 08 74 2a 48 63 35 70
----------
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-02-07 09:50 +0100 |
| Message-ID | <t89Dk-vJ-23@gated-at.bofh.it> |
| In reply to | #1575081 |
On Mon 06-02-17 20:13:35, Dmitry Vyukov wrote:
[...]
> Fuzzer now runs on 510948533b059f4f5033464f9f4a0c32d4ab0c08 of
> mmotm/auto-latest
> (git://git.kernel.org/pub/scm/linux/kernel/git/mhocko/mm.git):
>
> commit 510948533b059f4f5033464f9f4a0c32d4ab0c08
> Date: Thu Feb 2 10:08:47 2017 +0100
> mmotm: userfaultfd-non-cooperative-add-event-for-memory-unmaps-fix
>
> The commit you referenced is already there:
>
> commit 806b158031ca0b4714e775898396529a758ebc2c
> Date: Thu Feb 2 08:53:16 2017 +0100
> mm, page_alloc: use static global work_struct for draining per-cpu pages
>
> But I still got:
>
> [ INFO: possible circular locking dependency detected ]
> 4.9.0 #6 Not tainted
> -------------------------------------------------------
> syz-executor1/8199 is trying to acquire lock:
> (cpu_hotplug.dep_map){++++++}, at: [<ffffffff81422fe7>] get_online_cpus+0x37/0x90 kernel/cpu.c:246
> but task is already holding lock:
> (pcpu_alloc_mutex){+.+.+.}, at: [<ffffffff818f07ea>] pcpu_alloc+0xbda/0x1280 mm/percpu.c:896
> which lock already depends on the new lock.
>
the original was too hard to read so here is the reformated output.
> the existing dependency chain (in reverse order) is:
>
> [ 403.953319] [<ffffffff8156fc29>] validate_chain kernel/locking/lockdep.c:2265 [inline]
> [ 403.953319] [<ffffffff8156fc29>] __lock_acquire+0x2149/0x3430 kernel/locking/lockdep.c:3338
> [ 403.961232] [<ffffffff81571db1>] lock_acquire+0x2a1/0x630 kernel/locking/lockdep.c:3753
> [ 403.968788] [<ffffffff8436697e>] __mutex_lock_common kernel/locking/mutex.c:521 [inline]
> [ 403.968788] [<ffffffff8436697e>] mutex_lock_nested+0x24e/0xff0 kernel/locking/mutex.c:621
> [ 403.976782] [<ffffffff818f07ea>] pcpu_alloc+0xbda/0x1280 mm/percpu.c:896
> [ 403.984266] [<ffffffff818f0ee4>] __alloc_percpu+0x24/0x30 mm/percpu.c:1075
> [ 403.991873] [<ffffffff816543e3>] smpcfd_prepare_cpu+0x73/0xd0 kernel/smp.c:44
> [ 403.999799] [<ffffffff814240b4>] cpuhp_invoke_callback+0x254/0x1480 kernel/cpu.c:136
> [ 404.008253] [<ffffffff81425821>] cpuhp_up_callbacks+0x81/0x2a0 kernel/cpu.c:493
> [ 404.016365] [<ffffffff81427bf3>] _cpu_up+0x1e3/0x2a0 kernel/cpu.c:1057
> [ 404.023507] [<ffffffff81427d23>] do_cpu_up+0x73/0xa0 kernel/cpu.c:1087
> [ 404.030647] [<ffffffff81427d68>] cpu_up+0x18/0x20 kernel/cpu.c:1095
> [ 404.037523] [<ffffffff854ede84>] smp_init+0xe9/0xee kernel/smp.c:564
> [ 404.044559] [<ffffffff85482f81>] kernel_init_freeable+0x439/0x690 init/main.c:1010
> [ 404.052811] [<ffffffff84357083>] kernel_init+0x13/0x180 init/main.c:941
> [ 404.060198] [<ffffffff84377baa>] ret_from_fork+0x2a/0x40 arch/x86/entry/entry_64.S:433
cpu_hotplug_begin
cpu_hotplug.lock
pcpu_alloc
pcpu_alloc_mutex
> [ 404.072827] [<ffffffff8156fc29>] validate_chain kernel/locking/lockdep.c:2265 [inline]
> [ 404.072827] [<ffffffff8156fc29>] __lock_acquire+0x2149/0x3430 kernel/locking/lockdep.c:3338
> [ 404.080733] [<ffffffff81571db1>] lock_acquire+0x2a1/0x630 kernel/locking/lockdep.c:3753
> [ 404.088311] [<ffffffff8436697e>] __mutex_lock_common kernel/locking/mutex.c:521 [inline]
> [ 404.088311] [<ffffffff8436697e>] mutex_lock_nested+0x24e/0xff0 kernel/locking/mutex.c:621
> [ 404.096318] [<ffffffff81427876>] cpu_hotplug_begin+0x206/0x2e0 kernel/cpu.c:304
> [ 404.104321] [<ffffffff81427ada>] _cpu_up+0xca/0x2a0 kernel/cpu.c:1011
> [ 404.111357] [<ffffffff81427d23>] do_cpu_up+0x73/0xa0 kernel/cpu.c:1087
> [ 404.118480] [<ffffffff81427d68>] cpu_up+0x18/0x20 kernel/cpu.c:1095
> [ 404.125360] [<ffffffff854ede84>] smp_init+0xe9/0xee kernel/smp.c:564
> [ 404.132393] [<ffffffff85482f81>] kernel_init_freeable+0x439/0x690 init/main.c:1010
> [ 404.140668] [<ffffffff84357083>] kernel_init+0x13/0x180 init/main.c:941
> [ 404.148079] [<ffffffff84377baa>] ret_from_fork+0x2a/0x40 arch/x86/entry/entry_64.S:433
cpu_hotplug_begin
cpu_hotplug.lock
> [ 404.160977] [<ffffffff8156976d>] check_prev_add kernel/locking/lockdep.c:1828 [inline]
> [ 404.160977] [<ffffffff8156976d>] check_prevs_add+0xa8d/0x1c00 kernel/locking/lockdep.c:1938
> [ 404.168898] [<ffffffff8156fc29>] validate_chain kernel/locking/lockdep.c:2265 [inline]
> [ 404.168898] [<ffffffff8156fc29>] __lock_acquire+0x2149/0x3430 kernel/locking/lockdep.c:3338
> [ 404.176844] [<ffffffff81571db1>] lock_acquire+0x2a1/0x630 kernel/locking/lockdep.c:3753
> [ 404.184416] [<ffffffff81423012>] get_online_cpus+0x62/0x90 kernel/cpu.c:248
> [ 404.192103] [<ffffffff8185fcf8>] drain_all_pages+0xf8/0x710 mm/page_alloc.c:2385
> [ 404.199880] [<ffffffff81865e5d>] __alloc_pages_direct_reclaim mm/page_alloc.c:3440 [inline]
> [ 404.199880] [<ffffffff81865e5d>] __alloc_pages_slowpath+0x8fd/0x2370 mm/page_alloc.c:3778
> [ 404.208406] [<ffffffff818681c5>] __alloc_pages_nodemask+0x8f5/0xc60 mm/page_alloc.c:3980
> [ 404.216851] [<ffffffff818ed0c1>] __alloc_pages include/linux/gfp.h:426 [inline]
> [ 404.216851] [<ffffffff818ed0c1>] __alloc_pages_node include/linux/gfp.h:439 [inline]
> [ 404.216851] [<ffffffff818ed0c1>] alloc_pages_node include/linux/gfp.h:453 [inline]
> [ 404.216851] [<ffffffff818ed0c1>] pcpu_alloc_pages mm/percpu-vm.c:93 [inline]
> [ 404.216851] [<ffffffff818ed0c1>] pcpu_populate_chunk+0x1e1/0x900 mm/percpu-vm.c:282
> [ 404.225015] [<ffffffff818f0a11>] pcpu_alloc+0xe01/0x1280 mm/percpu.c:998
> [ 404.232482] [<ffffffff818f0eb7>] __alloc_percpu_gfp+0x27/0x30 mm/percpu.c:1062
> [ 404.240389] [<ffffffff817d25b2>] bpf_array_alloc_percpu kernel/bpf/arraymap.c:34 [inline]
> [ 404.240389] [<ffffffff817d25b2>] array_map_alloc+0x532/0x710 kernel/bpf/arraymap.c:99
> [ 404.248224] [<ffffffff817ba034>] find_and_alloc_map kernel/bpf/syscall.c:34 [inline]
> [ 404.248224] [<ffffffff817ba034>] map_create kernel/bpf/syscall.c:188 [inline]
> [ 404.248224] [<ffffffff817ba034>] SYSC_bpf kernel/bpf/syscall.c:870 [inline]
> [ 404.248224] [<ffffffff817ba034>] SyS_bpf+0xd64/0x2500 kernel/bpf/syscall.c:827
> [ 404.255434] [<ffffffff84377941>] entry_SYSCALL_64_fastpath+0x1f/0xc2
pcpu_alloc
pcpu_alloc_mutex
drain_all_pages
pcpu_drain_mutex
get_online_cpus
cpu_hotplug.lock
so the deadlock is real!
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-02-07 23:40 +0100 |
| Message-ID | <t8mAy-mn-5@gated-at.bofh.it> |
| In reply to | #1575081 |
On Mon, 6 Feb 2017, Dmitry Vyukov wrote:
> On Mon, Jan 30, 2017 at 4:48 PM, Dmitry Vyukov <dvyukov@google.com> wrote:
> Unfortunately it does not seem to help.
> Fuzzer now runs on 510948533b059f4f5033464f9f4a0c32d4ab0c08 of
> mmotm/auto-latest
> (git://git.kernel.org/pub/scm/linux/kernel/git/mhocko/mm.git):
>
> commit 510948533b059f4f5033464f9f4a0c32d4ab0c08
> Date: Thu Feb 2 10:08:47 2017 +0100
> mmotm: userfaultfd-non-cooperative-add-event-for-memory-unmaps-fix
>
> The commit you referenced is already there:
>
> commit 806b158031ca0b4714e775898396529a758ebc2c
> Date: Thu Feb 2 08:53:16 2017 +0100
> mm, page_alloc: use static global work_struct for draining per-cpu pages
<SNIP>
> Chain exists of:
> Possible unsafe locking scenario:
>
> CPU0 CPU1
> ---- ----
> lock(pcpu_alloc_mutex);
> lock(cpu_hotplug.lock);
> lock(pcpu_alloc_mutex);
> lock(cpu_hotplug.dep_map);
And that's exactly what happens:
cpu_up()
alloc_percpu() lock(hotplug.lock)
lock(&pcpu_alloc_mutex)
.. alloc_percpu()
drain_all_pages() lock(&pcpu_alloc_mutex)
get_online_cpus()
lock(hotplug.lock)
Classic deadlock, i.e. you _cannot_ call get_online_cpus() while holding
pcpu_alloc_mutex.
Alternatively you can forbid to do per cpu alloc/free while holding
hotplug.lock. I doubt that this will make people happy :)
Thanks,
tglx
[toc] | [prev] | [standalone]
Page 3 of 3 — ← Prev page 1 2 [3]
Back to top | Article view | linux.kernel
csiph-web