Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1617380 > unrolled thread
| Started by | Hugh Dickins <hughd@google.com> |
|---|---|
| First post | 2017-04-05 23:10 +0200 |
| Last post | 2017-04-08 20:10 +0200 |
| Articles | 11 — 3 participants |
Back to article view | Back to linux.kernel
Is it safe for kthreadd to drain_all_pages? Hugh Dickins <hughd@google.com> - 2017-04-05 23:10 +0200
Re: Is it safe for kthreadd to drain_all_pages? Michal Hocko <mhocko@kernel.org> - 2017-04-06 11:00 +0200
Re: Is it safe for kthreadd to drain_all_pages? Mel Gorman <mgorman@techsingularity.net> - 2017-04-06 15:10 +0200
Re: Is it safe for kthreadd to drain_all_pages? Hugh Dickins <hughd@google.com> - 2017-04-06 21:00 +0200
Re: Is it safe for kthreadd to drain_all_pages? Hugh Dickins <hughd@google.com> - 2017-04-07 18:30 +0200
Re: Is it safe for kthreadd to drain_all_pages? Michal Hocko <mhocko@kernel.org> - 2017-04-07 18:50 +0200
Re: Is it safe for kthreadd to drain_all_pages? Hugh Dickins <hughd@google.com> - 2017-04-07 19:00 +0200
Re: Is it safe for kthreadd to drain_all_pages? Michal Hocko <mhocko@kernel.org> - 2017-04-07 19:30 +0200
Re: Is it safe for kthreadd to drain_all_pages? Hugh Dickins <hughd@google.com> - 2017-04-07 20:50 +0200
Re: Is it safe for kthreadd to drain_all_pages? Hugh Dickins <hughd@google.com> - 2017-04-08 19:10 +0200
Re: Is it safe for kthreadd to drain_all_pages? Mel Gorman <mgorman@techsingularity.net> - 2017-04-08 20:10 +0200
| From | Hugh Dickins <hughd@google.com> |
|---|---|
| Date | 2017-04-05 23:10 +0200 |
| Subject | Is it safe for kthreadd to drain_all_pages? |
| Message-ID | <tt0lI-7Md-15@gated-at.bofh.it> |
Hi Mel,
I suspect that it's not safe for kthreadd to drain_all_pages();
but I haven't studied flush_work() etc, so don't really know what
I'm talking about: hoping that you will jump to a realization.
4.11-rc has been giving me hangs after hours of swapping load. At
first they looked like memory leaks ("fork: Cannot allocate memory");
but for no good reason I happened to do "cat /proc/sys/vm/stat_refresh"
before looking at /proc/meminfo one time, and the stat_refresh stuck
in D state, waiting for completion of flush_work like many kworkers.
kthreadd waiting for completion of flush_work in drain_all_pages().
But I only noticed that pattern later: originally tried to bisect
rc1 before rc2 came out, but underestimated how long to wait before
deciding a stage good - I thought 12 hours, but would now say 2 days.
Too late for bisection, I suspect your drain_all_pages() changes.
(I've also found order:0 page allocation stalls in /var/log/messages,
148804ms a nice example: which suggest that these hangs are perhaps a
condition it can sometimes get out of itself. None with the patch.)
Patch below has been running well for 36 hours now:
a bit too early to be sure, but I think it's time to turn to you.
[PATCH] mm: don't let kthreadd drain_all_pages
4.11-rc has been giving me hangs after many hours of swapping load: most
kworkers waiting for completion of a flush_work, kthreadd waiting for
completion of flush_work in drain_all_pages (while doing copy_process).
I suspect that kthreadd should not be allowed to drain_all_pages().
Signed-off-by: Hugh Dickins <hughd@google.com>
---
mm/page_alloc.c | 2 ++
1 file changed, 2 insertions(+)
--- 4.11-rc5/mm/page_alloc.c 2017-03-13 09:08:37.743209168 -0700
+++ linux/mm/page_alloc.c 2017-04-04 00:33:44.086867413 -0700
@@ -2376,6 +2376,8 @@ void drain_all_pages(struct zone *zone)
/* Workqueues cannot recurse */
if (current->flags & PF_WQ_WORKER)
return;
+ if (current == kthreadd_task)
+ return;
/*
* Do not drain if one is already in progress unless it's specific to
[toc] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-04-06 11:00 +0200 |
| Message-ID | <ttbqP-6ei-67@gated-at.bofh.it> |
| In reply to | #1617380 |
On Wed 05-04-17 13:59:49, Hugh Dickins wrote:
> Hi Mel,
>
> I suspect that it's not safe for kthreadd to drain_all_pages();
> but I haven't studied flush_work() etc, so don't really know what
> I'm talking about: hoping that you will jump to a realization.
>
> 4.11-rc has been giving me hangs after hours of swapping load. At
> first they looked like memory leaks ("fork: Cannot allocate memory");
> but for no good reason I happened to do "cat /proc/sys/vm/stat_refresh"
> before looking at /proc/meminfo one time, and the stat_refresh stuck
> in D state, waiting for completion of flush_work like many kworkers.
> kthreadd waiting for completion of flush_work in drain_all_pages().
>
> But I only noticed that pattern later: originally tried to bisect
> rc1 before rc2 came out, but underestimated how long to wait before
> deciding a stage good - I thought 12 hours, but would now say 2 days.
> Too late for bisection, I suspect your drain_all_pages() changes.
Yes, this is a fallout from Mel's changes. I was about to say that
my follow up fixes which made this flushing to the single WQ with rescuer
fixed that but it seems that
http://www.ozlabs.org/~akpm/mmotm/broken-out/mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch
didn't make it to the Linus tree. Could you re-test with this one?
While your change is obviously correct I think the above should address
it as well and it is more generic. If it works then I will ask Andrew to
send the above to Linus (along with its follow up
mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch)
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2017-04-06 15:10 +0200 |
| Message-ID | <ttfkK-ux-23@gated-at.bofh.it> |
| In reply to | #1617380 |
On Wed, Apr 05, 2017 at 01:59:49PM -0700, Hugh Dickins wrote:
> Hi Mel,
>
> I suspect that it's not safe for kthreadd to drain_all_pages();
> but I haven't studied flush_work() etc, so don't really know what
> I'm talking about: hoping that you will jump to a realization.
>
You're right, it's not safe. If kthreadd is creating the workqueue
thread to do the drain and it'll recurse into itself.
> 4.11-rc has been giving me hangs after hours of swapping load. At
> first they looked like memory leaks ("fork: Cannot allocate memory");
> but for no good reason I happened to do "cat /proc/sys/vm/stat_refresh"
> before looking at /proc/meminfo one time, and the stat_refresh stuck
> in D state, waiting for completion of flush_work like many kworkers.
> kthreadd waiting for completion of flush_work in drain_all_pages().
>
It's asking itself to do work in all likelihood.
> Patch below has been running well for 36 hours now:
> a bit too early to be sure, but I think it's time to turn to you.
>
I think the patch is valid but like Michal, would appreciate if you
could run the patch he linked to see if it also side-steps the same
problem.
Good spot!
--
Mel Gorman
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Hugh Dickins <hughd@google.com> |
|---|---|
| Date | 2017-04-06 21:00 +0200 |
| Message-ID | <ttkNs-4Ff-33@gated-at.bofh.it> |
| In reply to | #1617995 |
On Thu, 6 Apr 2017, Mel Gorman wrote:
> On Wed, Apr 05, 2017 at 01:59:49PM -0700, Hugh Dickins wrote:
> > Hi Mel,
> >
> > I suspect that it's not safe for kthreadd to drain_all_pages();
> > but I haven't studied flush_work() etc, so don't really know what
> > I'm talking about: hoping that you will jump to a realization.
> >
>
> You're right, it's not safe. If kthreadd is creating the workqueue
> thread to do the drain and it'll recurse into itself.
>
> > 4.11-rc has been giving me hangs after hours of swapping load. At
> > first they looked like memory leaks ("fork: Cannot allocate memory");
> > but for no good reason I happened to do "cat /proc/sys/vm/stat_refresh"
> > before looking at /proc/meminfo one time, and the stat_refresh stuck
> > in D state, waiting for completion of flush_work like many kworkers.
> > kthreadd waiting for completion of flush_work in drain_all_pages().
> >
>
> It's asking itself to do work in all likelihood.
>
> > Patch below has been running well for 36 hours now:
> > a bit too early to be sure, but I think it's time to turn to you.
> >
>
> I think the patch is valid but like Michal, would appreciate if you
> could run the patch he linked to see if it also side-steps the same
> problem.
>
> Good spot!
Thank you both for explanations, and direction to the two "drainging"
patches. I've put those on to 4.11-rc5 (and double-checked that I've
taken mine off), and set it going. Fine so far but much too soon to
tell - mine did 56 hours with clean /var/log/messages before I switched,
so I demand no less of Michal's :). I'll report back tomorrow and the
day after (unless badness appears sooner once I'm home).
Hugh
[toc] | [prev] | [next] | [standalone]
| From | Hugh Dickins <hughd@google.com> |
|---|---|
| Date | 2017-04-07 18:30 +0200 |
| Message-ID | <ttEVP-1dU-5@gated-at.bofh.it> |
| In reply to | #1618278 |
On Thu, 6 Apr 2017, Hugh Dickins wrote:
> On Thu, 6 Apr 2017, Mel Gorman wrote:
> > On Wed, Apr 05, 2017 at 01:59:49PM -0700, Hugh Dickins wrote:
> > > Hi Mel,
> > >
> > > I suspect that it's not safe for kthreadd to drain_all_pages();
> > > but I haven't studied flush_work() etc, so don't really know what
> > > I'm talking about: hoping that you will jump to a realization.
> > >
> >
> > You're right, it's not safe. If kthreadd is creating the workqueue
> > thread to do the drain and it'll recurse into itself.
> >
> > > 4.11-rc has been giving me hangs after hours of swapping load. At
> > > first they looked like memory leaks ("fork: Cannot allocate memory");
> > > but for no good reason I happened to do "cat /proc/sys/vm/stat_refresh"
> > > before looking at /proc/meminfo one time, and the stat_refresh stuck
> > > in D state, waiting for completion of flush_work like many kworkers.
> > > kthreadd waiting for completion of flush_work in drain_all_pages().
> > >
> >
> > It's asking itself to do work in all likelihood.
> >
> > > Patch below has been running well for 36 hours now:
> > > a bit too early to be sure, but I think it's time to turn to you.
> > >
> >
> > I think the patch is valid but like Michal, would appreciate if you
> > could run the patch he linked to see if it also side-steps the same
> > problem.
> >
> > Good spot!
>
> Thank you both for explanations, and direction to the two "drainging"
> patches. I've put those on to 4.11-rc5 (and double-checked that I've
> taken mine off), and set it going. Fine so far but much too soon to
> tell - mine did 56 hours with clean /var/log/messages before I switched,
> so I demand no less of Michal's :). I'll report back tomorrow and the
> day after (unless badness appears sooner once I'm home).
24 hours so far, and with a clean /var/log/messages. Not conclusive
yet, and of course I'll leave it running another couple of days, but
I'm increasingly sure that it works as you intended: I agree that
mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch
mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch
should go to Linus as soon as convenient. Though I think the commit
message needs something a bit stronger than "Quite annoying though".
Maybe add a line:
Fixes serious hang under load, observed repeatedly on 4.11-rc.
Thanks!
Hugh
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-04-07 18:50 +0200 |
| Message-ID | <ttFfc-1ms-25@gated-at.bofh.it> |
| In reply to | #1618947 |
On Fri 07-04-17 09:25:33, Hugh Dickins wrote: [...] > 24 hours so far, and with a clean /var/log/messages. Not conclusive > yet, and of course I'll leave it running another couple of days, but > I'm increasingly sure that it works as you intended: I agree that > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch > mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch > > should go to Linus as soon as convenient. Though I think the commit > message needs something a bit stronger than "Quite annoying though". > Maybe add a line: > > Fixes serious hang under load, observed repeatedly on 4.11-rc. Yeah, it is much less theoretical now. I will rephrase and ask Andrew to update the chagelog and send it to Linus once I've got your final go. Thanks! -- Michal Hocko SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Hugh Dickins <hughd@google.com> |
|---|---|
| Date | 2017-04-07 19:00 +0200 |
| Message-ID | <ttFoS-1r2-15@gated-at.bofh.it> |
| In reply to | #1618963 |
On Fri, 7 Apr 2017, Michal Hocko wrote: > On Fri 07-04-17 09:25:33, Hugh Dickins wrote: > [...] > > 24 hours so far, and with a clean /var/log/messages. Not conclusive > > yet, and of course I'll leave it running another couple of days, but > > I'm increasingly sure that it works as you intended: I agree that > > > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch > > > > should go to Linus as soon as convenient. Though I think the commit > > message needs something a bit stronger than "Quite annoying though". > > Maybe add a line: > > > > Fixes serious hang under load, observed repeatedly on 4.11-rc. > > Yeah, it is much less theoretical now. I will rephrase and ask Andrew to > update the chagelog and send it to Linus once I've got your final go. I don't know akpm's timetable, but your fix being more than a two-liner, I think it would be better if it could get into rc6, than wait another week for rc7, just in case others then find problems with it. So I think it's safer *not* to wait for my final go, but proceed on the assumption that it will follow a day later. Hugh
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2017-04-07 19:30 +0200 |
| Message-ID | <ttFRU-1SP-11@gated-at.bofh.it> |
| In reply to | #1618967 |
On Fri 07-04-17 09:58:17, Hugh Dickins wrote:
> On Fri, 7 Apr 2017, Michal Hocko wrote:
> > On Fri 07-04-17 09:25:33, Hugh Dickins wrote:
> > [...]
> > > 24 hours so far, and with a clean /var/log/messages. Not conclusive
> > > yet, and of course I'll leave it running another couple of days, but
> > > I'm increasingly sure that it works as you intended: I agree that
> > >
> > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch
> > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch
> > >
> > > should go to Linus as soon as convenient. Though I think the commit
> > > message needs something a bit stronger than "Quite annoying though".
> > > Maybe add a line:
> > >
> > > Fixes serious hang under load, observed repeatedly on 4.11-rc.
> >
> > Yeah, it is much less theoretical now. I will rephrase and ask Andrew to
> > update the chagelog and send it to Linus once I've got your final go.
>
> I don't know akpm's timetable, but your fix being more than a two-liner,
> I think it would be better if it could get into rc6, than wait another
> week for rc7, just in case others then find problems with it. So I
> think it's safer *not* to wait for my final go, but proceed on the
> assumption that it will follow a day later.
Fair enough. Andrew, could you update the changelog of
mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch
and send it to Linus along with
mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch before rc6?
I would add your Teste-by Hugh but I guess you want to give your testing
more time before feeling comfortable to give it.
---
mm: move pcp and lru-pcp draining into single wq
We currently have 2 specific WQ_RECLAIM workqueues in the mm code.
vmstat_wq for updating pcp stats and lru_add_drain_wq dedicated to drain
per cpu lru caches. This seems more than necessary because both can run
on a single WQ. Both do not block on locks requiring a memory allocation
nor perform any allocations themselves. We will save one rescuer thread
this way.
On the other hand drain_all_pages() queues work on the system wq which
doesn't have rescuer and so this depend on memory allocation (when all
workers are stuck allocating and new ones cannot be created). Initially
we thought this would be more of a theoretical problem but Hugh Dickins
has reported:
: 4.11-rc has been giving me hangs after hours of swapping load. At
: first they looked like memory leaks ("fork: Cannot allocate memory");
: but for no good reason I happened to do "cat /proc/sys/vm/stat_refresh"
: before looking at /proc/meminfo one time, and the stat_refresh stuck
: in D state, waiting for completion of flush_work like many kworkers.
: kthreadd waiting for completion of flush_work in drain_all_pages().
This worker should be using WQ_RECLAIM as well in order to guarantee
a forward progress. We can reuse the same one as for lru draining and
vmstat.
Link: http://lkml.kernel.org/r/20170307131751.24936-1-mhocko@kernel.org
Fixes: 0ccce3b92421 ("mm, page_alloc: drain per-cpu pages from workqueue context")
Signed-off-by: Michal Hocko <mhocko@suse.com>
Suggested-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
Acked-by: Vlastimil Babka <vbabka@suse.cz>
Acked-by: Mel Gorman <mgorman@suse.de>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Hugh Dickins <hughd@google.com> |
|---|---|
| Date | 2017-04-07 20:50 +0200 |
| Message-ID | <ttH7j-2Ox-1@gated-at.bofh.it> |
| In reply to | #1618991 |
On Fri, 7 Apr 2017, Michal Hocko wrote:
> On Fri 07-04-17 09:58:17, Hugh Dickins wrote:
> > On Fri, 7 Apr 2017, Michal Hocko wrote:
> > > On Fri 07-04-17 09:25:33, Hugh Dickins wrote:
> > > [...]
> > > > 24 hours so far, and with a clean /var/log/messages. Not conclusive
> > > > yet, and of course I'll leave it running another couple of days, but
> > > > I'm increasingly sure that it works as you intended: I agree that
> > > >
> > > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch
> > > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch
> > > >
> > > > should go to Linus as soon as convenient. Though I think the commit
> > > > message needs something a bit stronger than "Quite annoying though".
> > > > Maybe add a line:
> > > >
> > > > Fixes serious hang under load, observed repeatedly on 4.11-rc.
> > >
> > > Yeah, it is much less theoretical now. I will rephrase and ask Andrew to
> > > update the chagelog and send it to Linus once I've got your final go.
> >
> > I don't know akpm's timetable, but your fix being more than a two-liner,
> > I think it would be better if it could get into rc6, than wait another
> > week for rc7, just in case others then find problems with it. So I
> > think it's safer *not* to wait for my final go, but proceed on the
> > assumption that it will follow a day later.
>
> Fair enough. Andrew, could you update the changelog of
> mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch
> and send it to Linus along with
> mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch before rc6?
>
> I would add your Teste-by Hugh but I guess you want to give your testing
> more time before feeling comfortable to give it.
Yes, fair enough: at the moment it's just
Half-Tested-by: Hugh Dickins <hughd@google.com>
and I hope to take the Half- off in about 21 hours.
But I certainly wouldn't mind if it found its way to Linus without my
final seal of approval.
> ---
> mm: move pcp and lru-pcp draining into single wq
>
> We currently have 2 specific WQ_RECLAIM workqueues in the mm code.
> vmstat_wq for updating pcp stats and lru_add_drain_wq dedicated to drain
> per cpu lru caches. This seems more than necessary because both can run
> on a single WQ. Both do not block on locks requiring a memory allocation
> nor perform any allocations themselves. We will save one rescuer thread
> this way.
>
> On the other hand drain_all_pages() queues work on the system wq which
> doesn't have rescuer and so this depend on memory allocation (when all
> workers are stuck allocating and new ones cannot be created). Initially
> we thought this would be more of a theoretical problem but Hugh Dickins
> has reported:
> : 4.11-rc has been giving me hangs after hours of swapping load. At
> : first they looked like memory leaks ("fork: Cannot allocate memory");
> : but for no good reason I happened to do "cat /proc/sys/vm/stat_refresh"
> : before looking at /proc/meminfo one time, and the stat_refresh stuck
> : in D state, waiting for completion of flush_work like many kworkers.
> : kthreadd waiting for completion of flush_work in drain_all_pages().
>
> This worker should be using WQ_RECLAIM as well in order to guarantee
> a forward progress. We can reuse the same one as for lru draining and
> vmstat.
>
> Link: http://lkml.kernel.org/r/20170307131751.24936-1-mhocko@kernel.org
> Fixes: 0ccce3b92421 ("mm, page_alloc: drain per-cpu pages from workqueue context")
> Signed-off-by: Michal Hocko <mhocko@suse.com>
> Suggested-by: Tetsuo Handa <penguin-kernel@I-love.SAKURA.ne.jp>
> Acked-by: Vlastimil Babka <vbabka@suse.cz>
> Acked-by: Mel Gorman <mgorman@suse.de>
> Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
> --
> Michal Hocko
> SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Hugh Dickins <hughd@google.com> |
|---|---|
| Date | 2017-04-08 19:10 +0200 |
| Message-ID | <tu225-85S-1@gated-at.bofh.it> |
| In reply to | #1619042 |
On Fri, 7 Apr 2017, Hugh Dickins wrote: > On Fri, 7 Apr 2017, Michal Hocko wrote: > > On Fri 07-04-17 09:58:17, Hugh Dickins wrote: > > > On Fri, 7 Apr 2017, Michal Hocko wrote: > > > > On Fri 07-04-17 09:25:33, Hugh Dickins wrote: > > > > [...] > > > > > 24 hours so far, and with a clean /var/log/messages. Not conclusive > > > > > yet, and of course I'll leave it running another couple of days, but > > > > > I'm increasingly sure that it works as you intended: I agree that > > > > > > > > > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch > > > > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch > > > > > > > > > > should go to Linus as soon as convenient. Though I think the commit > > > > > message needs something a bit stronger than "Quite annoying though". > > > > > Maybe add a line: > > > > > > > > > > Fixes serious hang under load, observed repeatedly on 4.11-rc. > > > > > > > > Yeah, it is much less theoretical now. I will rephrase and ask Andrew to > > > > update the chagelog and send it to Linus once I've got your final go. > > > > > > I don't know akpm's timetable, but your fix being more than a two-liner, > > > I think it would be better if it could get into rc6, than wait another > > > week for rc7, just in case others then find problems with it. So I > > > think it's safer *not* to wait for my final go, but proceed on the > > > assumption that it will follow a day later. > > > > Fair enough. Andrew, could you update the changelog of > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch > > and send it to Linus along with > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch before rc6? > > > > I would add your Teste-by Hugh but I guess you want to give your testing > > more time before feeling comfortable to give it. > > Yes, fair enough: at the moment it's just > Half-Tested-by: Hugh Dickins <hughd@google.com> > and I hope to take the Half- off in about 21 hours. > But I certainly wouldn't mind if it found its way to Linus without my > final seal of approval. 48 hours and still going well: I declare it good, and thanks to Andrew, Linus has ce612879ddc7 "mm: move pcp and lru-pcp draining into single wq" already in for rc6. Hugh
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2017-04-08 20:10 +0200 |
| Message-ID | <tu2Y9-g7-7@gated-at.bofh.it> |
| In reply to | #1619332 |
On Sat, Apr 08, 2017 at 10:04:20AM -0700, Hugh Dickins wrote: > On Fri, 7 Apr 2017, Hugh Dickins wrote: > > On Fri, 7 Apr 2017, Michal Hocko wrote: > > > On Fri 07-04-17 09:58:17, Hugh Dickins wrote: > > > > On Fri, 7 Apr 2017, Michal Hocko wrote: > > > > > On Fri 07-04-17 09:25:33, Hugh Dickins wrote: > > > > > [...] > > > > > > 24 hours so far, and with a clean /var/log/messages. Not conclusive > > > > > > yet, and of course I'll leave it running another couple of days, but > > > > > > I'm increasingly sure that it works as you intended: I agree that > > > > > > > > > > > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch > > > > > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch > > > > > > > > > > > > should go to Linus as soon as convenient. Though I think the commit > > > > > > message needs something a bit stronger than "Quite annoying though". > > > > > > Maybe add a line: > > > > > > > > > > > > Fixes serious hang under load, observed repeatedly on 4.11-rc. > > > > > > > > > > Yeah, it is much less theoretical now. I will rephrase and ask Andrew to > > > > > update the chagelog and send it to Linus once I've got your final go. > > > > > > > > I don't know akpm's timetable, but your fix being more than a two-liner, > > > > I think it would be better if it could get into rc6, than wait another > > > > week for rc7, just in case others then find problems with it. So I > > > > think it's safer *not* to wait for my final go, but proceed on the > > > > assumption that it will follow a day later. > > > > > > Fair enough. Andrew, could you update the changelog of > > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq.patch > > > and send it to Linus along with > > > mm-move-pcp-and-lru-pcp-drainging-into-single-wq-fix.patch before rc6? > > > > > > I would add your Teste-by Hugh but I guess you want to give your testing > > > more time before feeling comfortable to give it. > > > > Yes, fair enough: at the moment it's just > > Half-Tested-by: Hugh Dickins <hughd@google.com> > > and I hope to take the Half- off in about 21 hours. > > But I certainly wouldn't mind if it found its way to Linus without my > > final seal of approval. > > 48 hours and still going well: I declare it good, and thanks to Andrew, > Linus has ce612879ddc7 "mm: move pcp and lru-pcp draining into single wq" > already in for rc6. > Excellent, thanks for that testing. -- Mel Gorman SUSE Labs
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web