Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1711598 > unrolled thread
| Started by | Tim Chen <tim.c.chen@linux.intel.com> |
|---|---|
| First post | 2017-08-15 03:20 +0200 |
| Last post | 2017-08-18 15:10 +0200 |
| Articles | 20 on this page of 63 — 9 participants |
Back to article view | Back to linux.kernel
[PATCH 1/2] sched/wait: Break up long wake list walk Tim Chen <tim.c.chen@linux.intel.com> - 2017-08-15 03:20 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-15 03:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Andi Kleen <ak@linux.intel.com> - 2017-08-15 04:30 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-15 05:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Andi Kleen <ak@linux.intel.com> - 2017-08-15 05:20 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-15 05:30 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Tim Chen <tim.c.chen@linux.intel.com> - 2017-08-15 21:10 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-15 21:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-15 21:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Davidlohr Bueso <dave@stgolabs.net> - 2017-08-16 00:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-16 01:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-16 02:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk ebiederm@xmission.com (Eric W. Biederman) - 2017-08-17 01:30 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-16 01:00 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-17 18:20 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-17 18:30 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-17 22:20 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-17 22:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Mel Gorman <mgorman@techsingularity.net> - 2017-08-18 14:30 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-18 16:30 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Mel Gorman <mgorman@techsingularity.net> - 2017-08-18 16:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Tim Chen <tim.c.chen@linux.intel.com> - 2017-08-18 18:40 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Andi Kleen <ak@linux.intel.com> - 2017-08-18 18:50 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-18 19:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-18 19:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Mel Gorman <mgorman@techsingularity.net> - 2017-08-18 21:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-18 21:20 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Andi Kleen <ak@linux.intel.com> - 2017-08-18 22:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-18 22:20 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Mel Gorman <mgorman@techsingularity.net> - 2017-08-21 20:40 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-21 21:00 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-22 19:30 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-22 20:20 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-22 20:30 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Peter Zijlstra <peterz@infradead.org> - 2017-08-22 21:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-22 21:20 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Peter Zijlstra <peterz@infradead.org> - 2017-08-22 21:10 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Andi Kleen <ak@linux.intel.com> - 2017-08-22 21:40 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Christopher Lameter <cl@linux.com> - 2017-08-22 23:10 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Andi Kleen <ak@linux.intel.com> - 2017-08-22 23:30 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-23 01:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-23 01:20 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-23 17:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-22 21:40 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-22 22:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-22 22:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-22 23:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Peter Zijlstra <peterz@infradead.org> - 2017-08-22 23:00 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-23 16:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Tim Chen <tim.c.chen@linux.intel.com> - 2017-08-23 18:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-23 20:20 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-23 23:00 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-24 01:40 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Tim Chen <tim.c.chen@linux.intel.com> - 2017-08-24 19:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-24 20:20 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Mel Gorman <mgorman@techsingularity.net> - 2017-08-24 22:50 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Mel Gorman <mgorman@techsingularity.net> - 2017-08-23 18:10 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Andi Kleen <ak@linux.intel.com> - 2017-08-18 22:10 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-18 22:40 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-18 22:30 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-18 22:40 +0200
Re: [PATCH 1/2] sched/wait: Break up long wake list walk Linus Torvalds <torvalds@linux-foundation.org> - 2017-08-18 19:00 +0200
RE: [PATCH 1/2] sched/wait: Break up long wake list walk "Liang, Kan" <kan.liang@intel.com> - 2017-08-18 15:10 +0200
Page 2 of 4 — ← Prev page 1 [2] 3 4 Next page →
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2017-08-18 16:50 +0200 |
| Message-ID | <ufQKZ-55f-7@gated-at.bofh.it> |
| In reply to | #1715187 |
On Fri, Aug 18, 2017 at 02:20:38PM +0000, Liang, Kan wrote: > > Nothing fancy other than needing a comment if it works. > > > > No, the patch doesn't work. > That indicates that it may be a hot page and it's possible that the page is locked for a short time but waiters accumulate. What happens if you leave NUMA balancing enabled but disable THP? Waiting on migration entries also uses wait_on_page_locked so it would be interesting to know if the problem is specific to THP. Can you tell me what this workload is doing? I want to see if it's something like many threads pounding on a limited number of pages very quickly. If it's many threads working on private data, it would also be important to know how each buffers threads are aligned, particularly if the buffers are smaller than a THP or base page size. For example, if each thread is operating on a base page sized buffer then disabling THP would side-step the problem but THP would be false sharing between multiple threads. -- Mel Gorman SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Tim Chen <tim.c.chen@linux.intel.com> |
|---|---|
| Date | 2017-08-18 18:40 +0200 |
| Message-ID | <ufSts-6iz-21@gated-at.bofh.it> |
| In reply to | #1715209 |
On 08/18/2017 07:46 AM, Mel Gorman wrote: > On Fri, Aug 18, 2017 at 02:20:38PM +0000, Liang, Kan wrote: >>> Nothing fancy other than needing a comment if it works. >>> >> >> No, the patch doesn't work. >> > > That indicates that it may be a hot page and it's possible that the page is > locked for a short time but waiters accumulate. What happens if you leave > NUMA balancing enabled but disable THP? Waiting on migration entries also > uses wait_on_page_locked so it would be interesting to know if the problem > is specific to THP. > > Can you tell me what this workload is doing? I want to see if it's something > like many threads pounding on a limited number of pages very quickly. If It is a customer workload so we have limited visibility. But we believe there are some pages that are frequently accessed by all threads. > it's many threads working on private data, it would also be important to > know how each buffers threads are aligned, particularly if the buffers > are smaller than a THP or base page size. For example, if each thread is > operating on a base page sized buffer then disabling THP would side-step > the problem but THP would be false sharing between multiple threads. > Still, I don't think this problem is THP specific. If there is a hot regular page getting migrated, we'll also see many threads get queued up quickly. THP may have made the problem worse as migrating it takes a longer time, meaning more threads could get queued up. Thanks. Tim
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <ak@linux.intel.com> |
|---|---|
| Date | 2017-08-18 18:50 +0200 |
| Message-ID | <ufSD7-6m1-1@gated-at.bofh.it> |
| In reply to | #1715311 |
> Still, I don't think this problem is THP specific. If there is a hot > regular page getting migrated, we'll also see many threads get > queued up quickly. THP may have made the problem worse as migrating > it takes a longer time, meaning more threads could get queued up. Also THP probably makes more threads collide because the pages are larger. But still it can all happen even without THP. -Andi
[toc] | [prev] | [next] | [standalone]
| From | "Liang, Kan" <kan.liang@intel.com> |
|---|---|
| Date | 2017-08-18 19:00 +0200 |
| Message-ID | <ufSMN-6pz-11@gated-at.bofh.it> |
| In reply to | #1715209 |
> On Fri, Aug 18, 2017 at 02:20:38PM +0000, Liang, Kan wrote: > > > Nothing fancy other than needing a comment if it works. > > > > > > > No, the patch doesn't work. > > > > That indicates that it may be a hot page and it's possible that the page is > locked for a short time but waiters accumulate. What happens if you leave > NUMA balancing enabled but disable THP? No, disabling THP doesn't help the case. Thanks, Kan > Waiting on migration entries also > uses wait_on_page_locked so it would be interesting to know if the problem > is specific to THP. > > Can you tell me what this workload is doing? I want to see if it's something > like many threads pounding on a limited number of pages very quickly. If it's > many threads working on private data, it would also be important to know > how each buffers threads are aligned, particularly if the buffers are smaller > than a THP or base page size. For example, if each thread is operating on a > base page sized buffer then disabling THP would side-step the problem but > THP would be false sharing between multiple threads. > > > -- > Mel Gorman > SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-08-18 19:50 +0200 |
| Message-ID | <ufTzb-6V7-17@gated-at.bofh.it> |
| In reply to | #1715317 |
On Fri, Aug 18, 2017 at 9:53 AM, Liang, Kan <kan.liang@intel.com> wrote:
>
>> On Fri, Aug 18, 2017 Mel Gorman wrote:
>>
>> That indicates that it may be a hot page and it's possible that the page is
>> locked for a short time but waiters accumulate. What happens if you leave
>> NUMA balancing enabled but disable THP?
>
> No, disabling THP doesn't help the case.
Interesting. That particular code sequence should only be active for
THP. What does the profile look like with THP disabled but with NUMA
balancing still enabled?
Just asking because maybe that different call chain could give us some
other ideas of what the commonality here is that triggers out
behavioral problem.
I was really hoping that we'd root-cause this and have a solution (and
then apply Tim's patch as a "belt and suspenders" kind of thing), but
it's starting to smell like we may have to apply Tim's patch as a
band-aid, and try to figure out what the trigger is longer-term.
Linus
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2017-08-18 21:00 +0200 |
| Message-ID | <ufUEV-7yR-9@gated-at.bofh.it> |
| In reply to | #1715348 |
On Fri, Aug 18, 2017 at 10:48:23AM -0700, Linus Torvalds wrote: > On Fri, Aug 18, 2017 at 9:53 AM, Liang, Kan <kan.liang@intel.com> wrote: > > > >> On Fri, Aug 18, 2017 Mel Gorman wrote: > >> > >> That indicates that it may be a hot page and it's possible that the page is > >> locked for a short time but waiters accumulate. What happens if you leave > >> NUMA balancing enabled but disable THP? > > > > No, disabling THP doesn't help the case. > > Interesting. That particular code sequence should only be active for > THP. What does the profile look like with THP disabled but with NUMA > balancing still enabled? > While that specific code sequence is active in the example, the problem is fundamental to what NUMA balancing does. If many threads share a single page, base page or THP, then any thread accessing the data during migration will block on page lock. The symptoms will be difference but I am willing to bet it'll be a wake on a page lock either way. NUMA balancing is somewhat unique in that it's relatively easy to have lots of threads depend on a single pages lock. > Just asking because maybe that different call chain could give us some > other ideas of what the commonality here is that triggers out > behavioral problem. > I am reasonably confident that the commonality is multiple threads sharing a page. Maybe it's a single hot structure that is shared between threads. Maybe it's parallel buffers that are aligned on a sub-page boundary. Multiple threads accessing buffers aligned by cache lines would do it which is reasonable behaviour for a parallelised compute load for example. If I'm right, writing a test case for it is straight-forward and I'll get to it on Monday when I'm back near my normal work machine. The initial guess that it may be allocation latency was obviously way off. I didn't follow through properly but the fact it's not THP specific means the length of time it takes to migrate is possibly irrelevant. If the page is hot enough, threads will block once migration starts even if the migration completes quickly. > I was really hoping that we'd root-cause this and have a solution (and > then apply Tim's patch as a "belt and suspenders" kind of thing), but > it's starting to smell like we may have to apply Tim's patch as a > band-aid, and try to figure out what the trigger is longer-term. > I believe the trigger is simply because a hot page gets unmapped and then threads lock on it. One option to mitigate (but not eliminate) the problem is to record when the page lock is contended and pass in TNF_PAGE_CONTENDED (new flag) to task_numa_fault(). For each time it's passed in, shift numa_scan_period << 1 which will slow the scanner and reduce the frequency contention occurs at. If it's heavily contended, the period will quickly reach numa_scan_period_max. That is simple with the caveat that a single hot contended page will slow all balancing. The main problem is that this mitigates and not eliminates the problem. No matter how slow the scanner is, it'll still hit the case where many threads contend on a single page. An alternative is to set a VMA flag on VMAs if many contentions are detected and stop scanning that VMA entirely. It would need a VMA flag which right now might mean making vm_flags u64 and increasing the size of vm_area_struct on 32-bit. The downside is that it is permanent. If heavy contention happens then scanning that VMA is disabled for the lifetime of the process because there would no longer be a way to detect that re-enabling is appropriate. A variation would be to record contentions in struct numa_group and return false in should_numa_migrate_memory if contentions are high and scale it down over time. It wouldn't be perfect as processes sharing hot pages are not guaranteed to have a numa_group in common but it may be sufficient. -- Mel Gorman SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-08-18 21:20 +0200 |
| Message-ID | <ufUYi-7Um-3@gated-at.bofh.it> |
| In reply to | #1715373 |
[Multipart message — attachments visible in raw view] — view raw
On Fri, Aug 18, 2017 at 11:54 AM, Mel Gorman
<mgorman@techsingularity.net> wrote:
>
> One option to mitigate (but not eliminate) the problem is to record when
> the page lock is contended and pass in TNF_PAGE_CONTENDED (new flag) to
> task_numa_fault().
Well, finding it contended is fairly easy - just look at the page wait
queue, and if it's not empty, assume it's due to contention.
I also wonder if we could be even *more* hacky, and in the whole
__migration_entry_wait() path, change the logic from:
- wait on page lock before retrying the fault
to
- yield()
which is hacky, but there's a rationale for it:
(a) avoid the crazy long wait queues ;)
(b) we know that migration is *supposed* to be CPU-bound (not IO
bound), so yielding the CPU and retrying may just be the right thing
to do.
It's possible that we could just do a hybrid approach, and introduce a
"wait_on_page_lock_or_yield()", that does a sleeping wait if the
wait-queue is short, and a yield otherwise, but it might be worth just
testing the truly stupid patch.
Because that code sequence doesn't actually depend on
"wait_on_page_lock()" for _correctness_ anyway, afaik. Anybody who
does "migration_entry_wait()" _has_ to retry anyway, since the page
table contents may have changed by waiting.
So I'm not proud of the attached patch, and I don't think it's really
acceptable as-is, but maybe it's worth testing? And maybe it's
arguably no worse than what we have now?
Comments?
(Yeah, if we take this approach, we might even just say "screw the
spinlock - just do ACCESS_ONCE() and do a yield() if it looks like a
migration entry")
Linus
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <ak@linux.intel.com> |
|---|---|
| Date | 2017-08-18 22:00 +0200 |
| Message-ID | <ufVB0-89n-7@gated-at.bofh.it> |
| In reply to | #1715390 |
> which is hacky, but there's a rationale for it: > > (a) avoid the crazy long wait queues ;) > > (b) we know that migration is *supposed* to be CPU-bound (not IO > bound), so yielding the CPU and retrying may just be the right thing > to do. So this would degenerate into a spin when the contention is with other CPUs? But then if we guarantee that migration has flat latency curve and no long tail it may be reasonable. If the contention is with the local CPU it could cause some unfairness (and in theory priority inheritance issues with local CPU contenders?), but hopefully not too bad. -Andi
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-08-18 22:20 +0200 |
| Message-ID | <ufVUm-8uR-15@gated-at.bofh.it> |
| In reply to | #1715400 |
On Fri, Aug 18, 2017 at 12:58 PM, Andi Kleen <ak@linux.intel.com> wrote:
>> which is hacky, but there's a rationale for it:
>>
>> (a) avoid the crazy long wait queues ;)
>>
>> (b) we know that migration is *supposed* to be CPU-bound (not IO
>> bound), so yielding the CPU and retrying may just be the right thing
>> to do.
>
> So this would degenerate into a spin when the contention is with
> other CPUs?
>
> But then if we guarantee that migration has flat latency curve
> and no long tail it may be reasonable.
Honestly, right now I'd say it's more of a "poath meant purely for
testing with some weak-ass excuse for why it might not be broken".
Linus
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2017-08-21 20:40 +0200 |
| Message-ID | <ugZMf-7RS-31@gated-at.bofh.it> |
| In reply to | #1715390 |
On Fri, Aug 18, 2017 at 12:14:12PM -0700, Linus Torvalds wrote:
> On Fri, Aug 18, 2017 at 11:54 AM, Mel Gorman
> <mgorman@techsingularity.net> wrote:
> >
> > One option to mitigate (but not eliminate) the problem is to record when
> > the page lock is contended and pass in TNF_PAGE_CONTENDED (new flag) to
> > task_numa_fault().
>
> Well, finding it contended is fairly easy - just look at the page wait
> queue, and if it's not empty, assume it's due to contention.
>
Yes.
> I also wonder if we could be even *more* hacky, and in the whole
> __migration_entry_wait() path, change the logic from:
>
> - wait on page lock before retrying the fault
>
> to
>
> - yield()
>
> which is hacky, but there's a rationale for it:
>
> (a) avoid the crazy long wait queues ;)
>
> (b) we know that migration is *supposed* to be CPU-bound (not IO
> bound), so yielding the CPU and retrying may just be the right thing
> to do.
>
Potentially. I spent a few hours trying to construct a test case that
would migrate constantly that could be used as a basis for evaluating a
patch or alternative. Unfortunately it was not as easy as I thought and
I still have to construct a case that causes migration storms that would
result in multiple threads waiting on a single page.
> Because that code sequence doesn't actually depend on
> "wait_on_page_lock()" for _correctness_ anyway, afaik. Anybody who
> does "migration_entry_wait()" _has_ to retry anyway, since the page
> table contents may have changed by waiting.
>
> So I'm not proud of the attached patch, and I don't think it's really
> acceptable as-is, but maybe it's worth testing? And maybe it's
> arguably no worse than what we have now?
>
> Comments?
>
The transhuge migration path for numa balancing doesn't go through the
migration_entry_wait patch despite similarly named functions that suggest
it does so this may only has the most effect when THP is disabled. It's
worth trying anyway.
Covering both paths would be something like the patch below which spins
until the page is unlocked or it should reschedule. It's not even boot
tested as I spent what time I had on the test case that I hoped would be
able to prove it really works.
diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
index 79b36f57c3ba..31cda1288176 100644
--- a/include/linux/pagemap.h
+++ b/include/linux/pagemap.h
@@ -517,6 +517,13 @@ static inline void wait_on_page_locked(struct page *page)
wait_on_page_bit(compound_head(page), PG_locked);
}
+void __spinwait_on_page_locked(struct page *page);
+static inline void spinwait_on_page_locked(struct page *page)
+{
+ if (PageLocked(page))
+ __spinwait_on_page_locked(page);
+}
+
static inline int wait_on_page_locked_killable(struct page *page)
{
if (!PageLocked(page))
diff --git a/mm/filemap.c b/mm/filemap.c
index a49702445ce0..c9d6f49614bc 100644
--- a/mm/filemap.c
+++ b/mm/filemap.c
@@ -1210,6 +1210,15 @@ int __lock_page_or_retry(struct page *page, struct mm_struct *mm,
}
}
+void __spinwait_on_page_locked(struct page *page)
+{
+ do {
+ cpu_relax();
+ } while (PageLocked(page) && !cond_resched());
+
+ wait_on_page_locked(page);
+}
+
/**
* page_cache_next_hole - find the next hole (not-present entry)
* @mapping: mapping
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 90731e3b7e58..c7025c806420 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -1443,7 +1443,7 @@ int do_huge_pmd_numa_page(struct vm_fault *vmf, pmd_t pmd)
if (!get_page_unless_zero(page))
goto out_unlock;
spin_unlock(vmf->ptl);
- wait_on_page_locked(page);
+ spinwait_on_page_locked(page);
put_page(page);
goto out;
}
@@ -1480,7 +1480,7 @@ int do_huge_pmd_numa_page(struct vm_fault *vmf, pmd_t pmd)
if (!get_page_unless_zero(page))
goto out_unlock;
spin_unlock(vmf->ptl);
- wait_on_page_locked(page);
+ spinwait_on_page_locked(page);
put_page(page);
goto out;
}
diff --git a/mm/migrate.c b/mm/migrate.c
index e84eeb4e4356..9b6c3fc5beac 100644
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -308,7 +308,7 @@ void __migration_entry_wait(struct mm_struct *mm, pte_t *ptep,
if (!get_page_unless_zero(page))
goto out;
pte_unmap_unlock(ptep, ptl);
- wait_on_page_locked(page);
+ spinwait_on_page_locked(page);
put_page(page);
return;
out:
[toc] | [prev] | [next] | [standalone]
| From | "Liang, Kan" <kan.liang@intel.com> |
|---|---|
| Date | 2017-08-21 21:00 +0200 |
| Message-ID | <uh05z-814-15@gated-at.bofh.it> |
| In reply to | #1716784 |
> > Because that code sequence doesn't actually depend on
> > "wait_on_page_lock()" for _correctness_ anyway, afaik. Anybody who
> > does "migration_entry_wait()" _has_ to retry anyway, since the page
> > table contents may have changed by waiting.
> >
> > So I'm not proud of the attached patch, and I don't think it's really
> > acceptable as-is, but maybe it's worth testing? And maybe it's
> > arguably no worse than what we have now?
> >
> > Comments?
> >
>
> The transhuge migration path for numa balancing doesn't go through the
> migration_entry_wait patch despite similarly named functions that suggest
> it does so this may only has the most effect when THP is disabled. It's
> worth trying anyway.
I just finished the test of yield patch (only functionality not performance).
Yes, it works well with THP disabled.
With THP enabled, I observed one LOCKUP caused by long queue wait.
Here is the call stack with THP enabled.
#
100.00% (ffffffff9e1aefca)
|
---wait_on_page_bit
do_huge_pmd_numa_page
__handle_mm_fault
handle_mm_fault
__do_page_fault
do_page_fault
page_fault
|
|--60.39%--0x2b7b7
| |
| |--34.26%--0x127d8
| | start_thread
| |
| --25.95%--0x127a2
| start_thread
|
--39.25%--0x2b788
|
--38.81%--0x127a2
start_thread
>
> Covering both paths would be something like the patch below which spins
> until the page is unlocked or it should reschedule. It's not even boot
> tested as I spent what time I had on the test case that I hoped would be
> able to prove it really works.
I will give it a try.
Thanks,
Kan
>
> diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h
> index 79b36f57c3ba..31cda1288176 100644
> --- a/include/linux/pagemap.h
> +++ b/include/linux/pagemap.h
> @@ -517,6 +517,13 @@ static inline void wait_on_page_locked(struct page
> *page)
> wait_on_page_bit(compound_head(page), PG_locked);
> }
>
> +void __spinwait_on_page_locked(struct page *page);
> +static inline void spinwait_on_page_locked(struct page *page)
> +{
> + if (PageLocked(page))
> + __spinwait_on_page_locked(page);
> +}
> +
> static inline int wait_on_page_locked_killable(struct page *page)
> {
> if (!PageLocked(page))
> diff --git a/mm/filemap.c b/mm/filemap.c
> index a49702445ce0..c9d6f49614bc 100644
> --- a/mm/filemap.c
> +++ b/mm/filemap.c
> @@ -1210,6 +1210,15 @@ int __lock_page_or_retry(struct page *page,
> struct mm_struct *mm,
> }
> }
>
> +void __spinwait_on_page_locked(struct page *page)
> +{
> + do {
> + cpu_relax();
> + } while (PageLocked(page) && !cond_resched());
> +
> + wait_on_page_locked(page);
> +}
> +
> /**
> * page_cache_next_hole - find the next hole (not-present entry)
> * @mapping: mapping
> diff --git a/mm/huge_memory.c b/mm/huge_memory.c
> index 90731e3b7e58..c7025c806420 100644
> --- a/mm/huge_memory.c
> +++ b/mm/huge_memory.c
> @@ -1443,7 +1443,7 @@ int do_huge_pmd_numa_page(struct vm_fault
> *vmf, pmd_t pmd)
> if (!get_page_unless_zero(page))
> goto out_unlock;
> spin_unlock(vmf->ptl);
> - wait_on_page_locked(page);
> + spinwait_on_page_locked(page);
> put_page(page);
> goto out;
> }
> @@ -1480,7 +1480,7 @@ int do_huge_pmd_numa_page(struct vm_fault
> *vmf, pmd_t pmd)
> if (!get_page_unless_zero(page))
> goto out_unlock;
> spin_unlock(vmf->ptl);
> - wait_on_page_locked(page);
> + spinwait_on_page_locked(page);
> put_page(page);
> goto out;
> }
> diff --git a/mm/migrate.c b/mm/migrate.c
> index e84eeb4e4356..9b6c3fc5beac 100644
> --- a/mm/migrate.c
> +++ b/mm/migrate.c
> @@ -308,7 +308,7 @@ void __migration_entry_wait(struct mm_struct *mm,
> pte_t *ptep,
> if (!get_page_unless_zero(page))
> goto out;
> pte_unmap_unlock(ptep, ptl);
> - wait_on_page_locked(page);
> + spinwait_on_page_locked(page);
> put_page(page);
> return;
> out:
>
[toc] | [prev] | [next] | [standalone]
| From | "Liang, Kan" <kan.liang@intel.com> |
|---|---|
| Date | 2017-08-22 19:30 +0200 |
| Message-ID | <uhla1-5tg-1@gated-at.bofh.it> |
| In reply to | #1716786 |
> > Covering both paths would be something like the patch below which
> > spins until the page is unlocked or it should reschedule. It's not
> > even boot tested as I spent what time I had on the test case that I
> > hoped would be able to prove it really works.
>
> I will give it a try.
Although the patch doesn't trigger watchdog, the spin lock wait time
is not small (0.45s).
It may get worse again on larger systems.
Irqsoff ftrace result.
# tracer: irqsoff
#
# irqsoff latency trace v1.1.5 on 4.13.0-rc4+
# --------------------------------------------------------------------
# latency: 451753 us, #4/4, CPU#159 | (M:desktop VP:0, KP:0, SP:0 HP:0 #P:224)
# -----------------
# | task: fjsctest-233851 (uid:0 nice:0 policy:0 rt_prio:0)
# -----------------
# => started at: wake_up_page_bit
# => ended at: wake_up_page_bit
#
#
# _------=> CPU#
# / _-----=> irqs-off
# | / _----=> need-resched
# || / _---=> hardirq/softirq
# ||| / _--=> preempt-depth
# |||| / delay
# cmd pid ||||| time | caller
# \ / ||||| \ | /
<...>-233851 159d... 0us@: _raw_spin_lock_irqsave <-wake_up_page_bit
<...>-233851 159dN.. 451726us+: _raw_spin_unlock_irqrestore <-wake_up_page_bit
<...>-233851 159dN.. 451754us!: trace_hardirqs_on <-wake_up_page_bit
<...>-233851 159dN.. 451873us : <stack trace>
=> unlock_page
=> migrate_pages
=> migrate_misplaced_page
=> __handle_mm_fault
=> handle_mm_fault
=> __do_page_fault
=> do_page_fault
=> page_fault
The call stack of wait_on_page_bit_common
100.00% (ffffffff971b252b)
|
---__spinwait_on_page_locked
|
|--96.81%--__migration_entry_wait
| migration_entry_wait
| do_swap_page
| __handle_mm_fault
| handle_mm_fault
| __do_page_fault
| do_page_fault
| page_fault
| |
| |--22.49%--0x123a2
| | |
| | --22.34%--start_thread
| |
| |--15.69%--0x127bc
| | |
| | --13.20%--start_thread
| |
| |--13.48%--0x12352
| | |
| | --11.74%--start_thread
| |
| |--13.43%--0x127f2
| | |
| | --11.25%--start_thread
| |
| |--10.03%--0x1285e
| | |
| | --8.59%--start_thread
| |
| |--5.90%--0x12894
| | |
| | --5.03%--start_thread
| |
| |--5.66%--0x12828
| | |
| | --4.81%--start_thread
| |
| |--5.17%--0x1233c
| | |
| | --4.46%--start_thread
| |
| --4.72%--0x2b788
| |
| --4.72%--0x127a2
| start_thread
|
--3.19%--do_huge_pmd_numa_page
__handle_mm_fault
handle_mm_fault
__do_page_fault
do_page_fault
page_fault
0x2b788
0x127a2
start_thread
>
> >
> > diff --git a/include/linux/pagemap.h b/include/linux/pagemap.h index
> > 79b36f57c3ba..31cda1288176 100644
> > --- a/include/linux/pagemap.h
> > +++ b/include/linux/pagemap.h
> > @@ -517,6 +517,13 @@ static inline void wait_on_page_locked(struct
> > page
> > *page)
> > wait_on_page_bit(compound_head(page), PG_locked); }
> >
> > +void __spinwait_on_page_locked(struct page *page); static inline void
> > +spinwait_on_page_locked(struct page *page) {
> > + if (PageLocked(page))
> > + __spinwait_on_page_locked(page);
> > +}
> > +
> > static inline int wait_on_page_locked_killable(struct page *page) {
> > if (!PageLocked(page))
> > diff --git a/mm/filemap.c b/mm/filemap.c index
> > a49702445ce0..c9d6f49614bc 100644
> > --- a/mm/filemap.c
> > +++ b/mm/filemap.c
> > @@ -1210,6 +1210,15 @@ int __lock_page_or_retry(struct page *page,
> > struct mm_struct *mm,
> > }
> > }
> >
> > +void __spinwait_on_page_locked(struct page *page) {
> > + do {
> > + cpu_relax();
> > + } while (PageLocked(page) && !cond_resched());
> > +
> > + wait_on_page_locked(page);
> > +}
> > +
> > /**
> > * page_cache_next_hole - find the next hole (not-present entry)
> > * @mapping: mapping
> > diff --git a/mm/huge_memory.c b/mm/huge_memory.c index
> > 90731e3b7e58..c7025c806420 100644
> > --- a/mm/huge_memory.c
> > +++ b/mm/huge_memory.c
> > @@ -1443,7 +1443,7 @@ int do_huge_pmd_numa_page(struct vm_fault
> *vmf,
> > pmd_t pmd)
> > if (!get_page_unless_zero(page))
> > goto out_unlock;
> > spin_unlock(vmf->ptl);
> > - wait_on_page_locked(page);
> > + spinwait_on_page_locked(page);
> > put_page(page);
> > goto out;
> > }
> > @@ -1480,7 +1480,7 @@ int do_huge_pmd_numa_page(struct vm_fault
> *vmf,
> > pmd_t pmd)
> > if (!get_page_unless_zero(page))
> > goto out_unlock;
> > spin_unlock(vmf->ptl);
> > - wait_on_page_locked(page);
> > + spinwait_on_page_locked(page);
> > put_page(page);
> > goto out;
> > }
> > diff --git a/mm/migrate.c b/mm/migrate.c index
> > e84eeb4e4356..9b6c3fc5beac 100644
> > --- a/mm/migrate.c
> > +++ b/mm/migrate.c
> > @@ -308,7 +308,7 @@ void __migration_entry_wait(struct mm_struct
> *mm,
> > pte_t *ptep,
> > if (!get_page_unless_zero(page))
> > goto out;
> > pte_unmap_unlock(ptep, ptl);
> > - wait_on_page_locked(page);
> > + spinwait_on_page_locked(page);
> > put_page(page);
> > return;
> > out:
> >
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-08-22 20:20 +0200 |
| Message-ID | <uhlWp-637-5@gated-at.bofh.it> |
| In reply to | #1717606 |
[Multipart message — attachments visible in raw view] — view raw
On Tue, Aug 22, 2017 at 10:23 AM, Liang, Kan <kan.liang@intel.com> wrote:
>
> Although the patch doesn't trigger watchdog, the spin lock wait time
> is not small (0.45s).
> It may get worse again on larger systems.
Yeah, I don't think Mel's patch is great - because I think we could do
so much better.
What I like about Mel's patch is that it recognizes that
"wait_on_page_locked()" there is special, and replaces it with
something else. I think that "something else" is worse than my
"yield()" call, though.
In particular, it wastes CPU time even in the good case, and the
process that will unlock the page may actually be waiting for us to
reschedule. It may be CPU bound, but it might well have just been
preempted out.
So if we do busy loops, I really think we should also make sure that
the thing we're waiting for is not preempted.
HOWEVER, I'm actually starting to think that there is perhaps
something else going on.
Let me walk you through my thinking:
This is the migration logic:
(a) migration locks the page
(b) migration is supposedly CPU-limited
(c) migration then unlocks the page.
Ignore all the details, that's the 10.000 ft view. Right?
Now, if the above is right, then I have a question for people:
HOW IN THE HELL DO WE HAVE TIME FOR THOUSANDS OF THREADS TO HIT THAT ONE PAGE?
That just sounds really sketchy to me. Even if all those thousands of
threads are runnable, we need to schedule into them just to get them
to wait on that one page.
So that sounds really quite odd when migration is supposed to hold the
page lock for a relatively short time and get out. Don't you agree?
Which is why I started thinking of what the hell could go on for that
long wait-queue to happen.
One thing that strikes me is that the way wait_on_page_bit() works is
that it will NOT wait until the next bit clearing, it will wait until
it actively *sees* the page bit being clear.
Now, work with me on that. What's the difference?
What we could have is some bad NUMA balancing pattern that actually
has a page that everybody touches.
And hey, we pretty much know that everybody touches that page, since
people get stuck on that wait-queue, right?
And since everybody touches it, as a result everybody eventually
thinks that page should be migrated to their NUMA node.
But for all we know, the migration keeps on failing, because one of
the points of that "lock page - try to move - unlock page" is that
*TRY* in "try to move". There's a number of things that makes it not
actually migrate. Like not being movable, or failing to isolate the
page, or whatever.
So we could have some situation where we end up locking and unlocking
the page over and over again (which admittedly is already a sign of
something wrong in the NUMA balancing, but that's a separate issue).
And if we get into that situation, where everybody wants that one hot
page, what happens to the waiters?
One of the thousands of waiters is unlucky (remember, this argument
started with the whole "you shouldn't get that many waiters on one
single page that isn't even locked for that long"), and goes:
(a) Oh, the page is locked, I will wait for the lock bit to clear
(b) go to sleep
(c) the migration fails, the lock bit is cleared, the waiter is woken
up but doesn't get the CPU immediately, and one of the other
*thousands* of threads decides to also try to migrate (see above),
(d) the guy waiting for the lock bit to clear will see the page
"still" locked (really just "locked again") and continue to wait.
In the meantime, one of the other threads happens to be unlucky, also
hits the race, and now we have one more thread waiting for that page
lock. It keeps getting unlocked, but it also keeps on getting locked,
and so the queue can keep growing.
See where I'm going here? I think it's really odd how *thousands* of
threads can hit that locked window that is supposed to be pretty
small. But I think it's much more likely if we have some kind of
repeated event going on.
So I'm starting to think that part of the problem may be how stupid
that "wait_for_page_bit_common()" code is. It really shouldn't wait
until it sees that the bit is clear. It could have been cleared and
then re-taken.
And honestly, we actually have extra code for that "let's go round
again". That seems pointless. If the bit has been cleared, we've been
woken up, and nothing else would have done so anyway, so if we're not
interested in locking, we're simply *done* after we've done the
"io_scheduler()".
So I propose testing the attached trivial patch. It may not do
anything at all. But the existing code is actually doing extra work
just to be fragile, in case the scenario above can happen.
Comments?
Linus
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-08-22 20:30 +0200 |
| Message-ID | <uhm66-66D-11@gated-at.bofh.it> |
| In reply to | #1717645 |
On Tue, Aug 22, 2017 at 11:19 AM, Linus Torvalds
<torvalds@linux-foundation.org> wrote:
>
> So I propose testing the attached trivial patch. It may not do
> anything at all. But the existing code is actually doing extra work
> just to be fragile, in case the scenario above can happen.
Side note: the patch compiles for me. But that is literally ALL the
testing it has gotten. I spent more time writing that email trying to
explain what my thinking was about that patch, than I spent anywhere
else on that patch.
So it may be garbage. Caveat probator.
Linus
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-08-22 21:00 +0200 |
| Message-ID | <uhmz8-6kj-13@gated-at.bofh.it> |
| In reply to | #1717645 |
On Tue, Aug 22, 2017 at 11:19:12AM -0700, Linus Torvalds wrote:
> diff --git a/mm/filemap.c b/mm/filemap.c
> index a49702445ce0..75c29a1f90fb 100644
> --- a/mm/filemap.c
> +++ b/mm/filemap.c
> @@ -991,13 +991,11 @@ static inline int wait_on_page_bit_common(wait_queue_head_t *q,
> }
> }
>
> - if (lock) {
> - if (!test_and_set_bit_lock(bit_nr, &page->flags))
> - break;
> - } else {
> - if (!test_bit(bit_nr, &page->flags))
> - break;
> - }
> + if (!lock)
> + break;
> +
> + if (!test_and_set_bit_lock(bit_nr, &page->flags))
> + break;
> }
Won't we now prematurely terminate the wait when we get a spurious
wakeup?
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-08-22 21:20 +0200 |
| Message-ID | <uhmSx-6HP-81@gated-at.bofh.it> |
| In reply to | #1717666 |
On Tue, Aug 22, 2017 at 11:56 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>
> Won't we now prematurely terminate the wait when we get a spurious
> wakeup?
I think there's two answers to that:
(a) do we even care?
(b) what spurious wakeup?
The "do we even care" quesiton is because wait_on_page_bit by
definition isn't really serializing. And I'm not even talking about
memory ordering, altough that is true too - I'm talking just
fundamentally, that by definition when we're not locking, by the time
wait_on_page_bit() returns to the caller, it could obviously have
changed again.
So I think wait_on_page_bit() is by definition not really guaranteeing
that the bit really is clear. And I don't think we have really have
cases that matter.
But if we do - say, 'fsync()' waiting for a page to wait for
writeback, where would you get spurious wakeups from? They normally
happen either when we have nested waiting (eg a page fault happens
while we have other wait queues active), and I'm not seeing that being
an issue here.
That said, I do think we might want to perhaps make a "careful" vs
"just wait a bit" version of this if the patch works out.
The patch is primarily for testing this particular case. I actually
think it's probably ok in general, but maybe there really is some
special case that could have multiple wakeup sources and it needs to
see *this* particular one.
(We could perhaps handle that case by checking "is the wait-queue
empty now" instead, and just get rid of the re-arming, not break out
of the loop immediately after the io_schedule()).
Linus
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2017-08-22 21:10 +0200 |
| Message-ID | <uhmIO-6Dr-25@gated-at.bofh.it> |
| In reply to | #1717645 |
On Tue, Aug 22, 2017 at 11:19:12AM -0700, Linus Torvalds wrote: > And since everybody touches it, as a result everybody eventually > thinks that page should be migrated to their NUMA node. So that migration stuff has a filter on, we need two consecutive numa faults from the same page_cpupid 'hash', see should_numa_migrate_memory(). And since this appears to be anonymous memory (no THP) this is all a single address space. However, we don't appear to invalidate TLBs when we upgrade the PTE protection bits (not strictly required of course), so we can have multiple CPUs trip over the same 'old' NUMA PTE. Still, generating such a migration storm would be fairly tricky I think.
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <ak@linux.intel.com> |
|---|---|
| Date | 2017-08-22 21:40 +0200 |
| Message-ID | <uhnbP-6Qf-9@gated-at.bofh.it> |
| In reply to | #1717677 |
> > Still, generating such a migration storm would be fairly tricky I think. > > Well, Mel seems to have been unable to generate a load that reproduces > the long page waitqueues. And I don't think we've had any other > reports of this either. It could be that it requires a fairly large system. On large systems under load a lot of things take much longer, so what's a tiny window on Mel's system may suddenly be very large, and with much more threads they have a higher chance of bad interactions anyways. We only see it on 4S+ today. But systems are always getting larger, so what's a large system today, will be a normal medium scale system tomorrow. BTW we also collected PT traces for the long hang cases, but it was hard to find a consistent pattern in them. -Andi
[toc] | [prev] | [next] | [standalone]
| From | Christopher Lameter <cl@linux.com> |
|---|---|
| Date | 2017-08-22 23:10 +0200 |
| Message-ID | <uhoAW-7V4-27@gated-at.bofh.it> |
| In reply to | #1717769 |
On Tue, 22 Aug 2017, Andi Kleen wrote: > We only see it on 4S+ today. But systems are always getting larger, > so what's a large system today, will be a normal medium scale system > tomorrow. > > BTW we also collected PT traces for the long hang cases, but it was > hard to find a consistent pattern in them. Hmmm... Maybe it would be wise to limit the pages autonuma can migrate? If a page has more than 50 refcounts or so then dont migrate it. I think high number of refcounts and a high frequewncy of calls are reached in particular for pages of the c library. Attempting to migrate those does not make much sense anyways because the load may shift and another function may become popular. We may end up shifting very difficult to migrate pages back and forth.
[toc] | [prev] | [next] | [standalone]
| From | Andi Kleen <ak@linux.intel.com> |
|---|---|
| Date | 2017-08-22 23:30 +0200 |
| Message-ID | <uhoUi-82K-25@gated-at.bofh.it> |
| In reply to | #1717842 |
On Tue, Aug 22, 2017 at 04:08:52PM -0500, Christopher Lameter wrote: > On Tue, 22 Aug 2017, Andi Kleen wrote: > > > We only see it on 4S+ today. But systems are always getting larger, > > so what's a large system today, will be a normal medium scale system > > tomorrow. > > > > BTW we also collected PT traces for the long hang cases, but it was > > hard to find a consistent pattern in them. > > Hmmm... Maybe it would be wise to limit the pages autonuma can migrate? > > If a page has more than 50 refcounts or so then dont migrate it. I think > high number of refcounts and a high frequewncy of calls are reached in > particular for pages of the c library. Attempting to migrate those does > not make much sense anyways because the load may shift and another > function may become popular. We may end up shifting very difficult to > migrate pages back and forth. I believe in this case it's used by threads, so a reference count limit wouldn't help. If migrating code was a problem I would probably rather just disable migration of read-only pages. -Andi
[toc] | [prev] | [next] | [standalone]
Page 2 of 4 — ← Prev page 1 [2] 3 4 Next page →
Back to top | Article view | linux.kernel
csiph-web