Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1223064 > unrolled thread
| Started by | Chris Mason <clm@fb.com> |
|---|---|
| First post | 2015-09-11 21:40 +0200 |
| Last post | 2015-09-12 01:20 +0200 |
| Articles | 13 on this page of 53 — 9 participants |
Back to article view | Back to linux.kernel
[PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-11 21:40 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-11 22:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-11 22:40 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Josef Bacik <jbacik@fb.com> - 2015-09-11 22:50 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-11 23:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-12 00:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-12 01:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-12 01:40 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-12 03:00 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-12 04:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-12 04:30 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-13 01:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-13 01:30 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-13 01:50 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-13 15:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Dave Chinner <david@fromorbit.com> - 2015-09-14 01:00 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Dave Chinner <david@fromorbit.com> - 2015-09-14 01:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-14 22:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Jan Kara <jack@suse.cz> - 2015-09-16 22:00 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-16 22:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Dave Chinner <david@fromorbit.com> - 2015-09-17 00:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Dave Chinner <david@fromorbit.com> - 2015-09-17 02:40 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-17 03:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Dave Chinner <david@fromorbit.com> - 2015-09-17 04:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-17 21:40 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-18 00:50 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-18 01:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-18 02:00 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Dave Chinner <david@fromorbit.com> - 2015-09-18 02:40 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-18 04:00 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Dave Chinner <david@fromorbit.com> - 2015-09-18 07:50 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-18 08:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-18 08:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Jens Axboe <axboe@fb.com> - 2015-09-18 16:30 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-18 15:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Jens Axboe <axboe@fb.com> - 2015-09-18 16:30 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-18 17:40 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Peter Zijlstra <peterz@infradead.org> - 2015-09-18 18:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Peter Zijlstra <peterz@infradead.org> - 2015-09-18 18:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-18 18:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Peter Zijlstra <peterz@infradead.org> - 2015-09-28 16:50 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-28 18:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Ingo Molnar <mingo@kernel.org> - 2015-09-29 10:00 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Dave Chinner <david@fromorbit.com> - 2015-09-19 01:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Jan Kara <jack@suse.cz> - 2015-09-21 11:30 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Andrew Morton <akpm@linux-foundation.org> - 2015-09-21 22:30 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Dave Chinner <david@fromorbit.com> - 2015-09-18 01:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-18 01:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-17 05:50 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Dave Chinner <david@fromorbit.com> - 2015-09-17 06:40 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-17 14:20 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Chris Mason <clm@fb.com> - 2015-09-12 01:10 +0200
Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() Linus Torvalds <torvalds@linux-foundation.org> - 2015-09-12 01:20 +0200
Page 3 of 3 — ← Prev page 1 2 [3]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2015-09-28 16:50 +0200 |
| Message-ID | <qdHUB-2nN-7@gated-at.bofh.it> |
| In reply to | #1228130 |
On Fri, Sep 18, 2015 at 09:12:38AM -0700, Linus Torvalds wrote: > > So I disagree with your notion that it's a recursion flag. It is > absolutely nothing of the sort. OK, agreed. I had it classed under recursion in my head, clearly I indexed it sloppily. In any case I have a patch that kills off PREEMPT_ACTIVE entirely. I just have to clean it up, benchmark, split and write changelogs. But it should be forthcoming 'soon'. As is, it boots.. > It gets set by preemption - and, > somewhat illogically, by cond_resched(). I suspect that was done to make cond_resched() (voluntary preemption) more robust and only have a single preemption path/logic. But all that was done well before I got involved. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2015-09-28 18:10 +0200 |
| Message-ID | <qdJa1-4l8-7@gated-at.bofh.it> |
| In reply to | #1234220 |
On Mon, Sep 28, 2015 at 10:47 AM, Peter Zijlstra <peterz@infradead.org> wrote:
>
>> It gets set by preemption - and,
>> somewhat illogically, by cond_resched().
>
> I suspect that was done to make cond_resched() (voluntary preemption)
> more robust and only have a single preemption path/logic. But all that
> was done well before I got involved.
So I think it's actually the name that is bad, not necessarily the behavior.
We tend to put "cond_resched()" (and particularly
"cond_resched_lock()") in some fairly awkward places, and it's not
always entirely clear that task->state == TASK_RUNNING there.
So the preemptive behavior of not *really* putting the task to sleep
may actually be the right one. But it is rather non-intuitive given
the name - because "cond_resched()" basically is not at all equivalent
to "if (need_resched()) schedule()", which you'd kind of expect.
An explicit schedule will actually act on the task->state, and make us
go to sleep. "cond_resched()" really is just a "voluntary preemption
point". And I think it would be better if it got named that way.
Linus
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2015-09-29 10:00 +0200 |
| Message-ID | <qdXZo-ds-11@gated-at.bofh.it> |
| In reply to | #1234269 |
* Linus Torvalds <torvalds@linux-foundation.org> wrote: > On Mon, Sep 28, 2015 at 10:47 AM, Peter Zijlstra <peterz@infradead.org> wrote: > > > >> It gets set by preemption - and, > >> somewhat illogically, by cond_resched(). > > > > I suspect that was done to make cond_resched() (voluntary preemption) > > more robust and only have a single preemption path/logic. But all that > > was done well before I got involved. > > So I think it's actually the name that is bad, not necessarily the behavior. > > We tend to put "cond_resched()" (and particularly > "cond_resched_lock()") in some fairly awkward places, and it's not > always entirely clear that task->state == TASK_RUNNING there. > > So the preemptive behavior of not *really* putting the task to sleep > may actually be the right one. But it is rather non-intuitive given > the name - because "cond_resched()" basically is not at all equivalent > to "if (need_resched()) schedule()", which you'd kind of expect. > > An explicit schedule will actually act on the task->state, and make us > go to sleep. "cond_resched()" really is just a "voluntary preemption > point". And I think it would be better if it got named that way. cond_preempt() perhaps? That would allude to preempt_schedule() and such, and would make it clearer that it's supposed to be an invariant on the sleep state (which schedule() is not). Thanks, Ingo -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2015-09-19 01:20 +0200 |
| Message-ID | <qad6F-ub-1@gated-at.bofh.it> |
| In reply to | #1227583 |
On Thu, Sep 17, 2015 at 11:04:03PM -0700, Linus Torvalds wrote: > On Thu, Sep 17, 2015 at 10:40 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > Ok, makes sense - the plug is not being flushed as we switch away, > > but Chris' patch makes it do that. > > Yup. > > And I actually think Chris' patch is better than the one I sent out > (but maybe the scheduler people should take a look at the behavior of > cond_resched()), I just wanted you to test that to verify the > behavior. > > The fact that Chris' patch ends up lowering the context switches > (because it does the unplugging directly) is also an argument for his > approach. > > I just wanted to understand the oddity with kblockd_workqueue. And I > think that's solved. > > > Context switches go back to the 4-4500/sec range. Otherwise > > behaviour and performance is indistinguishable from Chris' patch. > > .. this was exactly what I wanted to hear. So it sounds like we have > no odd unexplained behavior left in this area. > > Which is not to say that there wouldn't be room for improvement, but > it just makes me much happier about the state of these patches to feel > like we understand what was going on. > > > PS: just hit another "did this just get broken in 4.3-rc1" issue - I > > can't run blktrace while there's a IO load because: > > > > $ sudo blktrace -d /dev/vdc > > BLKTRACESETUP(2) /dev/vdc failed: 5/Input/output error > > Thread 1 failed open /sys/kernel/debug/block/(null)/trace1: 2/No such file or directory > > .... > > > > [ 641.424618] blktrace: page allocation failure: order:5, mode:0x2040d0 > > [ 641.438933] [<ffffffff811c1569>] kmem_cache_alloc_trace+0x129/0x400 > > [ 641.440240] [<ffffffff811424f8>] relay_open+0x68/0x2c0 > > [ 641.441299] [<ffffffff8115deb1>] do_blk_trace_setup+0x191/0x2d0 > > > > gdb) l *(relay_open+0x68) > > 0xffffffff811424f8 is in relay_open (kernel/relay.c:582). > > 577 return NULL; > > 578 if (subbuf_size > UINT_MAX / n_subbufs) > > 579 return NULL; > > 580 > > 581 chan = kzalloc(sizeof(struct rchan), GFP_KERNEL); > > 582 if (!chan) > > 583 return NULL; > > 584 > > 585 chan->version = RELAYFS_CHANNEL_VERSION; > > 586 chan->n_subbufs = n_subbufs; > > > > and struct rchan has a member struct rchan_buf *buf[NR_CPUS]; > > and CONFIG_NR_CPUS=8192, hence the attempt at an order 5 allocation > > that fails here.... > > Hm. Have you always had MAX_SMP (and the NR_CPU==8192 that it causes)? > From a quick check, none of this code seems to be new. Yes, I always build MAX_SMP kernels for testing, because XFS is often used on such machines and so I want to find issues exactly like this in my testing rather than on customer machines... :/ > That said, having that > > struct rchan_buf *buf[NR_CPUS]; > > in "struct rchan" really is something we should fix. We really should > strive to not allocate things by CONFIG_NR_CPU's, but by the actual > real CPU count. *nod*. But it doesn't fix the problem of the memory allocation failing when there's still gigabytes of immediately reclaimable memory available in the page cache. If this is failing under page cache memory pressure, then we're going to be doing an awful lot more falling back to vmalloc in the filesystem code where large allocations like this are done e.g. extended attribute buffers are order-5, and used a lot when doing things like backups which tend to also produce significant page cache memory pressure. Hence I'm tending towards there being a memory reclaim behaviour regression, not so much worrying about whether this specific allocation is optimal or not. Cheers, Dave. -- Dave Chinner david@fromorbit.com -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jan Kara <jack@suse.cz> |
|---|---|
| Date | 2015-09-21 11:30 +0200 |
| Message-ID | <qb5A8-2vD-41@gated-at.bofh.it> |
| In reply to | #1228339 |
On Sat 19-09-15 08:17:14, Dave Chinner wrote: > On Thu, Sep 17, 2015 at 11:04:03PM -0700, Linus Torvalds wrote: > > On Thu, Sep 17, 2015 at 10:40 PM, Dave Chinner <david@fromorbit.com> wrote: > > > PS: just hit another "did this just get broken in 4.3-rc1" issue - I > > > can't run blktrace while there's a IO load because: > > > > > > $ sudo blktrace -d /dev/vdc > > > BLKTRACESETUP(2) /dev/vdc failed: 5/Input/output error > > > Thread 1 failed open /sys/kernel/debug/block/(null)/trace1: 2/No such file or directory > > > .... > > > > > > [ 641.424618] blktrace: page allocation failure: order:5, mode:0x2040d0 > > > [ 641.438933] [<ffffffff811c1569>] kmem_cache_alloc_trace+0x129/0x400 > > > [ 641.440240] [<ffffffff811424f8>] relay_open+0x68/0x2c0 > > > [ 641.441299] [<ffffffff8115deb1>] do_blk_trace_setup+0x191/0x2d0 > > > > > > gdb) l *(relay_open+0x68) > > > 0xffffffff811424f8 is in relay_open (kernel/relay.c:582). > > > 577 return NULL; > > > 578 if (subbuf_size > UINT_MAX / n_subbufs) > > > 579 return NULL; > > > 580 > > > 581 chan = kzalloc(sizeof(struct rchan), GFP_KERNEL); > > > 582 if (!chan) > > > 583 return NULL; > > > 584 > > > 585 chan->version = RELAYFS_CHANNEL_VERSION; > > > 586 chan->n_subbufs = n_subbufs; > > > > > > and struct rchan has a member struct rchan_buf *buf[NR_CPUS]; > > > and CONFIG_NR_CPUS=8192, hence the attempt at an order 5 allocation > > > that fails here.... > > > > Hm. Have you always had MAX_SMP (and the NR_CPU==8192 that it causes)? > > From a quick check, none of this code seems to be new. > > Yes, I always build MAX_SMP kernels for testing, because XFS is > often used on such machines and so I want to find issues exactly > like this in my testing rather than on customer machines... :/ > > > That said, having that > > > > struct rchan_buf *buf[NR_CPUS]; > > > > in "struct rchan" really is something we should fix. We really should > > strive to not allocate things by CONFIG_NR_CPU's, but by the actual > > real CPU count. > > *nod*. But it doesn't fix the problem of the memory allocation > failing when there's still gigabytes of immediately reclaimable > memory available in the page cache. If this is failing under page > cache memory pressure, then we're going to be doing an awful lot > more falling back to vmalloc in the filesystem code where large > allocations like this are done e.g. extended attribute buffers are > order-5, and used a lot when doing things like backups which tend to > also produce significant page cache memory pressure. > > Hence I'm tending towards there being a memory reclaim behaviour > regression, not so much worrying about whether this specific > allocation is optimal or not. Yup, looks like a regression in reclaim. Added linux-mm folks to CC. Honza -- Jan Kara <jack@suse.com> SUSE Labs, CR -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Andrew Morton <akpm@linux-foundation.org> |
|---|---|
| Date | 2015-09-21 22:30 +0200 |
| Subject | Re: [PATCH] fs-writeback: drop wb->list_lock during blk_finish_plug() |
| Message-ID | <qbfSO-rd-25@gated-at.bofh.it> |
| In reply to | #1229143 |
On Mon, 21 Sep 2015 11:24:29 +0200 Jan Kara <jack@suse.cz> wrote: > On Sat 19-09-15 08:17:14, Dave Chinner wrote: > > On Thu, Sep 17, 2015 at 11:04:03PM -0700, Linus Torvalds wrote: > > > On Thu, Sep 17, 2015 at 10:40 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > PS: just hit another "did this just get broken in 4.3-rc1" issue - I > > > > can't run blktrace while there's a IO load because: > > > > > > > > $ sudo blktrace -d /dev/vdc > > > > BLKTRACESETUP(2) /dev/vdc failed: 5/Input/output error > > > > Thread 1 failed open /sys/kernel/debug/block/(null)/trace1: 2/No such file or directory > > > > .... > > > > > > > > [ 641.424618] blktrace: page allocation failure: order:5, mode:0x2040d0 > > > > [ 641.438933] [<ffffffff811c1569>] kmem_cache_alloc_trace+0x129/0x400 > > > > [ 641.440240] [<ffffffff811424f8>] relay_open+0x68/0x2c0 > > > > [ 641.441299] [<ffffffff8115deb1>] do_blk_trace_setup+0x191/0x2d0 > > > > > > > > gdb) l *(relay_open+0x68) > > > > 0xffffffff811424f8 is in relay_open (kernel/relay.c:582). > > > > 577 return NULL; > > > > 578 if (subbuf_size > UINT_MAX / n_subbufs) > > > > 579 return NULL; > > > > 580 > > > > 581 chan = kzalloc(sizeof(struct rchan), GFP_KERNEL); > > > > 582 if (!chan) > > > > 583 return NULL; > > > > 584 > > > > 585 chan->version = RELAYFS_CHANNEL_VERSION; > > > > 586 chan->n_subbufs = n_subbufs; > > > > > > > > and struct rchan has a member struct rchan_buf *buf[NR_CPUS]; > > > > and CONFIG_NR_CPUS=8192, hence the attempt at an order 5 allocation > > > > that fails here.... > > > > > > Hm. Have you always had MAX_SMP (and the NR_CPU==8192 that it causes)? > > > From a quick check, none of this code seems to be new. > > > > Yes, I always build MAX_SMP kernels for testing, because XFS is > > often used on such machines and so I want to find issues exactly > > like this in my testing rather than on customer machines... :/ > > > > > That said, having that > > > > > > struct rchan_buf *buf[NR_CPUS]; > > > > > > in "struct rchan" really is something we should fix. We really should > > > strive to not allocate things by CONFIG_NR_CPU's, but by the actual > > > real CPU count. > > > > *nod*. But it doesn't fix the problem of the memory allocation > > failing when there's still gigabytes of immediately reclaimable > > memory available in the page cache. If this is failing under page > > cache memory pressure, then we're going to be doing an awful lot > > more falling back to vmalloc in the filesystem code where large > > allocations like this are done e.g. extended attribute buffers are > > order-5, and used a lot when doing things like backups which tend to > > also produce significant page cache memory pressure. > > > > Hence I'm tending towards there being a memory reclaim behaviour > > regression, not so much worrying about whether this specific > > allocation is optimal or not. > > Yup, looks like a regression in reclaim. Added linux-mm folks to CC. That's going to be hard to find. Possibly Vlastimil's 5-patch series "mm, compaction: more robust check for scanners meeting", possibly Joonsoo's "mm/compaction: correct to flush migrated pages if pageblock skip happens". But probably something else :( Teach relay.c about alloc_percpu()? -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2015-09-18 01:10 +0200 |
| Message-ID | <q9Qts-1ra-11@gated-at.bofh.it> |
| In reply to | #1226629 |
On Thu, Sep 17, 2015 at 12:14:53PM +1000, Dave Chinner wrote: > On Wed, Sep 16, 2015 at 06:12:29PM -0700, Linus Torvalds wrote: > > On Wed, Sep 16, 2015 at 5:37 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > > > TL;DR: Results look really bad - not only is the plugging > > > problematic, baseline writeback performance has regressed > > > significantly. > > > > Dave, if you're testing my current -git, the other performance issue > > might still be the spinlock thing. > > I have the fix as the first commit in my local tree - it'll remain > there until I get a conflict after an update. :) > > > The plugging IO pauses are interesting, though. Plugging really > > *shouldn't* cause that kind of pauses, _regardless_ of what level it > > happens on, so I wonder if the patch ends up just exposing some really > > basic problem that just normally goes hidden. > > Right, that's what I suspect - it didn't happen on older kernels, > but we've just completely reworked the writeback code for the > control group awareness since I last looked really closely at > this... > > > Can you match up the IO wait times with just *where* it is > > waiting? Is it waiting for that inode I_SYNC thing in > > inode_sleep_on_writeback()? > > I'll do some more investigation. Ok, I'm happy to report there is actually nothing wrong with the plugging code that is your tree. I finally tracked the problem I was seeing down to a misbehaving RAID controller.[*] With that problem sorted: kernel files/s wall time 3.17 32500 5m54s 4.3-noplug 34400 5m25s 3.17-plug 52900 3m19s 4.3-badplug 60540 3m24s 4.3-rc1 56600 3m23s So the 3.17/4.3-noplug baselines so no regression - 4.3 is slightly faster. All the plugging variants show roughly the same improvement and IO behaviour. These numbers are reproducable and there are no weird performance inconsistencies during any of the 4.3-rc1 kernel runs. Hence my numbers and observed behaviour now aligns with Chris' results and so I think we can say the reworked high level plugging is behaving as we expected it to. Cheers, Dave. [*] It seems to have a dodgy battery connector, and so has been "losing" battery backup and changing the cache mode of the HBA from write back to write through. This results in changing from NVRAM performance to SSD native performance and back again. A small vibration would cause the connection to the battery to reconnect and the controller would switch back to writeback mode. The few log entries in the bios showed changes in status between a few seconds apart to minutes apart - enough for the cache status to change several times a 5-10 minute benchmark run. I didn't notice the hardware was playing up because it wasn't triggering the machine alert indicator through the bios like it's supposed to and so the visible and audible alarms were not being triggered, nor was the BMC logging the raid controller cache status changes. In the end, I noticed it by chance - during a low level test the behaviour changed very obviously as one of my dogs ran past the rack. I unplugged everything inside the server, plugged it all back in, powered it back up and fiddled with cables until I found what was causing the problem. And having done this, the BMC is now sending warnings and the audible alarm is working when the battery is disconnected... :/ -- Dave Chinner david@fromorbit.com -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2015-09-18 01:20 +0200 |
| Message-ID | <q9QD7-1CN-3@gated-at.bofh.it> |
| In reply to | #1227477 |
On Thu, Sep 17, 2015 at 4:03 PM, Dave Chinner <david@fromorbit.com> wrote:
>
> Ok, I'm happy to report there is actually nothing wrong with the
> plugging code that is your tree. I finally tracked the problem I
> was seeing down to a misbehaving RAID controller.[*]
Hey, that's great.
Just out of interest, try the patch that Chris just sent that turns
the unplugging synchronous for the special "cond_resched()" case.
If that helps your case too, let's just do it, even while we wonder
why kblockd_workqueue messes up so noticeable.
Because regardless of the kblockd_workqueue questions, I think it's
nicer to avoid the cond_resched_lock(), and do the rescheduling while
we're not holding any locks.
Linus
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Chris Mason <clm@fb.com> |
|---|---|
| Date | 2015-09-17 05:50 +0200 |
| Message-ID | <q9ymR-8es-7@gated-at.bofh.it> |
| In reply to | #1226603 |
On Thu, Sep 17, 2015 at 10:37:38AM +1000, Dave Chinner wrote:
> [cc Tejun]
>
> On Thu, Sep 17, 2015 at 08:07:04AM +1000, Dave Chinner wrote:
> > On Wed, Sep 16, 2015 at 04:00:12PM -0400, Chris Mason wrote:
> > > On Wed, Sep 16, 2015 at 09:58:06PM +0200, Jan Kara wrote:
> > > > On Wed 16-09-15 11:16:21, Chris Mason wrote:
> > > > > Short version, Linus' patch still gives bigger IOs and similar perf to
> > > > > Dave's original. I should have done the blktrace runs for 60 seconds
> > > > > instead of 30, I suspect that would even out the average sizes between
> > > > > the three patches.
> > > >
> > > > Thanks for the data Chris. So I guess we are fine with what's currently in,
> > > > right?
> > >
> > > Looks like it works well to me.
> >
> > Graph looks good, though I'll confirm it on my test rig once I get
> > out from under the pile of email and other stuff that is queued up
> > after being away for a week...
>
> I ran some tests in the background while reading other email.....
>
> TL;DR: Results look really bad - not only is the plugging
> problematic, baseline writeback performance has regressed
> significantly. We need to revert the plugging changes until the
> underlying writeback performance regressions are sorted out.
>
> In more detail, these tests were run on my usual 16p/16GB RAM
> performance test VM with storage set up as described here:
>
> https://urldefense.proofpoint.com/v1/url?u=http://permalink.gmane.org/gmane.linux.kernel/1768786&k=ZVNjlDMF0FElm4dQtryO4A%3D%3D%0A&r=6%2FL0lzzDhu0Y1hL9xm%2BQyA%3D%3D%0A&m=4Qwp5Zj8CpoMb6vOcz%2FNMQ%2Fsb0%2FamLUP1vqWgedxJL0%3D%0A&s=90b54e35a4a7fcc4bcab9e15e22c025c7c9e045541e4923500f2e3258fc1952b
>
> The test:
>
> $ ~/tests/fsmark-10-4-test-xfs.sh
> meta-data=/dev/vdc isize=512 agcount=500, agsize=268435455 blks
> = sectsz=512 attr=2, projid32bit=1
> = crc=1 finobt=1, sparse=0
> data = bsize=4096 blocks=134217727500, imaxpct=1
> = sunit=0 swidth=0 blks
> naming =version 2 bsize=4096 ascii-ci=0 ftype=1
> log =internal log bsize=4096 blocks=131072, version=2
> = sectsz=512 sunit=1 blks, lazy-count=1
> realtime =none extsz=4096 blocks=0, rtextents=0
>
> # ./fs_mark -D 10000 -S0 -n 10000 -s 4096 -L 120 -d /mnt/scratch/0 -d /mnt/scratch/1 -d /mnt/scratch/2 -d /mnt/scratch/3 -d /mnt/scratch/4 -d /mnt/scratch/5 -d /mnt/scratch/6 -d /mnt/scratch/7
> # Version 3.3, 8 thread(s) starting at Thu Sep 17 08:08:36 2015
> # Sync method: NO SYNC: Test does not issue sync() or fsync() calls.
> # Directories: Time based hash between directories across 10000 subdirectories with 180 seconds per subdirectory.
> # File names: 40 bytes long, (16 initial bytes of time stamp with 24 random bytes at end of name)
> # Files info: size 4096 bytes, written with an IO size of 16384 bytes per write
> # App overhead is time in microseconds spent in the test not doing file writing related system calls.
>
> FSUse% Count Size Files/sec App Overhead
> 0 80000 4096 106938.0 543310
> 0 160000 4096 102922.7 476362
> 0 240000 4096 107182.9 538206
> 0 320000 4096 107871.7 619821
> 0 400000 4096 99255.6 622021
> 0 480000 4096 103217.8 609943
> 0 560000 4096 96544.2 640988
> 0 640000 4096 100347.3 676237
> 0 720000 4096 87534.8 483495
> 0 800000 4096 72577.5 2556920
> 0 880000 4096 97569.0 646996
>
> <RAM fills here, sustained performance is now dependent on writeback>
I think too many variables have changed here.
My numbers:
FSUse% Count Size Files/sec App Overhead
0 160000 4096 356407.1 1458461
0 320000 4096 368755.1 1030047
0 480000 4096 358736.8 992123
0 640000 4096 361912.5 1009566
0 800000 4096 342851.4 1004152
0 960000 4096 358357.2 996014
0 1120000 4096 338025.8 1004412
0 1280000 4096 354440.3 997380
0 1440000 4096 335225.9 1000222
0 1600000 4096 278786.1 1164962
0 1760000 4096 268161.4 1205255
0 1920000 4096 259158.0 1298054
0 2080000 4096 276939.1 1219411
0 2240000 4096 252385.1 1245496
0 2400000 4096 280674.1 1189161
0 2560000 4096 290155.4 1141941
0 2720000 4096 280842.2 1179964
0 2880000 4096 272446.4 1155527
0 3040000 4096 268827.4 1235095
0 3200000 4096 251767.1 1250006
0 3360000 4096 248339.8 1235471
0 3520000 4096 267129.9 1200834
0 3680000 4096 257320.7 1244854
0 3840000 4096 233540.8 1267764
0 4000000 4096 269237.0 1216324
0 4160000 4096 249787.6 1291767
0 4320000 4096 256185.7 1253776
0 4480000 4096 257849.7 1212953
0 4640000 4096 253933.9 1181216
0 4800000 4096 263567.2 1233937
0 4960000 4096 255666.4 1231802
0 5120000 4096 257083.2 1282893
0 5280000 4096 254285.0 1229031
0 5440000 4096 265561.6 1219472
0 5600000 4096 266374.1 1229886
0 5760000 4096 241003.7 1257064
0 5920000 4096 245047.4 1298330
0 6080000 4096 254771.7 1257241
0 6240000 4096 254355.2 1261006
0 6400000 4096 254800.4 1201074
0 6560000 4096 262794.5 1234816
0 6720000 4096 248103.0 1287921
0 6880000 4096 231397.3 1291224
0 7040000 4096 227898.0 1285359
0 7200000 4096 227279.6 1296340
0 7360000 4096 232561.5 1748248
0 7520000 4096 231055.3 1169373
0 7680000 4096 245738.5 1121856
0 7840000 4096 234961.7 1147035
0 8000000 4096 243973.0 1152202
0 8160000 4096 246292.6 1169527
0 8320000 4096 249433.2 1197921
0 8480000 4096 222576.0 1253650
0 8640000 4096 239407.5 1263257
0 8800000 4096 246037.1 1218109
0 8960000 4096 242306.5 1293567
0 9120000 4096 238525.9 3745133
0 9280000 4096 269869.5 1159541
0 9440000 4096 266447.1 4794719
0 9600000 4096 265748.9 1161584
0 9760000 4096 269067.8 1149918
0 9920000 4096 248896.2 1164112
0 10080000 4096 261342.9 1174536
0 10240000 4096 254778.3 1225425
0 10400000 4096 257702.2 1211634
0 10560000 4096 233972.5 1203665
0 10720000 4096 232647.1 1197486
0 10880000 4096 242320.6 1203984
I can push the dirty threshold lower to try and make sure we end up in
the hard dirty limits but none of this is going to be related to the
plugging patch. I do see lower numbers if I let the test run even
longer, but there are a lot of things in the way that can slow it down
as the filesystem gets that big.
I'll try again with lower ratios.
[ ... ]
> The baseline of no plugging is a full 3 minutes faster than the
> plugging behaviour of Linus' patch. The IO behaviour demonstrates
> that, sustaining between 25-30,000 IOPS and throughput of
> 130-150MB/s. Hence, while Linus' patch does change the IO patterns,
> it does not result in a performance improvement like the original
> plugging patch did.
>
How consistent is this across runs?
> So I went back and had a look at my original patch, which I've been
> using locally for a couple of years and was similar to the original
> commit. It has this description from when I last updated the perf
> numbers from testing done on 3.17:
>
> | Test VM: 16p, 16GB RAM, 2xSSD in RAID0, 500TB sparse XFS filesystem,
> | metadata CRCs enabled.
> |
> | Test:
> |
> | $ ./fs_mark -D 10000 -S0 -n 10000 -s 4096 -L 120 -d
> | /mnt/scratch/0 -d /mnt/scratch/1 -d /mnt/scratch/2 -d
> | /mnt/scratch/3 -d /mnt/scratch/4 -d /mnt/scratch/5 -d
> | /mnt/scratch/6 -d /mnt/scratch/7
> |
> | Result:
> | wall sys create rate Physical write IO
> | time CPU (avg files/s) IOPS Bandwidth
> | ----- ----- ------------- ------ ---------
> | unpatched 5m54s 15m32s 32,500+/-2200 28,000 150MB/s
> | patched 3m19s 13m28s 52,900+/-1800 1,500 280MB/s
> | improvement -43.8% -13.3% +62.7% -94.6% +86.6%
>
> IOWs, what we are seeing here is that the baseline writeback
> performance has regressed quite significantly since I took these
> numbers back on 3.17. I'm running on exactly the same test setup;
> the only difference is the kernel and so the current kernel baseline
> is ~20% slower than the baseline numbers I have in my patch.
All of this in a VM, I'd much rather see this reproduced on bare metal.
I've had really consistent results with VMs in the past, but there is a
huge amount of code between 3.17 and now.
-chris
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2015-09-17 06:40 +0200 |
| Message-ID | <q9z9f-ZU-7@gated-at.bofh.it> |
| In reply to | #1226649 |
On Wed, Sep 16, 2015 at 11:48:59PM -0400, Chris Mason wrote: > On Thu, Sep 17, 2015 at 10:37:38AM +1000, Dave Chinner wrote: > > [cc Tejun] > > > > On Thu, Sep 17, 2015 at 08:07:04AM +1000, Dave Chinner wrote: > > # ./fs_mark -D 10000 -S0 -n 10000 -s 4096 -L 120 -d /mnt/scratch/0 -d /mnt/scratch/1 -d /mnt/scratch/2 -d /mnt/scratch/3 -d /mnt/scratch/4 -d /mnt/scratch/5 -d /mnt/scratch/6 -d /mnt/scratch/7 > > # Version 3.3, 8 thread(s) starting at Thu Sep 17 08:08:36 2015 > > # Sync method: NO SYNC: Test does not issue sync() or fsync() calls. > > # Directories: Time based hash between directories across 10000 subdirectories with 180 seconds per subdirectory. > > # File names: 40 bytes long, (16 initial bytes of time stamp with 24 random bytes at end of name) > > # Files info: size 4096 bytes, written with an IO size of 16384 bytes per write > > # App overhead is time in microseconds spent in the test not doing file writing related system calls. > > > > FSUse% Count Size Files/sec App Overhead > > 0 80000 4096 106938.0 543310 > > 0 160000 4096 102922.7 476362 > > 0 240000 4096 107182.9 538206 > > 0 320000 4096 107871.7 619821 > > 0 400000 4096 99255.6 622021 > > 0 480000 4096 103217.8 609943 > > 0 560000 4096 96544.2 640988 > > 0 640000 4096 100347.3 676237 > > 0 720000 4096 87534.8 483495 > > 0 800000 4096 72577.5 2556920 > > 0 880000 4096 97569.0 646996 > > > > <RAM fills here, sustained performance is now dependent on writeback> > > I think too many variables have changed here. > > My numbers: > > FSUse% Count Size Files/sec App Overhead > 0 160000 4096 356407.1 1458461 > 0 320000 4096 368755.1 1030047 > 0 480000 4096 358736.8 992123 > 0 640000 4096 361912.5 1009566 > 0 800000 4096 342851.4 1004152 <snip> > I can push the dirty threshold lower to try and make sure we end up in > the hard dirty limits but none of this is going to be related to the > plugging patch. The point of this test is to drive writeback as hard as possible, not to measure how fast we can create files in memory. i.e. if the test isn't pushing the dirty limits on your machines, then it really isn't putting a meaningful load on writeback, and so the plugging won't make significant difference because writeback isn't IO bound.... > I do see lower numbers if I let the test run even > longer, but there are a lot of things in the way that can slow it down > as the filesystem gets that big. Sure, that's why I hit the dirty limits early in the test - so it measures steady state performance before the fs gets to any significant scalability limits.... > > The baseline of no plugging is a full 3 minutes faster than the > > plugging behaviour of Linus' patch. The IO behaviour demonstrates > > that, sustaining between 25-30,000 IOPS and throughput of > > 130-150MB/s. Hence, while Linus' patch does change the IO patterns, > > it does not result in a performance improvement like the original > > plugging patch did. > > How consistent is this across runs? That's what I'm trying to work out. I didn't report it until I got consistently bad results - the numbers I reported were from the third time I ran the comparison, and they were representative and reproducable. I also ran my inode creation workload that is similar (but has not data writeback so doesn't go through writeback paths at all) and that shows no change in performance, so this problem (whatever it is) is only manifesting itself through data writeback.... The only measurable change I've noticed in my monitoring graphs is that there is a lot more iowait time than I normally see, even when the plugging appears to be working as desired. That's what I'm trying to track down now, and once I've got to the bottom of that I should have some idea of where the performance has gone.... As it is, there are a bunch of other things going wrong with 4.3-rc1+ right now that I'm working through - I haven't updated my kernel tree for 10 days because I've been away on holidays so I'm doing my usual "-rc1 is broken again" dance that I do every release cycle. (e.g every second boot hangs because systemd appears to be waiting for iscsi devices to appear without first starting the iscsi target daemon. Never happened before today, every new kernel I've booted today has hung on the first cold boot of the VM). > > IOWs, what we are seeing here is that the baseline writeback > > performance has regressed quite significantly since I took these > > numbers back on 3.17. I'm running on exactly the same test setup; > > the only difference is the kernel and so the current kernel baseline > > is ~20% slower than the baseline numbers I have in my patch. > > All of this in a VM, I'd much rather see this reproduced on bare metal. > I've had really consistent results with VMs in the past, but there is a > huge amount of code between 3.17 and now. I'm pretty sure it's not the VM - with the locking fix in place, everything else I've looked at (since applying the locking fix Linus mentioned) is within measurement error compared to 4.2. The only thing that is out of whack from a performance POV is data writeback.... Cheers, Dave. -- Dave Chinner david@fromorbit.com -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Chris Mason <clm@fb.com> |
|---|---|
| Date | 2015-09-17 14:20 +0200 |
| Message-ID | <q9Gkq-3fX-3@gated-at.bofh.it> |
| In reply to | #1226655 |
On Thu, Sep 17, 2015 at 02:30:08PM +1000, Dave Chinner wrote:
> On Wed, Sep 16, 2015 at 11:48:59PM -0400, Chris Mason wrote:
> > On Thu, Sep 17, 2015 at 10:37:38AM +1000, Dave Chinner wrote:
> > > [cc Tejun]
> > >
> > > On Thu, Sep 17, 2015 at 08:07:04AM +1000, Dave Chinner wrote:
> > > # ./fs_mark -D 10000 -S0 -n 10000 -s 4096 -L 120 -d /mnt/scratch/0 -d /mnt/scratch/1 -d /mnt/scratch/2 -d /mnt/scratch/3 -d /mnt/scratch/4 -d /mnt/scratch/5 -d /mnt/scratch/6 -d /mnt/scratch/7
> > > # Version 3.3, 8 thread(s) starting at Thu Sep 17 08:08:36 2015
> > > # Sync method: NO SYNC: Test does not issue sync() or fsync() calls.
> > > # Directories: Time based hash between directories across 10000 subdirectories with 180 seconds per subdirectory.
> > > # File names: 40 bytes long, (16 initial bytes of time stamp with 24 random bytes at end of name)
> > > # Files info: size 4096 bytes, written with an IO size of 16384 bytes per write
> > > # App overhead is time in microseconds spent in the test not doing file writing related system calls.
> > >
> > > FSUse% Count Size Files/sec App Overhead
> > > 0 80000 4096 106938.0 543310
> > > 0 160000 4096 102922.7 476362
> > > 0 240000 4096 107182.9 538206
> > > 0 320000 4096 107871.7 619821
> > > 0 400000 4096 99255.6 622021
> > > 0 480000 4096 103217.8 609943
> > > 0 560000 4096 96544.2 640988
> > > 0 640000 4096 100347.3 676237
> > > 0 720000 4096 87534.8 483495
> > > 0 800000 4096 72577.5 2556920
> > > 0 880000 4096 97569.0 646996
> > >
> > > <RAM fills here, sustained performance is now dependent on writeback>
> >
> > I think too many variables have changed here.
> >
> > My numbers:
> >
> > FSUse% Count Size Files/sec App Overhead
> > 0 160000 4096 356407.1 1458461
> > 0 320000 4096 368755.1 1030047
> > 0 480000 4096 358736.8 992123
> > 0 640000 4096 361912.5 1009566
> > 0 800000 4096 342851.4 1004152
>
> <snip>
>
> > I can push the dirty threshold lower to try and make sure we end up in
> > the hard dirty limits but none of this is going to be related to the
> > plugging patch.
>
> The point of this test is to drive writeback as hard as possible,
> not to measure how fast we can create files in memory. i.e. if the
> test isn't pushing the dirty limits on your machines, then it really
> isn't putting a meaningful load on writeback, and so the plugging
> won't make significant difference because writeback isn't IO
> bound....
It does end up IO bound on my rig, just because we do eventually hit the
dirty limits. Otherwise there would be zero benefits in fs_mark from
any patches vs plain v4.2
But I setup a run last night with a dirty_ratio_bytes at 3G and
dirty_background_ratio_bytes at 1.5G.
There is definitely variation, but nothing like what you saw:
FSUse% Count Size Files/sec App Overhead
0 160000 4096 317427.9 1524951
0 320000 4096 319723.9 1023874
0 480000 4096 336696.4 1053884
0 640000 4096 257113.1 1190851
0 800000 4096 257644.2 1198054
0 960000 4096 254896.6 1225610
0 1120000 4096 241052.6 1203227
0 1280000 4096 214961.2 1386236
0 1440000 4096 239985.7 1264659
0 1600000 4096 232174.3 1310018
0 1760000 4096 250477.9 1227289
0 1920000 4096 221500.9 1276223
0 2080000 4096 235212.1 1284989
0 2240000 4096 238580.2 1257260
0 2400000 4096 224182.6 1326821
0 2560000 4096 234628.7 1236402
0 2720000 4096 244675.3 1228400
0 2880000 4096 234364.0 1268408
0 3040000 4096 229712.6 1306148
0 3200000 4096 241170.5 1254490
0 3360000 4096 220487.8 1331456
0 3520000 4096 215831.7 1313682
0 3680000 4096 210934.7 1235750
0 3840000 4096 218435.4 1258077
0 4000000 4096 232127.7 1271555
0 4160000 4096 212017.6 1381525
0 4320000 4096 216309.3 1370558
0 4480000 4096 239072.4 1269086
0 4640000 4096 221959.1 1333164
0 4800000 4096 228396.8 1213160
0 4960000 4096 225747.5 1318503
0 5120000 4096 115727.0 1237327
0 5280000 4096 184171.4 1547357
0 5440000 4096 209917.8 1380510
0 5600000 4096 181074.7 1391764
0 5760000 4096 263516.7 1155172
0 5920000 4096 236405.8 1239719
0 6080000 4096 231587.2 1221408
0 6240000 4096 237118.8 1244272
0 6400000 4096 236773.2 1201428
0 6560000 4096 243987.5 1240527
0 6720000 4096 232428.0 1283265
0 6880000 4096 234839.9 1209152
0 7040000 4096 234947.3 1223456
0 7200000 4096 231463.1 1260628
0 7360000 4096 226750.3 1290098
0 7520000 4096 213632.0 1236409
0 7680000 4096 194710.2 1411595
0 7840000 4096 213963.1 4146893
0 8000000 4096 225109.8 1323573
0 8160000 4096 251322.1 1380271
0 8320000 4096 220167.2 1159390
0 8480000 4096 210991.2 1110593
0 8640000 4096 197922.8 1126072
0 8800000 4096 203539.3 1143501
0 8960000 4096 193041.7 1134329
0 9120000 4096 184667.9 1119222
0 9280000 4096 165968.7 1172738
0 9440000 4096 192767.3 1098361
0 9600000 4096 227115.7 1158097
0 9760000 4096 232139.8 1264245
0 9920000 4096 213320.5 1270505
0 10080000 4096 217013.4 1324569
0 10240000 4096 227171.6 1308668
0 10400000 4096 208591.4 1392098
0 10560000 4096 212006.0 1359188
0 10720000 4096 213449.3 1352084
0 10880000 4096 219890.1 1326240
0 11040000 4096 215907.7 1239180
0 11200000 4096 214207.2 1334846
0 11360000 4096 212875.2 1338429
0 11520000 4096 211690.0 1249519
0 11680000 4096 217013.0 1262050
0 11840000 4096 204730.1 1205087
0 12000000 4096 191146.9 1188635
0 12160000 4096 207844.6 1157033
0 12320000 4096 208857.7 1168111
0 12480000 4096 198256.4 1388368
0 12640000 4096 214996.1 1305412
0 12800000 4096 212332.9 1357814
0 12960000 4096 210325.8 1336127
0 13120000 4096 200292.1 1282419
0 13280000 4096 202030.2 1412105
0 13440000 4096 216553.7 1424076
0 13600000 4096 218721.7 1298149
0 13760000 4096 202037.4 1266877
0 13920000 4096 224032.3 1198159
0 14080000 4096 206105.6 1336489
0 14240000 4096 227540.3 1160841
0 14400000 4096 236921.7 1190394
0 14560000 4096 229343.3 1147451
0 14720000 4096 199435.1 1284374
0 14880000 4096 215177.3 1178542
0 15040000 4096 206194.1 1170832
0 15200000 4096 215762.3 1125633
0 15360000 4096 194511.0 1122947
0 15520000 4096 179008.5 1292603
0 15680000 4096 208636.9 1094960
0 15840000 4096 192173.1 1237891
0 16000000 4096 212888.9 1111551
0 16160000 4096 218403.0 1143400
0 16320000 4096 207260.5 1233526
0 16480000 4096 202123.2 1151509
0 16640000 4096 191033.0 1257706
0 16800000 4096 196865.4 1154520
0 16960000 4096 210361.2 1128930
0 17120000 4096 201755.2 1160469
0 17280000 4096 196946.6 1173529
0 17440000 4096 199677.8 1165750
0 17600000 4096 194248.4 1234944
0 17760000 4096 200027.9 1256599
0 17920000 4096 206507.0 1166820
0 18080000 4096 215082.7 1167599
0 18240000 4096 201475.5 1212202
0 18400000 4096 208247.6 1252255
0 18560000 4096 205482.7 1311436
0 18720000 4096 200111.9 1358784
0 18880000 4096 200028.3 1351332
0 19040000 4096 198873.4 1287400
0 19200000 4096 209609.3 1268400
0 19360000 4096 203538.6 1249787
0 19520000 4096 203427.9 1294105
0 19680000 4096 201905.3 1280714
0 19840000 4096 209642.9 1283281
0 20000000 4096 203438.9 1315427
0 20160000 4096 199690.7 1252267
0 20320000 4096 185965.2 1398905
0 20480000 4096 203221.6 1214029
0 20640000 4096 208654.8 1232679
0 20800000 4096 212488.6 1298458
0 20960000 4096 189701.1 1356640
0 21120000 4096 198522.1 1361240
0 21280000 4096 203857.3 1263402
0 21440000 4096 204616.8 1362853
0 21600000 4096 196310.6 1266710
0 21760000 4096 203275.4 1391150
0 21920000 4096 205998.5 1378741
0 22080000 4096 205434.2 1283787
0 22240000 4096 195918.0 1415912
0 22400000 4096 186193.0 1413623
0 22560000 4096 192911.3 1393471
0 22720000 4096 203726.3 1264281
0 22880000 4096 204853.4 1221048
0 23040000 4096 222803.2 1153031
0 23200000 4096 198558.6 1346256
0 23360000 4096 201001.4 1278817
0 23520000 4096 206225.2 1270440
0 23680000 4096 190894.2 1425299
0 23840000 4096 198555.6 1334122
0 24000000 4096 202386.4 1332157
0 24160000 4096 205103.1 1313607
>
> > I do see lower numbers if I let the test run even
> > longer, but there are a lot of things in the way that can slow it down
> > as the filesystem gets that big.
>
> Sure, that's why I hit the dirty limits early in the test - so it
> measures steady state performance before the fs gets to any
> significant scalability limits....
>
> > > The baseline of no plugging is a full 3 minutes faster than the
> > > plugging behaviour of Linus' patch. The IO behaviour demonstrates
> > > that, sustaining between 25-30,000 IOPS and throughput of
> > > 130-150MB/s. Hence, while Linus' patch does change the IO patterns,
> > > it does not result in a performance improvement like the original
> > > plugging patch did.
> >
> > How consistent is this across runs?
>
> That's what I'm trying to work out. I didn't report it until I got
> consistently bad results - the numbers I reported were from the
> third time I ran the comparison, and they were representative and
> reproducable. I also ran my inode creation workload that is similar
> (but has not data writeback so doesn't go through writeback paths at
> all) and that shows no change in performance, so this problem
> (whatever it is) is only manifesting itself through data
> writeback....
The big change between Linus' patch and your patch is with Linus kblockd
is probably doing most of the actual unplug work (except for the last
super block in the list). If a process is waiting for dirty writeout
progress, it has to wait for that context switch to kblockd.
In the VM, that's going to hurt more then my big two socket mostly idle
machine.
>
> The only measurable change I've noticed in my monitoring graphs is
> that there is a lot more iowait time than I normally see, even when
> the plugging appears to be working as desired. That's what I'm
> trying to track down now, and once I've got to the bottom of that I
> should have some idea of where the performance has gone....
>
> As it is, there are a bunch of other things going wrong with
> 4.3-rc1+ right now that I'm working through - I haven't updated my
> kernel tree for 10 days because I've been away on holidays so I'm
> doing my usual "-rc1 is broken again" dance that I do every release
> cycle. (e.g every second boot hangs because systemd appears to be
> waiting for iscsi devices to appear without first starting the iscsi
> target daemon. Never happened before today, every new kernel I've
> booted today has hung on the first cold boot of the VM).
I've been doing 4.2 plus patches because rc1 didn't boot on this strange
box. Let me nail that down and rerun.
-chris
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Chris Mason <clm@fb.com> |
|---|---|
| Date | 2015-09-12 01:10 +0200 |
| Message-ID | <q7FCb-7II-51@gated-at.bofh.it> |
| In reply to | #1223101 |
On Fri, Sep 11, 2015 at 02:04:13PM -0700, Linus Torvalds wrote: > On Fri, Sep 11, 2015 at 1:40 PM, Josef Bacik <jbacik@fb.com> wrote: > > > > So we talked about this when we were trying to figure out a solution. The > > problem with this approach is now we have a plug that covers multiple super > > blocks (__writeback_inodes_wb loops through the sb's starts writeback), > > which is likely to give us crappier performance than no plug at all. > > Why would that be? Either they are on separate disks, and the IO is > all independent anyway, and at most it got started by some (small) > CPU-amount later. Actual throughput should be the same. No? > > Or the filesystems are on the same disk, in which case it should > presumably be a win to submit the IO together. > > Of course, actual numbers would be the deciding factor if this really > is noticeable. But "cleaner code and saner locking" is definitely an > issue at least for me. Originally I was worried about the latency impact of holding the plugs over more than one super with high end flash. I just didn't want to hold onto the IO for longer than we had to. But, since this isn't really latency sensitive anyway, if we find we're not keeping the flash pipelines full the right answer is to short circuit the plugging in general. I'd agree actual throughput should be the same. But benchmarking is the best choice, I'll be able to reproduce Dave's original results without too much trouble. Our thinking for this duct tape patch was the lock wasn't very hot and this was the best immediate compromise between the bug and the perf improvement. Happy to sign up for pushing the lock higher if that's what people would rather see. -chris -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2015-09-12 01:20 +0200 |
| Message-ID | <q7FLP-7Un-5@gated-at.bofh.it> |
| In reply to | #1223207 |
On Fri, Sep 11, 2015 at 4:06 PM, Chris Mason <clm@fb.com> wrote:
>
> Originally I was worried about the latency impact of holding the
> plugs over more than one super with high end flash. I just didn't want
> to hold onto the IO for longer than we had to.
>
> But, since this isn't really latency sensitive anyway, if we find we're
> not keeping the flash pipelines full the right answer is to short
> circuit the plugging in general. I'd agree actual throughput should be
> the same.
Yeah, this only triggers for system-wide writeback, so I don't seer
that it should be latency-sensitive at that level, afaik.
But hey, it's IO, and I've been surprised by magic pattern
sensitivites before...
Linus
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Page 3 of 3 — ← Prev page 1 2 [3]
Back to top | Article view | linux.kernel
csiph-web