Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1367290 > unrolled thread
| Started by | Jens Axboe <axboe@fb.com> |
|---|---|
| First post | 2016-03-30 17:20 +0200 |
| Last post | 2016-04-01 19:10 +0200 |
| Articles | 20 — 3 participants |
Back to article view | Back to linux.kernel
[PATCHSET v3][RFC] Make background writeback not suck Jens Axboe <axboe@fb.com> - 2016-03-30 17:20 +0200
[PATCH 3/9] writeback: use WRITE_SYNC for reclaim or sync writeback Jens Axboe <axboe@fb.com> - 2016-03-30 17:20 +0200
[PATCH 6/9] sd: inform block layer of write cache state Jens Axboe <axboe@fb.com> - 2016-03-30 17:20 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Dave Chinner <david@fromorbit.com> - 2016-03-31 10:30 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Jens Axboe <axboe@fb.com> - 2016-03-31 16:30 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Jens Axboe <axboe@fb.com> - 2016-03-31 18:30 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Dave Chinner <david@fromorbit.com> - 2016-04-01 03:00 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Jens Axboe <axboe@fb.com> - 2016-04-01 05:30 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Jens Axboe <axboe@fb.com> - 2016-04-01 05:40 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Jens Axboe <axboe@fb.com> - 2016-04-01 05:40 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Dave Chinner <david@fromorbit.com> - 2016-04-01 08:20 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Jens Axboe <axboe@fb.com> - 2016-04-01 16:40 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Dave Chinner <david@fromorbit.com> - 2016-04-01 07:10 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Dave Chinner <david@fromorbit.com> - 2016-04-01 02:50 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Jens Axboe <axboe@fb.com> - 2016-04-01 05:30 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Dave Chinner <david@fromorbit.com> - 2016-04-01 08:30 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Jens Axboe <axboe@fb.com> - 2016-04-01 16:40 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Holger Hoffstätte <holger.hoffstaette@googlemail.com> - 2016-04-01 00:20 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Dave Chinner <david@fromorbit.com> - 2016-04-01 03:10 +0200
Re: [PATCHSET v3][RFC] Make background writeback not suck Holger Hoffstätte <holger.hoffstaette@googlemail.com> - 2016-04-01 19:10 +0200
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-03-30 17:20 +0200 |
| Subject | [PATCHSET v3][RFC] Make background writeback not suck |
| Message-ID | <ripUS-7fj-3@gated-at.bofh.it> |
Hi,
This patchset isn't as much a final solution, as it's demonstration
of what I believe is a huge issue. Since the dawn of time, our
background buffered writeback has sucked. When we do background
buffered writeback, it should have little impact on foreground
activity. That's the definition of background activity... But for as
long as I can remember, heavy buffered writers has not behaved like
that. For instance, if I do something like this:
$ dd if=/dev/zero of=foo bs=1M count=10k
on my laptop, and then try and start chrome, it basically won't start
before the buffered writeback is done. Or, for server oriented
workloads, where installation of a big RPM (or similar) adversely
impacts data base reads or sync writes. When that happens, I get people
yelling at me.
Last time I posted this, I used flash storage as the example. But
this works equally well on rotating storage. Let's run a test case
that writes a lot. This test writes 50 files, each 100M, on XFS on
a regular hard drive. While this happens, we attempt to read
another file with fio.
Writers:
$ time (./write-files ; sync)
real 1m6.304s
user 0m0.020s
sys 0m12.210s
Fio reader:
read : io=35580KB, bw=550868B/s, iops=134, runt= 66139msec
clat (usec): min=40, max=654204, avg=7432.37, stdev=43872.83
lat (usec): min=40, max=654204, avg=7432.70, stdev=43872.83
clat percentiles (usec):
| 1.00th=[ 41], 5.00th=[ 41], 10.00th=[ 41], 20.00th=[ 42],
| 30.00th=[ 42], 40.00th=[ 42], 50.00th=[ 43], 60.00th=[ 52],
| 70.00th=[ 59], 80.00th=[ 65], 90.00th=[ 87], 95.00th=[ 1192],
| 99.00th=[254976], 99.50th=[358400], 99.90th=[444416], 99.95th=[468992],
| 99.99th=[651264]
Let's run the same test, but with the patches applied, and wb_percent
set to 10%:
Writers:
$ time (./write-files ; sync)
real 1m29.384s
user 0m0.040s
sys 0m10.810s
Fio reader:
read : io=1024.0MB, bw=18640KB/s, iops=4660, runt= 56254msec
clat (usec): min=39, max=408400, avg=212.05, stdev=2982.44
lat (usec): min=39, max=408400, avg=212.30, stdev=2982.44
clat percentiles (usec):
| 1.00th=[ 40], 5.00th=[ 41], 10.00th=[ 41], 20.00th=[ 41],
| 30.00th=[ 42], 40.00th=[ 42], 50.00th=[ 42], 60.00th=[ 42],
| 70.00th=[ 43], 80.00th=[ 45], 90.00th=[ 56], 95.00th=[ 60],
| 99.00th=[ 454], 99.50th=[ 8768], 99.90th=[36608], 99.95th=[43264],
| 99.99th=[69120]
Much better, looking at the P99.x percentiles, and of course on
the bandwidth front as well. It's the difference between this:
---io---- -system-- ------cpu-----
bi bo in cs us sy id wa st
20636 45056 5593 10833 0 0 94 6 0
16416 46080 4484 8666 0 0 94 6 0
16960 47104 5183 8936 0 0 94 6 0
and this
---io---- -system-- ------cpu-----
bi bo in cs us sy id wa st
384 73728 571 558 0 0 95 5 0
384 73728 548 545 0 0 95 5 0
388 73728 575 763 0 0 96 4 0
in the vmstat output. It's not quite as bad as deeper queue depth
devices, where we have hugely bursty IO, but it's still very slow.
If we don't run the competing reader, the dirty data writeback proceeds
at normal rates:
# time (./write-files ; sync)
real 1m6.919s
user 0m0.010s
sys 0m10.900s
The above was run without scsi-mq, and with using the deadline scheduler,
results with CFQ are similary depressing for this test. So IO scheduling
is in place for this test, it's not pure blk-mq without scheduling.
The above was the why. The how is basically throttling background
writeback. We still want to issue big writes from the vm side of things,
so we get nice and big extents on the file system end. But we don't need
to flood the device with THOUSANDS of requests for background writeback.
For most devices, we don't need a whole lot to get decent throughput.
This adds some simple blk-wb code that keeps limits how much buffered
writeback we keep in flight on the device end. The default is pretty
low. If we end up switching to WB_SYNC_ALL, we up the limits. If the
dirtying task ends up being throttled in balance_dirty_pages(), we up
the limit. If we need to reclaim memory, we up the limit. The cases
that need to clean memory at or near device speeds, they get to do
that. We still don't need thousands of requests to accomplish that.
And for the cases where we don't need to be near device limits, we
can clean at a more reasonable pace. See the last patch in the series
for a more detailed description of the change, and the tunable.
I welcome testing. If you are sick of Linux bogging down when buffered
writes are happening, then this is for you, laptop or server. The
patchset is fully stable, I have not observed problems. It passes full
xfstest runs, and a variety of benchmarks as well. It works equally well
on blk-mq/scsi-mq, and "classic" setups.
You can also find this in a branch in the block git repo:
git://git.kernel.dk/linux-block.git wb-buf-throttle
Note that I rebase this branch when I collapse patches. Patches are
against current Linus' git, 4.6.0-rc1, I can make them available
against 4.5 as well, if there's any interest in that for test
purposes.
Changes since v2
- Switch from wb_depth to wb_percent, as that's an easier tunable.
- Add the patch to track device depth on the block layer side.
- Cleanup the limiting code.
- Don't use a fixed limit in the wb wait, since it can change
between wakeups.
- Minor tweaks, fixups, cleanups.
Changes since v1
- Drop sync() WB_SYNC_NONE -> WB_SYNC_ALL change
- wb_start_writeback() fills in background/reclaim/sync info in
the writeback work, based on writeback reason.
- Use WRITE_SYNC for reclaim/sync IO
- Split balance_dirty_pages() sleep change into separate patch
- Drop get_request() u64 flag change, set the bit on the request
directly after-the-fact.
- Fix wrong sysfs return value
- Various small cleanups
block/Makefile | 2
block/blk-core.c | 15 ++
block/blk-mq.c | 31 ++++-
block/blk-settings.c | 20 +++
block/blk-sysfs.c | 128 ++++++++++++++++++++
block/blk-wb.c | 238 +++++++++++++++++++++++++++++++++++++++
block/blk-wb.h | 33 +++++
drivers/nvme/host/core.c | 1
drivers/scsi/scsi.c | 3
drivers/scsi/sd.c | 5
fs/block_dev.c | 2
fs/buffer.c | 2
fs/f2fs/data.c | 2
fs/f2fs/node.c | 2
fs/fs-writeback.c | 13 ++
fs/gfs2/meta_io.c | 3
fs/mpage.c | 9 -
fs/xfs/xfs_aops.c | 2
include/linux/backing-dev-defs.h | 2
include/linux/blk_types.h | 2
include/linux/blkdev.h | 18 ++
include/linux/writeback.h | 8 +
mm/page-writeback.c | 2
23 files changed, 527 insertions(+), 16 deletions(-)
--
Jens Axboe
[toc] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-03-30 17:20 +0200 |
| Subject | [PATCH 3/9] writeback: use WRITE_SYNC for reclaim or sync writeback |
| Message-ID | <riq4y-7jq-21@gated-at.bofh.it> |
| In reply to | #1367290 |
If we're doing reclaim or sync IO, use WRITE_SYNC to inform the lower
levels of the importance of this IO.
Signed-off-by: Jens Axboe <axboe@fb.com>
---
include/linux/writeback.h | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/include/linux/writeback.h b/include/linux/writeback.h
index 719c255e105a..b2c75b8901da 100644
--- a/include/linux/writeback.h
+++ b/include/linux/writeback.h
@@ -102,7 +102,7 @@ struct writeback_control {
static inline int wbc_to_write(struct writeback_control *wbc)
{
- if (wbc->sync_mode == WB_SYNC_ALL)
+ if (wbc->sync_mode == WB_SYNC_ALL || wbc->for_reclaim || wbc->for_sync)
return WRITE_SYNC;
return WRITE;
--
2.8.0.rc4.6.g7e4ba36
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-03-30 17:20 +0200 |
| Subject | [PATCH 6/9] sd: inform block layer of write cache state |
| Message-ID | <riq4y-7jq-25@gated-at.bofh.it> |
| In reply to | #1367290 |
Signed-off-by: Jens Axboe <axboe@fb.com> --- drivers/scsi/sd.c | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/drivers/scsi/sd.c b/drivers/scsi/sd.c index 5a5457ac9cdb..049f424fb4ad 100644 --- a/drivers/scsi/sd.c +++ b/drivers/scsi/sd.c @@ -192,6 +192,7 @@ cache_type_store(struct device *dev, struct device_attribute *attr, sdkp->WCE = wce; sdkp->RCD = rcd; sd_set_flush_flag(sdkp); + blk_queue_write_cache(sdp->request_queue, wce != 0); return count; } @@ -2571,7 +2572,7 @@ sd_read_cache_type(struct scsi_disk *sdkp, unsigned char *buffer) sdkp->DPOFUA ? "supports DPO and FUA" : "doesn't support DPO or FUA"); - return; + goto done; } bad_sense: @@ -2596,6 +2597,8 @@ defaults: } sdkp->RCD = 0; sdkp->DPOFUA = 0; +done: + blk_queue_write_cache(sdp->request_queue, sdkp->WCE != 0); } /* -- 2.8.0.rc4.6.g7e4ba36
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-03-31 10:30 +0200 |
| Message-ID | <riG9l-29W-35@gated-at.bofh.it> |
| In reply to | #1367290 |
On Wed, Mar 30, 2016 at 09:07:48AM -0600, Jens Axboe wrote:
> Hi,
>
> This patchset isn't as much a final solution, as it's demonstration
> of what I believe is a huge issue. Since the dawn of time, our
> background buffered writeback has sucked. When we do background
> buffered writeback, it should have little impact on foreground
> activity. That's the definition of background activity... But for as
> long as I can remember, heavy buffered writers has not behaved like
> that. For instance, if I do something like this:
>
> $ dd if=/dev/zero of=foo bs=1M count=10k
>
> on my laptop, and then try and start chrome, it basically won't start
> before the buffered writeback is done. Or, for server oriented
> workloads, where installation of a big RPM (or similar) adversely
> impacts data base reads or sync writes. When that happens, I get people
> yelling at me.
>
> Last time I posted this, I used flash storage as the example. But
> this works equally well on rotating storage. Let's run a test case
> that writes a lot. This test writes 50 files, each 100M, on XFS on
> a regular hard drive. While this happens, we attempt to read
> another file with fio.
>
> Writers:
>
> $ time (./write-files ; sync)
> real 1m6.304s
> user 0m0.020s
> sys 0m12.210s
Great. So a basic IO tests looks good - let's through something more
complex at it. Say, a benchmark I've been using for years to stress
the Io subsystem, the filesystem and memory reclaim all at the same
time: a concurent fsmark inode creation test.
(first google hit https://lkml.org/lkml/2013/9/10/46)
This generates thousands of REQ_WRITE metadata IOs every second, so
iif I understand how the throttle works correctly, these would be
classified as background writeback by the block layer throttle.
And....
FSUse% Count Size Files/sec App Overhead
0 1600000 0 255845.0 10796891
0 3200000 0 261348.8 10842349
0 4800000 0 249172.3 14121232
0 6400000 0 245172.8 12453759
0 8000000 0 201249.5 14293100
0 9600000 0 200417.5 29496551
>>>> 0 11200000 0 90399.6 40665397
0 12800000 0 212265.6 21839031
0 14400000 0 206398.8 32598378
0 16000000 0 197589.7 26266552
0 17600000 0 206405.2 16447795
>>>> 0 19200000 0 99189.6 87650540
0 20800000 0 249720.8 12294862
0 22400000 0 138523.8 47330007
>>>> 0 24000000 0 85486.2 14271096
0 25600000 0 157538.1 64430611
0 27200000 0 109677.8 47835961
0 28800000 0 207230.5 31301031
0 30400000 0 188739.6 33750424
0 32000000 0 174197.9 41402526
0 33600000 0 139152.0 100838085
0 35200000 0 203729.7 34833764
0 36800000 0 228277.4 12459062
>>>> 0 38400000 0 94962.0 30189182
0 40000000 0 166221.9 40564922
>>>> 0 41600000 0 62902.5 80098461
0 43200000 0 217932.6 22539354
0 44800000 0 189594.6 24692209
0 46400000 0 137834.1 39822038
0 48000000 0 240043.8 12779453
0 49600000 0 176830.8 16604133
0 51200000 0 180771.8 32860221
real 5m35.967s
user 3m57.054s
sys 48m53.332s
In those highlighted report points, the performance has dropped
significantly. The typical range I expect to see ionce memory has
filled (a bit over 8m inodes) is 180k-220k. Runtime on a vanilla
kernel was 4m40s and there were no performance drops, so this
workload runs almost a minute slower with the block layer throttling
code.
What I see in these performance dips is the XFS transaction
subsystem stalling *completely* - instead of running at a steady
state of around 350,000 transactions/s, there are *zero*
transactions running for periods of up to ten seconds. This
co-incides with the CPU usage falling to almost zero as well.
AFAICT, the only thing that is running when the filesystem stalls
like this is memory reclaim.
Without the block throttling patches, the workload quickly finds a
steady state of around 7.5-8.5 million cached inodes, and it doesn't
vary much outside those bounds. With the block throttling patches,
on every transaction subsystem stall that occurs, the inode cache
gets 3-4 million inodes trimmed out of it (i.e. half the
cache), and in a couple of cases I saw it trim 6+ million inodes from
the cache before the transactions started up and the cache started
growing again.
> The above was run without scsi-mq, and with using the deadline scheduler,
> results with CFQ are similary depressing for this test. So IO scheduling
> is in place for this test, it's not pure blk-mq without scheduling.
virtio in guest, XFS direct IO -> no-op -> scsi in host.
Cheers,
Dave.
--
Dave Chinner
david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-03-31 16:30 +0200 |
| Message-ID | <riLLH-6lZ-1@gated-at.bofh.it> |
| In reply to | #1367970 |
On 03/31/2016 02:24 AM, Dave Chinner wrote: > On Wed, Mar 30, 2016 at 09:07:48AM -0600, Jens Axboe wrote: >> Hi, >> >> This patchset isn't as much a final solution, as it's demonstration >> of what I believe is a huge issue. Since the dawn of time, our >> background buffered writeback has sucked. When we do background >> buffered writeback, it should have little impact on foreground >> activity. That's the definition of background activity... But for as >> long as I can remember, heavy buffered writers has not behaved like >> that. For instance, if I do something like this: >> >> $ dd if=/dev/zero of=foo bs=1M count=10k >> >> on my laptop, and then try and start chrome, it basically won't start >> before the buffered writeback is done. Or, for server oriented >> workloads, where installation of a big RPM (or similar) adversely >> impacts data base reads or sync writes. When that happens, I get people >> yelling at me. >> >> Last time I posted this, I used flash storage as the example. But >> this works equally well on rotating storage. Let's run a test case >> that writes a lot. This test writes 50 files, each 100M, on XFS on >> a regular hard drive. While this happens, we attempt to read >> another file with fio. >> >> Writers: >> >> $ time (./write-files ; sync) >> real 1m6.304s >> user 0m0.020s >> sys 0m12.210s > > Great. So a basic IO tests looks good - let's through something more > complex at it. Say, a benchmark I've been using for years to stress > the Io subsystem, the filesystem and memory reclaim all at the same > time: a concurent fsmark inode creation test. > (first google hit https://lkml.org/lkml/2013/9/10/46) Is that how you are invoking it as well same arguments? > This generates thousands of REQ_WRITE metadata IOs every second, so > iif I understand how the throttle works correctly, these would be > classified as background writeback by the block layer throttle. > And.... > > FSUse% Count Size Files/sec App Overhead > 0 1600000 0 255845.0 10796891 > 0 3200000 0 261348.8 10842349 > 0 4800000 0 249172.3 14121232 > 0 6400000 0 245172.8 12453759 > 0 8000000 0 201249.5 14293100 > 0 9600000 0 200417.5 29496551 >>>>> 0 11200000 0 90399.6 40665397 > 0 12800000 0 212265.6 21839031 > 0 14400000 0 206398.8 32598378 > 0 16000000 0 197589.7 26266552 > 0 17600000 0 206405.2 16447795 >>>>> 0 19200000 0 99189.6 87650540 > 0 20800000 0 249720.8 12294862 > 0 22400000 0 138523.8 47330007 >>>>> 0 24000000 0 85486.2 14271096 > 0 25600000 0 157538.1 64430611 > 0 27200000 0 109677.8 47835961 > 0 28800000 0 207230.5 31301031 > 0 30400000 0 188739.6 33750424 > 0 32000000 0 174197.9 41402526 > 0 33600000 0 139152.0 100838085 > 0 35200000 0 203729.7 34833764 > 0 36800000 0 228277.4 12459062 >>>>> 0 38400000 0 94962.0 30189182 > 0 40000000 0 166221.9 40564922 >>>>> 0 41600000 0 62902.5 80098461 > 0 43200000 0 217932.6 22539354 > 0 44800000 0 189594.6 24692209 > 0 46400000 0 137834.1 39822038 > 0 48000000 0 240043.8 12779453 > 0 49600000 0 176830.8 16604133 > 0 51200000 0 180771.8 32860221 > > real 5m35.967s > user 3m57.054s > sys 48m53.332s > > In those highlighted report points, the performance has dropped > significantly. The typical range I expect to see ionce memory has > filled (a bit over 8m inodes) is 180k-220k. Runtime on a vanilla > kernel was 4m40s and there were no performance drops, so this > workload runs almost a minute slower with the block layer throttling > code. > > What I see in these performance dips is the XFS transaction > subsystem stalling *completely* - instead of running at a steady > state of around 350,000 transactions/s, there are *zero* > transactions running for periods of up to ten seconds. This > co-incides with the CPU usage falling to almost zero as well. > AFAICT, the only thing that is running when the filesystem stalls > like this is memory reclaim. I'll take a look at this, stalls should definitely not be occurring. How much memory does the box have? > Without the block throttling patches, the workload quickly finds a > steady state of around 7.5-8.5 million cached inodes, and it doesn't > vary much outside those bounds. With the block throttling patches, > on every transaction subsystem stall that occurs, the inode cache > gets 3-4 million inodes trimmed out of it (i.e. half the > cache), and in a couple of cases I saw it trim 6+ million inodes from > the cache before the transactions started up and the cache started > growing again. > >> The above was run without scsi-mq, and with using the deadline scheduler, >> results with CFQ are similary depressing for this test. So IO scheduling >> is in place for this test, it's not pure blk-mq without scheduling. > > virtio in guest, XFS direct IO -> no-op -> scsi in host. That has write back caching enabled on the guest, correct? -- Jens Axboe
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-03-31 18:30 +0200 |
| Message-ID | <riNDP-7LE-1@gated-at.bofh.it> |
| In reply to | #1368347 |
On 03/31/2016 08:29 AM, Jens Axboe wrote: >> What I see in these performance dips is the XFS transaction >> subsystem stalling *completely* - instead of running at a steady >> state of around 350,000 transactions/s, there are *zero* >> transactions running for periods of up to ten seconds. This >> co-incides with the CPU usage falling to almost zero as well. >> AFAICT, the only thing that is running when the filesystem stalls >> like this is memory reclaim. > > I'll take a look at this, stalls should definitely not be occurring. How > much memory does the box have? I can't seem to reproduce this at all. On an nvme device, I get a fairly steady 60K/sec file creation rate, and we're nowhere near being IO bound. So the throttling has no effect at all. On a raid0 on 4 flash devices, I get something that looks more IO bound, for some reason. Still no impact of the throttling, however. But given that your setup is this: virtio in guest, XFS direct IO -> no-op -> scsi in host. we do potentially have two throttling points, which we don't want. Is both the guest and the host running the new code, or just the guest? In any case, can I talk you into trying with two patches on top of the current code? It's the two newest patches here: http://git.kernel.dk/cgit/linux-block/log/?h=wb-buf-throttle The first treats REQ_META|REQ_PRIO like they should be treated, like high priority IO. The second disables throttling for virtual devices, so we only throttle on the backend. The latter should probably be the other way around, but we need some way of conveying that information to the backend. -- Jens Axboe
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-04-01 03:00 +0200 |
| Message-ID | <riVBo-4M2-7@gated-at.bofh.it> |
| In reply to | #1368429 |
On Thu, Mar 31, 2016 at 10:21:04AM -0600, Jens Axboe wrote: > On 03/31/2016 08:29 AM, Jens Axboe wrote: > >>What I see in these performance dips is the XFS transaction > >>subsystem stalling *completely* - instead of running at a steady > >>state of around 350,000 transactions/s, there are *zero* > >>transactions running for periods of up to ten seconds. This > >>co-incides with the CPU usage falling to almost zero as well. > >>AFAICT, the only thing that is running when the filesystem stalls > >>like this is memory reclaim. > > > >I'll take a look at this, stalls should definitely not be occurring. How > >much memory does the box have? > > I can't seem to reproduce this at all. On an nvme device, I get a > fairly steady 60K/sec file creation rate, and we're nowhere near > being IO bound. So the throttling has no effect at all. That's too slow to show the stalls - your likely concurrency bound in allocation by the default AG count (4) from mkfs. Use mkfs.xfs -d agcount=32 so that every thread works in it's own AG. > On a raid0 on 4 flash devices, I get something that looks more IO > bound, for some reason. Still no impact of the throttling, however. > But given that your setup is this: > > virtio in guest, XFS direct IO -> no-op -> scsi in host. > > we do potentially have two throttling points, which we don't want. > Is both the guest and the host running the new code, or just the > guest? Just the guest. Host is running a 4.2.x kernel, IIRC. > In any case, can I talk you into trying with two patches on top of > the current code? It's the two newest patches here: > > http://git.kernel.dk/cgit/linux-block/log/?h=wb-buf-throttle > > The first treats REQ_META|REQ_PRIO like they should be treated, like > high priority IO. The second disables throttling for virtual > devices, so we only throttle on the backend. The latter should > probably be the other way around, but we need some way of conveying > that information to the backend. I'm not changing the host kernels - it's a production machine and so it runs long uptime testing of stable kernels. (e.g. catch slow memory leaks, etc). So if you've disabled throttling in the guest, I can't test the throttling changes. Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-04-01 05:30 +0200 |
| Message-ID | <riXWx-6Md-5@gated-at.bofh.it> |
| In reply to | #1368919 |
On 03/31/2016 06:56 PM, Dave Chinner wrote: > On Thu, Mar 31, 2016 at 10:21:04AM -0600, Jens Axboe wrote: >> On 03/31/2016 08:29 AM, Jens Axboe wrote: >>>> What I see in these performance dips is the XFS transaction >>>> subsystem stalling *completely* - instead of running at a steady >>>> state of around 350,000 transactions/s, there are *zero* >>>> transactions running for periods of up to ten seconds. This >>>> co-incides with the CPU usage falling to almost zero as well. >>>> AFAICT, the only thing that is running when the filesystem stalls >>>> like this is memory reclaim. >>> >>> I'll take a look at this, stalls should definitely not be occurring. How >>> much memory does the box have? >> >> I can't seem to reproduce this at all. On an nvme device, I get a >> fairly steady 60K/sec file creation rate, and we're nowhere near >> being IO bound. So the throttling has no effect at all. > > That's too slow to show the stalls - your likely concurrency bound > in allocation by the default AG count (4) from mkfs. Use mkfs.xfs -d > agcount=32 so that every thread works in it's own AG. That's the key, with that I get 300-400K ops/sec instead. I'll run some testing with this tomorrow and see what I can find, it did one full run now and I didn't see any issues, but I need to run it at various settings and see if I can find the issue. >> On a raid0 on 4 flash devices, I get something that looks more IO >> bound, for some reason. Still no impact of the throttling, however. >> But given that your setup is this: >> >> virtio in guest, XFS direct IO -> no-op -> scsi in host. >> >> we do potentially have two throttling points, which we don't want. >> Is both the guest and the host running the new code, or just the >> guest? > > Just the guest. Host is running a 4.2.x kernel, IIRC. OK >> In any case, can I talk you into trying with two patches on top of >> the current code? It's the two newest patches here: >> >> https://urldefense.proofpoint.com/v2/url?u=http-3A__git.kernel.dk_cgit_linux-2Dblock_log_-3Fh-3Dwb-2Dbuf-2Dthrottle&d=CwIBAg&c=5VD0RTtNlTh3ycd41b3MUw&r=cK1a7KivzZRh1fKQMjSm2A&m=68CEi93IKLje5aOoxk1y9HMe_HF9pAhzxJGTmTZ7_DY&s=NeYNPvJa3VdF_EEsL8VqAQzJ4UycbXZ5PzHihwZAc_A&e= >> >> The first treats REQ_META|REQ_PRIO like they should be treated, like >> high priority IO. The second disables throttling for virtual >> devices, so we only throttle on the backend. The latter should >> probably be the other way around, but we need some way of conveying >> that information to the backend. > > I'm not changing the host kernels - it's a production machine and so > it runs long uptime testing of stable kernels. (e.g. catch slow > memory leaks, etc). So if you've disabled throttling in the guest, I > can't test the throttling changes. Right, that'd definitely hide the problem for you. I'll see if I can get it in a reproducible state and take it from there. On your host, you said it's SCSI backed, but what does the device look like? -- Jens Axboe
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-04-01 05:40 +0200 |
| Message-ID | <riY6e-6RZ-5@gated-at.bofh.it> |
| In reply to | #1368981 |
On 03/31/2016 09:29 PM, Jens Axboe wrote: >> I'm not changing the host kernels - it's a production machine and so >> it runs long uptime testing of stable kernels. (e.g. catch slow >> memory leaks, etc). So if you've disabled throttling in the guest, I >> can't test the throttling changes. > > Right, that'd definitely hide the problem for you. I'll see if I can get > it in a reproducible state and take it from there. Though on the guest, if you could try with just this one applied: http://git.kernel.dk/cgit/linux-block/commit/?h=wb-buf-throttle&id=f21fb0e42c7347bd639a17341dcd3f72c1a30d29 I'd appreciate it. It won't disable the throttling in the guest, just treat META and PRIO a bit differently. -- Jens Axboe
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-04-01 05:40 +0200 |
| Message-ID | <riY6e-6RZ-9@gated-at.bofh.it> |
| In reply to | #1368981 |
On 03/31/2016 09:29 PM, Jens Axboe wrote: >>> I can't seem to reproduce this at all. On an nvme device, I get a >>> fairly steady 60K/sec file creation rate, and we're nowhere near >>> being IO bound. So the throttling has no effect at all. >> >> That's too slow to show the stalls - your likely concurrency bound >> in allocation by the default AG count (4) from mkfs. Use mkfs.xfs -d >> agcount=32 so that every thread works in it's own AG. > > That's the key, with that I get 300-400K ops/sec instead. I'll run some > testing with this tomorrow and see what I can find, it did one full run > now and I didn't see any issues, but I need to run it at various > settings and see if I can find the issue. No stalls seen, I get the same performance with it disabled and with it enabled, at both default settings, and lower ones (wb_percent=20). Looking at iostat, we don't drive a lot of depth, so it makes sense, even with the throttling we're doing essentially the same amount of IO. What does 'nr_requests' say for your virtio_blk device? Looks like virtio_blk has a queue_depth setting, but it's not set by default, and then it uses the free entries in the ring. But I don't know what that is... -- Jens Axboe
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-04-01 08:20 +0200 |
| Message-ID | <rj0B4-jC-5@gated-at.bofh.it> |
| In reply to | #1368988 |
On Thu, Mar 31, 2016 at 09:39:25PM -0600, Jens Axboe wrote: > On 03/31/2016 09:29 PM, Jens Axboe wrote: > >>>I can't seem to reproduce this at all. On an nvme device, I get a > >>>fairly steady 60K/sec file creation rate, and we're nowhere near > >>>being IO bound. So the throttling has no effect at all. > >> > >>That's too slow to show the stalls - your likely concurrency bound > >>in allocation by the default AG count (4) from mkfs. Use mkfs.xfs -d > >>agcount=32 so that every thread works in it's own AG. > > > >That's the key, with that I get 300-400K ops/sec instead. I'll run some > >testing with this tomorrow and see what I can find, it did one full run > >now and I didn't see any issues, but I need to run it at various > >settings and see if I can find the issue. > > No stalls seen, I get the same performance with it disabled and with > it enabled, at both default settings, and lower ones > (wb_percent=20). Looking at iostat, we don't drive a lot of depth, > so it makes sense, even with the throttling we're doing essentially > the same amount of IO. Try appending numa=fake=4 to your guest's kernel command line. (that's what I'm using) > > What does 'nr_requests' say for your virtio_blk device? Looks like > virtio_blk has a queue_depth setting, but it's not set by default, > and then it uses the free entries in the ring. But I don't know what > that is... $ cat /sys/block/vdc/queue/nr_requests 128 $ Without the block throttling, guest IO (measured within the guest) looks like this over a fair proportion of the test (5s sample time) # iostat -d -x -m 5 /dev/vdc Device: rrqm/s wrqm/s r/s w/s rMB/s wMB/s avgrq-sz avgqu-sz await r_await w_await svctm %util vdc 0.00 20443.00 6.20 436.60 0.05 269.89 1248.48 73.83 146.11 486.58 141.27 1.64 72.40 vdc 0.00 11567.60 19.20 161.40 0.05 146.08 1657.12 119.17 704.57 707.25 704.25 5.34 96.48 vdc 0.00 12723.20 3.20 437.40 0.05 193.65 900.38 29.46 57.12 1.75 57.52 0.78 34.56 vdc 0.00 1739.80 22.40 426.80 0.05 123.62 563.86 23.44 62.51 79.89 61.59 1.01 45.28 vdc 0.00 12553.80 0.00 521.20 0.00 210.86 828.54 34.38 65.96 0.00 65.96 0.97 50.80 vdc 0.00 12523.60 25.60 529.60 0.10 201.94 745.29 52.24 77.73 0.41 81.47 1.14 63.20 vdc 0.00 5419.80 22.40 502.60 0.05 158.34 617.90 24.42 63.81 30.96 65.27 1.31 68.80 vdc 0.00 12059.00 0.00 439.60 0.00 174.85 814.59 30.91 70.27 0.00 70.27 0.72 31.76 vdc 0.00 7578.00 25.60 397.00 0.10 139.18 675.00 15.72 37.26 61.19 35.72 0.73 30.72 vdc 0.00 9156.00 0.00 537.40 0.00 173.57 661.45 17.08 29.62 0.00 29.62 0.53 28.72 vdc 0.00 5274.80 22.40 377.60 0.05 136.42 698.77 26.17 68.33 186.96 61.30 1.53 61.36 vdc 0.00 9407.00 3.20 541.00 0.05 174.28 656.05 36.10 66.33 3.00 66.71 0.87 47.60 vdc 0.00 8687.20 22.40 410.40 0.05 150.98 714.70 39.91 92.21 93.82 92.12 1.39 60.32 vdc 0.00 8872.80 0.00 422.60 0.00 139.28 674.96 25.01 33.03 0.00 33.03 0.91 38.40 vdc 0.00 1081.60 22.40 241.00 0.05 68.88 535.97 10.78 82.89 137.86 77.79 2.25 59.20 vdc 0.00 9826.80 0.00 445.00 0.00 167.42 770.49 45.16 101.49 0.00 101.49 1.80 79.92 vdc 0.00 7394.00 22.40 447.60 0.05 157.34 685.83 18.06 38.42 77.64 36.46 1.46 68.48 vdc 0.00 9984.80 3.20 252.00 0.05 108.46 870.82 85.68 293.73 16.75 297.24 3.00 76.64 vdc 0.00 0.00 22.40 454.20 0.05 117.67 505.86 8.11 39.51 35.71 39.70 1.17 55.76 vdc 0.00 10273.20 0.00 418.80 0.00 156.76 766.57 90.52 179.40 0.00 179.40 1.85 77.52 vdc 0.00 5650.00 22.40 185.00 0.05 84.12 831.20 103.90 575.15 60.82 637.42 4.21 87.36 vdc 0.00 7193.00 0.00 308.80 0.00 120.71 800.56 63.77 194.35 0.00 194.35 2.24 69.12 vdc 0.00 4460.80 9.80 211.00 0.03 69.52 645.07 72.35 154.81 269.39 149.49 4.42 97.60 vdc 0.00 683.00 14.00 374.60 0.05 99.13 522.69 25.38 167.61 603.14 151.33 1.45 56.24 vdc 0.00 7140.20 1.80 275.20 0.03 104.53 773.06 85.25 202.67 32.44 203.79 2.80 77.68 vdc 0.00 6916.00 0.00 164.00 0.00 82.59 1031.33 126.20 813.60 0.00 813.60 6.10 100.00 vdc 0.00 2255.60 22.40 359.00 0.05 107.41 577.06 42.97 170.03 92.79 174.85 2.17 82.64 vdc 0.00 7580.40 3.20 370.40 0.05 128.32 703.70 60.19 134.11 15.00 135.14 1.64 61.36 vdc 0.00 6438.40 18.80 159.20 0.04 78.04 898.36 126.80 706.27 639.15 714.19 5.62 100.00 vdc 0.00 5420.00 3.60 315.40 0.01 108.87 699.07 20.80 78.54 580.00 72.81 1.03 32.72 vdc 0.00 9444.00 2.60 242.40 0.00 118.72 992.38 126.21 488.66 146.15 492.33 4.08 100.00 vdc 0.00 0.00 19.80 434.60 0.05 110.14 496.65 12.74 57.56 313.78 45.89 1.10 49.84 vdc 0.00 14108.20 3.20 549.60 0.05 207.84 770.17 42.32 69.66 72.75 69.64 1.40 77.20 vdc 0.00 1306.40 35.20 268.20 0.08 78.74 532.08 30.84 114.22 175.07 106.24 2.02 61.20 vdc 0.00 14999.40 0.00 458.60 0.00 192.03 857.57 61.48 134.02 0.00 134.02 1.67 76.80 vdc 0.00 1.40 22.40 331.80 0.05 82.11 475.11 1.74 4.87 22.68 3.66 0.76 26.96 vdc 0.00 13971.80 0.00 670.20 0.00 248.26 758.63 34.45 51.37 0.00 51.37 1.04 69.52 vdc 0.00 7033.00 22.60 205.80 0.06 87.81 787.86 40.95 128.53 244.64 115.78 2.90 66.24 vdc 0.00 1282.00 3.20 456.00 0.05 123.21 549.74 14.56 46.99 21.00 47.17 1.42 65.20 vdc 0.00 9475.80 22.40 248.60 0.05 107.66 814.02 123.94 412.61 376.64 415.86 3.69 100.00 vdc 0.00 3603.60 0.00 418.80 0.00 133.32 651.94 71.28 210.08 0.00 210.08 1.77 74.00 You can see hat there are periods where it drives the request queue depth to congestion, but most of the time the device is only 60-70% utilised and the queue depths are only 30-40 deep. THere's quite a lot of idle time in the request queue. Note that there are a couple of points where merging stops completely - that's when memory reclaim is directly flushing dirty inodes because if all we have is cached inodes then we have toi throttle memory allocation back to the rate we at which we can clean dirty inodes. Throughput does drop when this happens, but because the device has idle overhead and spare request queue space, these less than optimal IO dispatch spikes don't really affect throughput because the device has the capacity available to soak them up without dropping performance. An equivalent trace from the middle of a run with block throttling enabled: Device: rrqm/s wrqm/s r/s w/s rMB/s wMB/s avgrq-sz avgqu-sz await r_await w_await svctm %util vdc 0.00 5143.40 22.40 188.00 0.05 81.04 789.38 19.17 89.09 237.86 71.37 4.75 99.92 vdc 0.00 7182.60 0.00 272.60 0.00 116.74 877.06 15.50 57.91 0.00 57.91 3.66 99.84 vdc 0.00 3732.80 11.60 102.20 0.01 53.17 957.05 19.35 151.47 514.00 110.32 8.79 100.00 vdc 0.00 0.00 10.80 1007.20 0.04 104.33 209.98 10.22 12.49 457.85 7.71 0.92 93.44 vdc 0.00 0.00 0.40 822.80 0.01 47.58 118.40 9.39 11.21 10.00 11.21 1.22 100.24 vdc 0.00 0.00 1.00 227.00 0.02 3.55 32.00 0.22 1.54 224.00 0.56 0.48 11.04 vdc 0.00 11100.40 3.20 437.40 0.05 125.91 585.49 7.95 17.77 47.25 17.55 1.56 68.56 vdc 0.00 14134.20 22.40 453.20 0.05 191.26 823.85 15.99 32.91 73.07 30.92 2.09 99.28 vdc 0.00 6667.00 0.00 265.20 0.00 105.11 811.70 15.71 58.98 0.00 58.98 3.77 100.00 vdc 0.00 6243.40 22.40 259.00 0.05 101.23 737.11 17.53 62.97 115.21 58.45 3.55 99.92 vdc 0.00 5590.80 0.00 278.20 0.00 105.30 775.18 18.09 65.55 0.00 65.55 3.59 100.00 vdc 0.00 0.00 14.20 714.80 0.02 97.81 274.86 11.61 12.86 260.85 7.93 1.23 89.44 vdc 0.00 0.00 9.80 1555.00 0.05 126.19 165.22 5.41 4.91 267.02 3.26 0.53 82.96 vdc 0.00 0.00 3.00 816.80 0.05 22.32 55.89 6.07 7.39 256.00 6.48 1.05 85.84 vdc 0.00 11172.80 0.20 463.00 0.00 125.77 556.10 6.13 13.23 260.00 13.13 0.93 43.28 vdc 0.00 9563.00 22.40 324.60 0.05 119.66 706.55 15.50 38.45 10.39 40.39 2.88 99.84 vdc 0.00 5333.60 0.00 218.00 0.00 83.71 786.46 15.57 80.46 0.00 80.46 4.59 100.00 vdc 0.00 5128.00 24.80 216.60 0.06 85.31 724.28 19.18 79.84 193.13 66.87 4.12 99.52 vdc 0.00 2746.40 0.00 257.40 0.00 81.13 645.49 11.16 43.70 0.00 43.70 3.87 99.68 vdc 0.00 0.00 0.00 418.80 0.00 104.68 511.92 5.33 12.74 0.00 12.74 1.93 80.96 vdc 0.00 8102.00 0.20 291.60 0.00 108.79 763.59 3.09 10.60 20.00 10.59 0.87 25.44 The first thing to note is the device utilisation is almost always above 80%, and often at 100%, meaning with throttling the device always has IO in flight. It's got no real idle time to soak up peaks of IO activity - throttling means the device is running at close to 100% utilisation all the time under worklaods like this, and there's not elasticity in the pipeline to handle changes in IO dispatch behaviour. So when memory reclaim does direct inode writeback, we see merging stop, but the request queue is not able to soak up all the IO being dispatched, even though there is very little read IO demand. hence changes in the dispatch patterns that would drive deeper queues and maintain performance will now get throttled, resulting in things like memory reclaim backing up a lot and everything on the machine suffering. I'll try the "don't throttle REQ_META" patch, but this seems like a fragile way to solve this problem - it shuts up the messenger, but doesn't solve the problem for any other subsystem that might have a similer issue. e.g. next we're going to have to make sure direct IO (which is also REQ_WRITE dispatch) does not get throttled, and so on.... It seems to me that the right thing to do here is add a separate classification flag for IO that can be throttled. e.g. as REQ_WRITEBACK and only background writeback work sets this flag. That would ensure that when the IO is being dispatched from other sources (e.g. fsync, sync_file_range(), direct IO, filesystem metadata, etc) it is clear that it is not a target for throttling. This would also allow us to easily switch off throttling if writeback is occurring for memory reclaim reasons, and so on. Throttling policy decisions belong above the block layer, even though the throttle mechanism itself is in the block layer. FWIW, this is analogous to REQ_READA, which tells the block layer that a read is not important and can be discarded if there is too much load. Policy is set at the layer that knows whether the IO can be discarded safely, the mechanism is implemented at a lower layer that knows about load, scheduling and other things the higher layers know nothing about. Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-04-01 16:40 +0200 |
| Message-ID | <rj8oW-5Jd-9@gated-at.bofh.it> |
| In reply to | #1369013 |
On 04/01/2016 12:16 AM, Dave Chinner wrote: > On Thu, Mar 31, 2016 at 09:39:25PM -0600, Jens Axboe wrote: >> On 03/31/2016 09:29 PM, Jens Axboe wrote: >>>>> I can't seem to reproduce this at all. On an nvme device, I get a >>>>> fairly steady 60K/sec file creation rate, and we're nowhere near >>>>> being IO bound. So the throttling has no effect at all. >>>> >>>> That's too slow to show the stalls - your likely concurrency bound >>>> in allocation by the default AG count (4) from mkfs. Use mkfs.xfs -d >>>> agcount=32 so that every thread works in it's own AG. >>> >>> That's the key, with that I get 300-400K ops/sec instead. I'll run some >>> testing with this tomorrow and see what I can find, it did one full run >>> now and I didn't see any issues, but I need to run it at various >>> settings and see if I can find the issue. >> >> No stalls seen, I get the same performance with it disabled and with >> it enabled, at both default settings, and lower ones >> (wb_percent=20). Looking at iostat, we don't drive a lot of depth, >> so it makes sense, even with the throttling we're doing essentially >> the same amount of IO. > > Try appending numa=fake=4 to your guest's kernel command line. > > (that's what I'm using) Sure, I can give that a go. >> What does 'nr_requests' say for your virtio_blk device? Looks like >> virtio_blk has a queue_depth setting, but it's not set by default, >> and then it uses the free entries in the ring. But I don't know what >> that is... > > $ cat /sys/block/vdc/queue/nr_requests > 128 OK, so that would put you in the 16/32/64 category for idle/normal/high priority writeback. Which fits with the iostat below, which is in the ~16 range. So the META thing should help, it'll bump it up a bit. But we're also seeing smaller requests, and I think that could be because after we do throttle, we could potentially have a merge candidate. The code doesn't check post-sleeping, it'll allow any merges before though. Though that part is a little harder to read from the iostat numbers, but there does seem to be a correlation between your higher depths and bigger request sizes. > I'll try the "don't throttle REQ_META" patch, but this seems like a > fragile way to solve this problem - it shuts up the messenger, but > doesn't solve the problem for any other subsystem that might have a > similer issue. e.g. next we're going to have to make sure direct IO > (which is also REQ_WRITE dispatch) does not get throttled, and so > on.... I don't think there's anything wrong with the REQ_META patch. Sure, we could have better classifications (like discussed below), but that's mainly tweaking. As long as we get the same answers, it's fine. There's no throttling of O_DIRECT writes in the current code, it specifically doesn't include those. It's only for the unbounded writes, which writeback tends to be. > It seems to me that the right thing to do here is add a separate > classification flag for IO that can be throttled. e.g. as > REQ_WRITEBACK and only background writeback work sets this flag. > That would ensure that when the IO is being dispatched from other > sources (e.g. fsync, sync_file_range(), direct IO, filesystem > metadata, etc) it is clear that it is not a target for throttling. > This would also allow us to easily switch off throttling if > writeback is occurring for memory reclaim reasons, and so on. > Throttling policy decisions belong above the block layer, even > though the throttle mechanism itself is in the block layer. We're already doing all of that, it's just doesn't include a specific REQ_WRITEBACK flag. And yeah, that would clean up the checking for request type, but functionally it should be the same as it is now. It'll be a bit more robust and easier to read if we just have a REQ_WRITEBACK, right now it's WRITE_SYNC vs WRITE for important vs not-important, with a check for write vs O_DIRECT write as well. -- Jens Axboe
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-04-01 07:10 +0200 |
| Message-ID | <riZvj-84e-1@gated-at.bofh.it> |
| In reply to | #1368981 |
On Thu, Mar 31, 2016 at 09:29:30PM -0600, Jens Axboe wrote: > On 03/31/2016 06:56 PM, Dave Chinner wrote: > >I'm not changing the host kernels - it's a production machine and so > >it runs long uptime testing of stable kernels. (e.g. catch slow > >memory leaks, etc). So if you've disabled throttling in the guest, I > >can't test the throttling changes. > > Right, that'd definitely hide the problem for you. I'll see if I can > get it in a reproducible state and take it from there. > > On your host, you said it's SCSI backed, but what does the device look like? HW RAID 0 w/ 1GB FBWC (dell h710, IIRC) of 2x200GB SATA SSDs (actually 256GB, but 25% of each is left as spare, unused space). Sustains about 35,000 random 4k write IOPS, up to 70k read IOPS. Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-04-01 02:50 +0200 |
| Message-ID | <riVrI-4Iz-5@gated-at.bofh.it> |
| In reply to | #1368347 |
On Thu, Mar 31, 2016 at 08:29:35AM -0600, Jens Axboe wrote:
> On 03/31/2016 02:24 AM, Dave Chinner wrote:
> >On Wed, Mar 30, 2016 at 09:07:48AM -0600, Jens Axboe wrote:
> >>Hi,
> >>
> >>This patchset isn't as much a final solution, as it's demonstration
> >>of what I believe is a huge issue. Since the dawn of time, our
> >>background buffered writeback has sucked. When we do background
> >>buffered writeback, it should have little impact on foreground
> >>activity. That's the definition of background activity... But for as
> >>long as I can remember, heavy buffered writers has not behaved like
> >>that. For instance, if I do something like this:
> >>
> >>$ dd if=/dev/zero of=foo bs=1M count=10k
> >>
> >>on my laptop, and then try and start chrome, it basically won't start
> >>before the buffered writeback is done. Or, for server oriented
> >>workloads, where installation of a big RPM (or similar) adversely
> >>impacts data base reads or sync writes. When that happens, I get people
> >>yelling at me.
> >>
> >>Last time I posted this, I used flash storage as the example. But
> >>this works equally well on rotating storage. Let's run a test case
> >>that writes a lot. This test writes 50 files, each 100M, on XFS on
> >>a regular hard drive. While this happens, we attempt to read
> >>another file with fio.
> >>
> >>Writers:
> >>
> >>$ time (./write-files ; sync)
> >>real 1m6.304s
> >>user 0m0.020s
> >>sys 0m12.210s
> >
> >Great. So a basic IO tests looks good - let's through something more
> >complex at it. Say, a benchmark I've been using for years to stress
> >the Io subsystem, the filesystem and memory reclaim all at the same
> >time: a concurent fsmark inode creation test.
> >(first google hit https://lkml.org/lkml/2013/9/10/46)
>
> Is that how you are invoking it as well same arguments?
Yes. And the VM is exactly the same, too - 16p/16GB RAM. Cut down
version of the script I use:
#!/bin/bash
QUOTA=
MKFSOPTS=
NFILES=100000
DEV=/dev/vdc
LOGBSIZE=256k
FSMARK=/home/dave/src/fs_mark-3.3/fs_mark
MNT=/mnt/scratch
while [ $# -gt 0 ]; do
case "$1" in
-q) QUOTA="uquota,gquota,pquota" ;;
-N) NFILES=$2 ; shift ;;
-d) DEV=$2 ; shift ;;
-l) LOGBSIZE=$2; shift ;;
--) shift ; break ;;
esac
shift
done
MKFSOPTS="$MKFSOPTS $*"
echo QUOTA=$QUOTA
echo MKFSOPTS=$MKFSOPTS
echo DEV=$DEV
sudo umount $MNT > /dev/null 2>&1
sudo mkfs.xfs -f $MKFSOPTS $DEV
sudo mount -o nobarrier,logbsize=$LOGBSIZE,$QUOTA $DEV $MNT
sudo chmod 777 $MNT
sudo sh -c "echo 1 > /proc/sys/fs/xfs/stats_clear"
time $FSMARK -D 10000 -S0 -n $NFILES -s 0 -L 32 \
-d $MNT/0 -d $MNT/1 \
-d $MNT/2 -d $MNT/3 \
-d $MNT/4 -d $MNT/5 \
-d $MNT/6 -d $MNT/7 \
-d $MNT/8 -d $MNT/9 \
-d $MNT/10 -d $MNT/11 \
-d $MNT/12 -d $MNT/13 \
-d $MNT/14 -d $MNT/15 \
| tee >(stats --trim-outliers | tail -1 1>&2)
sync
sudo umount /mnt/scratch
$
> >>The above was run without scsi-mq, and with using the deadline scheduler,
> >>results with CFQ are similary depressing for this test. So IO scheduling
> >>is in place for this test, it's not pure blk-mq without scheduling.
> >
> >virtio in guest, XFS direct IO -> no-op -> scsi in host.
>
> That has write back caching enabled on the guest, correct?
No. It uses virtio,cache=none (that's the "XFS Direct IO" bit above).
Sorry for not being clear about that.
Cheers,
Dave.
--
Dave Chinner
david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-04-01 05:30 +0200 |
| Message-ID | <riXWx-6Md-1@gated-at.bofh.it> |
| In reply to | #1368915 |
On 03/31/2016 06:46 PM, Dave Chinner wrote: > On Thu, Mar 31, 2016 at 08:29:35AM -0600, Jens Axboe wrote: >> On 03/31/2016 02:24 AM, Dave Chinner wrote: >>> On Wed, Mar 30, 2016 at 09:07:48AM -0600, Jens Axboe wrote: >>>> Hi, >>>> >>>> This patchset isn't as much a final solution, as it's demonstration >>>> of what I believe is a huge issue. Since the dawn of time, our >>>> background buffered writeback has sucked. When we do background >>>> buffered writeback, it should have little impact on foreground >>>> activity. That's the definition of background activity... But for as >>>> long as I can remember, heavy buffered writers has not behaved like >>>> that. For instance, if I do something like this: >>>> >>>> $ dd if=/dev/zero of=foo bs=1M count=10k >>>> >>>> on my laptop, and then try and start chrome, it basically won't start >>>> before the buffered writeback is done. Or, for server oriented >>>> workloads, where installation of a big RPM (or similar) adversely >>>> impacts data base reads or sync writes. When that happens, I get people >>>> yelling at me. >>>> >>>> Last time I posted this, I used flash storage as the example. But >>>> this works equally well on rotating storage. Let's run a test case >>>> that writes a lot. This test writes 50 files, each 100M, on XFS on >>>> a regular hard drive. While this happens, we attempt to read >>>> another file with fio. >>>> >>>> Writers: >>>> >>>> $ time (./write-files ; sync) >>>> real 1m6.304s >>>> user 0m0.020s >>>> sys 0m12.210s >>> >>> Great. So a basic IO tests looks good - let's through something more >>> complex at it. Say, a benchmark I've been using for years to stress >>> the Io subsystem, the filesystem and memory reclaim all at the same >>> time: a concurent fsmark inode creation test. >>> (first google hit https://lkml.org/lkml/2013/9/10/46) >> >> Is that how you are invoking it as well same arguments? > > Yes. And the VM is exactly the same, too - 16p/16GB RAM. Cut down > version of the script I use: > > #!/bin/bash > > QUOTA= > MKFSOPTS= > NFILES=100000 > DEV=/dev/vdc > LOGBSIZE=256k > FSMARK=/home/dave/src/fs_mark-3.3/fs_mark > MNT=/mnt/scratch > > while [ $# -gt 0 ]; do > case "$1" in > -q) QUOTA="uquota,gquota,pquota" ;; > -N) NFILES=$2 ; shift ;; > -d) DEV=$2 ; shift ;; > -l) LOGBSIZE=$2; shift ;; > --) shift ; break ;; > esac > shift > done > MKFSOPTS="$MKFSOPTS $*" > > echo QUOTA=$QUOTA > echo MKFSOPTS=$MKFSOPTS > echo DEV=$DEV > > sudo umount $MNT > /dev/null 2>&1 > sudo mkfs.xfs -f $MKFSOPTS $DEV > sudo mount -o nobarrier,logbsize=$LOGBSIZE,$QUOTA $DEV $MNT > sudo chmod 777 $MNT > sudo sh -c "echo 1 > /proc/sys/fs/xfs/stats_clear" > time $FSMARK -D 10000 -S0 -n $NFILES -s 0 -L 32 \ > -d $MNT/0 -d $MNT/1 \ > -d $MNT/2 -d $MNT/3 \ > -d $MNT/4 -d $MNT/5 \ > -d $MNT/6 -d $MNT/7 \ > -d $MNT/8 -d $MNT/9 \ > -d $MNT/10 -d $MNT/11 \ > -d $MNT/12 -d $MNT/13 \ > -d $MNT/14 -d $MNT/15 \ > | tee >(stats --trim-outliers | tail -1 1>&2) > sync > sudo umount /mnt/scratch Perfect, thanks! >>>> The above was run without scsi-mq, and with using the deadline scheduler, >>>> results with CFQ are similary depressing for this test. So IO scheduling >>>> is in place for this test, it's not pure blk-mq without scheduling. >>> >>> virtio in guest, XFS direct IO -> no-op -> scsi in host. >> >> That has write back caching enabled on the guest, correct? > > No. It uses virtio,cache=none (that's the "XFS Direct IO" bit above). > Sorry for not being clear about that. That's fine, it's one less worry if that's not the case. So if you cat the 'write_cache' file in the virtioblk sysfs block queue/ directory, it says 'write through'? Just want to confirm that we got that propagated correctly. -- Jens Axboe
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-04-01 08:30 +0200 |
| Message-ID | <rj0KK-nf-3@gated-at.bofh.it> |
| In reply to | #1368980 |
On Thu, Mar 31, 2016 at 09:25:33PM -0600, Jens Axboe wrote: > On 03/31/2016 06:46 PM, Dave Chinner wrote: > >>>virtio in guest, XFS direct IO -> no-op -> scsi in host. > >> > >>That has write back caching enabled on the guest, correct? > > > >No. It uses virtio,cache=none (that's the "XFS Direct IO" bit above). > >Sorry for not being clear about that. > > That's fine, it's one less worry if that's not the case. So if you > cat the 'write_cache' file in the virtioblk sysfs block queue/ > directory, it says 'write through'? Just want to confirm that we got > that propagated correctly. No such file. But I did find: $ cat /sys/block/vdc/cache_type write back Which is what I'd expect it to safe given the man page description of cache=none: Note that this is considered a writeback mode and the guest OS must handle the disk write cache correctly in order to avoid data corruption on host crashes. To make it say "write through" I need to use cache=directsync, but I have no need for such integrity guarantees on a volatile test device... Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@fb.com> |
|---|---|
| Date | 2016-04-01 16:40 +0200 |
| Message-ID | <rj8oX-5Jd-35@gated-at.bofh.it> |
| In reply to | #1369014 |
On 04/01/2016 12:27 AM, Dave Chinner wrote: > On Thu, Mar 31, 2016 at 09:25:33PM -0600, Jens Axboe wrote: >> On 03/31/2016 06:46 PM, Dave Chinner wrote: >>>>> virtio in guest, XFS direct IO -> no-op -> scsi in host. >>>> >>>> That has write back caching enabled on the guest, correct? >>> >>> No. It uses virtio,cache=none (that's the "XFS Direct IO" bit above). >>> Sorry for not being clear about that. >> >> That's fine, it's one less worry if that's not the case. So if you >> cat the 'write_cache' file in the virtioblk sysfs block queue/ >> directory, it says 'write through'? Just want to confirm that we got >> that propagated correctly. > > No such file. But I did find: > > $ cat /sys/block/vdc/cache_type > write back > > Which is what I'd expect it to safe given the man page description > of cache=none: > > Note that this is considered a writeback mode and the guest > OS must handle the disk write cache correctly in order to > avoid data corruption on host crashes. > > To make it say "write through" I need to use cache=directsync, but > I have no need for such integrity guarantees on a volatile test > device... I wasn't as concerned about the integrity side, more if it's flagged as write back then we induce further throttling. But I'll see if I can get your test case reproduced, then I don't see why it can't get fixed. I'm off all of next week though, so probably won't be until the week after... -- Jens Axboe
[toc] | [prev] | [next] | [standalone]
| From | Holger Hoffstätte <holger.hoffstaette@googlemail.com> |
|---|---|
| Date | 2016-04-01 00:20 +0200 |
| Message-ID | <riT6y-3a6-15@gated-at.bofh.it> |
| In reply to | #1367290 |
Hi, Jens mentioned on Twitter I should post my experience here as well, so here we go. I've backported this series (incl. updates) to stable-4.4.x - not too difficult, minus the NVM part which I don't need anyway - and have been running it for the past few days without any problem whatsoever, with GREAT success. My use case is primarily larger amounts of stuff (transcoded movies, finished downloads, built Gentoo packages) that gets copied from tmpfs to SSD (or disk) and every time that happens, the system noticeably strangles readers (desktop, interactive shell). It does not really matter how I tune writeback via the write_expire/dirty_bytes knobs or the scheduler (and yes, I understand how they work); lowering the writeback limits helped a bit but the system is still overwhelmed. Jacking up deadline's writes_starved to unreasonable levels helps a bit, but in turn makes all writes suffer. Anything else - even tried BFQ for a while, which has its own unrelated problems - didn't really help either. With this patchset the buffered writeback in these situations is much improved, and copying several GBs at once to a SATA-3 SSD (or even an external USB-2 disk with measly 40 MB/s) doodles along in the background like it always should have, and desktop work is not noticeably affected. I guess the effect will be even more noticeable on slower block devices (laptops, old SSDs or disks). So: +1 would apply again! cheers Holger
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-04-01 03:10 +0200 |
| Message-ID | <riVL5-57m-17@gated-at.bofh.it> |
| In reply to | #1368859 |
On Thu, Mar 31, 2016 at 10:09:56PM +0000, Holger Hoffstätte wrote: > > Hi, > > Jens mentioned on Twitter I should post my experience here as well, > so here we go. > > I've backported this series (incl. updates) to stable-4.4.x - not too > difficult, minus the NVM part which I don't need anyway - and have been > running it for the past few days without any problem whatsoever, with > GREAT success. > > My use case is primarily larger amounts of stuff (transcoded movies, > finished downloads, built Gentoo packages) that gets copied from tmpfs > to SSD (or disk) and every time that happens, the system noticeably > strangles readers (desktop, interactive shell). It does not really matter > how I tune writeback via the write_expire/dirty_bytes knobs or the > scheduler (and yes, I understand how they work); lowering the writeback > limits helped a bit but the system is still overwhelmed. Jacking up > deadline's writes_starved to unreasonable levels helps a bit, but in turn > makes all writes suffer. Anything else - even tried BFQ for a while, > which has its own unrelated problems - didn't really help either. Can you go back to your original kernel, and lower nr_requests to 8? Essentially all I see the block throttle doing is keeping the request queue depth to somewhere between 8-12 requests, rather than letting it blow out to near nr_requests (around 105-115), so it would be interesting to note whether the block throttling has any noticable difference in behaviour when compared to just having a very shallow request queue.... Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Holger Hoffstätte <holger.hoffstaette@googlemail.com> |
|---|---|
| Date | 2016-04-01 19:10 +0200 |
| Message-ID | <rjaK5-7AB-17@gated-at.bofh.it> |
| In reply to | #1368922 |
On 04/01/16 03:01, Dave Chinner wrote:
> Can you go back to your original kernel, and lower nr_requests to 8?
Sure, did that and as expected it didn't help much. Under prolonged stress
it was actually even a bit worse than writeback throttling. IMHO that's not
really surprising either, since small queues now punish everyone and
in interactive mode I really want to e.g. start loading hundreds of small
thumbnails at once, or du a directory.
Instead of randomized aka manual/interactive testing I created a simple
stress tester:
#!/bin/sh
while [[ true ]]
do
cp bigfile bigfile.out
done
and running that in the background turns the system into a tar pit,
which is laughable when you consider that I have 24G and 8 cores.
With the writeback patchset and wb_percent=1 (yes, really!) it is almost
unnoticeable, but according to nmon still writes ~250-280 MB/s.
This is with deadline on ext4 on an older SATA-3 SSD that can still
do peak ~465 MB/s (with dd).
cheers,
Holger
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web