Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1354530 > unrolled thread
| Started by | Gregory Farnum <greg@gregs42.com> |
|---|---|
| First post | 2016-03-09 23:30 +0100 |
| Last post | 2016-03-12 01:40 +0100 |
| Articles | 4 on this page of 44 — 13 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Gregory Farnum <greg@gregs42.com> - 2016-03-09 23:30 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-10 00:10 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Ric Wheeler <ricwheeler@gmail.com> - 2016-03-10 16:00 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Linus Torvalds <torvalds@linux-foundation.org> - 2016-03-10 19:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-10 22:50 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Ric Wheeler <rwheeler@redhat.com> - 2016-03-11 05:50 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks One Thousand Gnomes <gnomes@lxorguk.ukuu.org.uk> - 2016-03-11 15:10 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-11 16:30 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Linus Torvalds <torvalds@linux-foundation.org> - 2016-03-11 18:30 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Andy Lutomirski <luto@amacapital.net> - 2016-03-11 18:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Linus Torvalds <torvalds@linux-foundation.org> - 2016-03-11 19:30 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Dave Chinner <david@fromorbit.com> - 2016-03-11 23:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-12 01:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Linus Torvalds <torvalds@linux-foundation.org> - 2016-03-12 01:50 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-12 08:30 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Thomas Schoebel-Theuer <tst@schoebel-theuer.de> - 2016-03-12 11:20 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Dave Chinner <david@fromorbit.com> - 2016-03-14 00:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Ric Wheeler <rwheeler@redhat.com> - 2016-03-14 11:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-14 15:50 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Dave Chinner <david@fromorbit.com> - 2016-03-15 21:20 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Linus Torvalds <torvalds@linux-foundation.org> - 2016-03-15 21:50 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-15 22:30 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Dave Chinner <david@fromorbit.com> - 2016-03-15 23:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-16 00:00 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks "Darrick J. Wong" <darrick.wong@oracle.com> - 2016-03-16 03:00 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Andreas Dilger <adilger@dilger.ca> - 2016-03-16 22:50 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-17 01:20 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Eric Sandeen <esandeen@redhat.com> - 2016-03-17 01:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-17 02:00 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Gregory Farnum <greg@gregs42.com> - 2016-03-17 06:20 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Theodore Ts'o <tytso@mit.edu> - 2016-03-17 13:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Dave Chinner <david@fromorbit.com> - 2016-03-17 02:10 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks "Darrick J. Wong" <darrick.wong@oracle.com> - 2016-03-17 03:50 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Linus Torvalds <torvalds@linux-foundation.org> - 2016-03-16 00:10 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Linus Torvalds <torvalds@linux-foundation.org> - 2016-03-16 00:20 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Dave Chinner <david@fromorbit.com> - 2016-03-16 01:10 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Dave Chinner <david@fromorbit.com> - 2016-03-16 01:00 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Linus Torvalds <torvalds@linux-foundation.org> - 2016-03-16 01:10 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Eric Sandeen <esandeen@redhat.com> - 2016-03-16 01:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Chris Mason <clm@fb.com> - 2016-03-16 02:00 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Chris Mason <clm@fb.com> - 2016-03-16 23:30 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Ric Wheeler <rwheeler@redhat.com> - 2016-03-17 14:50 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Eric Sandeen <esandeen@redhat.com> - 2016-03-15 23:40 +0100
Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks Linus Torvalds <torvalds@linux-foundation.org> - 2016-03-12 01:40 +0100
Page 3 of 3 — ← Prev page 1 2 [3]
| From | Chris Mason <clm@fb.com> |
|---|---|
| Date | 2016-03-16 23:30 +0100 |
| Subject | Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks |
| Message-ID | <rds6Z-6Wr-1@gated-at.bofh.it> |
| In reply to | #1358458 |
On Tue, Mar 15, 2016 at 05:51:17PM -0700, Chris Mason wrote: > On Tue, Mar 15, 2016 at 07:30:14PM -0500, Eric Sandeen wrote: > > On 3/15/16 7:06 PM, Linus Torvalds wrote: > > > On Tue, Mar 15, 2016 at 4:52 PM, Dave Chinner <david@fromorbit.com> wrote: > > >> > > > >> > It is pretty clear that the onus is on the patch submitter to > > >> > provide justification for inclusion, not for the reviewer/Maintainer > > >> > to have to prove that the solution is unworkable. > > > I agree, but quite frankly, performance is a good justification. > > > > > > So if Ted can give performance numbers, that's justification enough. > > > We've certainly taken changes with less. > > > > I've been away from ext4 for a while, so I'm really not on top of the > > mechanics of the underlying problem at the moment. > > > > But I would say that in addition to numbers showing that ext4 has trouble > > with unwritten extent conversion, we should have an explanation of > > why it can't be solved in a way that doesn't open up these concerns. > > > > XFS certainly has different mechanisms, but is the demonstrated workload > > problematic on XFS (or btrfs) as well? If not, can ext4 adopt any of the > > solutions that make the workload perform better on other filesystems? > > When I've benchmarked this in the past, doing small random buffered writes > into an preallocated extent was dramatically (3x or more) slower on xfs > than doing them into a fully written extent. That was two years ago, > but I can redo it. So I re-ran some benchmarks, with 4K O_DIRECT random ios on nvme (4.5 kernel). This is O_DIRECT without O_SYNC. I don't think xfs will do commits for each IO into the prealloc file? O_SYNC makes it much slower, so hopefully I've got this right. The test runs for 60 seconds, and I used an iodepth of 4: prealloc file: 32,000 iops overwrite: 121,000 iops If I bump the iodepth up to 512: prealloc file: 33,000 iops overwrite: 279,000 iops For streaming writes, XFS converts prealloc to written much better when the IO isn't random. You can start seeing the difference at 16K sequential O_DIRECT writes, but really its not a huge impact. The worst case is 4K: prealloc file: 227MB/s overwrite: 340MB/s I can't think of sequential workloads where this will matter, since they will either end up with bigger IO or the performance impact won't get noticed. -chris
[toc] | [prev] | [next] | [standalone]
| From | Ric Wheeler <rwheeler@redhat.com> |
|---|---|
| Date | 2016-03-17 14:50 +0100 |
| Subject | Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks |
| Message-ID | <rdGtj-844-5@gated-at.bofh.it> |
| In reply to | #1359399 |
On 03/16/2016 06:23 PM, Chris Mason wrote: > On Tue, Mar 15, 2016 at 05:51:17PM -0700, Chris Mason wrote: >> On Tue, Mar 15, 2016 at 07:30:14PM -0500, Eric Sandeen wrote: >>> On 3/15/16 7:06 PM, Linus Torvalds wrote: >>>> On Tue, Mar 15, 2016 at 4:52 PM, Dave Chinner <david@fromorbit.com> wrote: >>>>>> It is pretty clear that the onus is on the patch submitter to >>>>>> provide justification for inclusion, not for the reviewer/Maintainer >>>>>> to have to prove that the solution is unworkable. >>>> I agree, but quite frankly, performance is a good justification. >>>> >>>> So if Ted can give performance numbers, that's justification enough. >>>> We've certainly taken changes with less. >>> I've been away from ext4 for a while, so I'm really not on top of the >>> mechanics of the underlying problem at the moment. >>> >>> But I would say that in addition to numbers showing that ext4 has trouble >>> with unwritten extent conversion, we should have an explanation of >>> why it can't be solved in a way that doesn't open up these concerns. >>> >>> XFS certainly has different mechanisms, but is the demonstrated workload >>> problematic on XFS (or btrfs) as well? If not, can ext4 adopt any of the >>> solutions that make the workload perform better on other filesystems? >> When I've benchmarked this in the past, doing small random buffered writes >> into an preallocated extent was dramatically (3x or more) slower on xfs >> than doing them into a fully written extent. That was two years ago, >> but I can redo it. > So I re-ran some benchmarks, with 4K O_DIRECT random ios on nvme (4.5 > kernel). This is O_DIRECT without O_SYNC. I don't think xfs will do > commits for each IO into the prealloc file? O_SYNC makes it much > slower, so hopefully I've got this right. > > The test runs for 60 seconds, and I used an iodepth of 4: > > prealloc file: 32,000 iops > overwrite: 121,000 iops > > If I bump the iodepth up to 512: > > prealloc file: 33,000 iops > overwrite: 279,000 iops > > For streaming writes, XFS converts prealloc to written much better when > the IO isn't random. You can start seeing the difference at 16K > sequential O_DIRECT writes, but really its not a huge impact. The worst > case is 4K: > > prealloc file: 227MB/s > overwrite: 340MB/s > > I can't think of sequential workloads where this will matter, since they > will either end up with bigger IO or the performance impact won't get > noticed. > > -chris I think that these numbers are the interesting ones, see a 4x slow down is certainly significant. If you do the same patch after hacking XFS preallocation as Dave suggested with xfs_db, do we get most of the performance back? Ric
[toc] | [prev] | [next] | [standalone]
| From | Eric Sandeen <esandeen@redhat.com> |
|---|---|
| Date | 2016-03-15 23:40 +0100 |
| Subject | Re: [PATCH 2/2] block: create ioctl to discard-or-zeroout a range of blocks |
| Message-ID | <rd5N7-8tc-5@gated-at.bofh.it> |
| In reply to | #1358224 |
On 3/15/16 3:14 PM, Dave Chinner wrote: > What we are missing is actual numbers that show that exposing stale > data is a /significant/ win for these applications that are > demanding it. And then we need evidence proving that the problem is > actually systemic and not just a hack around a bad implementation of > a feature... Thanks Dave; I totally agree on this point. We've spent more than enough time talking about how and if to implement stale data exposure, but nowhere in this thread has there been any actual performance data indicating why we should do it at all. -Eric
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2016-03-12 01:40 +0100 |
| Message-ID | <rbFL3-63Y-3@gated-at.bofh.it> |
| In reply to | #1356243 |
On Fri, Mar 11, 2016 at 2:30 PM, Dave Chinner <david@fromorbit.com> wrote:
> On Fri, Mar 11, 2016 at 10:25:30AM -0800, Linus Torvalds wrote:
>>
>> So you'd have to explicitly say "my setup is ok with hole punching".
>
> Except it's not hole punching that is the problem. [..]
> The problem here is
> preallocation of unwritten blocks that expose the stale data if the
> filesystem skips marking those blocks as unwritten.
Right you are.
> It's all well and good to restrict access to the fallocate() call to
> limit who can expose stale data, but it doesn't remove the fact it
> is easy for stale data to unintentionally escape the privileged
> group once it has been exposed because there is no record of the
> fact the file contains uninitialised blocks....
Good point.
It's not just that the user in question *couldn't* have exposed it
other ways by reading the raw device and then writing that info into a
file. You're right that the "don't initialize unwritten blocks" thing
has a more insidious problem of making it easy to unintentionally
expose such data just because you missed an error path (or the process
died without ever doing the proper over-write)..
It would be much better if we could somehow mitigate just _how_ easy
it is to screw up.
One way to do that would be to not just require that the user that
discards the initializing writes have read access to the underlying
device, but perhaps also have some strict requirement that you can
discard only if the file you are working with is legible only to you?
That would limit the damage, and keep the stale data "private" to the
user who is already able to read the raw data off the device. Sure,
you can then mark the file read-by-world by others later, but at that
point you're kind of *consciously* exposing that stale data (and at
that point, you have hopefully cleaned it all up and replaced the
stale data with real data).
But would that perhaps not be reasonable for the kind of use cases
that google has now?
Linus
[toc] | [prev] | [standalone]
Page 3 of 3 — ← Prev page 1 2 [3]
Back to top | Article view | linux.kernel
csiph-web