Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1471004 > unrolled thread
| Started by | Ross Zwisler <ross.zwisler@linux.intel.com> |
|---|---|
| First post | 2016-08-26 23:40 +0200 |
| Last post | 2016-08-30 09:30 +0200 |
| Articles | 5 — 4 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH v2 2/9] ext2: tell DAX the size of allocation holes Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-08-26 23:40 +0200
Re: [PATCH v2 2/9] ext2: tell DAX the size of allocation holes Dave Chinner <david@fromorbit.com> - 2016-08-29 02:50 +0200
Re: [PATCH v2 2/9] ext2: tell DAX the size of allocation holes Christoph Hellwig <hch@infradead.org> - 2016-08-29 10:10 +0200
Re: [PATCH v2 2/9] ext2: tell DAX the size of allocation holes Theodore Ts'o <tytso@mit.edu> - 2016-08-29 15:00 +0200
Re: [PATCH v2 2/9] ext2: tell DAX the size of allocation holes Christoph Hellwig <hch@infradead.org> - 2016-08-30 09:30 +0200
| From | Ross Zwisler <ross.zwisler@linux.intel.com> |
|---|---|
| Date | 2016-08-26 23:40 +0200 |
| Subject | Re: [PATCH v2 2/9] ext2: tell DAX the size of allocation holes |
| Message-ID | <sax0Z-41h-11@gated-at.bofh.it> |
On Thu, Aug 25, 2016 at 12:57:28AM -0700, Christoph Hellwig wrote: > Hi Ross, > > can you take at my (fully working, but not fully cleaned up) version > of the iomap based DAX code here: > > http://git.infradead.org/users/hch/vfs.git/shortlog/refs/heads/iomap-dax > > By using iomap we don't even have the size hole problem and totally > get out of the reverse-engineer what buffer_heads are trying to tell > us business. It also gets rid of the other warts of the DAX path > due to pretending to be like direct I/O, so this might be a better > way forward also for ext2/4. In general I agree that the usage of struct iomap seems more straightforward than the old way of using struct buffer_head + get_block_t. I really don't think we want to have two competing DAX I/O and fault paths, though, which I assume everyone else agrees with as well. These changes don't remove the things in XFS needed by the old I/O and fault paths (e.g. xfs_get_blocks_direct() is still there an unchanged). Is the correct way forward to get buy-in from ext2/ext4 so that they also move to supporting an iomap based I/O path (xfs_file_iomap_begin(), xfs_iomap_write_direct(), etc?). That would allow us to have parallel I/O and fault paths for a while, then remove the old buffer_head based versions when the three supported filesystems have moved to iomap. If ext2 and ext4 don't choose to move to iomap, though, I don't think we want to have a separate I/O & fault path for iomap/XFS. That seems too painful, and the old buffer_head version should continue to work, ugly as it may be. Assuming we can get buy-in from ext4/ext2, I can work on a PMD version of the iomap based fault path that is equivalent to the buffer_head based one I sent out in my series, and we can all eventually move to that. A few comments/questions on the implementation: 1) In your mail above you say "It also gets rid of the other warts of the DAX path due to pretending to be like direct I/O". I assume by this you mean the code in dax_do_io() around DIO_LOCKING, inode_dio_begin(), etc? Perhaps there are other things as well in XFS, but this is what I see in the DAX code. If so, yep, this seems like a win. I don't understand how DIO_LOCKING is relevant to the DAX I/O path, as we never mix buffered and direct access. The comment in dax_do_io() for the inode_dio_begin() call says that it prevents the I/O from races with truncate. Am I correct that we now get this protection via the xfs_rw_ilock()/xfs_rw_iunlock() calls in xfs_file_dax_write()? 2) Just a nit, I noticed that you used "~(PAGE_SIZE - 1)" in several places in iomap_dax_actor() and iomap_dax_fault() instead of PAGE_MASK. Was this intentional? 3) It's kind of weird having iomap_dax_fault() in fs/dax.c but having iomap_dax_actor() and iomap_dax_rw() in fs/iomap.c? I'm guessing the latter is placed where it is because it uses iomap_apply(), which is local to fs/iomap.c? Anyway, it would be nice if we could keep them together, if possible. 4) In iomap_dax_actor() you do this check: WARN_ON_ONCE(iomap->type != IOMAP_MAPPED); If we hit this we should bail with -EIO, yea? Otherwise we could write to unmapped space or something horrible. 5) In iomap_dax_fault, I think the "I/O beyond the end of the file" check might have been broken. Take for example an I/O to the second page of a file, where the file has size one page. So: vmf->pgoff = 1 i_size_read(inode) = 4096 Here's the old code in dax_fault(): size = (i_size_read(inode) + PAGE_SIZE - 1) >> PAGE_SHIFT; if (vmf->pgoff >= size) return VM_FAULT_SIGBUS; size = (4096 + 4096 - 1) >> PAGE_SHIFT = 1 vmf->pgoff is 1 and size is 1, so we return SIGBUS Here's the new code: if (pos >= i_size_read(inode) + PAGE_SIZE - 1) return VM_FAULT_SIGBUS; pos = vmf->pgoff << PAGE_SHIFT = 4096 i_size_read(inode) + PAGE_SIZE - 1 = 8193 so, 'pos' isn't >= where we calculate the end of the file to be, so we do I/O Basically the old check did the "+ PAGE_SIZE - 1" so that the >> PAGE_SHIFT was sure to round up to the next full page. You don't need this with your current logic, so I think the test should just be: if (pos >= i_size_read(inode)) return VM_FAULT_SIGBUS; Right? 6) Regarding the "we don't even have the size hole problem" comment in your mail, the current PMD logic requires us to know the size of the hole. This is important so that we can fault in a huge zero page if we have a 2 MiB hole. It's fine if that 2 MiB page then gets fragmented into 4k DAX allocations when we start to do writes, but the path the other way doesn't work. If we don't know the size of holes then we can't fault in a 2 MiB zero page, so we'll use 4k zero pages to satisfy reads. This means that if later we want to fault in a 2MiB DAX allocation, we don't have a single entry that we can use to lock the entire 2MiB range while we clean the radix tree an unmap the range from all the user processes. With the current PMD logic this will mean that if someone does a 4k read that faults in a 4k zero page, we will only use 4k faults for that range and won't use PMDs. The current XFS code in the v4.8 tree tells me the size of the hole, and I think we need to keep this functionality.
[toc] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-08-29 02:50 +0200 |
| Message-ID | <sbiVY-kL-3@gated-at.bofh.it> |
| In reply to | #1471004 |
On Fri, Aug 26, 2016 at 03:29:34PM -0600, Ross Zwisler wrote: > On Thu, Aug 25, 2016 at 12:57:28AM -0700, Christoph Hellwig wrote: > > Hi Ross, > > > > can you take at my (fully working, but not fully cleaned up) version > > of the iomap based DAX code here: > > > > http://git.infradead.org/users/hch/vfs.git/shortlog/refs/heads/iomap-dax > > > > By using iomap we don't even have the size hole problem and totally > > get out of the reverse-engineer what buffer_heads are trying to tell > > us business. It also gets rid of the other warts of the DAX path > > due to pretending to be like direct I/O, so this might be a better > > way forward also for ext2/4. > > In general I agree that the usage of struct iomap seems more straightforward > than the old way of using struct buffer_head + get_block_t. I really don't > think we want to have two competing DAX I/O and fault paths, though, which I > assume everyone else agrees with as well. We'll be moving XFS this way, regardless of whether the generic DAX code goes that way or not. iomap is a much cleaner, more efficient interface than get_blocks via bufferheads. We are slowly removing bufferheads from XFS so anything that uses them or depends on them that Xfs requires is going to have an iomap-based variant written for it. Christoph is doing the hard yards to make iomap a VFS level interface because that's a) the most efficient way to implement it, and b) it's the right place for IO path extent mapping abstractions. So there will be a iomap path for DAX, just like there will be a iomap path for direct IO, regardless of what other filesystems implement. i.e. other filesystems can move to the more efficient iomap infrastructure if they want, but we can't force them to do so. As such, the generic DAX path can either remain as it is, or we can move to iomap and use wrappers for converting get_block() + bufferehead to iomaps on non-iomap filesystems. (i.e. similar to the existing iomap_to_bh() for allowing iomap lookups to be used to replace bufferheads returned by get_block().) I'd much prefer we move DAX to iomaps before there is wider uptake of it in other filesystems - I've been saying we should use iomaps for DAX right from the start. Now we have the iomap infrastructure in place we should jump straight to it. If we have to drag ext4 kicking and screaming into the 1990s to get there then so be it - it won't be the first time... > These changes don't remove the things in XFS needed by the old I/O and fault > paths (e.g. xfs_get_blocks_direct() is still there an unchanged). Is the Yes, they'll remain until their functionality has been replaced by iomap functions. e.g. xfs_get_blocks_direct() can't be removed until the direct IO path has an iomap interface. .... > 6) Regarding the "we don't even have the size hole problem" comment in your > mail, the current PMD logic requires us to know the size of the hole. This .... > The current XFS code in the v4.8 tree tells me the size of the hole, and I > think we need to keep this functionality. IOMAP_HOLE extents. It's a requirement of the iomap infrastructure that the filesystem reports hole extents in full for the range being mapped. Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-08-29 10:10 +0200 |
| Message-ID | <sbpNM-4Om-31@gated-at.bofh.it> |
| In reply to | #1471004 |
On Fri, Aug 26, 2016 at 03:29:34PM -0600, Ross Zwisler wrote: > These changes don't remove the things in XFS needed by the old I/O and fault > paths (e.g. xfs_get_blocks_direct() is still there an unchanged). Is the > correct way forward to get buy-in from ext2/ext4 so that they also move to > supporting an iomap based I/O path (xfs_file_iomap_begin(), > xfs_iomap_write_direct(), etc?). That would allow us to have parallel I/O and > fault paths for a while, then remove the old buffer_head based versions when > the three supported filesystems have moved to iomap. > > If ext2 and ext4 don't choose to move to iomap, though, I don't think we want > to have a separate I/O & fault path for iomap/XFS. That seems too painful, > and the old buffer_head version should continue to work, ugly as it may be. We're going to move forward killing buffer_heads in XFS. I think ext4 would dramatically benefit from this a well, as would ext2 (although I think all that DAX work in ext2 is a horrible idea to start with). If I don't get buy-in for the iomap DAX work in the dax code we'll just have to keep it separate. That buffer_head mess just isn't maintainable the long run. > 1) In your mail above you say "It also gets rid of the other warts of the DAX > path due to pretending to be like direct I/O". I assume by this you mean > the code in dax_do_io() around DIO_LOCKING, inode_dio_begin(), etc? Yes. > Perhaps there are other things as well in XFS, but this is what I see in > the DAX code. If so, yep, this seems like a win. I don't understand how > DIO_LOCKING is relevant to the DAX I/O path, as we never mix buffered and > direct access. It's related to doing stupid copy and paste from direct I/O in the DAX code. > The comment in dax_do_io() for the inode_dio_begin() call says that it > prevents the I/O from races with truncate. Am I correct that we now get > this protection via the xfs_rw_ilock()/xfs_rw_iunlock() calls in > xfs_file_dax_write()? Yes, XFS always has a lock over reads that serializes with truncate. Currenrly it's the XFS i_iolock, but I'll remove that soon and use the VFS i_rwsem instead. For ext2/4 we could go straight to i_rwsem in shared mode. > 2) Just a nit, I noticed that you used "~(PAGE_SIZE - 1)" in several places in > iomap_dax_actor() and iomap_dax_fault() instead of PAGE_MASK. Was this > intentional? Mostly because that's how I think. I'm fine using PAGE_MASK, though. > 3) It's kind of weird having iomap_dax_fault() in fs/dax.c but having > iomap_dax_actor() and iomap_dax_rw() in fs/iomap.c? I'm guessing the > latter is placed where it is because it uses iomap_apply(), which is local > to fs/iomap.c? Anyway, it would be nice if we could keep them together, if > possible. It's still work in progress and could use a few cleanups. > > 4) In iomap_dax_actor() you do this check: > > WARN_ON_ONCE(iomap->type != IOMAP_MAPPED); > > If we hit this we should bail with -EIO, yea? Otherwise we could write to > unmapped space or something horrible. Fine with me. > 5) In iomap_dax_fault, I think the "I/O beyond the end of the file" check > might have been broken. Take for example an I/O to the second page of a > file, where the file has size one page. So: sure, I can fix this up. > 6) Regarding the "we don't even have the size hole problem" comment in your > mail, the current PMD logic requires us to know the size of the hole. And a big part of the iomap interface is proper reporting of holes.
[toc] | [prev] | [next] | [standalone]
| From | Theodore Ts'o <tytso@mit.edu> |
|---|---|
| Date | 2016-08-29 15:00 +0200 |
| Message-ID | <sbukq-7qb-19@gated-at.bofh.it> |
| In reply to | #1471627 |
On Mon, Aug 29, 2016 at 12:41:16AM -0700, Christoph Hellwig wrote: > > We're going to move forward killing buffer_heads in XFS. I think ext4 > would dramatically benefit from this a well, as would ext2 (although I > think all that DAX work in ext2 is a horrible idea to start with). It's been on my todo list. The only reason why I haven't done it yet is because I knew you were working on a solution, and I didn't want to do things one way for buffered I/O, and a different way for Direct I/O, and disentangling the DIO code and the different assumptions of how different file systems interact with the DIO code is a *mess*. It may have gotten better more recently, but a few years ago I took a look at it and backed slowly away..... - Ted
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-08-30 09:30 +0200 |
| Message-ID | <sbLEB-1Kt-5@gated-at.bofh.it> |
| In reply to | #1471815 |
On Mon, Aug 29, 2016 at 08:57:41AM -0400, Theodore Ts'o wrote: > It's been on my todo list. The only reason why I haven't done it yet > is because I knew you were working on a solution, and I didn't want to > do things one way for buffered I/O, and a different way for Direct > I/O, and disentangling the DIO code and the different assumptions of > how different file systems interact with the DIO code is a *mess*. It is. I have an almost working iomap direct I/O implementation, which simplifies a lot of this. Partially because it expects sane locking (a shared lock should be held over read, fortunately we now have i_rwsem for that) and allows less opt-in/out behavior, and partially just because we have so much better infrastructure now (iomap, bio chaining, iov_iters, ...). The iomap version is still work in progress, but I'm about to post a cut down version for block devices which reduces I/O latency by 20%.
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web