Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1320964 > unrolled thread
| Started by | Ross Zwisler <ross.zwisler@linux.intel.com> |
|---|---|
| First post | 2016-01-28 20:40 +0100 |
| Last post | 2016-02-02 11:40 +0100 |
| Articles | 20 on this page of 61 — 8 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
[PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-01-28 20:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-01-28 21:30 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Christoph Hellwig <hch@infradead.org> - 2016-01-28 22:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-01-29 19:30 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-01-30 00:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-01-30 01:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dave Chinner <david@fromorbit.com> - 2016-01-31 23:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Matthew Wilcox <willy@linux.intel.com> - 2016-01-30 06:30 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-01-30 07:10 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jared Hulbert <jaredeh@gmail.com> - 2016-01-30 08:10 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Matthew Wilcox <willy@linux.intel.com> - 2016-01-31 03:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <zwisler@gmail.com> - 2016-01-31 07:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Matthew Wilcox <willy@linux.intel.com> - 2016-01-31 12:00 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-01-31 17:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Matthew Wilcox <willy@linux.intel.com> - 2016-01-31 19:10 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-01-31 19:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Matthew Wilcox <willy@linux.intel.com> - 2016-01-31 19:30 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-01-31 20:00 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-01-31 21:00 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jan Kara <jack@suse.cz> - 2016-02-01 16:00 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Matthew Wilcox <willy@linux.intel.com> - 2016-02-01 21:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dave Chinner <david@fromorbit.com> - 2016-02-01 22:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jared Hulbert <jaredeh@gmail.com> - 2016-02-02 07:10 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-02-02 07:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jared Hulbert <jaredeh@gmail.com> - 2016-02-02 09:10 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-02-02 18:00 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jared Hulbert <jaredeh@gmail.com> - 2016-02-02 22:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Matthew Wilcox <willy@linux.intel.com> - 2016-02-03 01:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jared Hulbert <jaredeh@gmail.com> - 2016-02-03 02:30 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jan Kara <jack@suse.cz> - 2016-02-02 12:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-02-02 17:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jan Kara <jack@suse.cz> - 2016-02-02 17:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-02-02 18:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-02 18:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-02-02 18:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-02 19:30 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-02-02 18:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Matthew Wilcox <willy@linux.intel.com> - 2016-02-02 19:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dan Williams <dan.j.williams@intel.com> - 2016-02-02 20:10 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Matthew Wilcox <willy@linux.intel.com> - 2016-02-02 21:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jan Kara <jack@suse.cz> - 2016-02-03 12:10 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jan Kara <jack@suse.cz> - 2016-02-03 11:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-03 21:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jan Kara <jack@suse.cz> - 2016-02-04 10:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-05 00:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dave Chinner <david@fromorbit.com> - 2016-02-07 00:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-07 06:30 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-04 21:00 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jan Kara <jack@suse.cz> - 2016-02-04 21:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-04 23:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-05 23:30 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dave Chinner <david@fromorbit.com> - 2016-02-07 00:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-07 07:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jan Kara <jack@suse.cz> - 2016-02-08 14:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Christoph Hellwig <hch@infradead.org> - 2016-02-07 09:40 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-08 17:00 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-02 19:50 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-02 20:00 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-02-02 01:10 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Dave Chinner <david@fromorbit.com> - 2016-02-02 08:20 +0100
Re: [PATCH 2/2] dax: fix bdev NULL pointer dereferences Jan Kara <jack@suse.cz> - 2016-02-02 11:40 +0100
Page 1 of 4 [1] 2 3 4 Next page →
| From | Ross Zwisler <ross.zwisler@linux.intel.com> |
|---|---|
| Date | 2016-01-28 20:40 +0100 |
| Subject | [PATCH 2/2] dax: fix bdev NULL pointer dereferences |
| Message-ID | <qW0A9-87S-1@gated-at.bofh.it> |
There are a number of places in dax.c that look up the struct block_device
associated with an inode. Previously this was done by just using
inode->i_sb->s_bdev. This is correct for inodes that exist within the
filesystems supported by DAX (ext2, ext4 & XFS), but when running DAX
against raw block devices this value is NULL. This causes NULL pointer
dereferences when these block_device pointers are used.
Instead, for raw block devices we need to look up the struct block_device
using I_BDEV(). This patch fixes all the block_device lookups in dax.c so
that they work properly for both filesystems and raw block devices.
Signed-off-by: Ross Zwisler <ross.zwisler@linux.intel.com>
---
fs/dax.c | 15 +++++++++------
1 file changed, 9 insertions(+), 6 deletions(-)
diff --git a/fs/dax.c b/fs/dax.c
index 4fd6b0c..e60a5a7 100644
--- a/fs/dax.c
+++ b/fs/dax.c
@@ -32,6 +32,9 @@
#include <linux/pfn_t.h>
#include <linux/sizes.h>
+#define DAX_BDEV(inode) (S_ISBLK(inode->i_mode) ? I_BDEV(inode) \
+ : inode->i_sb->s_bdev)
+
static long dax_map_atomic(struct block_device *bdev, struct blk_dax_ctl *dax)
{
struct request_queue *q = bdev->bd_queue;
@@ -65,7 +68,7 @@ static void dax_unmap_atomic(struct block_device *bdev,
*/
int dax_clear_blocks(struct inode *inode, sector_t block, long _size)
{
- struct block_device *bdev = inode->i_sb->s_bdev;
+ struct block_device *bdev = DAX_BDEV(inode);
struct blk_dax_ctl dax = {
.sector = block << (inode->i_blkbits - 9),
.size = _size,
@@ -246,7 +249,7 @@ ssize_t dax_do_io(struct kiocb *iocb, struct inode *inode,
loff_t end = pos + iov_iter_count(iter);
memset(&bh, 0, sizeof(bh));
- bh.b_bdev = inode->i_sb->s_bdev;
+ bh.b_bdev = DAX_BDEV(inode);
if ((flags & DIO_LOCKING) && iov_iter_rw(iter) == READ) {
struct address_space *mapping = inode->i_mapping;
@@ -468,7 +471,7 @@ int dax_writeback_mapping_range(struct address_space *mapping, loff_t start,
loff_t end)
{
struct inode *inode = mapping->host;
- struct block_device *bdev = inode->i_sb->s_bdev;
+ struct block_device *bdev = DAX_BDEV(inode);
pgoff_t start_index, end_index, pmd_index;
pgoff_t indices[PAGEVEC_SIZE];
struct pagevec pvec;
@@ -608,7 +611,7 @@ int __dax_fault(struct vm_area_struct *vma, struct vm_fault *vmf,
memset(&bh, 0, sizeof(bh));
block = (sector_t)vmf->pgoff << (PAGE_SHIFT - blkbits);
- bh.b_bdev = inode->i_sb->s_bdev;
+ bh.b_bdev = DAX_BDEV(inode);
bh.b_size = PAGE_SIZE;
repeat:
@@ -827,7 +830,7 @@ int __dax_pmd_fault(struct vm_area_struct *vma, unsigned long address,
}
memset(&bh, 0, sizeof(bh));
- bh.b_bdev = inode->i_sb->s_bdev;
+ bh.b_bdev = DAX_BDEV(inode);
block = (sector_t)pgoff << (PAGE_SHIFT - blkbits);
bh.b_size = PMD_SIZE;
@@ -1080,7 +1083,7 @@ int dax_zero_page_range(struct inode *inode, loff_t from, unsigned length,
BUG_ON((offset + length) > PAGE_CACHE_SIZE);
memset(&bh, 0, sizeof(bh));
- bh.b_bdev = inode->i_sb->s_bdev;
+ bh.b_bdev = DAX_BDEV(inode);
bh.b_size = PAGE_CACHE_SIZE;
err = get_block(inode, index, &bh, 0);
if (err < 0)
--
2.5.0
[toc] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-01-28 21:30 +0100 |
| Message-ID | <qW1my-jc-9@gated-at.bofh.it> |
| In reply to | #1320964 |
On Thu, Jan 28, 2016 at 11:35 AM, Ross Zwisler <ross.zwisler@linux.intel.com> wrote: > There are a number of places in dax.c that look up the struct block_device > associated with an inode. Previously this was done by just using > inode->i_sb->s_bdev. This is correct for inodes that exist within the > filesystems supported by DAX (ext2, ext4 & XFS), but when running DAX > against raw block devices this value is NULL. This causes NULL pointer > dereferences when these block_device pointers are used. > > Instead, for raw block devices we need to look up the struct block_device > using I_BDEV(). This patch fixes all the block_device lookups in dax.c so > that they work properly for both filesystems and raw block devices. > > Signed-off-by: Ross Zwisler <ross.zwisler@linux.intel.com> It's a bit odd to check if it is a raw device inode in dax_clear_blocks() since there's no use case to clear blocks in that case, but I can't think of a better alternative. Acked-by: Dan Williams <dan.j.williams@intel.com>
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-01-28 22:40 +0100 |
| Message-ID | <qW2sh-Yt-1@gated-at.bofh.it> |
| In reply to | #1320964 |
On Thu, Jan 28, 2016 at 12:35:04PM -0700, Ross Zwisler wrote: > There are a number of places in dax.c that look up the struct block_device > associated with an inode. Previously this was done by just using > inode->i_sb->s_bdev. This is correct for inodes that exist within the > filesystems supported by DAX (ext2, ext4 & XFS), but when running DAX > against raw block devices this value is NULL. This causes NULL pointer > dereferences when these block_device pointers are used. It's also wrong for an XFS file system with a RT device.. > +#define DAX_BDEV(inode) (S_ISBLK(inode->i_mode) ? I_BDEV(inode) \ > + : inode->i_sb->s_bdev) .. but this isn't going to fix it. You must use a bdev returned by get_blocks or a similar file system method.
[toc] | [prev] | [next] | [standalone]
| From | Ross Zwisler <ross.zwisler@linux.intel.com> |
|---|---|
| Date | 2016-01-29 19:30 +0100 |
| Message-ID | <qWlXY-79P-15@gated-at.bofh.it> |
| In reply to | #1321075 |
On Thu, Jan 28, 2016 at 01:38:58PM -0800, Christoph Hellwig wrote: > On Thu, Jan 28, 2016 at 12:35:04PM -0700, Ross Zwisler wrote: > > There are a number of places in dax.c that look up the struct block_device > > associated with an inode. Previously this was done by just using > > inode->i_sb->s_bdev. This is correct for inodes that exist within the > > filesystems supported by DAX (ext2, ext4 & XFS), but when running DAX > > against raw block devices this value is NULL. This causes NULL pointer > > dereferences when these block_device pointers are used. > > It's also wrong for an XFS file system with a RT device.. > > > +#define DAX_BDEV(inode) (S_ISBLK(inode->i_mode) ? I_BDEV(inode) \ > > + : inode->i_sb->s_bdev) > > .. but this isn't going to fix it. You must use a bdev returned by > get_blocks or a similar file system method. I guess I need to go off and understand if we can have DAX mappings on such a device. If we can, we may have a problem - we can get the block_device from get_block() in I/O path and the various fault paths, but we don't have access to get_block() when flushing via dax_writeback_mapping_range(). We avoid needing it the normal case by storing the sector results from get_block() in the radix tree. /me is off to play with RT devices...
[toc] | [prev] | [next] | [standalone]
| From | Ross Zwisler <ross.zwisler@linux.intel.com> |
|---|---|
| Date | 2016-01-30 00:40 +0100 |
| Message-ID | <qWqNY-29Q-5@gated-at.bofh.it> |
| In reply to | #1321966 |
On Fri, Jan 29, 2016 at 11:28:15AM -0700, Ross Zwisler wrote: > On Thu, Jan 28, 2016 at 01:38:58PM -0800, Christoph Hellwig wrote: > > On Thu, Jan 28, 2016 at 12:35:04PM -0700, Ross Zwisler wrote: > > > There are a number of places in dax.c that look up the struct block_device > > > associated with an inode. Previously this was done by just using > > > inode->i_sb->s_bdev. This is correct for inodes that exist within the > > > filesystems supported by DAX (ext2, ext4 & XFS), but when running DAX > > > against raw block devices this value is NULL. This causes NULL pointer > > > dereferences when these block_device pointers are used. > > > > It's also wrong for an XFS file system with a RT device.. > > > > > +#define DAX_BDEV(inode) (S_ISBLK(inode->i_mode) ? I_BDEV(inode) \ > > > + : inode->i_sb->s_bdev) > > > > .. but this isn't going to fix it. You must use a bdev returned by > > get_blocks or a similar file system method. > > I guess I need to go off and understand if we can have DAX mappings on such a > device. If we can, we may have a problem - we can get the block_device from > get_block() in I/O path and the various fault paths, but we don't have access > to get_block() when flushing via dax_writeback_mapping_range(). We avoid > needing it the normal case by storing the sector results from get_block() in > the radix tree. > > /me is off to play with RT devices... Well, RT devices are completely broken as far as I can see. I've reported the breakage to the XFS list. Anything I do that triggers a RT block allocation in XFS causes a lockdep splat + a kernel BUG - I've tried regular pwrite(), xfs_rtcp and mmap() + write to address. Not a new bug either - happens just the same with v4.4. Happens with both PMEM and BRD, and has no relationship to whether I'm using DAX or not. Does it work for this patch to go in as-is since it fixes an immediate OOPS with raw block devices + DAX, and when RT devices are alive again I'll figure out how to make them work too?
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-01-30 01:20 +0100 |
| Message-ID | <qWrqG-2GF-5@gated-at.bofh.it> |
| In reply to | #1322176 |
On Fri, Jan 29, 2016 at 3:34 PM, Ross Zwisler <ross.zwisler@linux.intel.com> wrote: > On Fri, Jan 29, 2016 at 11:28:15AM -0700, Ross Zwisler wrote: >> On Thu, Jan 28, 2016 at 01:38:58PM -0800, Christoph Hellwig wrote: >> > On Thu, Jan 28, 2016 at 12:35:04PM -0700, Ross Zwisler wrote: >> > > There are a number of places in dax.c that look up the struct block_device >> > > associated with an inode. Previously this was done by just using >> > > inode->i_sb->s_bdev. This is correct for inodes that exist within the >> > > filesystems supported by DAX (ext2, ext4 & XFS), but when running DAX >> > > against raw block devices this value is NULL. This causes NULL pointer >> > > dereferences when these block_device pointers are used. >> > >> > It's also wrong for an XFS file system with a RT device.. >> > >> > > +#define DAX_BDEV(inode) (S_ISBLK(inode->i_mode) ? I_BDEV(inode) \ >> > > + : inode->i_sb->s_bdev) >> > >> > .. but this isn't going to fix it. You must use a bdev returned by >> > get_blocks or a similar file system method. >> >> I guess I need to go off and understand if we can have DAX mappings on such a >> device. If we can, we may have a problem - we can get the block_device from >> get_block() in I/O path and the various fault paths, but we don't have access >> to get_block() when flushing via dax_writeback_mapping_range(). We avoid >> needing it the normal case by storing the sector results from get_block() in >> the radix tree. >> >> /me is off to play with RT devices... > > Well, RT devices are completely broken as far as I can see. I've reported the > breakage to the XFS list. Anything I do that triggers a RT block allocation > in XFS causes a lockdep splat + a kernel BUG - I've tried regular pwrite(), > xfs_rtcp and mmap() + write to address. Not a new bug either - happens just > the same with v4.4. Happens with both PMEM and BRD, and has no relationship > to whether I'm using DAX or not. > > Does it work for this patch to go in as-is since it fixes an immediate OOPS > with raw block devices + DAX, and when RT devices are alive again I'll figure > out how to make them work too? Can we step back and be clear about which lookups should be coming from get_blocks(). Which ones are critical vs ones we just opportunistically lookup for a debug print. Right now xfs and ext4 are basically disagreeing on whether get_blocks() reliably sets ->bh_bdev, and checking for a raw block-device inode in dax_clear_blocks() does not make sense. So this all seems a bit confused.
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-01-31 23:50 +0100 |
| Message-ID | <qX8YG-1Bq-19@gated-at.bofh.it> |
| In reply to | #1322176 |
On Fri, Jan 29, 2016 at 04:34:30PM -0700, Ross Zwisler wrote: > On Fri, Jan 29, 2016 at 11:28:15AM -0700, Ross Zwisler wrote: > > On Thu, Jan 28, 2016 at 01:38:58PM -0800, Christoph Hellwig wrote: > > > On Thu, Jan 28, 2016 at 12:35:04PM -0700, Ross Zwisler wrote: > > > > There are a number of places in dax.c that look up the struct block_device > > > > associated with an inode. Previously this was done by just using > > > > inode->i_sb->s_bdev. This is correct for inodes that exist within the > > > > filesystems supported by DAX (ext2, ext4 & XFS), but when running DAX > > > > against raw block devices this value is NULL. This causes NULL pointer > > > > dereferences when these block_device pointers are used. > > > > > > It's also wrong for an XFS file system with a RT device.. > > > > > > > +#define DAX_BDEV(inode) (S_ISBLK(inode->i_mode) ? I_BDEV(inode) \ > > > > + : inode->i_sb->s_bdev) > > > > > > .. but this isn't going to fix it. You must use a bdev returned by > > > get_blocks or a similar file system method. > > > > I guess I need to go off and understand if we can have DAX mappings on such a > > device. If we can, we may have a problem - we can get the block_device from > > get_block() in I/O path and the various fault paths, but we don't have access > > to get_block() when flushing via dax_writeback_mapping_range(). We avoid > > needing it the normal case by storing the sector results from get_block() in > > the radix tree. > > > > /me is off to play with RT devices... > > Well, RT devices are completely broken as far as I can see. I've reported the > breakage to the XFS list. Anything I do that triggers a RT block allocation > in XFS causes a lockdep splat + a kernel BUG - I've tried regular pwrite(), Set CONFIG_XFS_DEBUG=n (assert failure that can be ignored causing the bug, and lockdep simply has an annotation problem) and it should work. Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Matthew Wilcox <willy@linux.intel.com> |
|---|---|
| Date | 2016-01-30 06:30 +0100 |
| Message-ID | <qWwgF-6ls-1@gated-at.bofh.it> |
| In reply to | #1321966 |
On Fri, Jan 29, 2016 at 11:28:15AM -0700, Ross Zwisler wrote: > I guess I need to go off and understand if we can have DAX mappings on such a > device. If we can, we may have a problem - we can get the block_device from > get_block() in I/O path and the various fault paths, but we don't have access > to get_block() when flushing via dax_writeback_mapping_range(). We avoid > needing it the normal case by storing the sector results from get_block() in > the radix tree. I think we're doing it wrong by storing the sector in the radix tree; we'd really need to store both the sector and the bdev which is too much data. If we store the PFN of the underlying page instead, we don't have this problem. Instead, we have a different problem; of the device going away under us. I'm trying to find the code which tears down PTEs when the device goes away, and I'm not seeing it. What do we do about user mappings of the device?
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-01-30 07:10 +0100 |
| Message-ID | <qWwTn-6Ud-1@gated-at.bofh.it> |
| In reply to | #1322260 |
On Fri, Jan 29, 2016 at 9:28 PM, Matthew Wilcox <willy@linux.intel.com> wrote: > On Fri, Jan 29, 2016 at 11:28:15AM -0700, Ross Zwisler wrote: >> I guess I need to go off and understand if we can have DAX mappings on such a >> device. If we can, we may have a problem - we can get the block_device from >> get_block() in I/O path and the various fault paths, but we don't have access >> to get_block() when flushing via dax_writeback_mapping_range(). We avoid >> needing it the normal case by storing the sector results from get_block() in >> the radix tree. > > I think we're doing it wrong by storing the sector in the radix tree; we'd > really need to store both the sector and the bdev which is too much data. > > If we store the PFN of the underlying page instead, we don't have this > problem. Instead, we have a different problem; of the device going > away under us. I'm trying to find the code which tears down PTEs when > the device goes away, and I'm not seeing it. What do we do about user > mappings of the device? > I deferred the dax tear down code until next cycle as Al rightly pointed out some needed re-works: https://lists.01.org/pipermail/linux-nvdimm/2016-January/003995.html
[toc] | [prev] | [next] | [standalone]
| From | Jared Hulbert <jaredeh@gmail.com> |
|---|---|
| Date | 2016-01-30 08:10 +0100 |
| Message-ID | <qWxPr-7BW-3@gated-at.bofh.it> |
| In reply to | #1322263 |
On Fri, Jan 29, 2016 at 10:01 PM, Dan Williams <dan.j.williams@intel.com> wrote: > On Fri, Jan 29, 2016 at 9:28 PM, Matthew Wilcox <willy@linux.intel.com> wrote: >> On Fri, Jan 29, 2016 at 11:28:15AM -0700, Ross Zwisler wrote: >>> I guess I need to go off and understand if we can have DAX mappings on such a >>> device. If we can, we may have a problem - we can get the block_device from >>> get_block() in I/O path and the various fault paths, but we don't have access >>> to get_block() when flushing via dax_writeback_mapping_range(). We avoid >>> needing it the normal case by storing the sector results from get_block() in >>> the radix tree. >> >> I think we're doing it wrong by storing the sector in the radix tree; we'd >> really need to store both the sector and the bdev which is too much data. >> >> If we store the PFN of the underlying page instead, we don't have this >> problem. Instead, we have a different problem; of the device going >> away under us. I'm trying to find the code which tears down PTEs when >> the device goes away, and I'm not seeing it. What do we do about user >> mappings of the device? >> > > I deferred the dax tear down code until next cycle as Al rightly > pointed out some needed re-works: > > https://lists.01.org/pipermail/linux-nvdimm/2016-January/003995.html If you store sectors in the radix and the device gets removed you still have to unmap user mappings of PFNs. So why is the device remove harder with the PFN vs bdev+sector radix entry? Either way you need a list of PFNs and their corresponding PTE's, right? And are we just talking graceful removal? Any plans for device failures?
[toc] | [prev] | [next] | [standalone]
| From | Matthew Wilcox <willy@linux.intel.com> |
|---|---|
| Date | 2016-01-31 03:40 +0100 |
| Message-ID | <qWQ5H-4S7-3@gated-at.bofh.it> |
| In reply to | #1322263 |
On Fri, Jan 29, 2016 at 10:01:13PM -0800, Dan Williams wrote:
> On Fri, Jan 29, 2016 at 9:28 PM, Matthew Wilcox <willy@linux.intel.com> wrote:
> > If we store the PFN of the underlying page instead, we don't have this
> > problem. Instead, we have a different problem; of the device going
> > away under us. I'm trying to find the code which tears down PTEs when
> > the device goes away, and I'm not seeing it. What do we do about user
> > mappings of the device?
>
> I deferred the dax tear down code until next cycle as Al rightly
> pointed out some needed re-works:
>
> https://lists.01.org/pipermail/linux-nvdimm/2016-January/003995.html
Thanks; I eventually found it in my email somewhere over the Pacific.
I did probably 70% of the work needed to switch the radix tree over to
storing PFNs instead of sectors. It seems viable, though it's a big
change from where we are today:
fs/dax.c | 415 +++++++++++++++++++++++----------------------
include/linux/dax.h | 3 +-
include/linux/pfn_t.h | 33 +++-
include/linux/radix-tree.h | 9 -
4 files changed, 236 insertions(+), 224 deletions(-)
I'll try and get that finished off this week.
One concrete and easily-separable piece is that dax_clear_blocks() has
the wrong signature. It currently takes an inode & block as parameters;
it has no way of finding out the correct block device. It's only two
callers are filesystems (ext2 and xfs). Those filesystems should be
passing the block_device instead of the inode. But without the inode,
we can't convert a block number to a sector number, so we also need
to pass the sector number, not the block number. It still has type
sector_t, annoyingly.
@@ -63,12 +238,11 @@ static void dax_unmap_atomic(struct block_device *bdev,
* and hence this means the stack from this point must follow GFP_NOFS
* semantics for all operations.
*/
-int dax_clear_blocks(struct inode *inode, sector_t block, long _size)
+int dax_clear_blocks(struct block_device *bdev, sector_t sector, long size)
{
- struct block_device *bdev = inode->i_sb->s_bdev;
struct blk_dax_ctl dax = {
- .sector = block << (inode->i_blkbits - 9),
- .size = _size,
+ .sector = sector,
+ .size = size,
};
might_sleep();
but I haven't looked at doing the conversion of xfs or ext2 to use that
new interface.
[toc] | [prev] | [next] | [standalone]
| From | Ross Zwisler <zwisler@gmail.com> |
|---|---|
| Date | 2016-01-31 07:20 +0100 |
| Message-ID | <qWTwB-7Ha-3@gated-at.bofh.it> |
| In reply to | #1322560 |
> On Jan 30, 2016, at 7:32 PM, Matthew Wilcox <willy@linux.intel.com> wrote:
>
>> On Fri, Jan 29, 2016 at 10:01:13PM -0800, Dan Williams wrote:
>>> On Fri, Jan 29, 2016 at 9:28 PM, Matthew Wilcox <willy@linux.intel.com> wrote:
>>> If we store the PFN of the underlying page instead, we don't have this
>>> problem. Instead, we have a different problem; of the device going
>>> away under us. I'm trying to find the code which tears down PTEs when
>>> the device goes away, and I'm not seeing it. What do we do about user
>>> mappings of the device?
>>
>> I deferred the dax tear down code until next cycle as Al rightly
>> pointed out some needed re-works:
>>
>> https://lists.01.org/pipermail/linux-nvdimm/2016-January/003995.html
>
> Thanks; I eventually found it in my email somewhere over the Pacific.
>
> I did probably 70% of the work needed to switch the radix tree over to
> storing PFNs instead of sectors. It seems viable, though it's a big
> change from where we are today:
At one point I had kaddrs in the radix tree, so I could just pull the addresses out
and flush them. That would save us a pfn -> kaddrs conversion before flush.
Is there a reason to store pnfs instead of kaddrs in the radix tree?
>
> fs/dax.c | 415 +++++++++++++++++++++++----------------------
> include/linux/dax.h | 3 +-
> include/linux/pfn_t.h | 33 +++-
> include/linux/radix-tree.h | 9 -
> 4 files changed, 236 insertions(+), 224 deletions(-)
>
> I'll try and get that finished off this week.
>
> One concrete and easily-separable piece is that dax_clear_blocks() has
> the wrong signature. It currently takes an inode & block as parameters;
> it has no way of finding out the correct block device. It's only two
> callers are filesystems (ext2 and xfs). Those filesystems should be
> passing the block_device instead of the inode. But without the inode,
> we can't convert a block number to a sector number, so we also need
> to pass the sector number, not the block number. It still has type
> sector_t, annoyingly.
>
> @@ -63,12 +238,11 @@ static void dax_unmap_atomic(struct block_device *bdev,
> * and hence this means the stack from this point must follow GFP_NOFS
> * semantics for all operations.
> */
> -int dax_clear_blocks(struct inode *inode, sector_t block, long _size)
> +int dax_clear_blocks(struct block_device *bdev, sector_t sector, long size)
> {
> - struct block_device *bdev = inode->i_sb->s_bdev;
> struct blk_dax_ctl dax = {
> - .sector = block << (inode->i_blkbits - 9),
> - .size = _size,
> + .sector = sector,
> + .size = size,
> };
>
> might_sleep();
>
> but I haven't looked at doing the conversion of xfs or ext2 to use that
> new interface.
> _______________________________________________
> Linux-nvdimm mailing list
> Linux-nvdimm@lists.01.org
> https://lists.01.org/mailman/listinfo/linux-nvdimm
[toc] | [prev] | [next] | [standalone]
| From | Matthew Wilcox <willy@linux.intel.com> |
|---|---|
| Date | 2016-01-31 12:00 +0100 |
| Message-ID | <qWXTA-24A-9@gated-at.bofh.it> |
| In reply to | #1322579 |
On Sat, Jan 30, 2016 at 11:12:12PM -0700, Ross Zwisler wrote:
> > I did probably 70% of the work needed to switch the radix tree over to
> > storing PFNs instead of sectors. It seems viable, though it's a big
> > change from where we are today:
>
> At one point I had kaddrs in the radix tree, so I could just pull the addresses out
> and flush them. That would save us a pfn -> kaddrs conversion before flush.
>
> Is there a reason to store pnfs instead of kaddrs in the radix tree?
Once ARM, MIPS and SPARC get supported, they're going to need temporary
kernel addresses assigned to PFNs rather than permanent ones. Also,
it'll be easier for teardown to delete PFNs associated with a particular
device than kaddrs associated with a particular device. And it lets
us support more persistent memory on a 32-bit machine (also on a 64-bit
machine, but that's mostly theoretical)
+/*
+ * DAX uses the 'exceptional' entries to store PFNs in the radix tree.
+ * Bit 0 is clear (the radix tree uses this for its own purposes). Bit
+ * 1 is set (to indicate an exceptional entry). Bits 2 & 3 are PFN_DEV
+ * and PFN_MAP. The top two bits denote the size of the entry (PTE, PMD,
+ * PUD, one reserved). That leaves us 26 bits on 32-bit systems and 58
+ * bits on 64-bit systems, able to address 256GB and 1024EB respectively.
+ */
It's also pretty cheap to look up the kaddr from the pfn, at least on
64-bit architectures without cache aliasing problems:
+static void *dax_map_pfn(pfn_t pfn, unsigned long index)
+{
+ preempt_disable();
+ pagefault_disable();
+ return pfn_to_kaddr(pfn_t_to_pfn(pfn));
+}
+
+static void dax_unmap_pfn(void *vaddr)
+{
+ pagefault_enable();
+ preempt_enable();
+}
32-bit x86 is going to want to do something similar to
iomap_atomic_prot_pfn(). ARM/SPARC/MIPS will want something in the
kmap_atomic family.
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-01-31 17:40 +0100 |
| Message-ID | <qX3cC-5Xg-11@gated-at.bofh.it> |
| In reply to | #1322602 |
On Sun, Jan 31, 2016 at 2:55 AM, Matthew Wilcox <willy@linux.intel.com> wrote:
> On Sat, Jan 30, 2016 at 11:12:12PM -0700, Ross Zwisler wrote:
>> > I did probably 70% of the work needed to switch the radix tree over to
>> > storing PFNs instead of sectors. It seems viable, though it's a big
>> > change from where we are today:
>>
>> At one point I had kaddrs in the radix tree, so I could just pull the addresses out
>> and flush them. That would save us a pfn -> kaddrs conversion before flush.
>>
>> Is there a reason to store pnfs instead of kaddrs in the radix tree?
>
> Once ARM, MIPS and SPARC get supported, they're going to need temporary
> kernel addresses assigned to PFNs rather than permanent ones. Also,
> it'll be easier for teardown to delete PFNs associated with a particular
> device than kaddrs associated with a particular device. And it lets
> us support more persistent memory on a 32-bit machine (also on a 64-bit
> machine, but that's mostly theoretical)
>
> +/*
> + * DAX uses the 'exceptional' entries to store PFNs in the radix tree.
> + * Bit 0 is clear (the radix tree uses this for its own purposes). Bit
> + * 1 is set (to indicate an exceptional entry). Bits 2 & 3 are PFN_DEV
> + * and PFN_MAP. The top two bits denote the size of the entry (PTE, PMD,
> + * PUD, one reserved). That leaves us 26 bits on 32-bit systems and 58
> + * bits on 64-bit systems, able to address 256GB and 1024EB respectively.
> + */
>
> It's also pretty cheap to look up the kaddr from the pfn, at least on
> 64-bit architectures without cache aliasing problems:
>
> +static void *dax_map_pfn(pfn_t pfn, unsigned long index)
> +{
> + preempt_disable();
> + pagefault_disable();
> + return pfn_to_kaddr(pfn_t_to_pfn(pfn));
pfn_to_kaddr() assumes persistent memory is direct mapped which is not
always the case.
[toc] | [prev] | [next] | [standalone]
| From | Matthew Wilcox <willy@linux.intel.com> |
|---|---|
| Date | 2016-01-31 19:10 +0100 |
| Message-ID | <qX4BJ-76A-43@gated-at.bofh.it> |
| In reply to | #1322680 |
On Sun, Jan 31, 2016 at 08:38:20AM -0800, Dan Williams wrote:
> On Sun, Jan 31, 2016 at 2:55 AM, Matthew Wilcox <willy@linux.intel.com> wrote:
> > On Sat, Jan 30, 2016 at 11:12:12PM -0700, Ross Zwisler wrote:
> >> Is there a reason to store pnfs instead of kaddrs in the radix tree?
> >
> > Once ARM, MIPS and SPARC get supported, they're going to need temporary
> > kernel addresses assigned to PFNs rather than permanent ones. Also,
> > it'll be easier for teardown to delete PFNs associated with a particular
> > device than kaddrs associated with a particular device. And it lets
> > us support more persistent memory on a 32-bit machine (also on a 64-bit
> > machine, but that's mostly theoretical)
> >
> > +/*
> > + * DAX uses the 'exceptional' entries to store PFNs in the radix tree.
> > + * Bit 0 is clear (the radix tree uses this for its own purposes). Bit
> > + * 1 is set (to indicate an exceptional entry). Bits 2 & 3 are PFN_DEV
> > + * and PFN_MAP. The top two bits denote the size of the entry (PTE, PMD,
> > + * PUD, one reserved). That leaves us 26 bits on 32-bit systems and 58
> > + * bits on 64-bit systems, able to address 256GB and 1024EB respectively.
> > + */
> >
> > It's also pretty cheap to look up the kaddr from the pfn, at least on
> > 64-bit architectures without cache aliasing problems:
> >
> > +static void *dax_map_pfn(pfn_t pfn, unsigned long index)
> > +{
> > + preempt_disable();
> > + pagefault_disable();
> > + return pfn_to_kaddr(pfn_t_to_pfn(pfn));
>
> pfn_to_kaddr() assumes persistent memory is direct mapped which is not
> always the case.
Yes. This is just the default implementation of dax_map_pfn() which works
for most situations. We can introduce more complex implementations of
dax_map_pfn() as necessary. You make another excellent point for why
we should store PFNs in the radix tree instead of kaddrs :-)
One option that I've been looking at (primarily for x86-32) is
having an rbtree of PFN ranges that drivers add to when they register
peristent memory. That would let us use the io_mapping_create_wc() /
io_mapping_map_atomic_wc() API. But having great support for persistent
memory with 32-bit x86 kernels is very very low on my priority list.
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-01-31 19:20 +0100 |
| Message-ID | <qX4Ln-7a2-1@gated-at.bofh.it> |
| In reply to | #1322697 |
On Sun, Jan 31, 2016 at 10:07 AM, Matthew Wilcox <willy@linux.intel.com> wrote:
> On Sun, Jan 31, 2016 at 08:38:20AM -0800, Dan Williams wrote:
>> On Sun, Jan 31, 2016 at 2:55 AM, Matthew Wilcox <willy@linux.intel.com> wrote:
>> > On Sat, Jan 30, 2016 at 11:12:12PM -0700, Ross Zwisler wrote:
>> >> Is there a reason to store pnfs instead of kaddrs in the radix tree?
>> >
>> > Once ARM, MIPS and SPARC get supported, they're going to need temporary
>> > kernel addresses assigned to PFNs rather than permanent ones. Also,
>> > it'll be easier for teardown to delete PFNs associated with a particular
>> > device than kaddrs associated with a particular device. And it lets
>> > us support more persistent memory on a 32-bit machine (also on a 64-bit
>> > machine, but that's mostly theoretical)
>> >
>> > +/*
>> > + * DAX uses the 'exceptional' entries to store PFNs in the radix tree.
>> > + * Bit 0 is clear (the radix tree uses this for its own purposes). Bit
>> > + * 1 is set (to indicate an exceptional entry). Bits 2 & 3 are PFN_DEV
>> > + * and PFN_MAP. The top two bits denote the size of the entry (PTE, PMD,
>> > + * PUD, one reserved). That leaves us 26 bits on 32-bit systems and 58
>> > + * bits on 64-bit systems, able to address 256GB and 1024EB respectively.
>> > + */
>> >
>> > It's also pretty cheap to look up the kaddr from the pfn, at least on
>> > 64-bit architectures without cache aliasing problems:
>> >
>> > +static void *dax_map_pfn(pfn_t pfn, unsigned long index)
>> > +{
>> > + preempt_disable();
>> > + pagefault_disable();
>> > + return pfn_to_kaddr(pfn_t_to_pfn(pfn));
>>
>> pfn_to_kaddr() assumes persistent memory is direct mapped which is not
>> always the case.
>
> Yes. This is just the default implementation of dax_map_pfn() which works
> for most situations. We can introduce more complex implementations of
> dax_map_pfn() as necessary. You make another excellent point for why
> we should store PFNs in the radix tree instead of kaddrs :-)
How much complexity do we want to add in support of an fsync/msync
mechanism that is not the recommended way to use DAX?
[toc] | [prev] | [next] | [standalone]
| From | Matthew Wilcox <willy@linux.intel.com> |
|---|---|
| Date | 2016-01-31 19:30 +0100 |
| Message-ID | <qX4V4-7h4-17@gated-at.bofh.it> |
| In reply to | #1322701 |
On Sun, Jan 31, 2016 at 10:18:46AM -0800, Dan Williams wrote: > On Sun, Jan 31, 2016 at 10:07 AM, Matthew Wilcox <willy@linux.intel.com> wrote: > > Yes. This is just the default implementation of dax_map_pfn() which works > > for most situations. We can introduce more complex implementations of > > dax_map_pfn() as necessary. You make another excellent point for why > > we should store PFNs in the radix tree instead of kaddrs :-) > > How much complexity do we want to add in support of an fsync/msync > mechanism that is not the recommended way to use DAX? It actually makes the dax_io path much, much simpler. And it's not primarily about fixing fsync/msync. It also makes the fault path cheaper in the case where we're refaulting a page that's already been faulted by another process (or was previously faulted by this process and now needs to be faulted at a different address). And it fixes the problem with filesystems that use multiple block_devices. It also makes DAX much less reliant on buffer heads, which is good for the problem that Jared raised where he doesn't have a block_device in an embedded system.
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-01-31 20:00 +0100 |
| Message-ID | <qX5o6-7sc-5@gated-at.bofh.it> |
| In reply to | #1322702 |
On Sun, Jan 31, 2016 at 10:27 AM, Matthew Wilcox <willy@linux.intel.com> wrote: > On Sun, Jan 31, 2016 at 10:18:46AM -0800, Dan Williams wrote: >> On Sun, Jan 31, 2016 at 10:07 AM, Matthew Wilcox <willy@linux.intel.com> wrote: >> > Yes. This is just the default implementation of dax_map_pfn() which works >> > for most situations. We can introduce more complex implementations of >> > dax_map_pfn() as necessary. You make another excellent point for why >> > we should store PFNs in the radix tree instead of kaddrs :-) >> >> How much complexity do we want to add in support of an fsync/msync >> mechanism that is not the recommended way to use DAX? > > It actually makes the dax_io path much, much simpler. And it's not > primarily about fixing fsync/msync. It also makes the fault path cheaper > in the case where we're refaulting a page that's already been faulted > by another process (or was previously faulted by this process and now > needs to be faulted at a different address). > > And it fixes the problem with filesystems that use multiple block_devices. > It also makes DAX much less reliant on buffer heads, which is good for > the problem that Jared raised where he doesn't have a block_device in > an embedded system. Oh I thought we were talking about what goes in the radix. Sure, de-emphasizing the usage of a block_device throughout the dax implementation is interesting. It also has some synergy with the LSF/MM topic I'm writing up "pmem as storage device vs pmem as memory".
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-01-31 21:00 +0100 |
| Message-ID | <qX6ka-899-9@gated-at.bofh.it> |
| In reply to | #1322697 |
On Sun, Jan 31, 2016 at 10:07 AM, Matthew Wilcox <willy@linux.intel.com> wrote: > One option that I've been looking at (primarily for x86-32) is > having an rbtree of PFN ranges that drivers add to when they register > peristent memory. On this specific point we do already have find_dev_pagemap().
[toc] | [prev] | [next] | [standalone]
| From | Jan Kara <jack@suse.cz> |
|---|---|
| Date | 2016-02-01 16:00 +0100 |
| Message-ID | <qXo7o-4jK-9@gated-at.bofh.it> |
| In reply to | #1322260 |
On Sat 30-01-16 00:28:33, Matthew Wilcox wrote: > On Fri, Jan 29, 2016 at 11:28:15AM -0700, Ross Zwisler wrote: > > I guess I need to go off and understand if we can have DAX mappings on such a > > device. If we can, we may have a problem - we can get the block_device from > > get_block() in I/O path and the various fault paths, but we don't have access > > to get_block() when flushing via dax_writeback_mapping_range(). We avoid > > needing it the normal case by storing the sector results from get_block() in > > the radix tree. > > I think we're doing it wrong by storing the sector in the radix tree; we'd > really need to store both the sector and the bdev which is too much data. > > If we store the PFN of the underlying page instead, we don't have this > problem. Instead, we have a different problem; of the device going > away under us. I'm trying to find the code which tears down PTEs when > the device goes away, and I'm not seeing it. What do we do about user > mappings of the device? So I don't have a strong opinion whether storing PFN or sector is better. Maybe PFN is somewhat more generic but OTOH turning DAX off for special cases like inodes on XFS RT devices would be IMHO fine. I'm somewhat concerned that there are several things in flight (page fault rework, invalidation on device removal, issues with DAX access to block devices Ross found) and this is IMHO the smallest trouble we have and changing this seems relatively invasive. So could we settle the fault code and similar stuff first and look into this somewhat later? Because frankly I have some trouble following how all the pieces are going to fit together and I'm afraid we'll introduce some non-trivial bugs when several fundamental things are in flux in parallel. Honza -- Jan Kara <jack@suse.com> SUSE Labs, CR
[toc] | [prev] | [next] | [standalone]
Page 1 of 4 [1] 2 3 4 Next page →
Back to top | Article view | linux.kernel
csiph-web