Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1709309 > unrolled thread
| Started by | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| First post | 2017-08-11 08:50 +0200 |
| Last post | 2017-08-11 08:50 +0200 |
| Articles | 3 — 1 participant |
Back to article view | Back to linux.kernel
[PATCH v3 0/6] fs, xfs: block map immutable files Dan Williams <dan.j.williams@intel.com> - 2017-08-11 08:50 +0200
[PATCH v3 3/6] fs, xfs: introduce FALLOC_FL_UNSEAL_BLOCK_MAP Dan Williams <dan.j.williams@intel.com> - 2017-08-11 08:50 +0200
[PATCH v3 4/6] xfs: introduce XFS_DIFLAG2_IOMAP_IMMUTABLE Dan Williams <dan.j.williams@intel.com> - 2017-08-11 08:50 +0200
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-08-11 08:50 +0200 |
| Subject | [PATCH v3 0/6] fs, xfs: block map immutable files |
| Message-ID | <udbVE-7y-3@gated-at.bofh.it> |
Changes since v2 [1]:
* Rather than have an IS_IOMAP_IMMUTABLE() check in
xfs_alloc_file_space(), place one centrally in xfs_bmapi_write() to
catch all attempts to write the block allocation map. (Dave)
* Make sealing an already sealed file, or unsealing an already unsealed
file return success (Darrick)
* Set S_IOMAP_IMMUTABLE along with the transaction that sets
XFS_DIFLAG2_IOMAP_IMMUTABLE (Darrick)
* Round the range of the allocation and extent conversion performed by
FALLOC_FL_SEAL_BLOCK_MAP up to the filesystem block size.
* Add a proof-of-concept patch for the use of immutable files with swap.
[1]: https://lkml.org/lkml/2017/8/3/996
---
The ability to make the physical block-allocation map of a file
immutable is a powerful mechanism that allows userspace to have
predictable dax-fault latencies, flush dax mappings to persistent memory
without a syscall, and otherwise enable access to storage directly
without ongoing mediation from the filesystem.
This last aspect of direct storage addressability has been called a
"horrible abuse" [2], but the reality is quite the reverse. Enabling
files to be block-map immutable allows applications that would otherwise
need to rely on dangerous raw device access to instead use a filesystem.
Security, naming, re-provisioning capacity between usages are all better
supported with safe semantics in a filesystem compared to a device file.
It is time to "give up the idea that only the filesystem can access the
storage underlying the filesystem" [3] to enable a better / safer
alternative to using a raw device for userpace block servers, dax
hypervisors, and peer-to-peer transfers to name a few use cases.
[2]: https://lkml.org/lkml/2017/8/5/56
[3]: https://lkml.org/lkml/2017/8/6/299
---
Dan Williams (6):
fs, xfs: introduce S_IOMAP_IMMUTABLE
fs, xfs: introduce FALLOC_FL_SEAL_BLOCK_MAP
fs, xfs: introduce FALLOC_FL_UNSEAL_BLOCK_MAP
xfs: introduce XFS_DIFLAG2_IOMAP_IMMUTABLE
xfs: toggle XFS_DIFLAG2_IOMAP_IMMUTABLE in response to fallocate
mm, xfs: protect swapfile contents with immutable + unwritten extents
fs/attr.c | 10 +++
fs/nfs/file.c | 7 ++
fs/open.c | 24 +++++++
fs/read_write.c | 3 +
fs/xfs/libxfs/xfs_bmap.c | 6 ++
fs/xfs/libxfs/xfs_bmap.h | 12 +++-
fs/xfs/libxfs/xfs_format.h | 5 +-
fs/xfs/xfs_aops.c | 54 ++++++++++++++++
fs/xfs/xfs_bmap_util.c | 142 +++++++++++++++++++++++++++++++++++++++++++
fs/xfs/xfs_bmap_util.h | 5 ++
fs/xfs/xfs_file.c | 16 ++++-
fs/xfs/xfs_inode.c | 2 +
fs/xfs/xfs_ioctl.c | 7 ++
fs/xfs/xfs_iops.c | 8 ++
include/linux/falloc.h | 4 +
include/linux/fs.h | 2 +
include/uapi/linux/falloc.h | 18 +++++
include/uapi/linux/fs.h | 1
mm/filemap.c | 5 ++
mm/page_io.c | 1
mm/swapfile.c | 20 ++----
21 files changed, 328 insertions(+), 24 deletions(-)
[toc] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-08-11 08:50 +0200 |
| Subject | [PATCH v3 3/6] fs, xfs: introduce FALLOC_FL_UNSEAL_BLOCK_MAP |
| Message-ID | <udbVE-7y-19@gated-at.bofh.it> |
| In reply to | #1709309 |
Provide an explicit fallocate operation type for clearing the
S_IOMAP_IMMUTABLE flag. Like the enable case it requires CAP_IMMUTABLE
and it can only be performed while no process has the file mapped.
Cc: Jan Kara <jack@suse.cz>
Cc: Jeff Moyer <jmoyer@redhat.com>
Cc: Christoph Hellwig <hch@lst.de>
Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
Cc: Alexander Viro <viro@zeniv.linux.org.uk>
Cc: "Darrick J. Wong" <darrick.wong@oracle.com>
Suggested-by: Dave Chinner <david@fromorbit.com>
Signed-off-by: Dan Williams <dan.j.williams@intel.com>
---
fs/open.c | 20 ++++++++++++------
fs/xfs/xfs_bmap_util.c | 47 +++++++++++++++++++++++++++++++++++++++++++
fs/xfs/xfs_bmap_util.h | 3 +++
fs/xfs/xfs_file.c | 4 +++-
include/linux/falloc.h | 3 ++-
include/uapi/linux/falloc.h | 1 +
6 files changed, 69 insertions(+), 9 deletions(-)
diff --git a/fs/open.c b/fs/open.c
index 76f57f7465c4..3075599f1c55 100644
--- a/fs/open.c
+++ b/fs/open.c
@@ -274,13 +274,17 @@ int vfs_fallocate(struct file *file, int mode, loff_t offset, loff_t len)
return -EINVAL;
/*
- * Seal block map operation should only be used exclusively, and
- * with the IMMUTABLE capability.
+ * Seal/unseal block map operations should only be used
+ * exclusively, and with the IMMUTABLE capability.
*/
- if (mode & FALLOC_FL_SEAL_BLOCK_MAP) {
+ if (mode & (FALLOC_FL_SEAL_BLOCK_MAP | FALLOC_FL_UNSEAL_BLOCK_MAP)) {
if (!capable(CAP_LINUX_IMMUTABLE))
return -EPERM;
- if (mode & ~FALLOC_FL_SEAL_BLOCK_MAP)
+ if (mode == (FALLOC_FL_SEAL_BLOCK_MAP
+ | FALLOC_FL_UNSEAL_BLOCK_MAP))
+ return -EINVAL;
+ if (mode & ~(FALLOC_FL_SEAL_BLOCK_MAP
+ | FALLOC_FL_UNSEAL_BLOCK_MAP))
return -EINVAL;
}
@@ -303,10 +307,12 @@ int vfs_fallocate(struct file *file, int mode, loff_t offset, loff_t len)
return -ETXTBSY;
/*
- * We cannot allow any allocation changes on an iomap immutable file,
- * but we can allow the fs to validate if this request is redundant.
+ * We cannot allow any allocation changes on an iomap immutable
+ * file, but we can allow the fs to validate if this request is
+ * redundant, or unseal the block map.
*/
- if (IS_IOMAP_IMMUTABLE(inode) && !(mode & FALLOC_FL_SEAL_BLOCK_MAP))
+ if (IS_IOMAP_IMMUTABLE(inode) && !(mode & (FALLOC_FL_SEAL_BLOCK_MAP
+ | FALLOC_FL_UNSEAL_BLOCK_MAP)))
return -ETXTBSY;
/*
diff --git a/fs/xfs/xfs_bmap_util.c b/fs/xfs/xfs_bmap_util.c
index 2ac8f4ed5723..888bae801961 100644
--- a/fs/xfs/xfs_bmap_util.c
+++ b/fs/xfs/xfs_bmap_util.c
@@ -1462,6 +1462,53 @@ xfs_seal_file_space(
return error;
}
+int
+xfs_unseal_file_space(
+ struct xfs_inode *ip,
+ xfs_off_t offset,
+ xfs_off_t len)
+{
+ struct inode *inode = VFS_I(ip);
+ struct address_space *mapping = inode->i_mapping;
+ int error;
+
+ ASSERT(xfs_isilocked(ip, XFS_IOLOCK_EXCL | XFS_MMAPLOCK_EXCL));
+
+ if (offset)
+ return -EINVAL;
+
+ xfs_ilock(ip, XFS_ILOCK_EXCL);
+ /*
+ * It does not make sense to unseal less than the full range of
+ * the file.
+ */
+ error = -EINVAL;
+ if (len != i_size_read(inode))
+ goto out_unlock;
+
+ /* Are we already unsealed? */
+ error = 0;
+ if (!IS_IOMAP_IMMUTABLE(inode))
+ goto out_unlock;
+
+ /*
+ * Provide safety against one thread changing the policy of not
+ * requiring fsync/msync (for block allocations) behind another
+ * thread's back.
+ */
+ error = -EBUSY;
+ if (mapping_mapped(mapping))
+ goto out_unlock;
+
+ error = 0;
+ inode->i_flags &= ~S_IOMAP_IMMUTABLE;
+
+out_unlock:
+ xfs_iunlock(ip, XFS_ILOCK_EXCL);
+
+ return error;
+}
+
/*
* @next_fsb will keep track of the extent currently undergoing shift.
* @stop_fsb will keep track of the extent at which we have to stop.
diff --git a/fs/xfs/xfs_bmap_util.h b/fs/xfs/xfs_bmap_util.h
index 5115a32a2483..b64653a75942 100644
--- a/fs/xfs/xfs_bmap_util.h
+++ b/fs/xfs/xfs_bmap_util.h
@@ -62,6 +62,9 @@ int xfs_insert_file_space(struct xfs_inode *, xfs_off_t offset,
xfs_off_t len);
int xfs_seal_file_space(struct xfs_inode *, xfs_off_t offset,
xfs_off_t len);
+int xfs_unseal_file_space(struct xfs_inode *, xfs_off_t offset,
+ xfs_off_t len);
+
/* EOF block manipulation functions */
bool xfs_can_free_eofblocks(struct xfs_inode *ip, bool force);
diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c
index e21121530a90..833f77700be2 100644
--- a/fs/xfs/xfs_file.c
+++ b/fs/xfs/xfs_file.c
@@ -740,7 +740,7 @@ xfs_file_write_iter(
(FALLOC_FL_KEEP_SIZE | FALLOC_FL_PUNCH_HOLE | \
FALLOC_FL_COLLAPSE_RANGE | FALLOC_FL_ZERO_RANGE | \
FALLOC_FL_INSERT_RANGE | FALLOC_FL_UNSHARE_RANGE | \
- FALLOC_FL_SEAL_BLOCK_MAP)
+ FALLOC_FL_SEAL_BLOCK_MAP | FALLOC_FL_UNSEAL_BLOCK_MAP)
STATIC long
xfs_file_fallocate(
@@ -840,6 +840,8 @@ xfs_file_fallocate(
XFS_BMAPI_PREALLOC);
} else if (mode & FALLOC_FL_SEAL_BLOCK_MAP) {
error = xfs_seal_file_space(ip, offset, len);
+ } else if (mode & FALLOC_FL_UNSEAL_BLOCK_MAP) {
+ error = xfs_unseal_file_space(ip, offset, len);
} else
error = xfs_alloc_file_space(ip, offset, len,
XFS_BMAPI_PREALLOC);
diff --git a/include/linux/falloc.h b/include/linux/falloc.h
index 48546c6fbec7..b22c1368ed1e 100644
--- a/include/linux/falloc.h
+++ b/include/linux/falloc.h
@@ -27,6 +27,7 @@ struct space_resv {
FALLOC_FL_ZERO_RANGE | \
FALLOC_FL_INSERT_RANGE | \
FALLOC_FL_UNSHARE_RANGE | \
- FALLOC_FL_SEAL_BLOCK_MAP)
+ FALLOC_FL_SEAL_BLOCK_MAP | \
+ FALLOC_FL_UNSEAL_BLOCK_MAP)
#endif /* _FALLOC_H_ */
diff --git a/include/uapi/linux/falloc.h b/include/uapi/linux/falloc.h
index e3867cfe31d5..5509e6216448 100644
--- a/include/uapi/linux/falloc.h
+++ b/include/uapi/linux/falloc.h
@@ -93,4 +93,5 @@
* with the punch, zero, collapse, or insert range modes.
*/
#define FALLOC_FL_SEAL_BLOCK_MAP 0x080
+#define FALLOC_FL_UNSEAL_BLOCK_MAP 0x100
#endif /* _UAPI_FALLOC_H_ */
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2017-08-11 08:50 +0200 |
| Subject | [PATCH v3 4/6] xfs: introduce XFS_DIFLAG2_IOMAP_IMMUTABLE |
| Message-ID | <udbVE-7y-21@gated-at.bofh.it> |
| In reply to | #1709309 |
Add an on-disk inode flag to record the state of the S_IOMAP_IMMUTABLE
in-memory vfs inode flags. This allows the protections against reflink
and hole punch to be automatically restored on a sub-sequent boot when
the in-memory inode is established.
The FS_XFLAG_IOMAP_IMMUTABLE is introduced to allow xfs_io to read the
state of the flag, but toggling the flag requires going through
fallocate(FALLOC_FL_[UN]SEAL_BLOCK_MAP). Support for toggling this
on-disk state is saved for a later patch.
Cc: Jan Kara <jack@suse.cz>
Cc: Jeff Moyer <jmoyer@redhat.com>
Cc: Christoph Hellwig <hch@lst.de>
Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
Suggested-by: Dave Chinner <david@fromorbit.com>
Suggested-by: "Darrick J. Wong" <darrick.wong@oracle.com>
Signed-off-by: Dan Williams <dan.j.williams@intel.com>
---
fs/xfs/libxfs/xfs_format.h | 5 ++++-
fs/xfs/xfs_inode.c | 2 ++
fs/xfs/xfs_ioctl.c | 1 +
fs/xfs/xfs_iops.c | 8 +++++---
include/uapi/linux/fs.h | 1 +
5 files changed, 13 insertions(+), 4 deletions(-)
diff --git a/fs/xfs/libxfs/xfs_format.h b/fs/xfs/libxfs/xfs_format.h
index d4d9bef20c3a..9e720e55776b 100644
--- a/fs/xfs/libxfs/xfs_format.h
+++ b/fs/xfs/libxfs/xfs_format.h
@@ -1063,12 +1063,15 @@ static inline void xfs_dinode_put_rdev(struct xfs_dinode *dip, xfs_dev_t rdev)
#define XFS_DIFLAG2_DAX_BIT 0 /* use DAX for this inode */
#define XFS_DIFLAG2_REFLINK_BIT 1 /* file's blocks may be shared */
#define XFS_DIFLAG2_COWEXTSIZE_BIT 2 /* copy on write extent size hint */
+#define XFS_DIFLAG2_IOMAP_IMMUTABLE_BIT 3 /* set S_IOMAP_IMMUTABLE for this inode */
#define XFS_DIFLAG2_DAX (1 << XFS_DIFLAG2_DAX_BIT)
#define XFS_DIFLAG2_REFLINK (1 << XFS_DIFLAG2_REFLINK_BIT)
#define XFS_DIFLAG2_COWEXTSIZE (1 << XFS_DIFLAG2_COWEXTSIZE_BIT)
+#define XFS_DIFLAG2_IOMAP_IMMUTABLE (1 << XFS_DIFLAG2_IOMAP_IMMUTABLE_BIT)
#define XFS_DIFLAG2_ANY \
- (XFS_DIFLAG2_DAX | XFS_DIFLAG2_REFLINK | XFS_DIFLAG2_COWEXTSIZE)
+ (XFS_DIFLAG2_DAX | XFS_DIFLAG2_REFLINK | XFS_DIFLAG2_COWEXTSIZE | \
+ XFS_DIFLAG2_IOMAP_IMMUTABLE)
/*
* Inode number format:
diff --git a/fs/xfs/xfs_inode.c b/fs/xfs/xfs_inode.c
index ceef77c0416a..4ca22e272ce6 100644
--- a/fs/xfs/xfs_inode.c
+++ b/fs/xfs/xfs_inode.c
@@ -674,6 +674,8 @@ _xfs_dic2xflags(
flags |= FS_XFLAG_DAX;
if (di_flags2 & XFS_DIFLAG2_COWEXTSIZE)
flags |= FS_XFLAG_COWEXTSIZE;
+ if (di_flags2 & XFS_DIFLAG2_IOMAP_IMMUTABLE)
+ flags |= FS_XFLAG_IOMAP_IMMUTABLE;
}
if (has_attr)
diff --git a/fs/xfs/xfs_ioctl.c b/fs/xfs/xfs_ioctl.c
index 2e64488bc4de..df2eef0f9d45 100644
--- a/fs/xfs/xfs_ioctl.c
+++ b/fs/xfs/xfs_ioctl.c
@@ -978,6 +978,7 @@ xfs_set_diflags(
return;
di_flags2 = (ip->i_d.di_flags2 & XFS_DIFLAG2_REFLINK);
+ di_flags2 |= (ip->i_d.di_flags2 & XFS_DIFLAG2_IOMAP_IMMUTABLE);
if (xflags & FS_XFLAG_DAX)
di_flags2 |= XFS_DIFLAG2_DAX;
if (xflags & FS_XFLAG_COWEXTSIZE)
diff --git a/fs/xfs/xfs_iops.c b/fs/xfs/xfs_iops.c
index 469c9fa4c178..174ef95453f5 100644
--- a/fs/xfs/xfs_iops.c
+++ b/fs/xfs/xfs_iops.c
@@ -1186,9 +1186,10 @@ xfs_diflags_to_iflags(
struct xfs_inode *ip)
{
uint16_t flags = ip->i_d.di_flags;
+ uint64_t flags2 = ip->i_d.di_flags2;
inode->i_flags &= ~(S_IMMUTABLE | S_APPEND | S_SYNC |
- S_NOATIME | S_DAX);
+ S_NOATIME | S_DAX | S_IOMAP_IMMUTABLE);
if (flags & XFS_DIFLAG_IMMUTABLE)
inode->i_flags |= S_IMMUTABLE;
@@ -1201,9 +1202,10 @@ xfs_diflags_to_iflags(
if (S_ISREG(inode->i_mode) &&
ip->i_mount->m_sb.sb_blocksize == PAGE_SIZE &&
!xfs_is_reflink_inode(ip) &&
- (ip->i_mount->m_flags & XFS_MOUNT_DAX ||
- ip->i_d.di_flags2 & XFS_DIFLAG2_DAX))
+ (ip->i_mount->m_flags & XFS_MOUNT_DAX || flags2 & XFS_DIFLAG2_DAX))
inode->i_flags |= S_DAX;
+ if (flags2 & XFS_DIFLAG2_IOMAP_IMMUTABLE)
+ inode->i_flags |= S_IOMAP_IMMUTABLE;
}
/*
diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h
index b7495d05e8de..4765e024ad74 100644
--- a/include/uapi/linux/fs.h
+++ b/include/uapi/linux/fs.h
@@ -182,6 +182,7 @@ struct fsxattr {
#define FS_XFLAG_FILESTREAM 0x00004000 /* use filestream allocator */
#define FS_XFLAG_DAX 0x00008000 /* use DAX for IO */
#define FS_XFLAG_COWEXTSIZE 0x00010000 /* CoW extent size allocator hint */
+#define FS_XFLAG_IOMAP_IMMUTABLE 0x00020000 /* block map immutable */
#define FS_XFLAG_HASATTR 0x80000000 /* no DIFLAG for this */
/* the read-only stuff doesn't really belong here, but any other place is
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web