Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1568313

[PATCH 3.12 003/235] ext4: fix data exposure after a crash

From Jiri Slaby <jslaby@suse.cz>
Newsgroups linux.kernel
Subject [PATCH 3.12 003/235] ext4: fix data exposure after a crash
Date 2017-01-27 13:10 +0100
Message-ID <t4dvQ-3YT-23@gated-at.bofh.it> (permalink)
References <t4cq5-32y-3@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


From: Jan Kara <jack@suse.cz>

3.12-stable review patch.  If anyone has any objections, please let me know.

===============

commit 06bd3c36a733ac27962fea7d6f47168841376824 upstream.

Huang has reported that in his powerfail testing he is seeing stale
block contents in some of recently allocated blocks although he mounts
ext4 in data=ordered mode. After some investigation I have found out
that indeed when delayed allocation is used, we don't add inode to
transaction's list of inodes needing flushing before commit. Originally
we were doing that but commit f3b59291a69d removed the logic with a
flawed argument that it is not needed.

The problem is that although for delayed allocated blocks we write their
contents immediately after allocating them, there is no guarantee that
the IO scheduler or device doesn't reorder things and thus transaction
allocating blocks and attaching them to inode can reach stable storage
before actual block contents. Actually whenever we attach freshly
allocated blocks to inode using a written extent, we should add inode to
transaction's ordered inode list to make sure we properly wait for block
contents to be written before committing the transaction. So that is
what we do in this patch. This also handles other cases where stale data
exposure was possible - like filling hole via mmap in
data=ordered,nodelalloc mode.

The only exception to the above rule are extending direct IO writes where
blkdev_direct_IO() waits for IO to complete before increasing i_size and
thus stale data exposure is not possible. For now we don't complicate
the code with optimizing this special case since the overhead is pretty
low. In case this is observed to be a performance problem we can always
handle it using a special flag to ext4_map_blocks().

Fixes: f3b59291a69d0b734be1fc8be489fef2dd846d3d
Reported-by: "HUANG Weller (CM/ESW12-CN)" <Weller.Huang@cn.bosch.com>
Tested-by: "HUANG Weller (CM/ESW12-CN)" <Weller.Huang@cn.bosch.com>
Signed-off-by: Jan Kara <jack@suse.cz>
Signed-off-by: Theodore Ts'o <tytso@mit.edu>
Signed-off-by: Jiri Slaby <jslaby@suse.cz>
---
 fs/ext4/inode.c | 23 ++++++++++++++---------
 1 file changed, 14 insertions(+), 9 deletions(-)

diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c
index 4a3735a795d0..3fa2da53400d 100644
--- a/fs/ext4/inode.c
+++ b/fs/ext4/inode.c
@@ -701,6 +701,20 @@ has_zeroout:
 		int ret = check_block_validity(inode, map);
 		if (ret != 0)
 			return ret;
+
+		/*
+		 * Inodes with freshly allocated blocks where contents will be
+		 * visible after transaction commit must be on transaction's
+		 * ordered data list.
+		 */
+		if (map->m_flags & EXT4_MAP_NEW &&
+		    !(map->m_flags & EXT4_MAP_UNWRITTEN) &&
+		    !IS_NOQUOTA(inode) &&
+		    ext4_should_order_data(inode)) {
+			ret = ext4_jbd2_file_inode(handle, inode);
+			if (ret)
+				return ret;
+		}
 	}
 	return retval;
 }
@@ -1065,15 +1079,6 @@ static int ext4_write_end(struct file *file,
 	int i_size_changed = 0;
 
 	trace_ext4_write_end(inode, pos, len, copied);
-	if (ext4_test_inode_state(inode, EXT4_STATE_ORDERED_MODE)) {
-		ret = ext4_jbd2_file_inode(handle, inode);
-		if (ret) {
-			unlock_page(page);
-			page_cache_release(page);
-			goto errout;
-		}
-	}
-
 	if (ext4_has_inline_data(inode)) {
 		ret = ext4_write_inline_data_end(inode, pos, len,
 						 copied, page);
-- 
2.11.0

Back to linux.kernel | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

[PATCH 3.12 001/235] driver core: Delete an unnecessary check before the function call "put_device" Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 015/235] USB: serial: kl5kusb105: fix open error path Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 018/235] usb: gadget: composite: correctly initialize ep->maxpacket Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 023/235] Btrfs: fix memory leak in reading btree blocks Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 012/235] Btrfs: fix tree search logic when replaying directory entry deletes Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 010/235] hotplug: Make register and unregister notifier API symmetric Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 003/235] ext4: fix data exposure after a crash Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 019/235] USB: UHCI: report non-PME wakeup signalling for Intel hardware Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 024/235] block_dev: don't test bdev->bd_contains when it is not stable Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 004/235] locking/rtmutex: Prevent dequeue vs. unlock race Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 030/235] ext4: add sanity checking to count_overhead() Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 009/235] m68k: Fix ndelay() macro Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 008/235] can: peak: fix bad memory access and free sequence Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 014/235] USB: serial: option: add dlink dwm-158 Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 013/235] USB: serial: option: add support for Telit LE922A PIDs 0x1040, 0x1041 Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100
  [PATCH 3.12 005/235] locking/rtmutex: Use READ_ONCE() in rt_mutex_owner() Jiri Slaby <jslaby@suse.cz> - 2017-01-27 13:10 +0100

csiph-web