Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #82724 > unrolled thread

Bug#1039883: linux: ext4 corruption with symlinks

Started byLuis Henriques <luis.henriques@linux.dev>
First post2024-06-14 18:30 +0200
Last post2024-06-18 15:30 +0200
Articles 3 — 1 participant

Back to article view | Back to linux.debian.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Bug#1039883: linux: ext4 corruption with symlinks Luis Henriques <luis.henriques@linux.dev> - 2024-06-14 18:30 +0200
    Bug#1039883: linux: ext4 corruption with symlinks Luis Henriques <luis.henriques@linux.dev> - 2024-06-18 12:10 +0200
      Bug#1039883: linux: ext4 corruption with symlinks Luis Henriques <luis.henriques@linux.dev> - 2024-06-18 15:30 +0200

#82724 — Bug#1039883: linux: ext4 corruption with symlinks

FromLuis Henriques <luis.henriques@linux.dev>
Date2024-06-14 18:30 +0200
SubjectBug#1039883: linux: ext4 corruption with symlinks
Message-ID<IPhYl-2V7L-3@gated-at.bofh.it>
On Mon 10 Jun 2024 06:03:58 PM +02, Ben Hutchings wrote;

> On Sun, 5 Nov 2023 16:12:41 +0000 Hervé Werner <dud225@hotmail.com>
> wrote:
>> Hello
>> 
>> I'm sorry for the delay.
>> 
>> > Are you able to reliably preoeduce the issue and can bisect it to
>> > the introducing commit?
>> I faced this issue on real data but I struggled to find a reliable
>> scenario to reproduce it. Here is what I just came up with:
>>   sudo mkfs -t ext4 -O fast_commit,inline_data /dev/sdb
>>   sudo mount /dev/sdb /mnt/
>>   sudo install -d -o myuser /mnt/annex
>>   cd /mnt/annex
>>   git init && git annex init
>>   for i in {1..2}; do
>>     for i in {1..10000}; do
>>       dd if=/dev/urandom of=file-${i} bs=1K count=1 2>/dev/null
>>     done
>>     git annex add -J cpus . >/dev/null && git annex sync -J cpus && git annex fsck -J cpus >/dev/null
>>     git rm * && git annex sync  && git annex dropunused all
>>   done
>> 
>> Then at some point the following error appears:
>>   EXT4-fs error (device sdb): ext4_map_blocks:577: inode #3942343: block 4: comm git-annex:w: lblock 1 mapped to illegal pblock 4 (length 1)
> [...]
>
> I can also reproduce this error message using the above script and:
>
> - Linux 6.10-rc2
> - A 2 GiB loopback devic instead of /dev/sdb
>
> I bisected this back to:
>
> commit 9725958bb75cdfa10f2ec11526fdb23e7485e8e4
> Author: Xin Yin <yinxin.x@bytedance.com>
> Date:   Thu Dec 23 11:23:37 2021 +0800
>  
>     ext4: fast commit may miss tracking unwritten range during ftruncate
>
> It is still possible to cleanly revert that commit from 6.10-rc2, and
> doing so removes the error message.

Because I recently fixed an issue in the fast commit code[1] I was hoping
that you were hitting the same bug.  I've executed the reproducer with the
fix (which hasn't been merged yet) and realised it's definitely a
different problem.

Debugged the issue a bit, it seems to be related with the fact that
ext4_fc_write_inode_data() isn't able to cope with the fact that
'ei->i_fc_lblk_len' is set to EXT_MAX_BLOCKS.

I'm CC'ing Harshad, maybe he has some idea.

[1] https://lore.kernel.org/all/20240529092030.9557-2-luis.henriques@linux.dev

Cheers,
-- 
Luís

[toc] | [next] | [standalone]


#82744

FromLuis Henriques <luis.henriques@linux.dev>
Date2024-06-18 12:10 +0200
Message-ID<IQDWO-3ML8-3@gated-at.bofh.it>
In reply to#82724
On Fri 14 Jun 2024 05:18:45 PM +01, Luis Henriques wrote;
[...}
>>
>> I can also reproduce this error message using the above script and:
>>
>> - Linux 6.10-rc2
>> - A 2 GiB loopback devic instead of /dev/sdb
>>
>> I bisected this back to:
>>
>> commit 9725958bb75cdfa10f2ec11526fdb23e7485e8e4
>> Author: Xin Yin <yinxin.x@bytedance.com>
>> Date:   Thu Dec 23 11:23:37 2021 +0800
>>  
>>     ext4: fast commit may miss tracking unwritten range during ftruncate
>>
>> It is still possible to cleanly revert that commit from 6.10-rc2, and
>> doing so removes the error message.
>
> Because I recently fixed an issue in the fast commit code[1] I was hoping
> that you were hitting the same bug.  I've executed the reproducer with the
> fix (which hasn't been merged yet) and realised it's definitely a
> different problem.
>
> Debugged the issue a bit, it seems to be related with the fact that
> ext4_fc_write_inode_data() isn't able to cope with the fact that
> 'ei->i_fc_lblk_len' is set to EXT_MAX_BLOCKS.

OK, I've looked into this again.  And something I didn't pay attention
before was that the filesystem was created with both fast_commit *and*
inline_data features.  And after some more debugging, I _think_ the patch
bellow should be the fix for this bug.

If I understand it correctly, when an inode has inlined data it means that
there's no inode data to be written and this case should be handled as if
the inode length was zero.

I'll send out a patch later after running a few more tests just to make
sure it doesn't break something else.  But it would awesome if you could
test it too.

Cheers,
-- 
Luís

diff --git a/fs/ext4/fast_commit.c b/fs/ext4/fast_commit.c
index 87c009e0c59a..c56b39a51865 100644
--- a/fs/ext4/fast_commit.c
+++ b/fs/ext4/fast_commit.c
@@ -897,7 +897,7 @@ static int ext4_fc_write_inode_data(struct inode *inode, u32 *crc)
 	int ret;
 
 	mutex_lock(&ei->i_fc_lock);
-	if (ei->i_fc_lblk_len == 0) {
+	if ((ei->i_fc_lblk_len == 0) || (ext4_has_inline_data(inode))) {
 		mutex_unlock(&ei->i_fc_lock);
 		return 0;
 	}

[toc] | [prev] | [next] | [standalone]


#82746

FromLuis Henriques <luis.henriques@linux.dev>
Date2024-06-18 15:30 +0200
Message-ID<IQH4l-3Oyr-1@gated-at.bofh.it>
In reply to#82744
On Tue 18 Jun 2024 10:52:55 AM +01, Luis Henriques wrote;

> On Fri 14 Jun 2024 05:18:45 PM +01, Luis Henriques wrote;
> [...}
>>>
>>> I can also reproduce this error message using the above script and:
>>>
>>> - Linux 6.10-rc2
>>> - A 2 GiB loopback devic instead of /dev/sdb
>>>
>>> I bisected this back to:
>>>
>>> commit 9725958bb75cdfa10f2ec11526fdb23e7485e8e4
>>> Author: Xin Yin <yinxin.x@bytedance.com>
>>> Date:   Thu Dec 23 11:23:37 2021 +0800
>>>  
>>>     ext4: fast commit may miss tracking unwritten range during ftruncate
>>>
>>> It is still possible to cleanly revert that commit from 6.10-rc2, and
>>> doing so removes the error message.
>>
>> Because I recently fixed an issue in the fast commit code[1] I was hoping
>> that you were hitting the same bug.  I've executed the reproducer with the
>> fix (which hasn't been merged yet) and realised it's definitely a
>> different problem.
>>
>> Debugged the issue a bit, it seems to be related with the fact that
>> ext4_fc_write_inode_data() isn't able to cope with the fact that
>> 'ei->i_fc_lblk_len' is set to EXT_MAX_BLOCKS.
>
> OK, I've looked into this again.  And something I didn't pay attention
> before was that the filesystem was created with both fast_commit *and*
> inline_data features.  And after some more debugging, I _think_ the patch
> bellow should be the fix for this bug.
>
> If I understand it correctly, when an inode has inlined data it means that
> there's no inode data to be written and this case should be handled as if
> the inode length was zero.
>
> I'll send out a patch later after running a few more tests just to make
> sure it doesn't break something else.  But it would awesome if you could
> test it too.

Hmm... looking closer, this patch seems to work with this specific test
script, but only because file data is probably small enough to fit in
inode->i_block.  However, it may actually truncate files that have inlined
data if the file data is also stored in the extended attribute space
(i.e. > 60 bytes).

So, the correct fix is probably something like the below patch (which I'll
send out soon).

Cheers,
-- 
Luís

diff --git a/fs/ext4/fast_commit.c b/fs/ext4/fast_commit.c
index 87c009e0c59a..d3a67bc06d10 100644
--- a/fs/ext4/fast_commit.c
+++ b/fs/ext4/fast_commit.c
@@ -649,6 +649,12 @@ void ext4_fc_track_range(handle_t *handle, struct inode *inode, ext4_lblk_t star
 	if (ext4_test_mount_flag(inode->i_sb, EXT4_MF_FC_INELIGIBLE))
 		return;
 
+	if (ext4_has_inline_data(inode)) {
+		ext4_fc_mark_ineligible(inode->i_sb, EXT4_FC_REASON_XATTR,
+					handle);
+		return;
+	}
+
 	args.start = start;
 	args.end = end;
 

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web