Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.kernel > #82724 > unrolled thread
| Started by | Luis Henriques <luis.henriques@linux.dev> |
|---|---|
| First post | 2024-06-14 18:30 +0200 |
| Last post | 2024-06-18 15:30 +0200 |
| Articles | 3 — 1 participant |
Back to article view | Back to linux.debian.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Bug#1039883: linux: ext4 corruption with symlinks Luis Henriques <luis.henriques@linux.dev> - 2024-06-14 18:30 +0200
Bug#1039883: linux: ext4 corruption with symlinks Luis Henriques <luis.henriques@linux.dev> - 2024-06-18 12:10 +0200
Bug#1039883: linux: ext4 corruption with symlinks Luis Henriques <luis.henriques@linux.dev> - 2024-06-18 15:30 +0200
| From | Luis Henriques <luis.henriques@linux.dev> |
|---|---|
| Date | 2024-06-14 18:30 +0200 |
| Subject | Bug#1039883: linux: ext4 corruption with symlinks |
| Message-ID | <IPhYl-2V7L-3@gated-at.bofh.it> |
On Mon 10 Jun 2024 06:03:58 PM +02, Ben Hutchings wrote;
> On Sun, 5 Nov 2023 16:12:41 +0000 Hervé Werner <dud225@hotmail.com>
> wrote:
>> Hello
>>
>> I'm sorry for the delay.
>>
>> > Are you able to reliably preoeduce the issue and can bisect it to
>> > the introducing commit?
>> I faced this issue on real data but I struggled to find a reliable
>> scenario to reproduce it. Here is what I just came up with:
>> sudo mkfs -t ext4 -O fast_commit,inline_data /dev/sdb
>> sudo mount /dev/sdb /mnt/
>> sudo install -d -o myuser /mnt/annex
>> cd /mnt/annex
>> git init && git annex init
>> for i in {1..2}; do
>> for i in {1..10000}; do
>> dd if=/dev/urandom of=file-${i} bs=1K count=1 2>/dev/null
>> done
>> git annex add -J cpus . >/dev/null && git annex sync -J cpus && git annex fsck -J cpus >/dev/null
>> git rm * && git annex sync && git annex dropunused all
>> done
>>
>> Then at some point the following error appears:
>> EXT4-fs error (device sdb): ext4_map_blocks:577: inode #3942343: block 4: comm git-annex:w: lblock 1 mapped to illegal pblock 4 (length 1)
> [...]
>
> I can also reproduce this error message using the above script and:
>
> - Linux 6.10-rc2
> - A 2 GiB loopback devic instead of /dev/sdb
>
> I bisected this back to:
>
> commit 9725958bb75cdfa10f2ec11526fdb23e7485e8e4
> Author: Xin Yin <yinxin.x@bytedance.com>
> Date: Thu Dec 23 11:23:37 2021 +0800
>
> ext4: fast commit may miss tracking unwritten range during ftruncate
>
> It is still possible to cleanly revert that commit from 6.10-rc2, and
> doing so removes the error message.
Because I recently fixed an issue in the fast commit code[1] I was hoping
that you were hitting the same bug. I've executed the reproducer with the
fix (which hasn't been merged yet) and realised it's definitely a
different problem.
Debugged the issue a bit, it seems to be related with the fact that
ext4_fc_write_inode_data() isn't able to cope with the fact that
'ei->i_fc_lblk_len' is set to EXT_MAX_BLOCKS.
I'm CC'ing Harshad, maybe he has some idea.
[1] https://lore.kernel.org/all/20240529092030.9557-2-luis.henriques@linux.dev
Cheers,
--
Luís
[toc] | [next] | [standalone]
| From | Luis Henriques <luis.henriques@linux.dev> |
|---|---|
| Date | 2024-06-18 12:10 +0200 |
| Message-ID | <IQDWO-3ML8-3@gated-at.bofh.it> |
| In reply to | #82724 |
On Fri 14 Jun 2024 05:18:45 PM +01, Luis Henriques wrote;
[...}
>>
>> I can also reproduce this error message using the above script and:
>>
>> - Linux 6.10-rc2
>> - A 2 GiB loopback devic instead of /dev/sdb
>>
>> I bisected this back to:
>>
>> commit 9725958bb75cdfa10f2ec11526fdb23e7485e8e4
>> Author: Xin Yin <yinxin.x@bytedance.com>
>> Date: Thu Dec 23 11:23:37 2021 +0800
>>
>> ext4: fast commit may miss tracking unwritten range during ftruncate
>>
>> It is still possible to cleanly revert that commit from 6.10-rc2, and
>> doing so removes the error message.
>
> Because I recently fixed an issue in the fast commit code[1] I was hoping
> that you were hitting the same bug. I've executed the reproducer with the
> fix (which hasn't been merged yet) and realised it's definitely a
> different problem.
>
> Debugged the issue a bit, it seems to be related with the fact that
> ext4_fc_write_inode_data() isn't able to cope with the fact that
> 'ei->i_fc_lblk_len' is set to EXT_MAX_BLOCKS.
OK, I've looked into this again. And something I didn't pay attention
before was that the filesystem was created with both fast_commit *and*
inline_data features. And after some more debugging, I _think_ the patch
bellow should be the fix for this bug.
If I understand it correctly, when an inode has inlined data it means that
there's no inode data to be written and this case should be handled as if
the inode length was zero.
I'll send out a patch later after running a few more tests just to make
sure it doesn't break something else. But it would awesome if you could
test it too.
Cheers,
--
Luís
diff --git a/fs/ext4/fast_commit.c b/fs/ext4/fast_commit.c
index 87c009e0c59a..c56b39a51865 100644
--- a/fs/ext4/fast_commit.c
+++ b/fs/ext4/fast_commit.c
@@ -897,7 +897,7 @@ static int ext4_fc_write_inode_data(struct inode *inode, u32 *crc)
int ret;
mutex_lock(&ei->i_fc_lock);
- if (ei->i_fc_lblk_len == 0) {
+ if ((ei->i_fc_lblk_len == 0) || (ext4_has_inline_data(inode))) {
mutex_unlock(&ei->i_fc_lock);
return 0;
}
[toc] | [prev] | [next] | [standalone]
| From | Luis Henriques <luis.henriques@linux.dev> |
|---|---|
| Date | 2024-06-18 15:30 +0200 |
| Message-ID | <IQH4l-3Oyr-1@gated-at.bofh.it> |
| In reply to | #82744 |
On Tue 18 Jun 2024 10:52:55 AM +01, Luis Henriques wrote;
> On Fri 14 Jun 2024 05:18:45 PM +01, Luis Henriques wrote;
> [...}
>>>
>>> I can also reproduce this error message using the above script and:
>>>
>>> - Linux 6.10-rc2
>>> - A 2 GiB loopback devic instead of /dev/sdb
>>>
>>> I bisected this back to:
>>>
>>> commit 9725958bb75cdfa10f2ec11526fdb23e7485e8e4
>>> Author: Xin Yin <yinxin.x@bytedance.com>
>>> Date: Thu Dec 23 11:23:37 2021 +0800
>>>
>>> ext4: fast commit may miss tracking unwritten range during ftruncate
>>>
>>> It is still possible to cleanly revert that commit from 6.10-rc2, and
>>> doing so removes the error message.
>>
>> Because I recently fixed an issue in the fast commit code[1] I was hoping
>> that you were hitting the same bug. I've executed the reproducer with the
>> fix (which hasn't been merged yet) and realised it's definitely a
>> different problem.
>>
>> Debugged the issue a bit, it seems to be related with the fact that
>> ext4_fc_write_inode_data() isn't able to cope with the fact that
>> 'ei->i_fc_lblk_len' is set to EXT_MAX_BLOCKS.
>
> OK, I've looked into this again. And something I didn't pay attention
> before was that the filesystem was created with both fast_commit *and*
> inline_data features. And after some more debugging, I _think_ the patch
> bellow should be the fix for this bug.
>
> If I understand it correctly, when an inode has inlined data it means that
> there's no inode data to be written and this case should be handled as if
> the inode length was zero.
>
> I'll send out a patch later after running a few more tests just to make
> sure it doesn't break something else. But it would awesome if you could
> test it too.
Hmm... looking closer, this patch seems to work with this specific test
script, but only because file data is probably small enough to fit in
inode->i_block. However, it may actually truncate files that have inlined
data if the file data is also stored in the extended attribute space
(i.e. > 60 bytes).
So, the correct fix is probably something like the below patch (which I'll
send out soon).
Cheers,
--
Luís
diff --git a/fs/ext4/fast_commit.c b/fs/ext4/fast_commit.c
index 87c009e0c59a..d3a67bc06d10 100644
--- a/fs/ext4/fast_commit.c
+++ b/fs/ext4/fast_commit.c
@@ -649,6 +649,12 @@ void ext4_fc_track_range(handle_t *handle, struct inode *inode, ext4_lblk_t star
if (ext4_test_mount_flag(inode->i_sb, EXT4_MF_FC_INELIGIBLE))
return;
+ if (ext4_has_inline_data(inode)) {
+ ext4_fc_mark_ineligible(inode->i_sb, EXT4_FC_REASON_XATTR,
+ handle);
+ return;
+ }
+
args.start = start;
args.end = end;
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.kernel
csiph-web