Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1549151 > unrolled thread
| Started by | Roman Pen <roman.penyaev@profitbricks.com> |
|---|---|
| First post | 2017-01-02 14:00 +0100 |
| Last post | 2017-01-02 14:00 +0100 |
| Articles | 9 — 3 participants |
Back to article view | Back to linux.kernel
[PATCH 0/3] ext4: fallocate insert/collapse range fixes Roman Pen <roman.penyaev@profitbricks.com> - 2017-01-02 14:00 +0100
[PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch Roman Pen <roman.penyaev@profitbricks.com> - 2017-01-02 14:00 +0100
Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch Theodore Ts'o <tytso@mit.edu> - 2017-01-03 15:50 +0100
Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch Roman Penyaev <roman.penyaev@profitbricks.com> - 2017-01-03 21:50 +0100
Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch Roman Penyaev <roman.penyaev@profitbricks.com> - 2017-01-04 19:40 +0100
Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch Theodore Ts'o <tytso@mit.edu> - 2017-01-05 01:00 +0100
Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch Roman Penyaev <roman.penyaev@profitbricks.com> - 2017-01-05 09:10 +0100
Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch Roman Penyaev <roman.penyaev@profitbricks.com> - 2017-01-06 21:30 +0100
[PATCH 2/3] ext4: Do not populate extents tree with outdated offsets while shifting extents Roman Pen <roman.penyaev@profitbricks.com> - 2017-01-02 14:00 +0100
| From | Roman Pen <roman.penyaev@profitbricks.com> |
|---|---|
| Date | 2017-01-02 14:00 +0100 |
| Subject | [PATCH 0/3] ext4: fallocate insert/collapse range fixes |
| Message-ID | <sVanv-6UV-3@gated-at.bofh.it> |
Hi all.
For couple of days I've played with inserting and collapsing range using
fallocate and have found two nasty bugs:
1. On right shift (insert range) start block is not included in the range
and hole appears at the wrong offset. The bug can be easily reproduced by
the following test:
ptr = malloc(4096);
assert(ptr);
fd = open("./ext4.file", O_CREAT | O_TRUNC | O_RDWR, 0600);
assert(fd >= 0);
rc = fallocate(fd, 0, 0, 8192);
assert(rc == 0);
for (i = 0; i < 2048; i++)
*((unsigned short *)ptr + i) = 0xbeef;
rc = pwrite(fd, ptr, 4096, 0);
assert(rc == 4096);
rc = pwrite(fd, ptr, 4096, 4096);
assert(rc == 4096);
for (block = 2; block < 1000; block++) {
rc = fallocate(fd, FALLOC_FL_INSERT_RANGE, 4096, 4096);
assert(rc == 0);
for (i = 0; i < 2048; i++)
*((unsigned short *)ptr + i) = block;
rc = pwrite(fd, ptr, 4096, 4096);
assert(rc == 4096);
}
After the test no zero blocks should appear (test always does pwrite() after
fallocate), but zero blocks do exist:
$ hexdump ./ext4.file | grep '0000 0000'
This bug is targeted by the first patch in the set.
2. Inside ext4_ext_shift_extents() function ext4_find_extent() is called
without EXT4_EX_NOCACHE flag, which should prevent cache population. This
leads to outdated offsets in the extents tree and wrong data blocks, which
can be observed doing read(). That is also quite well reproduced by the
test above.
This is fixed by the second patch.
3. Just a minor optimization: linear search of a extent inside a block is
replaced by a binsearch. This is the third patch.
Roman Pen (3):
ext4: Include forgotten start block on fallocate insert range
ext4: Do not populate extents tree with outdated offsets while
shifting extents
ext4: Find desired extent in ext4_ext_shift_extents() using binsearch
fs/ext4/extents.c | 40 +++++++++++++++++++++++++++-------------
1 file changed, 27 insertions(+), 13 deletions(-)
Signed-off-by: Roman Pen <roman.penyaev@profitbricks.com>
Cc: Namjae Jeon <namjae.jeon@samsung.com>
Cc: "Theodore Ts'o" <tytso@mit.edu>
Cc: Andreas Dilger <adilger.kernel@dilger.ca>
Cc: linux-ext4@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
--
2.10.2
[toc] | [next] | [standalone]
| From | Roman Pen <roman.penyaev@profitbricks.com> |
|---|---|
| Date | 2017-01-02 14:00 +0100 |
| Subject | [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch |
| Message-ID | <sVanv-6UV-15@gated-at.bofh.it> |
| In reply to | #1549151 |
The aim of this patch is to optimize a search of an extent while doing right shift using binsearch. Cc: Namjae Jeon <namjae.jeon@samsung.com> Cc: "Theodore Ts'o" <tytso@mit.edu> Cc: Andreas Dilger <adilger.kernel@dilger.ca> Cc: linux-ext4@vger.kernel.org Cc: linux-kernel@vger.kernel.org --- fs/ext4/extents.c | 13 +++++++++---- 1 file changed, 9 insertions(+), 4 deletions(-) diff --git a/fs/ext4/extents.c b/fs/ext4/extents.c index 9fbf92ca358c..f65cc2762780 100644 --- a/fs/ext4/extents.c +++ b/fs/ext4/extents.c @@ -5433,10 +5433,15 @@ ext4_ext_shift_extents(struct inode *inode, handle_t *handle, else /* Beginning is reached, end of the loop */ iterator = NULL; - /* Update path extent in case we need to stop */ - while (le32_to_cpu(extent->ee_block) < start) - extent++; - path[depth].p_ext = extent; + if (le32_to_cpu(extent->ee_block) < start) + /* + * Desired extent is somewhere in the middle, + * do binsearch and update a path with it. + */ + ext4_ext_binsearch(inode, &path[depth], start); + else + /* Set the first extent */ + path[depth].p_ext = extent; } ret = ext4_ext_shift_path_extents(path, shift, inode, handle, SHIFT); -- 2.10.2
[toc] | [prev] | [next] | [standalone]
| From | Theodore Ts'o <tytso@mit.edu> |
|---|---|
| Date | 2017-01-03 15:50 +0100 |
| Subject | Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch |
| Message-ID | <sVyzw-7Iz-21@gated-at.bofh.it> |
| In reply to | #1549152 |
On Mon, Jan 02, 2017 at 01:54:50PM +0100, Roman Pen wrote: > The aim of this patch is to optimize a search of an extent while > doing right shift using binsearch. > > Cc: Namjae Jeon <namjae.jeon@samsung.com> > Cc: "Theodore Ts'o" <tytso@mit.edu> > Cc: Andreas Dilger <adilger.kernel@dilger.ca> > Cc: linux-ext4@vger.kernel.org > Cc: linux-kernel@vger.kernel.org I would really appreciate it if patches that touch sensitive code (and the extents manipulation code is an example of code which is fairly subtle and fragile) are tested using xfstests first before you submit them for review. I've done a lot of work to make using xfstests simple and easy for ext4 developers. See: http://thunk.org/gce-xfstests (especially the last slide :-). BEGIN TEST 4k: Ext4 4k block Mon Jan 2 23:06:02 EST 2017 Failures: generic/061 generic/063 generic/075 generic/091 generic/112 generic/127 generic/231 generic/263 generic/389 BEGIN TEST 1k: Ext4 1k block Mon Jan 2 23:56:29 EST 2017 Failures: ext4/307 generic/013 generic/014 generic/016 generic/018 generic/020 generic/021 generic/022 generic/023 generic/024 generic/025 generic/028 generic/035 generic/036 generic/058 generic/060 generic/061 generic/063 generic/067 generic/070 generic/072 generic/074 generic/075 generic/077 generic/078 generic/080 generic/081 generic/082 generic/086 generic/087 generic/088 generic/089 generic/091 generic/092 generic/100 generic/112 generic/113 generic/114 generic/117 generic/123 generic/124 generic/126 generic/127 generic/131 generic/133 generic/184 generic/192 generic/193 generic/198 generic/207 generic/208 generic/209 generic/210 generic/211 generic/212 generic/213 generic/214 generic/215 generic/221 generic/228 generic/231 generic/233 generic/236 generic/237 generic/239 generic/240 generic/241 generic/245 generic/246 generic/247 generic/248 generic/249 generic/255 generic/256 generic/257 generic/258 generic/263 generic/269 generic/270 generic/273 generic/285 generic/286 generic/299 generic/300 generic/306 generic/308 generic/309 generic/310 generic/313 generic/314 generic/315 generic/316 generic/323 generic/355 generic/360 generic/361 generic/375 generic/378 generic/389 generic/391 shared/298 The 1k test failures look extremely scary, but that's because generic/013 corrupted the file system, and caused all of the subsequent tests using the test device to fail. Of course, patches _shouldn't_ be corrupting file systems. That's a regression which makes everyone said. :-) - Ted
[toc] | [prev] | [next] | [standalone]
| From | Roman Penyaev <roman.penyaev@profitbricks.com> |
|---|---|
| Date | 2017-01-03 21:50 +0100 |
| Subject | Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch |
| Message-ID | <sVEbU-3aT-15@gated-at.bofh.it> |
| In reply to | #1549833 |
On Tue, Jan 3, 2017 at 3:40 PM, Theodore Ts'o <tytso@mit.edu> wrote: > On Mon, Jan 02, 2017 at 01:54:50PM +0100, Roman Pen wrote: >> The aim of this patch is to optimize a search of an extent while >> doing right shift using binsearch. >> >> Cc: Namjae Jeon <namjae.jeon@samsung.com> >> Cc: "Theodore Ts'o" <tytso@mit.edu> >> Cc: Andreas Dilger <adilger.kernel@dilger.ca> >> Cc: linux-ext4@vger.kernel.org >> Cc: linux-kernel@vger.kernel.org > > I would really appreciate it if patches that touch sensitive code (and > the extents manipulation code is an example of code which is fairly > subtle and fragile) are tested using xfstests first before you submit > them for review. I've done a lot of work to make using xfstests > simple and easy for ext4 developers. See: > > http://thunk.org/gce-xfstests > > (especially the last slide :-). Thanks. Those slides that's exactly what I needed. > > BEGIN TEST 4k: Ext4 4k block Mon Jan 2 23:06:02 EST 2017 > Failures: generic/061 generic/063 generic/075 generic/091 generic/112 generic/127 generic/231 generic/263 generic/389 > BEGIN TEST 1k: Ext4 1k block Mon Jan 2 23:56:29 EST 2017 > Failures: ext4/307 generic/013 generic/014 generic/016 generic/018 generic/020 generic/021 generic/022 generic/023 > generic/024 generic/025 generic/028 generic/035 generic/036 generic/058 generic/060 generic/061 generic/063 generic/067 > generic/070 generic/072 generic/074 generic/075 generic/077 generic/078 generic/080 generic/081 generic/082 generic/086 > generic/087 generic/088 generic/089 generic/091 generic/092 generic/100 generic/112 generic/113 generic/114 generic/117 > generic/123 generic/124 generic/126 generic/127 generic/131 generic/133 generic/184 generic/192 generic/193 generic/198 > generic/207 generic/208 generic/209 generic/210 generic/211 generic/212 generic/213 generic/214 generic/215 generic/221 > generic/228 generic/231 generic/233 generic/236 generic/237 generic/239 generic/240 generic/241 generic/245 generic/246 > generic/247 generic/248 generic/249 generic/255 generic/256 generic/257 generic/258 generic/263 generic/269 generic/270 > generic/273 generic/285 generic/286 generic/299 generic/300 generic/306 generic/308 generic/309 generic/310 generic/313 > generic/314 generic/315 generic/316 generic/323 generic/355 generic/360 generic/361 generic/375 generic/378 generic/389 > generic/391 shared/298 > > The 1k test failures look extremely scary, but that's because > generic/013 corrupted the file system, and caused all of the > subsequent tests using the test device to fail. Of course, patches > _shouldn't_ be corrupting file systems. That's a regression which > makes everyone said. :-) > True, I ran only my own tests for inserting/collapsing big amount of blocks. My xfstests output is the following: (I had to say that right now I am testing on 4.4.28 kernel and testing on latest sources taken from linux-next will require some time, but of course I will retest and send up-to-date results) --- Failures: ext4/302 ext4/303 ext4/304 generic/061 generic/063 generic/075 generic/079 generic/091 generic/112 generic/127 generic/252 generic/263 Failed 12 of 200 tests --- Many of tests were ignored, because of reflink which is not supported, also there are some failures and I will try to figure out is that related to my changes, but still I do not see massive corruptions. Quick questions regarding xfstests: 1. What is the optimal size for $TEST_DEV ? I see these seconds numbers on a small Vm (disk is 16gb): ext4/007 22s and these numbers on a server with 2TB disk: ext4/007 547s 007 test does e2fsck, so it depends on $TEST_DEV size. So what optimal size to choose to cover all corner cases but not to wait forever? 2. I see many of the tests are ignored. Do we have some perfect run, reference, Standard, to see how many tests run on up-to-date kernel? E.g. I see generic/038 has this output: "[not run] This test requires at least 10GB free". Increasing size of $SCRATCH partition will bring the test to live. So would be nice to have some reference numbers that those tests are expected to run and to pass. -- Roman
[toc] | [prev] | [next] | [standalone]
| From | Roman Penyaev <roman.penyaev@profitbricks.com> |
|---|---|
| Date | 2017-01-04 19:40 +0100 |
| Subject | Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch |
| Message-ID | <sVYDE-8nf-19@gated-at.bofh.it> |
| In reply to | #1550171 |
On Tue, Jan 3, 2017 at 11:15 PM, Theodore Ts'o <tytso@mit.edu> wrote: > On Tue, Jan 03, 2017 at 09:44:15PM +0100, Roman Penyaev wrote: >> >> (I had to say that right now I am testing on 4.4.28 kernel and testing >> on latest sources taken from linux-next will require some time, but of >> course I will retest and send up-to-date results) >> >> --- >> Failures: ext4/302 ext4/303 ext4/304 generic/061 generic/063 >> generic/075 generic/079 generic/091 generic/112 generic/127 >> generic/252 generic/263 >> Failed 12 of 200 tests > > You didn't say what file system configuration you're using, but I'm > assuming it's a default ext4 4k configuration? Yep, that was the default run, 4K block. > One of the things > about my kvm-xfstests and gce-xfstests setup is that I test a range of > file system configurations, and there are specific exclude files to > skip certain known failures. Currently we are skipping the following > tests globally (from kvm-xfstests/test-appliances/files/root/fs/ext4/exclude): [ cut ] > You'll see that modulo the encryption failures and some failures on > the 1k config case which I need to track down and iron out, we're in > pretty good shape. Below please find the summary; attached please > find the full log file. Thanks a lot. Perfect reference. That is quite enough for a good start. This kvm-xfstests is a nice thingy. That was pretty annoying to rebuild xfsprogs on a debian jessie, which does not have 'xfs_io -c finsert' by default. Theodore, recently I sent an ext4/023 xfstest test. New test reproduces the problem with a wrong right shift. I tried to make the test explicit, but not sure succeeded or not. Probably, md5sum is not a good choice to make things clear, but hexdump of 1000 blocks is quite large (~80K). Could you please take a look? Now I am running kvm-xfstests, will figure out what I've done wrong with recent ext4 patches. -- Roman
[toc] | [prev] | [next] | [standalone]
| From | Theodore Ts'o <tytso@mit.edu> |
|---|---|
| Date | 2017-01-05 01:00 +0100 |
| Subject | Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch |
| Message-ID | <sW3Dk-3ex-29@gated-at.bofh.it> |
| In reply to | #1551089 |
It looks like the original (before your patch) 1k failures due to a bug introduced via the block.git tree, which has since been fixed in Linus's mainline tree as of today. It wouldn't surprise me if the bug interacted poorly your changes, so things will probably better with your patches applied directly on top of the tip of Linus's tree. That being said, it looks like there were still regressions introduced on the 4k configuration, so I'm in the middle of rerunning my baseline and trying out your patches as well. Cheers, - Ted
[toc] | [prev] | [next] | [standalone]
| From | Roman Penyaev <roman.penyaev@profitbricks.com> |
|---|---|
| Date | 2017-01-05 09:10 +0100 |
| Subject | Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch |
| Message-ID | <sWbhw-76-15@gated-at.bofh.it> |
| In reply to | #1551514 |
On Thu, Jan 5, 2017 at 12:58 AM, Theodore Ts'o <tytso@mit.edu> wrote: > It looks like the original (before your patch) 1k failures due to a > bug introduced via the block.git tree, which has since been fixed in > Linus's mainline tree as of today. It wouldn't surprise me if the bug > interacted poorly your changes, so things will probably better with > your patches applied directly on top of the tip of Linus's tree. > > That being said, it looks like there were still regressions introduced > on the 4k configuration, so I'm in the middle of rerunning my baseline > and trying out your patches as well. It seems that some of the "finsert" tests from xfstests are broken by my patches. Now I am able to run all configurations with a kvm-xfstests help, so will take a look deeply. -- Roman
[toc] | [prev] | [next] | [standalone]
| From | Roman Penyaev <roman.penyaev@profitbricks.com> |
|---|---|
| Date | 2017-01-06 21:30 +0100 |
| Subject | Re: [PATCH 3/3] ext4: Find desired extent in ext4_ext_shift_extents() using binsearch |
| Message-ID | <sWJjc-71s-9@gated-at.bofh.it> |
| In reply to | #1551514 |
On Thu, Jan 5, 2017 at 12:58 AM, Theodore Ts'o <tytso@mit.edu> wrote: > It looks like the original (before your patch) 1k failures due to a > bug introduced via the block.git tree, which has since been fixed in > Linus's mainline tree as of today. It wouldn't surprise me if the bug > interacted poorly your changes, so things will probably better with > your patches applied directly on top of the tip of Linus's tree. > > That being said, it looks like there were still regressions introduced > on the 4k configuration, so I'm in the middle of rerunning my baseline > and trying out your patches as well. I found a bug in my third patch, where I try to optimize linear search. I missed the fact, that ext4_ext_binsearch() searches for the closest extent from the left, but this linear search code does search for closest extent from the right. I ran the 'kvm-xfstests.sh -c 4k -g auto' against b25ead75e7b6 4.10-rc2 with and without my changes. No failures. 1k configuration was run against latest Linus's mainline. Also no regressions were found: these two "generic/270, generic/273" fail regardless my changes. I will resend the patchset without latest patch shortly. -- Roman
[toc] | [prev] | [next] | [standalone]
| From | Roman Pen <roman.penyaev@profitbricks.com> |
|---|---|
| Date | 2017-01-02 14:00 +0100 |
| Subject | [PATCH 2/3] ext4: Do not populate extents tree with outdated offsets while shifting extents |
| Message-ID | <sVanv-6UV-7@gated-at.bofh.it> |
| In reply to | #1549151 |
Inside ext4_ext_shift_extents() function ext4_find_extent() is called
without EXT4_EX_NOCACHE flag, which should prevent cache population.
This leads to oudated offsets in the extents tree and wrong blocks
afterwards.
Patch fixes the problem providing EXT4_EX_NOCACHE flag for each
ext4_find_extents() call inside ext4_ext_shift_extents function.
Signed-off-by: Roman Pen <roman.penyaev@profitbricks.com>
Cc: Namjae Jeon <namjae.jeon@samsung.com>
Cc: "Theodore Ts'o" <tytso@mit.edu>
Cc: Andreas Dilger <adilger.kernel@dilger.ca>
Cc: linux-ext4@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
---
fs/ext4/extents.c | 9 ++++++---
1 file changed, 6 insertions(+), 3 deletions(-)
diff --git a/fs/ext4/extents.c b/fs/ext4/extents.c
index b4987ea2ca79..9fbf92ca358c 100644
--- a/fs/ext4/extents.c
+++ b/fs/ext4/extents.c
@@ -5344,7 +5344,8 @@ ext4_ext_shift_extents(struct inode *inode, handle_t *handle,
ext4_lblk_t stop, *iterator, ex_start, ex_end;
/* Let path point to the last extent */
- path = ext4_find_extent(inode, EXT_MAX_BLOCKS - 1, NULL, 0);
+ path = ext4_find_extent(inode, EXT_MAX_BLOCKS - 1, NULL,
+ EXT4_EX_NOCACHE);
if (IS_ERR(path))
return PTR_ERR(path);
@@ -5360,7 +5361,8 @@ ext4_ext_shift_extents(struct inode *inode, handle_t *handle,
* sure the hole is big enough to accommodate the shift.
*/
if (SHIFT == SHIFT_LEFT) {
- path = ext4_find_extent(inode, start - 1, &path, 0);
+ path = ext4_find_extent(inode, start - 1, &path,
+ EXT4_EX_NOCACHE);
if (IS_ERR(path))
return PTR_ERR(path);
depth = path->p_depth;
@@ -5398,7 +5400,8 @@ ext4_ext_shift_extents(struct inode *inode, handle_t *handle,
* becomes NULL to indicate the end of the loop.
*/
while (iterator && start <= stop) {
- path = ext4_find_extent(inode, *iterator, &path, 0);
+ path = ext4_find_extent(inode, *iterator, &path,
+ EXT4_EX_NOCACHE);
if (IS_ERR(path))
return PTR_ERR(path);
depth = path->p_depth;
--
2.10.2
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web