Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1503409 > unrolled thread
| Started by | Dave Jones <davej@codemonkey.org.uk> |
|---|---|
| First post | 2016-10-19 00:50 +0200 |
| Last post | 2016-10-20 09:30 +0200 |
| Articles | 12 on this page of 52 — 9 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-19 00:50 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-19 01:40 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-19 02:20 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-19 02:30 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-21 00:50 +0200
Re: bio linked list corruption. Andy Lutomirski <luto@kernel.org> - 2016-10-19 03:10 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-21 01:00 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-21 01:10 +0200
Re: bio linked list corruption. Andy Lutomirski <luto@amacapital.net> - 2016-10-21 01:30 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-21 22:10 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-21 22:30 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-21 23:20 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-22 17:30 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-24 06:50 +0200
Re: bio linked list corruption. Andy Lutomirski <luto@amacapital.net> - 2016-10-24 22:10 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-24 22:50 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-24 23:20 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-25 00:00 +0200
Re: bio linked list corruption. Andy Lutomirski <luto@amacapital.net> - 2016-10-25 00:50 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-25 02:10 +0200
Re: bio linked list corruption. Andy Lutomirski <luto@amacapital.net> - 2016-10-25 03:20 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-26 02:30 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-26 03:40 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-26 03:40 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-26 18:40 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-26 18:50 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-26 20:20 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-26 20:50 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-26 21:10 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-27 00:10 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-27 00:30 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-27 00:50 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-27 01:00 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-27 01:00 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-27 01:10 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-27 01:50 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-31 20:00 +0100
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-31 20:40 +0100
Re: btrfs btree_ctree_super fault Dave Jones <davej@codemonkey.org.uk> - 2016-11-06 18:00 +0100
Re: btrfs btree_ctree_super fault Dave Jones <davej@codemonkey.org.uk> - 2016-11-08 16:00 +0100
Re: btrfs btree_ctree_super fault Dave Jones <davej@codemonkey.org.uk> - 2016-11-10 15:40 +0100
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-27 01:10 +0200
Re: bio linked list corruption. Christoph Hellwig <hch@infradead.org> - 2016-10-27 08:40 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-27 18:40 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-27 01:10 +0200
Re: bio linked list corruption. Dave Chinner <david@fromorbit.com> - 2016-10-27 07:50 +0200
Re: bio linked list corruption. Dave Jones <davej@codemonkey.org.uk> - 2016-10-27 19:30 +0200
Re: bio linked list corruption. Andy Lutomirski <luto@amacapital.net> - 2016-10-21 01:10 +0200
Re: bio linked list corruption. Philipp Hahn <pmhahn@pmhahn.de> - 2016-10-19 19:10 +0200
Re: bio linked list corruption. Linus Torvalds <torvalds@linux-foundation.org> - 2016-10-19 19:50 +0200
Re: bio linked list corruption. Ingo Molnar <mingo@kernel.org> - 2016-10-20 09:00 +0200
Re: bio linked list corruption. Thomas Gleixner <tglx@linutronix.de> - 2016-10-20 09:30 +0200
Page 3 of 3 — ← Prev page 1 2 [3]
| From | Dave Jones <davej@codemonkey.org.uk> |
|---|---|
| Date | 2016-11-10 15:40 +0100 |
| Subject | Re: btrfs btree_ctree_super fault |
| Message-ID | <sBYGe-7rk-29@gated-at.bofh.it> |
| In reply to | #1517264 |
On Tue, Nov 08, 2016 at 10:08:04AM -0500, Chris Mason wrote: > > And another new one: > > > > kernel BUG at fs/btrfs/ctree.c:3172! > > > > Call Trace: > > [<ffffffffa009cff0>] __btrfs_drop_extents+0xb00/0xe30 [btrfs] > > We've been hunting this one for at least two years. It's the white > whale of btrfs bugs. Josef has a semi-reliable reproducer now, but I > think it's not the same as the pagevec based problems you reported earlier. Great, now for whatever reason, I'm hitting this over and over. Even better, after the last time I hit it, it reboot and this happened during boot.. BTRFS info (device sda6): disk space caching is enabled BTRFS info (device sda6): has skinny extents BTRFS info (device sda3): disk space caching is enabled ------------[ cut here ]------------ WARNING: CPU: 1 PID: 443 at fs/btrfs/file.c:546 btrfs_drop_extent_cache+0x411/0x420 [btrfs] CPU: 1 PID: 443 Comm: mount Not tainted 4.9.0-rc4-think+ #1 ffffc90000c4b468 ffffffff813b66bc 0000000000000000 0000000000000000 ffffc90000c4b4a8 ffffffff81086d2b 0000022200c4b488 000000000002f265 40c8dded1afd6000 ffff8804ff5cddc8 ffff8804ef26f2b8 40c8dded1afd5000 Call Trace: [<ffffffff813b66bc>] dump_stack+0x4f/0x73 [<ffffffff81086d2b>] __warn+0xcb/0xf0 [<ffffffff81086e5d>] warn_slowpath_null+0x1d/0x20 [<ffffffffa009c0f1>] btrfs_drop_extent_cache+0x411/0x420 [btrfs] [<ffffffff81215923>] ? alloc_debug_processing+0x73/0x1b0 [<ffffffffa009c93f>] __btrfs_drop_extents+0x44f/0xe30 [btrfs] [<ffffffffa005426a>] ? btrfs_alloc_path+0x1a/0x20 [btrfs] [<ffffffffa005426a>] ? btrfs_alloc_path+0x1a/0x20 [btrfs] [<ffffffff8121842a>] ? kmem_cache_alloc+0x2aa/0x330 [<ffffffffa005426a>] ? btrfs_alloc_path+0x1a/0x20 [btrfs] [<ffffffffa009e399>] btrfs_drop_extents+0x79/0xa0 [btrfs] [<ffffffffa00ce4d1>] replay_one_extent+0x1e1/0x710 [btrfs] [<ffffffffa00cec6d>] replay_one_buffer+0x26d/0x7e0 [btrfs] [<ffffffff8121732c>] ? ___slab_alloc.constprop.83+0x27c/0x5c0 [<ffffffffa005426a>] ? btrfs_alloc_path+0x1a/0x20 [btrfs] [<ffffffff813d5b87>] ? debug_smp_processor_id+0x17/0x20 [<ffffffffa00ca3db>] walk_up_log_tree+0xeb/0x240 [btrfs] [<ffffffffa00ca5d6>] walk_log_tree+0xa6/0x1d0 [btrfs] [<ffffffffa00d32fc>] btrfs_recover_log_trees+0x1dc/0x460 [btrfs] [<ffffffffa00cea00>] ? replay_one_extent+0x710/0x710 [btrfs] [<ffffffffa0081f65>] open_ctree+0x2575/0x2670 [btrfs] [<ffffffffa005144b>] btrfs_mount+0xd0b/0xe10 [btrfs] [<ffffffff811da804>] ? pcpu_alloc+0x2d4/0x660 [<ffffffff810dce41>] ? lockdep_init_map+0x61/0x200 [<ffffffff810d39eb>] ? __init_waitqueue_head+0x3b/0x50 [<ffffffff81243794>] mount_fs+0x14/0xa0 [<ffffffff8126305b>] vfs_kern_mount+0x6b/0x150 [<ffffffffa0050a08>] btrfs_mount+0x2c8/0xe10 [btrfs] [<ffffffff811da804>] ? pcpu_alloc+0x2d4/0x660 [<ffffffff810dce41>] ? lockdep_init_map+0x61/0x200 [<ffffffff810dce41>] ? lockdep_init_map+0x61/0x200 [<ffffffff810d39eb>] ? __init_waitqueue_head+0x3b/0x50 [<ffffffff81243794>] mount_fs+0x14/0xa0 [<ffffffff8126305b>] vfs_kern_mount+0x6b/0x150 [<ffffffff81265bd2>] do_mount+0x1c2/0xda0 [<ffffffff811d41c0>] ? memdup_user+0x60/0x90 [<ffffffff81266ac3>] SyS_mount+0x83/0xd0 [<ffffffff81002d81>] do_syscall_64+0x61/0x170 [<ffffffff81894ccb>] entry_SYSCALL64_slow_path+0x25/0x25 ---[ end trace d3fa03bb9c115bbe ]--- BTRFS: error (device sda3) in btrfs_replay_log:2491: errno=-17 Object already exists (Failed to recover log tree) BTRFS error (device sda3): cleaner transaction attach returned -30 BTRFS error (device sda3): open_ctree failed Guess I'll hit it with btrfsck and hope for the best.. Dave
[toc] | [prev] | [next] | [standalone]
| From | Dave Jones <davej@codemonkey.org.uk> |
|---|---|
| Date | 2016-10-27 01:10 +0200 |
| Message-ID | <swFuy-226-13@gated-at.bofh.it> |
| In reply to | #1509917 |
On Wed, Oct 26, 2016 at 05:03:45PM -0600, Jens Axboe wrote: > On 10/26/2016 04:58 PM, Linus Torvalds wrote: > > On Wed, Oct 26, 2016 at 3:51 PM, Linus Torvalds > > <torvalds@linux-foundation.org> wrote: > >> > >> Dave: it might be a good idea to split that "WARN_ON_ONCE()" in > >> blk_mq_merge_queue_io() into two > > > > I did that myself too, since Dave sees this during boot. > > > > But I'm not getting the warning ;( > > > > Dave gets it with ext4, and thats' what I have too, so I'm not sure > > what the required trigger would be. > > Actually, I think I see what might trigger it. You are on nvme, iirc, > and that has a deep queue. Dave, are you testing on a sata drive or > something similar with a shallower queue depth? If we end up sleeping > for a request, I think we could trigger data->ctx being different. yeah, just regular sata. I've been meaning to put an ssd in that box for a while, but now it sounds like my procrastination may have paid off. > Dave, can you hit the warnings with this? Totally untested... Coming up.. Dave
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-10-27 08:40 +0200 |
| Message-ID | <swMw2-6AV-33@gated-at.bofh.it> |
| In reply to | #1509917 |
> Dave, can you hit the warnings with this? Totally untested...
Can we just kill off the unhelpful blk_map_ctx structure, e.g.:
diff --git a/block/blk-mq.c b/block/blk-mq.c
index ddc2eed..d74a74a 100644
--- a/block/blk-mq.c
+++ b/block/blk-mq.c
@@ -1190,21 +1190,15 @@ static inline bool blk_mq_merge_queue_io(struct blk_mq_hw_ctx *hctx,
}
}
-struct blk_map_ctx {
- struct blk_mq_hw_ctx *hctx;
- struct blk_mq_ctx *ctx;
-};
-
static struct request *blk_mq_map_request(struct request_queue *q,
struct bio *bio,
- struct blk_map_ctx *data)
+ struct blk_mq_alloc_data *data)
{
struct blk_mq_hw_ctx *hctx;
struct blk_mq_ctx *ctx;
struct request *rq;
int op = bio_data_dir(bio);
int op_flags = 0;
- struct blk_mq_alloc_data alloc_data;
blk_queue_enter_live(q);
ctx = blk_mq_get_ctx(q);
@@ -1214,12 +1208,10 @@ static struct request *blk_mq_map_request(struct request_queue *q,
op_flags |= REQ_SYNC;
trace_block_getrq(q, bio, op);
- blk_mq_set_alloc_data(&alloc_data, q, 0, ctx, hctx);
- rq = __blk_mq_alloc_request(&alloc_data, op, op_flags);
+ blk_mq_set_alloc_data(data, q, 0, ctx, hctx);
+ rq = __blk_mq_alloc_request(data, op, op_flags);
- hctx->queued++;
- data->hctx = hctx;
- data->ctx = ctx;
+ data->hctx->queued++;
return rq;
}
@@ -1267,7 +1259,7 @@ static blk_qc_t blk_mq_make_request(struct request_queue *q, struct bio *bio)
{
const int is_sync = rw_is_sync(bio_op(bio), bio->bi_opf);
const int is_flush_fua = bio->bi_opf & (REQ_PREFLUSH | REQ_FUA);
- struct blk_map_ctx data;
+ struct blk_mq_alloc_data data;
struct request *rq;
unsigned int request_count = 0;
struct blk_plug *plug;
@@ -1363,7 +1355,7 @@ static blk_qc_t blk_sq_make_request(struct request_queue *q, struct bio *bio)
const int is_flush_fua = bio->bi_opf & (REQ_PREFLUSH | REQ_FUA);
struct blk_plug *plug;
unsigned int request_count = 0;
- struct blk_map_ctx data;
+ struct blk_mq_alloc_data data;
struct request *rq;
blk_qc_t cookie;
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2016-10-27 18:40 +0200 |
| Message-ID | <swVSF-4lA-9@gated-at.bofh.it> |
| In reply to | #1510077 |
On Wed, Oct 26, 2016 at 11:33 PM, Christoph Hellwig <hch@infradead.org> wrote:
>> Dave, can you hit the warnings with this? Totally untested...
>
> Can we just kill off the unhelpful blk_map_ctx structure, e.g.:
Yeah, I found that hard to read too. The difference between
blk_map_ctx and blk_mq_alloc_data is minimal, might as well just use
the latter everywhere.
But please separate that patch from the patch that fixes the ctx list
corruption.
And Jens - can I have at least the corruption fix asap? Due to KS
travel, I'm going to do rc3 on Saturday, and I'd like to have the fix
in at least a day before that..
Linus
[toc] | [prev] | [next] | [standalone]
| From | Dave Jones <davej@codemonkey.org.uk> |
|---|---|
| Date | 2016-10-27 01:10 +0200 |
| Message-ID | <swFuy-226-5@gated-at.bofh.it> |
| In reply to | #1509914 |
On Wed, Oct 26, 2016 at 03:51:01PM -0700, Linus Torvalds wrote:
> Dave: it might be a good idea to split that "WARN_ON_ONCE()" in
> blk_mq_merge_queue_io() into two, since right now it can trigger both
> for the
>
> blk_mq_bio_to_request(rq, bio);
>
> path _and_ for the
>
> if (!blk_mq_attempt_merge(q, ctx, bio)) {
> blk_mq_bio_to_request(rq, bio);
> goto insert_rq;
>
> path. If you split it into two: one before that "insert_rq:" label,
> and one before the "goto insert_rq" thing, then we could see if it is
> just one of the blk_mq_merge_queue_io() cases (or both) that is
> broken..
It's the latter of the two.
[ 12.302392] WARNING: CPU: 3 PID: 272 at block/blk-mq.c:1191
blk_sq_make_request+0x320/0x4d0
Dave
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-10-27 07:50 +0200 |
| Message-ID | <swLJD-613-13@gated-at.bofh.it> |
| In reply to | #1508725 |
On Tue, Oct 25, 2016 at 08:27:52PM -0400, Dave Jones wrote: > DaveC: Do these look like real problems, or is this more "looks like > random memory corruption" ? It's been a while since I did some stress > testing on XFS, so these might not be new.. > > XFS: Assertion failed: oldlen > newlen, file: fs/xfs/libxfs/xfs_bmap.c, line: 2938 > ------------[ cut here ]------------ > kernel BUG at fs/xfs/xfs_message.c:113! > invalid opcode: 0000 [#1] PREEMPT SMP > CPU: 1 PID: 6227 Comm: trinity-c9 Not tainted 4.9.0-rc1-think+ #6 > task: ffff8804f4658040 task.stack: ffff88050568c000 > RIP: 0010:[<ffffffffa02d3e2b>] > [<ffffffffa02d3e2b>] assfail+0x1b/0x20 [xfs] > RSP: 0000:ffff88050568f9e8 EFLAGS: 00010282 > RAX: 00000000ffffffea RBX: 0000000000000046 RCX: 0000000000000001 > RDX: 00000000ffffffc0 RSI: 000000000000000a RDI: ffffffffa02fe34d > RBP: ffff88050568f9e8 R08: 0000000000000000 R09: 0000000000000000 > R10: 000000000000000a R11: f000000000000000 R12: ffff88050568fb44 > R13: 00000000000000f3 R14: ffff8804f292bf88 R15: 000ffffffffe0046 > FS: 00007fe2ddfdfb40(0000) GS:ffff88050a000000(0000) knlGS:0000000000000000 > CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 > CR2: 00007fe2dbabd000 CR3: 00000004f461f000 CR4: 00000000001406e0 > Stack: > ffff88050568fa88 ffffffffa027ccee fffffffffffffff9 ffff8804f16fd8b0 > 0000000000003ffa 0000000000000032 ffff8804f292bf40 0000000000004976 > 000ffffffffe0008 00000000000004fd ffff880400000000 0000000000005107 > Call Trace: > [<ffffffffa027ccee>] xfs_bmap_add_extent_hole_delay+0x54e/0x620 [xfs] > [<ffffffffa027f2d4>] xfs_bmapi_reserve_delalloc+0x2b4/0x400 [xfs] > [<ffffffffa02cadd7>] xfs_file_iomap_begin_delay.isra.12+0x247/0x3c0 [xfs] > [<ffffffffa02cb0d1>] xfs_file_iomap_begin+0x181/0x270 [xfs] > [<ffffffff8122e573>] iomap_apply+0x53/0x100 > [<ffffffff8122e68b>] iomap_file_buffered_write+0x6b/0x90 All this iomap code is new, so it's entirely possible that this is a new bug. The assert failure is indicative that the delalloc extent's metadata reservation grew when we expected it to stay the same or shrink. > XFS: Assertion failed: tp->t_blk_res_used <= tp->t_blk_res, file: fs/xfs/xfs_trans.c, line: 309 ... > [<ffffffffa057abe1>] xfs_trans_mod_sb+0x241/0x280 [xfs] > [<ffffffffa0548eff>] xfs_ag_resv_alloc_extent+0x4f/0xc0 [xfs] > [<ffffffffa050aa3d>] xfs_alloc_ag_vextent+0x23d/0x300 [xfs] > [<ffffffffa050bb1b>] xfs_alloc_vextent+0x5fb/0x6d0 [xfs] > [<ffffffffa051c1b4>] xfs_bmap_btalloc+0x304/0x8e0 [xfs] > [<ffffffffa054648e>] ? xfs_iext_bno_to_ext+0xee/0x170 [xfs] > [<ffffffffa051c8db>] xfs_bmap_alloc+0x2b/0x40 [xfs] > [<ffffffffa051dc30>] xfs_bmapi_write+0x640/0x1210 [xfs] > [<ffffffffa0569326>] xfs_iomap_write_allocate+0x166/0x350 [xfs] > [<ffffffffa05540b0>] xfs_map_blocks+0x1b0/0x260 [xfs] > [<ffffffffa0554beb>] xfs_do_writepage+0x23b/0x730 [xfs] And that's indicative of a delalloc metadata reservation being being too small and so we're allocating unreserved blocks. Different symptoms, same underlying cause, I think. I see the latter assert from time to time in my testing, but it's not common (maybe once a month) and I've never been able to track it down. However, it doesn't affect production systems unless they hit ENOSPC hard enough that it causes the critical reserve pool to be exhausted iand so the allocation fails. That's extremely rare - usually takes a several hundred processes all trying to write as had as they can concurrently and to all slip through the ENOSPC detection without the correct metadata reservations and all require multiple metadata blocks to be allocated durign writeback... If you've got a way to trigger it quickly and reliably, that would be helpful... Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Dave Jones <davej@codemonkey.org.uk> |
|---|---|
| Date | 2016-10-27 19:30 +0200 |
| Message-ID | <swWF3-4Rj-1@gated-at.bofh.it> |
| In reply to | #1510059 |
On Thu, Oct 27, 2016 at 04:41:33PM +1100, Dave Chinner wrote: > And that's indicative of a delalloc metadata reservation being > being too small and so we're allocating unreserved blocks. > > Different symptoms, same underlying cause, I think. > > I see the latter assert from time to time in my testing, but it's > not common (maybe once a month) and I've never been able to track it > down. However, it doesn't affect production systems unless they hit > ENOSPC hard enough that it causes the critical reserve pool to be > exhausted iand so the allocation fails. That's extremely rare - > usually takes a several hundred processes all trying to write as had > as they can concurrently and to all slip through the ENOSPC > detection without the correct metadata reservations and all require > multiple metadata blocks to be allocated durign writeback... > > If you've got a way to trigger it quickly and reliably, that would > be helpful... Seems pretty quickly reproducable for me in some shape or form. Run trinity with --enable-fds=testfile and create enough children to create a fair bit of contention, (for me -C64 seems a good fit on spinning rust, but if you're running on shiny nvme you might have to pump it up a bit). Dave
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2016-10-21 01:10 +0200 |
| Message-ID | <suuDg-5F6-9@gated-at.bofh.it> |
| In reply to | #1505314 |
On Thu, Oct 20, 2016 at 3:50 PM, Dave Jones <davej@codemonkey.org.uk> wrote: > On Tue, Oct 18, 2016 at 06:05:57PM -0700, Andy Lutomirski wrote: > > > One possible debugging approach would be to change: > > > > #define NR_CACHED_STACKS 2 > > > > to > > > > #define NR_CACHED_STACKS 0 > > > > in kernel/fork.c and to set CONFIG_DEBUG_PAGEALLOC=y. The latter will > > force an immediate TLB flush after vfree. > > I can give that idea some runtime, but it sounds like this a case where > we're trying to prove a negative, and that'll just run and run ? In which case I > might do this when I'm travelling on Sunday. The idea is that the stack will be free and unmapped immediately upon process exit if configured like this so that bogus stack accesses (by the CPU, not DMA) would OOPS immediately.
[toc] | [prev] | [next] | [standalone]
| From | Philipp Hahn <pmhahn@pmhahn.de> |
|---|---|
| Date | 2016-10-19 19:10 +0200 |
| Message-ID | <su2xj-4eW-5@gated-at.bofh.it> |
| In reply to | #1503450 |
Hello, Am 19.10.2016 um 01:42 schrieb Chris Mason: > On Tue, Oct 18, 2016 at 04:39:22PM -0700, Linus Torvalds wrote: >> On Tue, Oct 18, 2016 at 4:31 PM, Chris Mason <clm@fb.com> wrote: >>> >>> Jens, not sure if you saw the whole thread. This has triggered bad page >>> state errors, and also corrupted a btrfs list. It hurts me to say, >>> but it >>> might not actually be your fault. >> >> Where is that thread, and what is the "this" that triggers problems? >> >> Looking at the "->mq_list" users, I'm not seeing any changes there in >> the last year or so. So I don't think it's the list itself. > > Seems to be the whole thing: > > http://www.gossamer-threads.com/lists/linux/kernel/2545792 > > My guess is xattr, but I don't have a good reason for that. Nearly a month ago I reported also a "list_add corruption", but with 4.1.6: <http://marc.info/?l=linux-kernel&m=147508265316854&w=2> That server rungs Samba4, which also is a heavy user of xattr. Might be that it is related. Philipp Hahn
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2016-10-19 19:50 +0200 |
| Message-ID | <su3a2-4te-15@gated-at.bofh.it> |
| In reply to | #1504143 |
On Wed, Oct 19, 2016 at 10:09 AM, Philipp Hahn <pmhahn@pmhahn.de> wrote:
>
> Nearly a month ago I reported also a "list_add corruption", but with 4.1.6:
> <http://marc.info/?l=linux-kernel&m=147508265316854&w=2>
>
> That server rungs Samba4, which also is a heavy user of xattr.
That one looks very different. In fact, the list that got corrupted
for you has since been changed to a hlist (which is *similar* to our
doubly-linked list, but has a smaller head and does not allow adding
to the end of the list).
Also, the "should be" and "was" values are very close, and switched:
should be ffffffff81ab3ca8, but was ffffffff81ab3cc8
should be ffffffff81ab3cc8, but was ffffffff81ab3ca8
so it actually looks like it was the same data structure. In
particular, it looks like enqueue_timer() ended up racing on adding an
entry to one index in the "base->vectors[]" array, while hitting an
entry that was pointing to another index near-by.
So I don't think it's related. Yours looks like some subtle timer base
race. It smells like a locking problem with timers. I'm not seeing
what it might be, but it *might* have been fixed by doing the
TIMER_MIGRATING bit right in add_timer_on() (commit 22b886dd1018).
Adding some timer people just in case, but I don't think your 4.1
report is related.
Linus
[toc] | [prev] | [next] | [standalone]
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2016-10-20 09:00 +0200 |
| Message-ID | <sufuy-3ZE-1@gated-at.bofh.it> |
| In reply to | #1504183 |
* Linus Torvalds <torvalds@linux-foundation.org> wrote: > On Wed, Oct 19, 2016 at 10:09 AM, Philipp Hahn <pmhahn@pmhahn.de> wrote: > > > > Nearly a month ago I reported also a "list_add corruption", but with 4.1.6: > > <http://marc.info/?l=linux-kernel&m=147508265316854&w=2> > > > > That server rungs Samba4, which also is a heavy user of xattr. > > That one looks very different. In fact, the list that got corrupted > for you has since been changed to a hlist (which is *similar* to our > doubly-linked list, but has a smaller head and does not allow adding > to the end of the list). > > Also, the "should be" and "was" values are very close, and switched: > > should be ffffffff81ab3ca8, but was ffffffff81ab3cc8 > should be ffffffff81ab3cc8, but was ffffffff81ab3ca8 > > so it actually looks like it was the same data structure. In > particular, it looks like enqueue_timer() ended up racing on adding an > entry to one index in the "base->vectors[]" array, while hitting an > entry that was pointing to another index near-by. > > So I don't think it's related. Yours looks like some subtle timer base > race. It smells like a locking problem with timers. I'm not seeing > what it might be, but it *might* have been fixed by doing the > TIMER_MIGRATING bit right in add_timer_on() (commit 22b886dd1018). > > Adding some timer people just in case, but I don't think your 4.1 > report is related. Side note: in case timer callback related corruption is suspected, a very efficient debugging method is to enable debugobjects tracking+checking: CONFIG_DEBUG_OBJECTS=y CONFIG_DEBUG_OBJECTS_SELFTEST=y CONFIG_DEBUG_OBJECTS_FREE=y CONFIG_DEBUG_OBJECTS_TIMERS=y CONFIG_DEBUG_OBJECTS_WORK=y CONFIG_DEBUG_OBJECTS_RCU_HEAD=y CONFIG_DEBUG_OBJECTS_PERCPU_COUNTER=y CONFIG_DEBUG_KOBJECT_RELEASE=y ( Appending these to the .config and running 'make oldconfig' should enable all of these. ) If the problem is in any of these areas then a debug warning could trigger at a more convenient place than some later 'some unrelated bits got corrupted' warning. Thanks, Ingo
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-10-20 09:30 +0200 |
| Message-ID | <sufXz-4qj-9@gated-at.bofh.it> |
| In reply to | #1504549 |
On Thu, 20 Oct 2016, Ingo Molnar wrote: > * Linus Torvalds <torvalds@linux-foundation.org> wrote: > > So I don't think it's related. Yours looks like some subtle timer base > > race. It smells like a locking problem with timers. I'm not seeing > > what it might be, but it *might* have been fixed by doing the > > TIMER_MIGRATING bit right in add_timer_on() (commit 22b886dd1018). 4.1 does not have the new fangled timer wheel. So 22b886dd1018 is irrelevant. > > Adding some timer people just in case, but I don't think your 4.1 > > report is related. This looks like one of the classic issues where an enqueued timer gets freed/reinitialized or otherwise manipulated. So yes, Ingos recommendation for debugging is the right thing to do. Thanks, tglx
[toc] | [prev] | [standalone]
Page 3 of 3 — ← Prev page 1 2 [3]
Back to top | Article view | linux.kernel
csiph-web