Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1542254 > unrolled thread
| Started by | Dave Chinner <david@fromorbit.com> |
|---|---|
| First post | 2016-12-14 23:30 +0100 |
| Last post | 2016-12-22 07:40 +0100 |
| Articles | 15 on this page of 35 — 10 participants |
Back to article view | Back to linux.kernel
[4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-14 23:30 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-14 23:40 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Chris Leech <cleech@redhat.com> - 2016-12-16 20:00 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-21 23:30 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Linus Torvalds <torvalds@linux-foundation.org> - 2016-12-22 00:30 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Chris Leech <cleech@redhat.com> - 2016-12-22 01:20 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-22 06:20 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Linus Torvalds <torvalds@linux-foundation.org> - 2016-12-22 06:50 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-22 08:00 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Chris Leech <cleech@redhat.com> - 2016-12-22 20:00 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Ming Lei <tom.leiming@gmail.com> - 2016-12-23 01:00 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Chris Leech <cleech@redhat.com> - 2016-12-23 01:10 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Christoph Hellwig <hch@lst.de> - 2016-12-23 11:10 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Linus Torvalds <torvalds@linux-foundation.org> - 2016-12-23 20:50 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Christoph Hellwig <hch@lst.de> - 2016-12-24 10:50 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Christoph Hellwig <hch@lst.de> - 2016-12-24 11:10 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Christoph Hellwig <hch@lst.de> - 2016-12-24 14:20 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Hannes Reinecke <hare@suse.de> - 2016-12-24 14:20 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Hugh Dickins <hughd@google.com> - 2016-12-22 21:30 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Johannes Weiner <hannes@cmpxchg.org> - 2016-12-23 08:40 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Johannes Weiner <hannes@cmpxchg.org> - 2016-12-23 09:40 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Johannes Weiner <hannes@cmpxchg.org> - 2017-01-02 22:20 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Jan Kara <jack@suse.cz> - 2017-01-03 13:30 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-22 07:30 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Linus Torvalds <torvalds@linux-foundation.org> - 2016-12-22 18:30 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Thomas Gleixner <tglx@linutronix.de> - 2016-12-22 21:30 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-22 21:50 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-22 22:10 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Linus Torvalds <torvalds@linux-foundation.org> - 2016-12-22 22:20 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-22 23:20 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-22 23:40 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-23 05:00 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Christoph Hellwig <hch@lst.de> - 2016-12-22 07:20 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Dave Chinner <david@fromorbit.com> - 2016-12-22 07:40 +0100
Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 Christoph Hellwig <hch@lst.de> - 2016-12-22 07:40 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2016-12-23 09:40 +0100 |
| Message-ID | <sRtyp-4vW-1@gated-at.bofh.it> |
| In reply to | #1546771 |
On Fri, Dec 23, 2016 at 02:32:41AM -0500, Johannes Weiner wrote: > On Thu, Dec 22, 2016 at 12:22:27PM -0800, Hugh Dickins wrote: > > On Wed, 21 Dec 2016, Linus Torvalds wrote: > > > On Wed, Dec 21, 2016 at 9:13 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > I unmounted the fs, mkfs'd it again, ran the > > > > workload again and about a minute in this fired: > > > > > > > > [628867.607417] ------------[ cut here ]------------ > > > > [628867.608603] WARNING: CPU: 2 PID: 16925 at mm/workingset.c:461 shadow_lru_isolate+0x171/0x220 > > > > > > Well, part of the changes during the merge window were the shadow > > > entry tracking changes that came in through Andrew's tree. Adding > > > Johannes Weiner to the participants. > > > > > > > Now, this workload does not touch the page cache at all - it's > > > > entirely an XFS metadata workload, so it should not really be > > > > affecting the working set code. > > > > > > Well, I suspect that anything that creates memory pressure will end up > > > triggering the working set code, so .. > > > > > > That said, obviously memory corruption could be involved and result in > > > random issues too, but I wouldn't really expect that in this code. > > > > > > It would probably be really useful to get more data points - is the > > > problem reliably in this area, or is it going to be random and all > > > over the place. > > > > Data point: kswapd got WARNING on mm/workingset.c:457 in shadow_lru_isolate, > > soon followed by NULL pointer deref in list_lru_isolate, one time when > > I tried out Sunday's git tree. Not seen since, I haven't had time to > > investigate, just set it aside as something to worry about if it happens > > again. But it looks like shadow_lru_isolate() has issues beyond Dave's > > case (I've no XFS and no iscsi), suspect unrelated to his other problems. > > This seems consistent with what Dave observed: we encounter regular > pages in radix tree nodes on the shadow LRU that should only contain > nodes full of exceptional shadow entries. It could be an issue in the > new slot replacement code and the node tracking callback. Both encounters seem to indicate use-after-free. Dave's node didn't warn about an unexpected node->count / node->exceptional state, but had entries that were inconsistent with that. Hugh got the counter warning but crashed on a list_head that's not NULLed in a live node. workingset_update_node() should be called on page cache radix tree leaf nodes that go empty. I must be missing an update_node callback where a leaf node gets freed somewhere.
[toc] | [prev] | [next] | [standalone]
| From | Johannes Weiner <hannes@cmpxchg.org> |
|---|---|
| Date | 2017-01-02 22:20 +0100 |
| Message-ID | <sVibo-4qO-11@gated-at.bofh.it> |
| In reply to | #1546778 |
On Fri, Dec 23, 2016 at 03:33:29AM -0500, Johannes Weiner wrote: > On Fri, Dec 23, 2016 at 02:32:41AM -0500, Johannes Weiner wrote: > > On Thu, Dec 22, 2016 at 12:22:27PM -0800, Hugh Dickins wrote: > > > On Wed, 21 Dec 2016, Linus Torvalds wrote: > > > > On Wed, Dec 21, 2016 at 9:13 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > > I unmounted the fs, mkfs'd it again, ran the > > > > > workload again and about a minute in this fired: > > > > > > > > > > [628867.607417] ------------[ cut here ]------------ > > > > > [628867.608603] WARNING: CPU: 2 PID: 16925 at mm/workingset.c:461 shadow_lru_isolate+0x171/0x220 > > > > > > > > Well, part of the changes during the merge window were the shadow > > > > entry tracking changes that came in through Andrew's tree. Adding > > > > Johannes Weiner to the participants. > > > > > > > > > Now, this workload does not touch the page cache at all - it's > > > > > entirely an XFS metadata workload, so it should not really be > > > > > affecting the working set code. > > > > > > > > Well, I suspect that anything that creates memory pressure will end up > > > > triggering the working set code, so .. > > > > > > > > That said, obviously memory corruption could be involved and result in > > > > random issues too, but I wouldn't really expect that in this code. > > > > > > > > It would probably be really useful to get more data points - is the > > > > problem reliably in this area, or is it going to be random and all > > > > over the place. > > > > > > Data point: kswapd got WARNING on mm/workingset.c:457 in shadow_lru_isolate, > > > soon followed by NULL pointer deref in list_lru_isolate, one time when > > > I tried out Sunday's git tree. Not seen since, I haven't had time to > > > investigate, just set it aside as something to worry about if it happens > > > again. But it looks like shadow_lru_isolate() has issues beyond Dave's > > > case (I've no XFS and no iscsi), suspect unrelated to his other problems. > > > > This seems consistent with what Dave observed: we encounter regular > > pages in radix tree nodes on the shadow LRU that should only contain > > nodes full of exceptional shadow entries. It could be an issue in the > > new slot replacement code and the node tracking callback. > > Both encounters seem to indicate use-after-free. Dave's node didn't > warn about an unexpected node->count / node->exceptional state, but > had entries that were inconsistent with that. Hugh got the counter > warning but crashed on a list_head that's not NULLed in a live node. > > workingset_update_node() should be called on page cache radix tree > leaf nodes that go empty. I must be missing an update_node callback > where a leaf node gets freed somewhere. Sorry for dropping silent on this. I'm traveling over the holidays with sporadic access to my emails and no access to real equipment. The times I managed to sneak away to look at the code didn't turn up anything useful yet. Andrea encountered the warning as well and I gave him a debugging patch (attached below), but he hasn't been able to reproduce this condition. I've personally never seen the warning trigger, even though the patches have been running on my main development machine for quite a while now. Albeit against an older base; I've updated to Linus's master branch now in case it's an interaction with other new code. If anybody manages to reproduce this, that would be helpful. Any extra eyes on this would be much appreciated too until I'm back at my desk. Thanks diff --git a/lib/radix-tree.c b/lib/radix-tree.c index 6f382e07de77..0783af1c0ebb 100644 --- a/lib/radix-tree.c +++ b/lib/radix-tree.c @@ -640,6 +640,8 @@ static inline void radix_tree_shrink(struct radix_tree_root *root, update_node(node, private); } + WARN_ON_ONCE(!list_empty(&node->private_list)); + radix_tree_node_free(node); } } @@ -666,6 +668,8 @@ static void delete_node(struct radix_tree_root *root, root->rnode = NULL; } + WARN_ON_ONCE(!list_empty(&node->private_list)); + radix_tree_node_free(node); node = parent; @@ -767,6 +771,7 @@ static void radix_tree_free_nodes(struct radix_tree_node *node) struct radix_tree_node *old = child; offset = child->offset + 1; child = child->parent; + WARN_ON_ONCE(!list_empty(&node->private_list)); radix_tree_node_free(old); if (old == entry_to_node(node)) return;
[toc] | [prev] | [next] | [standalone]
| From | Jan Kara <jack@suse.cz> |
|---|---|
| Date | 2017-01-03 13:30 +0100 |
| Message-ID | <sVwo2-64j-23@gated-at.bofh.it> |
| In reply to | #1549399 |
On Mon 02-01-17 16:11:36, Johannes Weiner wrote: > On Fri, Dec 23, 2016 at 03:33:29AM -0500, Johannes Weiner wrote: > > On Fri, Dec 23, 2016 at 02:32:41AM -0500, Johannes Weiner wrote: > > > On Thu, Dec 22, 2016 at 12:22:27PM -0800, Hugh Dickins wrote: > > > > On Wed, 21 Dec 2016, Linus Torvalds wrote: > > > > > On Wed, Dec 21, 2016 at 9:13 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > > > I unmounted the fs, mkfs'd it again, ran the > > > > > > workload again and about a minute in this fired: > > > > > > > > > > > > [628867.607417] ------------[ cut here ]------------ > > > > > > [628867.608603] WARNING: CPU: 2 PID: 16925 at mm/workingset.c:461 shadow_lru_isolate+0x171/0x220 > > > > > > > > > > Well, part of the changes during the merge window were the shadow > > > > > entry tracking changes that came in through Andrew's tree. Adding > > > > > Johannes Weiner to the participants. > > > > > > > > > > > Now, this workload does not touch the page cache at all - it's > > > > > > entirely an XFS metadata workload, so it should not really be > > > > > > affecting the working set code. > > > > > > > > > > Well, I suspect that anything that creates memory pressure will end up > > > > > triggering the working set code, so .. > > > > > > > > > > That said, obviously memory corruption could be involved and result in > > > > > random issues too, but I wouldn't really expect that in this code. > > > > > > > > > > It would probably be really useful to get more data points - is the > > > > > problem reliably in this area, or is it going to be random and all > > > > > over the place. > > > > > > > > Data point: kswapd got WARNING on mm/workingset.c:457 in shadow_lru_isolate, > > > > soon followed by NULL pointer deref in list_lru_isolate, one time when > > > > I tried out Sunday's git tree. Not seen since, I haven't had time to > > > > investigate, just set it aside as something to worry about if it happens > > > > again. But it looks like shadow_lru_isolate() has issues beyond Dave's > > > > case (I've no XFS and no iscsi), suspect unrelated to his other problems. > > > > > > This seems consistent with what Dave observed: we encounter regular > > > pages in radix tree nodes on the shadow LRU that should only contain > > > nodes full of exceptional shadow entries. It could be an issue in the > > > new slot replacement code and the node tracking callback. > > > > Both encounters seem to indicate use-after-free. Dave's node didn't > > warn about an unexpected node->count / node->exceptional state, but > > had entries that were inconsistent with that. Hugh got the counter > > warning but crashed on a list_head that's not NULLed in a live node. > > > > workingset_update_node() should be called on page cache radix tree > > leaf nodes that go empty. I must be missing an update_node callback > > where a leaf node gets freed somewhere. > > Sorry for dropping silent on this. I'm traveling over the holidays > with sporadic access to my emails and no access to real equipment. > > The times I managed to sneak away to look at the code didn't turn up > anything useful yet. > > Andrea encountered the warning as well and I gave him a debugging > patch (attached below), but he hasn't been able to reproduce this > condition. I've personally never seen the warning trigger, even though > the patches have been running on my main development machine for quite > a while now. Albeit against an older base; I've updated to Linus's > master branch now in case it's an interaction with other new code. > > If anybody manages to reproduce this, that would be helpful. Any extra > eyes on this would be much appreciated too until I'm back at my desk. I was looking into this but I didn't find a way how we could possibly leave radix tree node on LRU. So your debug patch looks like a good way forward. Honza -- Jan Kara <jack@suse.com> SUSE Labs, CR
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-12-22 07:30 +0100 |
| Message-ID | <sR533-5Hu-5@gated-at.bofh.it> |
| In reply to | #1546160 |
On Thu, Dec 22, 2016 at 04:13:22PM +1100, Dave Chinner wrote:
> On Wed, Dec 21, 2016 at 04:13:03PM -0800, Chris Leech wrote:
> > On Wed, Dec 21, 2016 at 03:19:15PM -0800, Linus Torvalds wrote:
> > > Hi,
> > >
> > > On Wed, Dec 21, 2016 at 2:16 PM, Dave Chinner <david@fromorbit.com> wrote:
> > > > On Fri, Dec 16, 2016 at 10:59:06AM -0800, Chris Leech wrote:
> > > >> Thanks Dave,
> > > >>
> > > >> I'm hitting a bug at scatterlist.h:140 before I even get any iSCSI
> > > >> modules loaded (virtio block) so there's something else going on in the
> > > >> current merge window. I'll keep an eye on it and make sure there's
> > > >> nothing iSCSI needs fixing for.
> > > >
> > > > OK, so before this slips through the cracks.....
> > > >
> > > > Linus - your tree as of a few minutes ago still panics immediately
> > > > when starting xfstests on iscsi devices. It appears to be a
> > > > scatterlist corruption and not an iscsi problem, so the iscsi guys
> > > > seem to have bounced it and no-one is looking at it.
> > >
> > > Hmm. There's not much to go by.
> > >
> > > Can somebody in iscsi-land please try to just bisect it - I'm not
> > > seeing a lot of clues to where this comes from otherwise.
> >
> > Yeah, my hopes of this being quickly resolved by someone else didn't
> > work out and whatever is going on in that test VM is looking like a
> > different kind of odd. I'm saving that off for later, and seeing if I
> > can't be a bisect on the iSCSI issue.
>
> There may be deeper issues. I just started running scalability tests
> (e.g. 16-way fsmark create tests) and about a minute in I got a
> directory corruption reported - something I hadn't seen in the dev
> cycle at all. I unmounted the fs, mkfs'd it again, ran the
> workload again and about a minute in this fired:
>
> [628867.607417] ------------[ cut here ]------------
> [628867.608603] WARNING: CPU: 2 PID: 16925 at mm/workingset.c:461 shadow_lru_isolate+0x171/0x220
> [628867.610702] Modules linked in:
> [628867.611375] CPU: 2 PID: 16925 Comm: kworker/2:97 Tainted: G W 4.9.0-dgc #18
> [628867.613382] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS Debian-1.8.2-1 04/01/2014
> [628867.616179] Workqueue: events rht_deferred_worker
> [628867.632422] Call Trace:
> [628867.634691] dump_stack+0x63/0x83
> [628867.637937] __warn+0xcb/0xf0
> [628867.641359] warn_slowpath_null+0x1d/0x20
> [628867.643362] shadow_lru_isolate+0x171/0x220
> [628867.644627] __list_lru_walk_one.isra.11+0x79/0x110
> [628867.645780] ? __list_lru_init+0x70/0x70
> [628867.646628] list_lru_walk_one+0x17/0x20
> [628867.647488] scan_shadow_nodes+0x34/0x50
> [628867.648358] shrink_slab.part.65.constprop.86+0x1dc/0x410
> [628867.649506] shrink_node+0x57/0x90
> [628867.650233] do_try_to_free_pages+0xdd/0x230
> [628867.651157] try_to_free_pages+0xce/0x1a0
> [628867.652342] __alloc_pages_slowpath+0x2df/0x960
> [628867.653332] ? __might_sleep+0x4a/0x80
> [628867.654148] __alloc_pages_nodemask+0x24b/0x290
> [628867.655237] kmalloc_order+0x21/0x50
> [628867.656016] kmalloc_order_trace+0x24/0xc0
> [628867.656878] __kmalloc+0x17d/0x1d0
> [628867.657644] bucket_table_alloc+0x195/0x1d0
> [628867.658564] ? __might_sleep+0x4a/0x80
> [628867.659449] rht_deferred_worker+0x287/0x3c0
> [628867.660366] ? _raw_spin_unlock_irq+0xe/0x30
> [628867.661294] process_one_work+0x1de/0x4d0
> [628867.662208] worker_thread+0x4b/0x4f0
> [628867.662990] kthread+0x10c/0x140
> [628867.663687] ? process_one_work+0x4d0/0x4d0
> [628867.664564] ? kthread_create_on_node+0x40/0x40
> [628867.665523] ret_from_fork+0x25/0x30
> [628867.666317] ---[ end trace 7c38634006a9955e ]---
>
> Now, this workload does not touch the page cache at all - it's
> entirely an XFS metadata workload, so it should not really be
> affecting the working set code.
The system back up, and I haven't reproduced this problem yet.
However, benchmark results are way off where they should be, and at
times the performance is utterly abysmal. The XFS for-next tree
based on the 4.9 kernel shows none of these problems, so I don't
think there's an XFS problem here. Workload is the same 16-way
fsmark workload that I've been using for years as a performance
regression test.
The workload normally averages around 230k files/s - i'm seeing
and average of ~175k files/s on you current kernel. And there are
periods where performance just completely tanks:
# ./fs_mark -D 10000 -S0 -n 100000 -s 0 -L 32 -d /mnt/scratch/0 -d /mnt/scratch/1 -d /mnt/scratch/2 -d /mnt/scratch/3 -d /mnt/scratch/4 -d /mnt/scratch/5 -d /mnt/scratch/6 -d /mnt/scratch/7 -d /mnt/scratch/8 -d /mnt/scratch/9 -d /mnt/scratch/10 -d /mnt/scratch/11 -d /mnt/scratch/12 -d /mnt/scratch/13 -d /mnt/scratch/14 -d /mnt/scratch/15
# Version 3.3, 16 thread(s) starting at Thu Dec 22 16:29:20 2016
# Sync method: NO SYNC: Test does not issue sync() or fsync() calls.
# Directories: Time based hash between directories across 10000 subdirectories with 180 seconds per subdirectory.
# File names: 40 bytes long, (16 initial bytes of time stamp with 24 random bytes at end of name)
# Files info: size 0 bytes, written with an IO size of 16384 bytes per write
# App overhead is time in microseconds spent in the test not doing file writing related system calls.
FSUse% Count Size Files/sec App Overhead
0 1600000 0 256964.5 10696017
0 3200000 0 244151.9 12129380
0 4800000 0 239929.1 11724413
0 6400000 0 235646.5 11998954
0 8000000 0 162537.3 15027053
0 9600000 0 186597.4 17957243
.....
0 43200000 0 184972.6 27911543
0 44800000 0 141642.2 62862142
0 46400000 0 137700.8 39018008
0 48000000 0 92155.0 38277234
0 49600000 0 57053.8 27810215
0 51200000 0 178941.6 33321543
This sort of thing is normally indicative of a memory reclaim or
lock contention problem. Profile showed unusual spinlock contention,
but then I realised there was only one kswapd thread running.
Yup, sure enough, it's caused by a major change in memory reclaim
behaviour:
[ 0.000000] Zone ranges:
[ 0.000000] DMA [mem 0x0000000000001000-0x0000000000ffffff]
[ 0.000000] DMA32 [mem 0x0000000001000000-0x00000000ffffffff]
[ 0.000000] Normal [mem 0x0000000100000000-0x000000083fffffff]
[ 0.000000] Movable zone start for each node
[ 0.000000] Early memory node ranges
[ 0.000000] node 0: [mem 0x0000000000001000-0x000000000009efff]
[ 0.000000] node 0: [mem 0x0000000000100000-0x00000000bffdefff]
[ 0.000000] node 0: [mem 0x0000000100000000-0x00000003bfffffff]
[ 0.000000] node 0: [mem 0x00000005c0000000-0x00000005ffffffff]
[ 0.000000] node 0: [mem 0x0000000800000000-0x000000083fffffff]
[ 0.000000] Initmem setup node 0 [mem 0x0000000000001000-0x000000083fffffff]
the numa=fake=4 CLI option is broken.
And, just to pile another one in - there's a massive xfs_repair
performance regression, too. It normally takes 4m30s to run over the
50 million inode filesystem created by fsmark. on 4.10:
XFS_REPAIR Summary Thu Dec 22 16:43:58 2016
Phase Start End Duration
Phase 1: 12/22 16:35:36 12/22 16:35:38 2 seconds
Phase 2: 12/22 16:35:38 12/22 16:35:40 2 seconds
Phase 3: 12/22 16:35:40 12/22 16:39:00 3 minutes, 20 seconds
Phase 4: 12/22 16:39:00 12/22 16:40:40 1 minute, 40 seconds
Phase 5: 12/22 16:40:40 12/22 16:40:40
Phase 6: 12/22 16:40:40 12/22 16:43:57 3 minutes, 17 seconds
Phase 7: 12/22 16:43:57 12/22 16:43:57
Total run time: 8 minutes, 21 seconds
So, yeah, there's lots of bad stuff I'm tripping over right now.
This is all now somebody els's problem to deal with - by this time
tomorrow I'm outta here and won't be back for several months....
Cheers,
Dave.
--
Dave Chinner
david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2016-12-22 18:30 +0100 |
| Subject | Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 |
| Message-ID | <sRflM-3Vq-43@gated-at.bofh.it> |
| In reply to | #1546193 |
On Wed, Dec 21, 2016 at 10:28 PM, Dave Chinner <david@fromorbit.com> wrote:
>
> This sort of thing is normally indicative of a memory reclaim or
> lock contention problem. Profile showed unusual spinlock contention,
> but then I realised there was only one kswapd thread running.
> Yup, sure enough, it's caused by a major change in memory reclaim
> behaviour:
>
> [ 0.000000] Zone ranges:
> [ 0.000000] DMA [mem 0x0000000000001000-0x0000000000ffffff]
> [ 0.000000] DMA32 [mem 0x0000000001000000-0x00000000ffffffff]
> [ 0.000000] Normal [mem 0x0000000100000000-0x000000083fffffff]
> [ 0.000000] Movable zone start for each node
> [ 0.000000] Early memory node ranges
> [ 0.000000] node 0: [mem 0x0000000000001000-0x000000000009efff]
> [ 0.000000] node 0: [mem 0x0000000000100000-0x00000000bffdefff]
> [ 0.000000] node 0: [mem 0x0000000100000000-0x00000003bfffffff]
> [ 0.000000] node 0: [mem 0x00000005c0000000-0x00000005ffffffff]
> [ 0.000000] node 0: [mem 0x0000000800000000-0x000000083fffffff]
> [ 0.000000] Initmem setup node 0 [mem 0x0000000000001000-0x000000083fffffff]
>
> the numa=fake=4 CLI option is broken.
Ok, I think that is independent of anything else. Removing block
people and adding the x86 people.
I'm not seeing anything at all that would change the fake numa stuff,
but maybe the cpu hotplug changes?
Thomas/Ingo/Peter - Dave is going away for several months, so you
won't get feedback from him, but can you look at this? Or maybe point
me towards the right people - I'm seeing no possible relevant changes
at all fir x85 numa since 4.9, so it must be some indirect breakage.
Dave is using fake-numa to do performance testing in a VM, and it's a
big deal for the node optimizations for writeback etc. Do you have any
ideas?
Dave, if you're still around, can you send out the kernel config file
you used...
Linus
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2016-12-22 21:30 +0100 |
| Message-ID | <sRi9Y-5GP-45@gated-at.bofh.it> |
| In reply to | #1546511 |
On Thu, 22 Dec 2016, Linus Torvalds wrote: > On Wed, Dec 21, 2016 at 10:28 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > This sort of thing is normally indicative of a memory reclaim or > > lock contention problem. Profile showed unusual spinlock contention, > > but then I realised there was only one kswapd thread running. > > Yup, sure enough, it's caused by a major change in memory reclaim > > behaviour: > > > > [ 0.000000] Zone ranges: > > [ 0.000000] DMA [mem 0x0000000000001000-0x0000000000ffffff] > > [ 0.000000] DMA32 [mem 0x0000000001000000-0x00000000ffffffff] > > [ 0.000000] Normal [mem 0x0000000100000000-0x000000083fffffff] > > [ 0.000000] Movable zone start for each node > > [ 0.000000] Early memory node ranges > > [ 0.000000] node 0: [mem 0x0000000000001000-0x000000000009efff] > > [ 0.000000] node 0: [mem 0x0000000000100000-0x00000000bffdefff] > > [ 0.000000] node 0: [mem 0x0000000100000000-0x00000003bfffffff] > > [ 0.000000] node 0: [mem 0x00000005c0000000-0x00000005ffffffff] > > [ 0.000000] node 0: [mem 0x0000000800000000-0x000000083fffffff] > > [ 0.000000] Initmem setup node 0 [mem 0x0000000000001000-0x000000083fffffff] > > > > the numa=fake=4 CLI option is broken. > > Ok, I think that is independent of anything else. Removing block > people and adding the x86 people. > > I'm not seeing anything at all that would change the fake numa stuff, > but maybe the cpu hotplug changes? Nope. Double checked it for correctness. The cpu hotplug code there is not involved in setting up kswapd threads. They are created by the memory node stuff. > Thomas/Ingo/Peter - Dave is going away for several months, so you > won't get feedback from him, but can you look at this? Or maybe point > me towards the right people - I'm seeing no possible relevant changes > at all fir x85 numa since 4.9, so it must be some indirect breakage. > > Dave is using fake-numa to do performance testing in a VM, and it's a > big deal for the node optimizations for writeback etc. Do you have any > ideas? There was the nodeid -> cpuid mapping stuff, but that was in 4.9 so it's not a post 4.9 wreckage, but maybe it's related nevertheless. > Dave, if you're still around, can you send out the kernel config file > you used... Config file and kernel command line would be helpful. Thanks, tglx
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-12-22 21:50 +0100 |
| Message-ID | <sRitj-5Nr-3@gated-at.bofh.it> |
| In reply to | #1546511 |
On Thu, Dec 22, 2016 at 09:24:12AM -0800, Linus Torvalds wrote: > On Wed, Dec 21, 2016 at 10:28 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > This sort of thing is normally indicative of a memory reclaim or > > lock contention problem. Profile showed unusual spinlock contention, > > but then I realised there was only one kswapd thread running. > > Yup, sure enough, it's caused by a major change in memory reclaim > > behaviour: > > > > [ 0.000000] Zone ranges: > > [ 0.000000] DMA [mem 0x0000000000001000-0x0000000000ffffff] > > [ 0.000000] DMA32 [mem 0x0000000001000000-0x00000000ffffffff] > > [ 0.000000] Normal [mem 0x0000000100000000-0x000000083fffffff] > > [ 0.000000] Movable zone start for each node > > [ 0.000000] Early memory node ranges > > [ 0.000000] node 0: [mem 0x0000000000001000-0x000000000009efff] > > [ 0.000000] node 0: [mem 0x0000000000100000-0x00000000bffdefff] > > [ 0.000000] node 0: [mem 0x0000000100000000-0x00000003bfffffff] > > [ 0.000000] node 0: [mem 0x00000005c0000000-0x00000005ffffffff] > > [ 0.000000] node 0: [mem 0x0000000800000000-0x000000083fffffff] > > [ 0.000000] Initmem setup node 0 [mem 0x0000000000001000-0x000000083fffffff] > > > > the numa=fake=4 CLI option is broken. > > Ok, I think that is independent of anything else. Removing block > people and adding the x86 people. > > I'm not seeing anything at all that would change the fake numa stuff, > but maybe the cpu hotplug changes? > > Thomas/Ingo/Peter - Dave is going away for several months, so you > won't get feedback from him, but can you look at this? Or maybe point > me towards the right people - I'm seeing no possible relevant changes > at all fir x85 numa since 4.9, so it must be some indirect breakage. > > Dave is using fake-numa to do performance testing in a VM, and it's a > big deal for the node optimizations for writeback etc. Do you have any > ideas? > > Dave, if you're still around, can you send out the kernel config file > you used... Looking at this fresh this morning (i.e. not pissed off by having everything I tried to do fail in different ways all afternoon) I found this: $ grep NUMA .config CONFIG_ARCH_SUPPORTS_NUMA_BALANCING=y # CONFIG_NUMA is not set $ The .config I was using for 4.9 got 'make oldconfig' upgraded, and looking at it there's a bunch of stuff that has been turned off that I know was set: # CONFIG_EXPERT is not set # CONFIG_PARAVIRT_SPINLOCKS is not set # CONFIG_COMPACTION is not set and stuff I never use so don't set was set, like kernel crash dump, a bunch of stuff for AMD CPUs, susp/resume and power management debug, every partition type and filesystem under the sun was selected, heaps of network devices enabled, etc. So it looks like the problem has occurred during oldconfig, meaning I have no idea exactly WTF I was testing. Rebuilding now with a saner config, see what happens. Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-12-22 22:10 +0100 |
| Message-ID | <sRiMG-69S-9@gated-at.bofh.it> |
| In reply to | #1546587 |
On Fri, Dec 23, 2016 at 07:42:40AM +1100, Dave Chinner wrote: > On Thu, Dec 22, 2016 at 09:24:12AM -0800, Linus Torvalds wrote: > > On Wed, Dec 21, 2016 at 10:28 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > > > This sort of thing is normally indicative of a memory reclaim or > > > lock contention problem. Profile showed unusual spinlock contention, > > > but then I realised there was only one kswapd thread running. > > > Yup, sure enough, it's caused by a major change in memory reclaim > > > behaviour: > > > > > > [ 0.000000] Zone ranges: > > > [ 0.000000] DMA [mem 0x0000000000001000-0x0000000000ffffff] > > > [ 0.000000] DMA32 [mem 0x0000000001000000-0x00000000ffffffff] > > > [ 0.000000] Normal [mem 0x0000000100000000-0x000000083fffffff] > > > [ 0.000000] Movable zone start for each node > > > [ 0.000000] Early memory node ranges > > > [ 0.000000] node 0: [mem 0x0000000000001000-0x000000000009efff] > > > [ 0.000000] node 0: [mem 0x0000000000100000-0x00000000bffdefff] > > > [ 0.000000] node 0: [mem 0x0000000100000000-0x00000003bfffffff] > > > [ 0.000000] node 0: [mem 0x00000005c0000000-0x00000005ffffffff] > > > [ 0.000000] node 0: [mem 0x0000000800000000-0x000000083fffffff] > > > [ 0.000000] Initmem setup node 0 [mem 0x0000000000001000-0x000000083fffffff] > > > > > > the numa=fake=4 CLI option is broken. > > > > Ok, I think that is independent of anything else. Removing block > > people and adding the x86 people. > > > > I'm not seeing anything at all that would change the fake numa stuff, > > but maybe the cpu hotplug changes? > > > > Thomas/Ingo/Peter - Dave is going away for several months, so you > > won't get feedback from him, but can you look at this? Or maybe point > > me towards the right people - I'm seeing no possible relevant changes > > at all fir x85 numa since 4.9, so it must be some indirect breakage. > > > > Dave is using fake-numa to do performance testing in a VM, and it's a > > big deal for the node optimizations for writeback etc. Do you have any > > ideas? > > > > Dave, if you're still around, can you send out the kernel config file > > you used... > > Looking at this fresh this morning (i.e. not pissed off by having > everything I tried to do fail in different ways all afternoon) I > found this: > > $ grep NUMA .config > CONFIG_ARCH_SUPPORTS_NUMA_BALANCING=y > # CONFIG_NUMA is not set > $ > > The .config I was using for 4.9 got 'make oldconfig' upgraded, and > looking at it there's a bunch of stuff that has been turned off that > I know was set: > > # CONFIG_EXPERT is not set > # CONFIG_PARAVIRT_SPINLOCKS is not set > # CONFIG_COMPACTION is not set > > and stuff I never use so don't set was set, like kernel crash dump, > a bunch of stuff for AMD CPUs, susp/resume and power management > debug, every partition type and filesystem under the sun was > selected, heaps of network devices enabled, etc. > > So it looks like the problem has occurred during oldconfig, meaning > I have no idea exactly WTF I was testing. Rebuilding now with a > saner config, see what happens. Better, but still bad. average files/s is not up to 200k files/s, so still a good 10-15% off where it should be. xfs_repair is back down to 10-15% off where it should be, too. bulkstat still fires off a bad page reference count warning, iscsi still panics immediately. Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2016-12-22 22:20 +0100 |
| Subject | Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 |
| Message-ID | <sRiWm-6ev-41@gated-at.bofh.it> |
| In reply to | #1546593 |
Ok, so the numa issue was a red herring. With that fixed:
On Thu, Dec 22, 2016 at 1:06 PM, Dave Chinner <david@fromorbit.com> wrote:
>
> Better, but still bad. average files/s is not up to 200k files/s,
> so still a good 10-15% off where it should be. xfs_repair is back
> down to 10-15% off where it should be, too. bulkstat still fires off
> a bad page reference count warning, iscsi still panics immediately.
Do you have CONFIG_BLK_WBT enabled, perhaps?
It's new to this merge window, and I'm not convinced it's been tuned.
Particularly for your kinds of fairly extreme IO loads.
Linus
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-12-22 23:20 +0100 |
| Message-ID | <sRjSp-6Nc-3@gated-at.bofh.it> |
| In reply to | #1546606 |
On Thu, Dec 22, 2016 at 01:10:19PM -0800, Linus Torvalds wrote: > Ok, so the numa issue was a red herring. With that fixed: > > On Thu, Dec 22, 2016 at 1:06 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > Better, but still bad. average files/s is not up to 200k files/s, > > so still a good 10-15% off where it should be. xfs_repair is back > > down to 10-15% off where it should be, too. bulkstat still fires off > > a bad page reference count warning, iscsi still panics immediately. > > Do you have CONFIG_BLK_WBT enabled, perhaps? Ok, yes, that's enabled. Let me go turn it off and see what happens. > It's new to this merge window, and I'm not convinced it's been tuned. > Particularly for your kinds of fairly extreme IO loads. Well, these aren't extreme IO workloads. Yes, the filesystem does a lot of work on the CPU, but the IO loads XFS generates aren't particularly high - maybe 2-3000 IOPS at most, and mostly the bandwidth is below 100MB/s. Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-12-22 23:40 +0100 |
| Message-ID | <sRkbM-6TR-43@gated-at.bofh.it> |
| In reply to | #1546621 |
On Fri, Dec 23, 2016 at 09:15:00AM +1100, Dave Chinner wrote:
> On Thu, Dec 22, 2016 at 01:10:19PM -0800, Linus Torvalds wrote:
> > Ok, so the numa issue was a red herring. With that fixed:
> >
> > On Thu, Dec 22, 2016 at 1:06 PM, Dave Chinner <david@fromorbit.com> wrote:
> > >
> > > Better, but still bad. average files/s is not up to 200k files/s,
> > > so still a good 10-15% off where it should be. xfs_repair is back
> > > down to 10-15% off where it should be, too. bulkstat still fires off
> > > a bad page reference count warning, iscsi still panics immediately.
> >
> > Do you have CONFIG_BLK_WBT enabled, perhaps?
>
> Ok, yes, that's enabled. Let me go turn it off and see what happens.
Numbers are still all over the place.
FSUse% Count Size Files/sec App Overhead
.....
0 28800000 0 228175.5 21928812
0 30400000 0 167880.5 39628229
0 32000000 0 124289.5 41420925
0 33600000 0 150577.9 35382318
0 35200000 0 216535.4 16072628
0 36800000 0 233414.4 11846654
0 38400000 0 213812.0 13356633
0 40000000 0 175905.7 53012015
0 41600000 0 157028.7 34700794
0 43200000 0 138829.1 50282461
And the average is now back down to 185k files/s. repair runtime is
unchanged and still 10-15% off...
I've got to run away for a few hours right now, but I'll retest the
4.9 + xfs for-next branch when I get back to see if the problem is
my curent config or whether there really is a perf problem lurking
somewhere....
Cheers,
Dave.
--
Dave Chinner
david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-12-23 05:00 +0100 |
| Message-ID | <sRpbr-1Cs-1@gated-at.bofh.it> |
| In reply to | #1546627 |
On Fri, Dec 23, 2016 at 09:33:36AM +1100, Dave Chinner wrote: > On Fri, Dec 23, 2016 at 09:15:00AM +1100, Dave Chinner wrote: > > On Thu, Dec 22, 2016 at 01:10:19PM -0800, Linus Torvalds wrote: > > > Ok, so the numa issue was a red herring. With that fixed: > > > > > > On Thu, Dec 22, 2016 at 1:06 PM, Dave Chinner <david@fromorbit.com> wrote: > > > > > > > > Better, but still bad. average files/s is not up to 200k files/s, > > > > so still a good 10-15% off where it should be. xfs_repair is back > > > > down to 10-15% off where it should be, too. bulkstat still fires off > > > > a bad page reference count warning, iscsi still panics immediately. > > > > > > Do you have CONFIG_BLK_WBT enabled, perhaps? > > > > Ok, yes, that's enabled. Let me go turn it off and see what happens. > > Numbers are still all over the place. > > FSUse% Count Size Files/sec App Overhead > ..... > 0 28800000 0 228175.5 21928812 > 0 30400000 0 167880.5 39628229 > 0 32000000 0 124289.5 41420925 > 0 33600000 0 150577.9 35382318 > 0 35200000 0 216535.4 16072628 > 0 36800000 0 233414.4 11846654 > 0 38400000 0 213812.0 13356633 > 0 40000000 0 175905.7 53012015 > 0 41600000 0 157028.7 34700794 > 0 43200000 0 138829.1 50282461 > > And the average is now back down to 185k files/s. repair runtime is > unchanged and still 10-15% off... > > I've got to run away for a few hours right now, but I'll retest the > 4.9 + xfs for-next branch when I get back to see if the problem is > my curent config or whether there really is a perf problem lurking > somewhere.... Well, I'm not sure now. Taking that config back to 4.9 gave results of 210k files/s. A bit faster, but still not the 230k files/s I'm expecting. So I'm missing something in the .config at this point, though I'm not sure what. FWIW, updating from 4.9 to to the 4.10 tree, this happened: $ $ make oldconfig; make -j 32 scripts/kconfig/conf --oldconfig Kconfig * * Restart config... * * * General setup * Yup, it definitely did something unexpected. And almost silently, I might add - I didn't notice this the first time around, and wouldn't hav enoticed it this time if I wasn't looking for something strange to happen. As iit is, I still haven't found what magic config option is taking away that 10-15% of performance. There's no unexpected debug options set, and it's nothing obvious in the fs/block layer config. I'll keep looking for the moment... Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@lst.de> |
|---|---|
| Date | 2016-12-22 07:20 +0100 |
| Subject | Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 |
| Message-ID | <sR4Tn-5Ej-1@gated-at.bofh.it> |
| In reply to | #1546049 |
On Wed, Dec 21, 2016 at 03:19:15PM -0800, Linus Torvalds wrote:
> Looking around a bit, the only even halfway suspicious scatterlist
> initialization thing I see is commit f9d03f96b988 ("block: improve
> handling of the magic discard payload") which used to have a magic
> hack wrt !bio->bi_vcnt, and that got removed. See __blk_bios_map_sg(),
> now it does __blk_bvec_map_sg() instead.
But that check was only for discard (and discard-like) bios which
had the maic single page that sometimes was unused attached.
For "normal" bios the for_each_segment loop iterates over bi_vcnt,
so it will be ignored anyway. That being said both I and the lists
got CCed halfway through the thread and I haven't seen the original
report, so I'm not really sure what's going on here anyway.
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-12-22 07:40 +0100 |
| Message-ID | <sR5cJ-5KK-9@gated-at.bofh.it> |
| In reply to | #1546189 |
On Thu, Dec 22, 2016 at 07:18:27AM +0100, Christoph Hellwig wrote:
> On Wed, Dec 21, 2016 at 03:19:15PM -0800, Linus Torvalds wrote:
> > Looking around a bit, the only even halfway suspicious scatterlist
> > initialization thing I see is commit f9d03f96b988 ("block: improve
> > handling of the magic discard payload") which used to have a magic
> > hack wrt !bio->bi_vcnt, and that got removed. See __blk_bios_map_sg(),
> > now it does __blk_bvec_map_sg() instead.
>
> But that check was only for discard (and discard-like) bios which
> had the maic single page that sometimes was unused attached.
>
> For "normal" bios the for_each_segment loop iterates over bi_vcnt,
> so it will be ignored anyway. That being said both I and the lists
> got CCed halfway through the thread and I haven't seen the original
> report, so I'm not really sure what's going on here anyway.
http://www.gossamer-threads.com/lists/linux/kernel/2587485
Cheers,
Dave.
--
Dave Chinner
david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@lst.de> |
|---|---|
| Date | 2016-12-22 07:40 +0100 |
| Subject | Re: [4.10, panic, regression] iscsi: null pointer deref at iscsi_tcp_segment_done+0x20d/0x2e0 |
| Message-ID | <sR5cJ-5KK-13@gated-at.bofh.it> |
| In reply to | #1546195 |
On Thu, Dec 22, 2016 at 05:30:46PM +1100, Dave Chinner wrote: > > For "normal" bios the for_each_segment loop iterates over bi_vcnt, > > so it will be ignored anyway. That being said both I and the lists > > got CCed halfway through the thread and I haven't seen the original > > report, so I'm not really sure what's going on here anyway. > > http://www.gossamer-threads.com/lists/linux/kernel/2587485 This doesn't look like the discard changes, but if Chris wants to test without them f9d03f96b988 reverts cleanly.
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web