Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1215914 > unrolled thread
| Started by | Josh Boyer <jwboyer@fedoraproject.org> |
|---|---|
| First post | 2015-08-30 14:40 +0200 |
| Last post | 2015-09-06 10:40 +0200 |
| Articles | 11 — 3 participants |
Back to article view | Back to linux.kernel
__blkg_lookup oops with 4.2-rcX Josh Boyer <jwboyer@fedoraproject.org> - 2015-08-30 14:40 +0200
Re: __blkg_lookup oops with 4.2-rcX "Richard W.M. Jones" <rjones@redhat.com> - 2015-08-30 20:10 +0200
Re: __blkg_lookup oops with 4.2-rcX Tejun Heo <tj@kernel.org> - 2015-09-02 17:00 +0200
Re: __blkg_lookup oops with 4.2-rcX Tejun Heo <tj@kernel.org> - 2015-09-02 17:40 +0200
Re: __blkg_lookup oops with 4.2-rcX "Richard W.M. Jones" <rjones@redhat.com> - 2015-09-04 12:50 +0200
Re: __blkg_lookup oops with 4.2-rcX Tejun Heo <tj@kernel.org> - 2015-09-04 19:20 +0200
Re: __blkg_lookup oops with 4.2-rcX "Richard W.M. Jones" <rjones@redhat.com> - 2015-09-04 22:50 +0200
Re: __blkg_lookup oops with 4.2-rcX "Richard W.M. Jones" <rjones@redhat.com> - 2015-09-05 17:50 +0200
Re: __blkg_lookup oops with 4.2-rcX Tejun Heo <tj@kernel.org> - 2015-09-05 20:40 +0200
[PATCH block/for-linus] block: blkg_destroy_all() should clear q->root_blkg and ->root_rl.blkg Tejun Heo <tj@kernel.org> - 2015-09-05 21:50 +0200
Re: [PATCH block/for-linus] block: blkg_destroy_all() should clear q->root_blkg and ->root_rl.blkg "Richard W.M. Jones" <rjones@redhat.com> - 2015-09-06 10:40 +0200
| From | Josh Boyer <jwboyer@fedoraproject.org> |
|---|---|
| Date | 2015-08-30 14:40 +0200 |
| Subject | __blkg_lookup oops with 4.2-rcX |
| Message-ID | <q3a3T-EP-11@gated-at.bofh.it> |
Hi Tejun, Mike and Jeff suggested you be informed of the oops one of our community members is hitting in Fedora with 4.2-rcX. I thought they had already sent this upstream to you, but apparently they didn't. The latest oops is below. That is with 4.2-rc8. I believe the first report was against a merge window 4.2 kernel. The full bug report is here: https://bugzilla.redhat.com/show_bug.cgi?id=1237136 I believe Mike and Jeff suspected the cgroup writeback patches. josh lvm vgchange -a n /run/lvm/lvmetad.socket: connect failed: No such file or directory WARNING: Failed to connect to lvmetad. Falling back to internal scanning. [ 36.157672] BUG: unable to handle kernel NULL pointer dereference at 0000000000000558 [ 36.157672] IP: [<ffffffff81389746>] __blkg_lookup+0x26/0x70 [ 36.157672] PGD 0 [ 36.157672] Oops: 0000 [#1] SMP [ 36.157672] Modules linked in: kvm_amd kvm snd_pcsp snd_pcm snd_timer snd soundcore serio_raw ata_generic pata_acpi libcrc32c crc8 crc_itu_t crc_ccitt virtio_pci virtio_mmio virtio_input virtio_balloon virtio_scsi sym53c8xx scsi_transport_spi megaraid_sas megaraid_mbox megaraid_mm megaraid ideapad_laptop rfkill sparse_keymap video virtio_net virtio_gpu ttm drm_kms_helper drm virtio_console virtio_rng virtio_blk virtio_ring virtio crc32 [ 36.157672] CPU: 0 PID: 248 Comm: lvm Not tainted 4.2.0-0.rc8.git0.1.fc23.x86_64 #1 [ 36.157672] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.8.2-20150714_191134- 04/01/2014 [ 36.157672] task: ffff88001b4e2580 ti: ffff88001ac0c000 task.ti: ffff88001ac0c000 [ 36.157672] RIP: 0010:[<ffffffff81389746>] [<ffffffff81389746>] __blkg_lookup+0x26/0x70 [ 36.157672] RSP: 0018:ffff88001ac0fa58 EFLAGS: 00000046 [ 36.157672] RAX: 0000000000000000 RBX: 0000000000000000 RCX: 0000000000000000 [ 36.157672] RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffffffff820176a0 [ 36.157672] RBP: ffff88001ac0fa78 R08: ffff88001ac0c000 R09: ffff88001ce6e800 [ 36.157672] R10: 0000000000000002 R11: 000000000001eaa5 R12: ffff88001cc57000 [ 36.157672] R13: ffff88001ccf99c8 R14: ffff88001ccf9f38 R15: ffff88001ccba8d8 [ 36.157672] FS: 00007f9a21711880(0000) GS:ffff88001f000000(0000) knlGS:0000000000000000 [ 36.157672] CS: 0010 DS: 0000 ES: 0000 CR0: 000000008005003b [ 36.157672] CR2: 0000000000000558 CR3: 000000001b41d000 CR4: 00000000000006f0 [ 36.157672] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000 [ 36.157672] DR3: 0000000000000000 DR6: 0000000000000000 DR7: 0000000000000000 [ 36.157672] Stack: [ 36.157672] ffffffff82017700 ffffffff820176a0 ffff88001cc57000 ffff88001ccf99c8 [ 36.157672] ffff88001ac0faa8 ffffffff8138d14a ffff88001cc570b8 ffff88001ccf99c8 [ 36.157672] ffffffff81cb1c20 0000000000000000 ffff88001ac0fab8 ffffffff8138a108 [ 36.157672] Call Trace: [ 36.157672] [<ffffffff8138d14a>] blk_throtl_drain+0x5a/0x110 [ 36.157672] [<ffffffff8138a108>] blkcg_drain_queue+0x18/0x20 [ 36.157672] [<ffffffff81369a70>] __blk_drain_queue+0xc0/0x170 [ 36.157672] [<ffffffff8136a101>] blk_queue_bypass_start+0x61/0x80 [ 36.157672] [<ffffffff81388c59>] blkcg_deactivate_policy+0x39/0x100 [ 36.157672] [<ffffffff8138d328>] blk_throtl_exit+0x38/0x50 [ 36.157672] [<ffffffff8138a14e>] blkcg_exit_queue+0x3e/0x50 [ 36.157672] [<ffffffff8137016e>] blk_release_queue+0x1e/0xc0 [ 36.157672] [<ffffffff8139bcba>] kobject_release+0x7a/0x190 [ 36.157672] [<ffffffff8139bb6f>] kobject_put+0x2f/0x60 [ 36.157672] [<ffffffff8136a2b1>] blk_cleanup_queue+0x111/0x140 [ 36.157672] [<ffffffff815f13fc>] cleanup_mapped_device+0xdc/0x100 [ 36.157672] [<ffffffff815f2311>] __dm_destroy+0x161/0x260 [ 36.157672] [<ffffffff815f45d3>] dm_destroy+0x13/0x20 [ 36.157672] [<ffffffff815f9ebd>] dev_remove+0x10d/0x170 [ 36.157672] [<ffffffff815f9db0>] ? dev_suspend+0x280/0x280 [ 36.157672] [<ffffffff815fa572>] ctl_ioctl+0x232/0x4d0 [ 36.157672] [<ffffffff8130cd80>] ? SYSC_semtimedop+0x2b0/0xeb0 [ 36.157672] [<ffffffff810136f1>] ? __switch_to+0x261/0x4b0 [ 36.157672] [<ffffffff815fa823>] dm_ctl_ioctl+0x13/0x20 [ 36.157672] [<ffffffff8122ebd5>] do_vfs_ioctl+0x295/0x470 [ 36.157672] [<ffffffff8130b259>] ? sem_security+0x9/0x10 [ 36.157672] [<ffffffff8122ee29>] SyS_ioctl+0x79/0x90 [ 36.157672] [<ffffffff817750ae>] entry_SYSCALL_64_fastpath+0x12/0x71 [ 36.157672] Code: eb bf 0f 1f 00 66 66 66 66 90 55 48 89 e5 41 55 41 54 53 48 83 ec 08 48 8b 87 c8 00 00 00 48 85 c0 74 05 48 39 30 74 45 48 89 f3 <48> 63 b6 58 05 00 00 49 89 fd 48 8d bf b8 00 00 00 41 89 d4 e8 [ 36.157672] RIP [<ffffffff81389746>] __blkg_lookup+0x26/0x70 [ 36.157672] RSP <ffff88001ac0fa58> [ 36.157672] CR2: 0000000000000558 [ 36.157672] ---[ end trace a6310b2924d6c01e ]--- -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | "Richard W.M. Jones" <rjones@redhat.com> |
|---|---|
| Date | 2015-08-30 20:10 +0200 |
| Message-ID | <q3fdf-8dd-1@gated-at.bofh.it> |
| In reply to | #1215914 |
On Sun, Aug 30, 2015 at 08:30:41AM -0400, Josh Boyer wrote: > Hi Tejun, > > Mike and Jeff suggested you be informed of the oops one of our > community members is hitting in Fedora with 4.2-rcX. I thought they > had already sent this upstream to you, but apparently they didn't. > > The latest oops is below. That is with 4.2-rc8. I believe the first > report was against a merge window 4.2 kernel. The full bug report is > here: https://bugzilla.redhat.com/show_bug.cgi?id=1237136 > > I believe Mike and Jeff suspected the cgroup writeback patches. Thanks Josh. Also, I can test potential patches if you CC me on them. Rich. > josh > > lvm vgchange -a n > /run/lvm/lvmetad.socket: connect failed: No such file or directory > WARNING: Failed to connect to lvmetad. Falling back to internal scanning. > [ 36.157672] BUG: unable to handle kernel NULL pointer dereference > at 0000000000000558 > [ 36.157672] IP: [<ffffffff81389746>] __blkg_lookup+0x26/0x70 > [ 36.157672] PGD 0 > [ 36.157672] Oops: 0000 [#1] SMP > [ 36.157672] Modules linked in: kvm_amd kvm snd_pcsp snd_pcm > snd_timer snd soundcore serio_raw ata_generic pata_acpi libcrc32c crc8 > crc_itu_t crc_ccitt virtio_pci virtio_mmio virtio_input virtio_balloon > virtio_scsi sym53c8xx scsi_transport_spi megaraid_sas megaraid_mbox > megaraid_mm megaraid ideapad_laptop rfkill sparse_keymap video > virtio_net virtio_gpu ttm drm_kms_helper drm virtio_console virtio_rng > virtio_blk virtio_ring virtio crc32 > [ 36.157672] CPU: 0 PID: 248 Comm: lvm Not tainted > 4.2.0-0.rc8.git0.1.fc23.x86_64 #1 > [ 36.157672] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), > BIOS 1.8.2-20150714_191134- 04/01/2014 > [ 36.157672] task: ffff88001b4e2580 ti: ffff88001ac0c000 task.ti: > ffff88001ac0c000 > [ 36.157672] RIP: 0010:[<ffffffff81389746>] [<ffffffff81389746>] > __blkg_lookup+0x26/0x70 > [ 36.157672] RSP: 0018:ffff88001ac0fa58 EFLAGS: 00000046 > [ 36.157672] RAX: 0000000000000000 RBX: 0000000000000000 RCX: 0000000000000000 > [ 36.157672] RDX: 0000000000000000 RSI: 0000000000000000 RDI: ffffffff820176a0 > [ 36.157672] RBP: ffff88001ac0fa78 R08: ffff88001ac0c000 R09: ffff88001ce6e800 > [ 36.157672] R10: 0000000000000002 R11: 000000000001eaa5 R12: ffff88001cc57000 > [ 36.157672] R13: ffff88001ccf99c8 R14: ffff88001ccf9f38 R15: ffff88001ccba8d8 > [ 36.157672] FS: 00007f9a21711880(0000) GS:ffff88001f000000(0000) > knlGS:0000000000000000 > [ 36.157672] CS: 0010 DS: 0000 ES: 0000 CR0: 000000008005003b > [ 36.157672] CR2: 0000000000000558 CR3: 000000001b41d000 CR4: 00000000000006f0 > [ 36.157672] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000 > [ 36.157672] DR3: 0000000000000000 DR6: 0000000000000000 DR7: 0000000000000000 > [ 36.157672] Stack: > [ 36.157672] ffffffff82017700 ffffffff820176a0 ffff88001cc57000 > ffff88001ccf99c8 > [ 36.157672] ffff88001ac0faa8 ffffffff8138d14a ffff88001cc570b8 > ffff88001ccf99c8 > [ 36.157672] ffffffff81cb1c20 0000000000000000 ffff88001ac0fab8 > ffffffff8138a108 > [ 36.157672] Call Trace: > [ 36.157672] [<ffffffff8138d14a>] blk_throtl_drain+0x5a/0x110 > [ 36.157672] [<ffffffff8138a108>] blkcg_drain_queue+0x18/0x20 > [ 36.157672] [<ffffffff81369a70>] __blk_drain_queue+0xc0/0x170 > [ 36.157672] [<ffffffff8136a101>] blk_queue_bypass_start+0x61/0x80 > [ 36.157672] [<ffffffff81388c59>] blkcg_deactivate_policy+0x39/0x100 > [ 36.157672] [<ffffffff8138d328>] blk_throtl_exit+0x38/0x50 > [ 36.157672] [<ffffffff8138a14e>] blkcg_exit_queue+0x3e/0x50 > [ 36.157672] [<ffffffff8137016e>] blk_release_queue+0x1e/0xc0 > [ 36.157672] [<ffffffff8139bcba>] kobject_release+0x7a/0x190 > [ 36.157672] [<ffffffff8139bb6f>] kobject_put+0x2f/0x60 > [ 36.157672] [<ffffffff8136a2b1>] blk_cleanup_queue+0x111/0x140 > [ 36.157672] [<ffffffff815f13fc>] cleanup_mapped_device+0xdc/0x100 > [ 36.157672] [<ffffffff815f2311>] __dm_destroy+0x161/0x260 > [ 36.157672] [<ffffffff815f45d3>] dm_destroy+0x13/0x20 > [ 36.157672] [<ffffffff815f9ebd>] dev_remove+0x10d/0x170 > [ 36.157672] [<ffffffff815f9db0>] ? dev_suspend+0x280/0x280 > [ 36.157672] [<ffffffff815fa572>] ctl_ioctl+0x232/0x4d0 > [ 36.157672] [<ffffffff8130cd80>] ? SYSC_semtimedop+0x2b0/0xeb0 > [ 36.157672] [<ffffffff810136f1>] ? __switch_to+0x261/0x4b0 > [ 36.157672] [<ffffffff815fa823>] dm_ctl_ioctl+0x13/0x20 > [ 36.157672] [<ffffffff8122ebd5>] do_vfs_ioctl+0x295/0x470 > [ 36.157672] [<ffffffff8130b259>] ? sem_security+0x9/0x10 > [ 36.157672] [<ffffffff8122ee29>] SyS_ioctl+0x79/0x90 > [ 36.157672] [<ffffffff817750ae>] entry_SYSCALL_64_fastpath+0x12/0x71 > [ 36.157672] Code: eb bf 0f 1f 00 66 66 66 66 90 55 48 89 e5 41 55 > 41 54 53 48 83 ec 08 48 8b 87 c8 00 00 00 48 85 c0 74 05 48 39 30 74 > 45 48 89 f3 <48> 63 b6 58 05 00 00 49 89 fd 48 8d bf b8 00 00 00 41 89 > d4 e8 > [ 36.157672] RIP [<ffffffff81389746>] __blkg_lookup+0x26/0x70 > [ 36.157672] RSP <ffff88001ac0fa58> > [ 36.157672] CR2: 0000000000000558 > [ 36.157672] ---[ end trace a6310b2924d6c01e ]--- -- Richard Jones, Virtualization Group, Red Hat http://people.redhat.com/~rjones Read my programming and virtualization blog: http://rwmj.wordpress.com libguestfs lets you edit virtual machines. Supports shell scripting, bindings from many languages. http://libguestfs.org -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-02 17:00 +0200 |
| Message-ID | <q4hG2-8ct-3@gated-at.bofh.it> |
| In reply to | #1215914 |
Hello,
On Sun, Aug 30, 2015 at 08:30:41AM -0400, Josh Boyer wrote:
> Mike and Jeff suggested you be informed of the oops one of our
> community members is hitting in Fedora with 4.2-rcX. I thought they
> had already sent this upstream to you, but apparently they didn't.
>
> The latest oops is below. That is with 4.2-rc8. I believe the first
> report was against a merge window 4.2 kernel. The full bug report is
> here: https://bugzilla.redhat.com/show_bug.cgi?id=1237136
>
> I believe Mike and Jeff suspected the cgroup writeback patches.
>
> josh
>
> lvm vgchange -a n
> /run/lvm/lvmetad.socket: connect failed: No such file or directory
> WARNING: Failed to connect to lvmetad. Falling back to internal scanning.
> [ 36.157672] BUG: unable to handle kernel NULL pointer dereference
> at 0000000000000558
> [ 36.157672] IP: [<ffffffff81389746>] __blkg_lookup+0x26/0x70
...
> [ 36.157672] [<ffffffff8138d14a>] blk_throtl_drain+0x5a/0x110
> [ 36.157672] [<ffffffff8138a108>] blkcg_drain_queue+0x18/0x20
> [ 36.157672] [<ffffffff81369a70>] __blk_drain_queue+0xc0/0x170
> [ 36.157672] [<ffffffff8136a101>] blk_queue_bypass_start+0x61/0x80
> [ 36.157672] [<ffffffff81388c59>] blkcg_deactivate_policy+0x39/0x100
> [ 36.157672] [<ffffffff8138d328>] blk_throtl_exit+0x38/0x50
> [ 36.157672] [<ffffffff8138a14e>] blkcg_exit_queue+0x3e/0x50
> [ 36.157672] [<ffffffff8137016e>] blk_release_queue+0x1e/0xc0
> [ 36.157672] [<ffffffff8139bcba>] kobject_release+0x7a/0x190
> [ 36.157672] [<ffffffff8139bb6f>] kobject_put+0x2f/0x60
> [ 36.157672] [<ffffffff8136a2b1>] blk_cleanup_queue+0x111/0x140
> [ 36.157672] [<ffffffff815f13fc>] cleanup_mapped_device+0xdc/0x100
> [ 36.157672] [<ffffffff815f2311>] __dm_destroy+0x161/0x260
> [ 36.157672] [<ffffffff815f45d3>] dm_destroy+0x13/0x20
> [ 36.157672] [<ffffffff815f9ebd>] dev_remove+0x10d/0x170
> [ 36.157672] [<ffffffff815fa572>] ctl_ioctl+0x232/0x4d0
> [ 36.157672] [<ffffffff815fa823>] dm_ctl_ioctl+0x13/0x20
> [ 36.157672] [<ffffffff8122ebd5>] do_vfs_ioctl+0x295/0x470
> [ 36.157672] [<ffffffff8122ee29>] SyS_ioctl+0x79/0x90
> [ 36.157672] [<ffffffff817750ae>] entry_SYSCALL_64_fastpath+0x12/0x71
I think the offending commit is 776687bce42b ("block, blk-mq: draining
can't be skipped even if bypass_depth was non-zero"). It looks like
the patch makes shutdown path travel data structure which is already
destroyed. Will post the fix soon.
Thanks.
--
tejun
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-02 17:40 +0200 |
| Message-ID | <q4iiK-JG-5@gated-at.bofh.it> |
| In reply to | #1217683 |
Hello,
On Wed, Sep 02, 2015 at 10:53:07AM -0400, Tejun Heo wrote:
> On Sun, Aug 30, 2015 at 08:30:41AM -0400, Josh Boyer wrote:
> I think the offending commit is 776687bce42b ("block, blk-mq: draining
> can't be skipped even if bypass_depth was non-zero"). It looks like
> the patch makes shutdown path travel data structure which is already
> destroyed. Will post the fix soon.
Hmm... I can't reproduce it here or see how such oops would happen.
* Is the problem reproducible on v4.2? If so, can you please describe
the steps to reproduce? How is cgroup set up?
* Can you please run gdb or addr2line on it and report which line is
causing the oops?
Thanks.
--
tejun
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Richard W.M. Jones" <rjones@redhat.com> |
|---|---|
| Date | 2015-09-04 12:50 +0200 |
| Message-ID | <q4WJb-86i-11@gated-at.bofh.it> |
| In reply to | #1217704 |
On Wed, Sep 02, 2015 at 11:32:55AM -0400, Tejun Heo wrote:
> Hello,
>
> On Wed, Sep 02, 2015 at 10:53:07AM -0400, Tejun Heo wrote:
> > On Sun, Aug 30, 2015 at 08:30:41AM -0400, Josh Boyer wrote:
> > I think the offending commit is 776687bce42b ("block, blk-mq: draining
> > can't be skipped even if bypass_depth was non-zero"). It looks like
> > the patch makes shutdown path travel data structure which is already
> > destroyed. Will post the fix soon.
>
> Hmm... I can't reproduce it here or see how such oops would happen.
>
> * Is the problem reproducible on v4.2? If so, can you please describe
> the steps to reproduce? How is cgroup set up?
We have a test suite which does a lot of filesystem and device
operations, and this triggers it randomly (not reliably nor in the
same place every time, but still pretty frequently).
So .. I don't have steps that can reproduce it reliably unfortunately.
However I'm going to work on that now to see if I can create a
sequence of operations that triggers it some or all of the time.
> * Can you please run gdb or addr2line on it and report which line is
> causing the oops?
Below is another stack trace that I just collected. It came from a
test that does some hotplugging of a virtual machine. The kernel this
time is 4.2.0-0.rc3.git4.1.fc24.x86_64 (which is a bit old - am also
going to upgrade to the newest kernel soon).
The addr2line output from this one is:
$ addr2line -e /usr/lib/debug/lib/modules/4.2.0-0.rc3.git4.1.fc24.x86_64/vmlinux ffffffff814107a0
/usr/src/debug/kernel-4.1.fc24/linux-4.2.0-0.rc3.git4.1.fc24.x86_64/block/blk-throttle.c:1642
1636 /*
1637 * Drain each tg while doing post-order walk on the blkg tree, s 1637 o
1638 * that all bios are propagated to td->service_queue. It'd be
1639 * better to walk service_queue tree directly but blkg walk is
1640 * easier.
1641 */
1642 blkg_for_each_descendant_post(blkg, pos_css, td->queue->root_blkg)
1643 tg_drain_bios(&blkg_to_tg(blkg)->service_queue);
1644
Rich.
[ 6.784689] BUG: unable to handle kernel NULL pointer dereference at 0000000000000bb8
[ 6.787605] IP: [<ffffffff814107a0>] blk_throtl_drain+0x80/0x220
[ 6.789797] PGD 0
[ 6.790598] Oops: 0000 [#1] SMP
[ 6.791848] Modules linked in: kvm_intel kvm snd_pcsp snd_pcm snd_timer snd ghash_clmulni_intel soundcore joydev ata_generic serio_raw pata_acpi libcrc32c crc8 crc_itu_t crc_ccitt virtio_pci virtio_mmio virtio_input virtio_balloon virtio_scsi sym53c8xx scsi_transport_spi megaraid_sas megaraid_mbox megaraid_mm megaraid ideapad_laptop rfkill sparse_keymap video virtio_net virtio_gpu ttm drm_kms_helper drm virtio_console virtio_rng virtio_blk virtio_ring virtio crc32 crct10dif_pclmul crc32c_intel crc32_pclmul
[ 6.809710] CPU: 0 PID: 27 Comm: kworker/0:1 Not tainted 4.2.0-0.rc3.git4.1.fc24.x86_64 #1
[ 6.812650] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.8.2-20150714_191134- 04/01/2014
[ 6.816068] Workqueue: events_freezable virtscsi_handle_event [virtio_scsi]
[ 6.818588] task: ffff88001dfb3a00 ti: ffff88001d090000 task.ti: ffff88001d090000
[ 6.821252] RIP: 0010:[<ffffffff814107a0>] [<ffffffff814107a0>] blk_throtl_drain+0x80/0x220
[ 6.824302] RSP: 0000:ffff88001d0939d8 EFLAGS: 00010046
[ 6.826213] RAX: 0000000000000000 RBX: ffff88001b8f6698 RCX: 00000000000000e0
[ 6.828743] RDX: 31e18f88fc458000 RSI: 0000000000000000 RDI: 0000000000000000
[ 6.831292] RBP: ffff88001d093a08 R08: 0000000000000000 R09: 0000000000000000
[ 6.833835] R10: ffff88001dfb3a00 R11: ffffffff81e58200 R12: ffff88001ba67200
[ 6.836380] R13: ffff88001b8f6698 R14: ffff88001b9ee1f0 R15: ffff88001b9ee0d0
[ 6.838920] FS: 0000000000000000(0000) GS:ffff88001ee00000(0000) knlGS:0000000000000000
[ 6.841781] CS: 0010 DS: 0000 ES: 0000 CR0: 000000008005003b
[ 6.843838] CR2: 0000000000000bb8 CR3: 00000000180c4000 CR4: 00000000000006f0
[ 6.846383] Stack:
[ 6.847132] ffffffff81410756 ffff88001b9ee1f0 ffff88001d093a08 ffff88001b8f6698
[ 6.849950] ffffffff81ef5320 0000000000000000 ffff88001d093a28 ffffffff8140d5fd
[ 6.852746] ffff88001b8f6698 ffff88001b8f6698 ffff88001d093a58 ffffffff813e7839
[ 6.855562] Call Trace:
[ 6.856473] [<ffffffff81410756>] ? blk_throtl_drain+0x36/0x220
[ 6.858581] [<ffffffff8140d5fd>] blkcg_drain_queue+0x2d/0x60
[ 6.860639] [<ffffffff813e7839>] __blk_drain_queue+0xc9/0x1a0
[ 6.862741] [<ffffffff813e9218>] ? blk_queue_bypass_start+0x68/0xb0
[ 6.865029] [<ffffffff813e9222>] blk_queue_bypass_start+0x72/0xb0
[ 6.867236] [<ffffffff8140b539>] blkcg_deactivate_policy+0x39/0x100
[ 6.869513] [<ffffffff814173e0>] cfq_exit_queue+0xd0/0xf0
[ 6.871481] [<ffffffff813e5081>] elevator_exit+0x31/0x50
[ 6.873423] [<ffffffff813ef91e>] blk_release_queue+0x4e/0xc0
[ 6.875495] [<ffffffff814204aa>] kobject_release+0x7a/0x190
[ 6.877524] [<ffffffff8142035f>] kobject_put+0x2f/0x60
[ 6.879413] [<ffffffff813e7765>] blk_put_queue+0x15/0x20
[ 6.881351] [<ffffffff815bf324>] scsi_device_dev_release_usercontext+0xc4/0x120
[ 6.884010] [<ffffffff815bf260>] ? scsi_device_dev_release+0x20/0x20
[ 6.886297] [<ffffffff810cad3c>] execute_in_process_context+0x9c/0xb0
[ 6.888636] [<ffffffff815bf25c>] scsi_device_dev_release+0x1c/0x20
[ 6.890897] [<ffffffff81573706>] device_release+0x36/0xa0
[ 6.892867] [<ffffffff814204aa>] kobject_release+0x7a/0x190
[ 6.894901] [<ffffffff8142035f>] kobject_put+0x2f/0x60
[ 6.896772] [<ffffffff81573a47>] put_device+0x17/0x20
[ 6.898617] [<ffffffff815b050f>] scsi_device_put+0x2f/0x40
[ 6.900614] [<ffffffffa0155f61>] virtscsi_handle_event+0x101/0x1a0 [virtio_scsi]
[ 6.903284] [<ffffffff810cb3b2>] process_one_work+0x232/0x840
[ 6.905380] [<ffffffff810cb31b>] ? process_one_work+0x19b/0x840
[ 6.907522] [<ffffffff8112553d>] ? debug_lockdep_rcu_enabled+0x1d/0x20
[ 6.909893] [<ffffffff810cba95>] ? worker_thread+0xd5/0x450
[ 6.911921] [<ffffffff810cba0e>] worker_thread+0x4e/0x450
[ 6.913902] [<ffffffff810cb9c0>] ? process_one_work+0x840/0x840
[ 6.916066] [<ffffffff810cb9c0>] ? process_one_work+0x840/0x840
[ 6.918232] [<ffffffff810d2594>] kthread+0x104/0x120
[ 6.920059] [<ffffffff810d2490>] ? kthread_create_on_node+0x250/0x250
[ 6.922396] [<ffffffff8187105f>] ret_from_fork+0x3f/0x70
[ 6.924339] [<ffffffff810d2490>] ? kthread_create_on_node+0x250/0x250
[ 6.926663] Code: 04 24 56 07 41 81 e8 20 72 cf ff e8 9b 4d d1 ff 85 c0 74 0d 80 3d 64 04 b5 00 00 0f 84 19 01 00 00 49 8b 84 24 d0 00 00 00 31 ff <48> 8b 80 b8 0b 00 00 48 8b 70 28 e8 60 04 d5 ff 48 85 c0 48 89
[ 6.936207] RIP [<ffffffff814107a0>] blk_throtl_drain+0x80/0x220
[ 6.938432] RSP <ffff88001d0939d8>
[ 6.939692] CR2: 0000000000000bb8
[ 6.940915] ---[ end trace f1acb54c2a225dd4 ]---
Rich.
--
Richard Jones, Virtualization Group, Red Hat http://people.redhat.com/~rjones
Read my programming and virtualization blog: http://rwmj.wordpress.com
virt-p2v converts physical machines to virtual machines. Boot with a
live CD or over the network (PXE) and turn machines into KVM guests.
http://libguestfs.org/virt-v2v
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-04 19:20 +0200 |
| Message-ID | <q52OB-53-1@gated-at.bofh.it> |
| In reply to | #1218817 |
Hello,
On Fri, Sep 04, 2015 at 11:46:02AM +0100, Richard W.M. Jones wrote:
> $ addr2line -e /usr/lib/debug/lib/modules/4.2.0-0.rc3.git4.1.fc24.x86_64/vmlinux ffffffff814107a0
> /usr/src/debug/kernel-4.1.fc24/linux-4.2.0-0.rc3.git4.1.fc24.x86_64/block/blk-throttle.c:1642
>
> 1636 /*
> 1637 * Drain each tg while doing post-order walk on the blkg tree, s 1637 o
> 1638 * that all bios are propagated to td->service_queue. It'd be
> 1639 * better to walk service_queue tree directly but blkg walk is
> 1640 * easier.
> 1641 */
> 1642 blkg_for_each_descendant_post(blkg, pos_css, td->queue->root_blkg)
> 1643 tg_drain_bios(&blkg_to_tg(blkg)->service_queue);
> 1644
>
> Rich.
>
> [ 6.784689] BUG: unable to handle kernel NULL pointer dereference at 0000000000000bb8
> [ 6.787605] IP: [<ffffffff814107a0>] blk_throtl_drain+0x80/0x220
The only struct which is large enough for 0xbb8 offset is
request_queue. Hmm.... can you please try the brute force debug patch
below and report the kernel log after the crash?
Thanks.
diff --git a/block/blk-throttle.c b/block/blk-throttle.c
index b231935..09426e4 100644
--- a/block/blk-throttle.c
+++ b/block/blk-throttle.c
@@ -1639,8 +1639,22 @@ void blk_throtl_drain(struct request_queue *q)
* better to walk service_queue tree directly but blkg walk is
* easier.
*/
- blkg_for_each_descendant_post(blkg, pos_css, td->queue->root_blkg)
- tg_drain_bios(&blkg_to_tg(blkg)->service_queue);
+ printk("XXX blk_throtl_drain: td=%p ->queue=%p ->root_blkg=%p ->q/blkcg=%p/%p\n",
+ td, td ? td->queue : NULL,
+ (td && td->queue) ? td->queue->root_blkg : NULL,
+ (td && td->queue && td->queue->root_blkg) ? td->queue->root_blkg->q : NULL,
+ (td && td->queue && td->queue->root_blkg) ? td->queue->root_blkg->blkcg : NULL);
+
+ css_for_each_descendant_pre(pos_css, &td->queue->root_blkg->blkcg->css) {
+ printk("XXX pos_css=%p ", pos_css);
+ pr_cont_cgroup_path(pos_css->cgroup);
+ if ((blkg = __blkg_lookup(css_to_blkcg(pos_css),
+ td->queue->root_blkg->q, false))) {
+ pr_cont(" blkg=%p", blkg);
+ tg_drain_bios(&blkg_to_tg(blkg)->service_queue);
+ }
+ pr_cont("\n");
+ }
/* finally, transfer bios from top-level tg's into the td */
tg_drain_bios(&td->service_queue);
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Richard W.M. Jones" <rjones@redhat.com> |
|---|---|
| Date | 2015-09-04 22:50 +0200 |
| Message-ID | <q565Q-4De-13@gated-at.bofh.it> |
| In reply to | #1219215 |
On Fri, Sep 04, 2015 at 01:13:02PM -0400, Tejun Heo wrote: > The only struct which is large enough for 0xbb8 offset is > request_queue. Hmm.... can you please try the brute force debug patch > below and report the kernel log after the crash? So the good(?) news is this bug is not reproducible with the Fedora kernel 4.3.0-0.rc0.git7.1.fc24.x86_64 (tested both with and without your patch). I'll keep it running overnight just in case .. Rich. -- Richard Jones, Virtualization Group, Red Hat http://people.redhat.com/~rjones Read my programming and virtualization blog: http://rwmj.wordpress.com libguestfs lets you edit virtual machines. Supports shell scripting, bindings from many languages. http://libguestfs.org -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Richard W.M. Jones" <rjones@redhat.com> |
|---|---|
| Date | 2015-09-05 17:50 +0200 |
| Message-ID | <q5nT4-4L2-19@gated-at.bofh.it> |
| In reply to | #1219307 |
On Sat, Sep 05, 2015 at 04:34:39PM +0100, Richard W.M. Jones wrote:
> [ 52.259269] BUG: unable to handle kernel NULL pointer dereference at 00000000000009c8
> [ 52.259269] IP: [<ffffffff813f8b10>] __blkg_lookup+0x40/0xe0
And also:
$ addr2line -e /usr/lib/debug/lib/modules/4.3.0-0.rc0.git7.1.rwmj3.fc24.x86_64/vmlinux ffffffff813f8b10
/usr/src/debug/kernel-4.2.fc24/linux-4.3.0-0.rc0.git7.1.rwmj3.fc24.x86_64/block/blk-cgroup.c:158
152 /*
153 * Hint didn't match. Look up from the radix tree. Note that the
154 * hint can only be updated under queue_lock as otherwise @blkg
155 * could have already been removed from blkg_tree. The caller is
156 * responsible for grabbing queue_lock if @update_hint.
157 */
158 blkg = radix_tree_lookup(&blkcg->blkg_tree, q->id);
159 if (blkg && blkg->q == q) {
160 if (update_hint) {
161 lockdep_assert_held(q->queue_lock);
162 rcu_assign_pointer(blkcg->blkg_hint, blkg);
163 }
164 return blkg;
165 }
Rich.
--
Richard Jones, Virtualization Group, Red Hat http://people.redhat.com/~rjones
Read my programming and virtualization blog: http://rwmj.wordpress.com
libguestfs lets you edit virtual machines. Supports shell scripting,
bindings from many languages. http://libguestfs.org
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-05 20:40 +0200 |
| Message-ID | <q5qxz-6Y-7@gated-at.bofh.it> |
| In reply to | #1219571 |
On Sat, Sep 05, 2015 at 04:48:40PM +0100, Richard W.M. Jones wrote: > On Sat, Sep 05, 2015 at 04:34:39PM +0100, Richard W.M. Jones wrote: > > [ 52.259269] BUG: unable to handle kernel NULL pointer dereference at 00000000000009c8 > > [ 52.259269] IP: [<ffffffff813f8b10>] __blkg_lookup+0x40/0xe0 > > And also: > > $ addr2line -e /usr/lib/debug/lib/modules/4.3.0-0.rc0.git7.1.rwmj3.fc24.x86_64/vmlinux ffffffff813f8b10 > /usr/src/debug/kernel-4.2.fc24/linux-4.3.0-0.rc0.git7.1.rwmj3.fc24.x86_64/block/blk-cgroup.c:158 > > 152 /* > 153 * Hint didn't match. Look up from the radix tree. Note that the > 154 * hint can only be updated under queue_lock as otherwise @blkg > 155 * could have already been removed from blkg_tree. The caller is > 156 * responsible for grabbing queue_lock if @update_hint. > 157 */ > 158 blkg = radix_tree_lookup(&blkcg->blkg_tree, q->id); Okay, found the bug. It was an existing problem which got exposed by the recent bypass change. Will soon post a patch. Thanks a lot! -- tejun -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2015-09-05 21:50 +0200 |
| Subject | [PATCH block/for-linus] block: blkg_destroy_all() should clear q->root_blkg and ->root_rl.blkg |
| Message-ID | <q5rDk-1Dm-7@gated-at.bofh.it> |
| In reply to | #1219598 |
While making the root blkg unconditional, ec13b1d6f0a0 ("blkcg: always
create the blkcg_gq for the root blkcg") removed the part which clears
q->root_blkg and ->root_rl.blkg during q exit. This leaves the two
pointers dangling after blkg_destroy_all(). blk-throttle exit path
performs blkg traversals and dereferences ->root_blkg and can lead to
the following oops.
BUG: unable to handle kernel NULL pointer dereference at 0000000000000558
IP: [<ffffffff81389746>] __blkg_lookup+0x26/0x70
...
task: ffff88001b4e2580 ti: ffff88001ac0c000 task.ti: ffff88001ac0c000
RIP: 0010:[<ffffffff81389746>] [<ffffffff81389746>] __blkg_lookup+0x26/0x70
...
Call Trace:
[<ffffffff8138d14a>] blk_throtl_drain+0x5a/0x110
[<ffffffff8138a108>] blkcg_drain_queue+0x18/0x20
[<ffffffff81369a70>] __blk_drain_queue+0xc0/0x170
[<ffffffff8136a101>] blk_queue_bypass_start+0x61/0x80
[<ffffffff81388c59>] blkcg_deactivate_policy+0x39/0x100
[<ffffffff8138d328>] blk_throtl_exit+0x38/0x50
[<ffffffff8138a14e>] blkcg_exit_queue+0x3e/0x50
[<ffffffff8137016e>] blk_release_queue+0x1e/0xc0
...
While the bug is a straigh-forward use-after-free bug, it is tricky to
reproduce because blkg release is RCU protected and the rest of exit
path usually finishes before RCU grace period.
This patch fixes the bug by updating blkg_destro_all() to clear
q->root_blkg and ->root_rl.blkg.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reported-by: "Richard W.M. Jones" <rjones@redhat.com>
Reported-by: Josh Boyer <jwboyer@fedoraproject.org>
Link: http://lkml.kernel.org/g/CA+5PVA5rzQ0s4723n5rHBcxQa9t0cW8BPPBekr_9aMRoWt2aYg@mail.gmail.com
Fixes: ec13b1d6f0a0 ("blkcg: always create the blkcg_gq for the root blkcg")
Cc: stable@vger.kernel.org # v4.2+
---
Hello,
Richard, can you please verify that this patch fixes the bug?
Thanks a lot!
block/blk-cgroup.c | 3 +++
1 file changed, 3 insertions(+)
diff --git a/block/blk-cgroup.c b/block/blk-cgroup.c
index d6283b3..9cc48d1d 100644
--- a/block/blk-cgroup.c
+++ b/block/blk-cgroup.c
@@ -387,6 +387,9 @@ static void blkg_destroy_all(struct request_queue *q)
blkg_destroy(blkg);
spin_unlock(&blkcg->lock);
}
+
+ q->root_blkg = NULL;
+ q->root_rl.blkg = NULL;
}
/*
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Richard W.M. Jones" <rjones@redhat.com> |
|---|---|
| Date | 2015-09-06 10:40 +0200 |
| Subject | Re: [PATCH block/for-linus] block: blkg_destroy_all() should clear q->root_blkg and ->root_rl.blkg |
| Message-ID | <q5DEt-1Wj-7@gated-at.bofh.it> |
| In reply to | #1219626 |
On Sat, Sep 05, 2015 at 03:47:36PM -0400, Tejun Heo wrote:
> While making the root blkg unconditional, ec13b1d6f0a0 ("blkcg: always
> create the blkcg_gq for the root blkcg") removed the part which clears
> q->root_blkg and ->root_rl.blkg during q exit. This leaves the two
> pointers dangling after blkg_destroy_all(). blk-throttle exit path
> performs blkg traversals and dereferences ->root_blkg and can lead to
> the following oops.
>
> BUG: unable to handle kernel NULL pointer dereference at 0000000000000558
> IP: [<ffffffff81389746>] __blkg_lookup+0x26/0x70
> ...
> task: ffff88001b4e2580 ti: ffff88001ac0c000 task.ti: ffff88001ac0c000
> RIP: 0010:[<ffffffff81389746>] [<ffffffff81389746>] __blkg_lookup+0x26/0x70
> ...
> Call Trace:
> [<ffffffff8138d14a>] blk_throtl_drain+0x5a/0x110
> [<ffffffff8138a108>] blkcg_drain_queue+0x18/0x20
> [<ffffffff81369a70>] __blk_drain_queue+0xc0/0x170
> [<ffffffff8136a101>] blk_queue_bypass_start+0x61/0x80
> [<ffffffff81388c59>] blkcg_deactivate_policy+0x39/0x100
> [<ffffffff8138d328>] blk_throtl_exit+0x38/0x50
> [<ffffffff8138a14e>] blkcg_exit_queue+0x3e/0x50
> [<ffffffff8137016e>] blk_release_queue+0x1e/0xc0
> ...
>
> While the bug is a straigh-forward use-after-free bug, it is tricky to
> reproduce because blkg release is RCU protected and the rest of exit
> path usually finishes before RCU grace period.
>
> This patch fixes the bug by updating blkg_destro_all() to clear
> q->root_blkg and ->root_rl.blkg.
>
> Signed-off-by: Tejun Heo <tj@kernel.org>
> Reported-by: "Richard W.M. Jones" <rjones@redhat.com>
> Reported-by: Josh Boyer <jwboyer@fedoraproject.org>
> Link: http://lkml.kernel.org/g/CA+5PVA5rzQ0s4723n5rHBcxQa9t0cW8BPPBekr_9aMRoWt2aYg@mail.gmail.com
> Fixes: ec13b1d6f0a0 ("blkcg: always create the blkcg_gq for the root blkcg")
> Cc: stable@vger.kernel.org # v4.2+
> ---
> Hello,
>
> Richard, can you please verify that this patch fixes the bug?
This patch managed 477 iterations before dying from an unrelated
reason in the test harness. This is much better than before, so the
patch looks good to me.
Rich.
> Thanks a lot!
>
> block/blk-cgroup.c | 3 +++
> 1 file changed, 3 insertions(+)
>
> diff --git a/block/blk-cgroup.c b/block/blk-cgroup.c
> index d6283b3..9cc48d1d 100644
> --- a/block/blk-cgroup.c
> +++ b/block/blk-cgroup.c
> @@ -387,6 +387,9 @@ static void blkg_destroy_all(struct request_queue *q)
> blkg_destroy(blkg);
> spin_unlock(&blkcg->lock);
> }
> +
> + q->root_blkg = NULL;
> + q->root_rl.blkg = NULL;
> }
>
> /*
--
Richard Jones, Virtualization Group, Red Hat http://people.redhat.com/~rjones
Read my programming and virtualization blog: http://rwmj.wordpress.com
libguestfs lets you edit virtual machines. Supports shell scripting,
bindings from many languages. http://libguestfs.org
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web