Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1230900 > unrolled thread
| Started by | Alex Deucher <alexdeucher@gmail.com> |
|---|---|
| First post | 2015-09-22 22:00 +0200 |
| Last post | 2015-09-30 20:10 +0200 |
| Articles | 19 — 5 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() Alex Deucher <alexdeucher@gmail.com> - 2015-09-22 22:00 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() Alex Deucher <alexdeucher@gmail.com> - 2015-09-22 23:00 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() Daniel Vetter <daniel@ffwll.ch> - 2015-09-23 09:30 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() Borislav Petkov <bp@alien8.de> - 2015-09-23 11:10 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() Daniel Vetter <daniel@ffwll.ch> - 2015-09-23 16:50 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() Borislav Petkov <bp@alien8.de> - 2015-09-23 18:10 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() Borislav Petkov <bp@alien8.de> - 2015-09-23 18:20 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Borislav Petkov <bp@alien8.de> - 2015-09-26 18:50 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Jiang Liu <jiang.liu@linux.intel.com> - 2015-09-29 11:00 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Borislav Petkov <bp@alien8.de> - 2015-09-29 13:00 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Jiang Liu <jiang.liu@linux.intel.com> - 2015-09-30 09:50 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Joerg Roedel <joro@8bytes.org> - 2015-09-30 14:50 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Jiang Liu <jiang.liu@linux.intel.com> - 2015-09-30 19:10 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Borislav Petkov <bp@alien8.de> - 2015-09-30 19:40 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Joerg Roedel <joro@8bytes.org> - 2015-09-30 20:10 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Jiang Liu <jiang.liu@linux.intel.com> - 2015-10-03 09:40 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Borislav Petkov <bp@alien8.de> - 2015-10-03 11:40 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Joerg Roedel <joro@8bytes.org> - 2015-10-05 12:10 +0200
Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized Joerg Roedel <joro@8bytes.org> - 2015-09-30 20:10 +0200
| From | Alex Deucher <alexdeucher@gmail.com> |
|---|---|
| Date | 2015-09-22 22:00 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() |
| Message-ID | <qbBTk-6Qf-31@gated-at.bofh.it> |
On Mon, Sep 21, 2015 at 9:31 AM, Borislav Petkov <bp@alien8.de> wrote: > Hi guys, > > this assert_drm_connector_list_read_locked() thing fires here when > suspending to disk with Linus' master from around a week ago and > tip/master merged ontop. > > After I resume, box comes up but wedges solid. I've managed to capture > that splat in its whole glory too, see the end of this mail. > > Let me know if you need more info. > > Thanks. > What system is this? What GPU are you using? Can you bisect? Alex -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Alex Deucher <alexdeucher@gmail.com> |
|---|---|
| Date | 2015-09-22 23:00 +0200 |
| Message-ID | <qbCPo-8bs-21@gated-at.bofh.it> |
| In reply to | #1230900 |
On Tue, Sep 22, 2015 at 4:21 PM, Borislav Petkov <bp@alien8.de> wrote: > Hi Alex, > > On Tue, Sep 22, 2015 at 03:58:03PM -0400, Alex Deucher wrote: >> What system is this? > > my workstation - an > > "To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013" > > you gotta love the "To be filled" crap. In any case, it is an ASUS M5A97 > EVO R2.0. RD890 chip AFAICT. > >> What GPU are you using? > > RV635. Here's some dmesg: > > [ 6.489016] [drm] initializing kernel modesetting (RV635 0x1002:0x9598 0x1043:0x01DA). > [ 7.509177] radeon 0000:01:00.0: VRAM: 512M 0x0000000000000000 - 0x000000001FFFFFFF (512M used) > [ 7.518010] radeon 0000:01:00.0: GTT: 512M 0x0000000020000000 - 0x000000003FFFFFFF > [ 7.525724] [drm] Detected VRAM RAM=512M, BAR=256M > [ 7.530608] [drm] RAM width 128bits DDR > [ 7.535168] [TTM] Zone kernel: Available graphics memory: 8132226 kiB > [ 7.541779] [TTM] Zone dma32: Available graphics memory: 2097152 kiB > [ 7.548420] [TTM] Initializing pool allocator > [ 7.552896] [TTM] Initializing DMA pool allocator > [ 7.558176] [drm] radeon: 512M of VRAM memory ready > [ 7.563131] [drm] radeon: 512M of GTT memory ready. > [ 7.568151] [drm] Loading RV635 Microcode > [ 7.577382] [drm] Internal thermal controller without fan control > [ 7.584349] [drm] radeon: power management initialized > [ 7.590443] [drm] GART: num cpu pages 131072, num gpu pages 131072 > [ 7.597266] [drm] enabling PCIE gen 2 link speeds, disable with radeon.pcie_gen2=0 > [ 7.624386] [drm] PCIE GART of 512M enabled (table at 0x0000000000254000). > [ 7.631544] radeon 0000:01:00.0: WB enabled > [ 7.635794] radeon 0000:01:00.0: fence driver on ring 0 use gpu addr 0x0000000020000c00 and cpu addr 0xffff880427ef7c00 > [ 7.647039] radeon 0000:01:00.0: fence driver on ring 5 use gpu addr 0x00000000000521d0 and cpu addr 0xffffc900008121d0 > [ 7.657924] [drm] Supports vblank timestamp caching Rev 2 (21.10.2013). > [ 7.664601] [drm] Driver supports precise vblank timestamp query. > [ 7.670780] radeon 0000:01:00.0: radeon: MSI limited to 32-bit > [ 7.676801] radeon 0000:01:00.0: radeon: using MSI. > [ 7.681863] [drm] radeon: irq initialized. > [ 7.717757] [drm] ring test on 0 succeeded in 0 usecs > [ 7.897466] [drm] ring test on 5 succeeded in 1 usecs > [ 7.902585] [drm] UVD initialized successfully. > [ 7.908108] [drm] ib test on ring 0 succeeded in 0 usecs > [ 8.558968] [drm] ib test on ring 5 succeeded > [ 8.568734] [drm] Radeon Display Connectors > [ 8.573005] [drm] Connector 0: > [ 8.576189] [drm] DVI-I-1 > [ 8.579062] [drm] HPD1 > [ 8.581657] [drm] DDC: 0x7e50 0x7e50 0x7e54 0x7e54 0x7e58 0x7e58 0x7e5c 0x7e5c > [ 8.589172] [drm] Encoders: > [ 8.592234] [drm] DFP1: INTERNAL_UNIPHY > [ 8.596492] [drm] CRT2: INTERNAL_KLDSCP_DAC2 > [ 8.601182] [drm] Connector 1: > [ 8.604302] [drm] DIN-1 > [ 8.607012] [drm] Encoders: > [ 8.610043] [drm] TV1: INTERNAL_KLDSCP_DAC2 > [ 8.614642] [drm] Connector 2: > [ 8.617760] [drm] DVI-I-2 > [ 8.620621] [drm] HPD2 > [ 8.623226] [drm] DDC: 0x7e40 0x7e40 0x7e44 0x7e44 0x7e48 0x7e48 0x7e4c 0x7e4c > [ 8.630719] [drm] Encoders: > [ 8.633749] [drm] CRT1: INTERNAL_KLDSCP_DAC1 > [ 8.638436] [drm] DFP2: INTERNAL_KLDSCP_LVTMA > [ 8.719815] [drm] fb mappable at 0xC0355000 > [ 8.724089] [drm] vram apper at 0xC0000000 > [ 8.728243] [drm] size 9216000 > [ 8.731371] [drm] fb depth is 24 > [ 8.734664] [drm] pitch is 7680 > [ 8.739009] fbcon: radeondrmfb (fb0) is primary device > [ 8.802887] Console: switching to colour frame buffer device 240x75 > [ 8.818487] radeon 0000:01:00.0: fb0: radeondrmfb frame buffer device > [ 8.824948] radeon 0000:01:00.0: registered panic notifier > [ 8.846452] [drm] Initialized radeon 2.42.0 20080528 for 0000:01:00.0 on minor 0 > >> Can you bisect? > > It is my workstation so it will take longer but I'll try. > > If you can think of some particular commits I should try, let me know. Sorry, I can't think of anything off hand. I suspect it was some change or cleanup in the core drm code. Alex -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Daniel Vetter <daniel@ffwll.ch> |
|---|---|
| Date | 2015-09-23 09:30 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() |
| Message-ID | <qbMF3-5Lu-3@gated-at.bofh.it> |
| In reply to | #1230947 |
On Tue, Sep 22, 2015 at 04:54:54PM -0400, Alex Deucher wrote:
> On Tue, Sep 22, 2015 at 4:21 PM, Borislav Petkov <bp@alien8.de> wrote:
> > Hi Alex,
> >
> > On Tue, Sep 22, 2015 at 03:58:03PM -0400, Alex Deucher wrote:
> >> What system is this?
> >
> > my workstation - an
> >
> > "To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013"
> >
> > you gotta love the "To be filled" crap. In any case, it is an ASUS M5A97
> > EVO R2.0. RD890 chip AFAICT.
> >
> >> What GPU are you using?
> >
> > RV635. Here's some dmesg:
> >
> > [ 6.489016] [drm] initializing kernel modesetting (RV635 0x1002:0x9598 0x1043:0x01DA).
> > [ 7.509177] radeon 0000:01:00.0: VRAM: 512M 0x0000000000000000 - 0x000000001FFFFFFF (512M used)
> > [ 7.518010] radeon 0000:01:00.0: GTT: 512M 0x0000000020000000 - 0x000000003FFFFFFF
> > [ 7.525724] [drm] Detected VRAM RAM=512M, BAR=256M
> > [ 7.530608] [drm] RAM width 128bits DDR
> > [ 7.535168] [TTM] Zone kernel: Available graphics memory: 8132226 kiB
> > [ 7.541779] [TTM] Zone dma32: Available graphics memory: 2097152 kiB
> > [ 7.548420] [TTM] Initializing pool allocator
> > [ 7.552896] [TTM] Initializing DMA pool allocator
> > [ 7.558176] [drm] radeon: 512M of VRAM memory ready
> > [ 7.563131] [drm] radeon: 512M of GTT memory ready.
> > [ 7.568151] [drm] Loading RV635 Microcode
> > [ 7.577382] [drm] Internal thermal controller without fan control
> > [ 7.584349] [drm] radeon: power management initialized
> > [ 7.590443] [drm] GART: num cpu pages 131072, num gpu pages 131072
> > [ 7.597266] [drm] enabling PCIE gen 2 link speeds, disable with radeon.pcie_gen2=0
> > [ 7.624386] [drm] PCIE GART of 512M enabled (table at 0x0000000000254000).
> > [ 7.631544] radeon 0000:01:00.0: WB enabled
> > [ 7.635794] radeon 0000:01:00.0: fence driver on ring 0 use gpu addr 0x0000000020000c00 and cpu addr 0xffff880427ef7c00
> > [ 7.647039] radeon 0000:01:00.0: fence driver on ring 5 use gpu addr 0x00000000000521d0 and cpu addr 0xffffc900008121d0
> > [ 7.657924] [drm] Supports vblank timestamp caching Rev 2 (21.10.2013).
> > [ 7.664601] [drm] Driver supports precise vblank timestamp query.
> > [ 7.670780] radeon 0000:01:00.0: radeon: MSI limited to 32-bit
> > [ 7.676801] radeon 0000:01:00.0: radeon: using MSI.
> > [ 7.681863] [drm] radeon: irq initialized.
> > [ 7.717757] [drm] ring test on 0 succeeded in 0 usecs
> > [ 7.897466] [drm] ring test on 5 succeeded in 1 usecs
> > [ 7.902585] [drm] UVD initialized successfully.
> > [ 7.908108] [drm] ib test on ring 0 succeeded in 0 usecs
> > [ 8.558968] [drm] ib test on ring 5 succeeded
> > [ 8.568734] [drm] Radeon Display Connectors
> > [ 8.573005] [drm] Connector 0:
> > [ 8.576189] [drm] DVI-I-1
> > [ 8.579062] [drm] HPD1
> > [ 8.581657] [drm] DDC: 0x7e50 0x7e50 0x7e54 0x7e54 0x7e58 0x7e58 0x7e5c 0x7e5c
> > [ 8.589172] [drm] Encoders:
> > [ 8.592234] [drm] DFP1: INTERNAL_UNIPHY
> > [ 8.596492] [drm] CRT2: INTERNAL_KLDSCP_DAC2
> > [ 8.601182] [drm] Connector 1:
> > [ 8.604302] [drm] DIN-1
> > [ 8.607012] [drm] Encoders:
> > [ 8.610043] [drm] TV1: INTERNAL_KLDSCP_DAC2
> > [ 8.614642] [drm] Connector 2:
> > [ 8.617760] [drm] DVI-I-2
> > [ 8.620621] [drm] HPD2
> > [ 8.623226] [drm] DDC: 0x7e40 0x7e40 0x7e44 0x7e44 0x7e48 0x7e48 0x7e4c 0x7e4c
> > [ 8.630719] [drm] Encoders:
> > [ 8.633749] [drm] CRT1: INTERNAL_KLDSCP_DAC1
> > [ 8.638436] [drm] DFP2: INTERNAL_KLDSCP_LVTMA
> > [ 8.719815] [drm] fb mappable at 0xC0355000
> > [ 8.724089] [drm] vram apper at 0xC0000000
> > [ 8.728243] [drm] size 9216000
> > [ 8.731371] [drm] fb depth is 24
> > [ 8.734664] [drm] pitch is 7680
> > [ 8.739009] fbcon: radeondrmfb (fb0) is primary device
> > [ 8.802887] Console: switching to colour frame buffer device 240x75
> > [ 8.818487] radeon 0000:01:00.0: fb0: radeondrmfb frame buffer device
> > [ 8.824948] radeon 0000:01:00.0: registered panic notifier
> > [ 8.846452] [drm] Initialized radeon 2.42.0 20080528 for 0000:01:00.0 on minor 0
> >
> >> Can you bisect?
> >
> > It is my workstation so it will take longer but I'll try.
> >
> > If you can think of some particular commits I should try, let me know.
>
> Sorry, I can't think of anything off hand. I suspect it was some
> change or cleanup in the core drm code.
The locking check is new, but I was only adding locking checks, not yet
reworking the locking itself. So the backtrace is likely (but not 100%
guaranteed) a red herring.
Strange thing is that I've tested this on a radeon over here and I don't
see this backtrace ... wut. Below diff should appease the backtraces at
least.
-Daniel
diff --git a/drivers/gpu/drm/radeon/radeon_device.c b/drivers/gpu/drm/radeon/radeon_device.c
index d8319dae8358..9f05de73ae97 100644
--- a/drivers/gpu/drm/radeon/radeon_device.c
+++ b/drivers/gpu/drm/radeon/radeon_device.c
@@ -1734,9 +1734,11 @@ int radeon_resume_kms(struct drm_device *dev, bool resume, bool fbcon)
if (fbcon) {
drm_helper_resume_force_mode(dev);
/* turn on display hw */
+ drm_modeset_lock_all(dev);
list_for_each_entry(connector, &dev->mode_config.connector_list, head) {
drm_helper_connector_dpms(connector, DRM_MODE_DPMS_ON);
}
+ drm_modeset_unlock_all(dev);
}
drm_kms_helper_poll_enable(dev);
--
Daniel Vetter
Software Engineer, Intel Corporation
http://blog.ffwll.ch
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2015-09-23 11:10 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() |
| Message-ID | <qbOdQ-84X-21@gated-at.bofh.it> |
| In reply to | #1231191 |
On Wed, Sep 23, 2015 at 09:25:23AM +0200, Daniel Vetter wrote:
> Strange thing is that I've tested this on a radeon over here and I don't
> see this backtrace ... wut. Below diff should appease the backtraces at
> least.
Doesn't look like it.
This is what it says when suspending:
[ 42.962275] hib.sh (3269): drop_caches: 3
[ 42.967671] PM: Hibernation mode set to 'shutdown'
[ 42.979329] PM: Syncing filesystems ... done.
[ 42.993401] Freezing user space processes ... (elapsed 0.002 seconds) done.
[ 43.003632] PM: Marking nosave pages: [mem 0x00000000-0x00000fff]
[ 43.009840] PM: Marking nosave pages: [mem 0x0009e000-0x000fffff]
[ 43.015991] PM: Marking nosave pages: [mem 0xba9b8000-0xbca4dfff]
[ 43.022241] PM: Marking nosave pages: [mem 0xbca4f000-0xbcc54fff]
[ 43.028357] PM: Marking nosave pages: [mem 0xbd083000-0xbd7f3fff]
[ 43.034500] PM: Marking nosave pages: [mem 0xbd800000-0x100000fff]
[ 43.041371] PM: Basic memory bitmaps created
[ 43.045656] PM: Preallocating image memory... done (allocated 128867 pages)
[ 43.346216] PM: Allocated 515468 kbytes in 0.29 seconds (1777.47 MB/s)
[ 43.352759] Freezing remaining freezable tasks ... (elapsed 0.001 seconds) done.
[ 43.366104] ------------[ cut here ]------------
[ 43.370746] WARNING: CPU: 4 PID: 55 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90()
[ 43.380681] Modules linked in: binfmt_misc ipv6 vfat fat fuse dm_crypt dm_mod kvm_amd kvm crc32_pclmul aesni_intel ae
s_x86_64 lrw gf128mul glue_helper ablk_helper cryptd amd64_edac_mod k10temp fam15h_power edac_core amdkfd amd_iommu_v2 r
adeon acpi_cpufreq
[ 43.403916] CPU: 4 PID: 55 Comm: kworker/u16:2 Not tainted 4.3.0-rc2+ #3
[ 43.410633] Hardware name: To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013
[ 43.420567] Workqueue: events_unbound async_run_entry_fn
[ 43.425919] ffffffff8194ff67 ffff88042a223b60 ffffffff812c758a 0000000000000000
[ 43.433424] ffff88042a223b98 ffffffff810534c1 ffff880429eca000 ffff880429fe9200
[ 43.440959] ffff880429de1000 0000000000000000 ffffffff819571c3 ffff88042a223ba8
[ 43.448461] Call Trace:
[ 43.450929] [<ffffffff812c758a>] dump_stack+0x4e/0x84
[ 43.456094] [<ffffffff810534c1>] warn_slowpath_common+0x91/0xd0
[ 43.462126] [<ffffffff810535ba>] warn_slowpath_null+0x1a/0x20
[ 43.467983] [<ffffffff813bdc58>] drm_helper_choose_encoder_dpms+0x88/0x90
[ 43.474884] [<ffffffff813be0b6>] drm_helper_connector_dpms+0x56/0x110
[ 43.481463] [<ffffffffa003346b>] radeon_suspend_kms+0x6b/0x380 [radeon]
[ 43.488196] [<ffffffff816c0f2b>] ? _raw_spin_unlock_irqrestore+0x4b/0x80
[ 43.495020] [<ffffffff8130fe50>] ? pci_pm_poweroff+0x100/0x100
[ 43.500976] [<ffffffffa00311ac>] radeon_pmops_freeze+0x1c/0x20 [radeon]
[ 43.507730] [<ffffffff8130feba>] pci_pm_freeze+0x6a/0x100
[ 43.513241] [<ffffffff8130fe50>] ? pci_pm_poweroff+0x100/0x100
[ 43.519187] [<ffffffff8146e7e7>] dpm_run_callback+0x77/0x2a0
[ 43.524959] [<ffffffff8146f4a4>] __device_suspend+0x104/0x2c0
[ 43.530818] [<ffffffff8146f67f>] async_suspend+0x1f/0xa0
[ 43.536242] [<ffffffff81078a06>] async_run_entry_fn+0x46/0xf0
[ 43.542127] [<ffffffff8106f548>] process_one_work+0x1f8/0x640
[ 43.547984] [<ffffffff8106f4a4>] ? process_one_work+0x154/0x640
[ 43.554019] [<ffffffff8106f9db>] worker_thread+0x4b/0x440
[ 43.559530] [<ffffffff8107e133>] ? preempt_count_sub+0xb3/0x110
[ 43.565562] [<ffffffff8106f990>] ? process_one_work+0x640/0x640
[ 43.571603] [<ffffffff81075e86>] kthread+0xf6/0x110
[ 43.576597] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 43.583152] [<ffffffff816c1e3f>] ret_from_fork+0x3f/0x70
[ 43.588574] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 43.595170] ---[ end trace aab225b93a6f1dcc ]---
[ 43.595235] ------------[ cut here ]------------
[ 43.595240] WARNING: CPU: 5 PID: 55 at include/drm/drm_crtc.h:1577 drm_helper_choose_crtc_dpms+0x91/0xa0()
[ 43.595267] Modules linked in: binfmt_misc ipv6 vfat fat fuse dm_crypt dm_mod kvm_amd kvm crc32_pclmul aesni_intel aes_x86_64 lrw gf128mul glue_helper ablk_helper cryptd amd64_edac_mod k10temp fam15h_power edac_core amdkfd amd_iommu_v2 radeon acpi_cpufreq
[ 43.595271] CPU: 5 PID: 55 Comm: kworker/u16:2 Tainted: G W 4.3.0-rc2+ #3
[ 43.595272] Hardware name: To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013
[ 43.595276] Workqueue: events_unbound async_run_entry_fn
[ 43.595281] ffffffff8194ff67 ffff88042a223b60 ffffffff812c758a 0000000000000000
[ 43.595286] ffff88042a223b98 ffffffff810534c1 ffff880429eca000 ffff880429de1000
[ 43.595291] ffff880429de1000 0000000000000000 0000000000000003 ffff88042a223ba8
[ 43.595293] Call Trace:
[ 43.595295] [<ffffffff812c758a>] dump_stack+0x4e/0x84
[ 43.595298] [<ffffffff810534c1>] warn_slowpath_common+0x91/0xd0
[ 43.595300] [<ffffffff810535ba>] warn_slowpath_null+0x1a/0x20
[ 43.595303] [<ffffffff813bdcf1>] drm_helper_choose_crtc_dpms+0x91/0xa0
[ 43.595315] [<ffffffffa003d860>] ? atombios_blank_crtc+0x140/0x140 [radeon]
[ 43.595322] [<ffffffff813be124>] drm_helper_connector_dpms+0xc4/0x110
[ 43.595331] [<ffffffffa003346b>] radeon_suspend_kms+0x6b/0x380 [radeon]
[ 43.595334] [<ffffffff816c0f2b>] ? _raw_spin_unlock_irqrestore+0x4b/0x80
[ 43.595337] [<ffffffff8130fe50>] ? pci_pm_poweroff+0x100/0x100
[ 43.595346] [<ffffffffa00311ac>] radeon_pmops_freeze+0x1c/0x20 [radeon]
[ 43.595349] [<ffffffff8130feba>] pci_pm_freeze+0x6a/0x100
[ 43.595351] [<ffffffff8130fe50>] ? pci_pm_poweroff+0x100/0x100
[ 43.595353] [<ffffffff8146e7e7>] dpm_run_callback+0x77/0x2a0
[ 43.595356] [<ffffffff8146f4a4>] __device_suspend+0x104/0x2c0
[ 43.595358] [<ffffffff8146f67f>] async_suspend+0x1f/0xa0
[ 43.595361] [<ffffffff81078a06>] async_run_entry_fn+0x46/0xf0
[ 43.595363] [<ffffffff8106f548>] process_one_work+0x1f8/0x640
[ 43.595366] [<ffffffff8106f4a4>] ? process_one_work+0x154/0x640
[ 43.595369] [<ffffffff8106f9db>] worker_thread+0x4b/0x440
[ 43.595373] [<ffffffff8107e133>] ? preempt_count_sub+0xb3/0x110
[ 43.595375] [<ffffffff8106f990>] ? process_one_work+0x640/0x640
[ 43.595378] [<ffffffff81075e86>] kthread+0xf6/0x110
[ 43.595382] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 43.595385] [<ffffffff816c1e3f>] ret_from_fork+0x3f/0x70
[ 43.595387] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 43.595390] ---[ end trace aab225b93a6f1dcd ]---
[ 43.611465] ------------[ cut here ]------------
[ 43.611469] WARNING: CPU: 5 PID: 55 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90()
[ 43.611494] Modules linked in: binfmt_misc ipv6 vfat fat fuse dm_crypt dm_mod kvm_amd kvm crc32_pclmul aesni_intel aes_x86_64 lrw gf128mul glue_helper ablk_helper cryptd amd64_edac_mod k10temp fam15h_power edac_core amdkfd amd_iommu_v2 radeon acpi_cpufreq
[ 43.611498] CPU: 5 PID: 55 Comm: kworker/u16:2 Tainted: G W 4.3.0-rc2+ #3
[ 43.611499] Hardware name: To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013
[ 43.611503] Workqueue: events_unbound async_run_entry_fn
[ 43.611508] ffffffff8194ff67 ffff88042a223b60 ffffffff812c758a 0000000000000000
[ 43.611516] ffff88042a223b98 ffffffff810534c1 ffff880429eca000 ffff880429fe9e00
[ 43.611520] ffff880429de7000 0000000000000000 ffffffff819571c3 ffff88042a223ba8
[ 43.611521] Call Trace:
[ 43.611525] [<ffffffff812c758a>] dump_stack+0x4e/0x84
[ 43.611527] [<ffffffff810534c1>] warn_slowpath_common+0x91/0xd0
[ 43.611530] [<ffffffff810535ba>] warn_slowpath_null+0x1a/0x20
[ 43.611534] [<ffffffff813bdc58>] drm_helper_choose_encoder_dpms+0x88/0x90
[ 43.611537] [<ffffffff813be0b6>] drm_helper_connector_dpms+0x56/0x110
[ 43.611547] [<ffffffffa003346b>] radeon_suspend_kms+0x6b/0x380 [radeon]
[ 43.611551] [<ffffffff816c0f2b>] ? _raw_spin_unlock_irqrestore+0x4b/0x80
[ 43.611554] [<ffffffff8130fe50>] ? pci_pm_poweroff+0x100/0x100
[ 43.611563] [<ffffffffa00311ac>] radeon_pmops_freeze+0x1c/0x20 [radeon]
[ 43.611566] [<ffffffff8130feba>] pci_pm_freeze+0x6a/0x100
[ 43.611569] [<ffffffff8130fe50>] ? pci_pm_poweroff+0x100/0x100
[ 43.611573] [<ffffffff8146e7e7>] dpm_run_callback+0x77/0x2a0
[ 43.611575] [<ffffffff8146f4a4>] __device_suspend+0x104/0x2c0
[ 43.611578] [<ffffffff8146f67f>] async_suspend+0x1f/0xa0
[ 43.611580] [<ffffffff81078a06>] async_run_entry_fn+0x46/0xf0
[ 43.611582] [<ffffffff8106f548>] process_one_work+0x1f8/0x640
[ 43.611585] [<ffffffff8106f4a4>] ? process_one_work+0x154/0x640
[ 43.611588] [<ffffffff8106f9db>] worker_thread+0x4b/0x440
[ 43.611592] [<ffffffff8107e133>] ? preempt_count_sub+0xb3/0x110
[ 43.611594] [<ffffffff8106f990>] ? process_one_work+0x640/0x640
[ 43.611599] [<ffffffff81075e86>] kthread+0xf6/0x110
[ 43.611604] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 43.611607] [<ffffffff816c1e3f>] ret_from_fork+0x3f/0x70
[ 43.611609] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 43.611612] ---[ end trace aab225b93a6f1dce ]---
[ 43.611638] ------------[ cut here ]------------
[ 43.611641] WARNING: CPU: 5 PID: 55 at include/drm/drm_crtc.h:1577 drm_helper_choose_crtc_dpms+0x91/0xa0()
[ 43.611665] Modules linked in: binfmt_misc ipv6 vfat fat fuse dm_crypt dm_mod kvm_amd kvm crc32_pclmul aesni_intel aes_x86_64 lrw gf128mul glue_helper ablk_helper cryptd amd64_edac_mod k10temp fam15h_power edac_core amdkfd amd_iommu_v2 radeon acpi_cpufreq
[ 43.611670] CPU: 5 PID: 55 Comm: kworker/u16:2 Tainted: G W 4.3.0-rc2+ #3
[ 43.611670] Hardware name: To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013
[ 43.611675] Workqueue: events_unbound async_run_entry_fn
[ 43.611711] ffffffff8194ff67 ffff88042a223b60 ffffffff812c758a 0000000000000000
[ 43.611716] ffff88042a223b98 ffffffff810534c1 ffff880429eca000 ffff880429de7000
[ 43.611720] ffff880429de7000 0000000000000000 0000000000000003 ffff88042a223ba8
[ 43.611721] Call Trace:
[ 43.611725] [<ffffffff812c758a>] dump_stack+0x4e/0x84
[ 43.611728] [<ffffffff810534c1>] warn_slowpath_common+0x91/0xd0
[ 43.611730] [<ffffffff810535ba>] warn_slowpath_null+0x1a/0x20
[ 43.611732] [<ffffffff813bdcf1>] drm_helper_choose_crtc_dpms+0x91/0xa0
[ 43.611742] [<ffffffffa003d860>] ? atombios_blank_crtc+0x140/0x140 [radeon]
[ 43.611747] [<ffffffff813be124>] drm_helper_connector_dpms+0xc4/0x110
[ 43.611756] [<ffffffffa003346b>] radeon_suspend_kms+0x6b/0x380 [radeon]
[ 43.611759] [<ffffffff816c0f2b>] ? _raw_spin_unlock_irqrestore+0x4b/0x80
[ 43.611761] [<ffffffff8130fe50>] ? pci_pm_poweroff+0x100/0x100
[ 43.611771] [<ffffffffa00311ac>] radeon_pmops_freeze+0x1c/0x20 [radeon]
[ 43.611774] [<ffffffff8130feba>] pci_pm_freeze+0x6a/0x100
[ 43.611776] [<ffffffff8130fe50>] ? pci_pm_poweroff+0x100/0x100
[ 43.611778] [<ffffffff8146e7e7>] dpm_run_callback+0x77/0x2a0
[ 43.611781] [<ffffffff8146f4a4>] __device_suspend+0x104/0x2c0
[ 43.611783] [<ffffffff8146f67f>] async_suspend+0x1f/0xa0
[ 43.611786] [<ffffffff81078a06>] async_run_entry_fn+0x46/0xf0
[ 43.611788] [<ffffffff8106f548>] process_one_work+0x1f8/0x640
[ 43.611790] [<ffffffff8106f4a4>] ? process_one_work+0x154/0x640
[ 43.611793] [<ffffffff8106f9db>] worker_thread+0x4b/0x440
[ 43.611797] [<ffffffff8107e133>] ? preempt_count_sub+0xb3/0x110
[ 43.611800] [<ffffffff8106f990>] ? process_one_work+0x640/0x640
[ 43.611802] [<ffffffff81075e86>] kthread+0xf6/0x110
[ 43.611807] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 43.611809] [<ffffffff816c1e3f>] ret_from_fork+0x3f/0x70
[ 43.611812] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 43.611814] ---[ end trace aab225b93a6f1dcf ]---
[ 44.409186] PM: freeze of devices complete after 1045.976 msecs
[ 44.417505] PM: late freeze of devices complete after 2.393 msecs
[ 44.427721] PM: noirq freeze of devices complete after 4.126 msecs
[ 44.433913] Disabling non-boot CPUs ...
[ 44.451184] smpboot: CPU 1 is now offline
[ 44.498597] smpboot: CPU 2 is now offline
[ 44.534426] smpboot: CPU 3 is now offline
[ 44.574909] smpboot: CPU 4 is now offline
[ 44.614270] smpboot: CPU 5 is now offline
[ 44.652566] smpboot: CPU 6 is now offline
[ 44.694825] smpboot: CPU 7 is now offline
[ 44.711829] PM: Creating hibernation image:
[ 45.179293] PM: Need to copy 138846 pages
[ 45.183311] PM: Normal pages needed: 138846 + 1024, available pages: 4029960
[ 45.898261] PM: Hibernation image created (138846 pages copied)
[ 45.100717] LVT offset 0 assigned for vector 0x400
[ 45.106190] Enabling non-boot CPUs ...
[ 45.110174] x86: Booting SMP configuration:
[ 45.114369] smpboot: Booting Node 0 Processor 1 APIC 0x11
[ 45.141886] cache: parent cpu1 should not be sleeping
[ 45.148488] CPU1 is up
[ 45.150958] smpboot: Booting Node 0 Processor 2 APIC 0x12
[ 45.178801] cache: parent cpu2 should not be sleeping
[ 45.185799] CPU2 is up
[ 45.188316] smpboot: Booting Node 0 Processor 3 APIC 0x13
[ 45.215892] cache: parent cpu3 should not be sleeping
[ 45.222959] CPU3 is up
[ 45.225467] smpboot: Booting Node 0 Processor 4 APIC 0x14
[ 45.244730] cache: parent cpu4 should not be sleeping
[ 45.250525] CPU4 is up
[ 45.252954] smpboot: Booting Node 0 Processor 5 APIC 0x15
[ 45.273487] cache: parent cpu5 should not be sleeping
[ 45.279302] CPU5 is up
[ 45.281722] smpboot: Booting Node 0 Processor 6 APIC 0x16
[ 45.309551] cache: parent cpu6 should not be sleeping
[ 45.316600] CPU6 is up
[ 45.319109] smpboot: Booting Node 0 Processor 7 APIC 0x17
[ 45.344059] cache: parent cpu7 should not be sleeping
[ 45.351114] CPU7 is up
[ 45.376664] PM: noirq thaw of devices complete after 2.185 msecs
[ 45.384992] PM: early thaw of devices complete after 2.249 msecs
[ 45.391875] rtc_cmos 00:03: System wakeup disabled by ACPI
[ 45.393086] [drm] PCIE gen 2 link speeds already enabled
[ 45.393444] serial 00:06: activated
[ 45.396643] [drm] PCIE GART of 512M enabled (table at 0x0000000000254000).
[ 45.396699] radeon 0000:01:00.0: WB enabled
[ 45.396704] radeon 0000:01:00.0: fence driver on ring 0 use gpu addr 0x0000000020000c00 and cpu addr 0xffff8804292eec00
[ 45.397110] radeon 0000:01:00.0: fence driver on ring 5 use gpu addr 0x00000000000521d0 and cpu addr 0xffffc900008121d0
[ 45.428076] [drm] ring test on 0 succeeded in 0 usecs
[ 45.510492] r8169 0000:02:00.0 eth0: link down
[ 45.602794] [drm] ring test on 5 succeeded in 1 usecs
[ 45.602803] [drm] UVD initialized successfully.
[ 45.603013] [drm] ib test on ring 0 succeeded in 0 usecs
[ 45.710859] ata5: SATA link down (SStatus 0 SControl 300)
[ 45.710907] ata6: SATA link down (SStatus 0 SControl 300)
[ 45.882925] ata1: SATA link up 6.0 Gbps (SStatus 133 SControl 300)
[ 45.882994] ata3: SATA link up 6.0 Gbps (SStatus 133 SControl 300)
[ 45.883059] ata2: SATA link up 6.0 Gbps (SStatus 133 SControl 300)
[ 45.883092] ata4: SATA link up 3.0 Gbps (SStatus 123 SControl 300)
[ 45.884972] ata1.00: supports DRM functions and may not be fully accessible
[ 45.885071] ata2.00: supports DRM functions and may not be fully accessible
[ 45.885123] ata1.00: failed to get NCQ Send/Recv Log Emask 0x1
[ 45.885218] ata2.00: failed to get NCQ Send/Recv Log Emask 0x1
[ 45.885907] ata1.00: supports DRM functions and may not be fully accessible
[ 45.886024] ata1.00: failed to get NCQ Send/Recv Log Emask 0x1
[ 45.886068] ata2.00: supports DRM functions and may not be fully accessible
[ 45.886110] ata1.00: configured for UDMA/133
[ 45.886207] ata2.00: failed to get NCQ Send/Recv Log Emask 0x1
[ 45.886417] ata2.00: configured for UDMA/133
[ 45.896091] ata4.00: configured for UDMA/133
[ 45.906938] ata3.00: configured for UDMA/133
[ 45.907075] sd 2:0:0:0: [sdc] 1220942646 4096-byte logical blocks: (4.88 TB/4.54 TiB)
[ 46.250991] [drm] ib test on ring 5 succeeded
[ 46.644971] PM: thaw of devices complete after 1253.970 msecs
[ 46.848000] usb 8-2: reset low-speed USB device number 2 using ohci-pci
[ 47.084410] r8169 0000:02:00.0 eth0: link up
[ 47.193411] PM: writing image.
[ 47.203074] PM: Using 3 thread(s) for compression.
[ 47.203074] PM: Compressing and saving image data (139118 pages)...
[ 47.220107] PM: Image saving progress: 0%
[ 47.356371] PM: Image saving progress: 10%
[ 47.444633] PM: Image saving progress: 20%
[ 47.586282] PM: Image saving progress: 30%
[ 47.676366] PM: Image saving progress: 40%
[ 47.751889] PM: Image saving progress: 50%
[ 47.838831] PM: Image saving progress: 60%
[ 47.927719] PM: Image saving progress: 70%
[ 48.020115] PM: Image saving progress: 80%
[ 48.107692] PM: Image saving progress: 90%
[ 48.193569] PM: Image saving progress: 100%
[ 48.199217] PM: Image saving done.
[ 48.203784] PM: Wrote 556472 kbytes in 0.97 seconds (573.68 MB/s)
[ 48.211456] PM: S|
[ 48.311190] kvm: exiting hardware virtualization
[ 48.321689] AMD-Vi: Event logged [IO_PAGE_FAULT device=01:00.0 domain=0x0010 address=0x0000000020001000 flags=0x0000]
[ 48.522749] sd 3:0:0:0: [sdd] Synchronizing SCSI cache
[ 48.531044] sd 3:0:0:0: [sdd] Stopping disk
[ 49.418426] sd 2:0:0:0: [sdc] Synchronizing SCSI cache
[ 49.424890] sd 2:0:0:0: [sdc] Stopping disk
[ 49.596437] sd 1:0:0:0: [sdb] Synchronizing SCSI cache
[ 49.605855] sd 1:0:0:0: [sdb] Stopping disk
[ 49.920447] sd 0:0:0:0: [sda] Synchronizing SCSI cache
[ 49.929890] sd 0:0:0:0: [sda] Stopping disk
[ 50.124196] pcieport 0000:00:04.0: System wakeup enabled by ACPI
[ 50.152897] ACPI: Preparing to enter system sleep state S5
[ 50.160275] [Firmware Bug]: ACPI: BIOS _OSI(Linux) query honored via cmdline
[ 50.170923] reboot: Power down
[ 50.176837] acpi_power_off called
Then the resume kernel starts and loads the suspended one, which bombs
out completely:
[ 5.758121] PM: Checking hibernation image partition /dev/sda1
[ 5.764067] PM: Hibernation image partition 8:1 present
[ 5.769372] PM: Looking for hibernation image.
[ 5.769846] hid-generic 0003:04B4:0101.0003: input,hidraw2: USB HID v1.00 Device [DATACOMP SteelS쀁̄Љ̒DATA] on usb-0000
:00:12.0-2/input1
[ 5.788988] PM: Image signature found, resuming
[ 5.795952] PM: Preparing processes for restore.
[ 5.800605] Freezing user space processes ... (elapsed 0.000 seconds) done.
[ 5.807811] PM: Loading hibernation image.
[ 5.812211] PM: Marking nosave pages: [mem 0x00000000-0x00000fff]
[ 5.818339] PM: Marking nosave pages: [mem 0x0009e000-0x000fffff]
[ 5.824462] PM: Marking nosave pages: [mem 0xba9b8000-0xbca4dfff]
[ 5.830712] PM: Marking nosave pages: [mem 0xbca4f000-0xbcc54fff]
[ 5.836844] PM: Marking nosave pages: [mem 0xbd083000-0xbd7f3fff]
[ 5.843010] PM: Marking nosave pages: [mem 0xbd800000-0x100000fff]
[ 5.849924] PM: Basic memory bitmaps created
[ 5.869152] PM: Using 3 thread(s) for decompression.
[ 5.869152] PM: Loading and decompressing image data (139118 pages)...
[ 5.946421] PM: Image loading progress: 0%
[ 6.241471] PM: Image loading progress: 10%
[ 6.321625] PM: Image loading progress: 20%
[ 6.362579] random: nonblocking pool is initialized
[ 6.414363] PM: Image loading progress: 30%
[ 6.495027] PM: Image loading progress: 40%
[ 6.573551] PM: Image loading progress: 50%
[ 6.654276] PM: Image loading progress: 60%
[ 6.735344] PM: Image loading progress: 70%
[ 6.810312] PM: Image loading progress: 80%
[ 6.885237] PM: Image loading progress: 90%
[ 6.963869] PM: Image loading progress: 100%
[ 6.968292] PM: Image loading done.
[ 6.971907] PM: Read 556472 kbytes in 1.08 seconds (515.25 MB/s)
[ 6.981245] PM: Image successfully loaded
[ 7.208118] PM: quiesce of devices complete after 221.517 msecs
[ 7.215932] PM: late quiesce of devices complete after 1.806 msecs
[ 7.236424] PM: noirq quiesce of devices complete after 14.231 msecs
[ 7.242867] Disabling non-boot CPUs ...
[ 44.968602] LVT offset 0 assigned for vector 0x400
[ 44.973810] Enabling non-boot CPUs ...
[ 44.977701] x86: Booting SMP configuration:
[ 44.981885] smpboot: Booting Node 0 Processor 1 APIC 0x11
[ 45.004260] cache: parent cpu1 should not be sleeping
[ 45.010101] CPU1 is up
[ 45.012520] smpboot: Booting Node 0 Processor 2 APIC 0x12
[ 45.036083] cache: parent cpu2 should not be sleeping
[ 45.041998] CPU2 is up
[ 45.044429] smpboot: Booting Node 0 Processor 3 APIC 0x13
[ 45.064873] cache: parent cpu3 should not be sleeping
[ 45.070712] CPU3 is up
[ 45.073177] smpboot: Booting Node 0 Processor 4 APIC 0x14
[ 45.092508] cache: parent cpu4 should not be sleeping
[ 45.098293] CPU4 is up
[ 45.100772] smpboot: Booting Node 0 Processor 5 APIC 0x15
[ 45.120788] cache: parent cpu5 should not be sleeping
[ 45.126577] CPU5 is up
[ 45.128997] smpboot: Booting Node 0 Processor 6 APIC 0x16
[ 45.149043] cache: parent cpu6 should not be sleeping
[ 45.154843] CPU6 is up
[ 45.157261] smpboot: Booting Node 0 Processor 7 APIC 0x17
[ 45.177371] cache: parent cpu7 should not be sleeping
[ 45.183182] CPU7 is up
[ 45.195686] BUG: unable to handle kernel NULL pointer dereference at 0000000000000034
[ 45.203696] IP: [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
[ 45.210379] PGD 418009067 PUD 41800a067 PMD 0
[ 45.215023] Oops: 0000 [#1] PREEMPT SMP
[ 45.219132] Modules linked in: binfmt_misc ipv6 vfat fat fuse dm_crypt dm_mod kvm_amd kvm crc32_pclmul aesni_intel ae
s_x86_64 lrw gf128mul glue_helper ablk_helper cryptd amd64_edac_mod k10temp fam15h_power edac_core amdkfd amd_iommu_v2 r
adeon acpi_cpufreq
[ 45.242447] CPU: 2 PID: 804 Comm: kworker/u16:5 Tainted: G W 4.3.0-rc2+ #3
[ 45.250584] Hardware name: To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013
[ 45.260630] Workqueue: events_unbound async_run_entry_fn
[ 45.266105] task: ffff88042983df00 ti: ffff880428a40000 task.ti: ffff880428a40000
[ 45.273724] RIP: 0010:[<ffffffff81321296>] [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
[ 45.282840] RSP: 0018:ffff880428a43c28 EFLAGS: 00010286
[ 45.288292] RAX: 0000000000000000 RBX: ffff880429c2d000 RCX: 0000000000000000
[ 45.295564] RDX: 0000000000000001 RSI: ffffffff81304448 RDI: ffffffff816c0f2b
[ 45.302835] RBP: ffff880428a43c40 R08: 0000000000000001 R09: 0000000000522000
[ 45.310114] R10: 0000000000000000 R11: 0000000000000000 R12: 0000000000000000
[ 45.317386] R13: ffff880429c2d7b0 R14: ffff880429c2d010 R15: ffff880429c2d038
[ 45.324666] FS: 00007f653bc43700(0000) GS:ffff88042ca00000(0000) knlGS:0000000000000000
[ 45.332898] CS: 0010 DS: 0000 ES: 0000 CR0: 000000008005003b
[ 45.338784] CR2: 0000000000000034 CR3: 0000000418053000 CR4: 00000000000406e0
[ 45.346063] Stack:
[ 45.348221] 0080002c29c2d7b0 0000000000000000 ffff880429c2d000 ffff880428a43c78
[ 45.355822] ffffffff8130c141 ffff880429c2d098 ffff880429c2d000 0000000000000000
[ 45.363424] ffff88042a1216e8 ffffffff81957114 ffff880428a43c88 ffffffff8130c2b8
[ 45.371041] Call Trace:
[ 45.373642] [<ffffffff8130c141>] pci_restore_state.part.18+0xf1/0x250
[ 45.380316] [<ffffffff8130c2b8>] pci_restore_state+0x18/0x20
[ 45.386209] [<ffffffff8130f7fc>] pci_pm_restore_noirq+0x4c/0xd0
[ 45.392367] [<ffffffff8130f7b0>] ? pci_pm_freeze_noirq+0xf0/0xf0
[ 45.398611] [<ffffffff8146e7e7>] dpm_run_callback+0x77/0x2a0
[ 45.404531] [<ffffffff8146eaa3>] device_resume_noirq+0x93/0x150
[ 45.410682] [<ffffffff8146eb7d>] async_resume_noirq+0x1d/0x50
[ 45.416693] [<ffffffff81078a06>] async_run_entry_fn+0x46/0xf0
[ 45.422669] [<ffffffff8106f548>] process_one_work+0x1f8/0x640
[ 45.428648] [<ffffffff8106f4a4>] ? process_one_work+0x154/0x640
[ 45.434804] [<ffffffff8106f9db>] worker_thread+0x4b/0x440
[ 45.440435] [<ffffffff8106f990>] ? process_one_work+0x640/0x640
[ 45.446590] [<ffffffff81075e86>] kthread+0xf6/0x110
[ 45.451704] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 45.458376] [<ffffffff816c1e3f>] ret_from_fork+0x3f/0x70
[ 45.463923] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 45.470587] Code: 66 89 4d ee 0f b7 c9 e8 79 41 fe ff 48 89 df e8 d1 7a ce ff 0f b6 53 4b 8b 73 38 48 8d 4d ee 48 8b 7b 10 83 c2 02 e8 1a 31 fe ff <41> 0f b6 4c 24 34 41 8b 54 24 30 be ff ff ff ff c0 e9 04 83 e1
[ 45.490744] RIP [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
[ 45.497522] RSP <ffff880428a43c28>
[ 45.501162] CR2: 0000000000000034
[ 45.504630] ---[ end trace aab225b93a6f1dd0 ]---
[ 45.509270] BUG: unable to handle kernel paging request at ffffffffffffff98
[ 45.516436] IP: [<ffffffff81076770>] kthread_data+0x10/0x20
[ 45.522180] PGD 19e5067 PUD 19e7067 PMD 0
[ 45.526512] Oops: 0000 [#2] PREEMPT SMP
[ 45.530628] Modules linked in: binfmt_misc ipv6 vfat fat fuse dm_crypt dm_mod kvm_amd kvm crc32_pclmul aesni_intel aes_x86_64 lrw gf128mul glue_helper ablk_helper cryptd amd64_edac_mod k10temp fam15h_power edac_core amdkfd amd_iommu_v2 radeon acpi_cpufreq
[ 45.553976] CPU: 2 PID: 804 Comm: kworker/u16:5 Tainted: G D W 4.3.0-rc2+ #3
[ 45.562132] Hardware name: To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013
[ 45.572199] task: ffff88042983df00 ti: ffff880428a40000 task.ti: ffff880428a40000
[ 45.579856] RIP: 0010:[<ffffffff81076770>] [<ffffffff81076770>] kthread_data+0x10/0x20
[ 45.588037] RSP: 0018:ffff880428a43928 EFLAGS: 00010002
[ 45.593505] RAX: 0000000000000000 RBX: 0000000000000002 RCX: 0000000000000000
[ 45.600794] RDX: 000000012b244210 RSI: 0000000000000002 RDI: ffff88042983df00
[ 45.608083] RBP: ffff880428a43928 R08: ffff88042983df88 R09: ffff88042cbd5cb0
[ 45.615371] R10: ffff88042983df60 R11: 0000000000000000 R12: 00000000001d5c00
[ 45.622661] R13: ffff88042cbd5c18 R14: ffff88042983df00 R15: 0000000000000002
[ 45.629949] FS: 00007f653bc43700(0000) GS:ffff88042ca00000(0000) knlGS:0000000000000000
[ 45.638191] CS: 0010 DS: 0000 ES: 0000 CR0: 000000008005003b
[ 45.644094] CR2: 0000000000000028 CR3: 0000000418053000 CR4: 00000000000406e0
[ 45.651389] Stack:
[ 45.653565] ffff880428a43940 ffffffff810707a1 ffff88042cbd5c00 ffff880428a43998
[ 45.661185] ffffffff816bb6e6 ffffffff81055b9b 0000000000000000 0000000000000000
[ 45.668835] ffff88042983df00 ffff880428a44000 ffff880428a439f0 0000000000000000
[ 45.676453] Call Trace:
[ 45.679064] [<ffffffff810707a1>] wq_worker_sleeping+0x11/0x90
[ 45.685062] [<ffffffff816bb6e6>] __schedule+0x796/0xec0
[ 45.690538] [<ffffffff81055b9b>] ? do_exit+0x63b/0xac0
[ 45.695930] [<ffffffff816bbe9d>] schedule+0x3d/0x90
[ 45.701060] [<ffffffff81055c58>] do_exit+0x6f8/0xac0
[ 45.706279] [<ffffffff81007d8c>] oops_end+0x6c/0x90
[ 45.711408] [<ffffffff81045c13>] no_context+0x153/0x360
[ 45.716875] [<ffffffff81045f2b>] __bad_area_nosemaphore+0x10b/0x210
[ 45.723385] [<ffffffff812f3b77>] ? debug_smp_processor_id+0x17/0x20
[ 45.729894] [<ffffffff81046043>] bad_area_nosemaphore+0x13/0x20
[ 45.736055] [<ffffffff81046487>] __do_page_fault+0x1e7/0x360
[ 45.741958] [<ffffffff81000f70>] ? trace_hardirqs_off_thunk+0x17/0x19
[ 45.748640] [<ffffffff8104663c>] do_page_fault+0xc/0x10
[ 45.754099] [<ffffffff816c38af>] page_fault+0x1f/0x30
[ 45.759386] [<ffffffff81304448>] ? pci_bus_read_config_word+0x98/0xa0
[ 45.766058] [<ffffffff816c0f2b>] ? _raw_spin_unlock_irqrestore+0x4b/0x80
[ 45.772992] [<ffffffff81321296>] ? pci_restore_msi_state+0x196/0x240
[ 45.779581] [<ffffffff81321296>] ? pci_restore_msi_state+0x196/0x240
[ 45.786165] [<ffffffff8130c141>] pci_restore_state.part.18+0xf1/0x250
[ 45.792840] [<ffffffff8130c2b8>] pci_restore_state+0x18/0x20
[ 45.798732] [<ffffffff8130f7fc>] pci_pm_restore_noirq+0x4c/0xd0
[ 45.804886] [<ffffffff8130f7b0>] ? pci_pm_freeze_noirq+0xf0/0xf0
[ 45.811127] [<ffffffff8146e7e7>] dpm_run_callback+0x77/0x2a0
[ 45.817019] [<ffffffff8146eaa3>] device_resume_noirq+0x93/0x150
[ 45.823173] [<ffffffff8146eb7d>] async_resume_noirq+0x1d/0x50
[ 45.829153] [<ffffffff81078a06>] async_run_entry_fn+0x46/0xf0
[ 45.835132] [<ffffffff8106f548>] process_one_work+0x1f8/0x640
[ 45.841104] [<ffffffff8106f4a4>] ? process_one_work+0x154/0x640
[ 45.847248] [<ffffffff8106f9db>] worker_thread+0x4b/0x440
[ 45.852873] [<ffffffff8106f990>] ? process_one_work+0x640/0x640
[ 45.859018] [<ffffffff81075e86>] kthread+0xf6/0x110
[ 45.864124] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 45.870796] [<ffffffff816c1e3f>] ret_from_fork+0x3f/0x70
[ 45.876333] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 45.882998] Code: 60 74 0a 48 89 df e8 10 a8 64 00 eb d3 48 8b 53 48 eb b1 e8 33 cb fd ff 0f 1f 00 0f 1f 44 00 00 48 8b 87 f8 03 00 00 55 48 89 e5 <48> 8b 40 98 5d c3 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00
[ 45.903148] RIP [<ffffffff81076770>] kthread_data+0x10/0x20
[ 45.908972] RSP <ffff880428a43928>
[ 45.912612] CR2: ffffffffffffff98
[ 45.916079] ---[ end trace aab225b93a6f1dd1 ]---
[ 45.920697] Fixing recursive fault but reboot is needed!
[ 45.926002] BUG: scheduling while atomic: kworker/u16:5/804/0x00000004
[ 45.932529] INFO: lockdep is turned off.
[ 45.936453] Modules linked in: binfmt_misc ipv6 vfat fat fuse dm_crypt dm_mod kvm_amd kvm crc32_pclmul aesni_intel aes_x86_64 lrw gf128mul glue_helper ablk_helper cryptd amd64_edac_mod k10temp fam15h_power edac_core amdkfd amd_iommu_v2 radeon acpi_cpufreq
[ 45.959455] irq event stamp: 2072
[ 45.962766] hardirqs last enabled at (2071): [<ffffffff816c0f45>] _raw_spin_unlock_irqrestore+0x65/0x80
[ 45.972248] hardirqs last disabled at (2072): [<ffffffff816c3a80>] error_entry+0x60/0xb0
[ 45.980342] softirqs last enabled at (1792): [<ffffffff81058547>] __do_softirq+0x3a7/0x480
[ 45.988705] softirqs last disabled at (1765): [<ffffffff81058798>] irq_exit+0x88/0xb0
[ 45.996539] Preemption disabled at:[<ffffffff81007d8c>] oops_end+0x6c/0x90
[ 46.003420]
[ 46.004913] CPU: 2 PID: 804 Comm: kworker/u16:5 Tainted: G D W 4.3.0-rc2+ #3
[ 46.012911] Hardware name: To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013
[ 46.022823] 00000000001d5c00 ffff880428a43628 ffffffff812c758a ffff88042983df00
[ 46.030280] ffff880428a43640 ffffffff8107d528 ffff88042cbd5c00 ffff880428a43698
[ 46.037740] ffffffff816bb84f ffff880428a436b0 ffffffff81126b16 0000000000000008
[ 46.045193] Call Trace:
[ 46.047641] [<ffffffff812c758a>] dump_stack+0x4e/0x84
[ 46.052779] [<ffffffff8107d528>] __schedule_bug+0x68/0xc0
[ 46.058264] [<ffffffff816bb84f>] __schedule+0x8ff/0xec0
[ 46.063569] [<ffffffff81126b16>] ? printk+0x48/0x50
[ 46.068535] [<ffffffff816bbe9d>] schedule+0x3d/0x90
[ 46.073500] [<ffffffff81055e38>] do_exit+0x8d8/0xac0
[ 46.078580] [<ffffffff810b6d75>] ? kmsg_dump+0x135/0x180
[ 46.083978] [<ffffffff810b6c62>] ? kmsg_dump+0x22/0x180
[ 46.089291] [<ffffffff81007d8c>] oops_end+0x6c/0x90
[ 46.094257] [<ffffffff81045c13>] no_context+0x153/0x360
[ 46.099571] [<ffffffff812d47a4>] ? delay_tsc+0x94/0xc0
[ 46.104796] [<ffffffff81045f2b>] __bad_area_nosemaphore+0x10b/0x210
[ 46.111149] [<ffffffff81046043>] bad_area_nosemaphore+0x13/0x20
[ 46.117154] [<ffffffff81046487>] __do_page_fault+0x1e7/0x360
[ 46.122901] [<ffffffff81000f70>] ? trace_hardirqs_off_thunk+0x17/0x19
[ 46.129427] [<ffffffff8104663c>] do_page_fault+0xc/0x10
[ 46.134740] [<ffffffff816c38af>] page_fault+0x1f/0x30
[ 46.139879] [<ffffffff81076770>] ? kthread_data+0x10/0x20
[ 46.145364] [<ffffffff810707a1>] wq_worker_sleeping+0x11/0x90
[ 46.151197] [<ffffffff816bb6e6>] __schedule+0x796/0xec0
[ 46.156510] [<ffffffff81055b9b>] ? do_exit+0x63b/0xac0
[ 46.161736] [<ffffffff816bbe9d>] schedule+0x3d/0x90
[ 46.166702] [<ffffffff81055c58>] do_exit+0x6f8/0xac0
[ 46.171754] [<ffffffff81007d8c>] oops_end+0x6c/0x90
[ 46.176721] [<ffffffff81045c13>] no_context+0x153/0x360
[ 46.182033] [<ffffffff81045f2b>] __bad_area_nosemaphore+0x10b/0x210
[ 46.188385] [<ffffffff812f3b77>] ? debug_smp_processor_id+0x17/0x20
[ 46.194730] [<ffffffff81046043>] bad_area_nosemaphore+0x13/0x20
[ 46.200736] [<ffffffff81046487>] __do_page_fault+0x1e7/0x360
[ 46.206481] [<ffffffff81000f70>] ? trace_hardirqs_off_thunk+0x17/0x19
[ 46.213008] [<ffffffff8104663c>] do_page_fault+0xc/0x10
[ 46.218320] [<ffffffff816c38af>] page_fault+0x1f/0x30
[ 46.223461] [<ffffffff81304448>] ? pci_bus_read_config_word+0x98/0xa0
[ 46.229986] [<ffffffff816c0f2b>] ? _raw_spin_unlock_irqrestore+0x4b/0x80
[ 46.236772] [<ffffffff81321296>] ? pci_restore_msi_state+0x196/0x240
[ 46.243212] [<ffffffff81321296>] ? pci_restore_msi_state+0x196/0x240
[ 46.249649] [<ffffffff8130c141>] pci_restore_state.part.18+0xf1/0x250
[ 46.256177] [<ffffffff8130c2b8>] pci_restore_state+0x18/0x20
[ 46.261922] [<ffffffff8130f7fc>] pci_pm_restore_noirq+0x4c/0xd0
[ 46.267928] [<ffffffff8130f7b0>] ? pci_pm_freeze_noirq+0xf0/0xf0
[ 46.274022] [<ffffffff8146e7e7>] dpm_run_callback+0x77/0x2a0
[ 46.279768] [<ffffffff8146eaa3>] device_resume_noirq+0x93/0x150
[ 46.285773] [<ffffffff8146eb7d>] async_resume_noirq+0x1d/0x50
[ 46.291606] [<ffffffff81078a06>] async_run_entry_fn+0x46/0xf0
[ 46.297439] [<ffffffff8106f548>] process_one_work+0x1f8/0x640
[ 46.303271] [<ffffffff8106f4a4>] ? process_one_work+0x154/0x640
[ 46.309278] [<ffffffff8106f9db>] worker_thread+0x4b/0x440
[ 46.314762] [<ffffffff8106f990>] ? process_one_work+0x640/0x640
[ 46.320761] [<ffffffff81075e86>] kthread+0xf6/0x110
[ 46.325726] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 46.332251] [<ffffffff816c1e3f>] ret_from_fork+0x3f/0x70
[ 46.337644] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
Thanks.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Daniel Vetter <daniel@ffwll.ch> |
|---|---|
| Date | 2015-09-23 16:50 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() |
| Message-ID | <qbTwS-78H-19@gated-at.bofh.it> |
| In reply to | #1231274 |
On Wed, Sep 23, 2015 at 10:59:51AM +0200, Borislav Petkov wrote:
> On Wed, Sep 23, 2015 at 09:25:23AM +0200, Daniel Vetter wrote:
> > Strange thing is that I've tested this on a radeon over here and I don't
> > see this backtrace ... wut. Below diff should appease the backtraces at
> > least.
>
> Doesn't look like it.
sorry I sprinkled the locking stuff in the wrong places. Still confused
why the resume side doesn't blow up anywhere ... Oh well. New patch below.
Thanks, Daniel
diff --git a/drivers/gpu/drm/radeon/radeon_device.c b/drivers/gpu/drm/radeon/radeon_device.c
index d8319dae8358..f3f562f6d848 100644
--- a/drivers/gpu/drm/radeon/radeon_device.c
+++ b/drivers/gpu/drm/radeon/radeon_device.c
@@ -1573,10 +1573,12 @@ int radeon_suspend_kms(struct drm_device *dev, bool suspend, bool fbcon)
drm_kms_helper_poll_disable(dev);
+ drm_modeset_lock_all(dev);
/* turn off display hw */
list_for_each_entry(connector, &dev->mode_config.connector_list, head) {
drm_helper_connector_dpms(connector, DRM_MODE_DPMS_OFF);
}
+ drm_modeset_unlock_all(dev);
/* unpin the front buffers and cursors */
list_for_each_entry(crtc, &dev->mode_config.crtc_list, head) {
@@ -1734,9 +1736,11 @@ int radeon_resume_kms(struct drm_device *dev, bool resume, bool fbcon)
if (fbcon) {
drm_helper_resume_force_mode(dev);
/* turn on display hw */
+ drm_modeset_lock_all(dev);
list_for_each_entry(connector, &dev->mode_config.connector_list, head) {
drm_helper_connector_dpms(connector, DRM_MODE_DPMS_ON);
}
+ drm_modeset_unlock_all(dev);
}
drm_kms_helper_poll_enable(dev);
--
Daniel Vetter
Software Engineer, Intel Corporation
http://blog.ffwll.ch
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2015-09-23 18:10 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() |
| Message-ID | <qbUMi-Eg-29@gated-at.bofh.it> |
| In reply to | #1231486 |
On Wed, Sep 23, 2015 at 04:44:50PM +0200, Daniel Vetter wrote:
> sorry I sprinkled the locking stuff in the wrong places. Still confused
> why the resume side doesn't blow up anywhere
But it does:
[ 69.394204] BUG: unable to handle kernel NULL pointer dereference at 0000000000000034
[ 69.402080] IP: [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
[ 69.408624] PGD 4162b8067 PUD 416581067 PMD 0
[ 69.413122] Oops: 0000 [#1] PREEMPT SMP
[ 69.417101] Modules linked in: tun sha256_ssse3 sha256_generic drbg binfmt_misc ipv6 vfat fat fuse dm_crypt dm_mod kv
m_amd kvm crc32_pclmul aesni_intel aes_x86_64 lrw gf128mul glue_helper ablk_helper cryptd amd64_edac_mod edac_mce_amd fa
m15h_power k10temp amdkfd amd_iommu_v2 radeon acpi_cpufreq
[ 69.443647] CPU: 4 PID: 814 Comm: kworker/u16:5 Not tainted 4.3.0-rc2+ #3
[ 69.450430] Hardware name: To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013
[ 69.460336] Workqueue: events_unbound async_run_entry_fn
[ 69.465667] task: ffff88042a255f00 ti: ffff880428a68000 task.ti: ffff880428a68000
[ 69.473145] RIP: 0010:[<ffffffff81321296>] [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
[ 69.482131] RSP: 0018:ffff880428a6bc28 EFLAGS: 00010286
[ 69.487436] RAX: 0000000000000000 RBX: ffff88042a308000 RCX: 0000000000000000
[ 69.494568] RDX: 0000000000000001 RSI: ffffffff81304448 RDI: ffffffff816c7a1b
[ 69.501700] RBP: ffff880428a6bc40 R08: 0000000000000001 R09: 0000000000522000
[ 69.508833] R10: 0000000000000000 R11: 0000000000000000 R12: 0000000000000000
[ 69.515965] R13: ffff88042a3087b0 R14: ffff88042a308010 R15: ffff88042a308038
[ 69.523097] FS: 00007fc91328a700(0000) GS:ffff88042ce00000(0000) knlGS:0000000000000000
[ 69.531185] CS: 0010 DS: 0000 ES: 0000 CR0: 000000008005003b
[ 69.536931] CR2: 0000000000000034 CR3: 00000004164c7000 CR4: 00000000000406e0
[ 69.544061] Stack:
[ 69.546073] 0080002c2a3087b0 0000000000000000 ffff88042a308000 ffff880428a6bc78
[ 69.553525] ffffffff8130c141 ffff88042a308098 ffff88042a308000 0000000000000000
[ 69.560996] ffff8804284e77a8 ffffffff81961ef1 ffff880428a6bc88 ffffffff8130c2b8
[ 69.568450] Call Trace:
[ 69.571044] [<ffffffff8130c141>] pci_restore_state.part.18+0xf1/0x250
[ 69.577706] [<ffffffff8130c2b8>] pci_restore_state+0x18/0x20
[ 69.583591] [<ffffffff8130f7fc>] pci_pm_restore_noirq+0x4c/0xd0
[ 69.589734] [<ffffffff8130f7b0>] ? pci_pm_freeze_noirq+0xf0/0xf0
[ 69.595966] [<ffffffff8146e847>] dpm_run_callback+0x77/0x2a0
[ 69.601850] [<ffffffff8146eb03>] device_resume_noirq+0x93/0x150
[ 69.607994] [<ffffffff8146ebdd>] async_resume_noirq+0x1d/0x50
[ 69.613967] [<ffffffff81078a06>] async_run_entry_fn+0x46/0xf0
[ 69.619939] [<ffffffff8106f548>] process_one_work+0x1f8/0x640
[ 69.625910] [<ffffffff8106f4a4>] ? process_one_work+0x154/0x640
[ 69.632054] [<ffffffff8106f9db>] worker_thread+0x4b/0x440
[ 69.637677] [<ffffffff8106f990>] ? process_one_work+0x640/0x640
[ 69.643822] [<ffffffff81075e86>] kthread+0xf6/0x110
[ 69.648927] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 69.655591] [<ffffffff816c893f>] ret_from_fork+0x3f/0x70
[ 69.661128] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
[ 69.667794] Code: 66 89 4d ee 0f b7 c9 e8 79 41 fe ff 48 89 df e8 d1 7a ce ff 0f b6 53 4b 8b 73 38 48 8d 4d ee 48 8b 7b 10 83 c2 02 e8 1a 31 fe ff <41> 0f b6 4c 24 34 41 8b 54 24 30 be ff ff ff ff c0 e9 04 83 e1
[ 69.687986] RIP [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
[ 69.694772] RSP <ffff880428a6bc28>
[ 69.698412] CR2: 0000000000000034
[ 69.701879] ---[ end trace 814dd8cc56e427ae ]---
This happens at resume - I caught the output over serial - screen is
dead, it doesn't show anything because it simply locks up/panics.
> ... Oh well. New patch below.
Yep, this one took care of the warning in
drm_helper_choose_encoder_dpms(). Thanks!
Now I need to go decypher that NULL ptr deref above.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2015-09-23 18:20 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() |
| Message-ID | <qbUVY-SH-19@gated-at.bofh.it> |
| In reply to | #1231527 |
On Wed, Sep 23, 2015 at 06:06:21PM +0200, Borislav Petkov wrote:
> On Wed, Sep 23, 2015 at 04:44:50PM +0200, Daniel Vetter wrote:
> > sorry I sprinkled the locking stuff in the wrong places. Still confused
> > why the resume side doesn't blow up anywhere
>
> But it does:
>
> [ 69.394204] BUG: unable to handle kernel NULL pointer dereference at 0000000000000034
> [ 69.402080] IP: [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
> [ 69.408624] PGD 4162b8067 PUD 416581067 PMD 0
> [ 69.413122] Oops: 0000 [#1] PREEMPT SMP
> [ 69.417101] Modules linked in: tun sha256_ssse3 sha256_generic drbg binfmt_misc ipv6 vfat fat fuse dm_crypt dm_mod kv
> m_amd kvm crc32_pclmul aesni_intel aes_x86_64 lrw gf128mul glue_helper ablk_helper cryptd amd64_edac_mod edac_mce_amd fa
> m15h_power k10temp amdkfd amd_iommu_v2 radeon acpi_cpufreq
> [ 69.443647] CPU: 4 PID: 814 Comm: kworker/u16:5 Not tainted 4.3.0-rc2+ #3
> [ 69.450430] Hardware name: To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013
> [ 69.460336] Workqueue: events_unbound async_run_entry_fn
> [ 69.465667] task: ffff88042a255f00 ti: ffff880428a68000 task.ti: ffff880428a68000
> [ 69.473145] RIP: 0010:[<ffffffff81321296>] [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
> [ 69.482131] RSP: 0018:ffff880428a6bc28 EFLAGS: 00010286
> [ 69.487436] RAX: 0000000000000000 RBX: ffff88042a308000 RCX: 0000000000000000
> [ 69.494568] RDX: 0000000000000001 RSI: ffffffff81304448 RDI: ffffffff816c7a1b
> [ 69.501700] RBP: ffff880428a6bc40 R08: 0000000000000001 R09: 0000000000522000
> [ 69.508833] R10: 0000000000000000 R11: 0000000000000000 R12: 0000000000000000
> [ 69.515965] R13: ffff88042a3087b0 R14: ffff88042a308010 R15: ffff88042a308038
> [ 69.523097] FS: 00007fc91328a700(0000) GS:ffff88042ce00000(0000) knlGS:0000000000000000
> [ 69.531185] CS: 0010 DS: 0000 ES: 0000 CR0: 000000008005003b
> [ 69.536931] CR2: 0000000000000034 CR3: 00000004164c7000 CR4: 00000000000406e0
> [ 69.544061] Stack:
> [ 69.546073] 0080002c2a3087b0 0000000000000000 ffff88042a308000 ffff880428a6bc78
> [ 69.553525] ffffffff8130c141 ffff88042a308098 ffff88042a308000 0000000000000000
> [ 69.560996] ffff8804284e77a8 ffffffff81961ef1 ffff880428a6bc88 ffffffff8130c2b8
> [ 69.568450] Call Trace:
> [ 69.571044] [<ffffffff8130c141>] pci_restore_state.part.18+0xf1/0x250
> [ 69.577706] [<ffffffff8130c2b8>] pci_restore_state+0x18/0x20
> [ 69.583591] [<ffffffff8130f7fc>] pci_pm_restore_noirq+0x4c/0xd0
> [ 69.589734] [<ffffffff8130f7b0>] ? pci_pm_freeze_noirq+0xf0/0xf0
> [ 69.595966] [<ffffffff8146e847>] dpm_run_callback+0x77/0x2a0
> [ 69.601850] [<ffffffff8146eb03>] device_resume_noirq+0x93/0x150
> [ 69.607994] [<ffffffff8146ebdd>] async_resume_noirq+0x1d/0x50
> [ 69.613967] [<ffffffff81078a06>] async_run_entry_fn+0x46/0xf0
> [ 69.619939] [<ffffffff8106f548>] process_one_work+0x1f8/0x640
> [ 69.625910] [<ffffffff8106f4a4>] ? process_one_work+0x154/0x640
> [ 69.632054] [<ffffffff8106f9db>] worker_thread+0x4b/0x440
> [ 69.637677] [<ffffffff8106f990>] ? process_one_work+0x640/0x640
> [ 69.643822] [<ffffffff81075e86>] kthread+0xf6/0x110
> [ 69.648927] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
> [ 69.655591] [<ffffffff816c893f>] ret_from_fork+0x3f/0x70
> [ 69.661128] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
> [ 69.667794] Code: 66 89 4d ee 0f b7 c9 e8 79 41 fe ff 48 89 df e8 d1 7a ce ff 0f b6 53 4b 8b 73 38 48 8d 4d ee 48 8b 7b 10 83 c2 02 e8 1a 31 fe ff <41> 0f b6 4c 24 34 41 8b 54 24 30 be ff ff ff ff c0 e9 04 83 e1
> [ 69.687986] RIP [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
> [ 69.694772] RSP <ffff880428a6bc28>
> [ 69.698412] CR2: 0000000000000034
> [ 69.701879] ---[ end trace 814dd8cc56e427ae ]---
Ok, after some quick staring, we're at __pci_restore_msi_state():
pci_read_config_word(dev, dev->msi_cap + PCI_MSI_FLAGS, &control);
msi_mask_irq(entry, msi_mask(entry->msi_attrib.multi_cap),
entry->masked);
which is:
.loc 1 411 0
movq %rbx, %rdi # dev,
call arch_restore_msi_irqs #
.LBB1921:
.LBB1922:
.loc 2 902 0
movzbl 75(%rbx), %edx # dev_2(D)->msi_cap, D.31945
movl 56(%rbx), %esi # MEM[(const struct pci_dev *)dev_2(D)].devfn, MEM[(const struct pci_dev *)dev_2(D)].devfn
leaq -18(%rbp), %rcx #, tmp266
movq 16(%rbx), %rdi # MEM[(const struct pci_dev *)dev_2(D)].bus, MEM[(const struct pci_dev *)dev_2(D)].bus
addl $2, %edx #, D.31945
call pci_bus_read_config_word #
.LBE1922:
.LBE1921:
.loc 1 414 0
movzbl 52(%r12), %ecx # *_85, tmp208 <--- faulting insn
movl 48(%r12), %edx # _85->D.27233.D.27231.masked, D.31946
.LBB1923:
.LBB1924:
.loc 1 176 0
movl $-1, %esi #, D.31951
and that %r12 is supposed to contain struct msi_desc *entry in
__pci_restore_msi_state():
entry = irq_get_msi_desc(dev->irq);
which is
.loc 4 654 0
movl 1340(%rdi), %edi # dev_2(D)->irq, dev_2(D)->irq
call irq_get_irq_data #
.loc 4 655 0
testq %rax, %rax # d
je .L405 #,
movq 16(%rax), %rax # d_62->common, d_62->common
movq 16(%rax), %r12 # _63->msi_desc, D.31954
but as we see above %r12 is 0.
For some reason that entry thing in __pci_restore_msi_state() is not
checked for NULL even though irq_get_msi_desc() can return NULL.
Maybe tglx would have an idea...
Hrmmm.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2015-09-26 18:50 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qd0PD-5Bg-9@gated-at.bofh.it> |
| In reply to | #1231539 |
On Wed, Sep 23, 2015 at 06:18:39PM +0200, Borislav Petkov wrote:
> On Wed, Sep 23, 2015 at 06:06:21PM +0200, Borislav Petkov wrote:
> > On Wed, Sep 23, 2015 at 04:44:50PM +0200, Daniel Vetter wrote:
> > > sorry I sprinkled the locking stuff in the wrong places. Still confused
> > > why the resume side doesn't blow up anywhere
> >
> > But it does:
> >
> > [ 69.394204] BUG: unable to handle kernel NULL pointer dereference at 0000000000000034
> > [ 69.402080] IP: [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
> > [ 69.408624] PGD 4162b8067 PUD 416581067 PMD 0
> > [ 69.413122] Oops: 0000 [#1] PREEMPT SMP
> > [ 69.417101] Modules linked in: tun sha256_ssse3 sha256_generic drbg binfmt_misc ipv6 vfat fat fuse dm_crypt dm_mod kv
> > m_amd kvm crc32_pclmul aesni_intel aes_x86_64 lrw gf128mul glue_helper ablk_helper cryptd amd64_edac_mod edac_mce_amd fa
> > m15h_power k10temp amdkfd amd_iommu_v2 radeon acpi_cpufreq
> > [ 69.443647] CPU: 4 PID: 814 Comm: kworker/u16:5 Not tainted 4.3.0-rc2+ #3
> > [ 69.450430] Hardware name: To be filled by O.E.M. To be filled by O.E.M./M5A97 EVO R2.0, BIOS 1503 01/16/2013
> > [ 69.460336] Workqueue: events_unbound async_run_entry_fn
> > [ 69.465667] task: ffff88042a255f00 ti: ffff880428a68000 task.ti: ffff880428a68000
> > [ 69.473145] RIP: 0010:[<ffffffff81321296>] [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
> > [ 69.482131] RSP: 0018:ffff880428a6bc28 EFLAGS: 00010286
> > [ 69.487436] RAX: 0000000000000000 RBX: ffff88042a308000 RCX: 0000000000000000
> > [ 69.494568] RDX: 0000000000000001 RSI: ffffffff81304448 RDI: ffffffff816c7a1b
> > [ 69.501700] RBP: ffff880428a6bc40 R08: 0000000000000001 R09: 0000000000522000
> > [ 69.508833] R10: 0000000000000000 R11: 0000000000000000 R12: 0000000000000000
> > [ 69.515965] R13: ffff88042a3087b0 R14: ffff88042a308010 R15: ffff88042a308038
> > [ 69.523097] FS: 00007fc91328a700(0000) GS:ffff88042ce00000(0000) knlGS:0000000000000000
> > [ 69.531185] CS: 0010 DS: 0000 ES: 0000 CR0: 000000008005003b
> > [ 69.536931] CR2: 0000000000000034 CR3: 00000004164c7000 CR4: 00000000000406e0
> > [ 69.544061] Stack:
> > [ 69.546073] 0080002c2a3087b0 0000000000000000 ffff88042a308000 ffff880428a6bc78
> > [ 69.553525] ffffffff8130c141 ffff88042a308098 ffff88042a308000 0000000000000000
> > [ 69.560996] ffff8804284e77a8 ffffffff81961ef1 ffff880428a6bc88 ffffffff8130c2b8
> > [ 69.568450] Call Trace:
> > [ 69.571044] [<ffffffff8130c141>] pci_restore_state.part.18+0xf1/0x250
> > [ 69.577706] [<ffffffff8130c2b8>] pci_restore_state+0x18/0x20
> > [ 69.583591] [<ffffffff8130f7fc>] pci_pm_restore_noirq+0x4c/0xd0
> > [ 69.589734] [<ffffffff8130f7b0>] ? pci_pm_freeze_noirq+0xf0/0xf0
> > [ 69.595966] [<ffffffff8146e847>] dpm_run_callback+0x77/0x2a0
> > [ 69.601850] [<ffffffff8146eb03>] device_resume_noirq+0x93/0x150
> > [ 69.607994] [<ffffffff8146ebdd>] async_resume_noirq+0x1d/0x50
> > [ 69.613967] [<ffffffff81078a06>] async_run_entry_fn+0x46/0xf0
> > [ 69.619939] [<ffffffff8106f548>] process_one_work+0x1f8/0x640
> > [ 69.625910] [<ffffffff8106f4a4>] ? process_one_work+0x154/0x640
> > [ 69.632054] [<ffffffff8106f9db>] worker_thread+0x4b/0x440
> > [ 69.637677] [<ffffffff8106f990>] ? process_one_work+0x640/0x640
> > [ 69.643822] [<ffffffff81075e86>] kthread+0xf6/0x110
> > [ 69.648927] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
> > [ 69.655591] [<ffffffff816c893f>] ret_from_fork+0x3f/0x70
> > [ 69.661128] [<ffffffff81075d90>] ? kthread_create_on_node+0x1f0/0x1f0
> > [ 69.667794] Code: 66 89 4d ee 0f b7 c9 e8 79 41 fe ff 48 89 df e8 d1 7a ce ff 0f b6 53 4b 8b 73 38 48 8d 4d ee 48 8b 7b 10 83 c2 02 e8 1a 31 fe ff <41> 0f b6 4c 24 34 41 8b 54 24 30 be ff ff ff ff c0 e9 04 83 e1
> > [ 69.687986] RIP [<ffffffff81321296>] pci_restore_msi_state+0x196/0x240
> > [ 69.694772] RSP <ffff880428a6bc28>
> > [ 69.698412] CR2: 0000000000000034
> > [ 69.701879] ---[ end trace 814dd8cc56e427ae ]---
>
> Ok, after some quick staring, we're at __pci_restore_msi_state():
>
> pci_read_config_word(dev, dev->msi_cap + PCI_MSI_FLAGS, &control);
> msi_mask_irq(entry, msi_mask(entry->msi_attrib.multi_cap),
> entry->masked);
>
> which is:
>
> .loc 1 411 0
> movq %rbx, %rdi # dev,
> call arch_restore_msi_irqs #
> .LBB1921:
> .LBB1922:
> .loc 2 902 0
> movzbl 75(%rbx), %edx # dev_2(D)->msi_cap, D.31945
> movl 56(%rbx), %esi # MEM[(const struct pci_dev *)dev_2(D)].devfn, MEM[(const struct pci_dev *)dev_2(D)].devfn
> leaq -18(%rbp), %rcx #, tmp266
> movq 16(%rbx), %rdi # MEM[(const struct pci_dev *)dev_2(D)].bus, MEM[(const struct pci_dev *)dev_2(D)].bus
> addl $2, %edx #, D.31945
> call pci_bus_read_config_word #
> .LBE1922:
> .LBE1921:
> .loc 1 414 0
> movzbl 52(%r12), %ecx # *_85, tmp208 <--- faulting insn
> movl 48(%r12), %edx # _85->D.27233.D.27231.masked, D.31946
> .LBB1923:
> .LBB1924:
> .loc 1 176 0
> movl $-1, %esi #, D.31951
>
> and that %r12 is supposed to contain struct msi_desc *entry in
> __pci_restore_msi_state():
>
> entry = irq_get_msi_desc(dev->irq);
>
> which is
>
> .loc 4 654 0
> movl 1340(%rdi), %edi # dev_2(D)->irq, dev_2(D)->irq
> call irq_get_irq_data #
> .loc 4 655 0
> testq %rax, %rax # d
> je .L405 #,
> movq 16(%rax), %rax # d_62->common, d_62->common
> movq 16(%rax), %r12 # _63->msi_desc, D.31954
>
> but as we see above %r12 is 0.
>
> For some reason that entry thing in __pci_restore_msi_state() is not
> checked for NULL even though irq_get_msi_desc() can return NULL.
Ok, I bisected it.
First of all, Daniel, you didn't see the resume side blow up because
of the NULL ptr deref f*cking up the box much earlier. Once I reverted
the bad commit by hand (it wouldn't revert cleanly) the resume splats
showed.
And in talking about the bad commit, it is this one:
991de2e59090e55c65a7f59a049142e3c480f7bd is the first bad commit
commit 991de2e59090e55c65a7f59a049142e3c480f7bd
Author: Jiang Liu <jiang.liu@linux.intel.com>
Date: Wed Jun 10 16:54:59 2015 +0800
PCI, x86: Implement pcibios_alloc_irq() and pcibios_free_irq()
To support IOAPIC hotplug, we need to allocate PCI IRQ resources on demand
and free them when not used anymore.
Implement pcibios_alloc_irq() and pcibios_free_irq() to dynamically
allocate and free PCI IRQs.
Remove mp_should_keep_irq(), which is no longer used.
[bhelgaas: changelog]
Signed-off-by: Jiang Liu <jiang.liu@linux.intel.com>
Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
Acked-by: Thomas Gleixner <tglx@linutronix.de>
:040000 040000 765e2d5232d53247ec260b34b51589c3bccb36ae f680234a27685e94b1a35ae2a7218f8eafa9071a M arch
:040000 040000 d55a682bcde72682e883365e88ad1df6186fd54d f82c470a04a6845fcf5e0aa934512c75628f798d M drivers
Jiang, you have to stop breaking my box with your changes. This is
maybe the third time I'm bisecting fallout from your patches. If you're
touching all x86, you need to test on an AMD box too. Like everyone else
testing on the hardware their changes affect. It is that simple.
Anyway, reverting that commit by hand fixes my resume splat.
Here's the partial revert I did by hand:
---
diff --git a/arch/x86/include/asm/pci_x86.h b/arch/x86/include/asm/pci_x86.h
index fa1195dae425..164e3f8d3c3d 100644
--- a/arch/x86/include/asm/pci_x86.h
+++ b/arch/x86/include/asm/pci_x86.h
@@ -93,6 +93,8 @@ extern raw_spinlock_t pci_config_lock;
extern int (*pcibios_enable_irq)(struct pci_dev *dev);
extern void (*pcibios_disable_irq)(struct pci_dev *dev);
+extern bool mp_should_keep_irq(struct device *dev);
+
struct pci_raw_ops {
int (*read)(unsigned int domain, unsigned int bus, unsigned int devfn,
int reg, int len, u32 *val);
diff --git a/arch/x86/pci/common.c b/arch/x86/pci/common.c
index 09d3afc0a181..3bff24438b00 100644
--- a/arch/x86/pci/common.c
+++ b/arch/x86/pci/common.c
@@ -672,20 +672,22 @@ int pcibios_add_device(struct pci_dev *dev)
return 0;
}
-int pcibios_alloc_irq(struct pci_dev *dev)
+int pcibios_enable_device(struct pci_dev *dev, int mask)
{
- return pcibios_enable_irq(dev);
-}
+ int err;
-void pcibios_free_irq(struct pci_dev *dev)
-{
- if (pcibios_disable_irq)
- pcibios_disable_irq(dev);
+ if ((err = pci_enable_resources(dev, mask)) < 0)
+ return err;
+
+ if (!pci_dev_msi_enabled(dev))
+ return pcibios_enable_irq(dev);
+ return 0;
}
-int pcibios_enable_device(struct pci_dev *dev, int mask)
+void pcibios_disable_device (struct pci_dev *dev)
{
- return pci_enable_resources(dev, mask);
+ if (!pci_dev_msi_enabled(dev) && pcibios_disable_irq)
+ pcibios_disable_irq(dev);
}
int pci_ext_cfg_avail(void)
diff --git a/arch/x86/pci/irq.c b/arch/x86/pci/irq.c
index 32e70343e6fd..f229834b36d4 100644
--- a/arch/x86/pci/irq.c
+++ b/arch/x86/pci/irq.c
@@ -1186,6 +1186,18 @@ void pcibios_penalize_isa_irq(int irq, int active)
pirq_penalize_isa_irq(irq, active);
}
+bool mp_should_keep_irq(struct device *dev)
+{
+ if (dev->power.is_prepared)
+ return true;
+#ifdef CONFIG_PM
+ if (dev->power.runtime_status == RPM_SUSPENDING)
+ return true;
+#endif
+
+ return false;
+}
+
static int pirq_enable_irq(struct pci_dev *dev)
{
u8 pin = 0;
@@ -1258,7 +1270,8 @@ static int pirq_enable_irq(struct pci_dev *dev)
static void pirq_disable_irq(struct pci_dev *dev)
{
- if (io_apic_assign_pci_irqs && pci_has_managed_irq(dev)) {
+ if (io_apic_assign_pci_irqs && !mp_should_keep_irq(&dev->dev) &&
+ dev->irq_managed && dev->irq) {
mp_unmap_irq(dev->irq);
pci_reset_managed_irq(dev);
}
diff --git a/drivers/acpi/pci_irq.c b/drivers/acpi/pci_irq.c
index 6da0f9beab19..d8a3f49a960c 100644
--- a/drivers/acpi/pci_irq.c
+++ b/drivers/acpi/pci_irq.c
@@ -479,6 +479,14 @@ void acpi_pci_irq_disable(struct pci_dev *dev)
if (!pin || !pci_has_managed_irq(dev))
return;
+ /* Keep IOAPIC pin configuration when suspending */
+ if (dev->dev.power.is_prepared)
+ return;
+#ifdef CONFIG_PM
+ if (dev->dev.power.runtime_status == RPM_SUSPENDING)
+ return;
+#endif
+
entry = acpi_pci_irq_lookup(dev, pin);
if (!entry)
return;
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jiang Liu <jiang.liu@linux.intel.com> |
|---|---|
| Date | 2015-09-29 11:00 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qdYVt-1z1-19@gated-at.bofh.it> |
| In reply to | #1233222 |
[Multipart message — attachments visible in raw view] — view raw
On 2015/9/27 0:46, Borislav Petkov wrote:
> On Wed, Sep 23, 2015 at 06:18:39PM +0200, Borislav Petkov wrote:
>> On Wed, Sep 23, 2015 at 06:06:21PM +0200, Borislav Petkov wrote:
>>> On Wed, Sep 23, 2015 at 04:44:50PM +0200, Daniel Vetter wrote:
>>>> sorry I sprinkled the locking stuff in the wrong places. Still confused
>>>> why the resume side doesn't blow up anywhere
>>>
>>> But it does:
<snit>
>
> Ok, I bisected it.
>
> First of all, Daniel, you didn't see the resume side blow up because
> of the NULL ptr deref f*cking up the box much earlier. Once I reverted
> the bad commit by hand (it wouldn't revert cleanly) the resume splats
> showed.
>
> And in talking about the bad commit, it is this one:
>
> 991de2e59090e55c65a7f59a049142e3c480f7bd is the first bad commit
> commit 991de2e59090e55c65a7f59a049142e3c480f7bd
> Author: Jiang Liu <jiang.liu@linux.intel.com>
> Date: Wed Jun 10 16:54:59 2015 +0800
>
> PCI, x86: Implement pcibios_alloc_irq() and pcibios_free_irq()
>
> To support IOAPIC hotplug, we need to allocate PCI IRQ resources on demand
> and free them when not used anymore.
>
> Implement pcibios_alloc_irq() and pcibios_free_irq() to dynamically
> allocate and free PCI IRQs.
>
> Remove mp_should_keep_irq(), which is no longer used.
>
> [bhelgaas: changelog]
> Signed-off-by: Jiang Liu <jiang.liu@linux.intel.com>
> Signed-off-by: Bjorn Helgaas <bhelgaas@google.com>
> Acked-by: Thomas Gleixner <tglx@linutronix.de>
>
> :040000 040000 765e2d5232d53247ec260b34b51589c3bccb36ae f680234a27685e94b1a35ae2a7218f8eafa9071a M arch
> :040000 040000 d55a682bcde72682e883365e88ad1df6186fd54d f82c470a04a6845fcf5e0aa934512c75628f798d M drivers
>
> Jiang, you have to stop breaking my box with your changes. This is
> maybe the third time I'm bisecting fallout from your patches. If you're
> touching all x86, you need to test on an AMD box too. Like everyone else
> testing on the hardware their changes affect. It is that simple.
Hi Boris and Daniel,
Sorry for the regression!
I have tried to reproduce the regression by doing
suspend/resume with a laptop, but failed. The PCI MSI suspend/resume
code work as expected. And I have checked msi.c and radeon driver,
but haven't gotten any hint about the cause.
So could you please help to apply the attached debug patch
to gather more information about the regression?
Thanks!
Gerry
>
> Anyway, reverting that commit by hand fixes my resume splat.
>
> Here's the partial revert I did by hand:
>
> ---
> diff --git a/arch/x86/include/asm/pci_x86.h b/arch/x86/include/asm/pci_x86.h
> index fa1195dae425..164e3f8d3c3d 100644
> --- a/arch/x86/include/asm/pci_x86.h
> +++ b/arch/x86/include/asm/pci_x86.h
> @@ -93,6 +93,8 @@ extern raw_spinlock_t pci_config_lock;
> extern int (*pcibios_enable_irq)(struct pci_dev *dev);
> extern void (*pcibios_disable_irq)(struct pci_dev *dev);
>
> +extern bool mp_should_keep_irq(struct device *dev);
> +
> struct pci_raw_ops {
> int (*read)(unsigned int domain, unsigned int bus, unsigned int devfn,
> int reg, int len, u32 *val);
> diff --git a/arch/x86/pci/common.c b/arch/x86/pci/common.c
> index 09d3afc0a181..3bff24438b00 100644
> --- a/arch/x86/pci/common.c
> +++ b/arch/x86/pci/common.c
> @@ -672,20 +672,22 @@ int pcibios_add_device(struct pci_dev *dev)
> return 0;
> }
>
> -int pcibios_alloc_irq(struct pci_dev *dev)
> +int pcibios_enable_device(struct pci_dev *dev, int mask)
> {
> - return pcibios_enable_irq(dev);
> -}
> + int err;
>
> -void pcibios_free_irq(struct pci_dev *dev)
> -{
> - if (pcibios_disable_irq)
> - pcibios_disable_irq(dev);
> + if ((err = pci_enable_resources(dev, mask)) < 0)
> + return err;
> +
> + if (!pci_dev_msi_enabled(dev))
> + return pcibios_enable_irq(dev);
> + return 0;
> }
>
> -int pcibios_enable_device(struct pci_dev *dev, int mask)
> +void pcibios_disable_device (struct pci_dev *dev)
> {
> - return pci_enable_resources(dev, mask);
> + if (!pci_dev_msi_enabled(dev) && pcibios_disable_irq)
> + pcibios_disable_irq(dev);
> }
>
> int pci_ext_cfg_avail(void)
> diff --git a/arch/x86/pci/irq.c b/arch/x86/pci/irq.c
> index 32e70343e6fd..f229834b36d4 100644
> --- a/arch/x86/pci/irq.c
> +++ b/arch/x86/pci/irq.c
> @@ -1186,6 +1186,18 @@ void pcibios_penalize_isa_irq(int irq, int active)
> pirq_penalize_isa_irq(irq, active);
> }
>
> +bool mp_should_keep_irq(struct device *dev)
> +{
> + if (dev->power.is_prepared)
> + return true;
> +#ifdef CONFIG_PM
> + if (dev->power.runtime_status == RPM_SUSPENDING)
> + return true;
> +#endif
> +
> + return false;
> +}
> +
> static int pirq_enable_irq(struct pci_dev *dev)
> {
> u8 pin = 0;
> @@ -1258,7 +1270,8 @@ static int pirq_enable_irq(struct pci_dev *dev)
>
> static void pirq_disable_irq(struct pci_dev *dev)
> {
> - if (io_apic_assign_pci_irqs && pci_has_managed_irq(dev)) {
> + if (io_apic_assign_pci_irqs && !mp_should_keep_irq(&dev->dev) &&
> + dev->irq_managed && dev->irq) {
> mp_unmap_irq(dev->irq);
> pci_reset_managed_irq(dev);
> }
> diff --git a/drivers/acpi/pci_irq.c b/drivers/acpi/pci_irq.c
> index 6da0f9beab19..d8a3f49a960c 100644
> --- a/drivers/acpi/pci_irq.c
> +++ b/drivers/acpi/pci_irq.c
> @@ -479,6 +479,14 @@ void acpi_pci_irq_disable(struct pci_dev *dev)
> if (!pin || !pci_has_managed_irq(dev))
> return;
>
> + /* Keep IOAPIC pin configuration when suspending */
> + if (dev->dev.power.is_prepared)
> + return;
> +#ifdef CONFIG_PM
> + if (dev->dev.power.runtime_status == RPM_SUSPENDING)
> + return;
> +#endif
> +
> entry = acpi_pci_irq_lookup(dev, pin);
> if (!entry)
> return;
>
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2015-09-29 13:00 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qe0NA-4fJ-17@gated-at.bofh.it> |
| In reply to | #1234869 |
On Tue, Sep 29, 2015 at 04:50:36PM +0800, Jiang Liu wrote:
> So could you please help to apply the attached debug patch to gather
> more information about the regression?
Sure, just did.
I'm sending you a full s/r cycle attempt caught over serial in a private
message.
Thanks.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jiang Liu <jiang.liu@linux.intel.com> |
|---|---|
| Date | 2015-09-30 09:50 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qekjf-6VR-1@gated-at.bofh.it> |
| In reply to | #1234969 |
[Multipart message — attachments visible in raw view] — view raw
n 2015/9/29 18:51, Borislav Petkov wrote: > On Tue, Sep 29, 2015 at 04:50:36PM +0800, Jiang Liu wrote: >> So could you please help to apply the attached debug patch to gather >> more information about the regression? > > Sure, just did. > > I'm sending you a full s/r cycle attempt caught over serial in a private > message. Hi Boris, From the log file, we got to know that the NULL pointer dereference was caused by AMD IOMMU device. For normal MSI-enabled PCI devices, we get valid irq numbers such as: [ 74.661170] ahci 0000:04:00.0: irqdomain: freeze msi 1 irq28 [ 74.661297] radeon 0000:01:00.0: irqdomain: freeze msi 1 irq47 But for AMD IOMMU device, we got an invalid irq number(0) after enabling MSI as: [ 74.662488] pci 0000:00:00.2: irqdomain: freeze msi 1 irq0 which then caused NULL pointer deference when __pci_restore_msi_state() gets called by system resume code. So we need to figure out why we got irq number 0 after enabling MSI for AMD IOMMU device. The only hint I got is that iommu driver just grabbing the PCI device without providing a PCI device driver for IOMMU PCI device, we have solved a similar case for eata driver. So could you please help to apply this debug patch to gather more info and send me /proc/interrupts? Thanks! Gerry O> > Thanks. >
[toc] | [prev] | [next] | [standalone]
| From | Joerg Roedel <joro@8bytes.org> |
|---|---|
| Date | 2015-09-30 14:50 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qeoZA-5g3-7@gated-at.bofh.it> |
| In reply to | #1235853 |
On Wed, Sep 30, 2015 at 03:45:39PM +0800, Jiang Liu wrote: > So we need to figure out why we got irq number 0 after enabling > MSI for AMD IOMMU device. The only hint I got is that iommu driver just > grabbing the PCI device without providing a PCI device driver for IOMMU > PCI device, we have solved a similar case for eata driver. So could you > please help to apply this debug patch to gather more info and send me > /proc/interrupts? I think I have an idea on how dev->irq got 0 after pci_enable_msi(). The PCI probe code calls pcibios_alloc_irq() and after a failed probe it calls pcibios_free_irq(), which sets dev->irq to 0. The AMD IOMMU driver does not register a pci_driver for itself, it just doesn't make sense for it. But the PCI device containing the IOMMU gets probed later, which fails because there is no driver for it. So the following call to pcibios_free_irq() clears dev->irq, so that it is 0 on the next resume. Does that make sense? Joerg -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jiang Liu <jiang.liu@linux.intel.com> |
|---|---|
| Date | 2015-09-30 19:10 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qet3c-2QO-19@gated-at.bofh.it> |
| In reply to | #1236252 |
On 2015/9/30 20:44, Joerg Roedel wrote: > On Wed, Sep 30, 2015 at 03:45:39PM +0800, Jiang Liu wrote: >> So we need to figure out why we got irq number 0 after enabling >> MSI for AMD IOMMU device. The only hint I got is that iommu driver just >> grabbing the PCI device without providing a PCI device driver for IOMMU >> PCI device, we have solved a similar case for eata driver. So could you >> please help to apply this debug patch to gather more info and send me >> /proc/interrupts? > > I think I have an idea on how dev->irq got 0 after pci_enable_msi(). The > PCI probe code calls pcibios_alloc_irq() and after a failed probe it calls > pcibios_free_irq(), which sets dev->irq to 0. > The AMD IOMMU driver does not register a pci_driver for itself, it just > doesn't make sense for it. But the PCI device containing the IOMMU gets > probed later, which fails because there is no driver for it. So the > following call to pcibios_free_irq() clears dev->irq, so that it is 0 on > the next resume. Does that make sense? Thanks Joerg, that makes sense. If some driver tries to binding to the IOMMU device, it will trigger the scenario as you described. For example, Xen backend driver will try to probe all PCI devices if enabled. I will do more investigation tomorrow. Thanks! Gerry -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2015-09-30 19:40 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qetwd-3ox-1@gated-at.bofh.it> |
| In reply to | #1236547 |
On Thu, Oct 01, 2015 at 01:00:44AM +0800, Jiang Liu wrote:
> Thanks Joerg, that makes sense. If some driver tries to binding to
> the IOMMU device, it will trigger the scenario as you described. For
> example, Xen backend driver will try to probe all PCI devices if
> enabled. I will do more investigation tomorrow.
Right, so this fixes the issue on my box, courtesy of Joerg. WE
basically don't disable the IRQ on MSI-enabled devices. The AMD IOMMU
uses a barebones PCI device but not a PCI driver, which would be an
overkill.
---
diff --git a/arch/x86/pci/common.c b/arch/x86/pci/common.c
index 09d3afc..29ec2eb 100644
--- a/arch/x86/pci/common.c
+++ b/arch/x86/pci/common.c
@@ -674,12 +674,15 @@ int pcibios_add_device(struct pci_dev *dev)
int pcibios_alloc_irq(struct pci_dev *dev)
{
+ if (pci_dev_msi_enabled(dev))
+ return 0;
+
return pcibios_enable_irq(dev);
}
void pcibios_free_irq(struct pci_dev *dev)
{
- if (pcibios_disable_irq)
+ if (!pci_dev_msi_enabled(dev) && pcibios_disable_irq)
pcibios_disable_irq(dev);
}
--
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Joerg Roedel <joro@8bytes.org> |
|---|---|
| Date | 2015-09-30 20:10 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qetZf-4bK-1@gated-at.bofh.it> |
| In reply to | #1236574 |
On Wed, Sep 30, 2015 at 07:36:19PM +0200, Borislav Petkov wrote: > Right, so this fixes the issue on my box, courtesy of Joerg. WE > basically don't disable the IRQ on MSI-enabled devices. The AMD IOMMU > uses a barebones PCI device but not a PCI driver, which would be an > overkill. Well, not only overkill, but actually harmful. As I just wrote to Jiang, a device can be forcibly unbound from its driver, which is something we don't want for the IOMMU. Joerg -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jiang Liu <jiang.liu@linux.intel.com> |
|---|---|
| Date | 2015-10-03 09:40 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qfpAd-3DX-17@gated-at.bofh.it> |
| In reply to | #1236574 |
On 2015/10/1 1:36, Borislav Petkov wrote:
> On Thu, Oct 01, 2015 at 01:00:44AM +0800, Jiang Liu wrote:
>> Thanks Joerg, that makes sense. If some driver tries to binding to
>> the IOMMU device, it will trigger the scenario as you described. For
>> example, Xen backend driver will try to probe all PCI devices if
>> enabled. I will do more investigation tomorrow.
>
> Right, so this fixes the issue on my box, courtesy of Joerg. WE
> basically don't disable the IRQ on MSI-enabled devices. The AMD IOMMU
> uses a barebones PCI device but not a PCI driver, which would be an
> overkill.
>
> ---
> diff --git a/arch/x86/pci/common.c b/arch/x86/pci/common.c
> index 09d3afc..29ec2eb 100644
> --- a/arch/x86/pci/common.c
> +++ b/arch/x86/pci/common.c
> @@ -674,12 +674,15 @@ int pcibios_add_device(struct pci_dev *dev)
>
> int pcibios_alloc_irq(struct pci_dev *dev)
> {
> + if (pci_dev_msi_enabled(dev))
> + return 0;
We may return -EBUSY here to reject the probe operation. It
doesn't make sense to continue the probe if MSI is already enabled,
tt also helps to avoid calling pcibios_free_irq() in function
pci_device_probe().
> +
> return pcibios_enable_irq(dev);
> }
>
> void pcibios_free_irq(struct pci_dev *dev)
> {
> - if (pcibios_disable_irq)
> + if (!pci_dev_msi_enabled(dev) && pcibios_disable_irq)
The above change is not needed, pcibios_disable_irq() will
first check !pci_has_managed_irq(dev) before actually freeing
PCI irq. pci_has_managed_irq(dev) only returns true if
pcibios_alloc_irq() succeeds.
So to summary, I think we only need following change to fix the
regression:
int pcibios_alloc_irq(struct pci_dev *dev)
{
+ if (pci_dev_msi_enabled(dev))
+ return -EBUSY;
What do you think?
Thanks!
Gerry
> pcibios_disable_irq(dev);
> }
> --
>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2015-10-03 11:40 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qfrsm-6jk-15@gated-at.bofh.it> |
| In reply to | #1238751 |
On Sat, Oct 03, 2015 at 03:36:35PM +0800, Jiang Liu wrote:
> The above change is not needed, pcibios_disable_irq() will
> first check !pci_has_managed_irq(dev) before actually freeing
> PCI irq. pci_has_managed_irq(dev) only returns true if
> pcibios_alloc_irq() succeeds.
>
> So to summary, I think we only need following change to fix the
> regression:
> int pcibios_alloc_irq(struct pci_dev *dev)
> {
> + if (pci_dev_msi_enabled(dev))
> + return -EBUSY;
>
> What do you think?
Yap, that works too. I've got only this ontop of 4.3+tip:
---
diff --git a/arch/x86/pci/common.c b/arch/x86/pci/common.c
index dc78a4a9a466..a4687aa6c1fb 100644
--- a/arch/x86/pci/common.c
+++ b/arch/x86/pci/common.c
@@ -675,6 +675,9 @@ int pcibios_add_device(struct pci_dev *dev)
int pcibios_alloc_irq(struct pci_dev *dev)
{
+ if (pci_dev_msi_enabled(dev))
+ return -EBUSY;
+
return pcibios_enable_irq(dev);
}
---
and it suspend+resumed fine.
I guess it is time for Joerg to write a proper patch. :-)
Thanks.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Joerg Roedel <joro@8bytes.org> |
|---|---|
| Date | 2015-10-05 12:10 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qgaSu-4bV-15@gated-at.bofh.it> |
| In reply to | #1238751 |
Hi Jiang,
On Sat, Oct 03, 2015 at 03:36:35PM +0800, Jiang Liu wrote:
> So to summary, I think we only need following change to fix the
> regression:
> int pcibios_alloc_irq(struct pci_dev *dev)
> {
> + if (pci_dev_msi_enabled(dev))
> + return -EBUSY;
>
> What do you think?
Yes, that works too and has the added benefit that no driver can attach
to the iommu device and get in the way of the driver.
Will you send the patch for this change or should I do it?
Joerg
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Joerg Roedel <joro@8bytes.org> |
|---|---|
| Date | 2015-09-30 20:10 +0200 |
| Subject | Re: WARNING: CPU: 4 PID: 863 at include/drm/drm_crtc.h:1577 drm_helper_choose_encoder_dpms+0x88/0x90() - evildoer found and neutralized |
| Message-ID | <qetZg-4bK-21@gated-at.bofh.it> |
| In reply to | #1236547 |
On Thu, Oct 01, 2015 at 01:00:44AM +0800, Jiang Liu wrote:
> Thanks Joerg, that makes sense. If some driver tries to binding to the
> IOMMU device, it will trigger the scenario as you described. For
> example, Xen backend driver will try to probe all PCI devices
> if enabled. I will do more investigation tomorrow.
Not only that, the probe code looks like this in __pci_device_probe:
error = -ENODEV;
id = pci_match_device(drv, pci_dev);
if (id)
error = pci_call_probe(drv, pci_dev, id);
if (error >= 0)
error = 0;
The pci_match_device() function will always return NULL for the iommu
pci_dev, because no driver matches the ids of it. So the function
returns -ENODEV, which will be handled in the caller (pci_device_probe):
error = pcibios_alloc_irq(pci_dev);
if (error < 0)
return error;
pci_dev_get(pci_dev);
error = __pci_device_probe(drv, pci_dev);
if (error) {
pcibios_free_irq(pci_dev);
pci_dev_put(pci_dev);
}
For the IOMMU pci_dev a pcibios-irq will be allocated (if there is one,
like on Boris' system) and because __pci_device_probe returns -ENODEV it
will be freed again with pcibios_free_irq().
The pcibios_free_irq() function will set dev->irq = 0, which overwrites
the value that pci_enable_msi() wrote there. So later in suspend/resume
code the msi-handling part tries to fetch the irq-descriptor for the
wrong irq (which is NULL) and causes the crash.
The issue got introduced because with your changes pci_enable_msi() is
only allowed after a pci-device was successfully probed by the driver.
But this assumption is not true, as the AMD IOMMU driver does not
register as a pci-driver.
Registering a pci-driver would actually be harmful, because a device can
be forcibly unbound from its driver, which would be pretty bad for an
IOMMU in the running system.
So the right fix is to allow pci_enable_msi() for pci-devices not
registered against a driver. The fix I sent Boris has issues (I think it
leaks pcibios irqs when MSI is in use), but was thinking about fixing it
in pci_device_probe by not allocating a pcibios-irq when MSI is already
active. What do you think?
Regards,
Joerg
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web