Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1556729 > unrolled thread
| Started by | Chris Wilson <chris@chris-wilson.co.uk> |
|---|---|
| First post | 2017-01-11 18:10 +0100 |
| Last post | 2017-01-13 15:50 +0100 |
| Articles | 4 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [Intel-gfx] GPU hang with kernel 4.10rc3 Chris Wilson <chris@chris-wilson.co.uk> - 2017-01-11 18:10 +0100
Re: [Intel-gfx] GPU hang with kernel 4.10rc3 Juergen Gross <jgross@suse.com> - 2017-01-12 07:10 +0100
Re: [Intel-gfx] GPU hang with kernel 4.10rc3 Chris Wilson <chris@chris-wilson.co.uk> - 2017-01-12 10:30 +0100
Re: [Intel-gfx] GPU hang with kernel 4.10rc3 Juergen Gross <jgross@suse.com> - 2017-01-13 15:50 +0100
| From | Chris Wilson <chris@chris-wilson.co.uk> |
|---|---|
| Date | 2017-01-11 18:10 +0100 |
| Subject | Re: [Intel-gfx] GPU hang with kernel 4.10rc3 |
| Message-ID | <sYuzn-1aF-3@gated-at.bofh.it> |
On Wed, Jan 11, 2017 at 05:33:34PM +0100, Juergen Gross wrote: > With kernel 4.10rc3 running as Xen dm0 I get at each boot: > > [ 49.213697] [drm] GPU HANG: ecode 7:0:0x3d1d3d3d, in gnome-shell > [1431], reason: Hang on render ring, action: reset > [ 49.213699] [drm] GPU hangs can indicate a bug anywhere in the entire > gfx stack, including userspace. > [ 49.213700] [drm] Please file a _new_ bug report on > bugs.freedesktop.org against DRI -> DRM/Intel > [ 49.213700] [drm] drm/i915 developers can then reassign to the right > component if it's not a kernel issue. > [ 49.213700] [drm] The gpu crash dump is required to analyze gpu > hangs, so please always attach it. > [ 49.213701] [drm] GPU crash dump saved to /sys/class/drm/card0/error > [ 49.213755] drm/i915: Resetting chip after gpu hang > [ 60.213769] drm/i915: Resetting chip after gpu hang > [ 71.189737] drm/i915: Resetting chip after gpu hang > [ 82.165747] drm/i915: Resetting chip after gpu hang > [ 93.205727] drm/i915: Resetting chip after gpu hang > > The dump is attached. That's a nasty one. The first couple of pages of the batchbuffer appear to be overwritten. (Full of 0xc2c2c2c2, i.e. probably pixel data.) That may be a concurrent write by either the GPU or CPU, or we may have incorrected mapped a set of pages. That it doesn't recovered suggests that the corruption occurs frequently, probably on every request/batch. Is this a new bug? Bisection would be the fastest way to triage it. -Chris -- Chris Wilson, Intel Open Source Technology Centre
[toc] | [next] | [standalone]
| From | Juergen Gross <jgross@suse.com> |
|---|---|
| Date | 2017-01-12 07:10 +0100 |
| Message-ID | <sYGKe-t5-17@gated-at.bofh.it> |
| In reply to | #1556729 |
On 11/01/17 18:08, Chris Wilson wrote: > On Wed, Jan 11, 2017 at 05:33:34PM +0100, Juergen Gross wrote: >> With kernel 4.10rc3 running as Xen dm0 I get at each boot: >> >> [ 49.213697] [drm] GPU HANG: ecode 7:0:0x3d1d3d3d, in gnome-shell >> [1431], reason: Hang on render ring, action: reset >> [ 49.213699] [drm] GPU hangs can indicate a bug anywhere in the entire >> gfx stack, including userspace. >> [ 49.213700] [drm] Please file a _new_ bug report on >> bugs.freedesktop.org against DRI -> DRM/Intel >> [ 49.213700] [drm] drm/i915 developers can then reassign to the right >> component if it's not a kernel issue. >> [ 49.213700] [drm] The gpu crash dump is required to analyze gpu >> hangs, so please always attach it. >> [ 49.213701] [drm] GPU crash dump saved to /sys/class/drm/card0/error >> [ 49.213755] drm/i915: Resetting chip after gpu hang >> [ 60.213769] drm/i915: Resetting chip after gpu hang >> [ 71.189737] drm/i915: Resetting chip after gpu hang >> [ 82.165747] drm/i915: Resetting chip after gpu hang >> [ 93.205727] drm/i915: Resetting chip after gpu hang >> >> The dump is attached. > > That's a nasty one. The first couple of pages of the batchbuffer appear > to be overwritten. (Full of 0xc2c2c2c2, i.e. probably pixel data.) That > may be a concurrent write by either the GPU or CPU, or we may have > incorrected mapped a set of pages. That it doesn't recovered suggests > that the corruption occurs frequently, probably on every request/batch. I hoped someone would have an idea already. > Is this a new bug? Bisection would be the fastest way to triage it. Commit 7453c549f was still okay. Starting bisect now (2882 commits, 12 steps) ... Juergen
[toc] | [prev] | [next] | [standalone]
| From | Chris Wilson <chris@chris-wilson.co.uk> |
|---|---|
| Date | 2017-01-12 10:30 +0100 |
| Message-ID | <sYJRM-2er-19@gated-at.bofh.it> |
| In reply to | #1557132 |
On Thu, Jan 12, 2017 at 07:03:25AM +0100, Juergen Gross wrote: > On 11/01/17 18:08, Chris Wilson wrote: > > On Wed, Jan 11, 2017 at 05:33:34PM +0100, Juergen Gross wrote: > >> With kernel 4.10rc3 running as Xen dm0 I get at each boot: > >> > >> [ 49.213697] [drm] GPU HANG: ecode 7:0:0x3d1d3d3d, in gnome-shell > >> [1431], reason: Hang on render ring, action: reset > >> [ 49.213699] [drm] GPU hangs can indicate a bug anywhere in the entire > >> gfx stack, including userspace. > >> [ 49.213700] [drm] Please file a _new_ bug report on > >> bugs.freedesktop.org against DRI -> DRM/Intel > >> [ 49.213700] [drm] drm/i915 developers can then reassign to the right > >> component if it's not a kernel issue. > >> [ 49.213700] [drm] The gpu crash dump is required to analyze gpu > >> hangs, so please always attach it. > >> [ 49.213701] [drm] GPU crash dump saved to /sys/class/drm/card0/error > >> [ 49.213755] drm/i915: Resetting chip after gpu hang > >> [ 60.213769] drm/i915: Resetting chip after gpu hang > >> [ 71.189737] drm/i915: Resetting chip after gpu hang > >> [ 82.165747] drm/i915: Resetting chip after gpu hang > >> [ 93.205727] drm/i915: Resetting chip after gpu hang > >> > >> The dump is attached. > > > > That's a nasty one. The first couple of pages of the batchbuffer appear > > to be overwritten. (Full of 0xc2c2c2c2, i.e. probably pixel data.) That > > may be a concurrent write by either the GPU or CPU, or we may have > > incorrected mapped a set of pages. That it doesn't recovered suggests > > that the corruption occurs frequently, probably on every request/batch. > > I hoped someone would have an idea already. Sorry, first report of something like this in a long time (that I can remember at least). And the problem is that it can be anything from a coherency to a concurrency issue, so no one patch springs to mind. Thankfully it appears to be kernel related. -Chris -- Chris Wilson, Intel Open Source Technology Centre
[toc] | [prev] | [next] | [standalone]
| From | Juergen Gross <jgross@suse.com> |
|---|---|
| Date | 2017-01-13 15:50 +0100 |
| Message-ID | <sZbkZ-23t-3@gated-at.bofh.it> |
| In reply to | #1557231 |
On 12/01/17 10:21, Chris Wilson wrote:
> On Thu, Jan 12, 2017 at 07:03:25AM +0100, Juergen Gross wrote:
>> On 11/01/17 18:08, Chris Wilson wrote:
>>> On Wed, Jan 11, 2017 at 05:33:34PM +0100, Juergen Gross wrote:
>>>> With kernel 4.10rc3 running as Xen dm0 I get at each boot:
>>>>
>>>> [ 49.213697] [drm] GPU HANG: ecode 7:0:0x3d1d3d3d, in gnome-shell
>>>> [1431], reason: Hang on render ring, action: reset
>>>> [ 49.213699] [drm] GPU hangs can indicate a bug anywhere in the entire
>>>> gfx stack, including userspace.
>>>> [ 49.213700] [drm] Please file a _new_ bug report on
>>>> bugs.freedesktop.org against DRI -> DRM/Intel
>>>> [ 49.213700] [drm] drm/i915 developers can then reassign to the right
>>>> component if it's not a kernel issue.
>>>> [ 49.213700] [drm] The gpu crash dump is required to analyze gpu
>>>> hangs, so please always attach it.
>>>> [ 49.213701] [drm] GPU crash dump saved to /sys/class/drm/card0/error
>>>> [ 49.213755] drm/i915: Resetting chip after gpu hang
>>>> [ 60.213769] drm/i915: Resetting chip after gpu hang
>>>> [ 71.189737] drm/i915: Resetting chip after gpu hang
>>>> [ 82.165747] drm/i915: Resetting chip after gpu hang
>>>> [ 93.205727] drm/i915: Resetting chip after gpu hang
>>>>
>>>> The dump is attached.
>>>
>>> That's a nasty one. The first couple of pages of the batchbuffer appear
>>> to be overwritten. (Full of 0xc2c2c2c2, i.e. probably pixel data.) That
>>> may be a concurrent write by either the GPU or CPU, or we may have
>>> incorrected mapped a set of pages. That it doesn't recovered suggests
>>> that the corruption occurs frequently, probably on every request/batch.
>>
>> I hoped someone would have an idea already.
>
> Sorry, first report of something like this in a long time (that I can
> remember at least). And the problem is that it can be anything from a
> coherency to a concurrency issue, so no one patch springs to mind.
> Thankfully it appears to be kernel related.
> -Chris
>
Bisecting took longer than I thought, but I had to cherry pick some
patches and rebase one of them multiple times...
Finally I found the commit to blame: 920cf4194954ec ("drm/i915:
Introduce an internal allocator for disposable private objects")
In case you need me to produce some more data or test a patch
feel free to reach out.
Juergen
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web