Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1471952 > unrolled thread
| Started by | Bjorn Helgaas <helgaas@kernel.org> |
|---|---|
| First post | 2016-08-29 18:10 +0200 |
| Last post | 2016-08-31 22:20 +0200 |
| Articles | 10 on this page of 30 — 5 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: Kernel Freeze with American Megatrends BIOS Bjorn Helgaas <helgaas@kernel.org> - 2016-08-29 18:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-29 21:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Bjorn Helgaas <helgaas@kernel.org> - 2016-08-29 21:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-29 22:00 +0200
Re: Kernel Freeze with American Megatrends BIOS Bjorn Helgaas <helgaas@kernel.org> - 2016-08-30 02:00 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-30 12:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Bjorn Helgaas <helgaas@kernel.org> - 2016-08-30 15:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Emil Velikov <emil.l.velikov@gmail.com> - 2016-08-30 16:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-30 17:30 +0200
Re: Kernel Freeze with American Megatrends BIOS Ilia Mirkin <imirkin@alum.mit.edu> - 2016-08-30 17:50 +0200
Re: Kernel Freeze with American Megatrends BIOS Ilia Mirkin <imirkin@alum.mit.edu> - 2016-08-30 17:50 +0200
Re: Kernel Freeze with American Megatrends BIOS Emil Velikov <emil.l.velikov@gmail.com> - 2016-08-30 17:50 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-30 19:40 +0200
Re: Kernel Freeze with American Megatrends BIOS Ilia Mirkin <imirkin@alum.mit.edu> - 2016-08-30 19:50 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-30 20:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Ilia Mirkin <imirkin@alum.mit.edu> - 2016-08-30 20:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Peter Wu <peter@lekensteyn.nl> - 2016-08-30 21:30 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 13:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 13:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Emil Velikov <emil.l.velikov@gmail.com> - 2016-08-30 20:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Emil Velikov <emil.l.velikov@gmail.com> - 2016-08-30 20:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 13:00 +0200
Re: Kernel Freeze with American Megatrends BIOS Peter Wu <peter@lekensteyn.nl> - 2016-08-30 22:00 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 13:30 +0200
Re: Kernel Freeze with American Megatrends BIOS Peter Wu <peter@lekensteyn.nl> - 2016-08-31 13:50 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 14:30 +0200
Re: Kernel Freeze with American Megatrends BIOS Peter Wu <peter@lekensteyn.nl> - 2016-08-31 14:40 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 15:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 22:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 22:20 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | Emil Velikov <emil.l.velikov@gmail.com> |
|---|---|
| Date | 2016-08-30 20:20 +0200 |
| Message-ID | <sbVNE-8n7-35@gated-at.bofh.it> |
| In reply to | #1472743 |
On 30 August 2016 at 19:09, Emil Velikov <emil.l.velikov@gmail.com> wrote:
> On 30 August 2016 at 18:37, Roland Singer <roland.singer@desertbit.com> wrote:
>> I am running 4.7.2, but I also just tried the 4.8.0-rc4 mainline kernel.
>> The result is the same. There is no difference if bbswitch of acpi_call
>> is used. However I noticed following:
>>
>> 1. The nouveau driver is broken in both kernel version and is responsible
>> for the freezes while gathering power state information with bbswitch.
>> Sometimes while shutting the system down, everything except the LCD
>> screen is switched off. This only happens with nouveau.
>> I noticed following error log messages:
>>
> I second Ilia here. Using bbswitch in conjunction with any driver (be
> that nouveau or the proprietary one) is a bad idea.
>
>> kernel: nouveau 0000:01:00.0: fb: 6144 MiB GDDR5
>> kernel: nouveau 0000:01:00.0: priv: HUB0: 10ecc0 ffffffff (1e40822c)
>> kernel: nouveau 0000:01:00.0: DRM: VRAM: 6144 MiB
>> kernel: nouveau 0000:01:00.0: DRM: GART: 1048576 MiB
>> kernel: nouveau 0000:01:00.0: DRM: Pointer to TMDS table invalid
>> kernel: nouveau 0000:01:00.0: DRM: DCB version 4.1
>> kernel: nouveau 0000:01:00.0: DRM: Pointer to flat panel table invalid
>>
>> 2. -> Boot with nouveau module loaded
>> -> switch off the discrete GPU with bbswitch or acpi_call
>> -> start X11
>> -> obtaining power state with bbswitch freezes the system
>> -> or working with the system for some minutes freezes the system
>>
> (If Ilia's suggestions does not help) Confirm if the freeze is due
> to/as the GPU is powered on or off.
>
>> 3. -> Boot with nouveau module blacklisted
>> -> switch off the discrete GPU
>> -> start X11
>> -> system immediately freezes
>>
> It's perfectly possible that the discrete GPU is set as boot one and X
> goes angry since there's no driver/way to bring it up.
>
>> 4. -> Boot with nouveau module blacklisted
>> -> switch off the discrete GPU
>> -> start Wayland
>> -> system runs - Note: I tried this for couple of days with 4.6 and 4.7 mainline
>> and the system freezed randomly after some time.
>> However I have to test if this is still present with 4.7.2
>> and 4.8 mainline. Right now it seams to be fine.
>> -> running Xwayland (does not depend on the GPU power state) kills performance!
>> the system freezes for several seconds...
>> So working with Wayland is also no solution.
>>
>> My conclusion:
>>
>> 1. Nouveau has couple of problems with GTX 9** M Nvidia GPUs.
>> I would love to help here.
>>
>> 2. X11 is just broken and is not capable to start the graphical session
>> if the nvidia GPU is not handled by any video driver (kernel module).
>> Even forcing X11 to ignore the discrete GPU doesn't help.
>>
> Out of curiosity: how did you force X to ignore the device ?
>
>> Setting the command line arguments to:
>>
>> acpi_osi=! acpi_osi="Windows 2009"
>>
>> fixes the issues with X11 but other things break...
>> What the hell is going on?! :/
>>
> You can check if it's the boot_vga assumption with
>
[Sorry about that] ...
cat /sys/class/drm/card*/device/{boot_vga,vendor}
If the output changes them my assumption holds true.
-Emil
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-31 13:00 +0200 |
| Message-ID | <scbpo-1nW-21@gated-at.bofh.it> |
| In reply to | #1472749 |
Am 30.08.2016 um 20:09 schrieb Emil Velikov:
> I second Ilia here. Using bbswitch in conjunction with any driver (be
> that nouveau or the proprietary one) is a bad idea.
>
I removed bbswitch from my system and will use vgaswitcheroo to check
the GPU power state from now.
> (If Ilia's suggestions does not help) Confirm if the freeze is due
> to/as the GPU is powered on or off.
>
Yeah, the freeze is caused by the switched off GPU.
Waited for the nouveau driver to switch it off, before starting
the graphical user interface...
> Out of curiosity: how did you force X to ignore the device ?
>
I tried to tell X11 to ignore the device with the following
configuration:
Section "Device"
Identifier "Nvidia"
VendorName "NVIDIA Corporation"
Option "Ignore" "true"
EndSection
> You can check if it's the boot_vga assumption with
> cat /sys/class/drm/card*/device/{boot_vga,vendor}
> If the output changes them my assumption holds true.
Output did not change:
1
0x8086
0x8086 is the vendor ID of intel. So that's ok...
[toc] | [prev] | [next] | [standalone]
| From | Peter Wu <peter@lekensteyn.nl> |
|---|---|
| Date | 2016-08-30 22:00 +0200 |
| Message-ID | <sbXmp-Og-1@gated-at.bofh.it> |
| In reply to | #1471952 |
On Mon, Aug 29, 2016 at 11:02:10AM -0500, Bjorn Helgaas wrote:
> [+cc linux-acpi, linux-kernel, dri-devel]
>
> Hi Roland,
>
> I have no idea how to debug this problem. Are you seeing something
> that suggests it may be a PCI problem?
Yes I suspect there is an ACPI and/ or PCI problem, possibly
device-specific. Steps to reproduce on the affected machines:
1. Load nouveau.
2. Wait for it to runtime suspend.
2. Invoke 'lspci', this resumes the Nvidia PCI device via nouveau.
3. lspci never returns, few moments later an AML_INFINITE_LOOP is
reported.
If you use the external bbswitch module, the effect is the same. I have
been trying to debug this for some time on nouveau with no luck. The
PCI/PM D3cold patches from Mika makes no difference.
Runtime resume via nouveau triggers some ACPI methods (I'll assume the
Windows 8-style PR method and take the Clevo P651 as example):
\_SB.PCI0.PEG0.PG00._ON () ->
\_SB.PCI0.PGON (0)
Then:
Method (PGON, 1, Serialized) {
PION = Arg0 // note: 0 for PG00
// ...
If ((OSYS != 0x07DF)) { /* Not Windows 2015 (Windows 10), see below */ }
Else {
LKEN (PION)
}
// this is the infinite loop: it tries to bring the PCIe link to
// full speed, but fails to do so.
While ((\_SB.PCI0.PEG0.LNKS < 0x07)) {
Local0 = 0x20
While (Local0) {
If ((\_SB.PCI0.PEG0.LNKS < 0x07)) {
Stall (0x64)
Local0--
} Else { Break }
}
If ((Local0 == Zero)) {
\_SB.PCI0.PEG0.RTLK = One
Stall (0x64)
}
}
// ...
}
Without any workaround, this piece of code is invoked:
Method (LKEN, 1, NotSerialized) {
Local3 = (CPEX & 0x0F) // CPEX at 0x5ff9be7f and has value 000506e3
If ((Local3 == Zero)) {
/* Similar to below, but with Q0L0 -> P0L0 (register 0xBC bit 6) */
} ElseIf ((Local3 != Zero)) {
If ((Arg0 == Zero)) {
/* Enter L0 Activate state.
* (LKDS tries to enter L2, deep-energy-saving state.) */
Q0L0 = One // register 0x249 bit 0; \_SB.PCI0.OPG0.Q0L0 00:01.0
Sleep (0x10)
Local0 = Zero
While (Q0L0) {
If ((Local0 > 0x04)) { Break }
Sleep (0x10)
Local0++
}
} else { /* other cases, but we are only interested in PGON(0) */ }
}
}
The acpi_osi="!Windows 2015" workaround will invoke this instead:
If ((OSYS != 0x07DF)) {
If ((PION == Zero)) {
P0AP = Zero /* PGOF writes 3 */
P0RM = Zero /* PGOF writes 1 */
}
If ((PBGE != Zero)) { /* Observed to be false (PBGE == 0) */
If (SBDL (PION)) {
PUAB (PION)
CBDL = GUBC (PION)
MBDL = GMXB (PION)
If ((CBDL > MBDL)) {
CBDL = MBDL /* \_SB_.PCI0.MBDL */
}
PDUB (PION, CBDL)
}
}
If ((PION == Zero)) {
P0LD = Zero /* Link Disable = 0, PGOF sets 1 instead. */
P0TR = One /* Train? (PGOF does not set this). */
TCNT = Zero
While ((TCNT < LDLY)) { /* LDLY = 300 */
If ((P0VC == Zero)) {
/* VC Negotiation Pending 0 means VC negotation is complete. */
Break
}
Sleep (0x10)
TCNT += 0x10 /* At most 19 iterations, sleeping for 304ms. */
}
}
}
The comments above are my own interpretation based on the acpidumps I
extracted from the machine. These notes and ACPI tables can be found at
https://github.com/Lekensteyn/acpi-stuff/blob/master/Clevo-P651RA/notes.txt
https://github.com/Lekensteyn/acpi-stuff/tree/master/dsl/Clevo_P651RA
Other affected devices have similar code, differences are small:
- No check for LNKS (avoids the infinite loop, but device is still off)
- Instead of a check for != "Windows 2015", they check for == "Windows
2009" or even for == "Windows 2009" || "Windows 2013" (Dell Inspiron
7559).
The tested kernels (with bbswitch or nouveau) were Linux 4.4.0, 4.6,
4.7 (nouveau + PCI/PM + nouveau PR patches). The PCIe device is
something from the GTX 9xxM family in all cases.
I have a bunch of PCI config dumps from Windows and Linux, but there is
nothing extraordinary. Also did an ACPI trace via a Checked/Debug build
of Windows, but it just confirms that the ACPI method we use for the
Nvidia device is the correct one.
Let me know if you need more information, I would be glad to provide.
Kind regards,
Peter
> On Tue, Aug 23, 2016 at 11:23:45AM +0200, Roland Singer wrote:
> > Hi,
> >
> > hope somebody can help me fix this kernel problem which affects the following machines:
> >
> > - Clevo P651RA (i7-6700HQ/GTX 965M, part of the P6xxRx family which are also affected)
> > - MSI GE62 Apache Pro (i7-6700HQ/GTX 960M)
> > - Gigabyte P35V5 (i7-6700HQ/GTX 970M)
> > - Razer Blade 14" (2016) (i7-6700HQ/GTX 970M) (BIOS 5.11, 04/07/2016)
> >
> >
> > The kernel freezes if the graphical user session (Xorg & Wayland) is
> > started with a switched off discrete GPU card (NVIDIA).
> > If the discrete GPU is switched off after the graphical session start,
> > then everything works as expected, until the graphical session is restarted.
> >
> > This problem seams to be linked to specific BIOS settings. If the computer
> > is started with the following command line:
> >
> > acpi_osi=! acpi_osi="Windows 2009"
> >
> > then the kernel freeze does not occur anymore. However this required a special
> > ACPI DSDT firmware patch for the Razer Blade 2016 laptop:
> >
> > https://github.com/m4ng0squ4sh/razer_blade_14_2016_acpi_dsdt
> >
> > I strongly recommend to fix this in the kernel and I am ready to help and solve
> > this problem with some help.
> >
> > Here is a link to the GitHub issue with further information:
> >
> > https://github.com/Bumblebee-Project/Bumblebee/issues/764#issuecomment-241212595
> >
> > Here are some more detailed information:
> >
> > https://github.com/Lekensteyn/acpi-stuff/blob/master/Clevo-P651RA/notes.txt
> >
> > Hope somebody can help.
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-31 13:30 +0200 |
| Message-ID | <scbSr-1Nd-37@gated-at.bofh.it> |
| In reply to | #1472820 |
Am 30.08.2016 um 21:53 schrieb Peter Wu:
> On Mon, Aug 29, 2016 at 11:02:10AM -0500, Bjorn Helgaas wrote:
>> [+cc linux-acpi, linux-kernel, dri-devel]
>>
>> Hi Roland,
>>
>> I have no idea how to debug this problem. Are you seeing something
>> that suggests it may be a PCI problem?
>
> Yes I suspect there is an ACPI and/ or PCI problem, possibly
> device-specific. Steps to reproduce on the affected machines:
>
> 1. Load nouveau.
> 2. Wait for it to runtime suspend.
> 2. Invoke 'lspci', this resumes the Nvidia PCI device via nouveau.
> 3. lspci never returns, few moments later an AML_INFINITE_LOOP is
> reported.
>
I can confirm this. Same result on my machine.
Here is a link to my ACPI tables:
https://bugs.launchpad.net/lpbugreporter/+bug/752542/+attachment/4722651/+files/Razer-Blade.tar.gz
The specific source for the NVIDIA card can be found in the ssdt5.dsl file.
Method (PGON, 1, Serialized)
{
/* ... */
GPPR (PION, One)
If ((OSYS == 0x07D9)) /* Is Windows 2009 - In my case, setting to Windows 2009 only works! */
{
If ((PION == Zero))
{
P0AP = Zero
P0RM = Zero
}
ElseIf ((PION == One))
{
P1AP = Zero
P1RM = Zero
}
ElseIf ((PION == 0x02))
{
P2AP = Zero
P2RM = Zero
}
If ((PBGE != Zero))
{
If (SBDL (PION))
{
PUAB (PION)
CBDL = GUBC (PION)
MBDL = GMXB (PION)
If ((CBDL > MBDL))
{
CBDL = MBDL /* \_SB_.PCI0.MBDL */
}
PDUB (PION, CBDL)
}
}
If ((PION == Zero))
{
P0LD = Zero
P0TR = One
TCNT = Zero
While ((TCNT < LDLY))
{
If ((P0VC == Zero))
{
Break
}
Sleep (0x10)
TCNT += 0x10
}
}
ElseIf ((PION == One))
{
P1LD = Zero
P1TR = One
TCNT = Zero
While ((TCNT < LDLY))
{
If ((P1VC == Zero))
{
Break
}
Sleep (0x10)
TCNT += 0x10
}
}
ElseIf ((PION == 0x02))
{
P2LD = Zero
P2TR = One
TCNT = Zero
While ((TCNT < LDLY))
{
If ((P2VC == Zero))
{
Break
}
Sleep (0x10)
TCNT += 0x10
}
}
}
Else
{
LKEN (PION)
}
/* ... */
Return (Zero)
}
If not set to Windows 2009, then this is triggered:
Method (LKEN, 1, NotSerialized)
{
Local3 = (CPEX & 0x0F)
If ((Local3 == Zero))
{
If ((Arg0 == Zero))
{
P0L0 = One
Sleep (0x10)
Local0 = Zero
While (P0L0)
{
If ((Local0 > 0x04))
{
Break
}
Sleep (0x10)
Local0++
}
}
ElseIf ((Arg0 == One))
{
P1L0 = One
Sleep (0x10)
Local0 = Zero
While (P0L0)
{
If ((Local0 > 0x04))
{
Break
}
Sleep (0x10)
Local0++
}
}
ElseIf ((Arg0 == 0x02))
{
P2L0 = One
Sleep (0x10)
Local0 = Zero
While (P0L0)
{
If ((Local0 > 0x04))
{
Break
}
Sleep (0x10)
Local0++
}
}
}
ElseIf ((Local3 != Zero))
{
If ((Arg0 == Zero))
{
Q0L0 = One
Sleep (0x10)
Local0 = Zero
While (Q0L0)
{
If ((Local0 > 0x04))
{
Break
}
Sleep (0x10)
Local0++
}
}
ElseIf ((Arg0 == One))
{
Q1L0 = One
Sleep (0x10)
Local0 = Zero
While (Q1L0)
{
If ((Local0 > 0x04))
{
Break
}
Sleep (0x10)
Local0++
}
}
ElseIf ((Arg0 == 0x02))
{
Q2L0 = One
Sleep (0x10)
Local0 = Zero
While (Q2L0)
{
If ((Local0 > 0x04))
{
Break
}
Sleep (0x10)
Local0++
}
}
}
}
Is it possible to override the specific ACPI table functions (SSDT) in the DSDT?
This way I could try to debug to find some more information...
[toc] | [prev] | [next] | [standalone]
| From | Peter Wu <peter@lekensteyn.nl> |
|---|---|
| Date | 2016-08-31 13:50 +0200 |
| Message-ID | <sccbL-1TJ-9@gated-at.bofh.it> |
| In reply to | #1473340 |
On Wed, Aug 31, 2016 at 01:27:36PM +0200, Roland Singer wrote:
> Am 30.08.2016 um 21:53 schrieb Peter Wu:
> > On Mon, Aug 29, 2016 at 11:02:10AM -0500, Bjorn Helgaas wrote:
> >> [+cc linux-acpi, linux-kernel, dri-devel]
> >>
> >> Hi Roland,
> >>
> >> I have no idea how to debug this problem. Are you seeing something
> >> that suggests it may be a PCI problem?
> >
> > Yes I suspect there is an ACPI and/ or PCI problem, possibly
> > device-specific. Steps to reproduce on the affected machines:
> >
> > 1. Load nouveau.
> > 2. Wait for it to runtime suspend.
> > 2. Invoke 'lspci', this resumes the Nvidia PCI device via nouveau.
> > 3. lspci never returns, few moments later an AML_INFINITE_LOOP is
> > reported.
> >
>
> I can confirm this. Same result on my machine.
>
> Here is a link to my ACPI tables:
> https://bugs.launchpad.net/lpbugreporter/+bug/752542/+attachment/4722651/+files/Razer-Blade.tar.gz
>
> The specific source for the NVIDIA card can be found in the ssdt5.dsl file.
>
>
> Method (PGON, 1, Serialized)
> {
> /* ... */
>
> GPPR (PION, One)
> If ((OSYS == 0x07D9)) /* Is Windows 2009 - In my case, setting to Windows 2009 only works! */
> {
[..]
> }
> Else
> {
> LKEN (PION)
> }
>
> /* ... */
>
> Return (Zero)
> }
>
>
>
> If not set to Windows 2009, then this is triggered:
>
>
> Method (LKEN, 1, NotSerialized)
> {
[..]
> }
Yep, this is the same code. I stripped out irrelevant parts from the
previous mail for brevity.
> Is it possible to override the specific ACPI table functions (SSDT) in the DSDT?
> This way I could try to debug to find some more information...
See Documentation/acpi/initrd_table_override.txt and note that it is
important that the tables are really located at /kernel/firmware/acpi/
in your initrd (which must be the first, even before any possible
microcode updates).
What are you trying to do? For ACPI method tracing, see
Documentation/acpi/method-tracing.txt
--
Kind regards,
Peter Wu
https://lekensteyn.nl
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-31 14:30 +0200 |
| Message-ID | <sccOt-2m4-11@gated-at.bofh.it> |
| In reply to | #1473361 |
Am 31.08.2016 um 13:46 schrieb Peter Wu:
> On Wed, Aug 31, 2016 at 01:27:36PM +0200, Roland Singer wrote:
>> Am 30.08.2016 um 21:53 schrieb Peter Wu:
>>> On Mon, Aug 29, 2016 at 11:02:10AM -0500, Bjorn Helgaas wrote:
>>>> [+cc linux-acpi, linux-kernel, dri-devel]
>>>>
>>>> Hi Roland,
>>>>
>>>> I have no idea how to debug this problem. Are you seeing something
>>>> that suggests it may be a PCI problem?
>>>
>>> Yes I suspect there is an ACPI and/ or PCI problem, possibly
>>> device-specific. Steps to reproduce on the affected machines:
>>>
>>> 1. Load nouveau.
>>> 2. Wait for it to runtime suspend.
>>> 2. Invoke 'lspci', this resumes the Nvidia PCI device via nouveau.
>>> 3. lspci never returns, few moments later an AML_INFINITE_LOOP is
>>> reported.
>>>
>>
>> I can confirm this. Same result on my machine.
>>
>> Here is a link to my ACPI tables:
>> https://bugs.launchpad.net/lpbugreporter/+bug/752542/+attachment/4722651/+files/Razer-Blade.tar.gz
>>
>> The specific source for the NVIDIA card can be found in the ssdt5.dsl file.
>>
>>
>> Method (PGON, 1, Serialized)
>> {
>> /* ... */
>>
>> GPPR (PION, One)
>> If ((OSYS == 0x07D9)) /* Is Windows 2009 - In my case, setting to Windows 2009 only works! */
>> {
> [..]
>> }
>> Else
>> {
>> LKEN (PION)
>> }
>>
>> /* ... */
>>
>> Return (Zero)
>> }
>>
>>
>>
>> If not set to Windows 2009, then this is triggered:
>>
>>
>> Method (LKEN, 1, NotSerialized)
>> {
> [..]
>> }
>
> Yep, this is the same code. I stripped out irrelevant parts from the
> previous mail for brevity.
>
>> Is it possible to override the specific ACPI table functions (SSDT) in the DSDT?
>> This way I could try to debug to find some more information...
>
> See Documentation/acpi/initrd_table_override.txt and note that it is
> important that the tables are really located at /kernel/firmware/acpi/
> in your initrd (which must be the first, even before any possible
> microcode updates).
>
> What are you trying to do? For ACPI method tracing, see
> Documentation/acpi/method-tracing.txt
>
Oh, you're right.
Thanks. Right now I am overriding the DSDT, but I am not able to override
the SSDT, because I have to fix and compile all the SSDT files. There
are too many compile errors... Wanted to find the exact line which
is responsible for the hickup.
>>> Yes I suspect there is an ACPI and/ or PCI problem, possibly
>>> device-specific. Steps to reproduce on the affected machines:
>>>
>>> 1. Load nouveau.
>>> 2. Wait for it to runtime suspend.
>>> 2. Invoke 'lspci', this resumes the Nvidia PCI device via nouveau.
>>> 3. lspci never returns, few moments later an AML_INFINITE_LOOP is
>>> reported.
I noticed following:
1. Blacklist nouveau
2. Boot to GDM login manager (Wayland)
3. Switch to TTY with CTRL+ALT+FN2
4. Load bbswitch
5. Switch off GPU
6. run lspci -> no freeze
7. Switch to GDM
8. Login to a Wayland session (X11 won't work)
9. run lspci in a GUI terminal -> system freezes
[toc] | [prev] | [next] | [standalone]
| From | Peter Wu <peter@lekensteyn.nl> |
|---|---|
| Date | 2016-08-31 14:40 +0200 |
| Message-ID | <sccYa-2pg-15@gated-at.bofh.it> |
| In reply to | #1473414 |
On Wed, Aug 31, 2016 at 02:21:31PM +0200, Roland Singer wrote:
> Am 31.08.2016 um 13:46 schrieb Peter Wu:
> > On Wed, Aug 31, 2016 at 01:27:36PM +0200, Roland Singer wrote:
> >> Am 30.08.2016 um 21:53 schrieb Peter Wu:
> >>> On Mon, Aug 29, 2016 at 11:02:10AM -0500, Bjorn Helgaas wrote:
> >>>> [+cc linux-acpi, linux-kernel, dri-devel]
> >>>>
> >>>> Hi Roland,
> >>>>
> >>>> I have no idea how to debug this problem. Are you seeing something
> >>>> that suggests it may be a PCI problem?
> >>>
> >>> Yes I suspect there is an ACPI and/ or PCI problem, possibly
> >>> device-specific. Steps to reproduce on the affected machines:
> >>>
> >>> 1. Load nouveau.
> >>> 2. Wait for it to runtime suspend.
> >>> 2. Invoke 'lspci', this resumes the Nvidia PCI device via nouveau.
> >>> 3. lspci never returns, few moments later an AML_INFINITE_LOOP is
> >>> reported.
> >>>
> >>
> >> I can confirm this. Same result on my machine.
> >>
> >> Here is a link to my ACPI tables:
> >> https://bugs.launchpad.net/lpbugreporter/+bug/752542/+attachment/4722651/+files/Razer-Blade.tar.gz
> >>
> >> The specific source for the NVIDIA card can be found in the ssdt5.dsl file.
> >>
> >>
> >> Method (PGON, 1, Serialized)
> >> {
> >> /* ... */
> >>
> >> GPPR (PION, One)
> >> If ((OSYS == 0x07D9)) /* Is Windows 2009 - In my case, setting to Windows 2009 only works! */
> >> {
> > [..]
> >> }
> >> Else
> >> {
> >> LKEN (PION)
> >> }
> >>
> >> /* ... */
> >>
> >> Return (Zero)
> >> }
> >>
> >>
> >>
> >> If not set to Windows 2009, then this is triggered:
> >>
> >>
> >> Method (LKEN, 1, NotSerialized)
> >> {
> > [..]
> >> }
> >
> > Yep, this is the same code. I stripped out irrelevant parts from the
> > previous mail for brevity.
> >
> >> Is it possible to override the specific ACPI table functions (SSDT) in the DSDT?
> >> This way I could try to debug to find some more information...
> >
> > See Documentation/acpi/initrd_table_override.txt and note that it is
> > important that the tables are really located at /kernel/firmware/acpi/
> > in your initrd (which must be the first, even before any possible
> > microcode updates).
> >
> > What are you trying to do? For ACPI method tracing, see
> > Documentation/acpi/method-tracing.txt
> >
>
> Oh, you're right.
>
> Thanks. Right now I am overriding the DSDT, but I am not able to override
> the SSDT, because I have to fix and compile all the SSDT files. There
> are too many compile errors... Wanted to find the exact line which
> is responsible for the hickup.
Have you disassembled with externs included? That is,
iasl -e *.dat -d ssdtX.dat
If you are sure that the remaining errors are harmless, you can use the
'-f' option to ignore errors. You can also use the `-ve` option to
suppress warnings and remarks so you can focus on the errors.
If you look at my notes.txt, you will see that _OFF always executes the
same code. PGON differs. When the problem occurs, "Q0L0" somehow always
reads back as non-zero and LNKS < 7.
> >>> Yes I suspect there is an ACPI and/ or PCI problem, possibly
> >>> device-specific. Steps to reproduce on the affected machines:
> >>>
> >>> 1. Load nouveau.
> >>> 2. Wait for it to runtime suspend.
> >>> 2. Invoke 'lspci', this resumes the Nvidia PCI device via nouveau.
> >>> 3. lspci never returns, few moments later an AML_INFINITE_LOOP is
> >>> reported.
>
> I noticed following:
>
> 1. Blacklist nouveau
> 2. Boot to GDM login manager (Wayland)
> 3. Switch to TTY with CTRL+ALT+FN2
> 4. Load bbswitch
> 5. Switch off GPU
> 6. run lspci -> no freeze
> 7. Switch to GDM
> 8. Login to a Wayland session (X11 won't work)
> 9. run lspci in a GUI terminal -> system freezes
Is nouveau somehow loaded anyway? All those extra components (X11,
Wayland, etc.) are unnecessary to reproduce the core problem. It occurs
whenever the device is being resumed (either via DSM/_PS0 or via power
resource PG00._ON).
--
Kind regards,
Peter Wu
https://lekensteyn.nl
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-31 15:20 +0200 |
| Message-ID | <scdAR-2Rr-3@gated-at.bofh.it> |
| In reply to | #1473419 |
>> >> Thanks. Right now I am overriding the DSDT, but I am not able to override >> the SSDT, because I have to fix and compile all the SSDT files. There >> are too many compile errors... Wanted to find the exact line which >> is responsible for the hickup. > > Have you disassembled with externs included? That is, > > iasl -e *.dat -d ssdtX.dat > > If you are sure that the remaining errors are harmless, you can use the > '-f' option to ignore errors. You can also use the `-ve` option to > suppress warnings and remarks so you can focus on the errors. > Thanks, I'll try that. > If you look at my notes.txt, you will see that _OFF always executes the > same code. PGON differs. When the problem occurs, "Q0L0" somehow always > reads back as non-zero and LNKS < 7. > Oh you're Lekensteyn ^^ I don't have LNKS and no while loop after calling LKEN ?! >> >> I noticed following: >> >> 1. Blacklist nouveau >> 2. Boot to GDM login manager (Wayland) >> 3. Switch to TTY with CTRL+ALT+FN2 >> 4. Load bbswitch >> 5. Switch off GPU >> 6. run lspci -> no freeze >> 7. Switch to GDM >> 8. Login to a Wayland session (X11 won't work) >> 9. run lspci in a GUI terminal -> system freezes > > Is nouveau somehow loaded anyway? All those extra components (X11, > Wayland, etc.) are unnecessary to reproduce the core problem. It occurs > whenever the device is being resumed (either via DSM/_PS0 or via power > resource PG00._ON). > Sorry that was nonsense. The steps to reproduce the problem are still valid. I didn't wait enough to power it down... But whats interesting: 1. Blacklist nouveau 2. Load bbswitch 3. Power off GPU with bbswitch 4. Power on GPU with bbswitch 5. Run lspci 6. Power off GPU with bbswitch 7. Run lspci -> freeze So setting the GPU power state with bbswitch works as expected. Powering it on is also fine. I did this a couple of times. But powering it off and letting lspci powering it on, ends in a race. It might be, that lspci does not only power the GPU on, but triggers another pci action which causes the race condition. Does this have something to do with your quote about the retrain bit?
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-31 22:10 +0200 |
| Message-ID | <scjZD-6XP-11@gated-at.bofh.it> |
| In reply to | #1473449 |
Here is Peter Wu's reply, which was not send to the mailing list, because
I had to resend my e-mail to him due to a failure...
-------- Forwarded Message --------
Subject: Re: Fwd: Re: Kernel Freeze with American Megatrends BIOS
Date: Wed, 31 Aug 2016 18:08:53 +0200
From: Peter Wu <peter@lekensteyn.nl>
To: Roland Singer <roland.singer@desertbit.com>
On Wed, Aug 31, 2016 at 05:56:18PM +0200, Roland Singer wrote:
> > If you look at my notes.txt, you will see that _OFF always executes the
> > same code. PGON differs. When the problem occurs, "Q0L0" somehow always
> > reads back as non-zero and LNKS < 7.
> >
>
> Oh you're Lekensteyn ^^
Yes, that's me :) I wrote bbswitch, did the Optimus and PR3 ACPI support
in nouveau so I am fairly certain what happens behind the scenes.
> I don't have LNKS and no while loop after calling LKEN ?!
Yes that is what I said in
https://www.spinics.net/lists/linux-pci/msg53694.html:
"Other affected devices have similar code, differences are small:
No check for LNKS (avoids the infinite loop, but device is still off)"
> >>
> >> I noticed following:
> >>
> >> 1. Blacklist nouveau
> >> 2. Boot to GDM login manager (Wayland)
> >> 3. Switch to TTY with CTRL+ALT+FN2
> >> 4. Load bbswitch
> >> 5. Switch off GPU
> >> 6. run lspci -> no freeze
> >> 7. Switch to GDM
> >> 8. Login to a Wayland session (X11 won't work)
> >> 9. run lspci in a GUI terminal -> system freezes
> >
> > Is nouveau somehow loaded anyway? All those extra components (X11,
> > Wayland, etc.) are unnecessary to reproduce the core problem. It occurs
> > whenever the device is being resumed (either via DSM/_PS0 or via power
> > resource PG00._ON).
> >
>
> Sorry that was nonsense. The steps to reproduce the problem are still valid.
> I didn't wait enough to power it down...
>
> But whats interesting:
>
> 1. Blacklist nouveau
> 2. Load bbswitch
> 3. Power off GPU with bbswitch
> 4. Power on GPU with bbswitch
> 5. Run lspci
> 6. Power off GPU with bbswitch
> 7. Run lspci -> freeze
>
> So setting the GPU power state with bbswitch works as expected.
> Powering it on is also fine. I did this a couple of times.
> But powering it off and letting lspci powering it on, ends in a race.
In some cases I also found that it does always happen at the first try,
but with nouveau it always seem to happen.
> It might be, that lspci does not only power the GPU on, but triggers
> another pci action which causes the race condition.
> Does this have something to do with your quote about the retrain bit?
That is an interesting hypothesis. Even if you invoke `lspci -s01:00.0`
for example, it will always probe for all devices. So maybe interaction
with its parent device (PCI root port 00:02.0) causes issues.
However I also tested without lspci before, and the problem still
exists. You can trigger runtime resume via (as root):
echo > /sys/bus/pci/0000:01:00.0/power/control on
Set it to "auto" to make it sleep again.
--
Kind regards,
Peter Wu
https://lekensteyn.nl
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-31 22:20 +0200 |
| Message-ID | <sck9j-718-9@gated-at.bofh.it> |
| In reply to | #1473821 |
On 08/31/16 22:06, Roland Singer wrote: > Here is Peter Wu's reply, which was not send to the mailing list, because > I had to resend my e-mail to him due to a failure... > > > -------- Forwarded Message -------- > Subject: Re: Fwd: Re: Kernel Freeze with American Megatrends BIOS > Date: Wed, 31 Aug 2016 18:08:53 +0200 > From: Peter Wu <peter@lekensteyn.nl> > To: Roland Singer <roland.singer@desertbit.com> > > On Wed, Aug 31, 2016 at 05:56:18PM +0200, Roland Singer wrote: > >>> If you look at my notes.txt, you will see that _OFF always executes the >>> same code. PGON differs. When the problem occurs, "Q0L0" somehow always >>> reads back as non-zero and LNKS < 7. >>> >> >> Oh you're Lekensteyn ^^ > > Yes, that's me :) I wrote bbswitch, did the Optimus and PR3 ACPI support > in nouveau so I am fairly certain what happens behind the scenes. > Awesome! Thanks for all your efforts! Great work :) >> I don't have LNKS and no while loop after calling LKEN ?! > > Yes that is what I said in > https://www.spinics.net/lists/linux-pci/msg53694.html: > > "Other affected devices have similar code, differences are small: > No check for LNKS (avoids the infinite loop, but device is still off)" > Ah ok, missed that. >> It might be, that lspci does not only power the GPU on, but triggers >> another pci action which causes the race condition. >> Does this have something to do with your quote about the retrain bit? > > That is an interesting hypothesis. Even if you invoke `lspci -s01:00.0` > for example, it will always probe for all devices. So maybe interaction > with its parent device (PCI root port 00:02.0) causes issues. > > However I also tested without lspci before, and the problem still > exists. You can trigger runtime resume via (as root): > > echo > /sys/bus/pci/0000:01:00.0/power/control on > > Set it to "auto" to make it sleep again. > Just tried it over and over again. I don't have any problems switching the GPU power state with bbswitch. So, switching the GPU on is just fine. There must be something else, which does not cooperate well while switching it on (lspci)... I can confirm,, that `lspci -s01:00.0` also freezes the system. Trying to trigger runtime resume with `/sys/bus/pci/0000:01:00.0/power/control` did not work for me. The GPU just stayed off. Any hints how to get some more information?
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web