Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1471952 > unrolled thread
| Started by | Bjorn Helgaas <helgaas@kernel.org> |
|---|---|
| First post | 2016-08-29 18:10 +0200 |
| Last post | 2016-08-31 22:20 +0200 |
| Articles | 20 on this page of 30 — 5 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: Kernel Freeze with American Megatrends BIOS Bjorn Helgaas <helgaas@kernel.org> - 2016-08-29 18:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-29 21:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Bjorn Helgaas <helgaas@kernel.org> - 2016-08-29 21:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-29 22:00 +0200
Re: Kernel Freeze with American Megatrends BIOS Bjorn Helgaas <helgaas@kernel.org> - 2016-08-30 02:00 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-30 12:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Bjorn Helgaas <helgaas@kernel.org> - 2016-08-30 15:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Emil Velikov <emil.l.velikov@gmail.com> - 2016-08-30 16:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-30 17:30 +0200
Re: Kernel Freeze with American Megatrends BIOS Ilia Mirkin <imirkin@alum.mit.edu> - 2016-08-30 17:50 +0200
Re: Kernel Freeze with American Megatrends BIOS Ilia Mirkin <imirkin@alum.mit.edu> - 2016-08-30 17:50 +0200
Re: Kernel Freeze with American Megatrends BIOS Emil Velikov <emil.l.velikov@gmail.com> - 2016-08-30 17:50 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-30 19:40 +0200
Re: Kernel Freeze with American Megatrends BIOS Ilia Mirkin <imirkin@alum.mit.edu> - 2016-08-30 19:50 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-30 20:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Ilia Mirkin <imirkin@alum.mit.edu> - 2016-08-30 20:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Peter Wu <peter@lekensteyn.nl> - 2016-08-30 21:30 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 13:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 13:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Emil Velikov <emil.l.velikov@gmail.com> - 2016-08-30 20:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Emil Velikov <emil.l.velikov@gmail.com> - 2016-08-30 20:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 13:00 +0200
Re: Kernel Freeze with American Megatrends BIOS Peter Wu <peter@lekensteyn.nl> - 2016-08-30 22:00 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 13:30 +0200
Re: Kernel Freeze with American Megatrends BIOS Peter Wu <peter@lekensteyn.nl> - 2016-08-31 13:50 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 14:30 +0200
Re: Kernel Freeze with American Megatrends BIOS Peter Wu <peter@lekensteyn.nl> - 2016-08-31 14:40 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 15:20 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 22:10 +0200
Re: Kernel Freeze with American Megatrends BIOS Roland Singer <roland.singer@desertbit.com> - 2016-08-31 22:20 +0200
Page 1 of 2 [1] 2 Next page →
| From | Bjorn Helgaas <helgaas@kernel.org> |
|---|---|
| Date | 2016-08-29 18:10 +0200 |
| Subject | Re: Kernel Freeze with American Megatrends BIOS |
| Message-ID | <sbxii-10O-17@gated-at.bofh.it> |
[+cc linux-acpi, linux-kernel, dri-devel] Hi Roland, I have no idea how to debug this problem. Are you seeing something that suggests it may be a PCI problem? On Tue, Aug 23, 2016 at 11:23:45AM +0200, Roland Singer wrote: > Hi, > > hope somebody can help me fix this kernel problem which affects the following machines: > > - Clevo P651RA (i7-6700HQ/GTX 965M, part of the P6xxRx family which are also affected) > - MSI GE62 Apache Pro (i7-6700HQ/GTX 960M) > - Gigabyte P35V5 (i7-6700HQ/GTX 970M) > - Razer Blade 14" (2016) (i7-6700HQ/GTX 970M) (BIOS 5.11, 04/07/2016) > > > The kernel freezes if the graphical user session (Xorg & Wayland) is > started with a switched off discrete GPU card (NVIDIA). > If the discrete GPU is switched off after the graphical session start, > then everything works as expected, until the graphical session is restarted. > > This problem seams to be linked to specific BIOS settings. If the computer > is started with the following command line: > > acpi_osi=! acpi_osi="Windows 2009" > > then the kernel freeze does not occur anymore. However this required a special > ACPI DSDT firmware patch for the Razer Blade 2016 laptop: > > https://github.com/m4ng0squ4sh/razer_blade_14_2016_acpi_dsdt > > I strongly recommend to fix this in the kernel and I am ready to help and solve > this problem with some help. > > Here is a link to the GitHub issue with further information: > > https://github.com/Bumblebee-Project/Bumblebee/issues/764#issuecomment-241212595 > > Here are some more detailed information: > > https://github.com/Lekensteyn/acpi-stuff/blob/master/Clevo-P651RA/notes.txt > > Hope somebody can help.
[toc] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-29 21:10 +0200 |
| Message-ID | <sbA6u-2Nk-35@gated-at.bofh.it> |
| In reply to | #1471952 |
Hi Bjorn,
I am using the bbswitch kernel module to switch off/on the GPU and
to obtain the GPU power state.
Obtaining the GPU state immediately after starting the graphical user
session freezes the system.
This code triggers something, which is responsible for the freeze.
---
// Returns 1 if the card is disabled, 0 if enabled
static int is_card_disabled(void) {
u32 cfg_word;
// read first config word which contains Vendor and Device ID. If all bits
// are enabled, the device is assumed to be off
pci_read_config_dword(dis_dev, 0, &cfg_word);
// if one of the bits is not enabled (the card is enabled), the inverted
// result will be non-zero and hence logical not will make it 0 ("false")
return !~cfg_word;
}
static int bbswitch_proc_show(struct seq_file *seqfp, void *p) {
// show the card state. Example output: 0000:01:00:00 ON
dis_dev_get();
seq_printf(seqfp, "%s %s\n", dev_name(&dis_dev->dev),
is_card_disabled() ? "OFF" : "ON");
dis_dev_put();
return 0;
}
---
Either dis_dev_get or pci_read_config_dword is the trigger.
Link to the bbswitch module source code:
https://github.com/Bumblebee-Project/bbswitch/blob/master/bbswitch.c#L333
Am 29.08.2016 um 18:02 schrieb Bjorn Helgaas:
> [+cc linux-acpi, linux-kernel, dri-devel]
>
> Hi Roland,
>
> I have no idea how to debug this problem. Are you seeing something
> that suggests it may be a PCI problem?
>
> On Tue, Aug 23, 2016 at 11:23:45AM +0200, Roland Singer wrote:
>> Hi,
>>
>> hope somebody can help me fix this kernel problem which affects the following machines:
>>
>> - Clevo P651RA (i7-6700HQ/GTX 965M, part of the P6xxRx family which are also affected)
>> - MSI GE62 Apache Pro (i7-6700HQ/GTX 960M)
>> - Gigabyte P35V5 (i7-6700HQ/GTX 970M)
>> - Razer Blade 14" (2016) (i7-6700HQ/GTX 970M) (BIOS 5.11, 04/07/2016)
>>
>>
>> The kernel freezes if the graphical user session (Xorg & Wayland) is
>> started with a switched off discrete GPU card (NVIDIA).
>> If the discrete GPU is switched off after the graphical session start,
>> then everything works as expected, until the graphical session is restarted.
>>
>> This problem seams to be linked to specific BIOS settings. If the computer
>> is started with the following command line:
>>
>> acpi_osi=! acpi_osi="Windows 2009"
>>
>> then the kernel freeze does not occur anymore. However this required a special
>> ACPI DSDT firmware patch for the Razer Blade 2016 laptop:
>>
>> https://github.com/m4ng0squ4sh/razer_blade_14_2016_acpi_dsdt
>>
>> I strongly recommend to fix this in the kernel and I am ready to help and solve
>> this problem with some help.
>>
>> Here is a link to the GitHub issue with further information:
>>
>> https://github.com/Bumblebee-Project/Bumblebee/issues/764#issuecomment-241212595
>>
>> Here are some more detailed information:
>>
>> https://github.com/Lekensteyn/acpi-stuff/blob/master/Clevo-P651RA/notes.txt
>>
>> Hope somebody can help.
[toc] | [prev] | [next] | [standalone]
| From | Bjorn Helgaas <helgaas@kernel.org> |
|---|---|
| Date | 2016-08-29 21:10 +0200 |
| Message-ID | <sbA6u-2Nk-53@gated-at.bofh.it> |
| In reply to | #1472060 |
On Mon, Aug 29, 2016 at 08:46:17PM +0200, Roland Singer wrote:
> Hi Bjorn,
>
> I am using the bbswitch kernel module to switch off/on the GPU and
> to obtain the GPU power state.
> Obtaining the GPU state immediately after starting the graphical user
> session freezes the system.
>
> This code triggers something, which is responsible for the freeze.
>
> ---
> // Returns 1 if the card is disabled, 0 if enabled
> static int is_card_disabled(void) {
> u32 cfg_word;
> // read first config word which contains Vendor and Device ID. If all bits
> // are enabled, the device is assumed to be off
> pci_read_config_dword(dis_dev, 0, &cfg_word);
> // if one of the bits is not enabled (the card is enabled), the inverted
> // result will be non-zero and hence logical not will make it 0 ("false")
> return !~cfg_word;
> }
>
> static int bbswitch_proc_show(struct seq_file *seqfp, void *p) {
> // show the card state. Example output: 0000:01:00:00 ON
> dis_dev_get();
> seq_printf(seqfp, "%s %s\n", dev_name(&dis_dev->dev),
> is_card_disabled() ? "OFF" : "ON");
> dis_dev_put();
> return 0;
> }
> ---
>
> Either dis_dev_get or pci_read_config_dword is the trigger.
What happens if you remove the call to is_card_disabled()? Does the
system still freeze if you only do the dis_dev_get()/dis_dev_put()?
> Link to the bbswitch module source code:
> https://github.com/Bumblebee-Project/bbswitch/blob/master/bbswitch.c#L333
>
>
> Am 29.08.2016 um 18:02 schrieb Bjorn Helgaas:
> > [+cc linux-acpi, linux-kernel, dri-devel]
> >
> > Hi Roland,
> >
> > I have no idea how to debug this problem. Are you seeing something
> > that suggests it may be a PCI problem?
> >
> > On Tue, Aug 23, 2016 at 11:23:45AM +0200, Roland Singer wrote:
> >> Hi,
> >>
> >> hope somebody can help me fix this kernel problem which affects the following machines:
> >>
> >> - Clevo P651RA (i7-6700HQ/GTX 965M, part of the P6xxRx family which are also affected)
> >> - MSI GE62 Apache Pro (i7-6700HQ/GTX 960M)
> >> - Gigabyte P35V5 (i7-6700HQ/GTX 970M)
> >> - Razer Blade 14" (2016) (i7-6700HQ/GTX 970M) (BIOS 5.11, 04/07/2016)
> >>
> >>
> >> The kernel freezes if the graphical user session (Xorg & Wayland) is
> >> started with a switched off discrete GPU card (NVIDIA).
> >> If the discrete GPU is switched off after the graphical session start,
> >> then everything works as expected, until the graphical session is restarted.
> >>
> >> This problem seams to be linked to specific BIOS settings. If the computer
> >> is started with the following command line:
> >>
> >> acpi_osi=! acpi_osi="Windows 2009"
> >>
> >> then the kernel freeze does not occur anymore. However this required a special
> >> ACPI DSDT firmware patch for the Razer Blade 2016 laptop:
> >>
> >> https://github.com/m4ng0squ4sh/razer_blade_14_2016_acpi_dsdt
> >>
> >> I strongly recommend to fix this in the kernel and I am ready to help and solve
> >> this problem with some help.
> >>
> >> Here is a link to the GitHub issue with further information:
> >>
> >> https://github.com/Bumblebee-Project/Bumblebee/issues/764#issuecomment-241212595
> >>
> >> Here are some more detailed information:
> >>
> >> https://github.com/Lekensteyn/acpi-stuff/blob/master/Clevo-P651RA/notes.txt
> >>
> >> Hope somebody can help.
>
> --
> To unsubscribe from this list: send the line "unsubscribe linux-pci" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-29 22:00 +0200 |
| Message-ID | <sbASR-37n-1@gated-at.bofh.it> |
| In reply to | #1472063 |
Just tried it and the system didn't freeze. However it will freeze
after some time (few minutes while working).
Seams to be pci_read_config_dword. Where is this exactly defined?
Am 29.08.2016 um 21:07 schrieb Bjorn Helgaas:
> On Mon, Aug 29, 2016 at 08:46:17PM +0200, Roland Singer wrote:
>> Hi Bjorn,
>>
>> I am using the bbswitch kernel module to switch off/on the GPU and
>> to obtain the GPU power state.
>> Obtaining the GPU state immediately after starting the graphical user
>> session freezes the system.
>>
>> This code triggers something, which is responsible for the freeze.
>>
>> ---
>> // Returns 1 if the card is disabled, 0 if enabled
>> static int is_card_disabled(void) {
>> u32 cfg_word;
>> // read first config word which contains Vendor and Device ID. If all bits
>> // are enabled, the device is assumed to be off
>> pci_read_config_dword(dis_dev, 0, &cfg_word);
>> // if one of the bits is not enabled (the card is enabled), the inverted
>> // result will be non-zero and hence logical not will make it 0 ("false")
>> return !~cfg_word;
>> }
>>
>> static int bbswitch_proc_show(struct seq_file *seqfp, void *p) {
>> // show the card state. Example output: 0000:01:00:00 ON
>> dis_dev_get();
>> seq_printf(seqfp, "%s %s\n", dev_name(&dis_dev->dev),
>> is_card_disabled() ? "OFF" : "ON");
>> dis_dev_put();
>> return 0;
>> }
>> ---
>>
>> Either dis_dev_get or pci_read_config_dword is the trigger.
>
> What happens if you remove the call to is_card_disabled()? Does the
> system still freeze if you only do the dis_dev_get()/dis_dev_put()?
>
>> Link to the bbswitch module source code:
>> https://github.com/Bumblebee-Project/bbswitch/blob/master/bbswitch.c#L333
>>
>>
>> Am 29.08.2016 um 18:02 schrieb Bjorn Helgaas:
>>> [+cc linux-acpi, linux-kernel, dri-devel]
>>>
>>> Hi Roland,
>>>
>>> I have no idea how to debug this problem. Are you seeing something
>>> that suggests it may be a PCI problem?
>>>
>>> On Tue, Aug 23, 2016 at 11:23:45AM +0200, Roland Singer wrote:
>>>> Hi,
>>>>
>>>> hope somebody can help me fix this kernel problem which affects the following machines:
>>>>
>>>> - Clevo P651RA (i7-6700HQ/GTX 965M, part of the P6xxRx family which are also affected)
>>>> - MSI GE62 Apache Pro (i7-6700HQ/GTX 960M)
>>>> - Gigabyte P35V5 (i7-6700HQ/GTX 970M)
>>>> - Razer Blade 14" (2016) (i7-6700HQ/GTX 970M) (BIOS 5.11, 04/07/2016)
>>>>
>>>>
>>>> The kernel freezes if the graphical user session (Xorg & Wayland) is
>>>> started with a switched off discrete GPU card (NVIDIA).
>>>> If the discrete GPU is switched off after the graphical session start,
>>>> then everything works as expected, until the graphical session is restarted.
>>>>
>>>> This problem seams to be linked to specific BIOS settings. If the computer
>>>> is started with the following command line:
>>>>
>>>> acpi_osi=! acpi_osi="Windows 2009"
>>>>
>>>> then the kernel freeze does not occur anymore. However this required a special
>>>> ACPI DSDT firmware patch for the Razer Blade 2016 laptop:
>>>>
>>>> https://github.com/m4ng0squ4sh/razer_blade_14_2016_acpi_dsdt
>>>>
>>>> I strongly recommend to fix this in the kernel and I am ready to help and solve
>>>> this problem with some help.
>>>>
>>>> Here is a link to the GitHub issue with further information:
>>>>
>>>> https://github.com/Bumblebee-Project/Bumblebee/issues/764#issuecomment-241212595
>>>>
>>>> Here are some more detailed information:
>>>>
>>>> https://github.com/Lekensteyn/acpi-stuff/blob/master/Clevo-P651RA/notes.txt
>>>>
>>>> Hope somebody can help.
>>
>> --
>> To unsubscribe from this list: send the line "unsubscribe linux-pci" in
>> the body of a message to majordomo@vger.kernel.org
>> More majordomo info at http://vger.kernel.org/majordomo-info.html
> --
> To unsubscribe from this list: send the line "unsubscribe linux-pci" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
>
[toc] | [prev] | [next] | [standalone]
| From | Bjorn Helgaas <helgaas@kernel.org> |
|---|---|
| Date | 2016-08-30 02:00 +0200 |
| Message-ID | <sbED7-5tL-19@gated-at.bofh.it> |
| In reply to | #1472085 |
On Mon, Aug 29, 2016 at 09:55:56PM +0200, Roland Singer wrote:
> Just tried it and the system didn't freeze. However it will freeze
> after some time (few minutes while working).
>
> Seams to be pci_read_config_dword. Where is this exactly defined?
pci_read_config_dword() is defined in include/linux/pci.h. It calls
pci_bus_read_config_dword() which is defined by the PCI_OP_READ() macro
in drivers/pci/access.c.
If I understand correctly, this:
dis_dev_get();
pci_read_config_dword(dis_dev, 0, &cfg_word);
dis_dev_put();
causes an immediate system hang, but if you only do this:
dis_dev_get();
dis_dev_put();
the system hangs a few minutes later. Right?
> Am 29.08.2016 um 21:07 schrieb Bjorn Helgaas:
> > On Mon, Aug 29, 2016 at 08:46:17PM +0200, Roland Singer wrote:
> >> Hi Bjorn,
> >>
> >> I am using the bbswitch kernel module to switch off/on the GPU and
> >> to obtain the GPU power state.
> >> Obtaining the GPU state immediately after starting the graphical user
> >> session freezes the system.
> >>
> >> This code triggers something, which is responsible for the freeze.
> >>
> >> ---
> >> // Returns 1 if the card is disabled, 0 if enabled
> >> static int is_card_disabled(void) {
> >> u32 cfg_word;
> >> // read first config word which contains Vendor and Device ID. If all bits
> >> // are enabled, the device is assumed to be off
> >> pci_read_config_dword(dis_dev, 0, &cfg_word);
> >> // if one of the bits is not enabled (the card is enabled), the inverted
> >> // result will be non-zero and hence logical not will make it 0 ("false")
> >> return !~cfg_word;
> >> }
> >>
> >> static int bbswitch_proc_show(struct seq_file *seqfp, void *p) {
> >> // show the card state. Example output: 0000:01:00:00 ON
> >> dis_dev_get();
> >> seq_printf(seqfp, "%s %s\n", dev_name(&dis_dev->dev),
> >> is_card_disabled() ? "OFF" : "ON");
> >> dis_dev_put();
> >> return 0;
> >> }
> >> ---
> >>
> >> Either dis_dev_get or pci_read_config_dword is the trigger.
> >
> > What happens if you remove the call to is_card_disabled()? Does the
> > system still freeze if you only do the dis_dev_get()/dis_dev_put()?
> >
> >> Link to the bbswitch module source code:
> >> https://github.com/Bumblebee-Project/bbswitch/blob/master/bbswitch.c#L333
> >>
> >>
> >> Am 29.08.2016 um 18:02 schrieb Bjorn Helgaas:
> >>> [+cc linux-acpi, linux-kernel, dri-devel]
> >>>
> >>> Hi Roland,
> >>>
> >>> I have no idea how to debug this problem. Are you seeing something
> >>> that suggests it may be a PCI problem?
> >>>
> >>> On Tue, Aug 23, 2016 at 11:23:45AM +0200, Roland Singer wrote:
> >>>> Hi,
> >>>>
> >>>> hope somebody can help me fix this kernel problem which affects the following machines:
> >>>>
> >>>> - Clevo P651RA (i7-6700HQ/GTX 965M, part of the P6xxRx family which are also affected)
> >>>> - MSI GE62 Apache Pro (i7-6700HQ/GTX 960M)
> >>>> - Gigabyte P35V5 (i7-6700HQ/GTX 970M)
> >>>> - Razer Blade 14" (2016) (i7-6700HQ/GTX 970M) (BIOS 5.11, 04/07/2016)
> >>>>
> >>>>
> >>>> The kernel freezes if the graphical user session (Xorg & Wayland) is
> >>>> started with a switched off discrete GPU card (NVIDIA).
> >>>> If the discrete GPU is switched off after the graphical session start,
> >>>> then everything works as expected, until the graphical session is restarted.
> >>>>
> >>>> This problem seams to be linked to specific BIOS settings. If the computer
> >>>> is started with the following command line:
> >>>>
> >>>> acpi_osi=! acpi_osi="Windows 2009"
> >>>>
> >>>> then the kernel freeze does not occur anymore. However this required a special
> >>>> ACPI DSDT firmware patch for the Razer Blade 2016 laptop:
> >>>>
> >>>> https://github.com/m4ng0squ4sh/razer_blade_14_2016_acpi_dsdt
> >>>>
> >>>> I strongly recommend to fix this in the kernel and I am ready to help and solve
> >>>> this problem with some help.
> >>>>
> >>>> Here is a link to the GitHub issue with further information:
> >>>>
> >>>> https://github.com/Bumblebee-Project/Bumblebee/issues/764#issuecomment-241212595
> >>>>
> >>>> Here are some more detailed information:
> >>>>
> >>>> https://github.com/Lekensteyn/acpi-stuff/blob/master/Clevo-P651RA/notes.txt
> >>>>
> >>>> Hope somebody can help.
> >>
> >> --
> >> To unsubscribe from this list: send the line "unsubscribe linux-pci" in
> >> the body of a message to majordomo@vger.kernel.org
> >> More majordomo info at http://vger.kernel.org/majordomo-info.html
> > --
> > To unsubscribe from this list: send the line "unsubscribe linux-pci" in
> > the body of a message to majordomo@vger.kernel.org
> > More majordomo info at http://vger.kernel.org/majordomo-info.html
> >
>
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-30 12:20 +0200 |
| Message-ID | <sbOj8-3sC-19@gated-at.bofh.it> |
| In reply to | #1472173 |
Thanks for pointing it out.
Yeah that's right. The system will hang randomly a few minutes later,
because some certain actions in the graphical user session will trigger
the freeze.
I had a look at the function body of pci_read_config_dword:
#define PCI_OP_READ(size, type, len) \
int pci_bus_read_config_##size \
(struct pci_bus *bus, unsigned int devfn, int pos, type *value) \
{ \
int res; \
unsigned long flags; \
u32 data = 0; \
if (PCI_##size##_BAD) return PCIBIOS_BAD_REGISTER_NUMBER; \
raw_spin_lock_irqsave(&pci_lock, flags); \
res = bus->ops->read(bus, devfn, pos, len, &data); \
*value = (type)data; \
raw_spin_unlock_irqrestore(&pci_lock, flags); \
return res; \
}
I guess, that bus->ops->read(...) might be the trigger.
Any hints how to continue debugging?
Cheers,
Roland
Am 30.08.2016 um 01:54 schrieb Bjorn Helgaas:
> On Mon, Aug 29, 2016 at 09:55:56PM +0200, Roland Singer wrote:
>> Just tried it and the system didn't freeze. However it will freeze
>> after some time (few minutes while working).
>>
>> Seams to be pci_read_config_dword. Where is this exactly defined?
>
> pci_read_config_dword() is defined in include/linux/pci.h. It calls
> pci_bus_read_config_dword() which is defined by the PCI_OP_READ() macro
> in drivers/pci/access.c.
>
> If I understand correctly, this:
>
> dis_dev_get();
> pci_read_config_dword(dis_dev, 0, &cfg_word);
> dis_dev_put();
>
> causes an immediate system hang, but if you only do this:
>
> dis_dev_get();
> dis_dev_put();
>
> the system hangs a few minutes later. Right?
>
>> Am 29.08.2016 um 21:07 schrieb Bjorn Helgaas:
>>> On Mon, Aug 29, 2016 at 08:46:17PM +0200, Roland Singer wrote:
>>>> Hi Bjorn,
>>>>
>>>> I am using the bbswitch kernel module to switch off/on the GPU and
>>>> to obtain the GPU power state.
>>>> Obtaining the GPU state immediately after starting the graphical user
>>>> session freezes the system.
>>>>
>>>> This code triggers something, which is responsible for the freeze.
>>>>
>>>> ---
>>>> // Returns 1 if the card is disabled, 0 if enabled
>>>> static int is_card_disabled(void) {
>>>> u32 cfg_word;
>>>> // read first config word which contains Vendor and Device ID. If all bits
>>>> // are enabled, the device is assumed to be off
>>>> pci_read_config_dword(dis_dev, 0, &cfg_word);
>>>> // if one of the bits is not enabled (the card is enabled), the inverted
>>>> // result will be non-zero and hence logical not will make it 0 ("false")
>>>> return !~cfg_word;
>>>> }
>>>>
>>>> static int bbswitch_proc_show(struct seq_file *seqfp, void *p) {
>>>> // show the card state. Example output: 0000:01:00:00 ON
>>>> dis_dev_get();
>>>> seq_printf(seqfp, "%s %s\n", dev_name(&dis_dev->dev),
>>>> is_card_disabled() ? "OFF" : "ON");
>>>> dis_dev_put();
>>>> return 0;
>>>> }
>>>> ---
>>>>
>>>> Either dis_dev_get or pci_read_config_dword is the trigger.
>>>
>>> What happens if you remove the call to is_card_disabled()? Does the
>>> system still freeze if you only do the dis_dev_get()/dis_dev_put()?
>>>
>>>> Link to the bbswitch module source code:
>>>> https://github.com/Bumblebee-Project/bbswitch/blob/master/bbswitch.c#L333
>>>>
>>>>
>>>> Am 29.08.2016 um 18:02 schrieb Bjorn Helgaas:
>>>>> [+cc linux-acpi, linux-kernel, dri-devel]
>>>>>
>>>>> Hi Roland,
>>>>>
>>>>> I have no idea how to debug this problem. Are you seeing something
>>>>> that suggests it may be a PCI problem?
>>>>>
>>>>> On Tue, Aug 23, 2016 at 11:23:45AM +0200, Roland Singer wrote:
>>>>>> Hi,
>>>>>>
>>>>>> hope somebody can help me fix this kernel problem which affects the following machines:
>>>>>>
>>>>>> - Clevo P651RA (i7-6700HQ/GTX 965M, part of the P6xxRx family which are also affected)
>>>>>> - MSI GE62 Apache Pro (i7-6700HQ/GTX 960M)
>>>>>> - Gigabyte P35V5 (i7-6700HQ/GTX 970M)
>>>>>> - Razer Blade 14" (2016) (i7-6700HQ/GTX 970M) (BIOS 5.11, 04/07/2016)
>>>>>>
>>>>>>
>>>>>> The kernel freezes if the graphical user session (Xorg & Wayland) is
>>>>>> started with a switched off discrete GPU card (NVIDIA).
>>>>>> If the discrete GPU is switched off after the graphical session start,
>>>>>> then everything works as expected, until the graphical session is restarted.
>>>>>>
>>>>>> This problem seams to be linked to specific BIOS settings. If the computer
>>>>>> is started with the following command line:
>>>>>>
>>>>>> acpi_osi=! acpi_osi="Windows 2009"
>>>>>>
>>>>>> then the kernel freeze does not occur anymore. However this required a special
>>>>>> ACPI DSDT firmware patch for the Razer Blade 2016 laptop:
>>>>>>
>>>>>> https://github.com/m4ng0squ4sh/razer_blade_14_2016_acpi_dsdt
>>>>>>
>>>>>> I strongly recommend to fix this in the kernel and I am ready to help and solve
>>>>>> this problem with some help.
>>>>>>
>>>>>> Here is a link to the GitHub issue with further information:
>>>>>>
>>>>>> https://github.com/Bumblebee-Project/Bumblebee/issues/764#issuecomment-241212595
>>>>>>
>>>>>> Here are some more detailed information:
>>>>>>
>>>>>> https://github.com/Lekensteyn/acpi-stuff/blob/master/Clevo-P651RA/notes.txt
>>>>>>
>>>>>> Hope somebody can help.
>>>>
>>>> --
>>>> To unsubscribe from this list: send the line "unsubscribe linux-pci" in
>>>> the body of a message to majordomo@vger.kernel.org
>>>> More majordomo info at http://vger.kernel.org/majordomo-info.html
>>> --
>>> To unsubscribe from this list: send the line "unsubscribe linux-pci" in
>>> the body of a message to majordomo@vger.kernel.org
>>> More majordomo info at http://vger.kernel.org/majordomo-info.html
>>>
>>
[toc] | [prev] | [next] | [standalone]
| From | Bjorn Helgaas <helgaas@kernel.org> |
|---|---|
| Date | 2016-08-30 15:10 +0200 |
| Message-ID | <sbQXD-5kV-17@gated-at.bofh.it> |
| In reply to | #1472390 |
On Tue, Aug 30, 2016 at 12:08:57PM +0200, Roland Singer wrote:
> Thanks for pointing it out.
>
> Yeah that's right. The system will hang randomly a few minutes later,
> because some certain actions in the graphical user session will trigger
> the freeze.
>
> I had a look at the function body of pci_read_config_dword:
>
> #define PCI_OP_READ(size, type, len) \
> int pci_bus_read_config_##size \
> (struct pci_bus *bus, unsigned int devfn, int pos, type *value) \
> { \
> int res; \
> unsigned long flags; \
> u32 data = 0; \
> if (PCI_##size##_BAD) return PCIBIOS_BAD_REGISTER_NUMBER; \
> raw_spin_lock_irqsave(&pci_lock, flags); \
> res = bus->ops->read(bus, devfn, pos, len, &data); \
> *value = (type)data; \
> raw_spin_unlock_irqrestore(&pci_lock, flags); \
> return res; \
> }
>
> I guess, that bus->ops->read(...) might be the trigger.
> Any hints how to continue debugging?
It's not likely that the problem is in the bus->ops->read() path. That
is used by every device driver, so a problem there would cause more
serious problems than what you're seeing.
My guess would be some problem in the video driver or the bbswitch
thing.
> Am 30.08.2016 um 01:54 schrieb Bjorn Helgaas:
> > On Mon, Aug 29, 2016 at 09:55:56PM +0200, Roland Singer wrote:
> >> Just tried it and the system didn't freeze. However it will freeze
> >> after some time (few minutes while working).
> >>
> >> Seams to be pci_read_config_dword. Where is this exactly defined?
> >
> > pci_read_config_dword() is defined in include/linux/pci.h. It calls
> > pci_bus_read_config_dword() which is defined by the PCI_OP_READ() macro
> > in drivers/pci/access.c.
> >
> > If I understand correctly, this:
> >
> > dis_dev_get();
> > pci_read_config_dword(dis_dev, 0, &cfg_word);
> > dis_dev_put();
> >
> > causes an immediate system hang, but if you only do this:
> >
> > dis_dev_get();
> > dis_dev_put();
> >
> > the system hangs a few minutes later. Right?
> >
> >> Am 29.08.2016 um 21:07 schrieb Bjorn Helgaas:
> >>> On Mon, Aug 29, 2016 at 08:46:17PM +0200, Roland Singer wrote:
> >>>> Hi Bjorn,
> >>>>
> >>>> I am using the bbswitch kernel module to switch off/on the GPU and
> >>>> to obtain the GPU power state.
> >>>> Obtaining the GPU state immediately after starting the graphical user
> >>>> session freezes the system.
> >>>>
> >>>> This code triggers something, which is responsible for the freeze.
> >>>>
> >>>> ---
> >>>> // Returns 1 if the card is disabled, 0 if enabled
> >>>> static int is_card_disabled(void) {
> >>>> u32 cfg_word;
> >>>> // read first config word which contains Vendor and Device ID. If all bits
> >>>> // are enabled, the device is assumed to be off
> >>>> pci_read_config_dword(dis_dev, 0, &cfg_word);
> >>>> // if one of the bits is not enabled (the card is enabled), the inverted
> >>>> // result will be non-zero and hence logical not will make it 0 ("false")
> >>>> return !~cfg_word;
> >>>> }
> >>>>
> >>>> static int bbswitch_proc_show(struct seq_file *seqfp, void *p) {
> >>>> // show the card state. Example output: 0000:01:00:00 ON
> >>>> dis_dev_get();
> >>>> seq_printf(seqfp, "%s %s\n", dev_name(&dis_dev->dev),
> >>>> is_card_disabled() ? "OFF" : "ON");
> >>>> dis_dev_put();
> >>>> return 0;
> >>>> }
> >>>> ---
> >>>>
> >>>> Either dis_dev_get or pci_read_config_dword is the trigger.
> >>>
> >>> What happens if you remove the call to is_card_disabled()? Does the
> >>> system still freeze if you only do the dis_dev_get()/dis_dev_put()?
> >>>
> >>>> Link to the bbswitch module source code:
> >>>> https://github.com/Bumblebee-Project/bbswitch/blob/master/bbswitch.c#L333
> >>>>
> >>>>
> >>>> Am 29.08.2016 um 18:02 schrieb Bjorn Helgaas:
> >>>>> [+cc linux-acpi, linux-kernel, dri-devel]
> >>>>>
> >>>>> Hi Roland,
> >>>>>
> >>>>> I have no idea how to debug this problem. Are you seeing something
> >>>>> that suggests it may be a PCI problem?
> >>>>>
> >>>>> On Tue, Aug 23, 2016 at 11:23:45AM +0200, Roland Singer wrote:
> >>>>>> Hi,
> >>>>>>
> >>>>>> hope somebody can help me fix this kernel problem which affects the following machines:
> >>>>>>
> >>>>>> - Clevo P651RA (i7-6700HQ/GTX 965M, part of the P6xxRx family which are also affected)
> >>>>>> - MSI GE62 Apache Pro (i7-6700HQ/GTX 960M)
> >>>>>> - Gigabyte P35V5 (i7-6700HQ/GTX 970M)
> >>>>>> - Razer Blade 14" (2016) (i7-6700HQ/GTX 970M) (BIOS 5.11, 04/07/2016)
> >>>>>>
> >>>>>>
> >>>>>> The kernel freezes if the graphical user session (Xorg & Wayland) is
> >>>>>> started with a switched off discrete GPU card (NVIDIA).
> >>>>>> If the discrete GPU is switched off after the graphical session start,
> >>>>>> then everything works as expected, until the graphical session is restarted.
> >>>>>>
> >>>>>> This problem seams to be linked to specific BIOS settings. If the computer
> >>>>>> is started with the following command line:
> >>>>>>
> >>>>>> acpi_osi=! acpi_osi="Windows 2009"
> >>>>>>
> >>>>>> then the kernel freeze does not occur anymore. However this required a special
> >>>>>> ACPI DSDT firmware patch for the Razer Blade 2016 laptop:
> >>>>>>
> >>>>>> https://github.com/m4ng0squ4sh/razer_blade_14_2016_acpi_dsdt
> >>>>>>
> >>>>>> I strongly recommend to fix this in the kernel and I am ready to help and solve
> >>>>>> this problem with some help.
> >>>>>>
> >>>>>> Here is a link to the GitHub issue with further information:
> >>>>>>
> >>>>>> https://github.com/Bumblebee-Project/Bumblebee/issues/764#issuecomment-241212595
> >>>>>>
> >>>>>> Here are some more detailed information:
> >>>>>>
> >>>>>> https://github.com/Lekensteyn/acpi-stuff/blob/master/Clevo-P651RA/notes.txt
> >>>>>>
> >>>>>> Hope somebody can help.
> >>>>
> >>>> --
> >>>> To unsubscribe from this list: send the line "unsubscribe linux-pci" in
> >>>> the body of a message to majordomo@vger.kernel.org
> >>>> More majordomo info at http://vger.kernel.org/majordomo-info.html
> >>> --
> >>> To unsubscribe from this list: send the line "unsubscribe linux-pci" in
> >>> the body of a message to majordomo@vger.kernel.org
> >>> More majordomo info at http://vger.kernel.org/majordomo-info.html
> >>>
> >>
>
[toc] | [prev] | [next] | [standalone]
| From | Emil Velikov <emil.l.velikov@gmail.com> |
|---|---|
| Date | 2016-08-30 16:10 +0200 |
| Message-ID | <sbRTH-5Ua-9@gated-at.bofh.it> |
| In reply to | #1472475 |
On 30 August 2016 at 14:06, Bjorn Helgaas <helgaas@kernel.org> wrote:
> On Tue, Aug 30, 2016 at 12:08:57PM +0200, Roland Singer wrote:
>> Thanks for pointing it out.
>>
>> Yeah that's right. The system will hang randomly a few minutes later,
>> because some certain actions in the graphical user session will trigger
>> the freeze.
>>
>> I had a look at the function body of pci_read_config_dword:
>>
>> #define PCI_OP_READ(size, type, len) \
>> int pci_bus_read_config_##size \
>> (struct pci_bus *bus, unsigned int devfn, int pos, type *value) \
>> { \
>> int res; \
>> unsigned long flags; \
>> u32 data = 0; \
>> if (PCI_##size##_BAD) return PCIBIOS_BAD_REGISTER_NUMBER; \
>> raw_spin_lock_irqsave(&pci_lock, flags); \
>> res = bus->ops->read(bus, devfn, pos, len, &data); \
>> *value = (type)data; \
>> raw_spin_unlock_irqrestore(&pci_lock, flags); \
>> return res; \
>> }
>>
>> I guess, that bus->ops->read(...) might be the trigger.
>> Any hints how to continue debugging?
>
> It's not likely that the problem is in the bus->ops->read() path. That
> is used by every device driver, so a problem there would cause more
> serious problems than what you're seeing.
>
> My guess would be some problem in the video driver or the bbswitch
> thing.
>
FWIW I'm inclined to call it a bbswitch bug. It can (and does when
needed) power off the dedicated GPU.
Depending on the platform different methods are used:
Sometimes the GPU driver will get 0xffffffff (or similar) when trying
to read from the device mmio space. While one can say that the driver
should attribute for this, IMHO it's a bad idea to have two drivers
controlling the same hardware, let alone without any coordination
between them.
IIRC in some cases the device can disappear from the PCI bus (not 100%
sure this one). In which case a simple read can lead to a wide range
of fireworks.
Disclaimer: it's been a while since I've looked into bbswitch so
things might have changed/improved.
Regards,
Emil
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-30 17:30 +0200 |
| Message-ID | <sbT97-6Di-13@gated-at.bofh.it> |
| In reply to | #1472498 |
I tried these scenarios:
1. Booted the system without the bbswitch module. The nouveau module
was loaded and is responsible for the power management of the GPU.
The graphical session freezes after some minutes...
2. Booted the system without bbswitch and with nouveau blacklisted.
Manually loaded bbswitch to switch off the discrete GPU.
Same freeze after a while or by explicitly obtaining the GPU state.
Is there a possibility to switch off the discrete card without bbswitch?
If this is possible, then I could test this without nouveau and bbswitch
at all. If the system hangs, then it is not the video driver nor bbswitch.
Am 30.08.2016 um 16:08 schrieb Emil Velikov:
> On 30 August 2016 at 14:06, Bjorn Helgaas <helgaas@kernel.org> wrote:
>> On Tue, Aug 30, 2016 at 12:08:57PM +0200, Roland Singer wrote:
>>> Thanks for pointing it out.
>>>
>>> Yeah that's right. The system will hang randomly a few minutes later,
>>> because some certain actions in the graphical user session will trigger
>>> the freeze.
>>>
>>> I had a look at the function body of pci_read_config_dword:
>>>
>>> #define PCI_OP_READ(size, type, len) \
>>> int pci_bus_read_config_##size \
>>> (struct pci_bus *bus, unsigned int devfn, int pos, type *value) \
>>> { \
>>> int res; \
>>> unsigned long flags; \
>>> u32 data = 0; \
>>> if (PCI_##size##_BAD) return PCIBIOS_BAD_REGISTER_NUMBER; \
>>> raw_spin_lock_irqsave(&pci_lock, flags); \
>>> res = bus->ops->read(bus, devfn, pos, len, &data); \
>>> *value = (type)data; \
>>> raw_spin_unlock_irqrestore(&pci_lock, flags); \
>>> return res; \
>>> }
>>>
>>> I guess, that bus->ops->read(...) might be the trigger.
>>> Any hints how to continue debugging?
>>
>> It's not likely that the problem is in the bus->ops->read() path. That
>> is used by every device driver, so a problem there would cause more
>> serious problems than what you're seeing.
>>
>> My guess would be some problem in the video driver or the bbswitch
>> thing.
>>
> FWIW I'm inclined to call it a bbswitch bug. It can (and does when
> needed) power off the dedicated GPU.
>
> Depending on the platform different methods are used:
>
> Sometimes the GPU driver will get 0xffffffff (or similar) when trying
> to read from the device mmio space. While one can say that the driver
> should attribute for this, IMHO it's a bad idea to have two drivers
> controlling the same hardware, let alone without any coordination
> between them.
>
> IIRC in some cases the device can disappear from the PCI bus (not 100%
> sure this one). In which case a simple read can lead to a wide range
> of fireworks.
>
> Disclaimer: it's been a while since I've looked into bbswitch so
> things might have changed/improved.
>
> Regards,
> Emil
> --
> To unsubscribe from this list: send the line "unsubscribe linux-pci" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
>
[toc] | [prev] | [next] | [standalone]
| From | Ilia Mirkin <imirkin@alum.mit.edu> |
|---|---|
| Date | 2016-08-30 17:50 +0200 |
| Message-ID | <sbTst-6JN-13@gated-at.bofh.it> |
| In reply to | #1472559 |
On Tue, Aug 30, 2016 at 11:25 AM, Roland Singer <roland.singer@desertbit.com> wrote: > I tried these scenarios: > > 1. Booted the system without the bbswitch module. The nouveau module > was loaded and is responsible for the power management of the GPU. > The graphical session freezes after some minutes... > > 2. Booted the system without bbswitch and with nouveau blacklisted. > Manually loaded bbswitch to switch off the discrete GPU. > Same freeze after a while or by explicitly obtaining the GPU state. > > Is there a possibility to switch off the discrete card without bbswitch? > If this is possible, then I could test this without nouveau and bbswitch > at all. If the system hangs, then it is not the video driver nor bbswitch. You can use acpi_call (a random search points to https://github.com/mkottman/acpi_call, but I don't know if that's the "official" version) - need to find the right method to call, but that's basically all it takes to acpi-suspend a gpu. Separately, there was a recent fix to ... something, including but not limited to nouveau, involving hangs on gpu suspend on newer laptops. I don't think it's upstream yet. Look for patches from Lukas Wunner. -ilia
[toc] | [prev] | [next] | [standalone]
| From | Ilia Mirkin <imirkin@alum.mit.edu> |
|---|---|
| Date | 2016-08-30 17:50 +0200 |
| Message-ID | <sbTst-6JN-25@gated-at.bofh.it> |
| In reply to | #1472568 |
On Tue, Aug 30, 2016 at 11:44 AM, Ilia Mirkin <imirkin@alum.mit.edu> wrote:
> On Tue, Aug 30, 2016 at 11:25 AM, Roland Singer
> <roland.singer@desertbit.com> wrote:
>> I tried these scenarios:
>>
>> 1. Booted the system without the bbswitch module. The nouveau module
>> was loaded and is responsible for the power management of the GPU.
>> The graphical session freezes after some minutes...
>>
>> 2. Booted the system without bbswitch and with nouveau blacklisted.
>> Manually loaded bbswitch to switch off the discrete GPU.
>> Same freeze after a while or by explicitly obtaining the GPU state.
>>
>> Is there a possibility to switch off the discrete card without bbswitch?
>> If this is possible, then I could test this without nouveau and bbswitch
>> at all. If the system hangs, then it is not the video driver nor bbswitch.
>
> You can use acpi_call (a random search points to
> https://github.com/mkottman/acpi_call, but I don't know if that's the
> "official" version) - need to find the right method to call, but
> that's basically all it takes to acpi-suspend a gpu.
>
> Separately, there was a recent fix to ... something, including but not
> limited to nouveau, involving hangs on gpu suspend on newer laptops. I
> don't think it's upstream yet. Look for patches from Lukas Wunner.
Er oops. Looks like I misremembered. Patches are from Peter Wu, and at
least one of them is in v4.8-rc1:
commit 692a17dcc2922a91c6bcf11b3321503a3377b1b1
Author: Peter Wu <peter@lekensteyn.nl>
Date: Fri Jul 15 15:12:18 2016 +0200
drm/nouveau/acpi: fix lockup with PCIe runtime PM
along with a number of other related patches. It's not clear which
kernel you were trying this with... can you give v4.8-rcN a shot?
-ilia
[toc] | [prev] | [next] | [standalone]
| From | Emil Velikov <emil.l.velikov@gmail.com> |
|---|---|
| Date | 2016-08-30 17:50 +0200 |
| Message-ID | <sbTsu-6JN-31@gated-at.bofh.it> |
| In reply to | #1472559 |
On 30 August 2016 at 16:25, Roland Singer <roland.singer@desertbit.com> wrote: > I tried these scenarios: > > 1. Booted the system without the bbswitch module. The nouveau module > was loaded and is responsible for the power management of the GPU. > The graphical session freezes after some minutes... > > 2. Booted the system without bbswitch and with nouveau blacklisted. > Manually loaded bbswitch to switch off the discrete GPU. > Same freeze after a while or by explicitly obtaining the GPU state. > > Is there a possibility to switch off the discrete card without bbswitch? > If this is possible, then I could test this without nouveau and bbswitch > at all. If the system hangs, then it is not the video driver nor bbswitch. > As Ilia mentioned acpi_call should do it. You can also check with the nouveau/bbwswitch code to see which ones they use in your case and bash it manually. It might be that the 'wrong one' gets used thus things going horribly wrong. Regards, Emil
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-30 19:40 +0200 |
| Message-ID | <sbVaX-7T7-63@gated-at.bofh.it> |
| In reply to | #1472577 |
I am running 4.7.2, but I also just tried the 4.8.0-rc4 mainline kernel.
The result is the same. There is no difference if bbswitch of acpi_call
is used. However I noticed following:
1. The nouveau driver is broken in both kernel version and is responsible
for the freezes while gathering power state information with bbswitch.
Sometimes while shutting the system down, everything except the LCD
screen is switched off. This only happens with nouveau.
I noticed following error log messages:
kernel: nouveau 0000:01:00.0: fb: 6144 MiB GDDR5
kernel: nouveau 0000:01:00.0: priv: HUB0: 10ecc0 ffffffff (1e40822c)
kernel: nouveau 0000:01:00.0: DRM: VRAM: 6144 MiB
kernel: nouveau 0000:01:00.0: DRM: GART: 1048576 MiB
kernel: nouveau 0000:01:00.0: DRM: Pointer to TMDS table invalid
kernel: nouveau 0000:01:00.0: DRM: DCB version 4.1
kernel: nouveau 0000:01:00.0: DRM: Pointer to flat panel table invalid
2. -> Boot with nouveau module loaded
-> switch off the discrete GPU with bbswitch or acpi_call
-> start X11
-> obtaining power state with bbswitch freezes the system
-> or working with the system for some minutes freezes the system
3. -> Boot with nouveau module blacklisted
-> switch off the discrete GPU
-> start X11
-> system immediately freezes
4. -> Boot with nouveau module blacklisted
-> switch off the discrete GPU
-> start Wayland
-> system runs - Note: I tried this for couple of days with 4.6 and 4.7 mainline
and the system freezed randomly after some time.
However I have to test if this is still present with 4.7.2
and 4.8 mainline. Right now it seams to be fine.
-> running Xwayland (does not depend on the GPU power state) kills performance!
the system freezes for several seconds...
So working with Wayland is also no solution.
My conclusion:
1. Nouveau has couple of problems with GTX 9** M Nvidia GPUs.
I would love to help here.
2. X11 is just broken and is not capable to start the graphical session
if the nvidia GPU is not handled by any video driver (kernel module).
Even forcing X11 to ignore the discrete GPU doesn't help.
Setting the command line arguments to:
acpi_osi=! acpi_osi="Windows 2009"
fixes the issues with X11 but other things break...
What the hell is going on?! :/
Am 30.08.2016 um 17:48 schrieb Emil Velikov:
> On 30 August 2016 at 16:25, Roland Singer <roland.singer@desertbit.com> wrote:
>> I tried these scenarios:
>>
>> 1. Booted the system without the bbswitch module. The nouveau module
>> was loaded and is responsible for the power management of the GPU.
>> The graphical session freezes after some minutes...
>>
>> 2. Booted the system without bbswitch and with nouveau blacklisted.
>> Manually loaded bbswitch to switch off the discrete GPU.
>> Same freeze after a while or by explicitly obtaining the GPU state.
>>
>> Is there a possibility to switch off the discrete card without bbswitch?
>> If this is possible, then I could test this without nouveau and bbswitch
>> at all. If the system hangs, then it is not the video driver nor bbswitch.
>>
> As Ilia mentioned acpi_call should do it. You can also check with the
> nouveau/bbwswitch code to see which ones they use in your case and
> bash it manually. It might be that the 'wrong one' gets used thus
> things going horribly wrong.
>
> Regards,
> Emil
>
[toc] | [prev] | [next] | [standalone]
| From | Ilia Mirkin <imirkin@alum.mit.edu> |
|---|---|
| Date | 2016-08-30 19:50 +0200 |
| Message-ID | <sbVkC-7WI-31@gated-at.bofh.it> |
| In reply to | #1472698 |
On Tue, Aug 30, 2016 at 1:37 PM, Roland Singer <roland.singer@desertbit.com> wrote: > My conclusion: > > 1. Nouveau has couple of problems with GTX 9** M Nvidia GPUs. > I would love to help here. nouveau + bbswitch will always end in tears. You're going behind the driver's back and messing around with state it believes it is managing. What if you just use nouveau and let it auto-power-off the GPU like it's designed to, with v4.8-rc? -ilia
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-30 20:10 +0200 |
| Message-ID | <sbVDX-8iA-1@gated-at.bofh.it> |
| In reply to | #1472719 |
I configured bbswitch to not set any states automatically...
So it's possible to obtain and verify the GPU power state.
However I removed the bbswitch module and booted with nouveau.
Kernel 4.7.2: nouveau switches the discrete GPU off.
I can't trigger the freeze, because bbswitch is missing.
I'll work with the system and see if it will freeze.
Kernel 4.8-rc4: nouveau does not care about the power state and
the discrete GPU is never switched off. I will notice
this, because the second cooling FAN will stop...
Same log messages as send before.
Am 30.08.2016 um 19:43 schrieb Ilia Mirkin:
> On Tue, Aug 30, 2016 at 1:37 PM, Roland Singer
> <roland.singer@desertbit.com> wrote:
>> My conclusion:
>>
>> 1. Nouveau has couple of problems with GTX 9** M Nvidia GPUs.
>> I would love to help here.
>
> nouveau + bbswitch will always end in tears. You're going behind the
> driver's back and messing around with state it believes it is
> managing. What if you just use nouveau and let it auto-power-off the
> GPU like it's designed to, with v4.8-rc?
>
> -ilia
>
[toc] | [prev] | [next] | [standalone]
| From | Ilia Mirkin <imirkin@alum.mit.edu> |
|---|---|
| Date | 2016-08-30 20:20 +0200 |
| Message-ID | <sbVNE-8n7-37@gated-at.bofh.it> |
| In reply to | #1472741 |
On Tue, Aug 30, 2016 at 2:02 PM, Roland Singer <roland.singer@desertbit.com> wrote: > I configured bbswitch to not set any states automatically... > So it's possible to obtain and verify the GPU power state. > > However I removed the bbswitch module and booted with nouveau. > > Kernel 4.7.2: nouveau switches the discrete GPU off. > I can't trigger the freeze, because bbswitch is missing. > I'll work with the system and see if it will freeze. > > Kernel 4.8-rc4: nouveau does not care about the power state and > the discrete GPU is never switched off. I will notice > this, because the second cooling FAN will stop... > Same log messages as send before. That's surprising. I believe there's an issue with the new logic when there's an HDMI audio subdevice. However that only appears if there's a cable plugged in, at least in the systems Peter tested. You should be able to see whether it's there or not with 'lspci'. You can check for sure by looking in the vgaswitcheroo state. It should say DynOff when it's powered off. Either way, I think using bbswitch + nouveau isn't supported by anyone, so if you want to use it that way, you're on your own. (You may want to load nouveau with runpm=0 so that nouveau doesn't try to manage the GPU suspend stuff.) -ilia
[toc] | [prev] | [next] | [standalone]
| From | Peter Wu <peter@lekensteyn.nl> |
|---|---|
| Date | 2016-08-30 21:30 +0200 |
| Message-ID | <sbWTn-Dd-13@gated-at.bofh.it> |
| In reply to | #1472748 |
On Tue, Aug 30, 2016 at 02:13:46PM -0400, Ilia Mirkin wrote: > On Tue, Aug 30, 2016 at 2:02 PM, Roland Singer > <roland.singer@desertbit.com> wrote: > > I configured bbswitch to not set any states automatically... > > So it's possible to obtain and verify the GPU power state. > > > > However I removed the bbswitch module and booted with nouveau. > > > > Kernel 4.7.2: nouveau switches the discrete GPU off. > > I can't trigger the freeze, because bbswitch is missing. > > I'll work with the system and see if it will freeze. > > > > Kernel 4.8-rc4: nouveau does not care about the power state and > > the discrete GPU is never switched off. I will notice > > this, because the second cooling FAN will stop... > > Same log messages as send before. > > That's surprising. I believe there's an issue with the new logic when > there's an HDMI audio subdevice. However that only appears if there's > a cable plugged in, at least in the systems Peter tested. You should > be able to see whether it's there or not with 'lspci'. I doubt that the audio device is responsible here, that should only show up after following very specific steps (runtime suspend/resume (PCI or ACPI magic), remove PCI device, rescan bus). > You can check for sure by looking in the vgaswitcheroo state. It > should say DynOff when it's powered off. > > Either way, I think using bbswitch + nouveau isn't supported by > anyone, so if you want to use it that way, you're on your own. (You > may want to load nouveau with runpm=0 so that nouveau doesn't try to > manage the GPU suspend stuff.) I understood that Roland's intent is to check the power state, not use the suspend functionality of bbswitch, if you load bbswitch without module options amd do not write to /proc/bbswitch, then it allows you to read out the actual status (you could also just use lspci -H1 for that though). -- Kind regards, Peter Wu https://lekensteyn.nl
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-31 13:20 +0200 |
| Message-ID | <scbIK-1JI-7@gated-at.bofh.it> |
| In reply to | #1472810 |
Am 30.08.2016 um 21:21 schrieb Peter Wu:
> On Tue, Aug 30, 2016 at 02:13:46PM -0400, Ilia Mirkin wrote:
>> On Tue, Aug 30, 2016 at 2:02 PM, Roland Singer
>> <roland.singer@desertbit.com> wrote:
>>> I configured bbswitch to not set any states automatically...
>>> So it's possible to obtain and verify the GPU power state.
>>>
>>> However I removed the bbswitch module and booted with nouveau.
>>>
>>> Kernel 4.7.2: nouveau switches the discrete GPU off.
>>> I can't trigger the freeze, because bbswitch is missing.
>>> I'll work with the system and see if it will freeze.
>>>
>>> Kernel 4.8-rc4: nouveau does not care about the power state and
>>> the discrete GPU is never switched off. I will notice
>>> this, because the second cooling FAN will stop...
>>> Same log messages as send before.
>>
>> That's surprising. I believe there's an issue with the new logic when
>> there's an HDMI audio subdevice. However that only appears if there's
>> a cable plugged in, at least in the systems Peter tested. You should
>> be able to see whether it's there or not with 'lspci'.
>
> I doubt that the audio device is responsible here, that should only show
> up after following very specific steps (runtime suspend/resume (PCI or
> ACPI magic), remove PCI device, rescan bus).
>
>> You can check for sure by looking in the vgaswitcheroo state. It
>> should say DynOff when it's powered off.
>>
>> Either way, I think using bbswitch + nouveau isn't supported by
>> anyone, so if you want to use it that way, you're on your own. (You
>> may want to load nouveau with runpm=0 so that nouveau doesn't try to
>> manage the GPU suspend stuff.)
>
> I understood that Roland's intent is to check the power state, not use
> the suspend functionality of bbswitch, if you load bbswitch without
> module options amd do not write to /proc/bbswitch, then it allows you to
> read out the actual status (you could also just use lspci -H1 for that
> though).
>
lspci -H1 works perfect. Thanks.
Just tried to verify the output with lspci -H1. I unloaded the nouveau
module and modprobe freezed with:
$ modprobe -r nouveau
nouveau 0000:01:00.0: pci: failed to adjust lnkctl speed
nouveau 0000:01:00.0: fb: init failed. -22
nouveau 0000:01:00.0: init failed with -22
nouveau: DRM:00000000:00000000: init failed with -22
nouveau: DRM:00000000:00000000: init failed with -22
[toc] | [prev] | [next] | [standalone]
| From | Roland Singer <roland.singer@desertbit.com> |
|---|---|
| Date | 2016-08-31 13:20 +0200 |
| Message-ID | <scbIK-1JI-9@gated-at.bofh.it> |
| In reply to | #1472748 |
Am 30.08.2016 um 20:13 schrieb Ilia Mirkin: > On Tue, Aug 30, 2016 at 2:02 PM, Roland Singer > <roland.singer@desertbit.com> wrote: >> I configured bbswitch to not set any states automatically... >> So it's possible to obtain and verify the GPU power state. >> >> However I removed the bbswitch module and booted with nouveau. >> >> Kernel 4.7.2: nouveau switches the discrete GPU off. >> I can't trigger the freeze, because bbswitch is missing. >> I'll work with the system and see if it will freeze. >> >> Kernel 4.8-rc4: nouveau does not care about the power state and >> the discrete GPU is never switched off. I will notice >> this, because the second cooling FAN will stop... >> Same log messages as send before. > > That's surprising. I believe there's an issue with the new logic when > there's an HDMI audio subdevice. However that only appears if there's > a cable plugged in, at least in the systems Peter tested. You should > be able to see whether it's there or not with 'lspci'. > > You can check for sure by looking in the vgaswitcheroo state. It > should say DynOff when it's powered off. > > Either way, I think using bbswitch + nouveau isn't supported by > anyone, so if you want to use it that way, you're on your own. (You > may want to load nouveau with runpm=0 so that nouveau doesn't try to > manage the GPU suspend stuff.) > > -ilia > Kernel 4.8-rc4: While running lspci, following kernel log message was printed on the TTY: nouveau: 0000:01:00:0: priv: HUB0: 6013d4 0000573f (1f408200) nouveau: 0000:01:00:0: priv: HUB0: 10ecc0 ffffffff (1940822c) This is my output of lspci: 00:00.0 Host bridge: Intel Corporation Skylake Host Bridge/DRAM Registers (rev 07) 00:01.0 PCI bridge: Intel Corporation Skylake PCIe Controller (x16) (rev 07) 00:02.0 VGA compatible controller: Intel Corporation HD Graphics 530 (rev 06) 00:08.0 System peripheral: Intel Corporation Skylake Gaussian Mixture Model 00:14.0 USB controller: Intel Corporation Sunrise Point-H USB 3.0 xHCI Controller (rev 31) 00:14.2 Signal processing controller: Intel Corporation Sunrise Point-H Thermal subsystem (rev 31) 00:15.0 Signal processing controller: Intel Corporation Sunrise Point-H Serial IO I2C Controller #0 (rev 31) 00:15.1 Signal processing controller: Intel Corporation Sunrise Point-H Serial IO I2C Controller #1 (rev 31) 00:16.0 Communication controller: Intel Corporation Sunrise Point-H CSME HECI #1 (rev 31) 00:1c.0 PCI bridge: Intel Corporation Sunrise Point-H PCI Express Root Port #1 (rev f1) 00:1c.5 PCI bridge: Intel Corporation Sunrise Point-H PCI Express Root Port #6 (rev f1) 00:1d.0 PCI bridge: Intel Corporation Sunrise Point-H PCI Express Root Port #9 (rev f1) 00:1d.4 PCI bridge: Intel Corporation Sunrise Point-H PCI Express Root Port #13 (rev f1) 00:1e.0 Signal processing controller: Intel Corporation Sunrise Point-H Serial IO UART #0 (rev 31) 00:1f.0 ISA bridge: Intel Corporation Sunrise Point-H LPC Controller (rev 31) 00:1f.2 Memory controller: Intel Corporation Sunrise Point-H PMC (rev 31) 00:1f.3 Audio device: Intel Corporation Sunrise Point-H HD Audio (rev 31) 00:1f.4 SMBus: Intel Corporation Sunrise Point-H SMBus (rev 31) 01:00.0 3D controller: NVIDIA Corporation GM204M [GeForce GTX 970M] (rev a1) 3b:00.0 Network controller: Qualcomm Atheros QCA6174 802.11ac Wireless Network Adapter (rev 32) 3d:00.0 Non-Volatile memory controller: Samsung Electronics Co Ltd NVMe SSD Controller (rev 01)
[toc] | [prev] | [next] | [standalone]
| From | Emil Velikov <emil.l.velikov@gmail.com> |
|---|---|
| Date | 2016-08-30 20:10 +0200 |
| Message-ID | <sbVDY-8iA-27@gated-at.bofh.it> |
| In reply to | #1472698 |
On 30 August 2016 at 18:37, Roland Singer <roland.singer@desertbit.com> wrote: > I am running 4.7.2, but I also just tried the 4.8.0-rc4 mainline kernel. > The result is the same. There is no difference if bbswitch of acpi_call > is used. However I noticed following: > > 1. The nouveau driver is broken in both kernel version and is responsible > for the freezes while gathering power state information with bbswitch. > Sometimes while shutting the system down, everything except the LCD > screen is switched off. This only happens with nouveau. > I noticed following error log messages: > I second Ilia here. Using bbswitch in conjunction with any driver (be that nouveau or the proprietary one) is a bad idea. > kernel: nouveau 0000:01:00.0: fb: 6144 MiB GDDR5 > kernel: nouveau 0000:01:00.0: priv: HUB0: 10ecc0 ffffffff (1e40822c) > kernel: nouveau 0000:01:00.0: DRM: VRAM: 6144 MiB > kernel: nouveau 0000:01:00.0: DRM: GART: 1048576 MiB > kernel: nouveau 0000:01:00.0: DRM: Pointer to TMDS table invalid > kernel: nouveau 0000:01:00.0: DRM: DCB version 4.1 > kernel: nouveau 0000:01:00.0: DRM: Pointer to flat panel table invalid > > 2. -> Boot with nouveau module loaded > -> switch off the discrete GPU with bbswitch or acpi_call > -> start X11 > -> obtaining power state with bbswitch freezes the system > -> or working with the system for some minutes freezes the system > (If Ilia's suggestions does not help) Confirm if the freeze is due to/as the GPU is powered on or off. > 3. -> Boot with nouveau module blacklisted > -> switch off the discrete GPU > -> start X11 > -> system immediately freezes > It's perfectly possible that the discrete GPU is set as boot one and X goes angry since there's no driver/way to bring it up. > 4. -> Boot with nouveau module blacklisted > -> switch off the discrete GPU > -> start Wayland > -> system runs - Note: I tried this for couple of days with 4.6 and 4.7 mainline > and the system freezed randomly after some time. > However I have to test if this is still present with 4.7.2 > and 4.8 mainline. Right now it seams to be fine. > -> running Xwayland (does not depend on the GPU power state) kills performance! > the system freezes for several seconds... > So working with Wayland is also no solution. > > My conclusion: > > 1. Nouveau has couple of problems with GTX 9** M Nvidia GPUs. > I would love to help here. > > 2. X11 is just broken and is not capable to start the graphical session > if the nvidia GPU is not handled by any video driver (kernel module). > Even forcing X11 to ignore the discrete GPU doesn't help. > Out of curiosity: how did you force X to ignore the device ? > Setting the command line arguments to: > > acpi_osi=! acpi_osi="Windows 2009" > > fixes the issues with X11 but other things break... > What the hell is going on?! :/ > You can check if it's the boot_vga assumption with Check wh You're a victum of the Windows specific fun (quirks?) in > Am 30.08.2016 um 17:48 schrieb Emil Velikov: >> On 30 August 2016 at 16:25, Roland Singer <roland.singer@desertbit.com> wrote: >>> I tried these scenarios: >>> >>> 1. Booted the system without the bbswitch module. The nouveau module >>> was loaded and is responsible for the power management of the GPU. >>> The graphical session freezes after some minutes... >>> >>> 2. Booted the system without bbswitch and with nouveau blacklisted. >>> Manually loaded bbswitch to switch off the discrete GPU. >>> Same freeze after a while or by explicitly obtaining the GPU state. >>> >>> Is there a possibility to switch off the discrete card without bbswitch? >>> If this is possible, then I could test this without nouveau and bbswitch >>> at all. If the system hangs, then it is not the video driver nor bbswitch. >>> >> As Ilia mentioned acpi_call should do it. You can also check with the >> nouveau/bbwswitch code to see which ones they use in your case and >> bash it manually. It might be that the 'wrong one' gets used thus >> things going horribly wrong. >> >> Regards, >> Emil >> >
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web