Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1553763 > unrolled thread

Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

Started by"Rafael J. Wysocki" <rafael.j.wysocki@intel.com>
First post2017-01-08 00:40 +0100
Last post2017-01-10 15:00 +0100
Articles 20 on this page of 39 — 7 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size()  and early_acpi_os_unmap_memory() from Linux kernel "Rafael J. Wysocki" <rafael.j.wysocki@intel.com> - 2017-01-08 00:40 +0100
    Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-08 01:10 +0100
      Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size()  and early_acpi_os_unmap_memory() from Linux kernel "Rafael J. Wysocki" <rafael@kernel.org> - 2017-01-08 01:30 +0100
        Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-08 01:40 +0100
          Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size()  and early_acpi_os_unmap_memory() from Linux kernel "Rafael J. Wysocki" <rafael@kernel.org> - 2017-01-08 02:00 +0100
            Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-08 02:10 +0100
              Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size()  and early_acpi_os_unmap_memory() from Linux kernel "Rafael J. Wysocki" <rafael@kernel.org> - 2017-01-08 02:50 +0100
                Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size()  and early_acpi_os_unmap_memory() from Linux kernel "Rafael J. Wysocki" <rafael@kernel.org> - 2017-01-08 03:30 +0100
                  Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-08 14:10 +0100
                    RE: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel "Zheng, Lv" <lv.zheng@intel.com> - 2017-01-09 03:00 +0100
                      RE: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel "Zheng, Lv" <lv.zheng@intel.com> - 2017-01-09 03:40 +0100
                        Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-09 10:40 +0100
                          Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-01-09 23:20 +0100
                            Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size()  and early_acpi_os_unmap_memory() from Linux kernel "Rafael J. Wysocki" <rafael@kernel.org> - 2017-01-09 23:30 +0100
                              Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-01-10 00:20 +0100
                            Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-10 00:20 +0100
                              Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-01-10 00:40 +0100
                                Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-10 00:50 +0100
                                  Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-10 01:00 +0100
                                    RE: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel "Zheng, Lv" <lv.zheng@intel.com> - 2017-01-10 01:50 +0100
                                    Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size()  and early_acpi_os_unmap_memory() from Linux kernel "Rafael J. Wysocki" <rafael@kernel.org> - 2017-01-10 02:30 +0100
                                      Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-01-10 03:30 +0100
                                        RE: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel "Zheng, Lv" <lv.zheng@intel.com> - 2017-01-10 06:50 +0100
                                          Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-01-10 07:00 +0100
                                            Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-11 10:30 +0100
                                              Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-01-11 11:00 +0100
                                                Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-11 11:10 +0100
                                                  Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-01-11 11:30 +0100
                                      Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-10 10:50 +0100
                                        Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size()  and early_acpi_os_unmap_memory() from Linux kernel "Rafael J. Wysocki" <rafael@kernel.org> - 2017-01-11 04:50 +0100
                                          Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-11 10:50 +0100
                                  Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-01-10 01:00 +0100
                                Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size()  and early_acpi_os_unmap_memory() from Linux kernel "Rafael J. Wysocki" <rafael@kernel.org> - 2017-01-10 00:50 +0100
                                  Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2017-01-10 01:30 +0100
                      RE: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel "Zheng, Lv" <lv.zheng@intel.com> - 2017-01-09 06:30 +0100
                  Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Jörg Rödel <joro@8bytes.org> - 2017-01-09 12:00 +0100
                    Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2017-01-09 23:50 +0100
                      Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Borislav Petkov <bp@alien8.de> - 2017-01-10 00:00 +0100
                      Re: 174cc7187e6f ACPICA: Tables: Back port  acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux  kernel Jörg Rödel <joro@8bytes.org> - 2017-01-10 15:00 +0100

Page 1 of 2  [1] 2  Next page →


#1553763 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

From"Rafael J. Wysocki" <rafael.j.wysocki@intel.com>
Date2017-01-08 00:40 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sX8KC-6UM-7@gated-at.bofh.it>
On 1/7/2017 9:42 PM, Borislav Petkov wrote:
> Hi,
>
> I'm bisecting a boot freeze with 4.10-rc2 on a laptop and the commit in
> $Subject is introducing a breakage, see attached splat. Unfortunately,
> it is not complete as I don't have any other means of logging dmesg on a
> laptop.
>
> A temporary workaround is to boot with "intremap=off".
>
> Unfortunately, I cannot revert
>
>    174cc7187e6f ("ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel")
>
> as it doesn't revert cleanly anymore. From looking at the callstack,
> though, it looks like AMD IOMMU is calling acpi_put_table() and patch in
> $Subject touches it so I'm guessing the AMD IOMMU needs to be updated to
> the new way of parsing the IRQ remapping tables or whatever is going on
> there - I'm just guessing.

Hi,

Please check if this helps:

https://git.kernel.org/cgit/linux/kernel/git/torvalds/linux.git/commit/?id=696c7f8e0373026e8bfb29b2d9ff2d0a92059d4d

Thanks,
Rafael


> Thanks.
>

[toc] | [next] | [standalone]


#1553770 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

FromBorislav Petkov <bp@alien8.de>
Date2017-01-08 01:10 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sX9dD-7kg-3@gated-at.bofh.it>
In reply to#1553763
On Sun, Jan 08, 2017 at 12:30:27AM +0100, Rafael J. Wysocki wrote:
> Please check if this helps:
> 
> https://git.kernel.org/cgit/linux/kernel/git/torvalds/linux.git/commit/?id=696c7f8e0373026e8bfb29b2d9ff2d0a92059d4d

Unfortunately no, still same early freeze. :-\

The splat happens when booting 6b11d1d67713 directly, 4.10-rc2 doesn't
splat but freezes simply very early during boot.

-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1553771

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2017-01-08 01:30 +0100
Message-ID<sX9wZ-7sS-5@gated-at.bofh.it>
In reply to#1553770
On Sun, Jan 8, 2017 at 1:07 AM, Borislav Petkov <bp@alien8.de> wrote:
> On Sun, Jan 08, 2017 at 12:30:27AM +0100, Rafael J. Wysocki wrote:
>> Please check if this helps:
>>
>> https://git.kernel.org/cgit/linux/kernel/git/torvalds/linux.git/commit/?id=696c7f8e0373026e8bfb29b2d9ff2d0a92059d4d
>
> Unfortunately no, still same early freeze. :-\
>
> The splat happens when booting 6b11d1d67713 directly, 4.10-rc2 doesn't
> splat but freezes simply very early during boot.

Is an IVRS table actually present on this machine?

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1553773 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

FromBorislav Petkov <bp@alien8.de>
Date2017-01-08 01:40 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sX9GF-7vM-3@gated-at.bofh.it>
In reply to#1553771
On Sun, Jan 08, 2017 at 01:22:55AM +0100, Rafael J. Wysocki wrote:
> Is an IVRS table actually present on this machine?

Like this?

[    0.000000] ACPI: IVRS 0x000000009CFD6000 0000D0 (v02 AMD    AGESA    00000001 AMD  00000000)

-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1553774

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2017-01-08 02:00 +0100
Message-ID<sXa01-7Ct-7@gated-at.bofh.it>
In reply to#1553773

[Multipart message — attachments visible in raw view] — view raw

On Sun, Jan 8, 2017 at 1:37 AM, Borislav Petkov <bp@alien8.de> wrote:
> On Sun, Jan 08, 2017 at 01:22:55AM +0100, Rafael J. Wysocki wrote:
>> Is an IVRS table actually present on this machine?
>
> Like this?
>
> [    0.000000] ACPI: IVRS 0x000000009CFD6000 0000D0 (v02 AMD    AGESA    00000001 AMD  00000000)

Yup.

So we get the table, but apparently we crash when we attempt to put it.

Let's try to check the obvious just to rule it out (see attached), but
honestly I'm not sure what's going on in there.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1553775 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

FromBorislav Petkov <bp@alien8.de>
Date2017-01-08 02:10 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXa9H-7Vs-1@gated-at.bofh.it>
In reply to#1553774
On Sun, Jan 08, 2017 at 01:52:50AM +0100, Rafael J. Wysocki wrote:
> So we get the table, but apparently we crash when we attempt to put it.

Right, except on 4.10-rc2 we don't crash but we freeze early. These are
the last lines:

...
[    0.004778] mce: CPU supports 7 MCE banks
[    0.004861] LVT offset 1 assigned for vector 0xf9
[    0.004945] Last level iTLB entries: 4KB 512, 2MB 1024, 4MB 512
[    0.005025] Last level dTLB entries: 4KB 1024, 2MB 1024, 4MB 512, 1GB 0
[    0.005165] Freeing SMP alternatives memory: 24K
[    0.211154] ftrace: allocating 25022 entries in 98 pages
[    0.219614] smpboot: Max logical packages: 2
<EOF>

> Let's try to check the obvious just to rule it out (see attached), but
> honestly I'm not sure what's going on in there.

No change, same freeze.

-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1553779

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2017-01-08 02:50 +0100
Message-ID<sXaMp-8bf-3@gated-at.bofh.it>
In reply to#1553775
On Sun, Jan 8, 2017 at 2:01 AM, Borislav Petkov <bp@alien8.de> wrote:
> On Sun, Jan 08, 2017 at 01:52:50AM +0100, Rafael J. Wysocki wrote:
>> So we get the table, but apparently we crash when we attempt to put it.
>
> Right, except on 4.10-rc2 we don't crash but we freeze early. These are
> the last lines:
>
> ...
> [    0.004778] mce: CPU supports 7 MCE banks
> [    0.004861] LVT offset 1 assigned for vector 0xf9
> [    0.004945] Last level iTLB entries: 4KB 512, 2MB 1024, 4MB 512
> [    0.005025] Last level dTLB entries: 4KB 1024, 2MB 1024, 4MB 512, 1GB 0
> [    0.005165] Freeing SMP alternatives memory: 24K
> [    0.211154] ftrace: allocating 25022 entries in 98 pages
> [    0.219614] smpboot: Max logical packages: 2
> <EOF>
>
>> Let's try to check the obvious just to rule it out (see attached), but
>> honestly I'm not sure what's going on in there.
>
> No change, same freeze.

I was afraid that that would be the case.

Can you try to comment out the acpi_put_table() in
early_amd_iommu_init() and see if that makes any difference?

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1553785

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2017-01-08 03:30 +0100
Message-ID<sXbp7-cV-1@gated-at.bofh.it>
In reply to#1553779

[Multipart message — attachments visible in raw view] — view raw

On Sun, Jan 8, 2017 at 2:45 AM, Rafael J. Wysocki <rafael@kernel.org> wrote:
> On Sun, Jan 8, 2017 at 2:01 AM, Borislav Petkov <bp@alien8.de> wrote:
>> On Sun, Jan 08, 2017 at 01:52:50AM +0100, Rafael J. Wysocki wrote:
>>> So we get the table, but apparently we crash when we attempt to put it.
>>
>> Right, except on 4.10-rc2 we don't crash but we freeze early. These are
>> the last lines:
>>
>> ...
>> [    0.004778] mce: CPU supports 7 MCE banks
>> [    0.004861] LVT offset 1 assigned for vector 0xf9
>> [    0.004945] Last level iTLB entries: 4KB 512, 2MB 1024, 4MB 512
>> [    0.005025] Last level dTLB entries: 4KB 1024, 2MB 1024, 4MB 512, 1GB 0
>> [    0.005165] Freeing SMP alternatives memory: 24K
>> [    0.211154] ftrace: allocating 25022 entries in 98 pages
>> [    0.219614] smpboot: Max logical packages: 2
>> <EOF>
>>
>>> Let's try to check the obvious just to rule it out (see attached), but
>>> honestly I'm not sure what's going on in there.
>>
>> No change, same freeze.
>
> I was afraid that that would be the case.
>
> Can you try to comment out the acpi_put_table() in
> early_amd_iommu_init() and see if that makes any difference?

Well, there is a bug in early_amd_iommu_init() that may matter in
theory if the table checksum is incorrect.

Please see if the attached makes any difference.

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1553869 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

FromBorislav Petkov <bp@alien8.de>
Date2017-01-08 14:10 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXlot-6N7-9@gated-at.bofh.it>
In reply to#1553785
On Sun, Jan 08, 2017 at 03:20:20AM +0100, Rafael J. Wysocki wrote:
>  drivers/iommu/amd_iommu_init.c |    2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
> 
> Index: linux-pm/drivers/iommu/amd_iommu_init.c
> ===================================================================
> --- linux-pm.orig/drivers/iommu/amd_iommu_init.c
> +++ linux-pm/drivers/iommu/amd_iommu_init.c
> @@ -2230,7 +2230,7 @@ static int __init early_amd_iommu_init(v
>  	 */
>  	ret = check_ivrs_checksum(ivrs_base);
>  	if (ret)
> -		return ret;
> +		goto out;
>  
>  	amd_iommu_target_ivhd_type = get_highest_supported_ivhd_type(ivrs_base);
>  	DUMP_printk("Using IVHD type %#x\n", amd_iommu_target_ivhd_type);

Good catch, this one needs to be applied regardless.

However, it doesn't fix my issue though.

But I think I have it - I went and applied the well-proven debugging
technique of sprinkling printks around. Here's what I'm seeing:

early_amd_iommu_init()
|-> acpi_put_table(ivrs_base);
|-> acpi_tb_put_table(table_desc);
|-> acpi_tb_invalidate_table(table_desc);
|-> acpi_tb_release_table(...)
|-> acpi_os_unmap_memory
|-> acpi_os_unmap_iomem
|-> acpi_os_map_cleanup
|-> synchronize_rcu_expedited	<-- the kernel/rcu/tree_exp.h version with CONFIG_PREEMPT_RCU=y

Now that function goes and sends IPIs, i.e., schedule_work()
but this is too early - we haven't even done workqueue_init().
Actually, from looking at the callstack, we do
kernel_init_freeable->native_smp_prepare_cpus() and workqueue_init()
comes next.

And this makes sense because the splat rIP points to __queue_work() but
we haven't done that yet.

So that acpi_put_table() is happening too early. Looks like AMD IOMMU
should not put the table but WTH do I know?!

In any case, commenting out:

        acpi_put_table(ivrs_base);
        ivrs_base = NULL;

and the end of early_amd_iommu_init() makes the box boot again.

-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1554008 — RE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

From"Zheng, Lv" <lv.zheng@intel.com>
Date2017-01-09 03:00 +0100
SubjectRE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXxpD-5ZY-5@gated-at.bofh.it>
In reply to#1553869
Hi,

> From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of Borislav
> Petkov
> Subject: Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
> early_acpi_os_unmap_memory() from Linux kernel
> 
> On Sun, Jan 08, 2017 at 03:20:20AM +0100, Rafael J. Wysocki wrote:
> >  drivers/iommu/amd_iommu_init.c |    2 +-
> >  1 file changed, 1 insertion(+), 1 deletion(-)
> >
> > Index: linux-pm/drivers/iommu/amd_iommu_init.c
> > ===================================================================
> > --- linux-pm.orig/drivers/iommu/amd_iommu_init.c
> > +++ linux-pm/drivers/iommu/amd_iommu_init.c
> > @@ -2230,7 +2230,7 @@ static int __init early_amd_iommu_init(v
> >  	 */
> >  	ret = check_ivrs_checksum(ivrs_base);
> >  	if (ret)
> > -		return ret;
> > +		goto out;
> >
> >  	amd_iommu_target_ivhd_type = get_highest_supported_ivhd_type(ivrs_base);
> >  	DUMP_printk("Using IVHD type %#x\n", amd_iommu_target_ivhd_type);
> 
> Good catch, this one needs to be applied regardless.
> 
> However, it doesn't fix my issue though.
> 
> But I think I have it - I went and applied the well-proven debugging
> technique of sprinkling printks around. Here's what I'm seeing:
> 
> early_amd_iommu_init()
> |-> acpi_put_table(ivrs_base);
> |-> acpi_tb_put_table(table_desc);
> |-> acpi_tb_invalidate_table(table_desc);
> |-> acpi_tb_release_table(...)
> |-> acpi_os_unmap_memory
> |-> acpi_os_unmap_iomem
> |-> acpi_os_map_cleanup
> |-> synchronize_rcu_expedited	<-- the kernel/rcu/tree_exp.h version with CONFIG_PREEMPT_RCU=y
> 
> Now that function goes and sends IPIs, i.e., schedule_work()
> but this is too early - we haven't even done workqueue_init().
> Actually, from looking at the callstack, we do
> kernel_init_freeable->native_smp_prepare_cpus() and workqueue_init()
> comes next.
> 
> And this makes sense because the splat rIP points to __queue_work() but
> we haven't done that yet.
> 
> So that acpi_put_table() is happening too early. Looks like AMD IOMMU
> should not put the table but WTH do I know?!
> 
> In any case, commenting out:
> 
>         acpi_put_table(ivrs_base);
>         ivrs_base = NULL;
> 
> and the end of early_amd_iommu_init() makes the box boot again.

So please help to comment out these 2 lines (with descriptions and do not delete them).
Until acpi_os_unmap_memory() is able to handle such an early case.

Thanks and best regards
Lv

> 
> --
> Regards/Gruss,
>     Boris.
> 
> Good mailing practices for 400: avoid top-posting and trim the reply.
> --
> To unsubscribe from this list: send the line "unsubscribe linux-acpi" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html

[toc] | [prev] | [next] | [standalone]


#1554019 — RE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

From"Zheng, Lv" <lv.zheng@intel.com>
Date2017-01-09 03:40 +0100
SubjectRE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXy2m-6t3-15@gated-at.bofh.it>
In reply to#1554008
Hi,

> From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of Zheng,
> Lv
> Subject: RE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
> early_acpi_os_unmap_memory() from Linux kernel
> 
> Hi,
> 
> > From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of
> Borislav
> > Petkov
> > Subject: Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
> > early_acpi_os_unmap_memory() from Linux kernel
> >
> > On Sun, Jan 08, 2017 at 03:20:20AM +0100, Rafael J. Wysocki wrote:
> > >  drivers/iommu/amd_iommu_init.c |    2 +-
> > >  1 file changed, 1 insertion(+), 1 deletion(-)
> > >
> > > Index: linux-pm/drivers/iommu/amd_iommu_init.c
> > > ===================================================================
> > > --- linux-pm.orig/drivers/iommu/amd_iommu_init.c
> > > +++ linux-pm/drivers/iommu/amd_iommu_init.c
> > > @@ -2230,7 +2230,7 @@ static int __init early_amd_iommu_init(v
> > >  	 */
> > >  	ret = check_ivrs_checksum(ivrs_base);
> > >  	if (ret)
> > > -		return ret;
> > > +		goto out;
> > >
> > >  	amd_iommu_target_ivhd_type = get_highest_supported_ivhd_type(ivrs_base);
> > >  	DUMP_printk("Using IVHD type %#x\n", amd_iommu_target_ivhd_type);
> >
> > Good catch, this one needs to be applied regardless.
> >
> > However, it doesn't fix my issue though.
> >
> > But I think I have it - I went and applied the well-proven debugging
> > technique of sprinkling printks around. Here's what I'm seeing:
> >
> > early_amd_iommu_init()
> > |-> acpi_put_table(ivrs_base);
> > |-> acpi_tb_put_table(table_desc);
> > |-> acpi_tb_invalidate_table(table_desc);
> > |-> acpi_tb_release_table(...)
> > |-> acpi_os_unmap_memory
> > |-> acpi_os_unmap_iomem
> > |-> acpi_os_map_cleanup
> > |-> synchronize_rcu_expedited	<-- the kernel/rcu/tree_exp.h version with CONFIG_PREEMPT_RCU=y
> >
> > Now that function goes and sends IPIs, i.e., schedule_work()
> > but this is too early - we haven't even done workqueue_init().
> > Actually, from looking at the callstack, we do
> > kernel_init_freeable->native_smp_prepare_cpus() and workqueue_init()
> > comes next.
> >
> > And this makes sense because the splat rIP points to __queue_work() but
> > we haven't done that yet.
> >
> > So that acpi_put_table() is happening too early. Looks like AMD IOMMU
> > should not put the table but WTH do I know?!
> >
> > In any case, commenting out:
> >
> >         acpi_put_table(ivrs_base);
> >         ivrs_base = NULL;
> >
> > and the end of early_amd_iommu_init() makes the box boot again.
> 
> So please help to comment out these 2 lines (with descriptions and do not delete them).
> Until acpi_os_unmap_memory() is able to handle such an early case.

IMO, synchronize_rcu_expedited() should be improved:
If rcu_init() isn't called or there is nothing to synchronize, schedule_work() shouldn't be invoked.

Thanks and best regards
Lv

> 
> Thanks and best regards
> Lv
> 
> >
> > --
> > Regards/Gruss,
> >     Boris.
> >
> > Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1554169 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

FromBorislav Petkov <bp@alien8.de>
Date2017-01-09 10:40 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXEAO-2pA-11@gated-at.bofh.it>
In reply to#1554019
+ Paul for comment.

Leaving in the rest for him.

On Mon, Jan 09, 2017 at 02:36:33AM +0000, Zheng, Lv wrote:
> Hi,
> 
> > From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of Zheng,
> > Lv
> > Subject: RE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
> > early_acpi_os_unmap_memory() from Linux kernel
> > 
> > Hi,
> > 
> > > From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of
> > Borislav
> > > Petkov
> > > Subject: Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
> > > early_acpi_os_unmap_memory() from Linux kernel
> > >
> > > On Sun, Jan 08, 2017 at 03:20:20AM +0100, Rafael J. Wysocki wrote:
> > > >  drivers/iommu/amd_iommu_init.c |    2 +-
> > > >  1 file changed, 1 insertion(+), 1 deletion(-)
> > > >
> > > > Index: linux-pm/drivers/iommu/amd_iommu_init.c
> > > > ===================================================================
> > > > --- linux-pm.orig/drivers/iommu/amd_iommu_init.c
> > > > +++ linux-pm/drivers/iommu/amd_iommu_init.c
> > > > @@ -2230,7 +2230,7 @@ static int __init early_amd_iommu_init(v
> > > >  	 */
> > > >  	ret = check_ivrs_checksum(ivrs_base);
> > > >  	if (ret)
> > > > -		return ret;
> > > > +		goto out;
> > > >
> > > >  	amd_iommu_target_ivhd_type = get_highest_supported_ivhd_type(ivrs_base);
> > > >  	DUMP_printk("Using IVHD type %#x\n", amd_iommu_target_ivhd_type);
> > >
> > > Good catch, this one needs to be applied regardless.
> > >
> > > However, it doesn't fix my issue though.
> > >
> > > But I think I have it - I went and applied the well-proven debugging
> > > technique of sprinkling printks around. Here's what I'm seeing:
> > >
> > > early_amd_iommu_init()
> > > |-> acpi_put_table(ivrs_base);
> > > |-> acpi_tb_put_table(table_desc);
> > > |-> acpi_tb_invalidate_table(table_desc);
> > > |-> acpi_tb_release_table(...)
> > > |-> acpi_os_unmap_memory
> > > |-> acpi_os_unmap_iomem
> > > |-> acpi_os_map_cleanup
> > > |-> synchronize_rcu_expedited	<-- the kernel/rcu/tree_exp.h version with CONFIG_PREEMPT_RCU=y
> > >
> > > Now that function goes and sends IPIs, i.e., schedule_work()
> > > but this is too early - we haven't even done workqueue_init().
> > > Actually, from looking at the callstack, we do
> > > kernel_init_freeable->native_smp_prepare_cpus() and workqueue_init()
> > > comes next.
> > >
> > > And this makes sense because the splat rIP points to __queue_work() but
> > > we haven't done that yet.
> > >
> > > So that acpi_put_table() is happening too early. Looks like AMD IOMMU
> > > should not put the table but WTH do I know?!
> > >
> > > In any case, commenting out:
> > >
> > >         acpi_put_table(ivrs_base);
> > >         ivrs_base = NULL;
> > >
> > > and the end of early_amd_iommu_init() makes the box boot again.
> > 
> > So please help to comment out these 2 lines (with descriptions and do not delete them).
> > Until acpi_os_unmap_memory() is able to handle such an early case.
> 
> IMO, synchronize_rcu_expedited() should be improved:
> If rcu_init() isn't called or there is nothing to synchronize, schedule_work() shouldn't be invoked.

-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1554757 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-01-09 23:20 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXQsi-1ot-19@gated-at.bofh.it>
In reply to#1554169
On Mon, Jan 09, 2017 at 10:33:29AM +0100, Borislav Petkov wrote:
> + Paul for comment.
> 
> Leaving in the rest for him.
> 
> On Mon, Jan 09, 2017 at 02:36:33AM +0000, Zheng, Lv wrote:
> > Hi,
> > 
> > > From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of Zheng,
> > > Lv
> > > Subject: RE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
> > > early_acpi_os_unmap_memory() from Linux kernel
> > > 
> > > Hi,
> > > 
> > > > From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of
> > > Borislav
> > > > Petkov
> > > > Subject: Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
> > > > early_acpi_os_unmap_memory() from Linux kernel
> > > >
> > > > On Sun, Jan 08, 2017 at 03:20:20AM +0100, Rafael J. Wysocki wrote:
> > > > >  drivers/iommu/amd_iommu_init.c |    2 +-
> > > > >  1 file changed, 1 insertion(+), 1 deletion(-)
> > > > >
> > > > > Index: linux-pm/drivers/iommu/amd_iommu_init.c
> > > > > ===================================================================
> > > > > --- linux-pm.orig/drivers/iommu/amd_iommu_init.c
> > > > > +++ linux-pm/drivers/iommu/amd_iommu_init.c
> > > > > @@ -2230,7 +2230,7 @@ static int __init early_amd_iommu_init(v
> > > > >  	 */
> > > > >  	ret = check_ivrs_checksum(ivrs_base);
> > > > >  	if (ret)
> > > > > -		return ret;
> > > > > +		goto out;
> > > > >
> > > > >  	amd_iommu_target_ivhd_type = get_highest_supported_ivhd_type(ivrs_base);
> > > > >  	DUMP_printk("Using IVHD type %#x\n", amd_iommu_target_ivhd_type);
> > > >
> > > > Good catch, this one needs to be applied regardless.
> > > >
> > > > However, it doesn't fix my issue though.
> > > >
> > > > But I think I have it - I went and applied the well-proven debugging
> > > > technique of sprinkling printks around. Here's what I'm seeing:
> > > >
> > > > early_amd_iommu_init()
> > > > |-> acpi_put_table(ivrs_base);
> > > > |-> acpi_tb_put_table(table_desc);
> > > > |-> acpi_tb_invalidate_table(table_desc);
> > > > |-> acpi_tb_release_table(...)
> > > > |-> acpi_os_unmap_memory
> > > > |-> acpi_os_unmap_iomem
> > > > |-> acpi_os_map_cleanup
> > > > |-> synchronize_rcu_expedited	<-- the kernel/rcu/tree_exp.h version with CONFIG_PREEMPT_RCU=y
> > > >
> > > > Now that function goes and sends IPIs, i.e., schedule_work()
> > > > but this is too early - we haven't even done workqueue_init().
> > > > Actually, from looking at the callstack, we do
> > > > kernel_init_freeable->native_smp_prepare_cpus() and workqueue_init()
> > > > comes next.
> > > >
> > > > And this makes sense because the splat rIP points to __queue_work() but
> > > > we haven't done that yet.
> > > >
> > > > So that acpi_put_table() is happening too early. Looks like AMD IOMMU
> > > > should not put the table but WTH do I know?!
> > > >
> > > > In any case, commenting out:
> > > >
> > > >         acpi_put_table(ivrs_base);
> > > >         ivrs_base = NULL;
> > > >
> > > > and the end of early_amd_iommu_init() makes the box boot again.
> > > 
> > > So please help to comment out these 2 lines (with descriptions and do not delete them).
> > > Until acpi_os_unmap_memory() is able to handle such an early case.
> > 
> > IMO, synchronize_rcu_expedited() should be improved:
> > If rcu_init() isn't called or there is nothing to synchronize, schedule_work() shouldn't be invoked.

Indeed it should!

Does the (untested) patch below fix things for you?

If so, does this need to go into 4.10?  (My default workflow would get
it into 4.11 or 4.12, so please speak up if you need it.)

							Thanx, Paul

------------------------------------------------------------------------

commit 1b7feb708241f1662cfd529118468c9f9c0b1449
Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Date:   Mon Jan 9 14:10:50 2017 -0800

    rcu: Make synchronize_rcu_expedited() safe for early boot
    
    The synchronize_rcu_expedited() function does not check for early-boot
    use, which can result in failures if it is invoked before the scheduler
    has started.  Given that the rcupdate.rcu_expedited kernel parameter
    causes all calls to synchronize_rcu() to be directed instead to
    synchronize_rcu_expedited(), a usage restriction does not make sense.
    
    This commit therefore adds a rcu_scheduler_active check to
    synchronize_rcu_expedited(), so that it is a no-op before the scheduler
    starts.  This behavior is correct because there is only a single CPU
    running during that time.
    
    Reported-by: Lv Zheng <lv.zheng@intel.com>
    Reported-by: Borislav Petkov <bp@alien8.de>
    Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>

diff --git a/kernel/rcu/tree_exp.h b/kernel/rcu/tree_exp.h
index dfc3ba5a429e..a6c3d86480de 100644
--- a/kernel/rcu/tree_exp.h
+++ b/kernel/rcu/tree_exp.h
@@ -690,6 +690,8 @@ void synchronize_rcu_expedited(void)
 {
 	struct rcu_state *rsp = rcu_state_p;
 
+	if (!rcu_scheduler_active)
+		return;
 	_synchronize_rcu_expedited(rsp, sync_rcu_exp_handler);
 }
 EXPORT_SYMBOL_GPL(synchronize_rcu_expedited);

[toc] | [prev] | [next] | [standalone]


#1554767

From"Rafael J. Wysocki" <rafael@kernel.org>
Date2017-01-09 23:30 +0100
Message-ID<sXQBY-1rD-17@gated-at.bofh.it>
In reply to#1554757
Hi Paul,

On Mon, Jan 9, 2017 at 11:18 PM, Paul E. McKenney
<paulmck@linux.vnet.ibm.com> wrote:
> On Mon, Jan 09, 2017 at 10:33:29AM +0100, Borislav Petkov wrote:
>> + Paul for comment.
>>
>> Leaving in the rest for him.
>>
>> On Mon, Jan 09, 2017 at 02:36:33AM +0000, Zheng, Lv wrote:
>> > Hi,
>> >
>> > > From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of Zheng,
>> > > Lv
>> > > Subject: RE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
>> > > early_acpi_os_unmap_memory() from Linux kernel
>> > >
>> > > Hi,
>> > >
>> > > > From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of
>> > > Borislav
>> > > > Petkov
>> > > > Subject: Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
>> > > > early_acpi_os_unmap_memory() from Linux kernel
>> > > >
>> > > > On Sun, Jan 08, 2017 at 03:20:20AM +0100, Rafael J. Wysocki wrote:
>> > > > >  drivers/iommu/amd_iommu_init.c |    2 +-
>> > > > >  1 file changed, 1 insertion(+), 1 deletion(-)
>> > > > >
>> > > > > Index: linux-pm/drivers/iommu/amd_iommu_init.c
>> > > > > ===================================================================
>> > > > > --- linux-pm.orig/drivers/iommu/amd_iommu_init.c
>> > > > > +++ linux-pm/drivers/iommu/amd_iommu_init.c
>> > > > > @@ -2230,7 +2230,7 @@ static int __init early_amd_iommu_init(v
>> > > > >        */
>> > > > >       ret = check_ivrs_checksum(ivrs_base);
>> > > > >       if (ret)
>> > > > > -             return ret;
>> > > > > +             goto out;
>> > > > >
>> > > > >       amd_iommu_target_ivhd_type = get_highest_supported_ivhd_type(ivrs_base);
>> > > > >       DUMP_printk("Using IVHD type %#x\n", amd_iommu_target_ivhd_type);
>> > > >
>> > > > Good catch, this one needs to be applied regardless.
>> > > >
>> > > > However, it doesn't fix my issue though.
>> > > >
>> > > > But I think I have it - I went and applied the well-proven debugging
>> > > > technique of sprinkling printks around. Here's what I'm seeing:
>> > > >
>> > > > early_amd_iommu_init()
>> > > > |-> acpi_put_table(ivrs_base);
>> > > > |-> acpi_tb_put_table(table_desc);
>> > > > |-> acpi_tb_invalidate_table(table_desc);
>> > > > |-> acpi_tb_release_table(...)
>> > > > |-> acpi_os_unmap_memory
>> > > > |-> acpi_os_unmap_iomem
>> > > > |-> acpi_os_map_cleanup
>> > > > |-> synchronize_rcu_expedited   <-- the kernel/rcu/tree_exp.h version with CONFIG_PREEMPT_RCU=y
>> > > >
>> > > > Now that function goes and sends IPIs, i.e., schedule_work()
>> > > > but this is too early - we haven't even done workqueue_init().
>> > > > Actually, from looking at the callstack, we do
>> > > > kernel_init_freeable->native_smp_prepare_cpus() and workqueue_init()
>> > > > comes next.
>> > > >
>> > > > And this makes sense because the splat rIP points to __queue_work() but
>> > > > we haven't done that yet.
>> > > >
>> > > > So that acpi_put_table() is happening too early. Looks like AMD IOMMU
>> > > > should not put the table but WTH do I know?!
>> > > >
>> > > > In any case, commenting out:
>> > > >
>> > > >         acpi_put_table(ivrs_base);
>> > > >         ivrs_base = NULL;
>> > > >
>> > > > and the end of early_amd_iommu_init() makes the box boot again.
>> > >
>> > > So please help to comment out these 2 lines (with descriptions and do not delete them).
>> > > Until acpi_os_unmap_memory() is able to handle such an early case.
>> >
>> > IMO, synchronize_rcu_expedited() should be improved:
>> > If rcu_init() isn't called or there is nothing to synchronize, schedule_work() shouldn't be invoked.
>
> Indeed it should!
>
> Does the (untested) patch below fix things for you?
>
> If so, does this need to go into 4.10?  (My default workflow would get
> it into 4.11 or 4.12, so please speak up if you need it.)

Yes it should go into 4.10 (if it fixes the problem) as the reported
regression was introduced in 4.10-rc1.

>
> ------------------------------------------------------------------------
>
> commit 1b7feb708241f1662cfd529118468c9f9c0b1449
> Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> Date:   Mon Jan 9 14:10:50 2017 -0800
>
>     rcu: Make synchronize_rcu_expedited() safe for early boot
>
>     The synchronize_rcu_expedited() function does not check for early-boot
>     use, which can result in failures if it is invoked before the scheduler
>     has started.  Given that the rcupdate.rcu_expedited kernel parameter
>     causes all calls to synchronize_rcu() to be directed instead to
>     synchronize_rcu_expedited(), a usage restriction does not make sense.
>
>     This commit therefore adds a rcu_scheduler_active check to
>     synchronize_rcu_expedited(), so that it is a no-op before the scheduler
>     starts.  This behavior is correct because there is only a single CPU
>     running during that time.
>
>     Reported-by: Lv Zheng <lv.zheng@intel.com>
>     Reported-by: Borislav Petkov <bp@alien8.de>
>     Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>

When applying this, please also add:

Fixes: 174cc7187e6f ("ACPICA: Tables: Back port
acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux
kernel")
Link: https://lkml.kernel.org/r/4034dde8-ffc1-18e2-f40c-00cf37471793@intel.com

>
> diff --git a/kernel/rcu/tree_exp.h b/kernel/rcu/tree_exp.h
> index dfc3ba5a429e..a6c3d86480de 100644
> --- a/kernel/rcu/tree_exp.h
> +++ b/kernel/rcu/tree_exp.h
> @@ -690,6 +690,8 @@ void synchronize_rcu_expedited(void)
>  {
>         struct rcu_state *rsp = rcu_state_p;
>
> +       if (!rcu_scheduler_active)
> +               return;
>         _synchronize_rcu_expedited(rsp, sync_rcu_exp_handler);
>  }
>  EXPORT_SYMBOL_GPL(synchronize_rcu_expedited);

Thanks,
Rafael

[toc] | [prev] | [next] | [standalone]


#1554781 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-01-10 00:20 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXRol-1WC-1@gated-at.bofh.it>
In reply to#1554767
On Mon, Jan 09, 2017 at 11:25:39PM +0100, Rafael J. Wysocki wrote:
> Hi Paul,
> 
> On Mon, Jan 9, 2017 at 11:18 PM, Paul E. McKenney
> <paulmck@linux.vnet.ibm.com> wrote:
> > On Mon, Jan 09, 2017 at 10:33:29AM +0100, Borislav Petkov wrote:
> >> + Paul for comment.
> >>
> >> Leaving in the rest for him.
> >>
> >> On Mon, Jan 09, 2017 at 02:36:33AM +0000, Zheng, Lv wrote:
> >> > Hi,
> >> >
> >> > > From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of Zheng,
> >> > > Lv
> >> > > Subject: RE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
> >> > > early_acpi_os_unmap_memory() from Linux kernel
> >> > >
> >> > > Hi,
> >> > >
> >> > > > From: linux-acpi-owner@vger.kernel.org [mailto:linux-acpi-owner@vger.kernel.org] On Behalf Of
> >> > > Borislav
> >> > > > Petkov
> >> > > > Subject: Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
> >> > > > early_acpi_os_unmap_memory() from Linux kernel
> >> > > >
> >> > > > On Sun, Jan 08, 2017 at 03:20:20AM +0100, Rafael J. Wysocki wrote:
> >> > > > >  drivers/iommu/amd_iommu_init.c |    2 +-
> >> > > > >  1 file changed, 1 insertion(+), 1 deletion(-)
> >> > > > >
> >> > > > > Index: linux-pm/drivers/iommu/amd_iommu_init.c
> >> > > > > ===================================================================
> >> > > > > --- linux-pm.orig/drivers/iommu/amd_iommu_init.c
> >> > > > > +++ linux-pm/drivers/iommu/amd_iommu_init.c
> >> > > > > @@ -2230,7 +2230,7 @@ static int __init early_amd_iommu_init(v
> >> > > > >        */
> >> > > > >       ret = check_ivrs_checksum(ivrs_base);
> >> > > > >       if (ret)
> >> > > > > -             return ret;
> >> > > > > +             goto out;
> >> > > > >
> >> > > > >       amd_iommu_target_ivhd_type = get_highest_supported_ivhd_type(ivrs_base);
> >> > > > >       DUMP_printk("Using IVHD type %#x\n", amd_iommu_target_ivhd_type);
> >> > > >
> >> > > > Good catch, this one needs to be applied regardless.
> >> > > >
> >> > > > However, it doesn't fix my issue though.
> >> > > >
> >> > > > But I think I have it - I went and applied the well-proven debugging
> >> > > > technique of sprinkling printks around. Here's what I'm seeing:
> >> > > >
> >> > > > early_amd_iommu_init()
> >> > > > |-> acpi_put_table(ivrs_base);
> >> > > > |-> acpi_tb_put_table(table_desc);
> >> > > > |-> acpi_tb_invalidate_table(table_desc);
> >> > > > |-> acpi_tb_release_table(...)
> >> > > > |-> acpi_os_unmap_memory
> >> > > > |-> acpi_os_unmap_iomem
> >> > > > |-> acpi_os_map_cleanup
> >> > > > |-> synchronize_rcu_expedited   <-- the kernel/rcu/tree_exp.h version with CONFIG_PREEMPT_RCU=y
> >> > > >
> >> > > > Now that function goes and sends IPIs, i.e., schedule_work()
> >> > > > but this is too early - we haven't even done workqueue_init().
> >> > > > Actually, from looking at the callstack, we do
> >> > > > kernel_init_freeable->native_smp_prepare_cpus() and workqueue_init()
> >> > > > comes next.
> >> > > >
> >> > > > And this makes sense because the splat rIP points to __queue_work() but
> >> > > > we haven't done that yet.
> >> > > >
> >> > > > So that acpi_put_table() is happening too early. Looks like AMD IOMMU
> >> > > > should not put the table but WTH do I know?!
> >> > > >
> >> > > > In any case, commenting out:
> >> > > >
> >> > > >         acpi_put_table(ivrs_base);
> >> > > >         ivrs_base = NULL;
> >> > > >
> >> > > > and the end of early_amd_iommu_init() makes the box boot again.
> >> > >
> >> > > So please help to comment out these 2 lines (with descriptions and do not delete them).
> >> > > Until acpi_os_unmap_memory() is able to handle such an early case.
> >> >
> >> > IMO, synchronize_rcu_expedited() should be improved:
> >> > If rcu_init() isn't called or there is nothing to synchronize, schedule_work() shouldn't be invoked.
> >
> > Indeed it should!
> >
> > Does the (untested) patch below fix things for you?
> >
> > If so, does this need to go into 4.10?  (My default workflow would get
> > it into 4.11 or 4.12, so please speak up if you need it.)
> 
> Yes it should go into 4.10 (if it fixes the problem) as the reported
> regression was introduced in 4.10-rc1.

OK, I will test with rcutorture, but I also need a Tested-by for the
specific problem at hand.

> > ------------------------------------------------------------------------
> >
> > commit 1b7feb708241f1662cfd529118468c9f9c0b1449
> > Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> > Date:   Mon Jan 9 14:10:50 2017 -0800
> >
> >     rcu: Make synchronize_rcu_expedited() safe for early boot
> >
> >     The synchronize_rcu_expedited() function does not check for early-boot
> >     use, which can result in failures if it is invoked before the scheduler
> >     has started.  Given that the rcupdate.rcu_expedited kernel parameter
> >     causes all calls to synchronize_rcu() to be directed instead to
> >     synchronize_rcu_expedited(), a usage restriction does not make sense.
> >
> >     This commit therefore adds a rcu_scheduler_active check to
> >     synchronize_rcu_expedited(), so that it is a no-op before the scheduler
> >     starts.  This behavior is correct because there is only a single CPU
> >     running during that time.
> >
> >     Reported-by: Lv Zheng <lv.zheng@intel.com>
> >     Reported-by: Borislav Petkov <bp@alien8.de>
> >     Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
> 
> When applying this, please also add:
> 
> Fixes: 174cc7187e6f ("ACPICA: Tables: Back port
> acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux
> kernel")
> Link: https://lkml.kernel.org/r/4034dde8-ffc1-18e2-f40c-00cf37471793@intel.com

Like this?

							Thanx, Paul

------------------------------------------------------------------------

commit 9ae87e58f7c40b39ff80737e3375672f16316d23
Author: Paul E. McKenney <paulmck@linux.vnet.ibm.com>
Date:   Mon Jan 9 14:10:50 2017 -0800

    rcu: Make synchronize_rcu_expedited() safe for early boot
    
    The synchronize_rcu_expedited() function does not check for early-boot
    use, which can result in failures if it is invoked before the scheduler
    has started.  Given that the rcupdate.rcu_expedited kernel parameter
    causes all calls to synchronize_rcu() to be directed instead to
    synchronize_rcu_expedited(), a usage restriction does not make sense.
    
    This commit therefore adds a rcu_scheduler_active check to
    synchronize_rcu_expedited(), so that it is a no-op before the scheduler
    starts.  This behavior is correct because there is only a single CPU
    running during that time.
    
    Reported-by: Lv Zheng <lv.zheng@intel.com>
    Reported-by: Borislav Petkov <bp@alien8.de>
    Fixes: 174cc7187e6f ("ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel")
    Link: Link: https://lkml.kernel.org/r/4034dde8-ffc1-18e2-f40c-00cf37471793@intel.com
    Signed-off-by: Paul E. McKenney <paulmck@linux.vnet.ibm.com>

diff --git a/kernel/rcu/tree_exp.h b/kernel/rcu/tree_exp.h
index dfc3ba5a429e..a6c3d86480de 100644
--- a/kernel/rcu/tree_exp.h
+++ b/kernel/rcu/tree_exp.h
@@ -690,6 +690,8 @@ void synchronize_rcu_expedited(void)
 {
 	struct rcu_state *rsp = rcu_state_p;
 
+	if (!rcu_scheduler_active)
+		return;
 	_synchronize_rcu_expedited(rsp, sync_rcu_exp_handler);
 }
 EXPORT_SYMBOL_GPL(synchronize_rcu_expedited);

[toc] | [prev] | [next] | [standalone]


#1554786 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

FromBorislav Petkov <bp@alien8.de>
Date2017-01-10 00:20 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXRom-1WC-19@gated-at.bofh.it>
In reply to#1554757
On Mon, Jan 09, 2017 at 02:18:31PM -0800, Paul E. McKenney wrote:
> @@ -690,6 +690,8 @@ void synchronize_rcu_expedited(void)
>  {
>  	struct rcu_state *rsp = rcu_state_p;
>  
> +	if (!rcu_scheduler_active)
> +		return;
>  	_synchronize_rcu_expedited(rsp, sync_rcu_exp_handler);
>  }
>  EXPORT_SYMBOL_GPL(synchronize_rcu_expedited);

That doesn't work and it is because of those damn what goes before what
boot sequence issues :-\

We have:

rest_init()
|-> rcu_scheduler_starting()  ---> that sets rcu_scheduler_active = 1;
|-> kernel_thread(kernel_init, NULL, CLONE_FS);
|-> kernel_init()
|-> kernel_init_freeable()
|-> native_smp_prepare_cpus(setup_max_cpus)
|-> default_setup_apic_routing
|-> enable_IR_x2apic
|-> irq_remapping_prepare
|-> amd_iommu_prepare
|-> iommu_go_to_state
|-> acpi_put_table(ivrs_base);
|-> acpi_tb_put_table(table_desc);
|-> acpi_tb_invalidate_table(table_desc);
|-> acpi_tb_release_table(...)
|-> acpi_os_unmap_memory
|-> acpi_os_unmap_iomem
|-> acpi_os_map_cleanup
|-> synchronize_rcu_expedited()

Now here we have rcu_scheduler_active already set so the test doesn't
hit and we hang.

So we must do it differently.

-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1554800 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

From"Paul E. McKenney" <paulmck@linux.vnet.ibm.com>
Date2017-01-10 00:40 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXRHH-231-9@gated-at.bofh.it>
In reply to#1554786
On Tue, Jan 10, 2017 at 12:15:01AM +0100, Borislav Petkov wrote:
> On Mon, Jan 09, 2017 at 02:18:31PM -0800, Paul E. McKenney wrote:
> > @@ -690,6 +690,8 @@ void synchronize_rcu_expedited(void)
> >  {
> >  	struct rcu_state *rsp = rcu_state_p;
> >  
> > +	if (!rcu_scheduler_active)
> > +		return;
> >  	_synchronize_rcu_expedited(rsp, sync_rcu_exp_handler);
> >  }
> >  EXPORT_SYMBOL_GPL(synchronize_rcu_expedited);
> 
> That doesn't work and it is because of those damn what goes before what
> boot sequence issues :-\
> 
> We have:
> 
> rest_init()
> |-> rcu_scheduler_starting()  ---> that sets rcu_scheduler_active = 1;
> |-> kernel_thread(kernel_init, NULL, CLONE_FS);
> |-> kernel_init()
> |-> kernel_init_freeable()
> |-> native_smp_prepare_cpus(setup_max_cpus)
> |-> default_setup_apic_routing
> |-> enable_IR_x2apic
> |-> irq_remapping_prepare
> |-> amd_iommu_prepare
> |-> iommu_go_to_state
> |-> acpi_put_table(ivrs_base);
> |-> acpi_tb_put_table(table_desc);
> |-> acpi_tb_invalidate_table(table_desc);
> |-> acpi_tb_release_table(...)
> |-> acpi_os_unmap_memory
> |-> acpi_os_unmap_iomem
> |-> acpi_os_map_cleanup
> |-> synchronize_rcu_expedited()
> 
> Now here we have rcu_scheduler_active already set so the test doesn't
> hit and we hang.
> 
> So we must do it differently.

Yeah, there is a window just as the scheduler is starting where things don't
work.

We could move rcu_scheduler_starting() later, as long as there
is no chance of preemption or context switch before it is invoked.
Would that help in this case, or are we already context switching before
acpi_os_map_cleanup() is invoked?  (If we are already context switching,
short-circuiting synchronize_rcu_expedited() would be a bug.)

							Thanx, Paul

[toc] | [prev] | [next] | [standalone]


#1554806 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

FromBorislav Petkov <bp@alien8.de>
Date2017-01-10 00:50 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXRRn-26v-9@gated-at.bofh.it>
In reply to#1554800
On Mon, Jan 09, 2017 at 03:32:04PM -0800, Paul E. McKenney wrote:
> We could move rcu_scheduler_starting() later, as long as there
> is no chance of preemption or context switch before it is invoked.
> Would that help in this case, or are we already context switching before
> acpi_os_map_cleanup() is invoked?  (If we are already context switching,
> short-circuiting synchronize_rcu_expedited() would be a bug.)

Hmm, how about the below?

It would still happen before

        /*
         * The boot idle thread must execute schedule()
         * at least once to get things moving:
         */
        init_idle_bootup_task(current);
        schedule_preempt_disabled();

in rest_init() and right after native_smp_prepare_cpus() which is where
we're splatting.

Lemme run it.

Even if it works, we would have to stress-test this seriously...

---
diff --git a/init/main.c b/init/main.c
index b0c9d6facef9..9be221cc87c3 100644
--- a/init/main.c
+++ b/init/main.c
@@ -385,7 +385,6 @@ static noinline void __ref rest_init(void)
 {
 	int pid;
 
-	rcu_scheduler_starting();
 	/*
 	 * We need to spawn init first so that it obtains pid 1, however
 	 * the init task will end up wanting to create kthreads, which, if
@@ -1019,6 +1018,8 @@ static noinline void __init kernel_init_freeable(void)
 
 	smp_prepare_cpus(setup_max_cpus);
 
+	rcu_scheduler_starting();
+
 	workqueue_init();
 
 	do_pre_smp_initcalls();


-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1554810 — Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

FromBorislav Petkov <bp@alien8.de>
Date2017-01-10 01:00 +0100
SubjectRe: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXS14-29K-11@gated-at.bofh.it>
In reply to#1554806
On Tue, Jan 10, 2017 at 12:40:39AM +0100, Borislav Petkov wrote:
> Lemme run it.

Well, it boots but I get:

[    0.291447] ------------[ cut here ]------------
[    0.291702] WARNING: CPU: 0 PID: 1 at kernel/rcu/tree.c:3993 rcu_scheduler_starting+0x5c/0x70
[    0.292107] Modules linked in:
[    0.292277] CPU: 0 PID: 1 Comm: swapper/0 Not tainted 4.10.0-rc3+ #21
[    0.292540] Hardware name: HP HP EliteBook 745 G3/807E, BIOS N73 Ver. 01.08 01/28/2016
[    0.292893] Call Trace:
[    0.293072]  ? dump_stack+0x46/0x63
[    0.293285]  ? __warn+0xec/0x110
[    0.293487]  ? rcu_scheduler_starting+0x5c/0x70
[    0.293735]  ? kernel_init_freeable+0x58/0x19a
[    0.293976]  ? rest_init+0x80/0x80
[    0.294153]  ? kernel_init+0xa/0x100
[    0.294334]  ? ret_from_fork+0x22/0x30
[    0.294525] ---[ end trace 4c0fe009ed4dc740 ]---

TBH, I like Rafael's suggestion in the other mail to stick with fixing
this in ACPI, especially this is an ACPI problem, not RCU. Well,
more or less: RCU could be taught to *not* do schedule_work() if
workqueue_init() hasn't happened yet but that's a tangential.

So, I'm going to bed. When I wake up, I want to see working fixes!

:-)))

Thanks dudes!

-- 
Regards/Gruss,
    Boris.

Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


#1554829 — RE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel

From"Zheng, Lv" <lv.zheng@intel.com>
Date2017-01-10 01:50 +0100
SubjectRE: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and early_acpi_os_unmap_memory() from Linux kernel
Message-ID<sXSNr-2EK-15@gated-at.bofh.it>
In reply to#1554810

[Multipart message — attachments visible in raw view] — view raw

Hi,

Can the attached patch makes something different?

Thanks and best regards
Lv

> From: Borislav Petkov [mailto:bp@alien8.de]
> Subject: Re: 174cc7187e6f ACPICA: Tables: Back port acpi_get_table_with_size() and
> early_acpi_os_unmap_memory() from Linux kernel
> 
> On Tue, Jan 10, 2017 at 12:40:39AM +0100, Borislav Petkov wrote:
> > Lemme run it.
> 
> Well, it boots but I get:
> 
> [    0.291447] ------------[ cut here ]------------
> [    0.291702] WARNING: CPU: 0 PID: 1 at kernel/rcu/tree.c:3993 rcu_scheduler_starting+0x5c/0x70
> [    0.292107] Modules linked in:
> [    0.292277] CPU: 0 PID: 1 Comm: swapper/0 Not tainted 4.10.0-rc3+ #21
> [    0.292540] Hardware name: HP HP EliteBook 745 G3/807E, BIOS N73 Ver. 01.08 01/28/2016
> [    0.292893] Call Trace:
> [    0.293072]  ? dump_stack+0x46/0x63
> [    0.293285]  ? __warn+0xec/0x110
> [    0.293487]  ? rcu_scheduler_starting+0x5c/0x70
> [    0.293735]  ? kernel_init_freeable+0x58/0x19a
> [    0.293976]  ? rest_init+0x80/0x80
> [    0.294153]  ? kernel_init+0xa/0x100
> [    0.294334]  ? ret_from_fork+0x22/0x30
> [    0.294525] ---[ end trace 4c0fe009ed4dc740 ]---
> 
> TBH, I like Rafael's suggestion in the other mail to stick with fixing
> this in ACPI, especially this is an ACPI problem, not RCU. Well,
> more or less: RCU could be taught to *not* do schedule_work() if
> workqueue_init() hasn't happened yet but that's a tangential.
> 
> So, I'm going to bed. When I wake up, I want to see working fixes!
> 
> :-)))
> 
> Thanks dudes!
> 
> --
> Regards/Gruss,
>     Boris.
> 
> Good mailing practices for 400: avoid top-posting and trim the reply.

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.kernel


csiph-web