Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1591447 > unrolled thread
| Started by | Fengguang Wu <fengguang.wu@intel.com> |
|---|---|
| First post | 2017-03-02 21:30 +0100 |
| Last post | 2017-03-09 14:50 +0100 |
| Articles | 10 on this page of 30 — 8 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Fengguang Wu <fengguang.wu@intel.com> - 2017-03-02 21:30 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Daniel Borkmann <daniel@iogearbox.net> - 2017-03-02 22:40 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Linus Torvalds <torvalds@linux-foundation.org> - 2017-03-08 20:30 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Kees Cook <keescook@chromium.org> - 2017-03-08 23:40 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Daniel Borkmann <daniel@iogearbox.net> - 2017-03-09 00:20 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Laura Abbott <labbott@redhat.com> - 2017-03-09 01:30 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Kees Cook <keescook@chromium.org> - 2017-03-09 06:40 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Daniel Borkmann <daniel@iogearbox.net> - 2017-03-09 14:10 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Thomas Gleixner <tglx@linutronix.de> - 2017-03-09 14:20 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Daniel Borkmann <daniel@iogearbox.net> - 2017-03-09 15:10 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Thomas Gleixner <tglx@linutronix.de> - 2017-03-09 16:00 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Daniel Borkmann <daniel@iogearbox.net> - 2017-03-09 19:00 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf David Miller <davem@davemloft.net> - 2017-03-09 19:10 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Linus Torvalds <torvalds@linux-foundation.org> - 2017-03-09 19:20 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Linus Torvalds <torvalds@linux-foundation.org> - 2017-03-09 19:20 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Daniel Borkmann <daniel@iogearbox.net> - 2017-03-09 19:40 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Daniel Borkmann <daniel@iogearbox.net> - 2017-03-09 22:40 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Borislav Petkov <bp@suse.de> - 2017-03-09 23:10 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Daniel Borkmann <daniel@iogearbox.net> - 2017-03-09 23:20 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Borislav Petkov <bp@suse.de> - 2017-03-09 23:50 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Linus Torvalds <torvalds@linux-foundation.org> - 2017-03-10 00:30 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Borislav Petkov <bp@suse.de> - 2017-03-10 00:50 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Daniel Borkmann <daniel@iogearbox.net> - 2017-03-10 01:20 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Borislav Petkov <bp@suse.de> - 2017-03-12 22:50 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Borislav Petkov <bp@suse.de> - 2017-03-09 23:20 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Daniel Borkmann <daniel@iogearbox.net> - 2017-03-09 16:00 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Linus Torvalds <torvalds@linux-foundation.org> - 2017-03-09 18:50 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Linus Torvalds <torvalds@linux-foundation.org> - 2017-03-08 23:50 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Fengguang Wu <fengguang.wu@intel.com> - 2017-03-09 02:40 +0100
Re: [net/bpf] 3051bf36c2 BUG: unable to handle kernel paging request at 0000a7cf Thomas Gleixner <tglx@linutronix.de> - 2017-03-09 14:50 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-03-10 00:30 +0100 |
| Message-ID | <tjfFo-3B9-27@gated-at.bofh.it> |
| In reply to | #1596446 |
On Thu, Mar 9, 2017 at 2:48 PM, Borislav Petkov <bp@suse.de> wrote:
>
> I guess we could return to doing boot_cpu_has() in __flush_tlb_all()
> then. I mean, the timing-sensitivity argument is meh - killing global
> TLB entries a bit faster doesn't bring me a whole lot when I have to go
> and walk pagetable and reestablish them, which is the real price to pay
> anyway.
So should all of commit ("c109bf95992b x86/cpufeature: Remove
cpu_has_pge") just be reverted (and then marked for stable)?
Or do we have some alternate plan?
This has apparently been going on for a long while (it got merged into
4.7), but presumably it only actually _matters_ if lguest is enabled
and used and we've triggered that lguest_arch_host_init() code.
Maybe it's the lguest games with PGE that need to be removed?
Linus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@suse.de> |
|---|---|
| Date | 2017-03-10 00:50 +0100 |
| Message-ID | <tjfYJ-3JY-7@gated-at.bofh.it> |
| In reply to | #1596466 |
On Thu, Mar 09, 2017 at 03:26:02PM -0800, Linus Torvalds wrote:
> So should all of commit ("c109bf95992b x86/cpufeature: Remove
> cpu_has_pge") just be reverted (and then marked for stable)?
>
> Or do we have some alternate plan?
I think we want to do this:
diff --git a/arch/x86/include/asm/tlbflush.h b/arch/x86/include/asm/tlbflush.h
index 6fa85944af83..fc5abff9b7fd 100644
--- a/arch/x86/include/asm/tlbflush.h
+++ b/arch/x86/include/asm/tlbflush.h
@@ -188,7 +188,7 @@ static inline void __native_flush_tlb_single(unsigned long addr)
static inline void __flush_tlb_all(void)
{
- if (static_cpu_has(X86_FEATURE_PGE))
+ if (boot_cpu_has(X86_FEATURE_PGE))
__flush_tlb_global();
else
__flush_tlb();
---
but it is late here so I'd prefer to do a real patch tomorrow when I'm
not almost sleeping on the keyboard. Unless Daniel wants to write one
and test it now.
> This has apparently been going on for a long while (it got merged into
> 4.7), but presumably it only actually _matters_ if lguest is enabled
> and used and we've triggered that lguest_arch_host_init() code.
That's what I gather too, yes.
What sane code would go and clear X86_FEATURE_PGE?!? :-)))
> Maybe it's the lguest games with PGE that need to be removed?
Well, as far as I can read the comment in lguest_arch_host_init(), it
does some monkey business with switching to the guest kernel where
global pages are not present anymore... or something. So it sounds to me
like lguest would break if we removed the games but I have no idea what
it does with that.
And besides, the small hunk above restores the situation before
("c109bf95992b x86/cpufeature: Remove cpu_has_pge") so applying it would
actually be a no-brainer.
Thanks.
--
Regards/Gruss,
Boris.
SUSE Linux GmbH, GF: Felix Imendörffer, Jane Smithard, Graham Norton, HRB 21284 (AG Nürnberg)
--
[toc] | [prev] | [next] | [standalone]
| From | Daniel Borkmann <daniel@iogearbox.net> |
|---|---|
| Date | 2017-03-10 01:20 +0100 |
| Message-ID | <tjgrM-4b6-3@gated-at.bofh.it> |
| In reply to | #1596488 |
On 03/10/2017 12:44 AM, Borislav Petkov wrote:
> On Thu, Mar 09, 2017 at 03:26:02PM -0800, Linus Torvalds wrote:
>> So should all of commit ("c109bf95992b x86/cpufeature: Remove
>> cpu_has_pge") just be reverted (and then marked for stable)?
>>
>> Or do we have some alternate plan?
>
> I think we want to do this:
>
> diff --git a/arch/x86/include/asm/tlbflush.h b/arch/x86/include/asm/tlbflush.h
> index 6fa85944af83..fc5abff9b7fd 100644
> --- a/arch/x86/include/asm/tlbflush.h
> +++ b/arch/x86/include/asm/tlbflush.h
> @@ -188,7 +188,7 @@ static inline void __native_flush_tlb_single(unsigned long addr)
>
> static inline void __flush_tlb_all(void)
> {
> - if (static_cpu_has(X86_FEATURE_PGE))
> + if (boot_cpu_has(X86_FEATURE_PGE))
> __flush_tlb_global();
> else
> __flush_tlb();
> ---
>
> but it is late here so I'd prefer to do a real patch tomorrow when I'm
> not almost sleeping on the keyboard. Unless Daniel wants to write one
> and test it now.
I think we're in the same time zone. ;) I could send something
official tomorrow cooking a changelog with analysis, but I don't
mind at all if you want to go ahead with that either. Feel free
to add my SoB or Tested-by to it.
>> This has apparently been going on for a long while (it got merged into
>> 4.7), but presumably it only actually _matters_ if lguest is enabled
>> and used and we've triggered that lguest_arch_host_init() code.
>
> That's what I gather too, yes.
>
> What sane code would go and clear X86_FEATURE_PGE?!? :-)))
>
>> Maybe it's the lguest games with PGE that need to be removed?
>
> Well, as far as I can read the comment in lguest_arch_host_init(), it
> does some monkey business with switching to the guest kernel where
> global pages are not present anymore... or something. So it sounds to me
> like lguest would break if we removed the games but I have no idea what
> it does with that.
>
> And besides, the small hunk above restores the situation before
> ("c109bf95992b x86/cpufeature: Remove cpu_has_pge") so applying it would
> actually be a no-brainer.
Agree, looks only that hunk changed in behavior from c109bf95992b
("x86/cpufeature: Remove cpu_has_pge").
> Thanks.
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@suse.de> |
|---|---|
| Date | 2017-03-12 22:50 +0100 |
| Message-ID | <tkjxf-7k3-3@gated-at.bofh.it> |
| In reply to | #1596466 |
On Thu, Mar 09, 2017 at 03:26:02PM -0800, Linus Torvalds wrote:
> Maybe it's the lguest games with PGE that need to be removed?
Btw, tglx suggested something else the other day: warn when we're
changing boot_cpu_data x86_capability bits *after* alternatives have
run. The reasoning behind it being that potentially some patching
static_cpu_has() has done won't be correct anymore.
And it is pretty cheap to do it, it fires nicely on the 32-bit config
with LGUEST=y.
---
diff --git a/arch/x86/include/asm/cpufeature.h b/arch/x86/include/asm/cpufeature.h
index d59c15c3defd..f06c3dc6db70 100644
--- a/arch/x86/include/asm/cpufeature.h
+++ b/arch/x86/include/asm/cpufeature.h
@@ -124,8 +124,18 @@ extern const char * const x86_bug_flags[NBUGINTS*32];
#define boot_cpu_has(bit) cpu_has(&boot_cpu_data, bit)
-#define set_cpu_cap(c, bit) set_bit(bit, (unsigned long *)((c)->x86_capability))
-#define clear_cpu_cap(c, bit) clear_bit(bit, (unsigned long *)((c)->x86_capability))
+#define set_cpu_cap(c, bit) \
+({ \
+ WARN_ON(c == &boot_cpu_data && alternatives_patched); \
+ set_bit(bit, (unsigned long *)((c)->x86_capability)); \
+})
+
+#define clear_cpu_cap(c, bit) \
+({ \
+ WARN_ON(c == &boot_cpu_data && alternatives_patched); \
+ clear_bit(bit, (unsigned long *)((c)->x86_capability)); \
+})
+
#define setup_clear_cpu_cap(bit) do { \
clear_cpu_cap(&boot_cpu_data, bit); \
set_bit(bit, (unsigned long *)cpu_caps_cleared); \
--
Regards/Gruss,
Boris.
SUSE Linux GmbH, GF: Felix Imendörffer, Jane Smithard, Graham Norton, HRB 21284 (AG Nürnberg)
--
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@suse.de> |
|---|---|
| Date | 2017-03-09 23:20 +0100 |
| Message-ID | <tjepY-2Qt-5@gated-at.bofh.it> |
| In reply to | #1596401 |
On Thu, Mar 09, 2017 at 10:32:12PM +0100, Daniel Borkmann wrote:
> get_online_cpus();
> if (boot_cpu_has(X86_FEATURE_PGE)) { /* We have a broader idea of "global". */
> /* Remember that this was originally set (for cleanup). */
> cpu_had_pge = 1;
> /*
> * adjust_pge is a helper function which sets or unsets the PGE
> * bit on its CPU, depending on the argument (0 == unset).
> */
> on_each_cpu(adjust_pge, (void *)0, 1);
> /* Turn off the feature in the global feature set. */
> clear_cpu_cap(&boot_cpu_data, X86_FEATURE_PGE);
Can you make that:
setup_clear_cpu_cap(X86_FEATURE_PGE);
and see if it fixes your issue?
Thanks.
--
Regards/Gruss,
Boris.
SUSE Linux GmbH, GF: Felix Imendörffer, Jane Smithard, Graham Norton, HRB 21284 (AG Nürnberg)
--
[toc] | [prev] | [next] | [standalone]
| From | Daniel Borkmann <daniel@iogearbox.net> |
|---|---|
| Date | 2017-03-09 16:00 +0100 |
| Message-ID | <tj7HQ-6AN-25@gated-at.bofh.it> |
| In reply to | #1596081 |
On 03/09/2017 02:25 PM, Daniel Borkmann wrote:
> On 03/09/2017 02:10 PM, Thomas Gleixner wrote:
>> On Thu, 9 Mar 2017, Daniel Borkmann wrote:
>>> With regard to CPA_FLUSHTLB that Linus mentioned, when I investigated
>>> code paths in change_page_attr_set_clr(), I did see that CPA_FLUSHTLB
>>> was set each time we switched attrs and a cpa_flush_range() was
>>> performed (with the correct number of pages and cache set to 0). That
>>> would be a __flush_tlb_all() eventually.
>>>
>>> Hmm, it indeed might seem likely that this could be an emulation bug.
>>
>> Which variant of __flush_tlb_all() is used when the test fails?
>>
>> Check for the following flags in /proc/cpuinfo: pge invpcid
>
> I added the following and booted with both variants:
>
> printk("X86_FEATURE_PGE:%u\n", static_cpu_has(X86_FEATURE_PGE));
> printk("X86_FEATURE_INVPCID:%u\n", static_cpu_has(X86_FEATURE_INVPCID));
>
> "-cpu host" gives:
>
> [ 8.326117] X86_FEATURE_PGE:1
> [ 8.326381] X86_FEATURE_INVPCID:1
>
> "-cpu kvm64" gives:
>
> [ 8.517069] X86_FEATURE_PGE:1
> [ 8.517393] X86_FEATURE_INVPCID:0
Fwiw, I tried switching from using cr4 (__native_flush_tlb_global_irq_disabled())
to slower cr3 (__native_flush_tlb()) in "-cpu kvm64" mode, and it looks like it
also lets all test cases pass (rodata_test, test_setmem, test_bpf), no corruption
happening, etc.
Test diff used:
diff --git a/arch/x86/include/asm/tlbflush.h b/arch/x86/include/asm/tlbflush.h
index 6fa8594..34f4582 100644
--- a/arch/x86/include/asm/tlbflush.h
+++ b/arch/x86/include/asm/tlbflush.h
@@ -188,9 +188,9 @@ static inline void __native_flush_tlb_single(unsigned long addr)
static inline void __flush_tlb_all(void)
{
- if (static_cpu_has(X86_FEATURE_PGE))
- __flush_tlb_global();
- else
+// if (static_cpu_has(X86_FEATURE_PGE))
+// __flush_tlb_global();
+// else
__flush_tlb();
}
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-03-09 18:50 +0100 |
| Message-ID | <tjamm-8rP-17@gated-at.bofh.it> |
| In reply to | #1596130 |
On Thu, Mar 9, 2017 at 6:53 AM, Daniel Borkmann <daniel@iogearbox.net> wrote:
>
> Fwiw, I tried switching from using cr4
> (__native_flush_tlb_global_irq_disabled())
> to slower cr3 (__native_flush_tlb()) in "-cpu kvm64" mode, and it looks like
> it also lets all test cases pass (rodata_test, test_setmem, test_bpf), no
> corruption happening, etc.
Ok. I think this is conclusive: the qemu "-cpu kvm64" case is
definitely broken, since changing CR4.PGE is definitely
architecturally defined to flush all TLB entries.
This is not a guest kernel bug.
Of course, the bug may still be in the *host* kernel. Maybe the
emulation does something wrong. I see
if (((cr4 ^ old_cr4) & pdptr_bits) ||
(!(cr4 & X86_CR4_PCIDE) && (old_cr4 & X86_CR4_PCIDE)))
kvm_mmu_reset_context(vcpu);
(where pdptr_bits includes the PGE bit), but I'm not sure if emulation
is supposed to do something else too.
Linus
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-03-08 23:50 +0100 |
| Message-ID | <tiSz8-4xu-25@gated-at.bofh.it> |
| In reply to | #1595460 |
On Wed, Mar 8, 2017 at 2:27 PM, Daniel Borkmann <daniel@iogearbox.net> wrote:
>
> The issue seems to be accessing buff first (can be read or write access)
> and then doing set_memory_ro() doesn't make it read-only immediately,
> meaning the subsequent call into probe_kernel_write() will succeed without
> error.
>
> Then, if I don't touch buff first and only do the set_memory_ro() seems
> to work and probe_kernel_write() will then fail as expected due to pages
> being read-only now.
Ok, that definitely sounds like a TLB invalidate didn't happen.
> Now, if I access buff, do the set_memory_ro() and then a msleep(0), for
> example, it "kind of" works most of the time (see last log extract below),
> and probe_kernel_write() will fail.
Yeah, very much consistent with a missing TLB invalidate. Scheduling
will end up invalidating it, although if it's a global page even that
might not do it (but eventually the entry will just get flushed due to
other activity).
> None of this seems an issue with x86_64 and the test_setmem runs fine all
> the time, same for the actual BPF stuff.
The code does look somewhat confused about when to actually flush
things - see my earlier note about NX - but it would seem to always do
__flush_tlb_all() unless I missed something. At least as long as
CPA_FLUSHTLB is set. Maybe some case forgets to set that..
Linus
[toc] | [prev] | [next] | [standalone]
| From | Fengguang Wu <fengguang.wu@intel.com> |
|---|---|
| Date | 2017-03-09 02:40 +0100 |
| Message-ID | <tiVdE-6ny-1@gated-at.bofh.it> |
| In reply to | #1595568 |
On Wed, Mar 08, 2017 at 02:43:44PM -0800, Linus Torvalds wrote: >On Wed, Mar 8, 2017 at 2:27 PM, Daniel Borkmann <daniel@iogearbox.net> wrote: >> >> The issue seems to be accessing buff first (can be read or write access) >> and then doing set_memory_ro() doesn't make it read-only immediately, >> meaning the subsequent call into probe_kernel_write() will succeed without >> error. >> >> Then, if I don't touch buff first and only do the set_memory_ro() seems >> to work and probe_kernel_write() will then fail as expected due to pages >> being read-only now. > >Ok, that definitely sounds like a TLB invalidate didn't happen. > >> Now, if I access buff, do the set_memory_ro() and then a msleep(0), for >> example, it "kind of" works most of the time (see last log extract below), >> and probe_kernel_write() will fail. > >Yeah, very much consistent with a missing TLB invalidate. Scheduling >will end up invalidating it, although if it's a global page even that >might not do it (but eventually the entry will just get flushed due to >other activity). > >> None of this seems an issue with x86_64 and the test_setmem runs fine all >> the time, same for the actual BPF stuff. > >The code does look somewhat confused about when to actually flush >things - see my earlier note about NX - but it would seem to always do >__flush_tlb_all() unless I missed something. At least as long as >CPA_FLUSHTLB is set. Maybe some case forgets to set that.. Not sure if it's relevant, but out of 189 boots there are 2 boots showing the below "CPA: called for zero pte." warning. [ 7.116932] random: trinity: uninitialized urandom read (4 bytes read) [ 16.366468] sock: process `trinity-main' is using obsolete setsockopt SO_BSDCOMPAT [ 17.202396] BUG: unable to handle kernel paging request at 655d9eb2 [ 17.204081] IP: __release_sock+0x6e/0x100 [ 17.205207] *pde = 00000000 [ 17.205208] [ 17.206755] Oops: 0000 [#1] [ 17.207686] CPU: 0 PID: 382 Comm: trinity-main Not tainted 4.10.0-rc8-02017-g9d876e7 #1 [ 17.209819] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.9.3-20161025_171302-gandalf 04/01/2014 [ 17.212431] task: d625d200 task.stack: d6222000 [ 17.213655] EIP: __release_sock+0x6e/0x100 [ 17.214833] EFLAGS: 00010246 CPU: 0 [ 17.215951] EAX: 00000000 EBX: 655d9eb2 ECX: 00000000 EDX: 00000201 [ 17.217587] ESI: 00000605 EDI: d6064800 EBP: d6223ef4 ESP: d6223ee8 [ 17.219185] DS: 007b ES: 007b FS: 0000 GS: 0033 SS: 0068 [ 17.220602] CR0: 80050033 CR2: 655d9eb2 CR3: 1610f000 CR4: 00000610 [ 17.221966] DR0: 080cb000 DR1: 00000000 DR2: 00000000 DR3: 00000000 [ 17.223444] DR6: ffff0ff0 DR7: 00000600 [ 17.224343] Call Trace: [ 17.225007] release_sock+0x2e/0x80 [ 17.225900] sock_setsockopt+0x8c/0x880 [ 17.226857] SyS_socketcall+0x658/0x6a0 [ 17.227804] do_fast_syscall_32+0x9a/0x160 [ 17.228765] entry_SYSENTER_32+0x4c/0x7b [ 17.229694] EIP: 0xb7777cc5 [ 17.230428] EFLAGS: 00000282 CPU: 0 [ 17.231263] EAX: ffffffda EBX: 0000000e ECX: bfedce00 EDX: bfedce80 [ 17.232582] ESI: 0000001a EDI: 000000ae EBP: b754f93c ESP: bfedcdec [ 17.233882] DS: 007b ES: 007b FS: 0000 GS: 0033 SS: 007b [ 17.235044] Code: eb 29 8d 76 00 89 da 89 f8 ff 97 98 01 00 00 31 c9 ba 06 08 00 00 b8 d8 19 b1 c1 e8 ed 3d 85 ff e8 e8 62 04 00 85 f6 89 f3 74 42 <8b> 33 0f 18 06 8b 43 48 a8 01 74 0e 83 e0 fe 74 09 80 3d 3d 9c [ 17.240429] EIP: __release_sock+0x6e/0x100 SS:ESP: 0068:d6223ee8 [ 17.241689] CR2: 00000000655d9eb2 [ 17.242509] ---[ end trace dc10480164c75444 ]--- [ 17.243569] ------------[ cut here ]------------ [ 17.243574] WARNING: CPU: 0 PID: 15 at arch/x86/mm/pageattr.c:1150 __cpa_process_fault+0x388/0x390 [ 17.243575] CPA: called for zero pte. vaddr = d7ab4000 cpa->vaddr = d7ab4000 [ 17.243577] CPU: 0 PID: 15 Comm: kworker/0:1 Tainted: G D 4.10.0-rc8-02017-g9d876e7 #1 [ 17.243578] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.9.3-20161025_171302-gandalf 04/01/2014 [ 17.243582] Workqueue: events bpf_prog_free_deferred [ 17.243583] Call Trace: [ 17.243588] dump_stack+0x16/0x25 [ 17.243588] dump_stack+0x16/0x25 [ 17.243590] __warn+0xd1/0xf0 [ 17.243592] ? __cpa_process_fault+0x388/0x390 [ 17.243593] warn_slowpath_fmt+0x3b/0x40 [ 17.243594] __cpa_process_fault+0x388/0x390 [ 17.243596] ? lookup_address_in_pgd+0xa/0x90 [ 17.243598] __change_page_attr+0x520/0x6c0 [ 17.243600] ? pfn_range_is_mapped+0xe/0x80 [ 17.243601] __change_page_attr_set_clr+0x38/0x180 [ 17.243603] change_page_attr_set_clr+0x107/0x3f0 [ 17.243605] ? dequeue_entity+0x86/0x230 [ 17.243607] set_memory_rw+0x3a/0x40 [ 17.243608] bpf_prog_free_deferred+0x16/0x30 [ 17.243612] process_one_work+0xfc/0x440 [ 17.243614] ? pick_next_task_fair+0x149/0x1d0 [ 17.243615] worker_thread+0x37/0x4e0 [ 17.243617] kthread+0xdd/0x110 [ 17.243618] ? process_one_work+0x440/0x440 [ 17.243620] ? __kthread_create_on_node+0x100/0x100 [ 17.243622] ret_from_fork+0x21/0x2c [ 17.243623] ---[ end trace dc10480164c75445 ]--- [ 17.243627] BUG: unable to handle kernel NULL pointer dereference at 00000007 [ 17.243630] IP: ___cache_free+0x14/0x140 [ 17.243631] *pde = 00000000 [ 17.243631] [ 17.243633] Oops: 0000 [#2] [ 17.243635] CPU: 0 PID: 15 Comm: kworker/0:1 Tainted: G D W 4.10.0-rc8-02017-g9d876e7 #1 [ 17.243635] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.9.3-20161025_171302-gandalf 04/01/2014 [ 17.243636] Workqueue: events bpf_prog_free_deferred [ 17.243637] task: d524e8c0 task.stack: d5254000 [ 17.243639] EIP: ___cache_free+0x14/0x140 [ 17.243640] EFLAGS: 00010046 CPU: 0 [ 17.243641] EAX: d6945af8 EBX: 00000003 ECX: c10dc08e EDX: 3beb072d [ 17.243642] ESI: d6945af8 EDI: 3beb072d EBP: d5255f00 ESP: d5255ed4 [ 17.243643] DS: 007b ES: 007b FS: 0000 GS: 0000 SS: 0068 [ 17.243644] CR0: 80050033 CR2: 00000007 CR3: 1610f000 CR4: 00000610 [ 17.243647] DR0: 080cb000 DR1: 00000000 DR2: 00000000 DR3: 00000000 [ 17.243648] DR6: ffff0ff0 DR7: 00000600 [ 17.243648] Call Trace: [ 17.243650] kfree+0x64/0xe0 [ 17.243652] ? bpf_prog_free_deferred+0x1e/0x30 [ 17.243653] bpf_prog_free_deferred+0x1e/0x30 [ 17.243654] process_one_work+0xfc/0x440 [ 17.243656] ? pick_next_task_fair+0x149/0x1d0 [ 17.243658] worker_thread+0x37/0x4e0 [ 17.243659] kthread+0xdd/0x110 [ 17.243661] ? process_one_work+0x440/0x440 [ 17.243662] ? __kthread_create_on_node+0x100/0x100 [ 17.243664] ret_from_fork+0x21/0x2c [ 17.243664] Code: 89 da 89 f0 ff 0d 64 3e b4 c1 e8 f8 fe ff ff 83 c4 14 5b 5e 5d c3 90 55 89 e5 57 56 53 83 ec 20 e8 d2 21 72 00 8b 18 89 c6 89 d7 <8b> 43 04 39 03 73 65 a1 a0 fa 4b c2 85 c0 7f 1c 8b 03 8d 50 01 [ 17.243684] EIP: ___cache_free+0x14/0x140 SS:ESP: 0068:d5255ed4 [ 17.243684] CR2: 0000000000000007 [ 17.243685] ---[ end trace dc10480164c75446 ]--- [ 17.243686] Kernel panic - not syncing: Fatal exception [ 17.243687] Kernel Offset: disabled Regards, Fengguang
[toc] | [prev] | [next] | [standalone]
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-03-09 14:50 +0100 |
| Message-ID | <tj6C6-5St-7@gated-at.bofh.it> |
| In reply to | #1595460 |
On Wed, 8 Mar 2017, Linus Torvalds wrote: > Adding x86 people too, since this seems to be something off about > ARCH_HAS_SET_MEMORY for x86-32. > > The code seems to be shared between x86-32 and 64, I'm not seeing why > set_memory_r[ow]() should fail on one but not the other. Indeed. > Considering that it seems to be flaky even on 32-bit, maybe it's > timing-related, or possibly related to TLB sizes or whatever (ie more > likely hidden by a larger TLB on more modern hardware?) The only difference I can see is the way how __tlb_flush_all() is happening. We have 3 variants: invpcid_flush_all() - depends on X86_FEATURE_INVPCID and X86_FEATURE_PGE cr4 based flush - depends on X86_FEATURE_PGE cr3 based flush No idea which variant is used in that failure case. > Anyway, just looking at change_page_attr_set_clr(), I notice that the > page alias checking treats NX specially: > > /* No alias checking for _NX bit modifications */ > checkalias = (pgprot_val(mask_set) | pgprot_val(mask_clr)) != _PAGE_NX; > > which seems insane. Why would NX be different from other protection > bits (like _PAGE_RW)? The reason is that the alias mapping should never be executable at all. Thanks, tglx
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web