Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1726476 > unrolled thread
| Started by | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| First post | 2017-09-05 09:40 +0200 |
| Last post | 2017-09-08 17:00 +0200 |
| Articles | 20 on this page of 60 — 9 participants |
Back to article view | Back to linux.kernel
Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-05 09:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Peter Zijlstra <peterz@infradead.org> - 2017-09-05 11:00 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-05 12:00 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Thomas Gleixner <tglx@linutronix.de> - 2017-09-06 15:00 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-06 15:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-07 08:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Thomas Gleixner <tglx@linutronix.de> - 2017-09-08 08:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-08 10:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-08 11:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 11:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Ingo Molnar <mingo@kernel.org> - 2017-09-08 12:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 12:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 13:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-08 18:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 19:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-08 23:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 00:00 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 01:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 01:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 02:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 03:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@amacapital.net> - 2017-09-09 03:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 20:00 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 20:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 08:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 12:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 13:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 15:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 15:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 15:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 16:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 16:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 16:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 16:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 18:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 19:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 19:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 19:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 20:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 20:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 20:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 20:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 20:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 20:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 21:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-10 06:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@amacapital.net> - 2017-09-10 22:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Peter Zijlstra <peterz@infradead.org> - 2017-09-10 22:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Ingo Molnar <mingo@kernel.org> - 2017-09-17 19:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Rik van Riel <riel@redhat.com> - 2017-09-11 03:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-11 03:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Rik van Riel <riel@redhat.com> - 2017-09-11 17:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 21:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-12 09:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 10:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-08 17:00 +0200
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-09-09 03:10 +0200 |
| Message-ID | <unCrv-6zX-1@gated-at.bofh.it> |
| In reply to | #1729358 |
On Fri, Sep 8, 2017 at 5:00 PM, Andy Lutomirski <luto@kernel.org> wrote:
>
> I'm not convinced. The SDM says (Vol 3, 11.3, under WC):
>
> If the WC buffer is partially filled, the writes may be delayed until
> the next occurrence of a serializing event; such as, an SFENCE or
> MFENCE instruction, CPUID execution, a read or write to uncached
> memory, an interrupt occurrence, or a LOCK instruction execution.
>
> Thanks, Intel, for definiing "serializing event" differently here than
> anywhere else in the whole manual.
Yeah, it's really badly defined. Ok, maybe a locked instruction does
actually wait for it.. It should be invisible to anything, regardless.
> 1. The kernel wants to reclaim a page of normal memory, so it unmaps
> it and flushes. Another CPU has an entry for that page in its WC
> buffer. I don't think we care whether the flush causes the WC write
> to really hit RAM because it's unobservable -- we just need to make
> sure it is ordered, as seen by software, before the flush operation
> completes. From the quote above, I think we're okay here.
Agreed.
> 2. The kernel is unmapping some IO memory (e.g. a GPU command buffer).
> It wants a guarantee that, when flush_tlb_mm_range returns, all CPUs
> are really done writing to it. Here I'm less convinced. The SDM
> quote certainly suggests to me that we have a promise that the WC
> write has *started* before flush_tlb_mm_range returns, but I'm not
> sure I believe that it's guaranteed to have retired.
If others have writable TLB entries, what keeps them from just
continuing to write for a long time afterwards?
> I'd prefer to leave it as is except on the buggy AMD CPUs, though,
> since the current code is nice and fast.
So is there a patch to detect the 383 erratum and serialize for those?
I may have missed that part.
Linus
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2017-09-09 03:40 +0200 |
| Message-ID | <unCUy-6Vk-7@gated-at.bofh.it> |
| In reply to | #1729362 |
> On Sep 8, 2017, at 6:05 PM, Linus Torvalds <torvalds@linux-foundation.org> wrote: > >> On Fri, Sep 8, 2017 at 5:00 PM, Andy Lutomirski <luto@kernel.org> wrote: >> >> I'm not convinced. The SDM says (Vol 3, 11.3, under WC): >> >> If the WC buffer is partially filled, the writes may be delayed until >> the next occurrence of a serializing event; such as, an SFENCE or >> MFENCE instruction, CPUID execution, a read or write to uncached >> memory, an interrupt occurrence, or a LOCK instruction execution. >> >> Thanks, Intel, for definiing "serializing event" differently here than >> anywhere else in the whole manual. > > Yeah, it's really badly defined. Ok, maybe a locked instruction does > actually wait for it.. It should be invisible to anything, regardless. > >> 1. The kernel wants to reclaim a page of normal memory, so it unmaps >> it and flushes. Another CPU has an entry for that page in its WC >> buffer. I don't think we care whether the flush causes the WC write >> to really hit RAM because it's unobservable -- we just need to make >> sure it is ordered, as seen by software, before the flush operation >> completes. From the quote above, I think we're okay here. > > Agreed. > >> 2. The kernel is unmapping some IO memory (e.g. a GPU command buffer). >> It wants a guarantee that, when flush_tlb_mm_range returns, all CPUs >> are really done writing to it. Here I'm less convinced. The SDM >> quote certainly suggests to me that we have a promise that the WC >> write has *started* before flush_tlb_mm_range returns, but I'm not >> sure I believe that it's guaranteed to have retired. > > If others have writable TLB entries, what keeps them from just > continuing to write for a long time afterwards? Whoever unmaps the resource by kicking out their drm fd? I admit I'm just trying to think of the worst case. > >> I'd prefer to leave it as is except on the buggy AMD CPUs, though, >> since the current code is nice and fast. > > So is there a patch to detect the 383 erratum and serialize for those? > I may have missed that part. > The patch is in my head. It's imaginarily attached to this email. > Linus
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@kernel.org> |
|---|---|
| Date | 2017-09-09 20:00 +0200 |
| Message-ID | <unScV-j0-3@gated-at.bofh.it> |
| In reply to | #1729366 |
On Fri, Sep 8, 2017 at 6:39 PM, Andy Lutomirski <luto@amacapital.net> wrote: > > >> On Sep 8, 2017, at 6:05 PM, Linus Torvalds <torvalds@linux-foundation.org> wrote: >> >>> On Fri, Sep 8, 2017 at 5:00 PM, Andy Lutomirski <luto@kernel.org> wrote: >>> >>> I'm not convinced. The SDM says (Vol 3, 11.3, under WC): >>> >>> If the WC buffer is partially filled, the writes may be delayed until >>> the next occurrence of a serializing event; such as, an SFENCE or >>> MFENCE instruction, CPUID execution, a read or write to uncached >>> memory, an interrupt occurrence, or a LOCK instruction execution. >>> >>> Thanks, Intel, for definiing "serializing event" differently here than >>> anywhere else in the whole manual. >> >> Yeah, it's really badly defined. Ok, maybe a locked instruction does >> actually wait for it.. It should be invisible to anything, regardless. >> >>> 1. The kernel wants to reclaim a page of normal memory, so it unmaps >>> it and flushes. Another CPU has an entry for that page in its WC >>> buffer. I don't think we care whether the flush causes the WC write >>> to really hit RAM because it's unobservable -- we just need to make >>> sure it is ordered, as seen by software, before the flush operation >>> completes. From the quote above, I think we're okay here. >> >> Agreed. >> >>> 2. The kernel is unmapping some IO memory (e.g. a GPU command buffer). >>> It wants a guarantee that, when flush_tlb_mm_range returns, all CPUs >>> are really done writing to it. Here I'm less convinced. The SDM >>> quote certainly suggests to me that we have a promise that the WC >>> write has *started* before flush_tlb_mm_range returns, but I'm not >>> sure I believe that it's guaranteed to have retired. >> >> If others have writable TLB entries, what keeps them from just >> continuing to write for a long time afterwards? > > Whoever unmaps the resource by kicking out their drm fd? I admit I'm just trying to think of the worst case. > >> >>> I'd prefer to leave it as is except on the buggy AMD CPUs, though, >>> since the current code is nice and fast. >> >> So is there a patch to detect the 383 erratum and serialize for those? >> I may have missed that part. >> > > The patch is in my head. It's imaginarily attached to this email. After contemplating the info from Boris and Markus, I think I need to add a #3 to the list of reasons my patch could be problematic: 3. If a CPU frees a page table (or PUD or PMD or whatever), that CPU will flush before the memory goes back to the system. If that flush is deferred on a different CPU that has the pointer to the freed table cached in its TLB, then that CPU can speculatively load complete garbage into its TLB. I don't think this should be observable, but I can easily imagine it triggering errata or weird ill-advised machine checks. Anyway, if I need change the behavior back, I can do it in one of two ways. I can just switch to init_mm instead of going lazy, which is expensive, but not *that* expensive on CPUs with PCID. Or I can do it the way we used to do it and send the flush IPI to lazy CPUs. The latter will only have a performance impact when a flush happens, but the performance hit is much higher when there's a flush. --Andy
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-09-09 20:10 +0200 |
| Message-ID | <unSmB-Bd-7@gated-at.bofh.it> |
| In reply to | #1729501 |
On Sat, Sep 9, 2017 at 10:49 AM, Andy Lutomirski <luto@kernel.org> wrote:
>
> Anyway, if I need change the behavior back, I can do it in one of two
> ways. I can just switch to init_mm instead of going lazy, which is
> expensive, but not *that* expensive on CPUs with PCID. Or I can do it
> the way we used to do it and send the flush IPI to lazy CPUs. The
> latter will only have a performance impact when a flush happens, but
> the performance hit is much higher when there's a flush.
Why not both?
Let's at least entertain the idea. In particular, we don't send IPI's
to *all* CPU's. We only send them to the set of CPU's that could have
that MM cached.
And that set _may_ be very limited. In the best case, it's just the
current CPU, and no IPI is needed at all.
Which means that maybe we can use that set of CPU's as guidance to how
we should treat lazy.
We can *also* take PCID support into account.
So what I would suggest is something like
- if we have PCID support, _and_ the set of CPU's is more than just
us, just switch to init_mm. The switch is cheaper than the IPI's.
- otherwise do what we used to do, with the IPI.
The exact heuristics could be tuned later, but considering Markus's
report, and considering that not so many people have really even
heavily tested the new code yet (so _one_ report now means that there
are probably a shitload of machines that would show it later), I
really think we need to steer back towards our old behavior. But at
the same time, I think we can take advantage of newer CPU's that _do_
have PCID.
Hmm?
Linus
[toc] | [prev] | [next] | [standalone]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 08:40 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unHAR-1A9-7@gated-at.bofh.it> |
| In reply to | #1729323 |
On 2017.09.08 at 23:56 +0200, Borislav Petkov wrote: > On Fri, Sep 08, 2017 at 02:47:00PM -0700, Andy Lutomirski wrote: > > Any chance you could test with CONFIG_DEBUG_VM=y? There are lots of > > potentially useful assertions in that code. > > > > Can you also post your /proc/cpuinfo? And can you re-confirm that a > > problematic guest kernel is causing problems in the *host*? > > Also, have you seen any MCEs during early boot, after the freezes? > > You probably wouldn't have because we don't log them on F10h due to > broken BIOSen. So add "mce=bootlog" to your grub and warm-reset your box > after one of those freezes and send me dmesg. It should have an MCE in > there, if it happens what I think it happens. Unfortunately the machine hangs in the BIOS after the first warm-reset. Probably when it encounters an MCE it doesn't expect. I have to warm-reset a second time to get to the boot-loader. So it is impossible for me to see any possible MCE. -- Markus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 12:20 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unL1M-43b-9@gated-at.bofh.it> |
| In reply to | #1729392 |
On Sat, Sep 09, 2017 at 08:39:08AM +0200, Markus Trippelsdorf wrote:
> Unfortunately the machine hangs in the BIOS after the first warm-reset.
> Probably when it encounters an MCE it doesn't expect. I have to
> warm-reset a second time to get to the boot-loader. So it is impossible
> for me to see any possible MCE.
Ok, let's try to disable the syncflood before the test. As root:
# V=$(setpci -s 18.3 0x44.l)
# echo $V
# V=$(printf "0x%x" $((0x$V & ~(1 << 21))))
# setpci -s 18.3 0x44.l=$V
# echo $V
I've added the echo $V so that you can paste them as a reply so that I
can see their values.
And then run the triggering sequence again, better not on an X terminal
but in the text console to see any MCEs when it freezes. I remember you
saying that you don't have serial connected to it so catching the MCE
would need more staring :)
Thx.
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 13:10 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unLO9-4CV-11@gated-at.bofh.it> |
| In reply to | #1729422 |
On 2017.09.09 at 12:18 +0200, Borislav Petkov wrote: > On Sat, Sep 09, 2017 at 08:39:08AM +0200, Markus Trippelsdorf wrote: > > Unfortunately the machine hangs in the BIOS after the first warm-reset. > > Probably when it encounters an MCE it doesn't expect. I have to > > warm-reset a second time to get to the boot-loader. So it is impossible > > for me to see any possible MCE. > > Ok, let's try to disable the syncflood before the test. As root: > > # V=$(setpci -s 18.3 0x44.l) > # echo $V 4a70005c > # V=$(printf "0x%x" $((0x$V & ~(1 << 21)))) > # setpci -s 18.3 0x44.l=$V > # echo $V 0x4a50005c > I've added the echo $V so that you can paste them as a reply so that I > can see their values. > > And then run the triggering sequence again, better not on an X terminal > but in the text console to see any MCEs when it freezes. I remember you > saying that you don't have serial connected to it so catching the MCE > would need more staring :) It doesn't work. Compiling in a text console just freezes the machine before any MCE gets printed. -- Markus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 15:10 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unNGi-5NP-23@gated-at.bofh.it> |
| In reply to | #1729427 |
On Sat, Sep 09, 2017 at 01:07:49PM +0200, Markus Trippelsdorf wrote:
> It doesn't work. Compiling in a text console just freezes the machine
> before any MCE gets printed.
Ok, let's turn off all syncflood bits. Hunk below. Do a
$ dmesg | grep syncflood
to check it worked. It says
[ 1.557017] quirk_syncflood: 0x44: 0xa900044
[ 1.561431] quirk_syncflood: 0x44: wrote 0xa800040
[ 1.566361] quirk_syncflood: 0x180: 0x700022
[ 1.570775] quirk_syncflood: 0x180: wrote 0x20
here.
Also, make sure you boot with "pci=check_enable_amd_mmconf" on the
kernel cmdline because someone broke extended PCI cfg space again on
those machines. At least on my test box here... But that's something
I'll deal with later. :-\
Thanks.
---
diff --git a/arch/x86/kernel/quirks.c b/arch/x86/kernel/quirks.c
index eaa591cfd98b..c6a4430d2222 100644
--- a/arch/x86/kernel/quirks.c
+++ b/arch/x86/kernel/quirks.c
@@ -626,6 +626,32 @@ static void amd_disable_seq_and_redirect_scrub(struct pci_dev *dev)
DECLARE_PCI_FIXUP_FINAL(PCI_VENDOR_ID_AMD, PCI_DEVICE_ID_AMD_16H_NB_F3,
amd_disable_seq_and_redirect_scrub);
+static void quirk_syncflood(struct pci_dev *misc)
+{
+ u32 val;
+
+ pci_read_config_dword(misc, 0x44, &val);
+ pr_info("%s: 0x44: 0x%x\n", __func__, val);
+
+ val &= ~(BIT(30) | BIT(21) | BIT(20) | BIT(2));
+
+ pci_write_config_dword(misc, 0x44, val);
+ pci_read_config_dword(misc, 0x44, &val);
+ pr_info("%s: 0x44: wrote 0x%x\n", __func__, val);
+
+ pci_read_config_dword(misc, 0x180, &val);
+ pr_info("%s: 0x180: 0x%x\n", __func__, val);
+
+ val &= ~(BIT(22) | BIT(21) | BIT(20) | BIT(9) | BIT(8) | BIT(7) | BIT(6) | BIT(1));
+
+ pci_write_config_dword(misc, 0x180, val);
+ pci_read_config_dword(misc, 0x180, &val);
+ pr_info("%s: 0x180: wrote 0x%x\n", __func__, val);
+}
+
+DECLARE_PCI_FIXUP_FINAL(PCI_VENDOR_ID_AMD, PCI_DEVICE_ID_AMD_10H_NB_MISC,
+ quirk_syncflood);
+
#if defined(CONFIG_X86_64) && defined(CONFIG_X86_MCE)
#include <linux/jump_label.h>
#include <asm/string_64.h>
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 15:40 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unO9k-60W-19@gated-at.bofh.it> |
| In reply to | #1729448 |
On 2017.09.09 at 15:07 +0200, Borislav Petkov wrote: > On Sat, Sep 09, 2017 at 01:07:49PM +0200, Markus Trippelsdorf wrote: > > It doesn't work. Compiling in a text console just freezes the machine > > before any MCE gets printed. > > Ok, let's turn off all syncflood bits. Hunk below. Do a > > $ dmesg | grep syncflood > > to check it worked. It says > > [ 1.557017] quirk_syncflood: 0x44: 0xa900044 > [ 1.561431] quirk_syncflood: 0x44: wrote 0xa800040 > [ 1.566361] quirk_syncflood: 0x180: 0x700022 > [ 1.570775] quirk_syncflood: 0x180: wrote 0x20 > > here. > > Also, make sure you boot with "pci=check_enable_amd_mmconf" on the > kernel cmdline because someone broke extended PCI cfg space again on > those machines. At least on my test box here... But that's something > I'll deal with later. :-\ Thanks. This one worked: mce: [Hardware Error]: CPU: 0 Machine Check Exception: 4 Bank 4: fa000010000b0c0f mce: [Hardware Error]: TSC b75d6ef4ad MISC c00a00001000000 mce: [Hardware Error]: PROCESSOR 2:100f42 TIME 1504963036 SOCKET 0 APIC 0 microcode 1000db (I had to copy the above by hand, so it may not be 100% accurate). -- Markus
[toc] | [prev] | [next] | [standalone]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 15:50 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unOj0-669-23@gated-at.bofh.it> |
| In reply to | #1729449 |
On 2017.09.09 at 15:37 +0200, Markus Trippelsdorf wrote: > On 2017.09.09 at 15:07 +0200, Borislav Petkov wrote: > > On Sat, Sep 09, 2017 at 01:07:49PM +0200, Markus Trippelsdorf wrote: > > > It doesn't work. Compiling in a text console just freezes the machine > > > before any MCE gets printed. > > > > Ok, let's turn off all syncflood bits. Hunk below. Do a > > > > $ dmesg | grep syncflood > > > > to check it worked. It says > > > > [ 1.557017] quirk_syncflood: 0x44: 0xa900044 > > [ 1.561431] quirk_syncflood: 0x44: wrote 0xa800040 > > [ 1.566361] quirk_syncflood: 0x180: 0x700022 > > [ 1.570775] quirk_syncflood: 0x180: wrote 0x20 > > > > here. > > > > Also, make sure you boot with "pci=check_enable_amd_mmconf" on the > > kernel cmdline because someone broke extended PCI cfg space again on > > those machines. At least on my test box here... But that's something > > I'll deal with later. :-\ > > Thanks. This one worked: > > mce: [Hardware Error]: CPU: 0 Machine Check Exception: 4 Bank 4: fa000010000b0c0f > mce: [Hardware Error]: TSC b75d6ef4ad MISC c00a00001000000 > mce: [Hardware Error]: PROCESSOR 2:100f42 TIME 1504963036 SOCKET 0 APIC 0 microcode 1000db Decoded: CPU: 0 Machine Check Exception: 4 Bank 4: fa000010000b0c0f Hardware event. This is not a software error. CPU 0 0 data cache TSC b75d6ef4ad TIME 1504963036 Sat Sep 9 15:17:16 2017 STATUS 0 MCGSTATUS 0 CPUID Vendor AMD Family 16 Model 4 SOCKET 0 APIC 0 microcode 1000db -- Markus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 16:10 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unOCl-6t7-5@gated-at.bofh.it> |
| In reply to | #1729453 |
On Sat, Sep 09, 2017 at 03:39:54PM +0200, Markus Trippelsdorf wrote:
> > mce: [Hardware Error]: CPU: 0 Machine Check Exception: 4 Bank 4: fa000010000b0c0f
> > mce: [Hardware Error]: TSC b75d6ef4ad MISC c00a00001000000
> > mce: [Hardware Error]: PROCESSOR 2:100f42 TIME 1504963036 SOCKET 0 APIC 0 microcode 1000db
>
> Decoded:
>
> CPU: 0 Machine Check Exception: 4 Bank 4: fa000010000b0c0f
> Hardware event. This is not a software error.
> CPU 0 0 data cache TSC b75d6ef4ad
> TIME 1504963036 Sat Sep 9 15:17:16 2017
> STATUS 0 MCGSTATUS 0
> CPUID Vendor AMD Family 16 Model 4
> SOCKET 0 APIC 0 microcode 1000db
Yeah, this is not really decoding it - I need to address that case of
uncorrectable MCE not being decoded too.
In any case, it is not E383:
MC4_STATUS[Val|Over|UC|EN|MiscV|PCC|UECC|EEC: GART cache table walk encountered an invalid PTE (0x05)|ET: TLB(tt:GEN;ll:LG)]: 0xfa0020000005001b
And those should actually be masked out:
"BIOS is recommended to mask GART table walk errors by setting the bit
in MSRC001_0048 corresponding to F3x40[GartTblWkEn]."
And we disable those but for some reason, it doesn't stick :-)
Do
# modprobe msr
# rdmsr -a 0x00000410
# rdmsr -a 0xc0010048
as root.
Thanks.
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 16:30 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unOVH-6zR-7@gated-at.bofh.it> |
| In reply to | #1729455 |
On 2017.09.09 at 16:07 +0200, Borislav Petkov wrote: > On Sat, Sep 09, 2017 at 03:39:54PM +0200, Markus Trippelsdorf wrote: > > > mce: [Hardware Error]: CPU: 0 Machine Check Exception: 4 Bank 4: fa000010000b0c0f > > > mce: [Hardware Error]: TSC b75d6ef4ad MISC c00a00001000000 > > > mce: [Hardware Error]: PROCESSOR 2:100f42 TIME 1504963036 SOCKET 0 APIC 0 microcode 1000db > > > > Decoded: > > > > CPU: 0 Machine Check Exception: 4 Bank 4: fa000010000b0c0f > > Hardware event. This is not a software error. > > CPU 0 0 data cache TSC b75d6ef4ad > > TIME 1504963036 Sat Sep 9 15:17:16 2017 > > STATUS 0 MCGSTATUS 0 > > CPUID Vendor AMD Family 16 Model 4 > > SOCKET 0 APIC 0 microcode 1000db > > Yeah, this is not really decoding it - I need to address that case of > uncorrectable MCE not being decoded too. > > In any case, it is not E383: > > MC4_STATUS[Val|Over|UC|EN|MiscV|PCC|UECC|EEC: GART cache table walk encountered an invalid PTE (0x05)|ET: TLB(tt:GEN;ll:LG)]: 0xfa0020000005001b > > And those should actually be masked out: > > "BIOS is recommended to mask GART table walk errors by setting the bit > in MSRC001_0048 corresponding to F3x40[GartTblWkEn]." > > And we disable those but for some reason, it doesn't stick :-) > > Do > > # rdmsr -a 0x00000410 3fffffff 0 0 0 > # rdmsr -a 0xc0010048 780400 780400 780400 780400 -- Markus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 16:40 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unP5n-6DD-11@gated-at.bofh.it> |
| In reply to | #1729460 |
On Sat, Sep 09, 2017 at 04:20:14PM +0200, Markus Trippelsdorf wrote:
> > # rdmsr -a 0x00000410
>
> 3fffffff
> 0
> 0
> 0
WTF?! Those should be equal on every CPU. Yikes, we need to pay
attention to those... Grrr.
# wrmsr -a 0x00000410 0x3ffffbff
should fix your issue.
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 16:50 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unPf3-6Js-1@gated-at.bofh.it> |
| In reply to | #1729462 |
On 2017.09.09 at 16:33 +0200, Borislav Petkov wrote: > On Sat, Sep 09, 2017 at 04:20:14PM +0200, Markus Trippelsdorf wrote: > > > # rdmsr -a 0x00000410 > > > > 3fffffff > > 0 > > 0 > > 0 > > WTF?! Those should be equal on every CPU. Yikes, we need to pay > attention to those... Grrr. > > # wrmsr -a 0x00000410 0x3ffffbff > > should fix your issue. No, it doesn't work: x4 ~ # rdmsr -a 0x00000410 3fffffff 0 0 0 x4 ~ # wrmsr -a 0x00000410 0x3ffffbff x4 ~ # rdmsr -a 0x00000410 3ffffbff 0 0 0 -- Markus
[toc] | [prev] | [next] | [standalone]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 18:40 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unQXw-81G-11@gated-at.bofh.it> |
| In reply to | #1729463 |
On 2017.09.09 at 16:43 +0200, Markus Trippelsdorf wrote:
> On 2017.09.09 at 16:33 +0200, Borislav Petkov wrote:
> > On Sat, Sep 09, 2017 at 04:20:14PM +0200, Markus Trippelsdorf wrote:
> > > > # rdmsr -a 0x00000410
> > >
> > > 3fffffff
> > > 0
> > > 0
> > > 0
> >
> > WTF?! Those should be equal on every CPU. Yikes, we need to pay
> > attention to those... Grrr.
> >
> > # wrmsr -a 0x00000410 0x3ffffbff
> >
> > should fix your issue.
>
> No, it doesn't work:
>
> x4 ~ # rdmsr -a 0x00000410
> 3fffffff
> 0
> 0
> 0
> x4 ~ # wrmsr -a 0x00000410 0x3ffffbff
> x4 ~ # rdmsr -a 0x00000410
> 3ffffbff
> 0
> 0
> 0
Also tried the following patch. It does not help.
diff --git a/arch/x86/kernel/cpu/mcheck/mce.c b/arch/x86/kernel/cpu/mcheck/mce.c
index 3b413065c613..9ee1edb0929f 100644
--- a/arch/x86/kernel/cpu/mcheck/mce.c
+++ b/arch/x86/kernel/cpu/mcheck/mce.c
@@ -1580,7 +1580,7 @@ static int __mcheck_cpu_apply_quirks(struct cpuinfo_x86 *c)
/* This should be disabled by the BIOS, but isn't always */
if (c->x86_vendor == X86_VENDOR_AMD) {
- if (c->x86 == 15 && cfg->banks > 4) {
+ if ((c->x86 == 15 || c->x86 == 16) && cfg->banks > 4) {
/*
* disable GART TBL walk error reporting, which
* trips off incorrectly with the IOMMU & 3ware
--
Markus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 19:10 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unRqx-8u8-5@gated-at.bofh.it> |
| In reply to | #1729481 |
On Sat, Sep 09, 2017 at 06:32:25PM +0200, Markus Trippelsdorf wrote:
> Also tried the following patch. It does not help.
Ok, another theory. This one still needs to be fixed properly but that
for later.
For some reason (insufficient coffee maybe), I have mistyped your
MCi_STATUS value earlier. Your mail says it is "fa000010000b0c0f". Do
you still have a screen photo to verify it?
Because if so, the correct error type is:
MC4_STATUS[Val|Over|UC|EN|MiscV|PCC|EEC: Protocol error (link, L3, probe filter) (0x0b)|ET: BUS(pp:OBS;t:NOTIMOUT;r4:GEN;ii:GEN;ll:LG)]: 0xfa000010000b0c0f
And for that I'd need the MC4_ADDR value too.
So can you please apply the patch below ontop of the syncflood quirk
patch and retrigger, make a photo of the MCE and send it to me?
Thanks.
---
commit e84e5ad290c7c26af69a721148f404766529509b
Author: Borislav Petkov <bp@suse.de>
Date: Sat Sep 9 00:55:50 2017 +0200
x86/MCE/AMD: Collect error info even if valid bits are not set
The MCA banks log error info into MCA_ADDR, MCA_MISC0, and MCA_SYND even
if the corresponding valid bits are not set:
"Error handlers should save the values in MCA_ADDR, MCA_MISC0,
and MCA_SYND even if MCA_STATUS[AddrV], MCA_STATUS[MiscV], and
MCA_STATUS[SyndV] are zero."
Do so by setting those bits so that code down the MCE processing path
doesn't need to be changed.
Signed-off-by: Borislav Petkov <bp@suse.de>
diff --git a/arch/x86/kernel/cpu/mcheck/mce.c b/arch/x86/kernel/cpu/mcheck/mce.c
index 3b413065c613..c63c7ef326c7 100644
--- a/arch/x86/kernel/cpu/mcheck/mce.c
+++ b/arch/x86/kernel/cpu/mcheck/mce.c
@@ -436,6 +436,20 @@ static inline void mce_gather_info(struct mce *m, struct pt_regs *regs)
if (mca_cfg.rip_msr)
m->ip = mce_rdmsrl(mca_cfg.rip_msr);
}
+
+ /*
+ * Error handlers should save the values in MCA_ADDR, MCA_MISC0, and
+ * MCA_SYND even if MCA_STATUS[AddrV], MCA_STATUS[MiscV], and
+ * MCA_STATUS[SyndV] are zero.
+ */
+ if (m->cpuvendor == X86_VENDOR_AMD) {
+ u64 status = MCI_STATUS_ADDRV | MCI_STATUS_MISCV;
+
+ if (mce_flags.smca)
+ status |= MCI_STATUS_SYNDV;
+
+ m->status |= status;
+ }
}
int mce_available(struct cpuinfo_x86 *c)
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 19:30 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unRJT-8V-1@gated-at.bofh.it> |
| In reply to | #1729492 |
On 2017.09.09 at 19:05 +0200, Borislav Petkov wrote: > On Sat, Sep 09, 2017 at 06:32:25PM +0200, Markus Trippelsdorf wrote: > > Also tried the following patch. It does not help. > > Ok, another theory. This one still needs to be fixed properly but that > for later. > > For some reason (insufficient coffee maybe), I have mistyped your > MCi_STATUS value earlier. Your mail says it is "fa000010000b0c0f". Do > you still have a screen photo to verify it? I double checked and the value is correct. > Because if so, the correct error type is: > > MC4_STATUS[Val|Over|UC|EN|MiscV|PCC|EEC: Protocol error (link, L3, probe filter) (0x0b)|ET: BUS(pp:OBS;t:NOTIMOUT;r4:GEN;ii:GEN;ll:LG)]: 0xfa000010000b0c0f > > And for that I'd need the MC4_ADDR value too. > > So can you please apply the patch below ontop of the syncflood quirk > patch and retrigger, make a photo of the MCE and send it to me? Hmm, the output is exactly the same as before your patch. -- Markus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 19:40 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unRTA-cn-29@gated-at.bofh.it> |
| In reply to | #1729497 |
On Sat, Sep 09, 2017 at 07:23:52PM +0200, Markus Trippelsdorf wrote:
> Hmm, the output is exactly the same as before your patch.
Bah, that patch doesn't account for the fact that we're rereading the
status field again in do_machine_check().
Ok, let's force MCi_ADDR out. Ontop:
---
diff --git a/arch/x86/kernel/cpu/mcheck/mce.c b/arch/x86/kernel/cpu/mcheck/mce.c
index c63c7ef326c7..e5580da2c491 100644
--- a/arch/x86/kernel/cpu/mcheck/mce.c
+++ b/arch/x86/kernel/cpu/mcheck/mce.c
@@ -240,8 +240,7 @@ static void __print_mce(struct mce *m)
}
pr_emerg(HW_ERR "TSC %llx ", m->tsc);
- if (m->addr)
- pr_cont("ADDR %llx ", m->addr);
+ pr_cont("ADDR %llx ", m->addr);
if (m->misc)
pr_cont("MISC %llx ", m->misc);
@@ -636,8 +635,9 @@ static void mce_read_aux(struct mce *m, int i)
if (m->status & MCI_STATUS_MISCV)
m->misc = mce_rdmsrl(msr_ops.misc(i));
+ m->addr = mce_rdmsrl(msr_ops.addr(i));
+
if (m->status & MCI_STATUS_ADDRV) {
- m->addr = mce_rdmsrl(msr_ops.addr(i));
/*
* Mask the reported address by the reported granularity.
---
Thanks.
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 20:20 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unSwh-F4-1@gated-at.bofh.it> |
| In reply to | #1729499 |
On 2017.09.09 at 19:36 +0200, Borislav Petkov wrote: > On Sat, Sep 09, 2017 at 07:23:52PM +0200, Markus Trippelsdorf wrote: > > Hmm, the output is exactly the same as before your patch. > > Bah, that patch doesn't account for the fact that we're rereading the > status field again in do_machine_check(). > > Ok, let's force MCi_ADDR out. Ontop: Thanks, will try it later. I think the issue gets fixed by: # wrmsr -a 0xc0010015 0x1000018 Setting bit 3 of the Hardware Configuration Register to 1. Quote for the docs: »TlbCacheDis: cacheable memory disable. Read-write. 0=Enables performance optimization that assumes PML4, PDP, PDE, and PTE entries are in cacheable WB-DRAM; memory type checks may be bypassed, and addresses outside of WB-DRAM may result in undefined behavior or NB protocol errors. 1=Disables performance optimization and allows PML4, PDP, PDE and PTE entries to be in any memory type. Operating systems that maintain page tables in memory types other than WB- DRAM must set TlbCacheDis to insure proper operation.« I've been successfully compiling for over 15 minutes now. -- Markus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 20:30 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unSFY-IW-15@gated-at.bofh.it> |
| In reply to | #1729503 |
On Sat, Sep 09, 2017 at 08:14:45PM +0200, Markus Trippelsdorf wrote:
> # wrmsr -a 0xc0010015 0x1000018
I know but I'd still like to see the exact error signature.
So please clear that bit 3 and try to catch that MCE together with the
ADDR.
Thanks.
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web