Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1726476 > unrolled thread
| Started by | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| First post | 2017-09-05 09:40 +0200 |
| Last post | 2017-09-08 17:00 +0200 |
| Articles | 13 on this page of 53 — 8 participants |
Back to article view | Back to linux.kernel
Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-05 09:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Peter Zijlstra <peterz@infradead.org> - 2017-09-05 11:00 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-05 12:00 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Thomas Gleixner <tglx@linutronix.de> - 2017-09-06 15:00 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-06 15:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-07 08:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Thomas Gleixner <tglx@linutronix.de> - 2017-09-08 08:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-08 10:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-08 11:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 11:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Ingo Molnar <mingo@kernel.org> - 2017-09-08 12:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 12:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 13:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-08 18:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-08 19:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-08 23:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 00:00 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 01:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 01:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 02:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 03:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@amacapital.net> - 2017-09-09 03:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 20:00 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 20:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 08:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 12:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 13:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 15:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 15:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 15:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 16:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 16:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 16:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 16:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 18:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 19:10 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 19:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 19:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 20:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 20:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 20:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 20:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 20:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 20:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-09 21:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-09 21:40 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Andy Lutomirski <luto@kernel.org> - 2017-09-10 06:50 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Linus Torvalds <torvalds@linux-foundation.org> - 2017-09-09 21:30 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Markus Trippelsdorf <markus@trippelsdorf.de> - 2017-09-09 10:20 +0200
Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf Borislav Petkov <bp@alien8.de> - 2017-09-08 17:00 +0200
Page 3 of 3 — ← Prev page 1 2 [3]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 20:50 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unSZk-R2-9@gated-at.bofh.it> |
| In reply to | #1729506 |
On 2017.09.09 at 20:26 +0200, Borislav Petkov wrote: > On Sat, Sep 09, 2017 at 08:14:45PM +0200, Markus Trippelsdorf wrote: > > # wrmsr -a 0xc0010015 0x1000018 > > I know but I'd still like to see the exact error signature. > > So please clear that bit 3 and try to catch that MCE together with the > ADDR. OK. ADDR is 12. The rest is the same (modulo time). -- Markus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 21:20 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unTsl-1hK-11@gated-at.bofh.it> |
| In reply to | #1729514 |
On Sat, Sep 09, 2017 at 09:11:33PM +0200, Borislav Petkov wrote:
> On Sat, Sep 09, 2017 at 08:46:38PM +0200, Markus Trippelsdorf wrote:
> > OK. ADDR is 12. The rest is the same (modulo time).
>
> I'm assuming that's 12 hex... yeah, "ADDR %llx ".
>
> Dammit, that should have "0x" prepended. Grrr, I'll fix all that next
> week.
Ok, that 0x12 looks like it fits the TlbCacheDis thing (bits [5:1]):
"0_1001b
Link: A specific coherent-only packet from a CPU was issued to an
IO link. This may be caused by software which addresses page table
structures in a memory type other than cacheable WB-DRAM without
properly configuring MSRC001_0015[TlbCacheDis]. This may occur, for
example, when page table structure addresses are above top of memory. In
such cases, the NB will generate an MCE if it sees a mismatch between
the memory operation generated by the core and the link type. See
2.9.3.1.2 [Determining The Access Destination for CPU Accesses]."
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 21:20 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unTsl-1hK-13@gated-at.bofh.it> |
| In reply to | #1729514 |
On Sat, Sep 09, 2017 at 08:46:38PM +0200, Markus Trippelsdorf wrote:
> OK. ADDR is 12. The rest is the same (modulo time).
I'm assuming that's 12 hex... yeah, "ADDR %llx ".
Dammit, that should have "0x" prepended. Grrr, I'll fix all that next
week.
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-09-09 20:30 +0200 |
| Message-ID | <unSFY-IW-21@gated-at.bofh.it> |
| In reply to | #1729503 |
On Sat, Sep 9, 2017 at 11:14 AM, Markus Trippelsdorf
<markus@trippelsdorf.de> wrote:
>
> I think the issue gets fixed by:
>
> # wrmsr -a 0xc0010015 0x1000018
>
> Setting bit 3 of the Hardware Configuration Register to 1.
>
> Quote for the docs:
> »TlbCacheDis: cacheable memory disable. Read-write. 0=Enables performance optimization that
> assumes PML4, PDP, PDE, and PTE entries are in cacheable WB-DRAM
Uhhuh.
The page directories should *definitely* always be in cacheable
memory, so it should be ok for that bit to be 0, and it's possible
that setting it to 1 will seriously screw up performance.
But the fact that that fixes it for you does indicate that it's not
just a stale TLB entry or something, it really is some CPU using page
tables after they have been free'd and been re-allocated to something
else (and *then* they may point to garbage).
So I do think it's a sign that we definitely need that IPI for you.
Linus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 20:40 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unSPE-Nx-5@gated-at.bofh.it> |
| In reply to | #1729507 |
On Sat, Sep 09, 2017 at 11:26:27AM -0700, Linus Torvalds wrote:
> But the fact that that fixes it for you does indicate that it's not
> just a stale TLB entry or something, it really is some CPU using page
> tables after they have been free'd and been re-allocated to something
> else (and *then* they may point to garbage).
Cool, I was trying to think of a good use case how we'd hit that. I
guess you just gave one. :)
Thx.
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-09-09 20:50 +0200 |
| Message-ID | <unSZj-R2-7@gated-at.bofh.it> |
| In reply to | #1729509 |
On Sat, Sep 9, 2017 at 11:29 AM, Borislav Petkov <bp@alien8.de> wrote:
> On Sat, Sep 09, 2017 at 11:26:27AM -0700, Linus Torvalds wrote:
>> But the fact that that fixes it for you does indicate that it's not
>> just a stale TLB entry or something, it really is some CPU using page
>> tables after they have been free'd and been re-allocated to something
>> else (and *then* they may point to garbage).
>
> Cool, I was trying to think of a good use case how we'd hit that. I
> guess you just gave one. :)
The thing is, even with the delayed TLB flushing, I don't think it
should be *so* delayed that we should be seeing a TLB fill from
garbage page tables.
But the part in Andy's patch that worries me the most is that
+ cpumask_clear_cpu(cpu, mm_cpumask(mm));
in enter_lazy_tlb(). It means that we won't be notified by peopel
invalidating the page tables, and while we then do re-validate the TLB
when we switch back from lazy mode, I still worry. I'm not entirely
convinced by that tlb_gen logic.
I can't actually see anything *wrong* in the tlb_gen logic, but it worries me.
Linus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 21:20 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unTsm-1hK-31@gated-at.bofh.it> |
| In reply to | #1729513 |
On Sat, Sep 09, 2017 at 11:47:33AM -0700, Linus Torvalds wrote:
> The thing is, even with the delayed TLB flushing, I don't think it
> should be *so* delayed that we should be seeing a TLB fill from
> garbage page tables.
Yeah, but we can't know what kind of speculative accesses happen between
the removal from the mask and the actual flushing.
> But the part in Andy's patch that worries me the most is that
>
> + cpumask_clear_cpu(cpu, mm_cpumask(mm));
>
> in enter_lazy_tlb(). It means that we won't be notified by peopel
> invalidating the page tables, and while we then do re-validate the TLB
> when we switch back from lazy mode, I still worry. I'm not entirely
> convinced by that tlb_gen logic.
>
> I can't actually see anything *wrong* in the tlb_gen logic, but it worries me.
Yeah, sounds like we're uncovering a situation of possibly stale
mappings which we haven't had before. Or at least widening that window.
And I still need to analyze what that MCE on Markus' machine is saying
exactly. The TlbCacheDis thing is an optimization which does away with
memory type checks. But we probably will have to disable it on those
boxes as we can't guarantee pagetable elements are all in WB mem...
Or we can guarantee them in WB but the lazy flushing delays the actual
clearing of the TLB entries so much so that they end up pointing to
garbage, as you say, which is not in WB mem and thus causes the protocol
error.
Hmm. All still wet.
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@kernel.org> |
|---|---|
| Date | 2017-09-09 21:30 +0200 |
| Message-ID | <unTC1-1le-3@gated-at.bofh.it> |
| In reply to | #1729522 |
On Sat, Sep 9, 2017 at 12:09 PM, Borislav Petkov <bp@alien8.de> wrote: > On Sat, Sep 09, 2017 at 11:47:33AM -0700, Linus Torvalds wrote: >> The thing is, even with the delayed TLB flushing, I don't think it >> should be *so* delayed that we should be seeing a TLB fill from >> garbage page tables. > > Yeah, but we can't know what kind of speculative accesses happen between > the removal from the mask and the actual flushing. > >> But the part in Andy's patch that worries me the most is that >> >> + cpumask_clear_cpu(cpu, mm_cpumask(mm)); >> >> in enter_lazy_tlb(). It means that we won't be notified by peopel >> invalidating the page tables, and while we then do re-validate the TLB >> when we switch back from lazy mode, I still worry. I'm not entirely >> convinced by that tlb_gen logic. >> >> I can't actually see anything *wrong* in the tlb_gen logic, but it worries me. > > Yeah, sounds like we're uncovering a situation of possibly stale > mappings which we haven't had before. Or at least widening that window. > > And I still need to analyze what that MCE on Markus' machine is saying > exactly. The TlbCacheDis thing is an optimization which does away with > memory type checks. But we probably will have to disable it on those > boxes as we can't guarantee pagetable elements are all in WB mem... > > Or we can guarantee them in WB but the lazy flushing delays the actual > clearing of the TLB entries so much so that they end up pointing to > garbage, as you say, which is not in WB mem and thus causes the protocol > error. > > Hmm. All still wet. > I think it's my theory #3. The CPU has a "paging-structure cache" (Intel lingo) that points to a freed page. The CPU speculatively follows it and gets complete garbage, triggering this MCE and who knows what else. I propose the following fix. If PCID is on, then, in enter_lazy_tlb(), we switch to init_mm with the no-flush flag set. (And we give init_mm its own dedicated ASID to keep it simple and fast -- no need to use the LRU ASID mapping to assign one dynamically.) We clear the bit in mm_cpumask. That is, we more or less just skip the whole lazy TLB optimization and rely on PCID CPUs having reasonably fast CR3 writes. No extra IPIs. I suppose I need to benchmark this. It will certainly slow down workloads that rapidly toggle between a user thread and a kernel thread because it forces serialization on each mm switch, but maybe that's not so bad. If PCID is off, then we leave the old CR3 value when we go lazy, and we also leave the flag in mm_cpumask set. When a flush is requested, we send out the IPI and switch to init_mm (and flush because we have no choice). IOW, the no-PCID behavior goes back to what it used to be. For the PCID case, I'm relying on this language in the SDM (vol 3, 4.10): When a logical processor creates entries in the TLBs (Section 4.10.2) and paging-structure caches (Section 4.10.3), it associates those entries with the current PCID. When using entries in the TLBs and paging-structure caches to translate a linear address, a logical processor uses only those entries associated with the current PCID (see Section 4.10.2.4 for an exception). This is also just common sense -- a CPU that makes any assumptions about a paging-structure cache for an inactive ASID is just nuts, especially if it assumes that the result of following it is at all sane. IOW, we really should be able to switch to ASID 1 and back to 0 without any flushes without worrying that the old page tables for ASID 1 might get freed afterwards. Obviously we need to flush if we switch back to PCID 1, but the code already does this. Also, sorry Rik, this means your old increased laziness optimization is dead in the water. It will have exactly the same speculative load problem.
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-09 21:40 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unTLH-1pv-13@gated-at.bofh.it> |
| In reply to | #1729525 |
On Sat, Sep 09, 2017 at 12:28:30PM -0700, Andy Lutomirski wrote:
> I propose the following fix. If PCID is on, then, in
> enter_lazy_tlb(), we switch to init_mm with the no-flush flag set.
> (And we give init_mm its own dedicated ASID to keep it simple and fast
> -- no need to use the LRU ASID mapping to assign one dynamically.) We
> clear the bit in mm_cpumask. That is, we more or less just skip the
> whole lazy TLB optimization and rely on PCID CPUs having reasonably
> fast CR3 writes. No extra IPIs. I suppose I need to benchmark this.
> It will certainly slow down workloads that rapidly toggle between a
> user thread and a kernel thread because it forces serialization on
> each mm switch, but maybe that's not so bad.
Sounds ok so far.
> If PCID is off, then we leave the old CR3 value when we go lazy, and
> we also leave the flag in mm_cpumask set. When a flush is requested,
> we send out the IPI and switch to init_mm (and flush because we have
> no choice). IOW, the no-PCID behavior goes back to what it used to
> be.
Ok, question: why can't we load the new CR3 value too, immediately? Or
are we saying, we might get to return to the same CR3 we had before we
were lazy so we won't need to do an unnecessary CR3 write with the same
value. A microoptimization, if you will.
Yes?
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@kernel.org> |
|---|---|
| Date | 2017-09-10 06:50 +0200 |
| Message-ID | <uo2lX-78W-1@gated-at.bofh.it> |
| In reply to | #1729528 |
On Sat, Sep 9, 2017 at 12:37 PM, Borislav Petkov <bp@alien8.de> wrote: > On Sat, Sep 09, 2017 at 12:28:30PM -0700, Andy Lutomirski wrote: >> I propose the following fix. If PCID is on, then, in >> enter_lazy_tlb(), we switch to init_mm with the no-flush flag set. >> (And we give init_mm its own dedicated ASID to keep it simple and fast >> -- no need to use the LRU ASID mapping to assign one dynamically.) We >> clear the bit in mm_cpumask. That is, we more or less just skip the >> whole lazy TLB optimization and rely on PCID CPUs having reasonably >> fast CR3 writes. No extra IPIs. I suppose I need to benchmark this. >> It will certainly slow down workloads that rapidly toggle between a >> user thread and a kernel thread because it forces serialization on >> each mm switch, but maybe that's not so bad. > > Sounds ok so far. > >> If PCID is off, then we leave the old CR3 value when we go lazy, and >> we also leave the flag in mm_cpumask set. When a flush is requested, >> we send out the IPI and switch to init_mm (and flush because we have >> no choice). IOW, the no-PCID behavior goes back to what it used to >> be. > > Ok, question: why can't we load the new CR3 value too, immediately? Or > are we saying, we might get to return to the same CR3 we had before we > were lazy so we won't need to do an unnecessary CR3 write with the same > value. A microoptimization, if you will. It is indeed a microoptimization, but it's a microoptimization that we've had in the kernel for a long, long time. But it may be an ill-advised microoptimization, or at least a poorly implemented one historically. The microoptimization mostly affects workloads that have a process on an otherwise idle CPU that frequently sleeps for very short times. With the optimization, we avoid two TLB flushes and two serializing instructions every time we sleep. Historically, we got a bunch of useless IPIs, too, depending on the workload. The problem is that the implementation, which lives in kernel/sched/core.c for the most part, involves some extra reference counting, and there are NUMA workloads with many cores all running the same mm that pay a *huge* cost in refcounting, since all the CPUs are hammering the same refcount. And this refcount is (I think) basically pointless on x86 and maybe on most architectures. PeterZ and Ingo, would you be okay with adding a define so arches can opt out of the task_struct::active_mm field entirely? That is, with the option set, task_struct wouldn't have an active_mm field, the core wouldn't call mmgrab and mmdrop, and the arch would be responsible for that bookkeeping instead? x86, and presumably all arches without cross-core invalidation, would probably prefer to just shoot down the old mm entirely in __mmput() rather than trying to figure out when do finish freeing old mms. After all, exit_mmap() is going to send an IPI regardless, so I see no reason to have the scheduler core pin an old dead mm just because some random kernel thread's active_mm field points to it. IOW, if I'm going to reintroduce something like what the old lazy mode did on x86, I'd rather do it right. --Andy
[toc] | [prev] | [next] | [standalone]
| From | Linus Torvalds <torvalds@linux-foundation.org> |
|---|---|
| Date | 2017-09-09 21:30 +0200 |
| Message-ID | <unTC1-1le-5@gated-at.bofh.it> |
| In reply to | #1729522 |
On Sat, Sep 9, 2017 at 12:09 PM, Borislav Petkov <bp@alien8.de> wrote:
>
> Yeah, but we can't know what kind of speculative accesses happen between
> the removal from the mask and the actual flushing.
Indeed. The speculative kernel thread accesses while lazy could easily
trigger this.
And I guess those are pretty fundamental. So..
Linus
[toc] | [prev] | [next] | [standalone]
| From | Markus Trippelsdorf <markus@trippelsdorf.de> |
|---|---|
| Date | 2017-09-09 10:20 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unJ9D-2Ii-3@gated-at.bofh.it> |
| In reply to | #1729310 |
On 2017.09.08 at 14:47 -0700, Andy Lutomirski wrote: > On Fri, Sep 8, 2017 at 10:16 AM, Markus Trippelsdorf > <markus@trippelsdorf.de> wrote: > > On 2017.09.08 at 09:12 -0700, Andy Lutomirski wrote: > >> On Fri, Sep 8, 2017 at 4:30 AM, Markus Trippelsdorf > >> <markus@trippelsdorf.de> wrote: > >> > On 2017.09.08 at 12:39 +0200, Markus Trippelsdorf wrote: > >> >> On 2017.09.08 at 12:35 +0200, Ingo Molnar wrote: > >> >> > > >> >> > * Markus Trippelsdorf <markus@trippelsdorf.de> wrote: > >> >> > > >> >> > > On 2017.09.08 at 11:16 +0200, Borislav Petkov wrote: > >> >> > > > On Fri, Sep 08, 2017 at 10:05:36AM +0200, Borislav Petkov wrote: > >> >> > > > > On Fri, Sep 08, 2017 at 08:26:44AM +0200, Thomas Gleixner wrote: > >> >> > > > > > On Fri, 8 Sep 2017, Markus Trippelsdorf wrote: > >> >> > > > > > > >> >> > > > > > CC+ Borislav. He might have access to such a beast > >> >> > > > > > >> >> > > > > Can I have /proc/cpuinfo and dmesg pls, in order to see whether I have > >> >> > > > > something similar? > >> >> > > > > > >> >> > > > > Private mail's fine too. > >> >> > > > > >> >> > > > So I don't have exactly your model - mine is model 2, stepping 3 but I see > >> >> > > > something strange too, in dmesg: > >> >> > > > >> >> > > I'm pretty sure the bug is in the merged 'x86-mm-for-linus' branch: > >> >> > > Either Andy's "PCID optimized TLB flushing" (would be my guess) or > >> >> > > 'encrypted memory' support by Tom Lendacky. > >> >> > > > >> >> > > (Bisecting is hard, because sometimes I can compile stuff for over 15 > >> >> > > minutes without hitting the bug. At other times the machine locks up > >> >> > > hard when starting X11 already.) > >> >> > > >> >> > Do you have the 72c0098d92ce fix? > >> >> > >> >> Yes. The bug still happens on the current git tree (which has the fix > >> >> already): > >> > > >> > The bug is definitely caused by Andy Lutomirski's PCID optimized TLB > >> > flushing" patches. Tom is off the hook. > >> > >> I'm pretty sure it can't be PCID per se, since these CPUs are way too > >> old and are very unlikely to have PCID. > > > > Yes, the CPU doesn't support PCID (,but it does support PGE). > > > >> It could plausibly be the lazy TLB flushing changes. > > > > Yes, I've narrowed it down to: > > > > commit 94b1b03b519b81c494900cb112aa00ed205cc2d9 > > Author: Andy Lutomirski <luto@kernel.org> > > Date: Thu Jun 29 08:53:17 2017 -0700 > > > > x86/mm: Rework lazy TLB mode and TLB freshness tracking > > > > > > Theoretically you guys should be able to reproduce the issue by using > > the "nopcid" boot option. > > > > Any chance you could test with CONFIG_DEBUG_VM=y? There are lots of > potentially useful assertions in that code. CONFIG_DEBUG_VM=y doesn't change anything. I still get the hard hang without anything in the logs. > Can you also post your /proc/cpuinfo? And can you re-confirm that a > problematic guest kernel is causing problems in the *host*? processor : 0 vendor_id : AuthenticAMD cpu family : 16 model : 4 model name : AMD Phenom(tm) II X4 955 Processor stepping : 2 microcode : 0x10000db cpu MHz : 3210.960 cache size : 512 KB physical id : 0 siblings : 4 core id : 0 cpu cores : 4 apicid : 0 initial apicid : 0 fpu : yes fpu_exception : yes cpuid level : 5 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm 3dnowext 3dnow constant_tsc rep_good nopl nonstop_tsc cpuid extd_apicid pni monitor cx16 popcnt lahf_lm cmp_legacy svm extapic cr8_legacy abm sse4a misalignsse 3dnowprefetch osvw ibs skinit wdt hw_pstate vmmcall npt lbrv svm_lock nrip_save bugs : tlb_mmatch apic_c1e fxsave_leak sysret_ss_attrs null_seg amd_e400 bogomips : 6424.50 TLB size : 1024 4K pages clflush size : 64 cache_alignment : 64 address sizes : 48 bits physical, 48 bits virtual power management: ts ttp tm stc 100mhzsteps hwpstate Unfortunately I cannot reproduce the qemu (kvm) problem anymore. (Perhaps I have not tried long enough). Anyway, kvm has code that should handle erratum_383. -- Markus
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@alien8.de> |
|---|---|
| Date | 2017-09-08 17:00 +0200 |
| Subject | Re: Current mainline git (24e700e291d52bd2) hangs when building e.g. perf |
| Message-ID | <unsVc-8sm-9@gated-at.bofh.it> |
| In reply to | #1728713 |
On Fri, Sep 08, 2017 at 11:16:14AM +0200, Borislav Petkov wrote:
...
> +blkid[1020]: segfault at 8 ip 00007fec7760c3fd sp 00007ffc9dd05890 error 4 in ld-linux-x86-64.so.2[7fec775fc000+22000]
> +blkid[1026]: segfault at 8 ip 00007f5a31ecc3fd sp 00007fffaf3604b0 error 4 in ld-linux-x86-64.so.2[7f5a31ebc000+22000]
> +logsave[1027]: segfault at 8 ip 00007f237d2033fd sp 00007fff53933e60 error 4 in ld-linux-x86-64.so.2[7f237d1f3000+22000]
Yap, definitely no segfaults with 4.13
--
Regards/Gruss,
Boris.
Good mailing practices for 400: avoid top-posting and trim the reply.
[toc] | [prev] | [standalone]
Page 3 of 3 — ← Prev page 1 2 [3]
Back to top | Article view | linux.kernel
csiph-web