Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1160706 > unrolled thread
| Started by | Dave Hansen <dave.hansen@intel.com> |
|---|---|
| First post | 2015-06-08 20:30 +0200 |
| Last post | 2015-06-08 22:10 +0200 |
| Articles | 3 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH 0/3] TLB flush multiple pages per IPI v5 Dave Hansen <dave.hansen@intel.com> - 2015-06-08 20:30 +0200
Re: [PATCH 0/3] TLB flush multiple pages per IPI v5 Ingo Molnar <mingo@kernel.org> - 2015-06-08 22:00 +0200
Re: [PATCH 0/3] TLB flush multiple pages per IPI v5 Ingo Molnar <mingo@kernel.org> - 2015-06-08 22:10 +0200
| From | Dave Hansen <dave.hansen@intel.com> |
|---|---|
| Date | 2015-06-08 20:30 +0200 |
| Subject | Re: [PATCH 0/3] TLB flush multiple pages per IPI v5 |
| Message-ID | <pz9Y6-6UC-21@gated-at.bofh.it> |
On 06/08/2015 10:45 AM, Ingo Molnar wrote: > As per my measurements the __flush_tlb_single() primitive (which you use in patch > #2) is very expensive on most Intel and AMD CPUs. It barely makes sense for a 2 > pages and gets exponentially worse. It's probably done in microcode and its > performance is horrible. I discussed this a bit in commit a5102476a2. I'd be curious what numbers you came up with. But, don't we have to take in to account the cost of refilling the TLB in addition to the cost of emptying it? The TLB size is historically increasing on a per-core basis, so isn't this refill cost only going to get worse? -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2015-06-08 22:00 +0200 |
| Message-ID | <pzbnc-pK-25@gated-at.bofh.it> |
| In reply to | #1160706 |
* Dave Hansen <dave.hansen@intel.com> wrote:
> On 06/08/2015 10:45 AM, Ingo Molnar wrote:
> > As per my measurements the __flush_tlb_single() primitive (which you use in patch
> > #2) is very expensive on most Intel and AMD CPUs. It barely makes sense for a 2
> > pages and gets exponentially worse. It's probably done in microcode and its
> > performance is horrible.
>
> I discussed this a bit in commit a5102476a2. I'd be curious what
> numbers you came up with.
... which for those of us who don't have sha1's cached in their brain is:
a5102476a24b ("x86/mm: Set TLB flush tunable to sane value (33)")
;-)
So what I measured agrees generally with the comment you added in the commit:
+ * Each single flush is about 100 ns, so this caps the maximum overhead at
+ * _about_ 3,000 ns.
Let that sink through: 3,000 nsecs = 3 usecs, that's like eternity!
A CR3 driven TLB flush takes less time than a single INVLPG (!):
[ 0.389028] x86/fpu: Cost of: __flush_tlb() fn : 96 cycles
[ 0.405885] x86/fpu: Cost of: __flush_tlb_one() fn : 260 cycles
[ 0.414302] x86/fpu: Cost of: __flush_tlb_range() fn : 404 cycles
it's true that a full flush has hidden costs not measured above, because it has
knock-on effects (because it drops non-global TLB entries), but it's not _that_
bad due to:
- there almost always being a L1 or L2 cache miss when a TLB miss occurs,
which latency can be overlaid
- global bit being held for kernel entries
- user-space with high memory pressure trashing through TLBs typically
... and especially with caches and Intel's historically phenomenally low TLB
refill latency it's difficult to measure the effects of local TLB refills, let
alone measure it in any macro benchmark.
Cross-CPU flushes are expensive, absolutely no argument about that - my suggestion
here is to keep the batching but simplify it: because I strongly suspect that the
biggest win is the batching, not the pfn queueing.
We might even win a bit more performance due to the simplification.
> But, don't we have to take in to account the cost of refilling the TLB in
> addition to the cost of emptying it? The TLB size is historically increasing on
> a per-core basis, so isn't this refill cost only going to get worse?
Only if TLB refill latency sucks - but Intel's is very good and AMD's is pretty
good as well.
Also, usually if you miss the TLB you miss the cache line as well (you definitely
miss the L1 cache, and TLB caches are sized to hold a fair chunk of your L2
cache), and the CPU can overlap the two latencies.
So while it might sound counter-intuitive, a full TLB flush might be faster than
trying to do software based TLB cache management ...
INVLPG really sucks. I can be convinced by numbers, but this isn't nearly as
clear-cut as it might look.
Thanks,
Ingo
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2015-06-08 22:10 +0200 |
| Message-ID | <pzbwS-QJ-25@gated-at.bofh.it> |
| In reply to | #1160750 |
* Ingo Molnar <mingo@kernel.org> wrote: > So what I measured agrees generally with the comment you added in the commit: > > + * Each single flush is about 100 ns, so this caps the maximum overhead at > + * _about_ 3,000 ns. > > Let that sink through: 3,000 nsecs = 3 usecs, that's like eternity! > > A CR3 driven TLB flush takes less time than a single INVLPG (!): > > [ 0.389028] x86/fpu: Cost of: __flush_tlb() fn : 96 cycles > [ 0.405885] x86/fpu: Cost of: __flush_tlb_one() fn : 260 cycles > [ 0.414302] x86/fpu: Cost of: __flush_tlb_range() fn : 404 cycles > > it's true that a full flush has hidden costs not measured above, because it has > knock-on effects (because it drops non-global TLB entries), but it's not _that_ > bad due to: > > - there almost always being a L1 or L2 cache miss when a TLB miss occurs, > which latency can be overlaid > > - global bit being held for kernel entries > > - user-space with high memory pressure trashing through TLBs typically I also have cache-cold numbers from another (Intel) system: [ 0.176473] x86/bench:########################################################################## [ 0.185656] x86/bench: Running x86 benchmarks: cache- hot / cold cycles [ 1.234448] x86/bench: Cost of: null : 35 / 73 cycles [ ........] [ 27.930451] x86/bench:######## MM instructions: ###################################### [ 28.979251] x86/bench: Cost of: __flush_tlb() fn : 251 / 366 cycles [ 30.028795] x86/bench: Cost of: __flush_tlb_global() fn : 746 / 1795 cycles [ 31.077862] x86/bench: Cost of: __flush_tlb_one() fn : 237 / 883 cycles [ 32.127371] x86/bench: Cost of: __flush_tlb_range() fn : 312 / 1603 cycles [ 35.254202] x86/bench: Cost of: wbinvd() insn : 2491761 / 2491922 cycles Note how the numbers are even worse in the cache-cold case: the algorithmic complexity of __flush_tlb_range() versus __flush_tlb() makes it run slower (because we miss the I$), while the TLB cache-preservation argument is probably weaker, because when we are cache cold then TLB refill latency probably matters less (as it can be overlapped). So __flush_tlb_range() is software trying to beat hardware, and that's almost always a bad idea on x86. Thanks, Ingo -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web