Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1304320 > unrolled thread
| Started by | Chris Wilson <chris@chris-wilson.co.uk> |
|---|---|
| First post | 2016-01-08 11:00 +0100 |
| Last post | 2016-01-08 19:40 +0100 |
| Articles | 4 — 4 participants |
Back to article view | Back to linux.kernel
[PATCH] x86: Micro-optimise clflush_cache_range() Chris Wilson <chris@chris-wilson.co.uk> - 2016-01-08 11:00 +0100
Re: [PATCH] x86: Micro-optimise clflush_cache_range() Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-01-08 18:00 +0100
Re: [PATCH] x86: Micro-optimise clflush_cache_range() Toshi Kani <toshi.kani@hpe.com> - 2016-01-08 19:30 +0100
[tip:x86/mm] x86/mm: Micro-optimise clflush_cache_range() tip-bot for Chris Wilson <tipbot@zytor.com> - 2016-01-08 19:40 +0100
| From | Chris Wilson <chris@chris-wilson.co.uk> |
|---|---|
| Date | 2016-01-08 11:00 +0100 |
| Subject | [PATCH] x86: Micro-optimise clflush_cache_range() |
| Message-ID | <qOBZV-4VO-39@gated-at.bofh.it> |
Whilst inspecting the asm for clflush_cache_range() and some perf profiles
that required extensive flushing of single cachelines (from part of the
intel-gpu-tools GPU benchmarks), we noticed that gcc was reloading
boot_cpu_data.x86_clflush_size on every iteration of the loop. We can
manually hoist that read which perf regarded as taking ~25% of the
function time for a single cacheline flush.
Signed-off-by: Chris Wilson <chris@chris-wilson.co.uk>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: "H. Peter Anvin" <hpa@zytor.com>
Cc: x86@kernel.org
Cc: Toshi Kani <toshi.kani@hpe.com>
Cc: Borislav Petkov <bp@suse.de>
Cc: "Luis R. Rodriguez" <mcgrof@suse.com>
Cc: Stephen Rothwell <sfr@canb.auug.org.au>
Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
Cc: Sai Praneeth <sai.praneeth.prakhya@intel.com>
Cc: linux-kernel@vger.kernel.org
Acked-by: "H. Peter Anvin" <hpa@zytor.com>
---
arch/x86/mm/pageattr.c | 10 ++++++----
1 file changed, 6 insertions(+), 4 deletions(-)
diff --git a/arch/x86/mm/pageattr.c b/arch/x86/mm/pageattr.c
index a3137a4feed1..6000ad7f560c 100644
--- a/arch/x86/mm/pageattr.c
+++ b/arch/x86/mm/pageattr.c
@@ -129,14 +129,16 @@ within(unsigned long addr, unsigned long start, unsigned long end)
*/
void clflush_cache_range(void *vaddr, unsigned int size)
{
- unsigned long clflush_mask = boot_cpu_data.x86_clflush_size - 1;
+ const unsigned long clflush_size = boot_cpu_data.x86_clflush_size;
+ void *p = (void *)((unsigned long)vaddr & ~(clflush_size - 1));
void *vend = vaddr + size;
- void *p;
+
+ if (p >= vend)
+ return;
mb();
- for (p = (void *)((unsigned long)vaddr & ~clflush_mask);
- p < vend; p += boot_cpu_data.x86_clflush_size)
+ for (; p < vend; p += clflush_size)
clflushopt(p);
mb();
--
2.7.0.rc3
[toc] | [next] | [standalone]
| From | Ross Zwisler <ross.zwisler@linux.intel.com> |
|---|---|
| Date | 2016-01-08 18:00 +0100 |
| Message-ID | <qOIyl-YW-5@gated-at.bofh.it> |
| In reply to | #1304320 |
On Fri, Jan 08, 2016 at 09:55:33AM +0000, Chris Wilson wrote:
> Whilst inspecting the asm for clflush_cache_range() and some perf profiles
> that required extensive flushing of single cachelines (from part of the
> intel-gpu-tools GPU benchmarks), we noticed that gcc was reloading
> boot_cpu_data.x86_clflush_size on every iteration of the loop. We can
> manually hoist that read which perf regarded as taking ~25% of the
> function time for a single cacheline flush.
>
> Signed-off-by: Chris Wilson <chris@chris-wilson.co.uk>
> Cc: Thomas Gleixner <tglx@linutronix.de>
> Cc: Ingo Molnar <mingo@redhat.com>
> Cc: "H. Peter Anvin" <hpa@zytor.com>
> Cc: x86@kernel.org
> Cc: Toshi Kani <toshi.kani@hpe.com>
> Cc: Borislav Petkov <bp@suse.de>
> Cc: "Luis R. Rodriguez" <mcgrof@suse.com>
> Cc: Stephen Rothwell <sfr@canb.auug.org.au>
> Cc: Ross Zwisler <ross.zwisler@linux.intel.com>
> Cc: Sai Praneeth <sai.praneeth.prakhya@intel.com>
> Cc: linux-kernel@vger.kernel.org
> Acked-by: "H. Peter Anvin" <hpa@zytor.com>
Looks good to me.
Reviewed-by: Ross Zwisler <ross.zwisler@linux.intel.com>
> ---
> arch/x86/mm/pageattr.c | 10 ++++++----
> 1 file changed, 6 insertions(+), 4 deletions(-)
>
> diff --git a/arch/x86/mm/pageattr.c b/arch/x86/mm/pageattr.c
> index a3137a4feed1..6000ad7f560c 100644
> --- a/arch/x86/mm/pageattr.c
> +++ b/arch/x86/mm/pageattr.c
> @@ -129,14 +129,16 @@ within(unsigned long addr, unsigned long start, unsigned long end)
> */
> void clflush_cache_range(void *vaddr, unsigned int size)
> {
> - unsigned long clflush_mask = boot_cpu_data.x86_clflush_size - 1;
> + const unsigned long clflush_size = boot_cpu_data.x86_clflush_size;
> + void *p = (void *)((unsigned long)vaddr & ~(clflush_size - 1));
> void *vend = vaddr + size;
> - void *p;
> +
> + if (p >= vend)
> + return;
>
> mb();
>
> - for (p = (void *)((unsigned long)vaddr & ~clflush_mask);
> - p < vend; p += boot_cpu_data.x86_clflush_size)
> + for (; p < vend; p += clflush_size)
> clflushopt(p);
>
> mb();
> --
> 2.7.0.rc3
>
[toc] | [prev] | [next] | [standalone]
| From | Toshi Kani <toshi.kani@hpe.com> |
|---|---|
| Date | 2016-01-08 19:30 +0100 |
| Message-ID | <qOJXs-29h-5@gated-at.bofh.it> |
| In reply to | #1304320 |
On Fri, 2016-01-08 at 09:55 +0000, Chris Wilson wrote: > Whilst inspecting the asm for clflush_cache_range() and some perf > profiles > that required extensive flushing of single cachelines (from part of the > intel-gpu-tools GPU benchmarks), we noticed that gcc was reloading > boot_cpu_data.x86_clflush_size on every iteration of the loop. We can > manually hoist that read which perf regarded as taking ~25% of the > function time for a single cacheline flush. > > Signed-off-by: Chris Wilson <chris@chris-wilson.co.uk> > Cc: Thomas Gleixner <tglx@linutronix.de> > Cc: Ingo Molnar <mingo@redhat.com> > Cc: "H. Peter Anvin" <hpa@zytor.com> > Cc: x86@kernel.org > Cc: Toshi Kani <toshi.kani@hpe.com> > Cc: Borislav Petkov <bp@suse.de> > Cc: "Luis R. Rodriguez" <mcgrof@suse.com> > Cc: Stephen Rothwell <sfr@canb.auug.org.au> > Cc: Ross Zwisler <ross.zwisler@linux.intel.com> > Cc: Sai Praneeth <sai.praneeth.prakhya@intel.com> > Cc: linux-kernel@vger.kernel.org > Acked-by: "H. Peter Anvin" <hpa@zytor.com> Thanks for the improvement! The change looks good to me. Reviewed-by: Toshi Kani <toshi.kani@hpe.com> -Toshi
[toc] | [prev] | [next] | [standalone]
| From | tip-bot for Chris Wilson <tipbot@zytor.com> |
|---|---|
| Date | 2016-01-08 19:40 +0100 |
| Subject | [tip:x86/mm] x86/mm: Micro-optimise clflush_cache_range() |
| Message-ID | <qOK7a-2eA-67@gated-at.bofh.it> |
| In reply to | #1304320 |
Commit-ID: 1f1a89ac05f6e88aa341e86e57435fdbb1177c0c
Gitweb: http://git.kernel.org/tip/1f1a89ac05f6e88aa341e86e57435fdbb1177c0c
Author: Chris Wilson <chris@chris-wilson.co.uk>
AuthorDate: Fri, 8 Jan 2016 09:55:33 +0000
Committer: Thomas Gleixner <tglx@linutronix.de>
CommitDate: Fri, 8 Jan 2016 19:27:39 +0100
x86/mm: Micro-optimise clflush_cache_range()
Whilst inspecting the asm for clflush_cache_range() and some perf profiles
that required extensive flushing of single cachelines (from part of the
intel-gpu-tools GPU benchmarks), we noticed that gcc was reloading
boot_cpu_data.x86_clflush_size on every iteration of the loop. We can
manually hoist that read which perf regarded as taking ~25% of the
function time for a single cacheline flush.
Signed-off-by: Chris Wilson <chris@chris-wilson.co.uk>
Reviewed-by: Ross Zwisler <ross.zwisler@linux.intel.com>
Acked-by: "H. Peter Anvin" <hpa@zytor.com>
Cc: Toshi Kani <toshi.kani@hpe.com>
Cc: Borislav Petkov <bp@suse.de>
Cc: Luis R. Rodriguez <mcgrof@suse.com>
Cc: Stephen Rothwell <sfr@canb.auug.org.au>
Cc: Sai Praneeth <sai.praneeth.prakhya@intel.com>
Link: http://lkml.kernel.org/r/1452246933-10890-1-git-send-email-chris@chris-wilson.co.uk
Signed-off-by: Thomas Gleixner <tglx@linutronix.de>
---
arch/x86/mm/pageattr.c | 10 ++++++----
1 file changed, 6 insertions(+), 4 deletions(-)
diff --git a/arch/x86/mm/pageattr.c b/arch/x86/mm/pageattr.c
index a3137a4..6000ad7 100644
--- a/arch/x86/mm/pageattr.c
+++ b/arch/x86/mm/pageattr.c
@@ -129,14 +129,16 @@ within(unsigned long addr, unsigned long start, unsigned long end)
*/
void clflush_cache_range(void *vaddr, unsigned int size)
{
- unsigned long clflush_mask = boot_cpu_data.x86_clflush_size - 1;
+ const unsigned long clflush_size = boot_cpu_data.x86_clflush_size;
+ void *p = (void *)((unsigned long)vaddr & ~(clflush_size - 1));
void *vend = vaddr + size;
- void *p;
+
+ if (p >= vend)
+ return;
mb();
- for (p = (void *)((unsigned long)vaddr & ~clflush_mask);
- p < vend; p += boot_cpu_data.x86_clflush_size)
+ for (; p < vend; p += clflush_size)
clflushopt(p);
mb();
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web