Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.programming.threads > #2529 > unrolled thread
| Started by | Ramine <ramine@1.1> |
|---|---|
| First post | 2014-06-28 15:51 -0700 |
| Last post | 2014-06-29 10:54 -0700 |
| Articles | 15 — 3 participants |
Back to article view | Back to comp.programming.threads
To Melzzzzz Ramine <ramine@1.1> - 2014-06-28 15:51 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.invalid> - 2014-06-29 04:10 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:33 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 16:39 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:47 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:00 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:14 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:20 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:23 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:31 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:32 -0700
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:38 -0700
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:41 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.invalid> - 2014-06-29 16:46 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:54 -0700
| From | Ramine <ramine@1.1> |
|---|---|
| Date | 2014-06-28 15:51 -0700 |
| Subject | To Melzzzzz |
| Message-ID | <longtm$bkp$1@dont-email.me> |
Hello, I have looked at your matrix multiplication using SSE2 and AVX that you have posted on assembler.x86 forum, i just wanted to tell you that you don't need to write it in assembler, cause GCC do auto-vectorize the floating point calculation and it auto-vectorize also the Matrix multiplication even at -O1 optimization... and GCC do auto-vectorize beautifully since the GCC auto-vectorization have given a good results on scimark2 benchmark too. Thank you, Amine Moulay Ramdane.
[toc] | [next] | [standalone]
| From | Melzzzzz <mel@zzzzz.invalid> |
|---|---|
| Date | 2014-06-29 04:10 +0200 |
| Message-ID | <lonsia$ea2$1@solani.org> |
| In reply to | #2529 |
On Sat, 28 Jun 2014 15:51:59 -0700 Ramine <ramine@1.1> wrote: > > Hello, > > I have looked at your matrix multiplication using > SSE2 and AVX that you have posted on assembler.x86 forum, > i just wanted to tell you that you don't need to write > it in assembler, cause GCC do auto-vectorize the floating > point calculation and it auto-vectorize also the Matrix > multiplication even at -O1 optimization... as I see it it don't.... and GCC do > auto-vectorize beautifully since the GCC auto-vectorization > have given a good results on scimark2 benchmark too. Care to share example? > > > Thank you, > Amine Moulay Ramdane. > > > > -- Click OK to continue...
[toc] | [prev] | [next] | [standalone]
| From | Ramine <ramine@1.1> |
|---|---|
| Date | 2014-06-29 10:33 -0700 |
| Message-ID | <lop839$ute$1@dont-email.me> |
| In reply to | #2531 |
On 6/28/2014 7:10 PM, Melzzzzz wrote: > On Sat, 28 Jun 2014 15:51:59 -0700 > Ramine <ramine@1.1> wrote: > >> >> Hello, >> >> I have looked at your matrix multiplication using >> SSE2 and AVX that you have posted on assembler.x86 forum, >> i just wanted to tell you that you don't need to write >> it in assembler, cause GCC do auto-vectorize the floating >> point calculation and it auto-vectorize also the Matrix >> multiplication even at -O1 optimization... > > as I see it it don't.... > I am using the newest version of tdm-gcc and auto-vectorization is working even at -O1 level optimization. You can download tdb-gcc from here and try it yourself: http://tdm-gcc.tdragon.net/ Thank you, Amine Moulay Ramdane. > > and GCC do >> auto-vectorize beautifully since the GCC auto-vectorization >> have given a good results on scimark2 benchmark too. > > Care to share example? > >> >> >> Thank you, >> Amine Moulay Ramdane. >> >> >> >> > > >
[toc] | [prev] | [next] | [standalone]
| From | Melzzzzz <mel@zzzzz.com> |
|---|---|
| Date | 2014-06-29 16:39 +0200 |
| Message-ID | <lop8fk$9ui$3@news.albasani.net> |
| In reply to | #2532 |
On Sun, 29 Jun 2014 10:33:41 -0700 Ramine <ramine@1.1> wrote: > On 6/28/2014 7:10 PM, Melzzzzz wrote: > > On Sat, 28 Jun 2014 15:51:59 -0700 > > Ramine <ramine@1.1> wrote: > > > >> > >> Hello, > >> > >> I have looked at your matrix multiplication using > >> SSE2 and AVX that you have posted on assembler.x86 forum, > >> i just wanted to tell you that you don't need to write > >> it in assembler, cause GCC do auto-vectorize the floating > >> point calculation and it auto-vectorize also the Matrix > >> multiplication even at -O1 optimization... > > > > as I see it it don't.... > > > > > I am using the newest version of tdm-gcc and auto-vectorization is > working even at -O1 level optimization. > > > You can download tdb-gcc from here and try it yourself: > > http://tdm-gcc.tdragon.net/ That is based on gcc 4.8.1 , Im using 4.9 and trunk versions and there is no vectorization of matrix multiplication (even for 4x4 matrices). Care to show compiler output of assembly and c source code so I can see what I am missing? > > > > > Thank you, > Amine Moulay Ramdane. > > > > > > > > > > > and GCC do > >> auto-vectorize beautifully since the GCC auto-vectorization > >> have given a good results on scimark2 benchmark too. > > > > Care to share example? > > > >> > >> > >> Thank you, > >> Amine Moulay Ramdane. > >> > >> > >> > >> > > > > > > > -- Click OK to continue...
[toc] | [prev] | [next] | [standalone]
| From | Ramine <ramine@1.1> |
|---|---|
| Date | 2014-06-29 10:47 -0700 |
| Message-ID | <lop8tr$7ra$1@dont-email.me> |
| In reply to | #2533 |
Hello, Look for example at the assembler source code of SparseCompRow.c that you will find inside the scimark2 source code benchmark, look carefully at the follwing assembler code, and you will notice that it is auto-vectorizing, i am using just using -O1 level optimization and the -S flag to generate the assembler code, here it is: === .file "SparseCompRow.c" .text .globl SparseCompRow_num_flops .def SparseCompRow_num_flops; .scl 2; .type 32; .endef .seh_proc SparseCompRow_num_flops SparseCompRow_num_flops: .seh_endprologue movl %edx, %eax sarl $31, %edx idivl %ecx imull %eax, %ecx cvtsi2sd %ecx, %xmm0 addsd %xmm0, %xmm0 cvtsi2sd %r8d, %xmm1 mulsd %xmm1, %xmm0 ret .seh_endproc .globl SparseCompRow_matmult .def SparseCompRow_matmult; .scl 2; .type 32; .endef .seh_proc SparseCompRow_matmult SparseCompRow_matmult: pushq %r13 .seh_pushreg %r13 pushq %r12 .seh_pushreg %r12 pushq %rbp .seh_pushreg %rbp pushq %rdi .seh_pushreg %rdi pushq %rsi .seh_pushreg %rsi pushq %rbx .seh_pushreg %rbx .seh_endprologue movl %ecx, %edi movq %rdx, %rbp movq %r8, %rdx movq %r9, %rsi movq 88(%rsp), %r9 movq 96(%rsp), %rcx movl 104(%rsp), %r13d movl $0, %r12d xorpd %xmm2, %xmm2 movapd %xmm2, %xmm3 testl %r13d, %r13d jg .L15 jmp .L2 .L12: movl (%rsi,%rbx,4), %eax movl 4(%rsi,%rbx,4), %r8d cmpl %r8d, %eax jge .L10 movapd %xmm2, %xmm0 .L6: movslq %eax, %r10 movslq (%r9,%r10,4), %r11 movsd (%rcx,%r11,8), %xmm1 mulsd (%rdx,%r10,8), %xmm1 addsd %xmm1, %xmm0 addl $1, %eax cmpl %r8d, %eax jne .L6 jmp .L5 .L10: movapd %xmm3, %xmm0 .L5: movsd %xmm0, 0(%rbp,%rbx,8) addq $1, %rbx cmpl %ebx, %edi jg .L12 .L8: addl $1, %r12d cmpl %r13d, %r12d jne .L15 jmp .L2 .L15: movl $0, %ebx testl %edi, %edi jg .L12 .p2align 4,,4 jmp .L8 .L2: popq %rbx popq %rsi popq %rdi popq %rbp popq %r12 popq %r13 ret .seh_endproc == Thank you, Amine Moulay Ramdane. On 6/29/2014 7:39 AM, Melzzzzz wrote: > On Sun, 29 Jun 2014 10:33:41 -0700 > Ramine <ramine@1.1> wrote: > >> On 6/28/2014 7:10 PM, Melzzzzz wrote: >>> On Sat, 28 Jun 2014 15:51:59 -0700 >>> Ramine <ramine@1.1> wrote: >>> >>>> >>>> Hello, >>>> >>>> I have looked at your matrix multiplication using >>>> SSE2 and AVX that you have posted on assembler.x86 forum, >>>> i just wanted to tell you that you don't need to write >>>> it in assembler, cause GCC do auto-vectorize the floating >>>> point calculation and it auto-vectorize also the Matrix >>>> multiplication even at -O1 optimization... >>> >>> as I see it it don't.... >>> >> >> >> I am using the newest version of tdm-gcc and auto-vectorization is >> working even at -O1 level optimization. >> >> >> You can download tdb-gcc from here and try it yourself: >> >> http://tdm-gcc.tdragon.net/ > > That is based on gcc 4.8.1 , Im using 4.9 and trunk versions > and there is no vectorization of matrix multiplication > (even for 4x4 matrices). > Care to show compiler output of assembly and c source code > so I can see what I am missing? > >> >> >> >> >> Thank you, >> Amine Moulay Ramdane. >> >> >> >> >> >> >> >>> >>> and GCC do >>>> auto-vectorize beautifully since the GCC auto-vectorization >>>> have given a good results on scimark2 benchmark too. >>> >>> Care to share example? >>> >>>> >>>> >>>> Thank you, >>>> Amine Moulay Ramdane. >>>> >>>> >>>> >>>> >>> >>> >>> >> > > >
[toc] | [prev] | [next] | [standalone]
| From | Melzzzzz <mel@zzzzz.com> |
|---|---|
| Date | 2014-06-29 17:00 +0200 |
| Message-ID | <lop9ml$9ui$4@news.albasani.net> |
| In reply to | #2536 |
On Sun, 29 Jun 2014 10:47:51 -0700 Ramine <ramine@1.1> wrote: > > Hello, > > Look for example at the assembler source code of > SparseCompRow.c that you will find inside the > scimark2 source code benchmark, look carefully > at the follwing assembler code, and you will notice > that it is auto-vectorizing, i am using just using > -O1 level optimization and the -S flag to generate > the assembler code, here it is: > > > === > > > .file "SparseCompRow.c" > .text > .globl SparseCompRow_num_flops > .def SparseCompRow_num_flops; .scl > 2; .type 32; .endef .seh_proc > SparseCompRow_num_flops SparseCompRow_num_flops: > .seh_endprologue > movl %edx, %eax > sarl $31, %edx > idivl %ecx > imull %eax, %ecx > cvtsi2sd %ecx, %xmm0 > addsd %xmm0, %xmm0 > cvtsi2sd %r8d, %xmm1 > mulsd %xmm1, %xmm0 > ret > .seh_endproc > .globl SparseCompRow_matmult > .def SparseCompRow_matmult; .scl > 2; .type 32; .endef .seh_proc > SparseCompRow_matmult SparseCompRow_matmult: > pushq %r13 > .seh_pushreg %r13 > pushq %r12 > .seh_pushreg %r12 > pushq %rbp > .seh_pushreg %rbp > pushq %rdi > .seh_pushreg %rdi > pushq %rsi > .seh_pushreg %rsi > pushq %rbx > .seh_pushreg %rbx > .seh_endprologue > movl %ecx, %edi > movq %rdx, %rbp > movq %r8, %rdx > movq %r9, %rsi > movq 88(%rsp), %r9 > movq 96(%rsp), %rcx > movl 104(%rsp), %r13d > movl $0, %r12d > xorpd %xmm2, %xmm2 > movapd %xmm2, %xmm3 > testl %r13d, %r13d > jg .L15 > jmp .L2 > .L12: > movl (%rsi,%rbx,4), %eax > movl 4(%rsi,%rbx,4), %r8d > cmpl %r8d, %eax > jge .L10 > movapd %xmm2, %xmm0 > .L6: > movslq %eax, %r10 > movslq (%r9,%r10,4), %r11 > movsd (%rcx,%r11,8), %xmm1 > mulsd (%rdx,%r10,8), %xmm1 > addsd %xmm1, %xmm0 > addl $1, %eax > cmpl %r8d, %eax > jne .L6 > jmp .L5 > .L10: > movapd %xmm3, %xmm0 > .L5: > movsd %xmm0, 0(%rbp,%rbx,8) > addq $1, %rbx > cmpl %ebx, %edi > jg .L12 > .L8: > addl $1, %r12d > cmpl %r13d, %r12d > jne .L15 > jmp .L2 > .L15: > movl $0, %ebx > testl %edi, %edi > jg .L12 > .p2align 4,,4 > jmp .L8 > .L2: > popq %rbx > popq %rsi > popq %rdi > popq %rbp > popq %r12 > popq %r13 > ret > .seh_endproc > > == No this is just scalar code. It does not use mulpd/addpd rather mulsd/addsd. It is scalar instruction for simd. -- Click OK to continue...
[toc] | [prev] | [next] | [standalone]
| From | Ramine <ramine@1.1> |
|---|---|
| Date | 2014-06-29 11:14 -0700 |
| Message-ID | <lopafm$jkg$1@dont-email.me> |
| In reply to | #2538 |
On 6/29/2014 8:00 AM, Melzzzzz wrote: > On Sun, 29 Jun 2014 10:47:51 -0700 > Ramine <ramine@1.1> wrote: > >> >> Hello, >> >> Look for example at the assembler source code of >> SparseCompRow.c that you will find inside the >> scimark2 source code benchmark, look carefully >> at the follwing assembler code, and you will notice >> that it is auto-vectorizing, i am using just using >> -O1 level optimization and the -S flag to generate >> the assembler code, here it is: >> >> >> === >> >> >> .file "SparseCompRow.c" >> .text >> .globl SparseCompRow_num_flops >> .def SparseCompRow_num_flops; .scl >> 2; .type 32; .endef .seh_proc >> SparseCompRow_num_flops SparseCompRow_num_flops: >> .seh_endprologue >> movl %edx, %eax >> sarl $31, %edx >> idivl %ecx >> imull %eax, %ecx >> cvtsi2sd %ecx, %xmm0 >> addsd %xmm0, %xmm0 >> cvtsi2sd %r8d, %xmm1 >> mulsd %xmm1, %xmm0 >> ret >> .seh_endproc >> .globl SparseCompRow_matmult >> .def SparseCompRow_matmult; .scl >> 2; .type 32; .endef .seh_proc >> SparseCompRow_matmult SparseCompRow_matmult: >> pushq %r13 >> .seh_pushreg %r13 >> pushq %r12 >> .seh_pushreg %r12 >> pushq %rbp >> .seh_pushreg %rbp >> pushq %rdi >> .seh_pushreg %rdi >> pushq %rsi >> .seh_pushreg %rsi >> pushq %rbx >> .seh_pushreg %rbx >> .seh_endprologue >> movl %ecx, %edi >> movq %rdx, %rbp >> movq %r8, %rdx >> movq %r9, %rsi >> movq 88(%rsp), %r9 >> movq 96(%rsp), %rcx >> movl 104(%rsp), %r13d >> movl $0, %r12d >> xorpd %xmm2, %xmm2 >> movapd %xmm2, %xmm3 >> testl %r13d, %r13d >> jg .L15 >> jmp .L2 >> .L12: >> movl (%rsi,%rbx,4), %eax >> movl 4(%rsi,%rbx,4), %r8d >> cmpl %r8d, %eax >> jge .L10 >> movapd %xmm2, %xmm0 >> .L6: >> movslq %eax, %r10 >> movslq (%r9,%r10,4), %r11 >> movsd (%rcx,%r11,8), %xmm1 >> mulsd (%rdx,%r10,8), %xmm1 >> addsd %xmm1, %xmm0 >> addl $1, %eax >> cmpl %r8d, %eax >> jne .L6 >> jmp .L5 >> .L10: >> movapd %xmm3, %xmm0 >> .L5: >> movsd %xmm0, 0(%rbp,%rbx,8) >> addq $1, %rbx >> cmpl %ebx, %edi >> jg .L12 >> .L8: >> addl $1, %r12d >> cmpl %r13d, %r12d >> jne .L15 >> jmp .L2 >> .L15: >> movl $0, %ebx >> testl %edi, %edi >> jg .L12 >> .p2align 4,,4 >> jmp .L8 >> .L2: >> popq %rbx >> popq %rsi >> popq %rdi >> popq %rbp >> popq %r12 >> popq %r13 >> ret >> .seh_endproc >> >> == > > No this is just scalar code. It does not use mulpd/addpd rather > mulsd/addsd. > It is scalar instruction for simd. > > > Sorry it's an error on my part, mulsd do in fact Multiplies the low double-precision floating-point value in the source operand (second operand) by the low double-precision floating-point value in the destination operand (first operand), and stores the double-precision floating-point result in the destination operand. so it is not auto-vectorizing the code... so that's amazing. how then gcc is generating 2X times faster code than Delphi and FreePascal ? i think that Delphi and FreePascal are slow compilers. Thank you, Amine Moulay Ramdane.
[toc] | [prev] | [next] | [standalone]
| From | Melzzzzz <mel@zzzzz.com> |
|---|---|
| Date | 2014-06-29 17:20 +0200 |
| Message-ID | <loparl$9ui$5@news.albasani.net> |
| In reply to | #2539 |
On Sun, 29 Jun 2014 11:14:31 -0700 Ramine <ramine@1.1> wrote: so it is not auto-vectorizing the code... so that's amazing. > how then gcc is generating 2X times faster code than Delphi and > FreePascal ? i think that Delphi and FreePascal are slow compilers. gcc - C and C++ is being developed at fast pace and improved constantly. Many many commits. Hack gcc is even better than intel icc now, I think. bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -Ofast -mavx *.c -o scimark2 -lm bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 ** ** ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark ** ** for details. (Results can be submitted to pozo@nist.gov) ** ** ** Using 2.00 seconds min time per kenel. Composite Score: 2147.56 FFT Mflops: 1810.43 (N=1024) SOR Mflops: 1848.91 (100 x 100) MonteCarlo: Mflops: 631.23 Sparse matmult Mflops: 2253.35 (N=1000, nz=5000) LU Mflops: 4193.91 (M=100, N=100) real 0m28.592s user 0m28.516s sys 0m0.050s bmaxa@maxa:~/examples/forth/sci$ icc -fast -mavx *.c -o scimark2 -lm bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 ** ** ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark ** ** for details. (Results can be submitted to pozo@nist.gov) ** ** ** Using 2.00 seconds min time per kenel. Composite Score: 2079.37 FFT Mflops: 1707.05 (N=1024) SOR Mflops: 1527.84 (100 x 100) MonteCarlo: Mflops: 1203.64 Sparse matmult Mflops: 1695.46 (N=1000, nz=5000) LU Mflops: 4262.85 (M=100, N=100) real 0m27.656s user 0m27.601s sys 0m0.034s (Intel seems that autovectorize MonteCarlo, which gcc don't) but overal intel is slower. > > > > Thank you, > Amine Moulay Ramdane. > > > > -- Click OK to continue...
[toc] | [prev] | [next] | [standalone]
| From | Ramine <ramine@1.1> |
|---|---|
| Date | 2014-06-29 11:23 -0700 |
| Message-ID | <lopb0g$nel$1@dont-email.me> |
| In reply to | #2539 |
Hello, Sorry Melzzzzz i have forgot: Look at the assembler code it is using many: xorpd and movpd that's also gives much better performance , no ? Thank you, Amine Moulay Rmdane.
[toc] | [prev] | [next] | [standalone]
| From | Melzzzzz <mel@zzzzz.com> |
|---|---|
| Date | 2014-06-29 17:31 +0200 |
| Message-ID | <lopbg0$9ui$6@news.albasani.net> |
| In reply to | #2541 |
On Sun, 29 Jun 2014 11:23:30 -0700 Ramine <ramine@1.1> wrote: > > Hello, > > Sorry Melzzzzz i have forgot: > > Look at the assembler code it is using many: > > xorpd and movpd > > that's also gives much better performance , no ? > there is no xorsd and movapd is only instruction to move from xmm register to xmm register. So, no. It does not give better perfromance. > > Thank you, > Amine Moulay Rmdane. > > -- Click OK to continue...
[toc] | [prev] | [next] | [standalone]
| From | Ramine <ramine@1.1> |
|---|---|
| Date | 2014-06-29 11:32 -0700 |
| Message-ID | <lopbhu$rca$1@dont-email.me> |
| In reply to | #2541 |
So when you say that gcc is not vectorising, that's not completly true cause if you look at the source code it is using many xorpd and movapd and that's also auto-vectorization... I think that's also why GCC is faster than FreePascal and Delphi on scimark2. Thank you, Amine Moulay Ramdane.
[toc] | [prev] | [next] | [standalone]
| From | Ramine <ramine@1.1> |
|---|---|
| Date | 2014-06-29 11:38 -0700 |
| Message-ID | <lopbsn$sui$1@dont-email.me> |
| In reply to | #2543 |
On 6/29/2014 11:32 AM, Ramine wrote: > > So when you say that gcc is not vectorising, that's not completly true > cause if you look at the source code it is using many xorpd and movapd > and that's also auto-vectorization... > > I think that's also why GCC is faster than FreePascal and Delphi on > scimark2. > > Thank you, > Amine Moulay Ramdane. > But i don't think so, those xorpd and movapd can not give double performance, so i think that gcc is a faster compiler than Delphi and FreePascal even without auto-vectorization. Thank you, Amine Moulay Ramdane.
[toc] | [prev] | [next] | [standalone]
| From | Ramine <ramine@1.1> |
|---|---|
| Date | 2014-06-29 10:41 -0700 |
| Message-ID | <lop8hu$59l$1@dont-email.me> |
| In reply to | #2531 |
On 6/28/2014 7:10 PM, Melzzzzz wrote: > On Sat, 28 Jun 2014 15:51:59 -0700 > Ramine <ramine@1.1> wrote: > >> >> Hello, >> >> I have looked at your matrix multiplication using >> SSE2 and AVX that you have posted on assembler.x86 forum, >> i just wanted to tell you that you don't need to write >> it in assembler, cause GCC do auto-vectorize the floating >> point calculation and it auto-vectorize also the Matrix >> multiplication even at -O1 optimization... > > as I see it it don't.... > > > and GCC do >> auto-vectorize beautifully since the GCC auto-vectorization >> have given a good results on scimark2 benchmark too. > > Care to share example? You can download the following scimark2 and compile it with the newest version of tdm-gcc with just -O1 level optimization and -S to generate the assembler code and after that look at the assembler code and you will see that it is auto-vectorizing the code: http://math.nist.gov/scimark2/ Thank you, Amine Moulay Ramdane. > >> >> >> Thank you, >> Amine Moulay Ramdane. >> >> >> >> > > >
[toc] | [prev] | [next] | [standalone]
| From | Melzzzzz <mel@zzzzz.invalid> |
|---|---|
| Date | 2014-06-29 16:46 +0200 |
| Message-ID | <lop8rt$834$3@solani.org> |
| In reply to | #2534 |
On Sun, 29 Jun 2014 10:41:30 -0700 Ramine <ramine@1.1> wrote: > On 6/28/2014 7:10 PM, Melzzzzz wrote: > > On Sat, 28 Jun 2014 15:51:59 -0700 > > Ramine <ramine@1.1> wrote: > > > >> > >> Hello, > >> > >> I have looked at your matrix multiplication using > >> SSE2 and AVX that you have posted on assembler.x86 forum, > >> i just wanted to tell you that you don't need to write > >> it in assembler, cause GCC do auto-vectorize the floating > >> point calculation and it auto-vectorize also the Matrix > >> multiplication even at -O1 optimization... > > > > as I see it it don't.... > > > > > > and GCC do > >> auto-vectorize beautifully since the GCC auto-vectorization > >> have given a good results on scimark2 benchmark too. > > > > Care to share example? > > > You can download the following scimark2 and compile it with > the newest version of tdm-gcc with just -O1 level optimization > and -S to generate the assembler code and after that look at the > assembler code and you will see that it is auto-vectorizing the code: > > http://math.nist.gov/scimark2/ I have this benchmark long ago and it seems to autovectorize onli LU and that at -O3. bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O2 *.c -o scimark2 -lm bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 ** ** ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark ** ** for details. (Results can be submitted to pozo@nist.gov) ** ** ** Using 2.00 seconds min time per kenel. Composite Score: 1739.60 FFT Mflops: 1782.75 (N=1024) SOR Mflops: 1305.62 (100 x 100) MonteCarlo: Mflops: 618.10 Sparse matmult Mflops: 2255.16 (N=1000, nz=5000) LU Mflops: 2736.35 (M=100, N=100) real 0m33.408s user 0m33.331s sys 0m0.052s bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O3 *.c -o scimark2 -lm bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 ** ** ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark ** ** for details. (Results can be submitted to pozo@nist.gov) ** ** ** Using 2.00 seconds min time per kenel. Composite Score: 2102.19 FFT Mflops: 1855.88 (N=1024) SOR Mflops: 1848.64 (100 x 100) MonteCarlo: Mflops: 621.80 Sparse matmult Mflops: 2275.65 (N=1000, nz=5000) LU Mflops: 3908.98 (M=100, N=100) real 0m28.933s user 0m28.849s sys 0m0.050s > > > Thank you, > Amine Moulay Ramdane. > > > > > >> > >> > >> Thank you, > >> Amine Moulay Ramdane. > >> > >> > >> > >> > > > > > > > -- Click OK to continue...
[toc] | [prev] | [next] | [standalone]
| From | Ramine <ramine@1.1> |
|---|---|
| Date | 2014-06-29 10:54 -0700 |
| Message-ID | <lop9ag$ai3$1@dont-email.me> |
| In reply to | #2535 |
Melzzzzz wrote: > I have this benchmark long ago and it seems to autovectorize onli LU > and that at -O3. No i don't think so, i am using the newest version of tdm-gcc and it is auto-vectorizing the FFT and the MonteCarlo and Sparse MatMult and LU too. Thank you, Amine Moulay Ramdane. On 6/29/2014 7:46 AM, Melzzzzz wrote: > On Sun, 29 Jun 2014 10:41:30 -0700 > Ramine <ramine@1.1> wrote: > >> On 6/28/2014 7:10 PM, Melzzzzz wrote: >>> On Sat, 28 Jun 2014 15:51:59 -0700 >>> Ramine <ramine@1.1> wrote: >>> >>>> >>>> Hello, >>>> >>>> I have looked at your matrix multiplication using >>>> SSE2 and AVX that you have posted on assembler.x86 forum, >>>> i just wanted to tell you that you don't need to write >>>> it in assembler, cause GCC do auto-vectorize the floating >>>> point calculation and it auto-vectorize also the Matrix >>>> multiplication even at -O1 optimization... >>> >>> as I see it it don't.... >>> >>> >>> and GCC do >>>> auto-vectorize beautifully since the GCC auto-vectorization >>>> have given a good results on scimark2 benchmark too. >>> >>> Care to share example? >> >> >> You can download the following scimark2 and compile it with >> the newest version of tdm-gcc with just -O1 level optimization >> and -S to generate the assembler code and after that look at the >> assembler code and you will see that it is auto-vectorizing the code: >> >> http://math.nist.gov/scimark2/ > > I have this benchmark long ago and it seems to autovectorize onli LU > and that at -O3. > bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O2 *.c -o scimark2 -lm > bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 > ** ** > ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark ** > ** for details. (Results can be submitted to pozo@nist.gov) ** > ** ** > Using 2.00 seconds min time per kenel. > Composite Score: 1739.60 > FFT Mflops: 1782.75 (N=1024) > SOR Mflops: 1305.62 (100 x 100) > MonteCarlo: Mflops: 618.10 > Sparse matmult Mflops: 2255.16 (N=1000, nz=5000) > LU Mflops: 2736.35 (M=100, N=100) > > real 0m33.408s > user 0m33.331s > sys 0m0.052s > bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O3 *.c -o scimark2 -lm > bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 > ** ** > ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark ** > ** for details. (Results can be submitted to pozo@nist.gov) ** > ** ** > Using 2.00 seconds min time per kenel. > Composite Score: 2102.19 > FFT Mflops: 1855.88 (N=1024) > SOR Mflops: 1848.64 (100 x 100) > MonteCarlo: Mflops: 621.80 > Sparse matmult Mflops: 2275.65 (N=1000, nz=5000) > LU Mflops: 3908.98 (M=100, N=100) > > real 0m28.933s > user 0m28.849s > sys 0m0.050s > > >> >> >> Thank you, >> Amine Moulay Ramdane. >> >> >>> >>>> >>>> >>>> Thank you, >>>> Amine Moulay Ramdane. >>>> >>>> >>>> >>>> >>> >>> >>> >> > > >
[toc] | [prev] | [standalone]
Back to top | Article view | comp.programming.threads
csiph-web