Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.programming.threads > #2529 > unrolled thread

To Melzzzzz

Started byRamine <ramine@1.1>
First post2014-06-28 15:51 -0700
Last post2014-06-29 10:54 -0700
Articles 15 — 3 participants

Back to article view | Back to comp.programming.threads


Contents

  To Melzzzzz Ramine <ramine@1.1> - 2014-06-28 15:51 -0700
    Re: To Melzzzzz Melzzzzz <mel@zzzzz.invalid> - 2014-06-29 04:10 +0200
      Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:33 -0700
        Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 16:39 +0200
          Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:47 -0700
            Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:00 +0200
              Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:14 -0700
                Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:20 +0200
                Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:23 -0700
                  Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:31 +0200
                  Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:32 -0700
                    Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:38 -0700
      Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:41 -0700
        Re: To Melzzzzz Melzzzzz <mel@zzzzz.invalid> - 2014-06-29 16:46 +0200
          Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:54 -0700

#2529 — To Melzzzzz

FromRamine <ramine@1.1>
Date2014-06-28 15:51 -0700
SubjectTo Melzzzzz
Message-ID<longtm$bkp$1@dont-email.me>
Hello,

I have looked at your matrix multiplication using
SSE2 and AVX that you have posted on assembler.x86 forum,
i just wanted to tell you that you don't need to write
it in assembler, cause GCC do auto-vectorize the floating
point calculation and it auto-vectorize also the Matrix
multiplication even at -O1 optimization... and GCC do
auto-vectorize beautifully since the GCC auto-vectorization
have  given a good results on scimark2 benchmark too.


Thank you,
Amine Moulay Ramdane.



[toc] | [next] | [standalone]


#2531

FromMelzzzzz <mel@zzzzz.invalid>
Date2014-06-29 04:10 +0200
Message-ID<lonsia$ea2$1@solani.org>
In reply to#2529
On Sat, 28 Jun 2014 15:51:59 -0700
Ramine <ramine@1.1> wrote:

> 
> Hello,
> 
> I have looked at your matrix multiplication using
> SSE2 and AVX that you have posted on assembler.x86 forum,
> i just wanted to tell you that you don't need to write
> it in assembler, cause GCC do auto-vectorize the floating
> point calculation and it auto-vectorize also the Matrix
> multiplication even at -O1 optimization...

as I see it it don't....


 and GCC do
> auto-vectorize beautifully since the GCC auto-vectorization
> have  given a good results on scimark2 benchmark too.

Care to share example?

> 
> 
> Thank you,
> Amine Moulay Ramdane.
> 
> 
> 
> 



-- 
Click OK to continue...

[toc] | [prev] | [next] | [standalone]


#2532

FromRamine <ramine@1.1>
Date2014-06-29 10:33 -0700
Message-ID<lop839$ute$1@dont-email.me>
In reply to#2531
On 6/28/2014 7:10 PM, Melzzzzz wrote:
> On Sat, 28 Jun 2014 15:51:59 -0700
> Ramine <ramine@1.1> wrote:
>
>>
>> Hello,
>>
>> I have looked at your matrix multiplication using
>> SSE2 and AVX that you have posted on assembler.x86 forum,
>> i just wanted to tell you that you don't need to write
>> it in assembler, cause GCC do auto-vectorize the floating
>> point calculation and it auto-vectorize also the Matrix
>> multiplication even at -O1 optimization...
>
> as I see it it don't....
>


I am using the newest version of tdm-gcc and auto-vectorization is 
working even at -O1 level optimization.


You can download tdb-gcc from here and try it yourself:

http://tdm-gcc.tdragon.net/




Thank you,
Amine Moulay Ramdane.







>
>   and GCC do
>> auto-vectorize beautifully since the GCC auto-vectorization
>> have  given a good results on scimark2 benchmark too.
>
> Care to share example?
>
>>
>>
>> Thank you,
>> Amine Moulay Ramdane.
>>
>>
>>
>>
>
>
>

[toc] | [prev] | [next] | [standalone]


#2533

FromMelzzzzz <mel@zzzzz.com>
Date2014-06-29 16:39 +0200
Message-ID<lop8fk$9ui$3@news.albasani.net>
In reply to#2532
On Sun, 29 Jun 2014 10:33:41 -0700
Ramine <ramine@1.1> wrote:

> On 6/28/2014 7:10 PM, Melzzzzz wrote:
> > On Sat, 28 Jun 2014 15:51:59 -0700
> > Ramine <ramine@1.1> wrote:
> >
> >>
> >> Hello,
> >>
> >> I have looked at your matrix multiplication using
> >> SSE2 and AVX that you have posted on assembler.x86 forum,
> >> i just wanted to tell you that you don't need to write
> >> it in assembler, cause GCC do auto-vectorize the floating
> >> point calculation and it auto-vectorize also the Matrix
> >> multiplication even at -O1 optimization...
> >
> > as I see it it don't....
> >
> 
> 
> I am using the newest version of tdm-gcc and auto-vectorization is 
> working even at -O1 level optimization.
> 
> 
> You can download tdb-gcc from here and try it yourself:
> 
> http://tdm-gcc.tdragon.net/

That is based on gcc 4.8.1 , Im using 4.9 and trunk versions
and there is no vectorization of matrix multiplication
(even for 4x4 matrices).
Care to show compiler output of assembly and c source code
so I can see what I am missing?

> 
> 
> 
> 
> Thank you,
> Amine Moulay Ramdane.
> 
> 
> 
> 
> 
> 
> 
> >
> >   and GCC do
> >> auto-vectorize beautifully since the GCC auto-vectorization
> >> have  given a good results on scimark2 benchmark too.
> >
> > Care to share example?
> >
> >>
> >>
> >> Thank you,
> >> Amine Moulay Ramdane.
> >>
> >>
> >>
> >>
> >
> >
> >
> 



-- 
Click OK to continue...

[toc] | [prev] | [next] | [standalone]


#2536

FromRamine <ramine@1.1>
Date2014-06-29 10:47 -0700
Message-ID<lop8tr$7ra$1@dont-email.me>
In reply to#2533
Hello,

Look for example at the assembler source code of
SparseCompRow.c that you will find inside the
scimark2 source code benchmark, look carefully
at the follwing assembler code, and you will notice
that it is auto-vectorizing, i am using just using
-O1 level optimization and the -S flag to generate
the assembler code, here it is:


===


.file	"SparseCompRow.c"
	.text
	.globl	SparseCompRow_num_flops
	.def	SparseCompRow_num_flops;	.scl	2;	.type	32;	.endef
	.seh_proc	SparseCompRow_num_flops
SparseCompRow_num_flops:
	.seh_endprologue
	movl	%edx, %eax
	sarl	$31, %edx
	idivl	%ecx
	imull	%eax, %ecx
	cvtsi2sd	%ecx, %xmm0
	addsd	%xmm0, %xmm0
	cvtsi2sd	%r8d, %xmm1
	mulsd	%xmm1, %xmm0
	ret
	.seh_endproc
	.globl	SparseCompRow_matmult
	.def	SparseCompRow_matmult;	.scl	2;	.type	32;	.endef
	.seh_proc	SparseCompRow_matmult
SparseCompRow_matmult:
	pushq	%r13
	.seh_pushreg	%r13
	pushq	%r12
	.seh_pushreg	%r12
	pushq	%rbp
	.seh_pushreg	%rbp
	pushq	%rdi
	.seh_pushreg	%rdi
	pushq	%rsi
	.seh_pushreg	%rsi
	pushq	%rbx
	.seh_pushreg	%rbx
	.seh_endprologue
	movl	%ecx, %edi
	movq	%rdx, %rbp
	movq	%r8, %rdx
	movq	%r9, %rsi
	movq	88(%rsp), %r9
	movq	96(%rsp), %rcx
	movl	104(%rsp), %r13d
	movl	$0, %r12d
	xorpd	%xmm2, %xmm2
	movapd	%xmm2, %xmm3
	testl	%r13d, %r13d
	jg	.L15
	jmp	.L2
.L12:
	movl	(%rsi,%rbx,4), %eax
	movl	4(%rsi,%rbx,4), %r8d
	cmpl	%r8d, %eax
	jge	.L10
	movapd	%xmm2, %xmm0
.L6:
	movslq	%eax, %r10
	movslq	(%r9,%r10,4), %r11
	movsd	(%rcx,%r11,8), %xmm1
	mulsd	(%rdx,%r10,8), %xmm1
	addsd	%xmm1, %xmm0
	addl	$1, %eax
	cmpl	%r8d, %eax
	jne	.L6
	jmp	.L5
.L10:
	movapd	%xmm3, %xmm0
.L5:
	movsd	%xmm0, 0(%rbp,%rbx,8)
	addq	$1, %rbx
	cmpl	%ebx, %edi
	jg	.L12
.L8:
	addl	$1, %r12d
	cmpl	%r13d, %r12d
	jne	.L15
	jmp	.L2
.L15:
	movl	$0, %ebx
	testl	%edi, %edi
	jg	.L12
	.p2align 4,,4
	jmp	.L8
.L2:
	popq	%rbx
	popq	%rsi
	popq	%rdi
	popq	%rbp
	popq	%r12
	popq	%r13
	ret
	.seh_endproc

==



Thank you,
Amine Moulay Ramdane.



On 6/29/2014 7:39 AM, Melzzzzz wrote:
> On Sun, 29 Jun 2014 10:33:41 -0700
> Ramine <ramine@1.1> wrote:
>
>> On 6/28/2014 7:10 PM, Melzzzzz wrote:
>>> On Sat, 28 Jun 2014 15:51:59 -0700
>>> Ramine <ramine@1.1> wrote:
>>>
>>>>
>>>> Hello,
>>>>
>>>> I have looked at your matrix multiplication using
>>>> SSE2 and AVX that you have posted on assembler.x86 forum,
>>>> i just wanted to tell you that you don't need to write
>>>> it in assembler, cause GCC do auto-vectorize the floating
>>>> point calculation and it auto-vectorize also the Matrix
>>>> multiplication even at -O1 optimization...
>>>
>>> as I see it it don't....
>>>
>>
>>
>> I am using the newest version of tdm-gcc and auto-vectorization is
>> working even at -O1 level optimization.
>>
>>
>> You can download tdb-gcc from here and try it yourself:
>>
>> http://tdm-gcc.tdragon.net/
>
> That is based on gcc 4.8.1 , Im using 4.9 and trunk versions
> and there is no vectorization of matrix multiplication
> (even for 4x4 matrices).
> Care to show compiler output of assembly and c source code
> so I can see what I am missing?
>
>>
>>
>>
>>
>> Thank you,
>> Amine Moulay Ramdane.
>>
>>
>>
>>
>>
>>
>>
>>>
>>>    and GCC do
>>>> auto-vectorize beautifully since the GCC auto-vectorization
>>>> have  given a good results on scimark2 benchmark too.
>>>
>>> Care to share example?
>>>
>>>>
>>>>
>>>> Thank you,
>>>> Amine Moulay Ramdane.
>>>>
>>>>
>>>>
>>>>
>>>
>>>
>>>
>>
>
>
>

[toc] | [prev] | [next] | [standalone]


#2538

FromMelzzzzz <mel@zzzzz.com>
Date2014-06-29 17:00 +0200
Message-ID<lop9ml$9ui$4@news.albasani.net>
In reply to#2536
On Sun, 29 Jun 2014 10:47:51 -0700
Ramine <ramine@1.1> wrote:

> 
> Hello,
> 
> Look for example at the assembler source code of
> SparseCompRow.c that you will find inside the
> scimark2 source code benchmark, look carefully
> at the follwing assembler code, and you will notice
> that it is auto-vectorizing, i am using just using
> -O1 level optimization and the -S flag to generate
> the assembler code, here it is:
> 
> 
> ===
> 
> 
> .file	"SparseCompRow.c"
> 	.text
> 	.globl	SparseCompRow_num_flops
> 	.def	SparseCompRow_num_flops;	.scl
> 2;	.type	32;	.endef .seh_proc
> SparseCompRow_num_flops SparseCompRow_num_flops:
> 	.seh_endprologue
> 	movl	%edx, %eax
> 	sarl	$31, %edx
> 	idivl	%ecx
> 	imull	%eax, %ecx
> 	cvtsi2sd	%ecx, %xmm0
> 	addsd	%xmm0, %xmm0
> 	cvtsi2sd	%r8d, %xmm1
> 	mulsd	%xmm1, %xmm0
> 	ret
> 	.seh_endproc
> 	.globl	SparseCompRow_matmult
> 	.def	SparseCompRow_matmult;	.scl
> 2;	.type	32;	.endef .seh_proc
> SparseCompRow_matmult SparseCompRow_matmult:
> 	pushq	%r13
> 	.seh_pushreg	%r13
> 	pushq	%r12
> 	.seh_pushreg	%r12
> 	pushq	%rbp
> 	.seh_pushreg	%rbp
> 	pushq	%rdi
> 	.seh_pushreg	%rdi
> 	pushq	%rsi
> 	.seh_pushreg	%rsi
> 	pushq	%rbx
> 	.seh_pushreg	%rbx
> 	.seh_endprologue
> 	movl	%ecx, %edi
> 	movq	%rdx, %rbp
> 	movq	%r8, %rdx
> 	movq	%r9, %rsi
> 	movq	88(%rsp), %r9
> 	movq	96(%rsp), %rcx
> 	movl	104(%rsp), %r13d
> 	movl	$0, %r12d
> 	xorpd	%xmm2, %xmm2
> 	movapd	%xmm2, %xmm3
> 	testl	%r13d, %r13d
> 	jg	.L15
> 	jmp	.L2
> .L12:
> 	movl	(%rsi,%rbx,4), %eax
> 	movl	4(%rsi,%rbx,4), %r8d
> 	cmpl	%r8d, %eax
> 	jge	.L10
> 	movapd	%xmm2, %xmm0
> .L6:
> 	movslq	%eax, %r10
> 	movslq	(%r9,%r10,4), %r11
> 	movsd	(%rcx,%r11,8), %xmm1
> 	mulsd	(%rdx,%r10,8), %xmm1
> 	addsd	%xmm1, %xmm0
> 	addl	$1, %eax
> 	cmpl	%r8d, %eax
> 	jne	.L6
> 	jmp	.L5
> .L10:
> 	movapd	%xmm3, %xmm0
> .L5:
> 	movsd	%xmm0, 0(%rbp,%rbx,8)
> 	addq	$1, %rbx
> 	cmpl	%ebx, %edi
> 	jg	.L12
> .L8:
> 	addl	$1, %r12d
> 	cmpl	%r13d, %r12d
> 	jne	.L15
> 	jmp	.L2
> .L15:
> 	movl	$0, %ebx
> 	testl	%edi, %edi
> 	jg	.L12
> 	.p2align 4,,4
> 	jmp	.L8
> .L2:
> 	popq	%rbx
> 	popq	%rsi
> 	popq	%rdi
> 	popq	%rbp
> 	popq	%r12
> 	popq	%r13
> 	ret
> 	.seh_endproc
> 
> ==

No this is just scalar code. It does not use mulpd/addpd rather
mulsd/addsd.
It is scalar instruction for simd.



-- 
Click OK to continue...

[toc] | [prev] | [next] | [standalone]


#2539

FromRamine <ramine@1.1>
Date2014-06-29 11:14 -0700
Message-ID<lopafm$jkg$1@dont-email.me>
In reply to#2538
On 6/29/2014 8:00 AM, Melzzzzz wrote:
> On Sun, 29 Jun 2014 10:47:51 -0700
> Ramine <ramine@1.1> wrote:
>
>>
>> Hello,
>>
>> Look for example at the assembler source code of
>> SparseCompRow.c that you will find inside the
>> scimark2 source code benchmark, look carefully
>> at the follwing assembler code, and you will notice
>> that it is auto-vectorizing, i am using just using
>> -O1 level optimization and the -S flag to generate
>> the assembler code, here it is:
>>
>>
>> ===
>>
>>
>> .file	"SparseCompRow.c"
>> 	.text
>> 	.globl	SparseCompRow_num_flops
>> 	.def	SparseCompRow_num_flops;	.scl
>> 2;	.type	32;	.endef .seh_proc
>> SparseCompRow_num_flops SparseCompRow_num_flops:
>> 	.seh_endprologue
>> 	movl	%edx, %eax
>> 	sarl	$31, %edx
>> 	idivl	%ecx
>> 	imull	%eax, %ecx
>> 	cvtsi2sd	%ecx, %xmm0
>> 	addsd	%xmm0, %xmm0
>> 	cvtsi2sd	%r8d, %xmm1
>> 	mulsd	%xmm1, %xmm0
>> 	ret
>> 	.seh_endproc
>> 	.globl	SparseCompRow_matmult
>> 	.def	SparseCompRow_matmult;	.scl
>> 2;	.type	32;	.endef .seh_proc
>> SparseCompRow_matmult SparseCompRow_matmult:
>> 	pushq	%r13
>> 	.seh_pushreg	%r13
>> 	pushq	%r12
>> 	.seh_pushreg	%r12
>> 	pushq	%rbp
>> 	.seh_pushreg	%rbp
>> 	pushq	%rdi
>> 	.seh_pushreg	%rdi
>> 	pushq	%rsi
>> 	.seh_pushreg	%rsi
>> 	pushq	%rbx
>> 	.seh_pushreg	%rbx
>> 	.seh_endprologue
>> 	movl	%ecx, %edi
>> 	movq	%rdx, %rbp
>> 	movq	%r8, %rdx
>> 	movq	%r9, %rsi
>> 	movq	88(%rsp), %r9
>> 	movq	96(%rsp), %rcx
>> 	movl	104(%rsp), %r13d
>> 	movl	$0, %r12d
>> 	xorpd	%xmm2, %xmm2
>> 	movapd	%xmm2, %xmm3
>> 	testl	%r13d, %r13d
>> 	jg	.L15
>> 	jmp	.L2
>> .L12:
>> 	movl	(%rsi,%rbx,4), %eax
>> 	movl	4(%rsi,%rbx,4), %r8d
>> 	cmpl	%r8d, %eax
>> 	jge	.L10
>> 	movapd	%xmm2, %xmm0
>> .L6:
>> 	movslq	%eax, %r10
>> 	movslq	(%r9,%r10,4), %r11
>> 	movsd	(%rcx,%r11,8), %xmm1
>> 	mulsd	(%rdx,%r10,8), %xmm1
>> 	addsd	%xmm1, %xmm0
>> 	addl	$1, %eax
>> 	cmpl	%r8d, %eax
>> 	jne	.L6
>> 	jmp	.L5
>> .L10:
>> 	movapd	%xmm3, %xmm0
>> .L5:
>> 	movsd	%xmm0, 0(%rbp,%rbx,8)
>> 	addq	$1, %rbx
>> 	cmpl	%ebx, %edi
>> 	jg	.L12
>> .L8:
>> 	addl	$1, %r12d
>> 	cmpl	%r13d, %r12d
>> 	jne	.L15
>> 	jmp	.L2
>> .L15:
>> 	movl	$0, %ebx
>> 	testl	%edi, %edi
>> 	jg	.L12
>> 	.p2align 4,,4
>> 	jmp	.L8
>> .L2:
>> 	popq	%rbx
>> 	popq	%rsi
>> 	popq	%rdi
>> 	popq	%rbp
>> 	popq	%r12
>> 	popq	%r13
>> 	ret
>> 	.seh_endproc
>>
>> ==
>
> No this is just scalar code. It does not use mulpd/addpd rather
> mulsd/addsd.
> It is scalar instruction for simd.
>
>
>


Sorry it's an error on my part, mulsd do in fact
Multiplies the low double-precision floating-point value in the source 
operand (second operand) by the low double-precision floating-point 
value in the destination operand (first operand), and stores the 
double-precision floating-point result in the destination operand. so it 
is not auto-vectorizing the code... so that's amazing. how then gcc is 
generating 2X times faster code than Delphi and FreePascal ? i think
that Delphi and FreePascal are slow compilers.



Thank you,
Amine Moulay Ramdane.



[toc] | [prev] | [next] | [standalone]


#2540

FromMelzzzzz <mel@zzzzz.com>
Date2014-06-29 17:20 +0200
Message-ID<loparl$9ui$5@news.albasani.net>
In reply to#2539
On Sun, 29 Jun 2014 11:14:31 -0700
Ramine <ramine@1.1> wrote:

 so it is not auto-vectorizing the code... so that's amazing.
> how then gcc is generating 2X times faster code than Delphi and
> FreePascal ? i think that Delphi and FreePascal are slow compilers.

gcc - C and C++ is being developed at fast pace and improved constantly.
Many many commits. 
Hack gcc is even better than intel icc now, I think.

bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -Ofast -mavx *.c -o scimark2
-lm bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 
**                                                              **
** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark **
** for details. (Results can be submitted to pozo@nist.gov)     **
**                                                              **
Using       2.00 seconds min time per kenel.
Composite Score:         2147.56
FFT             Mflops:  1810.43    (N=1024)
SOR             Mflops:  1848.91    (100 x 100)
MonteCarlo:     Mflops:   631.23
Sparse matmult  Mflops:  2253.35    (N=1000, nz=5000)
LU              Mflops:  4193.91    (M=100, N=100)

real	0m28.592s
user	0m28.516s
sys	0m0.050s
bmaxa@maxa:~/examples/forth/sci$ icc -fast -mavx *.c -o scimark2 -lm
bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 
**                                                              **
** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark **
** for details. (Results can be submitted to pozo@nist.gov)     **
**                                                              **
Using       2.00 seconds min time per kenel.
Composite Score:         2079.37
FFT             Mflops:  1707.05    (N=1024)
SOR             Mflops:  1527.84    (100 x 100)
MonteCarlo:     Mflops:  1203.64
Sparse matmult  Mflops:  1695.46    (N=1000, nz=5000)
LU              Mflops:  4262.85    (M=100, N=100)

real	0m27.656s
user	0m27.601s
sys	0m0.034s

(Intel seems that autovectorize MonteCarlo, which gcc don't)
but overal intel is slower.

> 
> 
> 
> Thank you,
> Amine Moulay Ramdane.
> 
> 
> 
> 



-- 
Click OK to continue...

[toc] | [prev] | [next] | [standalone]


#2541

FromRamine <ramine@1.1>
Date2014-06-29 11:23 -0700
Message-ID<lopb0g$nel$1@dont-email.me>
In reply to#2539
Hello,

Sorry Melzzzzz i have forgot:

Look at the assembler code it is using many:

xorpd and movpd

that's also gives much better performance , no ?


Thank you,
Amine Moulay Rmdane.

[toc] | [prev] | [next] | [standalone]


#2542

FromMelzzzzz <mel@zzzzz.com>
Date2014-06-29 17:31 +0200
Message-ID<lopbg0$9ui$6@news.albasani.net>
In reply to#2541
On Sun, 29 Jun 2014 11:23:30 -0700
Ramine <ramine@1.1> wrote:

> 
> Hello,
> 
> Sorry Melzzzzz i have forgot:
> 
> Look at the assembler code it is using many:
> 
> xorpd and movpd
> 
> that's also gives much better performance , no ?
> 

there is no xorsd and movapd is only instruction to
move from xmm register to xmm register. So, no.
It does not give better perfromance.


> 
> Thank you,
> Amine Moulay Rmdane.
> 
> 



-- 
Click OK to continue...

[toc] | [prev] | [next] | [standalone]


#2543

FromRamine <ramine@1.1>
Date2014-06-29 11:32 -0700
Message-ID<lopbhu$rca$1@dont-email.me>
In reply to#2541
So when you say that gcc is not vectorising, that's not completly true
cause if you look at the source code it is using many xorpd and movapd
and that's also auto-vectorization...

I think that's also why GCC is faster than FreePascal and Delphi on 
scimark2.

Thank you,
Amine Moulay Ramdane.

[toc] | [prev] | [next] | [standalone]


#2544

FromRamine <ramine@1.1>
Date2014-06-29 11:38 -0700
Message-ID<lopbsn$sui$1@dont-email.me>
In reply to#2543
On 6/29/2014 11:32 AM, Ramine wrote:
>
> So when you say that gcc is not vectorising, that's not completly true
> cause if you look at the source code it is using many xorpd and movapd
> and that's also auto-vectorization...
>
> I think that's also why GCC is faster than FreePascal and Delphi on
> scimark2.
>
> Thank you,
> Amine Moulay Ramdane.
>



But i don't think so, those xorpd and movapd can not give
double performance, so i think that gcc is a faster compiler
than Delphi and FreePascal even without auto-vectorization.



Thank you,
Amine Moulay Ramdane.

[toc] | [prev] | [next] | [standalone]


#2534

FromRamine <ramine@1.1>
Date2014-06-29 10:41 -0700
Message-ID<lop8hu$59l$1@dont-email.me>
In reply to#2531
On 6/28/2014 7:10 PM, Melzzzzz wrote:
> On Sat, 28 Jun 2014 15:51:59 -0700
> Ramine <ramine@1.1> wrote:
>
>>
>> Hello,
>>
>> I have looked at your matrix multiplication using
>> SSE2 and AVX that you have posted on assembler.x86 forum,
>> i just wanted to tell you that you don't need to write
>> it in assembler, cause GCC do auto-vectorize the floating
>> point calculation and it auto-vectorize also the Matrix
>> multiplication even at -O1 optimization...
>
> as I see it it don't....
>
>
>   and GCC do
>> auto-vectorize beautifully since the GCC auto-vectorization
>> have  given a good results on scimark2 benchmark too.
>
> Care to share example?


You can download the following scimark2 and compile it with
the newest version of tdm-gcc with just -O1 level optimization
and -S to generate the assembler code and after that look at the 
assembler code and you will see that it is auto-vectorizing the code:

http://math.nist.gov/scimark2/


Thank you,
Amine Moulay Ramdane.


>
>>
>>
>> Thank you,
>> Amine Moulay Ramdane.
>>
>>
>>
>>
>
>
>

[toc] | [prev] | [next] | [standalone]


#2535

FromMelzzzzz <mel@zzzzz.invalid>
Date2014-06-29 16:46 +0200
Message-ID<lop8rt$834$3@solani.org>
In reply to#2534
On Sun, 29 Jun 2014 10:41:30 -0700
Ramine <ramine@1.1> wrote:

> On 6/28/2014 7:10 PM, Melzzzzz wrote:
> > On Sat, 28 Jun 2014 15:51:59 -0700
> > Ramine <ramine@1.1> wrote:
> >
> >>
> >> Hello,
> >>
> >> I have looked at your matrix multiplication using
> >> SSE2 and AVX that you have posted on assembler.x86 forum,
> >> i just wanted to tell you that you don't need to write
> >> it in assembler, cause GCC do auto-vectorize the floating
> >> point calculation and it auto-vectorize also the Matrix
> >> multiplication even at -O1 optimization...
> >
> > as I see it it don't....
> >
> >
> >   and GCC do
> >> auto-vectorize beautifully since the GCC auto-vectorization
> >> have  given a good results on scimark2 benchmark too.
> >
> > Care to share example?
> 
> 
> You can download the following scimark2 and compile it with
> the newest version of tdm-gcc with just -O1 level optimization
> and -S to generate the assembler code and after that look at the 
> assembler code and you will see that it is auto-vectorizing the code:
> 
> http://math.nist.gov/scimark2/

I have this benchmark long ago and it seems to autovectorize onli LU
and that at -O3.
bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O2 *.c -o scimark2 -lm
bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 
**                                                              **
** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark **
** for details. (Results can be submitted to pozo@nist.gov)     **
**                                                              **
Using       2.00 seconds min time per kenel.
Composite Score:         1739.60
FFT             Mflops:  1782.75    (N=1024)
SOR             Mflops:  1305.62    (100 x 100)
MonteCarlo:     Mflops:   618.10
Sparse matmult  Mflops:  2255.16    (N=1000, nz=5000)
LU              Mflops:  2736.35    (M=100, N=100)

real	0m33.408s
user	0m33.331s
sys	0m0.052s
bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O3 *.c -o scimark2 -lm
bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 
**                                                              **
** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark **
** for details. (Results can be submitted to pozo@nist.gov)     **
**                                                              **
Using       2.00 seconds min time per kenel.
Composite Score:         2102.19
FFT             Mflops:  1855.88    (N=1024)
SOR             Mflops:  1848.64    (100 x 100)
MonteCarlo:     Mflops:   621.80
Sparse matmult  Mflops:  2275.65    (N=1000, nz=5000)
LU              Mflops:  3908.98    (M=100, N=100)

real	0m28.933s
user	0m28.849s
sys	0m0.050s


> 
> 
> Thank you,
> Amine Moulay Ramdane.
> 
> 
> >
> >>
> >>
> >> Thank you,
> >> Amine Moulay Ramdane.
> >>
> >>
> >>
> >>
> >
> >
> >
> 



-- 
Click OK to continue...

[toc] | [prev] | [next] | [standalone]


#2537

FromRamine <ramine@1.1>
Date2014-06-29 10:54 -0700
Message-ID<lop9ag$ai3$1@dont-email.me>
In reply to#2535
Melzzzzz wrote:
 > I have this benchmark long ago and it seems to autovectorize onli LU
 > and that at -O3.


No i don't think so, i am using the newest version of tdm-gcc
and it is auto-vectorizing  the FFT and the MonteCarlo and Sparse 
MatMult and LU too.



Thank you,
Amine Moulay Ramdane.



On 6/29/2014 7:46 AM, Melzzzzz wrote:
> On Sun, 29 Jun 2014 10:41:30 -0700
> Ramine <ramine@1.1> wrote:
>
>> On 6/28/2014 7:10 PM, Melzzzzz wrote:
>>> On Sat, 28 Jun 2014 15:51:59 -0700
>>> Ramine <ramine@1.1> wrote:
>>>
>>>>
>>>> Hello,
>>>>
>>>> I have looked at your matrix multiplication using
>>>> SSE2 and AVX that you have posted on assembler.x86 forum,
>>>> i just wanted to tell you that you don't need to write
>>>> it in assembler, cause GCC do auto-vectorize the floating
>>>> point calculation and it auto-vectorize also the Matrix
>>>> multiplication even at -O1 optimization...
>>>
>>> as I see it it don't....
>>>
>>>
>>>    and GCC do
>>>> auto-vectorize beautifully since the GCC auto-vectorization
>>>> have  given a good results on scimark2 benchmark too.
>>>
>>> Care to share example?
>>
>>
>> You can download the following scimark2 and compile it with
>> the newest version of tdm-gcc with just -O1 level optimization
>> and -S to generate the assembler code and after that look at the
>> assembler code and you will see that it is auto-vectorizing the code:
>>
>> http://math.nist.gov/scimark2/
>
> I have this benchmark long ago and it seems to autovectorize onli LU
> and that at -O3.
> bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O2 *.c -o scimark2 -lm
> bmaxa@maxa:~/examples/forth/sci$ time ./scimark2
> **                                                              **
> ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark **
> ** for details. (Results can be submitted to pozo@nist.gov)     **
> **                                                              **
> Using       2.00 seconds min time per kenel.
> Composite Score:         1739.60
> FFT             Mflops:  1782.75    (N=1024)
> SOR             Mflops:  1305.62    (100 x 100)
> MonteCarlo:     Mflops:   618.10
> Sparse matmult  Mflops:  2255.16    (N=1000, nz=5000)
> LU              Mflops:  2736.35    (M=100, N=100)
>
> real	0m33.408s
> user	0m33.331s
> sys	0m0.052s
> bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O3 *.c -o scimark2 -lm
> bmaxa@maxa:~/examples/forth/sci$ time ./scimark2
> **                                                              **
> ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark **
> ** for details. (Results can be submitted to pozo@nist.gov)     **
> **                                                              **
> Using       2.00 seconds min time per kenel.
> Composite Score:         2102.19
> FFT             Mflops:  1855.88    (N=1024)
> SOR             Mflops:  1848.64    (100 x 100)
> MonteCarlo:     Mflops:   621.80
> Sparse matmult  Mflops:  2275.65    (N=1000, nz=5000)
> LU              Mflops:  3908.98    (M=100, N=100)
>
> real	0m28.933s
> user	0m28.849s
> sys	0m0.050s
>
>
>>
>>
>> Thank you,
>> Amine Moulay Ramdane.
>>
>>
>>>
>>>>
>>>>
>>>> Thank you,
>>>> Amine Moulay Ramdane.
>>>>
>>>>
>>>>
>>>>
>>>
>>>
>>>
>>
>
>
>

[toc] | [prev] | [standalone]


Back to top | Article view | comp.programming.threads


csiph-web