Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.programming.threads > #2536

Re: To Melzzzzz

From Ramine <ramine@1.1>
Newsgroups comp.programming, comp.programming.threads
Subject Re: To Melzzzzz
Date 2014-06-29 10:47 -0700
Organization A noiseless patient Spider
Message-ID <lop8tr$7ra$1@dont-email.me> (permalink)
References <longtm$bkp$1@dont-email.me> <lonsia$ea2$1@solani.org> <lop839$ute$1@dont-email.me> <lop8fk$9ui$3@news.albasani.net>

Cross-posted to 2 groups.

Show all headers | View raw


Hello,

Look for example at the assembler source code of
SparseCompRow.c that you will find inside the
scimark2 source code benchmark, look carefully
at the follwing assembler code, and you will notice
that it is auto-vectorizing, i am using just using
-O1 level optimization and the -S flag to generate
the assembler code, here it is:


===


.file	"SparseCompRow.c"
	.text
	.globl	SparseCompRow_num_flops
	.def	SparseCompRow_num_flops;	.scl	2;	.type	32;	.endef
	.seh_proc	SparseCompRow_num_flops
SparseCompRow_num_flops:
	.seh_endprologue
	movl	%edx, %eax
	sarl	$31, %edx
	idivl	%ecx
	imull	%eax, %ecx
	cvtsi2sd	%ecx, %xmm0
	addsd	%xmm0, %xmm0
	cvtsi2sd	%r8d, %xmm1
	mulsd	%xmm1, %xmm0
	ret
	.seh_endproc
	.globl	SparseCompRow_matmult
	.def	SparseCompRow_matmult;	.scl	2;	.type	32;	.endef
	.seh_proc	SparseCompRow_matmult
SparseCompRow_matmult:
	pushq	%r13
	.seh_pushreg	%r13
	pushq	%r12
	.seh_pushreg	%r12
	pushq	%rbp
	.seh_pushreg	%rbp
	pushq	%rdi
	.seh_pushreg	%rdi
	pushq	%rsi
	.seh_pushreg	%rsi
	pushq	%rbx
	.seh_pushreg	%rbx
	.seh_endprologue
	movl	%ecx, %edi
	movq	%rdx, %rbp
	movq	%r8, %rdx
	movq	%r9, %rsi
	movq	88(%rsp), %r9
	movq	96(%rsp), %rcx
	movl	104(%rsp), %r13d
	movl	$0, %r12d
	xorpd	%xmm2, %xmm2
	movapd	%xmm2, %xmm3
	testl	%r13d, %r13d
	jg	.L15
	jmp	.L2
.L12:
	movl	(%rsi,%rbx,4), %eax
	movl	4(%rsi,%rbx,4), %r8d
	cmpl	%r8d, %eax
	jge	.L10
	movapd	%xmm2, %xmm0
.L6:
	movslq	%eax, %r10
	movslq	(%r9,%r10,4), %r11
	movsd	(%rcx,%r11,8), %xmm1
	mulsd	(%rdx,%r10,8), %xmm1
	addsd	%xmm1, %xmm0
	addl	$1, %eax
	cmpl	%r8d, %eax
	jne	.L6
	jmp	.L5
.L10:
	movapd	%xmm3, %xmm0
.L5:
	movsd	%xmm0, 0(%rbp,%rbx,8)
	addq	$1, %rbx
	cmpl	%ebx, %edi
	jg	.L12
.L8:
	addl	$1, %r12d
	cmpl	%r13d, %r12d
	jne	.L15
	jmp	.L2
.L15:
	movl	$0, %ebx
	testl	%edi, %edi
	jg	.L12
	.p2align 4,,4
	jmp	.L8
.L2:
	popq	%rbx
	popq	%rsi
	popq	%rdi
	popq	%rbp
	popq	%r12
	popq	%r13
	ret
	.seh_endproc

==



Thank you,
Amine Moulay Ramdane.



On 6/29/2014 7:39 AM, Melzzzzz wrote:
> On Sun, 29 Jun 2014 10:33:41 -0700
> Ramine <ramine@1.1> wrote:
>
>> On 6/28/2014 7:10 PM, Melzzzzz wrote:
>>> On Sat, 28 Jun 2014 15:51:59 -0700
>>> Ramine <ramine@1.1> wrote:
>>>
>>>>
>>>> Hello,
>>>>
>>>> I have looked at your matrix multiplication using
>>>> SSE2 and AVX that you have posted on assembler.x86 forum,
>>>> i just wanted to tell you that you don't need to write
>>>> it in assembler, cause GCC do auto-vectorize the floating
>>>> point calculation and it auto-vectorize also the Matrix
>>>> multiplication even at -O1 optimization...
>>>
>>> as I see it it don't....
>>>
>>
>>
>> I am using the newest version of tdm-gcc and auto-vectorization is
>> working even at -O1 level optimization.
>>
>>
>> You can download tdb-gcc from here and try it yourself:
>>
>> http://tdm-gcc.tdragon.net/
>
> That is based on gcc 4.8.1 , Im using 4.9 and trunk versions
> and there is no vectorization of matrix multiplication
> (even for 4x4 matrices).
> Care to show compiler output of assembly and c source code
> so I can see what I am missing?
>
>>
>>
>>
>>
>> Thank you,
>> Amine Moulay Ramdane.
>>
>>
>>
>>
>>
>>
>>
>>>
>>>    and GCC do
>>>> auto-vectorize beautifully since the GCC auto-vectorization
>>>> have  given a good results on scimark2 benchmark too.
>>>
>>> Care to share example?
>>>
>>>>
>>>>
>>>> Thank you,
>>>> Amine Moulay Ramdane.
>>>>
>>>>
>>>>
>>>>
>>>
>>>
>>>
>>
>
>
>

Back to comp.programming.threads | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread


Thread

To Melzzzzz Ramine <ramine@1.1> - 2014-06-28 15:51 -0700
  Re: To Melzzzzz Melzzzzz <mel@zzzzz.invalid> - 2014-06-29 04:10 +0200
    Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:33 -0700
      Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 16:39 +0200
        Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:47 -0700
          Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:00 +0200
            Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:14 -0700
              Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:20 +0200
              Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:23 -0700
                Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:31 +0200
                Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:32 -0700
                Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:38 -0700
    Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:41 -0700
      Re: To Melzzzzz Melzzzzz <mel@zzzzz.invalid> - 2014-06-29 16:46 +0200
        Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:54 -0700

csiph-web