Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.programming.threads > #2536
| Path | csiph.com!usenet.pasdenom.info!aioe.org!news.swapon.de!eternal-september.org!feeder.eternal-september.org!news.eternal-september.org!.POSTED!not-for-mail |
|---|---|
| From | Ramine <ramine@1.1> |
| Newsgroups | comp.programming, comp.programming.threads |
| Subject | Re: To Melzzzzz |
| Date | Sun, 29 Jun 2014 10:47:51 -0700 |
| Organization | A noiseless patient Spider |
| Lines | 187 |
| Message-ID | <lop8tr$7ra$1@dont-email.me> (permalink) |
| References | <longtm$bkp$1@dont-email.me> <lonsia$ea2$1@solani.org> <lop839$ute$1@dont-email.me> <lop8fk$9ui$3@news.albasani.net> |
| Mime-Version | 1.0 |
| Content-Type | text/plain; charset=ISO-8859-1; format=flowed |
| Content-Transfer-Encoding | 7bit |
| Injection-Date | Sun, 29 Jun 2014 14:47:23 +0000 (UTC) |
| Injection-Info | mx05.eternal-september.org; posting-host="518d70aeead21e0a1a15f1c91bb07e09"; logging-data="8042"; mail-complaints-to="abuse@eternal-september.org"; posting-account="U2FsdGVkX1/HwSBxac+HJ65uuqy+yp22" |
| User-Agent | Mozilla/5.0 (Windows NT 6.0; WOW64; rv:24.0) Gecko/20100101 Thunderbird/24.6.0 |
| In-Reply-To | <lop8fk$9ui$3@news.albasani.net> |
| Cancel-Lock | sha1:cZF28ApYCs3e1FpdtmVANqF8edo= |
| Xref | csiph.com comp.programming:4639 comp.programming.threads:2536 |
Cross-posted to 2 groups.
Show key headers only | View raw
Hello, Look for example at the assembler source code of SparseCompRow.c that you will find inside the scimark2 source code benchmark, look carefully at the follwing assembler code, and you will notice that it is auto-vectorizing, i am using just using -O1 level optimization and the -S flag to generate the assembler code, here it is: === .file "SparseCompRow.c" .text .globl SparseCompRow_num_flops .def SparseCompRow_num_flops; .scl 2; .type 32; .endef .seh_proc SparseCompRow_num_flops SparseCompRow_num_flops: .seh_endprologue movl %edx, %eax sarl $31, %edx idivl %ecx imull %eax, %ecx cvtsi2sd %ecx, %xmm0 addsd %xmm0, %xmm0 cvtsi2sd %r8d, %xmm1 mulsd %xmm1, %xmm0 ret .seh_endproc .globl SparseCompRow_matmult .def SparseCompRow_matmult; .scl 2; .type 32; .endef .seh_proc SparseCompRow_matmult SparseCompRow_matmult: pushq %r13 .seh_pushreg %r13 pushq %r12 .seh_pushreg %r12 pushq %rbp .seh_pushreg %rbp pushq %rdi .seh_pushreg %rdi pushq %rsi .seh_pushreg %rsi pushq %rbx .seh_pushreg %rbx .seh_endprologue movl %ecx, %edi movq %rdx, %rbp movq %r8, %rdx movq %r9, %rsi movq 88(%rsp), %r9 movq 96(%rsp), %rcx movl 104(%rsp), %r13d movl $0, %r12d xorpd %xmm2, %xmm2 movapd %xmm2, %xmm3 testl %r13d, %r13d jg .L15 jmp .L2 .L12: movl (%rsi,%rbx,4), %eax movl 4(%rsi,%rbx,4), %r8d cmpl %r8d, %eax jge .L10 movapd %xmm2, %xmm0 .L6: movslq %eax, %r10 movslq (%r9,%r10,4), %r11 movsd (%rcx,%r11,8), %xmm1 mulsd (%rdx,%r10,8), %xmm1 addsd %xmm1, %xmm0 addl $1, %eax cmpl %r8d, %eax jne .L6 jmp .L5 .L10: movapd %xmm3, %xmm0 .L5: movsd %xmm0, 0(%rbp,%rbx,8) addq $1, %rbx cmpl %ebx, %edi jg .L12 .L8: addl $1, %r12d cmpl %r13d, %r12d jne .L15 jmp .L2 .L15: movl $0, %ebx testl %edi, %edi jg .L12 .p2align 4,,4 jmp .L8 .L2: popq %rbx popq %rsi popq %rdi popq %rbp popq %r12 popq %r13 ret .seh_endproc == Thank you, Amine Moulay Ramdane. On 6/29/2014 7:39 AM, Melzzzzz wrote: > On Sun, 29 Jun 2014 10:33:41 -0700 > Ramine <ramine@1.1> wrote: > >> On 6/28/2014 7:10 PM, Melzzzzz wrote: >>> On Sat, 28 Jun 2014 15:51:59 -0700 >>> Ramine <ramine@1.1> wrote: >>> >>>> >>>> Hello, >>>> >>>> I have looked at your matrix multiplication using >>>> SSE2 and AVX that you have posted on assembler.x86 forum, >>>> i just wanted to tell you that you don't need to write >>>> it in assembler, cause GCC do auto-vectorize the floating >>>> point calculation and it auto-vectorize also the Matrix >>>> multiplication even at -O1 optimization... >>> >>> as I see it it don't.... >>> >> >> >> I am using the newest version of tdm-gcc and auto-vectorization is >> working even at -O1 level optimization. >> >> >> You can download tdb-gcc from here and try it yourself: >> >> http://tdm-gcc.tdragon.net/ > > That is based on gcc 4.8.1 , Im using 4.9 and trunk versions > and there is no vectorization of matrix multiplication > (even for 4x4 matrices). > Care to show compiler output of assembly and c source code > so I can see what I am missing? > >> >> >> >> >> Thank you, >> Amine Moulay Ramdane. >> >> >> >> >> >> >> >>> >>> and GCC do >>>> auto-vectorize beautifully since the GCC auto-vectorization >>>> have given a good results on scimark2 benchmark too. >>> >>> Care to share example? >>> >>>> >>>> >>>> Thank you, >>>> Amine Moulay Ramdane. >>>> >>>> >>>> >>>> >>> >>> >>> >> > > >
Back to comp.programming.threads | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
To Melzzzzz Ramine <ramine@1.1> - 2014-06-28 15:51 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.invalid> - 2014-06-29 04:10 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:33 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 16:39 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:47 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:00 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:14 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:20 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:23 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:31 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:32 -0700
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:38 -0700
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:41 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.invalid> - 2014-06-29 16:46 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:54 -0700
csiph-web