Path: csiph.com!v102.xanadu-bbs.net!xanadu-bbs.net!feeder.erje.net!eu.feeder.erje.net!newsfeed.datemas.de!weretis.net!feeder1.news.weretis.net!news.solani.org!.POSTED!not-for-mail From: Melzzzzz Newsgroups: comp.programming,comp.programming.threads Subject: Re: To Melzzzzz Date: Sun, 29 Jun 2014 16:46:21 +0200 Organization: None Lines: 97 Message-ID: References: Mime-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit X-Trace: solani.org 1404053181 8292 eJwNydEVADEEBMCWyFmkHLzVfwm5+R18rj5hDjcsNs5sX/OWmwyuukCYc4LNqgrWtOUfMhqLByuHEfQ= (29 Jun 2014 14:46:21 GMT) X-Complaints-To: abuse@news.solani.org NNTP-Posting-Date: Sun, 29 Jun 2014 14:46:21 +0000 (UTC) X-User-ID: eJwFwYEBgDAIA7CXZKxFzsEC/59gAqdRcQleLDb1jOSx4/6aDkQPT83mKIwDWvfXdsoKUz8oGhHA X-Newsreader: Claws Mail 3.10.1 (GTK+ 2.24.23; x86_64-pc-linux-gnu) Cancel-Lock: sha1:jBn+uulnBLgZkH9yY8iuZOqqsXM= X-NNTP-Posting-Host: eJwNxMEBADEEBMCWCLu0E0H/JdzNY2BUvnCCjsWe3doaMKEThLj2tNetCqa55FOxP63jfeMDKI4RJQ== Xref: csiph.com comp.programming:4638 comp.programming.threads:2535 On Sun, 29 Jun 2014 10:41:30 -0700 Ramine wrote: > On 6/28/2014 7:10 PM, Melzzzzz wrote: > > On Sat, 28 Jun 2014 15:51:59 -0700 > > Ramine wrote: > > > >> > >> Hello, > >> > >> I have looked at your matrix multiplication using > >> SSE2 and AVX that you have posted on assembler.x86 forum, > >> i just wanted to tell you that you don't need to write > >> it in assembler, cause GCC do auto-vectorize the floating > >> point calculation and it auto-vectorize also the Matrix > >> multiplication even at -O1 optimization... > > > > as I see it it don't.... > > > > > > and GCC do > >> auto-vectorize beautifully since the GCC auto-vectorization > >> have given a good results on scimark2 benchmark too. > > > > Care to share example? > > > You can download the following scimark2 and compile it with > the newest version of tdm-gcc with just -O1 level optimization > and -S to generate the assembler code and after that look at the > assembler code and you will see that it is auto-vectorizing the code: > > http://math.nist.gov/scimark2/ I have this benchmark long ago and it seems to autovectorize onli LU and that at -O3. bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O2 *.c -o scimark2 -lm bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 ** ** ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark ** ** for details. (Results can be submitted to pozo@nist.gov) ** ** ** Using 2.00 seconds min time per kenel. Composite Score: 1739.60 FFT Mflops: 1782.75 (N=1024) SOR Mflops: 1305.62 (100 x 100) MonteCarlo: Mflops: 618.10 Sparse matmult Mflops: 2255.16 (N=1000, nz=5000) LU Mflops: 2736.35 (M=100, N=100) real 0m33.408s user 0m33.331s sys 0m0.052s bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O3 *.c -o scimark2 -lm bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 ** ** ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark ** ** for details. (Results can be submitted to pozo@nist.gov) ** ** ** Using 2.00 seconds min time per kenel. Composite Score: 2102.19 FFT Mflops: 1855.88 (N=1024) SOR Mflops: 1848.64 (100 x 100) MonteCarlo: Mflops: 621.80 Sparse matmult Mflops: 2275.65 (N=1000, nz=5000) LU Mflops: 3908.98 (M=100, N=100) real 0m28.933s user 0m28.849s sys 0m0.050s > > > Thank you, > Amine Moulay Ramdane. > > > > > >> > >> > >> Thank you, > >> Amine Moulay Ramdane. > >> > >> > >> > >> > > > > > > > -- Click OK to continue...