Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > comp.programming > #4638
| From | Melzzzzz <mel@zzzzz.invalid> |
|---|---|
| Newsgroups | comp.programming, comp.programming.threads |
| Subject | Re: To Melzzzzz |
| Date | 2014-06-29 16:46 +0200 |
| Organization | None |
| Message-ID | <lop8rt$834$3@solani.org> (permalink) |
| References | <longtm$bkp$1@dont-email.me> <lonsia$ea2$1@solani.org> <lop8hu$59l$1@dont-email.me> |
Cross-posted to 2 groups.
On Sun, 29 Jun 2014 10:41:30 -0700 Ramine <ramine@1.1> wrote: > On 6/28/2014 7:10 PM, Melzzzzz wrote: > > On Sat, 28 Jun 2014 15:51:59 -0700 > > Ramine <ramine@1.1> wrote: > > > >> > >> Hello, > >> > >> I have looked at your matrix multiplication using > >> SSE2 and AVX that you have posted on assembler.x86 forum, > >> i just wanted to tell you that you don't need to write > >> it in assembler, cause GCC do auto-vectorize the floating > >> point calculation and it auto-vectorize also the Matrix > >> multiplication even at -O1 optimization... > > > > as I see it it don't.... > > > > > > and GCC do > >> auto-vectorize beautifully since the GCC auto-vectorization > >> have given a good results on scimark2 benchmark too. > > > > Care to share example? > > > You can download the following scimark2 and compile it with > the newest version of tdm-gcc with just -O1 level optimization > and -S to generate the assembler code and after that look at the > assembler code and you will see that it is auto-vectorizing the code: > > http://math.nist.gov/scimark2/ I have this benchmark long ago and it seems to autovectorize onli LU and that at -O3. bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O2 *.c -o scimark2 -lm bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 ** ** ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark ** ** for details. (Results can be submitted to pozo@nist.gov) ** ** ** Using 2.00 seconds min time per kenel. Composite Score: 1739.60 FFT Mflops: 1782.75 (N=1024) SOR Mflops: 1305.62 (100 x 100) MonteCarlo: Mflops: 618.10 Sparse matmult Mflops: 2255.16 (N=1000, nz=5000) LU Mflops: 2736.35 (M=100, N=100) real 0m33.408s user 0m33.331s sys 0m0.052s bmaxa@maxa:~/examples/forth/sci$ gcc-trunk -O3 *.c -o scimark2 -lm bmaxa@maxa:~/examples/forth/sci$ time ./scimark2 ** ** ** SciMark2 Numeric Benchmark, see http://math.nist.gov/scimark ** ** for details. (Results can be submitted to pozo@nist.gov) ** ** ** Using 2.00 seconds min time per kenel. Composite Score: 2102.19 FFT Mflops: 1855.88 (N=1024) SOR Mflops: 1848.64 (100 x 100) MonteCarlo: Mflops: 621.80 Sparse matmult Mflops: 2275.65 (N=1000, nz=5000) LU Mflops: 3908.98 (M=100, N=100) real 0m28.933s user 0m28.849s sys 0m0.050s > > > Thank you, > Amine Moulay Ramdane. > > > > > >> > >> > >> Thank you, > >> Amine Moulay Ramdane. > >> > >> > >> > >> > > > > > > > -- Click OK to continue...
Back to comp.programming | Previous | Next — Previous in thread | Next in thread | Find similar | Unroll thread
To Melzzzzz Ramine <ramine@1.1> - 2014-06-28 15:51 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.invalid> - 2014-06-29 04:10 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:33 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 16:39 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:47 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:00 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:14 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:20 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:23 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.com> - 2014-06-29 17:31 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:32 -0700
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 11:38 -0700
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:41 -0700
Re: To Melzzzzz Melzzzzz <mel@zzzzz.invalid> - 2014-06-29 16:46 +0200
Re: To Melzzzzz Ramine <ramine@1.1> - 2014-06-29 10:54 -0700
csiph-web