Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1381350 > unrolled thread
| Started by | Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> |
|---|---|
| First post | 2016-04-18 10:00 +0200 |
| Last post | 2016-04-27 11:10 +0200 |
| Articles | 10 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: zram: per-cpu compression streams Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-04-18 10:00 +0200
Re: zram: per-cpu compression streams Minchan Kim <minchan@kernel.org> - 2016-04-19 10:00 +0200
Re: zram: per-cpu compression streams Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-04-19 10:10 +0200
Re: zram: per-cpu compression streams Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-04-26 13:30 +0200
Re: zram: per-cpu compression streams Minchan Kim <minchan@kernel.org> - 2016-04-27 09:30 +0200
Re: zram: per-cpu compression streams Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-04-27 09:50 +0200
Re: zram: per-cpu compression streams Minchan Kim <minchan@kernel.org> - 2016-04-27 10:00 +0200
Re: zram: per-cpu compression streams Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-04-27 10:10 +0200
Re: zram: per-cpu compression streams Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-04-27 11:00 +0200
Re: zram: per-cpu compression streams Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-04-27 11:10 +0200
| From | Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> |
|---|---|
| Date | 2016-04-18 10:00 +0200 |
| Subject | Re: zram: per-cpu compression streams |
| Message-ID | <rpcga-1JC-23@gated-at.bofh.it> |
Hello Minchan,
sorry, it took me so long to return back to testing.
I collected extended stats (perf), just like you requested.
- 3G zram, lzo; 4 CPU x86_64 box.
- fio with perf stat
4 streams 8 streams per-cpu
===========================================================
#jobs1
READ: 2520.1MB/s 2566.5MB/s 2491.5MB/s
READ: 2102.7MB/s 2104.2MB/s 2091.3MB/s
WRITE: 1355.1MB/s 1320.2MB/s 1378.9MB/s
WRITE: 1103.5MB/s 1097.2MB/s 1122.5MB/s
READ: 434013KB/s 435153KB/s 439961KB/s
WRITE: 433969KB/s 435109KB/s 439917KB/s
READ: 403166KB/s 405139KB/s 403373KB/s
WRITE: 403223KB/s 405197KB/s 403430KB/s
#jobs2
READ: 7958.6MB/s 8105.6MB/s 8073.7MB/s
READ: 6864.9MB/s 6989.8MB/s 7021.8MB/s
WRITE: 2438.1MB/s 2346.9MB/s 3400.2MB/s
WRITE: 1994.2MB/s 1990.3MB/s 2941.2MB/s
READ: 981504KB/s 973906KB/s 1018.8MB/s
WRITE: 981659KB/s 974060KB/s 1018.1MB/s
READ: 937021KB/s 938976KB/s 987250KB/s
WRITE: 934878KB/s 936830KB/s 984993KB/s
#jobs3
READ: 13280MB/s 13553MB/s 13553MB/s
READ: 11534MB/s 11785MB/s 11755MB/s
WRITE: 3456.9MB/s 3469.9MB/s 4810.3MB/s
WRITE: 3029.6MB/s 3031.6MB/s 4264.8MB/s
READ: 1363.8MB/s 1362.6MB/s 1448.9MB/s
WRITE: 1361.9MB/s 1360.7MB/s 1446.9MB/s
READ: 1309.4MB/s 1310.6MB/s 1397.5MB/s
WRITE: 1307.4MB/s 1308.5MB/s 1395.3MB/s
#jobs4
READ: 20244MB/s 20177MB/s 20344MB/s
READ: 17886MB/s 17913MB/s 17835MB/s
WRITE: 4071.6MB/s 4046.1MB/s 6370.2MB/s
WRITE: 3608.9MB/s 3576.3MB/s 5785.4MB/s
READ: 1824.3MB/s 1821.6MB/s 1997.5MB/s
WRITE: 1819.8MB/s 1817.4MB/s 1992.5MB/s
READ: 1765.7MB/s 1768.3MB/s 1937.3MB/s
WRITE: 1767.5MB/s 1769.1MB/s 1939.2MB/s
#jobs5
READ: 18663MB/s 18986MB/s 18823MB/s
READ: 16659MB/s 16605MB/s 16954MB/s
WRITE: 3912.4MB/s 3888.7MB/s 6126.9MB/s
WRITE: 3506.4MB/s 3442.5MB/s 5519.3MB/s
READ: 1798.2MB/s 1746.5MB/s 1935.8MB/s
WRITE: 1792.7MB/s 1740.7MB/s 1929.1MB/s
READ: 1727.6MB/s 1658.2MB/s 1917.3MB/s
WRITE: 1726.5MB/s 1657.2MB/s 1916.6MB/s
#jobs6
READ: 21017MB/s 20922MB/s 21162MB/s
READ: 19022MB/s 19140MB/s 18770MB/s
WRITE: 3968.2MB/s 4037.7MB/s 6620.8MB/s
WRITE: 3643.5MB/s 3590.2MB/s 6027.5MB/s
READ: 1871.8MB/s 1880.5MB/s 2049.9MB/s
WRITE: 1867.8MB/s 1877.2MB/s 2046.2MB/s
READ: 1755.8MB/s 1710.3MB/s 1964.7MB/s
WRITE: 1750.5MB/s 1705.9MB/s 1958.8MB/s
#jobs7
READ: 21103MB/s 20677MB/s 21482MB/s
READ: 18522MB/s 18379MB/s 19443MB/s
WRITE: 4022.5MB/s 4067.4MB/s 6755.9MB/s
WRITE: 3691.7MB/s 3695.5MB/s 5925.6MB/s
READ: 1841.5MB/s 1933.9MB/s 2090.5MB/s
WRITE: 1842.7MB/s 1935.3MB/s 2091.9MB/s
READ: 1832.4MB/s 1856.4MB/s 1971.5MB/s
WRITE: 1822.3MB/s 1846.2MB/s 1960.6MB/s
#jobs8
READ: 20463MB/s 20194MB/s 20862MB/s
READ: 18178MB/s 17978MB/s 18299MB/s
WRITE: 4085.9MB/s 4060.2MB/s 7023.8MB/s
WRITE: 3776.3MB/s 3737.9MB/s 6278.2MB/s
READ: 1957.6MB/s 1944.4MB/s 2109.5MB/s
WRITE: 1959.2MB/s 1946.2MB/s 2111.4MB/s
READ: 1900.6MB/s 1885.7MB/s 2082.1MB/s
WRITE: 1896.2MB/s 1881.4MB/s 2078.3MB/s
#jobs9
READ: 19692MB/s 19734MB/s 19334MB/s
READ: 17678MB/s 18249MB/s 17666MB/s
WRITE: 4004.7MB/s 4064.8MB/s 6990.7MB/s
WRITE: 3724.7MB/s 3772.1MB/s 6193.6MB/s
READ: 1953.7MB/s 1967.3MB/s 2105.6MB/s
WRITE: 1953.4MB/s 1966.7MB/s 2104.1MB/s
READ: 1860.4MB/s 1897.4MB/s 2068.5MB/s
WRITE: 1858.9MB/s 1895.9MB/s 2066.8MB/s
#jobs10
READ: 19730MB/s 19579MB/s 19492MB/s
READ: 18028MB/s 18018MB/s 18221MB/s
WRITE: 4027.3MB/s 4090.6MB/s 7020.1MB/s
WRITE: 3810.5MB/s 3846.8MB/s 6426.8MB/s
READ: 1956.1MB/s 1994.6MB/s 2145.2MB/s
WRITE: 1955.9MB/s 1993.5MB/s 2144.8MB/s
READ: 1852.8MB/s 1911.6MB/s 2075.8MB/s
WRITE: 1855.7MB/s 1914.6MB/s 2078.1MB/s
perf stat
4 streams 8 streams per-cpu
====================================================================================================================
jobs1 ( ) ( ) ( )
stalled-cycles-frontend 23,174,811,209 ( 38.21%) 23,220,254,188 ( 38.25%) 23,061,406,918 ( 38.34%)
stalled-cycles-backend 11,514,174,638 ( 18.98%) 11,696,722,657 ( 19.27%) 11,370,852,810 ( 18.90%)
instructions 73,925,005,782 ( 1.22) 73,903,177,632 ( 1.22) 73,507,201,037 ( 1.22)
branches 14,455,124,835 ( 756.063) 14,455,184,779 ( 755.281) 14,378,599,509 ( 758.546)
branch-misses 69,801,336 ( 0.48%) 80,225,529 ( 0.55%) 72,044,726 ( 0.50%)
jobs2 ( ) ( ) ( )
stalled-cycles-frontend 49,912,741,782 ( 46.11%) 50,101,189,290 ( 45.95%) 32,874,195,633 ( 35.11%)
stalled-cycles-backend 27,080,366,230 ( 25.02%) 27,949,970,232 ( 25.63%) 16,461,222,706 ( 17.58%)
instructions 122,831,629,690 ( 1.13) 122,919,846,419 ( 1.13) 121,924,786,775 ( 1.30)
branches 23,725,889,239 ( 692.663) 23,733,547,140 ( 688.062) 23,553,950,311 ( 794.794)
branch-misses 90,733,041 ( 0.38%) 96,320,895 ( 0.41%) 84,561,092 ( 0.36%)
jobs3 ( ) ( ) ( )
stalled-cycles-frontend 66,437,834,608 ( 45.58%) 63,534,923,344 ( 43.69%) 42,101,478,505 ( 33.19%)
stalled-cycles-backend 34,940,799,661 ( 23.97%) 34,774,043,148 ( 23.91%) 21,163,324,388 ( 16.68%)
instructions 171,692,121,862 ( 1.18) 171,775,373,044 ( 1.18) 170,353,542,261 ( 1.34)
branches 32,968,962,622 ( 628.723) 32,987,739,894 ( 630.512) 32,729,463,918 ( 717.027)
branch-misses 111,522,732 ( 0.34%) 110,472,894 ( 0.33%) 99,791,291 ( 0.30%)
jobs4 ( ) ( ) ( )
stalled-cycles-frontend 98,741,701,675 ( 49.72%) 94,797,349,965 ( 47.59%) 54,535,655,381 ( 33.53%)
stalled-cycles-backend 54,642,609,615 ( 27.51%) 55,233,554,408 ( 27.73%) 27,882,323,541 ( 17.14%)
instructions 220,884,807,851 ( 1.11) 220,930,887,273 ( 1.11) 218,926,845,851 ( 1.35)
branches 42,354,518,180 ( 592.105) 42,362,770,587 ( 590.452) 41,955,552,870 ( 716.154)
branch-misses 138,093,449 ( 0.33%) 131,295,286 ( 0.31%) 121,794,771 ( 0.29%)
jobs5 ( ) ( ) ( )
stalled-cycles-frontend 116,219,747,212 ( 48.14%) 110,310,397,012 ( 46.29%) 66,373,082,723 ( 33.70%)
stalled-cycles-backend 66,325,434,776 ( 27.48%) 64,157,087,914 ( 26.92%) 32,999,097,299 ( 16.76%)
instructions 270,615,008,466 ( 1.12) 270,546,409,525 ( 1.14) 268,439,910,948 ( 1.36)
branches 51,834,046,557 ( 599.108) 51,811,867,722 ( 608.883) 51,412,576,077 ( 729.213)
branch-misses 158,197,086 ( 0.31%) 142,639,805 ( 0.28%) 133,425,455 ( 0.26%)
jobs6 ( ) ( ) ( )
stalled-cycles-frontend 138,009,414,492 ( 48.23%) 139,063,571,254 ( 48.80%) 75,278,568,278 ( 32.80%)
stalled-cycles-backend 79,211,949,650 ( 27.68%) 79,077,241,028 ( 27.75%) 37,735,797,899 ( 16.44%)
instructions 319,763,993,731 ( 1.12) 319,937,782,834 ( 1.12) 316,663,600,784 ( 1.38)
branches 61,219,433,294 ( 595.056) 61,250,355,540 ( 598.215) 60,523,446,617 ( 733.706)
branch-misses 169,257,123 ( 0.28%) 154,898,028 ( 0.25%) 141,180,587 ( 0.23%)
jobs7 ( ) ( ) ( )
stalled-cycles-frontend 162,974,812,119 ( 49.20%) 159,290,061,987 ( 48.43%) 88,046,641,169 ( 33.21%)
stalled-cycles-backend 92,223,151,661 ( 27.84%) 91,667,904,406 ( 27.87%) 44,068,454,971 ( 16.62%)
instructions 369,516,432,430 ( 1.12) 369,361,799,063 ( 1.12) 365,290,380,661 ( 1.38)
branches 70,795,673,950 ( 594.220) 70,743,136,124 ( 597.876) 69,803,996,038 ( 732.822)
branch-misses 181,708,327 ( 0.26%) 165,767,821 ( 0.23%) 150,109,797 ( 0.22%)
jobs8 ( ) ( ) ( )
stalled-cycles-frontend 185,000,017,027 ( 49.30%) 182,334,345,473 ( 48.37%) 99,980,147,041 ( 33.26%)
stalled-cycles-backend 105,753,516,186 ( 28.18%) 107,937,830,322 ( 28.63%) 51,404,177,181 ( 17.10%)
instructions 418,153,161,055 ( 1.11) 418,308,565,828 ( 1.11) 413,653,475,581 ( 1.38)
branches 80,035,882,398 ( 592.296) 80,063,204,510 ( 589.843) 79,024,105,589 ( 730.530)
branch-misses 199,764,528 ( 0.25%) 177,936,926 ( 0.22%) 160,525,449 ( 0.20%)
jobs9 ( ) ( ) ( )
stalled-cycles-frontend 210,941,799,094 ( 49.63%) 204,714,679,254 ( 48.55%) 114,251,113,756 ( 33.96%)
stalled-cycles-backend 122,640,849,067 ( 28.85%) 122,188,553,256 ( 28.98%) 58,360,041,127 ( 17.35%)
instructions 468,151,025,415 ( 1.10) 467,354,869,323 ( 1.11) 462,665,165,216 ( 1.38)
branches 89,657,067,510 ( 585.628) 89,411,550,407 ( 588.990) 88,360,523,943 ( 730.151)
branch-misses 218,292,301 ( 0.24%) 191,701,247 ( 0.21%) 178,535,678 ( 0.20%)
jobs10 ( ) ( ) ( )
stalled-cycles-frontend 233,595,958,008 ( 49.81%) 227,540,615,689 ( 49.11%) 160,341,979,938 ( 43.07%)
stalled-cycles-backend 136,153,676,021 ( 29.03%) 133,635,240,742 ( 28.84%) 65,909,135,465 ( 17.70%)
instructions 517,001,168,497 ( 1.10) 516,210,976,158 ( 1.11) 511,374,038,613 ( 1.37)
branches 98,911,641,329 ( 585.796) 98,700,069,712 ( 591.583) 97,646,761,028 ( 728.712)
branch-misses 232,341,823 ( 0.23%) 199,256,308 ( 0.20%) 183,135,268 ( 0.19%)
per-cpu streams tend to cause significantly less stalled cycles.
perf stat reported execution time
4 streams 8 streams per-cpu
====================================================================
jobs1
seconds elapsed 20.909073870 20.875670495 20.817838540
jobs2
seconds elapsed 18.529488399 18.720566469 16.356103108
jobs3
seconds elapsed 18.991159531 18.991340812 16.766216066
jobs4
seconds elapsed 19.560643828 19.551323547 16.246621715
jobs5
seconds elapsed 24.746498464 25.221646740 20.696112444
jobs6
seconds elapsed 28.258181828 28.289765505 22.885688857
jobs7
seconds elapsed 32.632490241 31.909125381 26.272753738
jobs8
seconds elapsed 35.651403851 36.027596308 29.108024711
jobs9
seconds elapsed 40.569362365 40.024227989 32.898204012
jobs10
seconds elapsed 44.673112304 43.874898137 35.632952191
quite interesting numbers.
NOTE:
-- fio seems does not attempt to write to device more than disk size, so
the test don't include 're-compresion path'.
-ss
[toc] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2016-04-19 10:00 +0200 |
| Message-ID | <rpyJI-3gW-7@gated-at.bofh.it> |
| In reply to | #1381350 |
On Mon, Apr 18, 2016 at 04:57:58PM +0900, Sergey Senozhatsky wrote: > Hello Minchan, > sorry, it took me so long to return back to testing. > > I collected extended stats (perf), just like you requested. > - 3G zram, lzo; 4 CPU x86_64 box. > - fio with perf stat > > 4 streams 8 streams per-cpu > =========================================================== > #jobs1 > READ: 2520.1MB/s 2566.5MB/s 2491.5MB/s > READ: 2102.7MB/s 2104.2MB/s 2091.3MB/s > WRITE: 1355.1MB/s 1320.2MB/s 1378.9MB/s > WRITE: 1103.5MB/s 1097.2MB/s 1122.5MB/s > READ: 434013KB/s 435153KB/s 439961KB/s > WRITE: 433969KB/s 435109KB/s 439917KB/s > READ: 403166KB/s 405139KB/s 403373KB/s > WRITE: 403223KB/s 405197KB/s 403430KB/s > #jobs2 > READ: 7958.6MB/s 8105.6MB/s 8073.7MB/s > READ: 6864.9MB/s 6989.8MB/s 7021.8MB/s > WRITE: 2438.1MB/s 2346.9MB/s 3400.2MB/s > WRITE: 1994.2MB/s 1990.3MB/s 2941.2MB/s > READ: 981504KB/s 973906KB/s 1018.8MB/s > WRITE: 981659KB/s 974060KB/s 1018.1MB/s > READ: 937021KB/s 938976KB/s 987250KB/s > WRITE: 934878KB/s 936830KB/s 984993KB/s > #jobs3 > READ: 13280MB/s 13553MB/s 13553MB/s > READ: 11534MB/s 11785MB/s 11755MB/s > WRITE: 3456.9MB/s 3469.9MB/s 4810.3MB/s > WRITE: 3029.6MB/s 3031.6MB/s 4264.8MB/s > READ: 1363.8MB/s 1362.6MB/s 1448.9MB/s > WRITE: 1361.9MB/s 1360.7MB/s 1446.9MB/s > READ: 1309.4MB/s 1310.6MB/s 1397.5MB/s > WRITE: 1307.4MB/s 1308.5MB/s 1395.3MB/s > #jobs4 > READ: 20244MB/s 20177MB/s 20344MB/s > READ: 17886MB/s 17913MB/s 17835MB/s > WRITE: 4071.6MB/s 4046.1MB/s 6370.2MB/s > WRITE: 3608.9MB/s 3576.3MB/s 5785.4MB/s > READ: 1824.3MB/s 1821.6MB/s 1997.5MB/s > WRITE: 1819.8MB/s 1817.4MB/s 1992.5MB/s > READ: 1765.7MB/s 1768.3MB/s 1937.3MB/s > WRITE: 1767.5MB/s 1769.1MB/s 1939.2MB/s > #jobs5 > READ: 18663MB/s 18986MB/s 18823MB/s > READ: 16659MB/s 16605MB/s 16954MB/s > WRITE: 3912.4MB/s 3888.7MB/s 6126.9MB/s > WRITE: 3506.4MB/s 3442.5MB/s 5519.3MB/s > READ: 1798.2MB/s 1746.5MB/s 1935.8MB/s > WRITE: 1792.7MB/s 1740.7MB/s 1929.1MB/s > READ: 1727.6MB/s 1658.2MB/s 1917.3MB/s > WRITE: 1726.5MB/s 1657.2MB/s 1916.6MB/s > #jobs6 > READ: 21017MB/s 20922MB/s 21162MB/s > READ: 19022MB/s 19140MB/s 18770MB/s > WRITE: 3968.2MB/s 4037.7MB/s 6620.8MB/s > WRITE: 3643.5MB/s 3590.2MB/s 6027.5MB/s > READ: 1871.8MB/s 1880.5MB/s 2049.9MB/s > WRITE: 1867.8MB/s 1877.2MB/s 2046.2MB/s > READ: 1755.8MB/s 1710.3MB/s 1964.7MB/s > WRITE: 1750.5MB/s 1705.9MB/s 1958.8MB/s > #jobs7 > READ: 21103MB/s 20677MB/s 21482MB/s > READ: 18522MB/s 18379MB/s 19443MB/s > WRITE: 4022.5MB/s 4067.4MB/s 6755.9MB/s > WRITE: 3691.7MB/s 3695.5MB/s 5925.6MB/s > READ: 1841.5MB/s 1933.9MB/s 2090.5MB/s > WRITE: 1842.7MB/s 1935.3MB/s 2091.9MB/s > READ: 1832.4MB/s 1856.4MB/s 1971.5MB/s > WRITE: 1822.3MB/s 1846.2MB/s 1960.6MB/s > #jobs8 > READ: 20463MB/s 20194MB/s 20862MB/s > READ: 18178MB/s 17978MB/s 18299MB/s > WRITE: 4085.9MB/s 4060.2MB/s 7023.8MB/s > WRITE: 3776.3MB/s 3737.9MB/s 6278.2MB/s > READ: 1957.6MB/s 1944.4MB/s 2109.5MB/s > WRITE: 1959.2MB/s 1946.2MB/s 2111.4MB/s > READ: 1900.6MB/s 1885.7MB/s 2082.1MB/s > WRITE: 1896.2MB/s 1881.4MB/s 2078.3MB/s > #jobs9 > READ: 19692MB/s 19734MB/s 19334MB/s > READ: 17678MB/s 18249MB/s 17666MB/s > WRITE: 4004.7MB/s 4064.8MB/s 6990.7MB/s > WRITE: 3724.7MB/s 3772.1MB/s 6193.6MB/s > READ: 1953.7MB/s 1967.3MB/s 2105.6MB/s > WRITE: 1953.4MB/s 1966.7MB/s 2104.1MB/s > READ: 1860.4MB/s 1897.4MB/s 2068.5MB/s > WRITE: 1858.9MB/s 1895.9MB/s 2066.8MB/s > #jobs10 > READ: 19730MB/s 19579MB/s 19492MB/s > READ: 18028MB/s 18018MB/s 18221MB/s > WRITE: 4027.3MB/s 4090.6MB/s 7020.1MB/s > WRITE: 3810.5MB/s 3846.8MB/s 6426.8MB/s > READ: 1956.1MB/s 1994.6MB/s 2145.2MB/s > WRITE: 1955.9MB/s 1993.5MB/s 2144.8MB/s > READ: 1852.8MB/s 1911.6MB/s 2075.8MB/s > WRITE: 1855.7MB/s 1914.6MB/s 2078.1MB/s > > > perf stat > > 4 streams 8 streams per-cpu > ==================================================================================================================== > jobs1 ( ) ( ) ( ) > stalled-cycles-frontend 23,174,811,209 ( 38.21%) 23,220,254,188 ( 38.25%) 23,061,406,918 ( 38.34%) > stalled-cycles-backend 11,514,174,638 ( 18.98%) 11,696,722,657 ( 19.27%) 11,370,852,810 ( 18.90%) > instructions 73,925,005,782 ( 1.22) 73,903,177,632 ( 1.22) 73,507,201,037 ( 1.22) > branches 14,455,124,835 ( 756.063) 14,455,184,779 ( 755.281) 14,378,599,509 ( 758.546) > branch-misses 69,801,336 ( 0.48%) 80,225,529 ( 0.55%) 72,044,726 ( 0.50%) > jobs2 ( ) ( ) ( ) > stalled-cycles-frontend 49,912,741,782 ( 46.11%) 50,101,189,290 ( 45.95%) 32,874,195,633 ( 35.11%) > stalled-cycles-backend 27,080,366,230 ( 25.02%) 27,949,970,232 ( 25.63%) 16,461,222,706 ( 17.58%) > instructions 122,831,629,690 ( 1.13) 122,919,846,419 ( 1.13) 121,924,786,775 ( 1.30) > branches 23,725,889,239 ( 692.663) 23,733,547,140 ( 688.062) 23,553,950,311 ( 794.794) > branch-misses 90,733,041 ( 0.38%) 96,320,895 ( 0.41%) 84,561,092 ( 0.36%) > jobs3 ( ) ( ) ( ) > stalled-cycles-frontend 66,437,834,608 ( 45.58%) 63,534,923,344 ( 43.69%) 42,101,478,505 ( 33.19%) > stalled-cycles-backend 34,940,799,661 ( 23.97%) 34,774,043,148 ( 23.91%) 21,163,324,388 ( 16.68%) > instructions 171,692,121,862 ( 1.18) 171,775,373,044 ( 1.18) 170,353,542,261 ( 1.34) > branches 32,968,962,622 ( 628.723) 32,987,739,894 ( 630.512) 32,729,463,918 ( 717.027) > branch-misses 111,522,732 ( 0.34%) 110,472,894 ( 0.33%) 99,791,291 ( 0.30%) > jobs4 ( ) ( ) ( ) > stalled-cycles-frontend 98,741,701,675 ( 49.72%) 94,797,349,965 ( 47.59%) 54,535,655,381 ( 33.53%) > stalled-cycles-backend 54,642,609,615 ( 27.51%) 55,233,554,408 ( 27.73%) 27,882,323,541 ( 17.14%) > instructions 220,884,807,851 ( 1.11) 220,930,887,273 ( 1.11) 218,926,845,851 ( 1.35) > branches 42,354,518,180 ( 592.105) 42,362,770,587 ( 590.452) 41,955,552,870 ( 716.154) > branch-misses 138,093,449 ( 0.33%) 131,295,286 ( 0.31%) 121,794,771 ( 0.29%) > jobs5 ( ) ( ) ( ) > stalled-cycles-frontend 116,219,747,212 ( 48.14%) 110,310,397,012 ( 46.29%) 66,373,082,723 ( 33.70%) > stalled-cycles-backend 66,325,434,776 ( 27.48%) 64,157,087,914 ( 26.92%) 32,999,097,299 ( 16.76%) > instructions 270,615,008,466 ( 1.12) 270,546,409,525 ( 1.14) 268,439,910,948 ( 1.36) > branches 51,834,046,557 ( 599.108) 51,811,867,722 ( 608.883) 51,412,576,077 ( 729.213) > branch-misses 158,197,086 ( 0.31%) 142,639,805 ( 0.28%) 133,425,455 ( 0.26%) > jobs6 ( ) ( ) ( ) > stalled-cycles-frontend 138,009,414,492 ( 48.23%) 139,063,571,254 ( 48.80%) 75,278,568,278 ( 32.80%) > stalled-cycles-backend 79,211,949,650 ( 27.68%) 79,077,241,028 ( 27.75%) 37,735,797,899 ( 16.44%) > instructions 319,763,993,731 ( 1.12) 319,937,782,834 ( 1.12) 316,663,600,784 ( 1.38) > branches 61,219,433,294 ( 595.056) 61,250,355,540 ( 598.215) 60,523,446,617 ( 733.706) > branch-misses 169,257,123 ( 0.28%) 154,898,028 ( 0.25%) 141,180,587 ( 0.23%) > jobs7 ( ) ( ) ( ) > stalled-cycles-frontend 162,974,812,119 ( 49.20%) 159,290,061,987 ( 48.43%) 88,046,641,169 ( 33.21%) > stalled-cycles-backend 92,223,151,661 ( 27.84%) 91,667,904,406 ( 27.87%) 44,068,454,971 ( 16.62%) > instructions 369,516,432,430 ( 1.12) 369,361,799,063 ( 1.12) 365,290,380,661 ( 1.38) > branches 70,795,673,950 ( 594.220) 70,743,136,124 ( 597.876) 69,803,996,038 ( 732.822) > branch-misses 181,708,327 ( 0.26%) 165,767,821 ( 0.23%) 150,109,797 ( 0.22%) > jobs8 ( ) ( ) ( ) > stalled-cycles-frontend 185,000,017,027 ( 49.30%) 182,334,345,473 ( 48.37%) 99,980,147,041 ( 33.26%) > stalled-cycles-backend 105,753,516,186 ( 28.18%) 107,937,830,322 ( 28.63%) 51,404,177,181 ( 17.10%) > instructions 418,153,161,055 ( 1.11) 418,308,565,828 ( 1.11) 413,653,475,581 ( 1.38) > branches 80,035,882,398 ( 592.296) 80,063,204,510 ( 589.843) 79,024,105,589 ( 730.530) > branch-misses 199,764,528 ( 0.25%) 177,936,926 ( 0.22%) 160,525,449 ( 0.20%) > jobs9 ( ) ( ) ( ) > stalled-cycles-frontend 210,941,799,094 ( 49.63%) 204,714,679,254 ( 48.55%) 114,251,113,756 ( 33.96%) > stalled-cycles-backend 122,640,849,067 ( 28.85%) 122,188,553,256 ( 28.98%) 58,360,041,127 ( 17.35%) > instructions 468,151,025,415 ( 1.10) 467,354,869,323 ( 1.11) 462,665,165,216 ( 1.38) > branches 89,657,067,510 ( 585.628) 89,411,550,407 ( 588.990) 88,360,523,943 ( 730.151) > branch-misses 218,292,301 ( 0.24%) 191,701,247 ( 0.21%) 178,535,678 ( 0.20%) > jobs10 ( ) ( ) ( ) > stalled-cycles-frontend 233,595,958,008 ( 49.81%) 227,540,615,689 ( 49.11%) 160,341,979,938 ( 43.07%) > stalled-cycles-backend 136,153,676,021 ( 29.03%) 133,635,240,742 ( 28.84%) 65,909,135,465 ( 17.70%) > instructions 517,001,168,497 ( 1.10) 516,210,976,158 ( 1.11) 511,374,038,613 ( 1.37) > branches 98,911,641,329 ( 585.796) 98,700,069,712 ( 591.583) 97,646,761,028 ( 728.712) > branch-misses 232,341,823 ( 0.23%) 199,256,308 ( 0.20%) 183,135,268 ( 0.19%) > > > per-cpu streams tend to cause significantly less stalled cycles. Great! So, based on your experiment, the reason I couldn't see such huge win in my mahcine is cache size difference(i.e., yours is twice than mine, IIRC.) and my perf stat didn't show such big difference. If I have a time, I will test it in bigger machine. > > > perf stat reported execution time > > 4 streams 8 streams per-cpu > ==================================================================== > jobs1 > seconds elapsed 20.909073870 20.875670495 20.817838540 > jobs2 > seconds elapsed 18.529488399 18.720566469 16.356103108 > jobs3 > seconds elapsed 18.991159531 18.991340812 16.766216066 > jobs4 > seconds elapsed 19.560643828 19.551323547 16.246621715 > jobs5 > seconds elapsed 24.746498464 25.221646740 20.696112444 > jobs6 > seconds elapsed 28.258181828 28.289765505 22.885688857 > jobs7 > seconds elapsed 32.632490241 31.909125381 26.272753738 > jobs8 > seconds elapsed 35.651403851 36.027596308 29.108024711 > jobs9 > seconds elapsed 40.569362365 40.024227989 32.898204012 > jobs10 > seconds elapsed 44.673112304 43.874898137 35.632952191 > > > quite interesting numbers. > > > > > NOTE: > -- fio seems does not attempt to write to device more than disk size, so > the test don't include 're-compresion path'. I'm convinced now with your data. Super thanks! However, as you know, we need data how bad it is in heavy memory pressure. Maybe, you can test it with fio and backgound memory hogger, Thanks for the test, Sergey!
[toc] | [prev] | [next] | [standalone]
| From | Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> |
|---|---|
| Date | 2016-04-19 10:10 +0200 |
| Message-ID | <rpyTp-3C4-15@gated-at.bofh.it> |
| In reply to | #1382248 |
Hello Minchan, On (04/19/16 17:00), Minchan Kim wrote: > Great! > > So, based on your experiment, the reason I couldn't see such huge win > in my mahcine is cache size difference(i.e., yours is twice than mine, > IIRC.) and my perf stat didn't show such big difference. > If I have a time, I will test it in bigger machine. quite possible it's due to the cache size. [..] > > NOTE: > > -- fio seems does not attempt to write to device more than disk size, so > > the test don't include 're-compresion path'. > > I'm convinced now with your data. Super thanks! > However, as you know, we need data how bad it is in heavy memory pressure. > Maybe, you can test it with fio and backgound memory hogger, yeah, sure, will work on it. > Thanks for the test, Sergey! thanks! -ss
[toc] | [prev] | [next] | [standalone]
| From | Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> |
|---|---|
| Date | 2016-04-26 13:30 +0200 |
| Message-ID | <rs9lM-6K3-7@gated-at.bofh.it> |
| In reply to | #1382248 |
Hello Minchan,
On (04/19/16 17:00), Minchan Kim wrote:
[..]
> I'm convinced now with your data. Super thanks!
> However, as you know, we need data how bad it is in heavy memory pressure.
> Maybe, you can test it with fio and backgound memory hogger,
it's really hard to produce stable test results when the system
is under mem pressure.
first, I modified zram to export the re-compression number
(put cpu stream and re-try handler allocation)
mm_stat for numjobs{1..10}. the number of re-compressions is in "< NUM>" format
3221225472 3221225472 3221225472 0 3221229568 0 0 < 6421>
3221225472 3221225472 3221225472 0 3221233664 0 0 < 6998>
3221225472 2912157607 2952802304 0 2952814592 0 84 < 7271>
3221225472 2893479936 2899120128 0 2899136512 0 156 < 8260>
3221217280 2886040814 2899099648 0 2899128320 0 78 < 8297>
3221225472 2880045056 2885693440 0 2885718016 0 54 < 7794>
3221213184 2877431364 2883756032 0 2883801088 0 144 < 7336>
3221225472 2873229312 2876096512 0 2876133376 0 28 < 8699>
3221213184 2870728008 2871693312 0 2871730176 0 30 < 8189>
2899095552 2899095552 2899095552 0 2899136512 78643 0 < 7485>
as we can see, the number of re-compressions can vary from 6421 to 8699.
the test:
-- 4 GB x86_64 box
-- zram 3GB, lzo
-- mem-hogger pre-faults 3GB of pages before the fio test
-- fio test has been modified to have 11% compression ratio (to increase the
chances of re-compressions)
-- buffer_compress_percentage=11
-- scramble_buffers=0
considering buffer_compress_percentage=11, the box was under somewhat
heavy pressure.
now, the results
fio stats
4 streams 8 streams per cpu
===========================================================
#jobs1
READ: 2411.4MB/s 2430.4MB/s 2440.4MB/s
READ: 2094.8MB/s 2002.7MB/s 2034.5MB/s
WRITE: 141571KB/s 140334KB/s 143542KB/s
WRITE: 712025KB/s 706111KB/s 745256KB/s
READ: 531014KB/s 525250KB/s 537547KB/s
WRITE: 530960KB/s 525197KB/s 537492KB/s
READ: 473577KB/s 470320KB/s 476880KB/s
WRITE: 473645KB/s 470387KB/s 476948KB/s
#jobs2
READ: 7897.2MB/s 8031.4MB/s 7968.9MB/s
READ: 6864.9MB/s 6803.2MB/s 6903.4MB/s
WRITE: 321386KB/s 314227KB/s 313101KB/s
WRITE: 1275.3MB/s 1245.6MB/s 1383.5MB/s
READ: 1035.5MB/s 1021.9MB/s 1098.4MB/s
WRITE: 1035.6MB/s 1021.1MB/s 1098.6MB/s
READ: 972014KB/s 952321KB/s 987.66MB/s
WRITE: 969792KB/s 950144KB/s 985.40MB/s
#jobs3
READ: 13260MB/s 13260MB/s 13222MB/s
READ: 11636MB/s 11636MB/s 11755MB/s
WRITE: 511500KB/s 507730KB/s 504959KB/s
WRITE: 1646.1MB/s 1673.9MB/s 1755.5MB/s
READ: 1389.5MB/s 1387.2MB/s 1479.6MB/s
WRITE: 1387.6MB/s 1385.3MB/s 1477.4MB/s
READ: 1286.8MB/s 1289.1MB/s 1377.3MB/s
WRITE: 1284.8MB/s 1287.1MB/s 1374.9MB/s
#jobs4
READ: 19851MB/s 20244MB/s 20344MB/s
READ: 17732MB/s 17835MB/s 18097MB/s
WRITE: 667776KB/s 655599KB/s 693464KB/s
WRITE: 2041.2MB/s 2072.6MB/s 2474.1MB/s
READ: 1770.1MB/s 1781.7MB/s 2035.5MB/s
WRITE: 1765.8MB/s 1777.3MB/s 2030.5MB/s
READ: 1641.6MB/s 1672.4MB/s 1892.5MB/s
WRITE: 1643.2MB/s 1674.2MB/s 1894.4MB/s
#jobs5
READ: 19468MB/s 18484MB/s 18439MB/s
READ: 17594MB/s 17757MB/s 17716MB/s
WRITE: 843266KB/s 859627KB/s 867928KB/s
WRITE: 1927.1MB/s 2041.8MB/s 2168.9MB/s
READ: 1718.6MB/s 1771.7MB/s 1963.5MB/s
WRITE: 1712.7MB/s 1765.6MB/s 1956.8MB/s
READ: 1705.3MB/s 1663.6MB/s 1767.3MB/s
WRITE: 1704.3MB/s 1662.6MB/s 1766.2MB/s
#jobs6
READ: 21583MB/s 21685MB/s 21483MB/s
READ: 19160MB/s 18432MB/s 18618MB/s
WRITE: 986276KB/s 1004.2MB/s 981.11MB/s
WRITE: 2013.6MB/s 1922.5MB/s 2429.1MB/s
READ: 1797.1MB/s 1678.9MB/s 2038.8MB/s
WRITE: 1794.8MB/s 1675.9MB/s 2035.2MB/s
READ: 1678.2MB/s 1632.5MB/s 1917.4MB/s
WRITE: 1673.9MB/s 1627.6MB/s 1911.6MB/s
#jobs7
READ: 20697MB/s 21677MB/s 21062MB/s
READ: 18781MB/s 18667MB/s 19338MB/s
WRITE: 1074.6MB/s 1099.8MB/s 1105.3MB/s
WRITE: 2100.7MB/s 2010.3MB/s 2598.7MB/s
READ: 1783.2MB/s 1710.2MB/s 2027.8MB/s
WRITE: 1784.3MB/s 1712.1MB/s 2029.6MB/s
READ: 1690.8MB/s 1620.6MB/s 1893.6MB/s
WRITE: 1681.4MB/s 1611.7MB/s 1883.7MB/s
#jobs8
READ: 19883MB/s 20827MB/s 20395MB/s
READ: 18562MB/s 18178MB/s 17822MB/s
WRITE: 1240.5MB/s 1307.3MB/s 1331.7MB/s
WRITE: 2132.1MB/s 2143.6MB/s 2564.9MB/s
READ: 1841.1MB/s 1831.1MB/s 2111.4MB/s
WRITE: 1843.1MB/s 1833.1MB/s 2113.4MB/s
READ: 1795.4MB/s 1778.6MB/s 2029.3MB/s
WRITE: 1791.4MB/s 1774.5MB/s 2024.5MB/s
#jobs9
READ: 18834MB/s 19470MB/s 19402MB/s
READ: 17988MB/s 18118MB/s 18531MB/s
WRITE: 1339.4MB/s 1441.2MB/s 1512.6MB/s
WRITE: 2102.4MB/s 2111.9MB/s 2478.8MB/s
READ: 1754.5MB/s 1777.3MB/s 2050.2MB/s
WRITE: 1753.9MB/s 1776.7MB/s 2049.5MB/s
READ: 1686.4MB/s 1698.2MB/s 1931.6MB/s
WRITE: 1684.1MB/s 1696.8MB/s 1929.1MB/s
#jobs10
READ: 19128MB/s 19517MB/s 19592MB/s
READ: 18177MB/s 17544MB/s 18221MB/s
WRITE: 1397.1MB/s 1567.4MB/s 1683.2MB/s
WRITE: 2151.9MB/s 2205.1MB/s 2642.6MB/s
READ: 1879.2MB/s 1907.3MB/s 2223.2MB/s
WRITE: 1878.5MB/s 1906.2MB/s 2222.8MB/s
READ: 1835.7MB/s 1837.9MB/s 2131.4MB/s
WRITE: 1838.6MB/s 1840.8MB/s 2134.8MB/s
perf stats
4 streams 8 streams per cpu
====================================================================================================================
jobs1
stalled-cycles-frontend 52,219,601,943 ( 55.87%) 53,406,899,652 ( 56.33%) 49,944,625,376 ( 56.27%)
stalled-cycles-backend 23,194,739,214 ( 24.82%) 24,397,423,796 ( 25.73%) 22,782,579,660 ( 25.67%)
instructions 86,078,512,819 ( 0.92) 86,235,354,709 ( 0.91) 80,378,845,354 ( 0.91)
branches 15,732,850,506 ( 532.108) 15,743,473,327 ( 522.592) 14,725,420,241 ( 523.425)
branch-misses 104,546,578 ( 0.66%) 107,847,818 ( 0.69%) 106,343,602 ( 0.72%)
jobs2
stalled-cycles-frontend 118,614,605,521 ( 59.74%) 113,520,838,279 ( 59.94%) 104,301,243,221 ( 59.06%)
stalled-cycles-backend 59,490,170,824 ( 29.96%) 56,518,872,622 ( 29.84%) 50,161,702,782 ( 28.40%)
instructions 169,663,993,572 ( 0.85) 160,959,388,344 ( 0.85) 153,541,182,646 ( 0.87)
branches 31,859,926,551 ( 497.945) 30,132,524,256 ( 494.660) 28,579,927,064 ( 503.079)
branch-misses 164,531,311 ( 0.52%) 163,509,596 ( 0.54%) 145,472,902 ( 0.51%)
jobs3
stalled-cycles-frontend 153,932,401,104 ( 60.86%) 158,470,334,291 ( 60.81%) 150,767,641,835 ( 59.21%)
stalled-cycles-backend 77,023,824,597 ( 30.45%) 79,673,952,089 ( 30.57%) 72,693,245,174 ( 28.55%)
instructions 197,452,119,661 ( 0.78) 204,116,060,906 ( 0.78) 207,832,729,315 ( 0.82)
branches 36,579,918,543 ( 404.660) 37,980,582,651 ( 406.326) 39,091,715,974 ( 428.559)
branch-misses 214,292,753 ( 0.59%) 215,861,282 ( 0.57%) 203,320,703 ( 0.52%)
jobs4
stalled-cycles-frontend 237,223,396,661 ( 64.22%) 227,572,336,186 ( 64.37%) 202,100,979,033 ( 61.41%)
stalled-cycles-backend 129,935,296,918 ( 35.17%) 124,957,172,193 ( 35.34%) 103,626,575,103 ( 31.49%)
instructions 270,083,196,348 ( 0.73) 257,652,752,109 ( 0.73) 259,773,237,031 ( 0.79)
branches 52,120,828,566 ( 391.426) 49,121,254,042 ( 385.647) 49,896,944,076 ( 420.532)
branch-misses 260,480,947 ( 0.50%) 254,957,745 ( 0.52%) 239,402,681 ( 0.48%)
jobs5
stalled-cycles-frontend 257,778,703,389 ( 64.89%) 265,688,762,182 ( 65.13%) 229,916,792,090 ( 61.41%)
stalled-cycles-backend 142,090,098,727 ( 35.77%) 147,101,411,510 ( 36.06%) 117,081,586,471 ( 31.27%)
instructions 291,859,438,730 ( 0.73) 298,380,653,546 ( 0.73) 302,840,047,693 ( 0.81)
branches 55,111,567,225 ( 385.905) 56,316,470,332 ( 383.545) 57,500,842,324 ( 428.083)
branch-misses 270,056,201 ( 0.49%) 269,400,845 ( 0.48%) 258,495,925 ( 0.45%)
jobs6
stalled-cycles-frontend 311,626,093,277 ( 65.61%) 314,291,595,576 ( 65.77%) 249,524,291,273 ( 61.39%)
stalled-cycles-backend 174,358,063,361 ( 36.71%) 177,312,195,233 ( 37.10%) 126,508,172,269 ( 31.13%)
instructions 345,271,436,105 ( 0.73) 346,679,577,246 ( 0.73) 333,258,054,473 ( 0.82)
branches 65,298,537,641 ( 381.664) 65,995,652,812 ( 383.717) 62,730,160,550 ( 428.999)
branch-misses 313,241,654 ( 0.48%) 307,876,772 ( 0.47%) 282,570,360 ( 0.45%)
jobs7
stalled-cycles-frontend 333,896,608,350 ( 64.68%) 349,165,441,969 ( 64.85%) 276,185,831,513 ( 59.95%)
stalled-cycles-backend 186,083,638,772 ( 36.05%) 197,000,957,906 ( 36.59%) 138,835,486,733 ( 30.14%)
instructions 388,707,023,219 ( 0.75) 404,347,465,692 ( 0.75) 394,078,203,426 ( 0.86)
branches 71,999,476,930 ( 387.008) 76,197,698,685 ( 392.759) 73,195,649,665 ( 440.914)
branch-misses 328,598,294 ( 0.46%) 323,895,230 ( 0.43%) 298,205,996 ( 0.41%)
jobs8
stalled-cycles-frontend 378,806,234,772 ( 66.73%) 369,453,970,323 ( 66.55%) 313,738,845,641 ( 62.55%)
stalled-cycles-backend 211,732,966,238 ( 37.30%) 207,691,463,546 ( 37.41%) 161,120,924,768 ( 32.12%)
instructions 406,674,721,912 ( 0.72) 401,922,649,599 ( 0.72) 405,830,823,213 ( 0.81)
branches 75,637,492,422 ( 369.371) 74,287,789,757 ( 371.226) 75,967,291,039 ( 420.260)
branch-misses 355,733,892 ( 0.47%) 328,972,387 ( 0.44%) 318,203,258 ( 0.42%)
jobs9
stalled-cycles-frontend 422,712,242,907 ( 66.39%) 417,293,429,710 ( 66.14%) 343,703,467,466 ( 61.35%)
stalled-cycles-backend 239,356,726,574 ( 37.59%) 231,725,068,834 ( 36.73%) 172,101,321,805 ( 30.72%)
instructions 465,964,470,967 ( 0.73) 468,561,486,803 ( 0.74) 474,119,504,255 ( 0.85)
branches 86,724,291,348 ( 377.755) 86,534,438,758 ( 380.374) 88,431,722,886 ( 437.939)
branch-misses 385,706,052 ( 0.44%) 360,946,347 ( 0.42%) 337,858,267 ( 0.38%)
jobs10
stalled-cycles-frontend 451,844,797,592 ( 67.24%) 435,099,070,573 ( 67.18%) 352,877,428,118 ( 62.18%)
stalled-cycles-backend 255,533,666,521 ( 38.03%) 249,295,276,734 ( 38.49%) 179,754,582,074 ( 31.67%)
instructions 472,331,884,636 ( 0.70) 458,948,698,965 ( 0.71) 464,131,768,633 ( 0.82)
branches 88,848,212,769 ( 366.556) 85,330,239,413 ( 365.282) 86,837,838,069 ( 424.329)
branch-misses 398,856,497 ( 0.45%) 359,532,394 ( 0.42%) 333,821,387 ( 0.38%)
perf reported execution time
4 streams 8 streams per cpu
====================================================================
seconds elapsed 41.359653597 43.131195776 40.961640812
seconds elapsed 37.778174380 38.681792299 38.368529861
seconds elapsed 38.367149768 39.368008799 37.687545579
seconds elapsed 40.402963748 39.177529033 36.205357101
seconds elapsed 44.145428970 43.251655348 41.810848146
seconds elapsed 49.344988495 49.951048242 44.270045250
seconds elapsed 53.865398777 54.271392367 48.824173559
seconds elapsed 57.028770416 56.228105290 51.332017545
seconds elapsed 62.931350164 61.251237873 55.977463074
seconds elapsed 67.088285633 63.544376242 57.690998344
-ss
[toc] | [prev] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2016-04-27 09:30 +0200 |
| Message-ID | <rss54-5GG-17@gated-at.bofh.it> |
| In reply to | #1387386 |
Hello Sergey,
On Tue, Apr 26, 2016 at 08:23:05PM +0900, Sergey Senozhatsky wrote:
> Hello Minchan,
>
> On (04/19/16 17:00), Minchan Kim wrote:
> [..]
> > I'm convinced now with your data. Super thanks!
> > However, as you know, we need data how bad it is in heavy memory pressure.
> > Maybe, you can test it with fio and backgound memory hogger,
>
> it's really hard to produce stable test results when the system
> is under mem pressure.
>
> first, I modified zram to export the re-compression number
> (put cpu stream and re-try handler allocation)
>
> mm_stat for numjobs{1..10}. the number of re-compressions is in "< NUM>" format
>
> 3221225472 3221225472 3221225472 0 3221229568 0 0 < 6421>
> 3221225472 3221225472 3221225472 0 3221233664 0 0 < 6998>
> 3221225472 2912157607 2952802304 0 2952814592 0 84 < 7271>
> 3221225472 2893479936 2899120128 0 2899136512 0 156 < 8260>
> 3221217280 2886040814 2899099648 0 2899128320 0 78 < 8297>
> 3221225472 2880045056 2885693440 0 2885718016 0 54 < 7794>
> 3221213184 2877431364 2883756032 0 2883801088 0 144 < 7336>
> 3221225472 2873229312 2876096512 0 2876133376 0 28 < 8699>
> 3221213184 2870728008 2871693312 0 2871730176 0 30 < 8189>
> 2899095552 2899095552 2899095552 0 2899136512 78643 0 < 7485>
It would be great when we see the below ratio for each test.
1-compression : 2(re)-compression
>
> as we can see, the number of re-compressions can vary from 6421 to 8699.
>
>
> the test:
>
> -- 4 GB x86_64 box
> -- zram 3GB, lzo
> -- mem-hogger pre-faults 3GB of pages before the fio test
> -- fio test has been modified to have 11% compression ratio (to increase the
> chances of re-compressions)
Could you test concurrent mem hogger with fio rather than pre-fault before fio test
in next submit?
> -- buffer_compress_percentage=11
> -- scramble_buffers=0
>
>
> considering buffer_compress_percentage=11, the box was under somewhat
> heavy pressure.
>
> now, the results
Yeb, Even, recompression case is fater than old but want to see more heavy memory
pressure case and the ratio I mentioned above.
If the result is still good, please send public patch with number.
Thanks for looking this, Sergey!
>
>
> fio stats
>
> 4 streams 8 streams per cpu
> ===========================================================
> #jobs1
> READ: 2411.4MB/s 2430.4MB/s 2440.4MB/s
> READ: 2094.8MB/s 2002.7MB/s 2034.5MB/s
> WRITE: 141571KB/s 140334KB/s 143542KB/s
> WRITE: 712025KB/s 706111KB/s 745256KB/s
> READ: 531014KB/s 525250KB/s 537547KB/s
> WRITE: 530960KB/s 525197KB/s 537492KB/s
> READ: 473577KB/s 470320KB/s 476880KB/s
> WRITE: 473645KB/s 470387KB/s 476948KB/s
> #jobs2
> READ: 7897.2MB/s 8031.4MB/s 7968.9MB/s
> READ: 6864.9MB/s 6803.2MB/s 6903.4MB/s
> WRITE: 321386KB/s 314227KB/s 313101KB/s
> WRITE: 1275.3MB/s 1245.6MB/s 1383.5MB/s
> READ: 1035.5MB/s 1021.9MB/s 1098.4MB/s
> WRITE: 1035.6MB/s 1021.1MB/s 1098.6MB/s
> READ: 972014KB/s 952321KB/s 987.66MB/s
> WRITE: 969792KB/s 950144KB/s 985.40MB/s
> #jobs3
> READ: 13260MB/s 13260MB/s 13222MB/s
> READ: 11636MB/s 11636MB/s 11755MB/s
> WRITE: 511500KB/s 507730KB/s 504959KB/s
> WRITE: 1646.1MB/s 1673.9MB/s 1755.5MB/s
> READ: 1389.5MB/s 1387.2MB/s 1479.6MB/s
> WRITE: 1387.6MB/s 1385.3MB/s 1477.4MB/s
> READ: 1286.8MB/s 1289.1MB/s 1377.3MB/s
> WRITE: 1284.8MB/s 1287.1MB/s 1374.9MB/s
> #jobs4
> READ: 19851MB/s 20244MB/s 20344MB/s
> READ: 17732MB/s 17835MB/s 18097MB/s
> WRITE: 667776KB/s 655599KB/s 693464KB/s
> WRITE: 2041.2MB/s 2072.6MB/s 2474.1MB/s
> READ: 1770.1MB/s 1781.7MB/s 2035.5MB/s
> WRITE: 1765.8MB/s 1777.3MB/s 2030.5MB/s
> READ: 1641.6MB/s 1672.4MB/s 1892.5MB/s
> WRITE: 1643.2MB/s 1674.2MB/s 1894.4MB/s
> #jobs5
> READ: 19468MB/s 18484MB/s 18439MB/s
> READ: 17594MB/s 17757MB/s 17716MB/s
> WRITE: 843266KB/s 859627KB/s 867928KB/s
> WRITE: 1927.1MB/s 2041.8MB/s 2168.9MB/s
> READ: 1718.6MB/s 1771.7MB/s 1963.5MB/s
> WRITE: 1712.7MB/s 1765.6MB/s 1956.8MB/s
> READ: 1705.3MB/s 1663.6MB/s 1767.3MB/s
> WRITE: 1704.3MB/s 1662.6MB/s 1766.2MB/s
> #jobs6
> READ: 21583MB/s 21685MB/s 21483MB/s
> READ: 19160MB/s 18432MB/s 18618MB/s
> WRITE: 986276KB/s 1004.2MB/s 981.11MB/s
> WRITE: 2013.6MB/s 1922.5MB/s 2429.1MB/s
> READ: 1797.1MB/s 1678.9MB/s 2038.8MB/s
> WRITE: 1794.8MB/s 1675.9MB/s 2035.2MB/s
> READ: 1678.2MB/s 1632.5MB/s 1917.4MB/s
> WRITE: 1673.9MB/s 1627.6MB/s 1911.6MB/s
> #jobs7
> READ: 20697MB/s 21677MB/s 21062MB/s
> READ: 18781MB/s 18667MB/s 19338MB/s
> WRITE: 1074.6MB/s 1099.8MB/s 1105.3MB/s
> WRITE: 2100.7MB/s 2010.3MB/s 2598.7MB/s
> READ: 1783.2MB/s 1710.2MB/s 2027.8MB/s
> WRITE: 1784.3MB/s 1712.1MB/s 2029.6MB/s
> READ: 1690.8MB/s 1620.6MB/s 1893.6MB/s
> WRITE: 1681.4MB/s 1611.7MB/s 1883.7MB/s
> #jobs8
> READ: 19883MB/s 20827MB/s 20395MB/s
> READ: 18562MB/s 18178MB/s 17822MB/s
> WRITE: 1240.5MB/s 1307.3MB/s 1331.7MB/s
> WRITE: 2132.1MB/s 2143.6MB/s 2564.9MB/s
> READ: 1841.1MB/s 1831.1MB/s 2111.4MB/s
> WRITE: 1843.1MB/s 1833.1MB/s 2113.4MB/s
> READ: 1795.4MB/s 1778.6MB/s 2029.3MB/s
> WRITE: 1791.4MB/s 1774.5MB/s 2024.5MB/s
> #jobs9
> READ: 18834MB/s 19470MB/s 19402MB/s
> READ: 17988MB/s 18118MB/s 18531MB/s
> WRITE: 1339.4MB/s 1441.2MB/s 1512.6MB/s
> WRITE: 2102.4MB/s 2111.9MB/s 2478.8MB/s
> READ: 1754.5MB/s 1777.3MB/s 2050.2MB/s
> WRITE: 1753.9MB/s 1776.7MB/s 2049.5MB/s
> READ: 1686.4MB/s 1698.2MB/s 1931.6MB/s
> WRITE: 1684.1MB/s 1696.8MB/s 1929.1MB/s
> #jobs10
> READ: 19128MB/s 19517MB/s 19592MB/s
> READ: 18177MB/s 17544MB/s 18221MB/s
> WRITE: 1397.1MB/s 1567.4MB/s 1683.2MB/s
> WRITE: 2151.9MB/s 2205.1MB/s 2642.6MB/s
> READ: 1879.2MB/s 1907.3MB/s 2223.2MB/s
> WRITE: 1878.5MB/s 1906.2MB/s 2222.8MB/s
> READ: 1835.7MB/s 1837.9MB/s 2131.4MB/s
> WRITE: 1838.6MB/s 1840.8MB/s 2134.8MB/s
>
>
> perf stats
>
> 4 streams 8 streams per cpu
> ====================================================================================================================
> jobs1
> stalled-cycles-frontend 52,219,601,943 ( 55.87%) 53,406,899,652 ( 56.33%) 49,944,625,376 ( 56.27%)
> stalled-cycles-backend 23,194,739,214 ( 24.82%) 24,397,423,796 ( 25.73%) 22,782,579,660 ( 25.67%)
> instructions 86,078,512,819 ( 0.92) 86,235,354,709 ( 0.91) 80,378,845,354 ( 0.91)
> branches 15,732,850,506 ( 532.108) 15,743,473,327 ( 522.592) 14,725,420,241 ( 523.425)
> branch-misses 104,546,578 ( 0.66%) 107,847,818 ( 0.69%) 106,343,602 ( 0.72%)
> jobs2
> stalled-cycles-frontend 118,614,605,521 ( 59.74%) 113,520,838,279 ( 59.94%) 104,301,243,221 ( 59.06%)
> stalled-cycles-backend 59,490,170,824 ( 29.96%) 56,518,872,622 ( 29.84%) 50,161,702,782 ( 28.40%)
> instructions 169,663,993,572 ( 0.85) 160,959,388,344 ( 0.85) 153,541,182,646 ( 0.87)
> branches 31,859,926,551 ( 497.945) 30,132,524,256 ( 494.660) 28,579,927,064 ( 503.079)
> branch-misses 164,531,311 ( 0.52%) 163,509,596 ( 0.54%) 145,472,902 ( 0.51%)
> jobs3
> stalled-cycles-frontend 153,932,401,104 ( 60.86%) 158,470,334,291 ( 60.81%) 150,767,641,835 ( 59.21%)
> stalled-cycles-backend 77,023,824,597 ( 30.45%) 79,673,952,089 ( 30.57%) 72,693,245,174 ( 28.55%)
> instructions 197,452,119,661 ( 0.78) 204,116,060,906 ( 0.78) 207,832,729,315 ( 0.82)
> branches 36,579,918,543 ( 404.660) 37,980,582,651 ( 406.326) 39,091,715,974 ( 428.559)
> branch-misses 214,292,753 ( 0.59%) 215,861,282 ( 0.57%) 203,320,703 ( 0.52%)
> jobs4
> stalled-cycles-frontend 237,223,396,661 ( 64.22%) 227,572,336,186 ( 64.37%) 202,100,979,033 ( 61.41%)
> stalled-cycles-backend 129,935,296,918 ( 35.17%) 124,957,172,193 ( 35.34%) 103,626,575,103 ( 31.49%)
> instructions 270,083,196,348 ( 0.73) 257,652,752,109 ( 0.73) 259,773,237,031 ( 0.79)
> branches 52,120,828,566 ( 391.426) 49,121,254,042 ( 385.647) 49,896,944,076 ( 420.532)
> branch-misses 260,480,947 ( 0.50%) 254,957,745 ( 0.52%) 239,402,681 ( 0.48%)
> jobs5
> stalled-cycles-frontend 257,778,703,389 ( 64.89%) 265,688,762,182 ( 65.13%) 229,916,792,090 ( 61.41%)
> stalled-cycles-backend 142,090,098,727 ( 35.77%) 147,101,411,510 ( 36.06%) 117,081,586,471 ( 31.27%)
> instructions 291,859,438,730 ( 0.73) 298,380,653,546 ( 0.73) 302,840,047,693 ( 0.81)
> branches 55,111,567,225 ( 385.905) 56,316,470,332 ( 383.545) 57,500,842,324 ( 428.083)
> branch-misses 270,056,201 ( 0.49%) 269,400,845 ( 0.48%) 258,495,925 ( 0.45%)
> jobs6
> stalled-cycles-frontend 311,626,093,277 ( 65.61%) 314,291,595,576 ( 65.77%) 249,524,291,273 ( 61.39%)
> stalled-cycles-backend 174,358,063,361 ( 36.71%) 177,312,195,233 ( 37.10%) 126,508,172,269 ( 31.13%)
> instructions 345,271,436,105 ( 0.73) 346,679,577,246 ( 0.73) 333,258,054,473 ( 0.82)
> branches 65,298,537,641 ( 381.664) 65,995,652,812 ( 383.717) 62,730,160,550 ( 428.999)
> branch-misses 313,241,654 ( 0.48%) 307,876,772 ( 0.47%) 282,570,360 ( 0.45%)
> jobs7
> stalled-cycles-frontend 333,896,608,350 ( 64.68%) 349,165,441,969 ( 64.85%) 276,185,831,513 ( 59.95%)
> stalled-cycles-backend 186,083,638,772 ( 36.05%) 197,000,957,906 ( 36.59%) 138,835,486,733 ( 30.14%)
> instructions 388,707,023,219 ( 0.75) 404,347,465,692 ( 0.75) 394,078,203,426 ( 0.86)
> branches 71,999,476,930 ( 387.008) 76,197,698,685 ( 392.759) 73,195,649,665 ( 440.914)
> branch-misses 328,598,294 ( 0.46%) 323,895,230 ( 0.43%) 298,205,996 ( 0.41%)
> jobs8
> stalled-cycles-frontend 378,806,234,772 ( 66.73%) 369,453,970,323 ( 66.55%) 313,738,845,641 ( 62.55%)
> stalled-cycles-backend 211,732,966,238 ( 37.30%) 207,691,463,546 ( 37.41%) 161,120,924,768 ( 32.12%)
> instructions 406,674,721,912 ( 0.72) 401,922,649,599 ( 0.72) 405,830,823,213 ( 0.81)
> branches 75,637,492,422 ( 369.371) 74,287,789,757 ( 371.226) 75,967,291,039 ( 420.260)
> branch-misses 355,733,892 ( 0.47%) 328,972,387 ( 0.44%) 318,203,258 ( 0.42%)
> jobs9
> stalled-cycles-frontend 422,712,242,907 ( 66.39%) 417,293,429,710 ( 66.14%) 343,703,467,466 ( 61.35%)
> stalled-cycles-backend 239,356,726,574 ( 37.59%) 231,725,068,834 ( 36.73%) 172,101,321,805 ( 30.72%)
> instructions 465,964,470,967 ( 0.73) 468,561,486,803 ( 0.74) 474,119,504,255 ( 0.85)
> branches 86,724,291,348 ( 377.755) 86,534,438,758 ( 380.374) 88,431,722,886 ( 437.939)
> branch-misses 385,706,052 ( 0.44%) 360,946,347 ( 0.42%) 337,858,267 ( 0.38%)
> jobs10
> stalled-cycles-frontend 451,844,797,592 ( 67.24%) 435,099,070,573 ( 67.18%) 352,877,428,118 ( 62.18%)
> stalled-cycles-backend 255,533,666,521 ( 38.03%) 249,295,276,734 ( 38.49%) 179,754,582,074 ( 31.67%)
> instructions 472,331,884,636 ( 0.70) 458,948,698,965 ( 0.71) 464,131,768,633 ( 0.82)
> branches 88,848,212,769 ( 366.556) 85,330,239,413 ( 365.282) 86,837,838,069 ( 424.329)
> branch-misses 398,856,497 ( 0.45%) 359,532,394 ( 0.42%) 333,821,387 ( 0.38%)
>
>
>
> perf reported execution time
>
> 4 streams 8 streams per cpu
> ====================================================================
> seconds elapsed 41.359653597 43.131195776 40.961640812
> seconds elapsed 37.778174380 38.681792299 38.368529861
> seconds elapsed 38.367149768 39.368008799 37.687545579
> seconds elapsed 40.402963748 39.177529033 36.205357101
> seconds elapsed 44.145428970 43.251655348 41.810848146
> seconds elapsed 49.344988495 49.951048242 44.270045250
> seconds elapsed 53.865398777 54.271392367 48.824173559
> seconds elapsed 57.028770416 56.228105290 51.332017545
> seconds elapsed 62.931350164 61.251237873 55.977463074
> seconds elapsed 67.088285633 63.544376242 57.690998344
>
>
> -ss
[toc] | [prev] | [next] | [standalone]
| From | Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> |
|---|---|
| Date | 2016-04-27 09:50 +0200 |
| Message-ID | <rssoq-5QU-13@gated-at.bofh.it> |
| In reply to | #1388720 |
Hello, On (04/27/16 16:29), Minchan Kim wrote: [..] > > the test: > > > > -- 4 GB x86_64 box > > -- zram 3GB, lzo > > -- mem-hogger pre-faults 3GB of pages before the fio test > > -- fio test has been modified to have 11% compression ratio (to increase the > > chances of re-compressions) > > Could you test concurrent mem hogger with fio rather than pre-fault before fio test > in next submit? this test will not prove anything, unfortunately. I performed it; and it's impossible to guarantee even remotely stable results. mem-hogger process can spend on pre-fault from 41 to 81 seconds; so I'm quite sceptical about the actual value of this test. > > considering buffer_compress_percentage=11, the box was under somewhat > > heavy pressure. > > > > now, the results > > Yeb, Even, recompression case is fater than old but want to see more heavy memory > pressure case and the ratio I mentioned above. I did quite heavy testing over the last 7 days, with numerous OOM kills and OOM panics. -ss
[toc] | [prev] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2016-04-27 10:00 +0200 |
| Message-ID | <rssy6-5VB-1@gated-at.bofh.it> |
| In reply to | #1388734 |
On Wed, Apr 27, 2016 at 04:43:35PM +0900, Sergey Senozhatsky wrote: > Hello, > > On (04/27/16 16:29), Minchan Kim wrote: > [..] > > > the test: > > > > > > -- 4 GB x86_64 box > > > -- zram 3GB, lzo > > > -- mem-hogger pre-faults 3GB of pages before the fio test > > > -- fio test has been modified to have 11% compression ratio (to increase the > > > chances of re-compressions) > > > > Could you test concurrent mem hogger with fio rather than pre-fault before fio test > > in next submit? > > this test will not prove anything, unfortunately. I performed it; > and it's impossible to guarantee even remotely stable results. > mem-hogger process can spend on pre-fault from 41 to 81 seconds; > so I'm quite sceptical about the actual value of this test. > > > > considering buffer_compress_percentage=11, the box was under somewhat > > > heavy pressure. > > > > > > now, the results > > > > Yeb, Even, recompression case is fater than old but want to see more heavy memory > > pressure case and the ratio I mentioned above. > > I did quite heavy testing over the last 7 days, with numerous OOM kills > and OOM panics. Okay, I think it's worth to merge enough and see the result. Please send formal patch which has recompression stat. ;-) Thanks.
[toc] | [prev] | [next] | [standalone]
| From | Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> |
|---|---|
| Date | 2016-04-27 10:10 +0200 |
| Message-ID | <rssHN-6hy-39@gated-at.bofh.it> |
| In reply to | #1388744 |
On (04/27/16 16:55), Minchan Kim wrote: [..] > > > Could you test concurrent mem hogger with fio rather than pre-fault before fio test > > > in next submit? > > > > this test will not prove anything, unfortunately. I performed it; > > and it's impossible to guarantee even remotely stable results. > > mem-hogger process can spend on pre-fault from 41 to 81 seconds; > > so I'm quite sceptical about the actual value of this test. > > > > > > considering buffer_compress_percentage=11, the box was under somewhat > > > > heavy pressure. > > > > > > > > now, the results > > > > > > Yeb, Even, recompression case is fater than old but want to see more heavy memory > > > pressure case and the ratio I mentioned above. > > > > I did quite heavy testing over the last 7 days, with numerous OOM kills > > and OOM panics. > > Okay, I think it's worth to merge enough and see the result. > Please send formal patch which has recompression stat. ;-) correction: those 41-81s spikes in mem-hogger were observed under different scenario: 10GB zram with 6GB mem-hogger on a 4GB system. I'll do another round of tests (with parallel mem-hogger pre-fault and 4GB/4GB zram/mem-hogger split) and collect the number that you asked for. thanks! -ss
[toc] | [prev] | [next] | [standalone]
| From | Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> |
|---|---|
| Date | 2016-04-27 11:00 +0200 |
| Message-ID | <rstua-6Jw-15@gated-at.bofh.it> |
| In reply to | #1388720 |
Hello,
more tests. I did only 8streams vs per-cpu this time. the changes
to the test are:
-- mem-hogger now per-faults pages in parallel with fio
-- mem-hogger alloc size increased from 3GB to 4GB.
the system couldn't survive 4GB/4GB zram(buffer_compress_percentage=11)/mem-hogger
split (OOM), so I executed the 3GB/4GB test (close to system's OOM edge).
-- 4 GB x86_64
-- 3 GB zram lzo
firts, the mm_stat. <num_writes / num_recompressions>
8 streams (base kernel):
3221225472 3221225472 3221225472 0 3221229568 0 0 < 2752460/ 0>
3221225472 3221225472 3221225472 0 3221233664 0 0 < 5504124/ 0>
3221225472 2912157607 2952802304 0 2952826880 0 81 < 8253369/ 0>
3221225472 2893479936 2899120128 0 2899136512 0 147 <11003056/ 0>
3221217280 2886040814 2899103744 0 2899128320 0 26 <13748450/ 0>
3221225472 2880045056 2885693440 0 2885718016 0 180 <16503120/ 0>
3221213184 2877431364 2883756032 0 2883809280 0 132 <19259891/ 0>
3221225472 2873229312 2876096512 0 2876133376 0 16 <22016512/ 0>
3221213184 2870728008 2871693312 0 2871726080 0 24 <24768909/ 0>
2899095552 2899095552 2899095552 0 2899132416 78643 0 <27523600/ 0>
per-cpu:
3221225472 3221225472 3221225472 0 3221229568 0 0 < 2752460/ 8180>
3221225472 3221225472 3221225472 0 3221233664 0 0 < 5504124/ 10523>
3221225472 2912157607 2952802304 0 2952814592 0 117 < 8253369/ 9451>
3221225472 2893479936 2899120128 0 2899136512 0 129 <11003056/ 9395>
3221217280 2886040814 2899103744 0 2899128320 0 51 <13748450/ 10879>
3221225472 2880045056 2885693440 0 2885718016 0 126 <16503120/ 10300>
3221213184 2877431364 2883772416 0 2883801088 0 252 <19259891/ 10509>
3221225472 2873229312 2876100608 0 2876133376 0 14 <22016512/ 11081>
3221213184 2870728008 2871693312 0 2871730176 0 54 <24768909/ 10770>
2899095552 2899095552 2899095552 0 2899136512 78643 0 <27523600/ 10231>
mem-hogger pre-fault times
8 streams (base kernel):
[431] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f3f5d38a010 <+ 6.031550428>
[470] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7fa29d414010 <+ 5.242295692>
[514] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f4a7eac8010 <+ 5.485469454>
[563] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f07da76b010 <+ 5.563647658>
[619] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7ff5efc26010 <+ 5.516866208>
[681] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f8fb896d010 <+ 5.535275748>
[751] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7fb2ac6fa010 <+ 4.594626366>
[825] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f355f9a0010 <+ 5.075849029>
[905] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7feb16715010 <+ 4.696363680>
[991] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f3a1b9f4010 <+ 5.292365453>
per-cpu:
[413] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7fe8058f5010 <+ 5.513944292>
[451] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f65fe753010 <+ 4.742384977>
[494] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7fb99a05c010 <+ 5.394711696>
[542] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f0d61c81010 <+ 5.021011664>
[598] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f9abdeb6010 <+ 5.094722019>
[660] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7fb192ae9010 <+ 4.943961060>
[728] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f7313aeb010 <+ 5.437872456>
[802] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f25ffdeb010 <+ 5.422829590>
[881] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f60daa8e010 <+ 4.806425351>
[970] single-alloc: INFO: Allocated 0x100000000 bytes at address 0x7f384cf04010 <+ 4.982513395>
so, pre-fault time range is somewhat big. for example, from 4.696363680 to 6.031550428 seconds.
fio
8 streams per-cpu
===========================================
#jobs1
READ: 2507.8MB/s 2526.4MB/s
READ: 2043.1MB/s 1970.6MB/s
WRITE: 127100KB/s 139160KB/s
WRITE: 724488KB/s 733440KB/s
READ: 534624KB/s 540967KB/s
WRITE: 534569KB/s 540912KB/s
READ: 471165KB/s 477459KB/s
WRITE: 471233KB/s 477527KB/s
#jobs2
READ: 8041.1MB/s 7866.9MB/s
READ: 6751.7MB/s 6692.9MB/s
WRITE: 268372KB/s 268269KB/s
WRITE: 1197.5MB/s 1331.3MB/s
READ: 997.75MB/s 1057.6MB/s
WRITE: 997.91MB/s 1057.7MB/s
READ: 934518KB/s 992852KB/s
WRITE: 932382KB/s 990582KB/s
#jobs3
READ: 13318MB/s 13454MB/s
READ: 11463MB/s 11491MB/s
WRITE: 449903KB/s 448791KB/s
WRITE: 1582.8MB/s 1782.3MB/s
READ: 1337.5MB/s 1449.1MB/s
WRITE: 1335.6MB/s 1447.2MB/s
READ: 1241.3MB/s 1345.7MB/s
WRITE: 1239.6MB/s 1343.6MB/s
#jobs4
READ: 19948MB/s 20013MB/s
READ: 17732MB/s 17479MB/s
WRITE: 630690KB/s 495078KB/s
WRITE: 1843.2MB/s 2226.9MB/s
READ: 1603.4MB/s 1846.8MB/s
WRITE: 1599.4MB/s 1842.2MB/s
READ: 1547.7MB/s 1740.7MB/s
WRITE: 1549.2MB/s 1742.4MB/s
#jobs5
READ: 18800MB/s 4792.6MB/s
READ: 16659MB/s 16898MB/s
WRITE: 777796KB/s 721363KB/s
WRITE: 1771.9MB/s 2138.7MB/s
READ: 1517.9MB/s 1837.8MB/s
WRITE: 1512.6MB/s 1831.5MB/s
READ: 1501.4MB/s 1784.1MB/s
WRITE: 1500.5MB/s 1783.9MB/s
#jobs6
READ: 20827MB/s 20571MB/s
READ: 19382MB/s 19505MB/s
WRITE: 850618KB/s 776148KB/s
WRITE: 1886.2MB/s 2127.7MB/s
READ: 1685.3MB/s 1864.8MB/s
WRITE: 1682.4MB/s 1860.8MB/s
READ: 1598.3MB/s 1727.9MB/s
WRITE: 1593.5MB/s 1722.6MB/s
#jobs7
READ: 21547MB/s 21000MB/s
READ: 18814MB/s 18715MB/s
WRITE: 1008.5MB/s 991.56MB/s
WRITE: 1922.5MB/s 2232.9MB/s
READ: 1640.3MB/s 1795.2MB/s
WRITE: 1641.3MB/s 1796.4MB/s
READ: 1578.2MB/s 1763.2MB/s
WRITE: 1569.5MB/s 1753.5MB/s
#jobs8
READ: 20277MB/s 20916MB/s
READ: 17952MB/s 18340MB/s
WRITE: 1186.1MB/s 1170.6MB/s
WRITE: 1955.8MB/s 2347.6MB/s
READ: 1686.5MB/s 1936.7MB/s
WRITE: 1688.3MB/s 1938.8MB/s
READ: 1610.3MB/s 1894.7MB/s
WRITE: 1606.6MB/s 1890.4MB/s
#jobs9
READ: 20108MB/s 19361MB/s
READ: 18012MB/s 18177MB/s
WRITE: 1355.1MB/s 1325.7MB/s
WRITE: 1948.5MB/s 2305.6MB/s
READ: 1662.4MB/s 1892.6MB/s
WRITE: 1661.9MB/s 1891.5MB/s
READ: 1605.5MB/s 1812.4MB/s
WRITE: 1604.2MB/s 1810.7MB/s
#jobs10
READ: 20039MB/s 19455MB/s
READ: 18028MB/s 17716MB/s
WRITE: 1465.3MB/s 1486.7MB/s
WRITE: 2007.2MB/s 2317.5MB/s
READ: 1755.3MB/s 2005.9MB/s
WRITE: 1754.3MB/s 2003.1MB/s
READ: 1691.2MB/s 1874.8MB/s
WRITE: 1694.7MB/s 1877.7MB/s
perf stat
8 streams per-cpu
====================================================================================
jobs1
stalled-cycles-frontend 56,052,305,338 ( 55.05%) 58,803,628,418 ( 55.33%)
stalled-cycles-backend 24,355,709,967 ( 23.92%) 25,413,107,301 ( 23.91%)
instructions 96,175,143,640 ( 0.94) 100,364,109,185 ( 0.94)
branches 18,009,998,853 ( 559.201) 18,513,273,860 ( 550.828)
branch-misses 111,500,123 ( 0.62%) 106,011,616 ( 0.57%)
jobs2
stalled-cycles-frontend 126,012,750,354 ( 59.31%) 123,465,054,991 ( 57.62%)
stalled-cycles-backend 61,866,277,568 ( 29.12%) 58,959,838,855 ( 27.52%)
instructions 183,736,567,135 ( 0.86) 193,832,846,843 ( 0.90)
branches 34,879,279,141 ( 506.137) 36,700,776,757 ( 530.136)
branch-misses 175,122,665 ( 0.50%) 165,057,491 ( 0.45%)
jobs3
stalled-cycles-frontend 175,428,933,301 ( 60.40%) 160,530,802,385 ( 58.59%)
stalled-cycles-backend 87,409,032,068 ( 30.10%) 76,093,143,994 ( 27.77%)
instructions 231,949,985,071 ( 0.80) 229,875,149,073 ( 0.84)
branches 44,034,175,160 ( 423.325) 43,578,595,543 ( 443.250)
branch-misses 237,974,300 ( 0.54%) 216,380,848 ( 0.50%)
jobs4
stalled-cycles-frontend 265,519,049,536 ( 64.46%) 221,049,841,649 ( 61.81%)
stalled-cycles-backend 146,538,881,296 ( 35.57%) 113,774,053,039 ( 31.82%)
instructions 298,241,854,695 ( 0.72) 278,000,866,874 ( 0.78)
branches 59,531,800,053 ( 400.919) 55,096,944,109 ( 427.816)
branch-misses 285,108,083 ( 0.48%) 260,972,185 ( 0.47%)
jobs5
stalled-cycles-frontend 290,281,266,141 ( 64.97%) 260,946,337,232 ( 61.99%)
stalled-cycles-backend 161,884,390,707 ( 36.23%) 137,154,776,973 ( 32.58%)
instructions 326,306,594,233 ( 0.73) 334,011,271,525 ( 0.79)
branches 63,904,071,806 ( 398.348) 65,664,365,815 ( 434.393)
branch-misses 293,793,049 ( 0.46%) 279,794,621 ( 0.43%)
jobs6
stalled-cycles-frontend 344,523,942,841 ( 65.53%) 287,955,119,151 ( 61.88%)
stalled-cycles-backend 193,660,445,380 ( 36.84%) 150,866,799,639 ( 32.42%)
instructions 381,794,200,792 ( 0.73) 375,547,185,965 ( 0.81)
branches 74,623,783,129 ( 394.258) 73,649,248,349 ( 441.644)
branch-misses 358,005,680 ( 0.48%) 306,143,187 ( 0.42%)
jobs7
stalled-cycles-frontend 369,290,213,422 ( 64.74%) 319,117,228,349 ( 61.21%)
stalled-cycles-backend 206,236,039,426 ( 36.16%) 168,934,948,019 ( 32.40%)
instructions 427,549,938,405 ( 0.75) 429,752,151,831 ( 0.82)
branches 82,174,130,236 ( 400.051) 83,110,458,913 ( 442.306)
branch-misses 354,517,174 ( 0.43%) 332,430,584 ( 0.40%)
jobs8
stalled-cycles-frontend 409,541,894,683 ( 66.76%) 349,581,824,315 ( 62.51%)
stalled-cycles-backend 229,256,571,129 ( 37.37%) 181,622,273,772 ( 32.47%)
instructions 437,816,833,182 ( 0.71) 450,389,502,564 ( 0.81)
branches 84,525,812,473 ( 382.128) 87,501,121,276 ( 434.210)
branch-misses 372,309,759 ( 0.44%) 349,523,647 ( 0.40%)
jobs9
stalled-cycles-frontend 442,628,204,560 ( 65.90%) 380,475,695,919 ( 61.52%)
stalled-cycles-backend 251,927,332,399 ( 37.51%) 199,303,426,179 ( 32.22%)
instructions 491,437,868,336 ( 0.73) 511,514,729,028 ( 0.83)
branches 93,730,386,271 ( 386.978) 98,645,937,110 ( 442.304)
branch-misses 401,101,757 ( 0.43%) 368,882,924 ( 0.37%)
jobs10
stalled-cycles-frontend 478,576,939,331 ( 67.41%) 408,498,109,428 ( 63.43%)
stalled-cycles-backend 274,043,625,756 ( 38.60%) 219,162,314,972 ( 34.03%)
instructions 495,607,125,031 ( 0.70) 505,149,872,644 ( 0.78)
branches 95,885,616,294 ( 374.483) 99,094,367,930 ( 426.624)
branch-misses 418,267,387 ( 0.44%) 392,516,508 ( 0.40%)
perf reported execution time
8 streams per-cpu
====================================================
seconds elapsed 51.128322985 47.600230868
seconds elapsed 49.309895468 48.291538090
seconds elapsed 47.075673742 46.068406557
seconds elapsed 47.816933840 52.966896478
seconds elapsed 54.345548549 50.918853799
seconds elapsed 58.613938093 58.130571913
seconds elapsed 62.799745992 60.086779664
seconds elapsed 65.664854260 61.414515686
seconds elapsed 71.340920175 67.717224950
seconds elapsed 74.169664807 69.485210016
-ss
[toc] | [prev] | [next] | [standalone]
| From | Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> |
|---|---|
| Date | 2016-04-27 11:10 +0200 |
| Message-ID | <rstDR-779-25@gated-at.bofh.it> |
| In reply to | #1388802 |
On (04/27/16 17:54), Sergey Senozhatsky wrote: > #jobs4 > READ: 19948MB/s 20013MB/s > READ: 17732MB/s 17479MB/s > WRITE: 630690KB/s 495078KB/s > WRITE: 1843.2MB/s 2226.9MB/s > READ: 1603.4MB/s 1846.8MB/s > WRITE: 1599.4MB/s 1842.2MB/s > READ: 1547.7MB/s 1740.7MB/s > WRITE: 1549.2MB/s 1742.4MB/s > jobs4 > stalled-cycles-frontend 265,519,049,536 ( 64.46%) 221,049,841,649 ( 61.81%) > stalled-cycles-backend 146,538,881,296 ( 35.57%) 113,774,053,039 ( 31.82%) > instructions 298,241,854,695 ( 0.72) 278,000,866,874 ( 0.78) > branches 59,531,800,053 ( 400.919) 55,096,944,109 ( 427.816) > branch-misses 285,108,083 ( 0.48%) 260,972,185 ( 0.47%) > seconds elapsed 47.816933840 52.966896478 per-cpu in general looks better in this test (jobs4): less stalls, less branches, less misses, better fio speeds (except for WRITE: 630690KB/s 495078KB/s). the system was under pressure, so quite possible that it took more time to kill the process, thus execution time is in favor of 8 streams test. -ss
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web