Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1302362 > unrolled thread
| Started by | kernel test robot <ying.huang@linux.intel.com> |
|---|---|
| First post | 2016-01-06 04:30 +0100 |
| Last post | 2016-01-08 12:20 +0100 |
| Articles | 4 — 3 participants |
Back to article view | Back to linux.kernel
[lkp] [mm/vmstat] 6cdb18ad98: -8.5% will-it-scale.per_thread_ops kernel test robot <ying.huang@linux.intel.com> - 2016-01-06 04:30 +0100
Re: [lkp] [mm/vmstat] 6cdb18ad98: -8.5% will-it-scale.per_thread_ops Heiko Carstens <heiko.carstens@de.ibm.com> - 2016-01-07 12:30 +0100
Re: [LKP] [lkp] [mm/vmstat] 6cdb18ad98: -8.5% will-it-scale.per_thread_ops "Huang\, Ying" <ying.huang@intel.com> - 2016-01-08 06:30 +0100
Re: [LKP] [lkp] [mm/vmstat] 6cdb18ad98: -8.5% will-it-scale.per_thread_ops Heiko Carstens <heiko.carstens@de.ibm.com> - 2016-01-08 12:20 +0100
| From | kernel test robot <ying.huang@linux.intel.com> |
|---|---|
| Date | 2016-01-06 04:30 +0100 |
| Subject | [lkp] [mm/vmstat] 6cdb18ad98: -8.5% will-it-scale.per_thread_ops |
| Message-ID | <qNMXn-3XE-7@gated-at.bofh.it> |
[Multipart message — attachments visible in raw view] — view raw
FYI, we noticed the below changes on
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git master
commit 6cdb18ad98a49f7e9b95d538a0614cde827404b8 ("mm/vmstat: fix overflow in mod_zone_page_state()")
=========================================================================================
compiler/cpufreq_governor/kconfig/rootfs/tbox_group/test/testcase:
gcc-4.9/performance/x86_64-rhel/debian-x86_64-2015-02-07.cgz/ivb42/pread1/will-it-scale
commit:
cc28d6d80f6ab494b10f0e2ec949eacd610f66e3
6cdb18ad98a49f7e9b95d538a0614cde827404b8
cc28d6d80f6ab494 6cdb18ad98a49f7e9b95d538a0
---------------- --------------------------
%stddev %change %stddev
\ | \
2733943 ± 0% -8.5% 2502129 ± 0% will-it-scale.per_thread_ops
3410 ± 0% -2.0% 3343 ± 0% will-it-scale.time.system_time
340.08 ± 0% +19.7% 406.99 ± 0% will-it-scale.time.user_time
69882822 ± 2% -24.3% 52926191 ± 5% cpuidle.C1-IVT.time
340.08 ± 0% +19.7% 406.99 ± 0% time.user_time
491.25 ± 6% -17.7% 404.25 ± 7% numa-vmstat.node0.nr_alloc_batch
2799 ± 20% -36.6% 1776 ± 0% numa-vmstat.node0.nr_mapped
630.00 ±140% +244.4% 2169 ± 1% numa-vmstat.node1.nr_inactive_anon
6440 ± 11% -15.5% 5440 ± 16% numa-vmstat.node1.nr_slab_reclaimable
11204 ± 20% -36.6% 7106 ± 0% numa-meminfo.node0.Mapped
1017 ±173% +450.3% 5598 ± 15% numa-meminfo.node1.AnonHugePages
2521 ±140% +244.1% 8678 ± 1% numa-meminfo.node1.Inactive(anon)
25762 ± 11% -15.5% 21764 ± 16% numa-meminfo.node1.SReclaimable
70103 ± 9% -9.8% 63218 ± 9% numa-meminfo.node1.Slab
2.29 ± 3% +32.8% 3.04 ± 4% perf-profile.cycles-pp.atime_needs_update.touch_atime.shmem_file_read_iter.__vfs_read.vfs_read
1.10 ± 3% -27.4% 0.80 ± 5% perf-profile.cycles-pp.current_fs_time.atime_needs_update.touch_atime.shmem_file_read_iter.__vfs_read
2.33 ± 2% -13.0% 2.02 ± 3% perf-profile.cycles-pp.fput.entry_SYSCALL_64_fastpath
0.89 ± 2% +29.6% 1.15 ± 7% perf-profile.cycles-pp.fsnotify.vfs_read.sys_pread64.entry_SYSCALL_64_fastpath
2.85 ± 2% +45.4% 4.14 ± 5% perf-profile.cycles-pp.touch_atime.shmem_file_read_iter.__vfs_read.vfs_read.sys_pread64
63939 ± 0% +17.9% 75370 ± 15% sched_debug.cfs_rq:/.exec_clock.25
72.50 ± 73% -63.1% 26.75 ± 19% sched_debug.cfs_rq:/.load_avg.1
34.00 ± 62% -61.8% 13.00 ± 12% sched_debug.cfs_rq:/.load_avg.14
18.00 ± 11% -11.1% 16.00 ± 10% sched_debug.cfs_rq:/.load_avg.20
14.75 ± 41% +122.0% 32.75 ± 26% sched_debug.cfs_rq:/.load_avg.25
278.88 ± 11% +18.8% 331.25 ± 7% sched_debug.cfs_rq:/.load_avg.max
51.89 ± 11% +13.6% 58.97 ± 4% sched_debug.cfs_rq:/.load_avg.stddev
7.25 ± 5% +255.2% 25.75 ± 53% sched_debug.cfs_rq:/.runnable_load_avg.25
28.50 ± 1% +55.3% 44.25 ± 46% sched_debug.cfs_rq:/.runnable_load_avg.7
72.50 ± 73% -63.1% 26.75 ± 19% sched_debug.cfs_rq:/.tg_load_avg_contrib.1
34.00 ± 62% -61.8% 13.00 ± 12% sched_debug.cfs_rq:/.tg_load_avg_contrib.14
18.00 ± 11% -11.1% 16.00 ± 10% sched_debug.cfs_rq:/.tg_load_avg_contrib.20
14.75 ± 41% +122.0% 32.75 ± 25% sched_debug.cfs_rq:/.tg_load_avg_contrib.25
279.29 ± 11% +19.1% 332.67 ± 7% sched_debug.cfs_rq:/.tg_load_avg_contrib.max
52.01 ± 11% +13.8% 59.18 ± 4% sched_debug.cfs_rq:/.tg_load_avg_contrib.stddev
359.50 ± 6% +41.5% 508.75 ± 22% sched_debug.cfs_rq:/.util_avg.25
206.25 ± 16% -13.1% 179.25 ± 11% sched_debug.cfs_rq:/.util_avg.40
688.75 ± 1% +18.5% 816.00 ± 1% sched_debug.cfs_rq:/.util_avg.7
953467 ± 1% -17.9% 782518 ± 10% sched_debug.cpu.avg_idle.5
9177 ± 43% +73.9% 15957 ± 29% sched_debug.cpu.nr_switches.13
7365 ± 19% -35.4% 4755 ± 11% sched_debug.cpu.nr_switches.20
12203 ± 28% -62.2% 4608 ± 9% sched_debug.cpu.nr_switches.22
1868 ± 49% -51.1% 913.50 ± 27% sched_debug.cpu.nr_switches.27
2546 ± 56% -70.0% 763.00 ± 18% sched_debug.cpu.nr_switches.28
3003 ± 78% -77.9% 663.00 ± 18% sched_debug.cpu.nr_switches.33
1820 ± 19% +68.0% 3058 ± 33% sched_debug.cpu.nr_switches.8
-4.00 ±-35% -156.2% 2.25 ± 85% sched_debug.cpu.nr_uninterruptible.11
4.00 ±133% -187.5% -3.50 ±-24% sched_debug.cpu.nr_uninterruptible.17
1.75 ± 74% -214.3% -2.00 ±-127% sched_debug.cpu.nr_uninterruptible.25
0.00 ± 2% +Inf% 4.00 ± 39% sched_debug.cpu.nr_uninterruptible.26
2.50 ± 44% -110.0% -0.25 ±-591% sched_debug.cpu.nr_uninterruptible.27
1.33 ±154% -287.5% -2.50 ±-72% sched_debug.cpu.nr_uninterruptible.32
-1.00 ±-244% -250.0% 1.50 ±251% sched_debug.cpu.nr_uninterruptible.45
3.50 ± 82% -135.7% -1.25 ±-66% sched_debug.cpu.nr_uninterruptible.46
-4.50 ±-40% -133.3% 1.50 ±242% sched_debug.cpu.nr_uninterruptible.6
-3.00 ±-78% -433.3% 10.00 ±150% sched_debug.cpu.nr_uninterruptible.7
10124 ± 39% +65.8% 16783 ± 23% sched_debug.cpu.sched_count.13
12833 ± 23% -54.6% 5823 ± 32% sched_debug.cpu.sched_count.22
1934 ± 48% -49.8% 971.00 ± 26% sched_debug.cpu.sched_count.27
3065 ± 76% -76.2% 728.25 ± 16% sched_debug.cpu.sched_count.33
2098 ± 24% +664.1% 16030 ±126% sched_debug.cpu.sched_count.5
4653 ± 33% +83.4% 8536 ± 25% sched_debug.cpu.sched_goidle.15
5061 ± 41% -61.1% 1968 ± 13% sched_debug.cpu.sched_goidle.22
834.75 ± 57% -60.2% 332.00 ± 35% sched_debug.cpu.sched_goidle.27
719.00 ± 71% -63.3% 264.00 ± 19% sched_debug.cpu.sched_goidle.28
943.25 ±115% -76.3% 223.25 ± 21% sched_debug.cpu.sched_goidle.33
2520 ± 26% +112.4% 5353 ± 19% sched_debug.cpu.ttwu_count.13
5324 ± 22% -49.7% 2679 ± 45% sched_debug.cpu.ttwu_count.22
2926 ± 38% +231.1% 9690 ± 37% sched_debug.cpu.ttwu_count.23
277.25 ± 18% +166.7% 739.50 ± 83% sched_debug.cpu.ttwu_count.27
1247 ± 61% -76.6% 292.25 ± 11% sched_debug.cpu.ttwu_count.28
751.75 ± 22% +183.9% 2134 ± 9% sched_debug.cpu.ttwu_count.3
6405 ± 97% -75.9% 1542 ± 48% sched_debug.cpu.ttwu_count.41
5582 ±104% -76.2% 1327 ± 55% sched_debug.cpu.ttwu_count.43
3201 ± 26% -75.1% 796.75 ± 18% sched_debug.cpu.ttwu_local.22
ivb42: Ivytown Ivy Bridge-EP
Memory: 64G
To reproduce:
git clone git://git.kernel.org/pub/scm/linux/kernel/git/wfg/lkp-tests.git
cd lkp-tests
bin/lkp install job.yaml # job file is attached in this email
bin/lkp run job.yaml
Disclaimer:
Results have been estimated based on internal Intel analysis and are provided
for informational purposes only. Any difference in system hardware or software
design or configuration may affect actual performance.
Thanks,
Ying Huang
[toc] | [next] | [standalone]
| From | Heiko Carstens <heiko.carstens@de.ibm.com> |
|---|---|
| Date | 2016-01-07 12:30 +0100 |
| Message-ID | <qOgVs-7pv-23@gated-at.bofh.it> |
| In reply to | #1302362 |
On Wed, Jan 06, 2016 at 11:20:55AM +0800, kernel test robot wrote:
> FYI, we noticed the below changes on
>
> https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git master
> commit 6cdb18ad98a49f7e9b95d538a0614cde827404b8 ("mm/vmstat: fix overflow in mod_zone_page_state()")
>
>
> =========================================================================================
> compiler/cpufreq_governor/kconfig/rootfs/tbox_group/test/testcase:
> gcc-4.9/performance/x86_64-rhel/debian-x86_64-2015-02-07.cgz/ivb42/pread1/will-it-scale
>
> commit:
> cc28d6d80f6ab494b10f0e2ec949eacd610f66e3
> 6cdb18ad98a49f7e9b95d538a0614cde827404b8
>
> cc28d6d80f6ab494 6cdb18ad98a49f7e9b95d538a0
> ---------------- --------------------------
> %stddev %change %stddev
> \ | \
> 2733943 ± 0% -8.5% 2502129 ± 0% will-it-scale.per_thread_ops
> 3410 ± 0% -2.0% 3343 ± 0% will-it-scale.time.system_time
> 340.08 ± 0% +19.7% 406.99 ± 0% will-it-scale.time.user_time
> 69882822 ± 2% -24.3% 52926191 ± 5% cpuidle.C1-IVT.time
> 340.08 ± 0% +19.7% 406.99 ± 0% time.user_time
> 491.25 ± 6% -17.7% 404.25 ± 7% numa-vmstat.node0.nr_alloc_batch
> 2799 ± 20% -36.6% 1776 ± 0% numa-vmstat.node0.nr_mapped
> 630.00 ±140% +244.4% 2169 ± 1% numa-vmstat.node1.nr_inactive_anon
Hmm... this is odd. I did review all callers of mod_zone_page_state() and
couldn't find anything obvious that would go wrong after the int -> long
change.
I also tried the "pread1_threads" test case from
https://github.com/antonblanchard/will-it-scale.git
However the results seem to vary a lot after a reboot(!), at least on s390.
So I'm not sure if this is really a regression.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Huang\, Ying" <ying.huang@intel.com> |
|---|---|
| Date | 2016-01-08 06:30 +0100 |
| Subject | Re: [LKP] [lkp] [mm/vmstat] 6cdb18ad98: -8.5% will-it-scale.per_thread_ops |
| Message-ID | <qOxMC-28c-13@gated-at.bofh.it> |
| In reply to | #1303494 |
Heiko Carstens <heiko.carstens@de.ibm.com> writes:
> On Wed, Jan 06, 2016 at 11:20:55AM +0800, kernel test robot wrote:
>> FYI, we noticed the below changes on
>>
>> https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git master
>> commit 6cdb18ad98a49f7e9b95d538a0614cde827404b8 ("mm/vmstat: fix overflow in mod_zone_page_state()")
>>
>>
>> =========================================================================================
>> compiler/cpufreq_governor/kconfig/rootfs/tbox_group/test/testcase:
>> gcc-4.9/performance/x86_64-rhel/debian-x86_64-2015-02-07.cgz/ivb42/pread1/will-it-scale
>>
>> commit:
>> cc28d6d80f6ab494b10f0e2ec949eacd610f66e3
>> 6cdb18ad98a49f7e9b95d538a0614cde827404b8
>>
>> cc28d6d80f6ab494 6cdb18ad98a49f7e9b95d538a0
>> ---------------- --------------------------
>> %stddev %change %stddev
>> \ | \
>> 2733943 0% -8.5% 2502129 0% will-it-scale.per_thread_ops
>> 3410 0% -2.0% 3343 0% will-it-scale.time.system_time
>> 340.08 0% +19.7% 406.99 0% will-it-scale.time.user_time
>> 69882822 2% -24.3% 52926191 5% cpuidle.C1-IVT.time
>> 340.08 0% +19.7% 406.99 0% time.user_time
>> 491.25 6% -17.7% 404.25 7% numa-vmstat.node0.nr_alloc_batch
>> 2799 20% -36.6% 1776 0% numa-vmstat.node0.nr_mapped
>> 630.00 140% +244.4% 2169 1% numa-vmstat.node1.nr_inactive_anon
>
> Hmm... this is odd. I did review all callers of mod_zone_page_state() and
> couldn't find anything obvious that would go wrong after the int -> long
> change.
>
> I also tried the "pread1_threads" test case from
> https://github.com/antonblanchard/will-it-scale.git
>
> However the results seem to vary a lot after a reboot(!), at least on s390.
>
> So I'm not sure if this is really a regression.
The test is quite stable for my side. We run the test case 7 times for
your commit and its parent. The standard variation is very low.
you commit:
[2493136, 2510964, 2508784, 2495632, 2506735, 2503016, 2510121]
parent commit:
[2735669, 2719566, 2739052, 2741485, 2735152, 2739356, 2739125]
The test result is stable for bisection too. The below figure show the
results of good commits and bad commits. The distance between is quite
big. And the variation is quite small.
will-it-scale.per_thread_ops
2.75e+06 ++--*---*--------------*---*------*---*---*-*-*-*----------*---*-+
*.* + + + .*. .*.*.* + + + .*.* + + + + *.*.**.* * *
2.7e+06 ++ * ** * * * * * |
| |
| |
2.65e+06 ++ |
| |
2.6e+06 ++ |
| |
2.55e+06 ++ |
| |
O O O O O O O O O O O O O O O |
2.5e+06 ++ O OO O O O O O O O |
| |
2.45e+06 ++---------------------------------------------------------------+
FYI, I test your patch on x86 platform. I have no s390 system.
Best Regards,
Huang, Ying
[toc] | [prev] | [next] | [standalone]
| From | Heiko Carstens <heiko.carstens@de.ibm.com> |
|---|---|
| Date | 2016-01-08 12:20 +0100 |
| Subject | Re: [LKP] [lkp] [mm/vmstat] 6cdb18ad98: -8.5% will-it-scale.per_thread_ops |
| Message-ID | <qODfk-5UZ-21@gated-at.bofh.it> |
| In reply to | #1304186 |
On Fri, Jan 08, 2016 at 01:24:30PM +0800, Huang, Ying wrote:
> Heiko Carstens <heiko.carstens@de.ibm.com> writes:
>
> > On Wed, Jan 06, 2016 at 11:20:55AM +0800, kernel test robot wrote:
> >> FYI, we noticed the below changes on
> >>
> >> https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git master
> >> commit 6cdb18ad98a49f7e9b95d538a0614cde827404b8 ("mm/vmstat: fix overflow in mod_zone_page_state()")
> >>
> >>
> >> =========================================================================================
> >> compiler/cpufreq_governor/kconfig/rootfs/tbox_group/test/testcase:
> >> gcc-4.9/performance/x86_64-rhel/debian-x86_64-2015-02-07.cgz/ivb42/pread1/will-it-scale
> >>
> >> commit:
> >> cc28d6d80f6ab494b10f0e2ec949eacd610f66e3
> >> 6cdb18ad98a49f7e9b95d538a0614cde827404b8
> >>
> >> cc28d6d80f6ab494 6cdb18ad98a49f7e9b95d538a0
> >> ---------------- --------------------------
> >> %stddev %change %stddev
> >> \ | \
> >> 2733943 0% -8.5% 2502129 0% will-it-scale.per_thread_ops
> >> 3410 0% -2.0% 3343 0% will-it-scale.time.system_time
> >> 340.08 0% +19.7% 406.99 0% will-it-scale.time.user_time
> >> 69882822 2% -24.3% 52926191 5% cpuidle.C1-IVT.time
> >> 340.08 0% +19.7% 406.99 0% time.user_time
> >> 491.25 6% -17.7% 404.25 7% numa-vmstat.node0.nr_alloc_batch
> >> 2799 20% -36.6% 1776 0% numa-vmstat.node0.nr_mapped
> >> 630.00 140% +244.4% 2169 1% numa-vmstat.node1.nr_inactive_anon
> >
> > Hmm... this is odd. I did review all callers of mod_zone_page_state() and
> > couldn't find anything obvious that would go wrong after the int -> long
> > change.
> >
> > I also tried the "pread1_threads" test case from
> > https://github.com/antonblanchard/will-it-scale.git
> >
> > However the results seem to vary a lot after a reboot(!), at least on s390.
> >
> > So I'm not sure if this is really a regression.
>
> The test is quite stable for my side. We run the test case 7 times for
> your commit and its parent. The standard variation is very low.
>
> you commit:
>
> [2493136, 2510964, 2508784, 2495632, 2506735, 2503016, 2510121]
>
> parent commit:
>
> [2735669, 2719566, 2739052, 2741485, 2735152, 2739356, 2739125]
>
> The test result is stable for bisection too. The below figure show the
> results of good commits and bad commits. The distance between is quite
> big. And the variation is quite small.
Ok, so it seems to be quite stable on your machine across reboots.
I have to admit I still cannot make much sense of this. Is the "pread1"
testcase the only one that performs worse, or are there more?
Also could you please provide the output of /proc/zoneinfo and the output
of "perf top" of the good/bad cases? Maybe that might help to figure out
what is happening.
> FYI, I test your patch on x86 platform. I have no s390 system.
Sure, I wouldn't expect that.
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web