Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1463938 > unrolled thread
| Started by | "H. Peter Anvin" <hpa@zytor.com> |
|---|---|
| First post | 2016-08-16 19:00 +0200 |
| Last post | 2016-08-17 08:50 +0200 |
| Articles | 10 — 4 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement "H. Peter Anvin" <hpa@zytor.com> - 2016-08-16 19:00 +0200
Re: [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement Borislav Petkov <bp@suse.de> - 2016-08-16 19:20 +0200
Re: [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement "H. Peter Anvin" <hpa@zytor.com> - 2016-08-17 01:20 +0200
Re: [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement Borislav Petkov <bp@suse.de> - 2016-08-17 07:50 +0200
Re: [LKP] [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement "Huang\, Ying" <ying.huang@intel.com> - 2016-08-18 00:30 +0200
Re: [LKP] [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement Borislav Petkov <bp@suse.de> - 2016-08-18 05:50 +0200
Re: [LKP] [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement "Huang\, Ying" <ying.huang@intel.com> - 2016-08-18 06:00 +0200
Re: [LKP] [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement Borislav Petkov <bp@suse.de> - 2016-08-18 06:20 +0200
Re: [LKP] [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement "H. Peter Anvin" <hpa@zytor.com> - 2016-08-18 06:00 +0200
Re: [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement Peter Zijlstra <peterz@infradead.org> - 2016-08-17 08:50 +0200
| From | "H. Peter Anvin" <hpa@zytor.com> |
|---|---|
| Date | 2016-08-16 19:00 +0200 |
| Subject | Re: [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement |
| Message-ID | <s6PSx-5xO-1@gated-at.bofh.it> |
On August 16, 2016 7:26:43 AM PDT, kernel test robot <xiaolong.ye@intel.com> wrote:
>
>FYI, we noticed a 9.3% improvement of will-it-scale.per_process_ops due
>to commit:
>
>commit 65ea11ec6a82b1d44aba62b59e9eb20247e57c6e ("x86/hweight: Don't
>clobber %rdi")
>https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
>master
>
>in testcase: will-it-scale
>on test machine: 32 threads Sandy Bridge-EP with 64G memory
>with following parameters:
>
> test: unix1
> cpufreq_governor: performance
>
>
>Disclaimer:
>Results have been estimated based on internal Intel analysis and are
>provided
>for informational purposes only. Any difference in system hardware or
>software
>design or configuration may affect actual performance.
>
>Details are as below:
>-------------------------------------------------------------------------------------------------->
>
>
>To reproduce:
>
>git clone
>git://git.kernel.org/pub/scm/linux/kernel/git/wfg/lkp-tests.git
> cd lkp-tests
> bin/lkp install job.yaml # job file is attached in this email
> bin/lkp run job.yaml
>
>=========================================================================================
>compiler/cpufreq_governor/kconfig/rootfs/tbox_group/test/testcase:
>gcc-6/performance/x86_64-rhel/debian-x86_64-2015-02-07.cgz/lkp-sb03/unix1/will-it-scale
>
>commit:
> v4.8-rc1
> 65ea11ec6a ("x86/hweight: Don't clobber %rdi")
>
> v4.8-rc1 65ea11ec6a82b1d44aba62b59e
>---------------- --------------------------
> fail:runs %reproduction fail:runs
> | | |
> 1:8 -12% :4 last_state.is_incomplete_run
>4:8 -50% :4
>kmsg.DHCP/BOOTP:Reply_not_for_us,op[#]xid[#]
>7:8 -88% :4
>kmsg.drm:drm_edid_block_valid[drm]]*ERROR*EDID_checksum_is_invalid,remainder_is
>7:8 -88% :4
>kmsg.i8042:Can't_read_CTR_while_initializing_i8042
> %stddev %change %stddev
> \ | \
>1063041 ± 0% +9.3% 1161810 ± 0%
>will-it-scale.per_process_ops
> 976004 ± 0% +9.0% 1063615 ± 0% will-it-scale.per_thread_ops
> 0.57 ± 0% -6.7% 0.53 ± 1% will-it-scale.scalability
> 175.96 ± 0% +8.0% 190.10 ± 0% will-it-scale.time.user_time
>0.00 ± 20% -31.5% 0.00 ± 26%
>sched_debug.cpu.next_balance.stddev
>101.14 ± 11% +9639.4% 9850 ±121%
>latency_stats.avg.rpc_wait_bit_killable.__rpc_execute.rpc_execute.rpc_run_task.nfs4_call_sync_sequence.[nfsv4]._nfs4_proc_getattr.[nfsv4].nfs4_proc_getattr.[nfsv4].__nfs_revalidate_inode.nfs_do_access.nfs_permission.__inode_permission.inode_permission
>148.57 ± 15% +57704.4% 85880 ±125%
>latency_stats.max.rpc_wait_bit_killable.__rpc_execute.rpc_execute.rpc_run_task.nfs4_call_sync_sequence.[nfsv4]._nfs4_proc_getattr.[nfsv4].nfs4_proc_getattr.[nfsv4].__nfs_revalidate_inode.nfs_do_access.nfs_permission.__inode_permission.inode_permission
>886.00 ± 14% +9757.0% 87333 ±123%
>latency_stats.sum.rpc_wait_bit_killable.__rpc_execute.rpc_execute.rpc_run_task.nfs4_call_sync_sequence.[nfsv4]._nfs4_proc_getattr.[nfsv4].nfs4_proc_getattr.[nfsv4].__nfs_revalidate_inode.nfs_do_access.nfs_permission.__inode_permission.inode_permission
>3.041e+12 ± 1% +7.4% 3.267e+12 ± 1%
>perf-stat.branch-instructions
> 0.31 ± 0% -86.6% 0.04 ± 4% perf-stat.branch-miss-rate
> 9.456e+09 ± 1% -85.6% 1.364e+09 ± 3% perf-stat.branch-misses
> 5.147e+12 ± 1% +5.4% 5.427e+12 ± 1% perf-stat.dTLB-loads
> 3.869e+12 ± 0% +6.7% 4.128e+12 ± 1% perf-stat.dTLB-stores
> 29.02 ± 13% +223.2% 93.80 ± 0% perf-stat.iTLB-load-miss-rate
>2.353e+08 ± 21% +733.0% 1.96e+09 ± 0% perf-stat.iTLB-load-misses
> 5.7e+08 ± 9% -77.2% 1.297e+08 ± 10% perf-stat.iTLB-loads
> 1.696e+13 ± 0% +6.9% 1.814e+13 ± 0% perf-stat.instructions
>75030 ± 18% -87.7% 9251 ± 1%
>perf-stat.instructions-per-iTLB-miss
> 1.04 ± 0% +7.6% 1.12 ± 1% perf-stat.ipc
> 24064971 ± 3% -6.6% 22469931 ± 3% perf-stat.node-load-misses
> 53705459 ± 1% -3.1% 52034054 ± 2% perf-stat.node-loads
>7.32 ± 5% +23.3% 9.03 ± 4%
>perf-profile.cycles.__alloc_skb.alloc_skb_with_frags.sock_alloc_send_pskb.unix_stream_sendmsg.sock_sendmsg
>1.29 ± 4% +11.7% 1.44 ± 5%
>perf-profile.cycles.__fdget_pos.sys_write.entry_SYSCALL_64_fastpath
>1.15 ± 4% +12.1% 1.29 ± 4%
>perf-profile.cycles.__fget.__fget_light.__fdget_pos.sys_write.entry_SYSCALL_64_fastpath
>1.22 ± 5% +11.7% 1.36 ± 5%
>perf-profile.cycles.__fget_light.__fdget_pos.sys_write.entry_SYSCALL_64_fastpath
>1.86 ± 4% -58.4% 0.77 ± 7%
>perf-profile.cycles.__inode_security_revalidate.selinux_file_permission.security_file_permission.rw_verify_area.vfs_write
>0.00 ± -1% +Inf% 2.65 ± 5%
>perf-profile.cycles.__kmalloc_node_track_caller.__kmalloc_reserve.isra.33.__alloc_skb.alloc_skb_with_frags.sock_alloc_send_pskb
>1.89 ± 8% -100.0% 0.00 ± -1%
>perf-profile.cycles.__kmalloc_node_track_caller.__kmalloc_reserve.isra.35.__alloc_skb.alloc_skb_with_frags.sock_alloc_send_pskb
>0.00 ± -1% +Inf% 3.55 ± 5%
>perf-profile.cycles.__kmalloc_reserve.isra.33.__alloc_skb.alloc_skb_with_frags.sock_alloc_send_pskb.unix_stream_sendmsg
>2.52 ± 8% -100.0% 0.00 ± -1%
>perf-profile.cycles.__kmalloc_reserve.isra.35.__alloc_skb.alloc_skb_with_frags.sock_alloc_send_pskb.unix_stream_sendmsg
>1.43 ± 4% -91.1% 0.13 ±173%
>perf-profile.cycles.__might_sleep.__inode_security_revalidate.selinux_file_permission.security_file_permission.rw_verify_area
>1.15 ± 5% -65.7% 0.40 ± 57%
>perf-profile.cycles.__might_sleep.mutex_lock.unix_stream_read_generic.unix_stream_recvmsg.sock_recvmsg
>1.33 ± 7% +14.0% 1.52 ± 2%
>perf-profile.cycles._raw_spin_lock_irqsave.skb_queue_tail.unix_stream_sendmsg.sock_sendmsg.sock_write_iter
>1.37 ± 6% +20.4% 1.65 ± 3%
>perf-profile.cycles._raw_spin_lock_irqsave.skb_unlink.unix_stream_read_generic.unix_stream_recvmsg.sock_recvmsg
>1.09 ± 9% +15.6% 1.26 ± 5%
>perf-profile.cycles._raw_spin_unlock_irqrestore.skb_queue_tail.unix_stream_sendmsg.sock_sendmsg.sock_write_iter
>1.01 ± 6% +15.4% 1.17 ± 7%
>perf-profile.cycles._raw_spin_unlock_irqrestore.skb_unlink.unix_stream_read_generic.unix_stream_recvmsg.sock_recvmsg
>8.01 ± 6% +22.5% 9.82 ± 4%
>perf-profile.cycles.alloc_skb_with_frags.sock_alloc_send_pskb.unix_stream_sendmsg.sock_sendmsg.sock_write_iter
>7.33 ± 6% +14.8% 8.42 ± 4%
>perf-profile.cycles.consume_skb.unix_stream_read_generic.unix_stream_recvmsg.sock_recvmsg.sock_read_iter
>0.98 ± 8% +15.0% 1.12 ± 4%
>perf-profile.cycles.consume_skb.unix_stream_recvmsg.sock_recvmsg.sock_read_iter.__vfs_read
>1.60 ± 5% +18.7% 1.91 ± 3%
>perf-profile.cycles.copy_from_iter.skb_copy_datagram_from_iter.unix_stream_sendmsg.sock_sendmsg.sock_write_iter
>2.30 ± 4% +11.5% 2.56 ± 6%
>perf-profile.cycles.entry_SYSCALL_64
>2.10 ± 3% +18.1% 2.48 ± 5%
>perf-profile.cycles.entry_SYSCALL_64_after_swapgs
>2.82 ± 7% -34.6% 1.85 ± 6%
>perf-profile.cycles.file_has_perm.selinux_file_permission.security_file_permission.rw_verify_area.vfs_read
>1.55 ± 6% +21.3% 1.89 ± 5%
>perf-profile.cycles.fput.entry_SYSCALL_64_fastpath
>1.13 ± 9% +17.0% 1.32 ± 3%
>perf-profile.cycles.kfree.skb_free_head.skb_release_data.skb_release_all.consume_skb
>0.76 ± 8% +21.9% 0.93 ± 5%
>perf-profile.cycles.kfree_skbmem.consume_skb.unix_stream_read_generic.unix_stream_recvmsg.sock_recvmsg
>0.77 ± 10% +27.0% 0.98 ± 5%
>perf-profile.cycles.ksize.__alloc_skb.alloc_skb_with_frags.sock_alloc_send_pskb.unix_stream_sendmsg
>2.08 ± 6% -31.5% 1.42 ± 6%
>perf-profile.cycles.mutex_lock.unix_stream_read_generic.unix_stream_recvmsg.sock_recvmsg.sock_read_iter
>0.89 ± 9% +18.8% 1.06 ± 6%
>perf-profile.cycles.mutex_unlock.unix_stream_recvmsg.sock_recvmsg.sock_read_iter.__vfs_read
>6.80 ± 3% -19.3% 5.49 ± 3%
>perf-profile.cycles.rw_verify_area.vfs_read.sys_read.entry_SYSCALL_64_fastpath
>5.54 ± 4% -23.5% 4.24 ± 5%
>perf-profile.cycles.rw_verify_area.vfs_write.sys_write.entry_SYSCALL_64_fastpath
>6.21 ± 4% -19.5% 5.00 ± 3%
>perf-profile.cycles.security_file_permission.rw_verify_area.vfs_read.sys_read.entry_SYSCALL_64_fastpath
>5.23 ± 4% -25.6% 3.89 ± 5%
>perf-profile.cycles.security_file_permission.rw_verify_area.vfs_write.sys_write.entry_SYSCALL_64_fastpath
>4.67 ± 4% -24.1% 3.55 ± 4%
>perf-profile.cycles.selinux_file_permission.security_file_permission.rw_verify_area.vfs_read.sys_read
>4.87 ± 5% -28.0% 3.51 ± 5%
>perf-profile.cycles.selinux_file_permission.security_file_permission.rw_verify_area.vfs_write.sys_write
>2.43 ± 5% +29.8% 3.15 ± 3%
>perf-profile.cycles.skb_copy_datagram_from_iter.unix_stream_sendmsg.sock_sendmsg.sock_write_iter.__vfs_write
>1.18 ± 8% +16.1% 1.36 ± 2%
>perf-profile.cycles.skb_free_head.skb_release_data.skb_release_all.consume_skb.unix_stream_read_generic
>2.60 ± 7% +15.4% 3.00 ± 3%
>perf-profile.cycles.skb_queue_tail.unix_stream_sendmsg.sock_sendmsg.sock_write_iter.__vfs_write
>6.30 ± 6% +15.2% 7.26 ± 4%
>perf-profile.cycles.skb_release_all.consume_skb.unix_stream_read_generic.unix_stream_recvmsg.sock_recvmsg
>1.45 ± 7% +19.4% 1.73 ± 2%
>perf-profile.cycles.skb_release_data.skb_release_all.consume_skb.unix_stream_read_generic.unix_stream_recvmsg
>4.63 ± 6% +14.4% 5.30 ± 5%
>perf-profile.cycles.skb_release_head_state.skb_release_all.consume_skb.unix_stream_read_generic.unix_stream_recvmsg
>1.01 ± 4% +16.7% 1.18 ± 5%
>perf-profile.cycles.skb_set_owner_w.sock_alloc_send_pskb.unix_stream_sendmsg.sock_sendmsg.sock_write_iter
>2.59 ± 6% +18.2% 3.07 ± 4%
>perf-profile.cycles.skb_unlink.unix_stream_read_generic.unix_stream_recvmsg.sock_recvmsg.sock_read_iter
>9.66 ± 5% +21.1% 11.70 ± 3%
>perf-profile.cycles.sock_alloc_send_pskb.unix_stream_sendmsg.sock_sendmsg.sock_write_iter.__vfs_write
>25.86 ± 5% +14.8% 29.68 ± 4%
>perf-profile.cycles.sock_sendmsg.sock_write_iter.__vfs_write.vfs_write.sys_write
>3.88 ± 7% +13.1% 4.38 ± 5%
>perf-profile.cycles.sock_wfree.unix_destruct_scm.skb_release_head_state.skb_release_all.consume_skb
>4.24 ± 7% +13.3% 4.80 ± 5%
>perf-profile.cycles.unix_destruct_scm.skb_release_head_state.skb_release_all.consume_skb.unix_stream_read_generic
>21.96 ± 5% +17.1% 25.71 ± 3%
>perf-profile.cycles.unix_stream_sendmsg.sock_sendmsg.sock_write_iter.__vfs_write.vfs_write
>1.20 ± 6% -100.0% 0.00 ± -1%
>perf-profile.cycles.unix_stream_sendmsg.sock_write_iter.__vfs_write.vfs_write.sys_write
>2.28 ± 6% +13.7% 2.60 ± 3%
>perf-profile.cycles.unix_write_space.sock_wfree.unix_destruct_scm.skb_release_head_state.skb_release_all
>3.84 ± 5% -16.8% 3.20 ± 2%
>perf-profile.func.cycles.___might_sleep
>1.96 ± 7% +20.8% 2.36 ± 4%
>perf-profile.func.cycles.__alloc_skb
>2.40 ± 4% +11.3% 2.67 ± 4% perf-profile.func.cycles.__fget
>1.30 ± 9% +48.7% 1.94 ± 4%
>perf-profile.func.cycles.__kmalloc_node_track_caller
>1.05 ± 5% +12.6% 1.19 ± 7%
>perf-profile.func.cycles.__vfs_read
>0.99 ± 7% +27.1% 1.26 ± 4%
>perf-profile.func.cycles.__vfs_write
>1.01 ± 5% -51.9% 0.48 ± 3%
>perf-profile.func.cycles._cond_resched
>2.78 ± 6% +17.0% 3.25 ± 2%
>perf-profile.func.cycles._raw_spin_lock_irqsave
>2.19 ± 8% +15.5% 2.53 ± 6%
>perf-profile.func.cycles._raw_spin_unlock_irqrestore
>1.10 ± 8% +11.2% 1.23 ± 4%
>perf-profile.func.cycles.consume_skb
>0.97 ± 5% +25.6% 1.22 ± 3%
>perf-profile.func.cycles.copy_from_iter
>2.30 ± 4% +11.5% 2.56 ± 6%
>perf-profile.func.cycles.entry_SYSCALL_64
>2.10 ± 3% +18.1% 2.48 ± 5%
>perf-profile.func.cycles.entry_SYSCALL_64_after_swapgs
>2.26 ± 4% -38.4% 1.39 ± 5%
>perf-profile.func.cycles.file_has_perm
> 1.55 ± 6% +21.3% 1.89 ± 5% perf-profile.func.cycles.fput
> 1.18 ± 8% +17.2% 1.38 ± 3% perf-profile.func.cycles.kfree
> 0.86 ± 10% +22.0% 1.05 ± 4% perf-profile.func.cycles.ksize
>0.90 ± 8% +18.7% 1.06 ± 5%
>perf-profile.func.cycles.mutex_unlock
>1.91 ± 6% -13.1% 1.66 ± 3%
>perf-profile.func.cycles.selinux_file_permission
>1.05 ± 5% +16.7% 1.23 ± 5%
>perf-profile.func.cycles.skb_set_owner_w
>1.66 ± 8% +16.3% 1.93 ± 7%
>perf-profile.func.cycles.sock_wfree
>2.44 ± 4% -39.7% 1.47 ± 2%
>perf-profile.func.cycles.sock_write_iter
>4.20 ± 6% -21.1% 3.32 ± 3%
>perf-profile.func.cycles.unix_stream_sendmsg
>2.35 ± 6% +14.3% 2.69 ± 3%
>perf-profile.func.cycles.unix_write_space
>
>
>
>Thanks,
>Xiaolong
Dang...
--
Sent from my Android device with K-9 Mail. Please excuse brevity and formatting.
[toc] | [next] | [standalone]
| From | Borislav Petkov <bp@suse.de> |
|---|---|
| Date | 2016-08-16 19:20 +0200 |
| Subject | Re: [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement |
| Message-ID | <s6QbU-5V1-33@gated-at.bofh.it> |
| In reply to | #1463938 |
On Tue, Aug 16, 2016 at 09:59:00AM -0700, H. Peter Anvin wrote:
> Dang...
Isn't 9.3% improvement a good thing(tm) ?
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
SUSE Linux GmbH, GF: Felix Imendörffer, Jane Smithard, Graham Norton, HRB 21284 (AG Nürnberg)
--
[toc] | [prev] | [next] | [standalone]
| From | "H. Peter Anvin" <hpa@zytor.com> |
|---|---|
| Date | 2016-08-17 01:20 +0200 |
| Message-ID | <s6VOh-161-1@gated-at.bofh.it> |
| In reply to | #1463954 |
On August 16, 2016 10:16:35 AM PDT, Borislav Petkov <bp@suse.de> wrote: >On Tue, Aug 16, 2016 at 09:59:00AM -0700, H. Peter Anvin wrote: >> Dang... > >Isn't 9.3% improvement a good thing(tm) ? Yes, it's huge. The only explanation I could imagine is that scrambling %rdi caused the scheduler to do completely the wrong thing. -- Sent from my Android device with K-9 Mail. Please excuse brevity and formatting.
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@suse.de> |
|---|---|
| Date | 2016-08-17 07:50 +0200 |
| Subject | Re: [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement |
| Message-ID | <s71TH-5aR-1@gated-at.bofh.it> |
| In reply to | #1464230 |
On Tue, Aug 16, 2016 at 04:09:19PM -0700, H. Peter Anvin wrote:
> On August 16, 2016 10:16:35 AM PDT, Borislav Petkov <bp@suse.de> wrote:
> >On Tue, Aug 16, 2016 at 09:59:00AM -0700, H. Peter Anvin wrote:
> >> Dang...
> >
> >Isn't 9.3% improvement a good thing(tm) ?
>
> Yes, it's huge. The only explanation I could imagine is that scrambling %rdi caused the scheduler to do completely the wrong thing.
I'm questioning the validity, actually. Report says test machine was
Sandy Bridge-EP and I'd bet good money this one has POPCNT support so
how are we even hitting that __sw_hweight64() path, at all?
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
SUSE Linux GmbH, GF: Felix Imendörffer, Jane Smithard, Graham Norton, HRB 21284 (AG Nürnberg)
--
[toc] | [prev] | [next] | [standalone]
| From | "Huang\, Ying" <ying.huang@intel.com> |
|---|---|
| Date | 2016-08-18 00:30 +0200 |
| Subject | Re: [LKP] [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement |
| Message-ID | <s7hvs-7wd-15@gated-at.bofh.it> |
| In reply to | #1464328 |
Borislav Petkov <bp@suse.de> writes:
> On Tue, Aug 16, 2016 at 04:09:19PM -0700, H. Peter Anvin wrote:
>> On August 16, 2016 10:16:35 AM PDT, Borislav Petkov <bp@suse.de> wrote:
>> >On Tue, Aug 16, 2016 at 09:59:00AM -0700, H. Peter Anvin wrote:
>> >> Dang...
>> >
>> >Isn't 9.3% improvement a good thing(tm) ?
>>
>> Yes, it's huge. The only explanation I could imagine is that scrambling %rdi caused the scheduler to do completely the wrong thing.
>
> I'm questioning the validity, actually. Report says test machine was
> Sandy Bridge-EP and I'd bet good money this one has POPCNT support so
> how are we even hitting that __sw_hweight64() path, at all?
We done 8 tests for the base and 4 tests for the head, and the result is
quite stable.
I found there is another change between the two comments,
base:
"perf-stat.branch-miss-rate": [
0.3089533646503185,
0.3099821038600304,
0.3123762964028104,
0.311511881793534,
0.31231973343587144,
0.3096478429327263,
0.31166037272389924,
0.3097364392684626
],
first bad commit:
"perf-stat.branch-miss-rate": [
0.039853905034485354,
0.0402472142423231,
0.04380682345704418,
0.04319082390667179
],
branch-miss-rate decreased from ~0.30% to ~0.043%.
So I guess there are some code alignment change, which caused decreased
branch miss rate.
Best Regards,
Huang, Ying
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@suse.de> |
|---|---|
| Date | 2016-08-18 05:50 +0200 |
| Subject | Re: [LKP] [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement |
| Message-ID | <s7mv8-2qJ-3@gated-at.bofh.it> |
| In reply to | #1464843 |
On Wed, Aug 17, 2016 at 03:29:04PM -0700, Huang, Ying wrote:
> branch-miss-rate decreased from ~0.30% to ~0.043%.
>
> So I guess there are some code alignment change, which caused decreased
> branch miss rate.
Hrrm, I still can't imagine how that would happen if the machine
supports POPCNT and we never call the __sw_hweight* variants. Or does
it?
Can you paste /proc/cpuinfo from that Sandy Bridge-EP box?
Thanks.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
SUSE Linux GmbH, GF: Felix Imendörffer, Jane Smithard, Graham Norton, HRB 21284 (AG Nürnberg)
--
[toc] | [prev] | [next] | [standalone]
| From | "Huang\, Ying" <ying.huang@intel.com> |
|---|---|
| Date | 2016-08-18 06:00 +0200 |
| Subject | Re: [LKP] [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement |
| Message-ID | <s7mEO-2uu-11@gated-at.bofh.it> |
| In reply to | #1464913 |
Borislav Petkov <bp@suse.de> writes: > On Wed, Aug 17, 2016 at 03:29:04PM -0700, Huang, Ying wrote: >> branch-miss-rate decreased from ~0.30% to ~0.043%. >> >> So I guess there are some code alignment change, which caused decreased >> branch miss rate. > > Hrrm, I still can't imagine how that would happen if the machine > supports POPCNT and we never call the __sw_hweight* variants. Or does > it? > > Can you paste /proc/cpuinfo from that Sandy Bridge-EP box? Here it is. processor : 31 vendor_id : GenuineIntel cpu family : 6 model : 45 model name : Intel(R) Xeon(R) CPU E5-2680 0 @ 2.70GHz stepping : 6 microcode : 0x603 cpu MHz : 3158.129 cache size : 20480 KB physical id : 1 siblings : 16 core id : 7 cpu cores : 8 apicid : 47 initial apicid : 47 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic popcnt tsc_deadline_timer aes xsave avx lahf_lm tpr_shadow vnmi flexpriority ept vpid xsaveopt dtherm ida arat pln pts bugs : bogomips : 5392.85 clflush size : 64 cache_alignment : 64 address sizes : 46 bits physical, 48 bits virtual power management: Best Regards, Huang, Ying
[toc] | [prev] | [next] | [standalone]
| From | Borislav Petkov <bp@suse.de> |
|---|---|
| Date | 2016-08-18 06:20 +0200 |
| Subject | Re: [LKP] [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement |
| Message-ID | <s7mY9-2RP-1@gated-at.bofh.it> |
| In reply to | #1464919 |
On Wed, Aug 17, 2016 at 08:54:11PM -0700, Huang, Ying wrote:
> flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 cx16 xtpr pdcm pcid dca sse4_1 sse4_2 x2apic
> popcnt tsc_deadline_timer aes xsave avx lahf_lm tpr_shadow vnmi flexpriority ept vpid xsaveopt dtherm ida arat pln pts
^^^^^^
There it is.
So if there's no bug, alternatives should replace all "call
__sw_hweightXX" calls with POPCNT. So you shouldn't be even calling
these functions and hitting that path.
Can you boot the kernel with "debug-alternative" and put that dmesg
somewhere along with vmlinux for me to stare at? Privately is fine too.
I'd like to make sure the alternatives application actually happens.
Thanks.
--
Regards/Gruss,
Boris.
ECO tip #101: Trim your mails when you reply.
SUSE Linux GmbH, GF: Felix Imendörffer, Jane Smithard, Graham Norton, HRB 21284 (AG Nürnberg)
--
[toc] | [prev] | [next] | [standalone]
| From | "H. Peter Anvin" <hpa@zytor.com> |
|---|---|
| Date | 2016-08-18 06:00 +0200 |
| Subject | Re: [LKP] [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement |
| Message-ID | <s7mEO-2uu-15@gated-at.bofh.it> |
| In reply to | #1464913 |
On August 17, 2016 8:45:13 PM PDT, Borislav Petkov <bp@suse.de> wrote: >On Wed, Aug 17, 2016 at 03:29:04PM -0700, Huang, Ying wrote: >> branch-miss-rate decreased from ~0.30% to ~0.043%. >> >> So I guess there are some code alignment change, which caused >decreased >> branch miss rate. > >Hrrm, I still can't imagine how that would happen if the machine >supports POPCNT and we never call the __sw_hweight* variants. Or does >it? > >Can you paste /proc/cpuinfo from that Sandy Bridge-EP box? > >Thanks. popcnt was introduced in Nehalem AFAIK, two generations before Sandy Bridge. -- Sent from my Android device with K-9 Mail. Please excuse brevity and formatting.
[toc] | [prev] | [next] | [standalone]
| From | Peter Zijlstra <peterz@infradead.org> |
|---|---|
| Date | 2016-08-17 08:50 +0200 |
| Subject | Re: [lkp] [x86/hweight] 65ea11ec6a: will-it-scale.per_process_ops 9.3% improvement |
| Message-ID | <s72PM-5MN-17@gated-at.bofh.it> |
| In reply to | #1464230 |
On Tue, Aug 16, 2016 at 04:09:19PM -0700, H. Peter Anvin wrote: > On August 16, 2016 10:16:35 AM PDT, Borislav Petkov <bp@suse.de> wrote: > >On Tue, Aug 16, 2016 at 09:59:00AM -0700, H. Peter Anvin wrote: > >> Dang... > > > >Isn't 9.3% improvement a good thing(tm) ? > > Yes, it's huge. The only explanation I could imagine is that scrambling %rdi caused the scheduler to do completely the wrong thing. Not entirely surprising. We have plenty bitmasks and if hweight is corrupting the source data instead of computing the weight then we end up having two bits of wrong information. After that, all we can do is more wrong of course...
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web