Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1672552 > unrolled thread
| Started by | root <yang.zhang.wz@gmail.com> |
|---|---|
| First post | 2017-06-22 13:30 +0200 |
| Last post | 2017-06-27 16:10 +0200 |
| Articles | 6 on this page of 26 — 7 participants |
Back to article view | Back to linux.kernel
[PATCH 0/2] x86/idle: add halt poll support root <yang.zhang.wz@gmail.com> - 2017-06-22 13:30 +0200
[PATCH 1/2] x86/idle: add halt poll for halt idle root <yang.zhang.wz@gmail.com> - 2017-06-22 13:30 +0200
Re: [PATCH 1/2] x86/idle: add halt poll for halt idle Thomas Gleixner <tglx@linutronix.de> - 2017-06-22 16:30 +0200
Re: [PATCH 1/2] x86/idle: add halt poll for halt idle Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-23 06:10 +0200
[PATCH 2/2] x86/idle: use dynamic halt poll root <yang.zhang.wz@gmail.com> - 2017-06-22 13:30 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Paolo Bonzini <pbonzini@redhat.com> - 2017-06-22 14:00 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-23 06:00 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-27 13:30 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Paolo Bonzini <pbonzini@redhat.com> - 2017-06-27 14:10 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Wanpeng Li <kernellwp@gmail.com> - 2017-06-27 14:30 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Paolo Bonzini <pbonzini@redhat.com> - 2017-06-27 14:30 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Radim Krčmář <rkrcmar@redhat.com> - 2017-06-27 15:50 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Paolo Bonzini <pbonzini@redhat.com> - 2017-06-27 16:00 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Paolo Bonzini <pbonzini@redhat.com> - 2017-06-27 16:30 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Radim Krčmář <rkrcmar@redhat.com> - 2017-06-27 16:30 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Yang Zhang <yang.zhang.wz@gmail.com> - 2017-07-03 11:30 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Thomas Gleixner <tglx@linutronix.de> - 2017-07-03 12:10 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Thomas Gleixner <tglx@linutronix.de> - 2017-06-22 16:40 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-23 06:10 +0200
Re: [PATCH 2/2] x86/idle: use dynamic halt poll kbuild test robot <lkp@intel.com> - 2017-06-23 00:50 +0200
Re: [PATCH 0/2] x86/idle: add halt poll support Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-22 13:40 +0200
Re: [PATCH 0/2] x86/idle: add halt poll support Wanpeng Li <kernellwp@gmail.com> - 2017-06-22 14:00 +0200
Re: [PATCH 0/2] x86/idle: add halt poll support Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-23 06:10 +0200
Re: [PATCH 0/2] x86/idle: add halt poll support Wanpeng Li <kernellwp@gmail.com> - 2017-06-23 06:40 +0200
Re: [PATCH 0/2] x86/idle: add halt poll support Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-23 08:50 +0200
Re: [PATCH 0/2] x86/idle: add halt poll support Radim Krčmář <rkrcmar@redhat.com> - 2017-06-27 16:10 +0200
Page 2 of 2 — ← Prev page 1 [2]
| From | Yang Zhang <yang.zhang.wz@gmail.com> |
|---|---|
| Date | 2017-06-22 13:40 +0200 |
| Message-ID | <tV8CS-42M-21@gated-at.bofh.it> |
| In reply to | #1672552 |
On 2017/6/22 19:22, root wrote: > From: Yang Zhang <yang.zhang.wz@gmail.com> Sorry to use wrong username to send patch because i am using a new machine which don't setup the git config well. > > Some latency-intensive workload will see obviously performance > drop when running inside VM. The main reason is that the overhead > is amplified when running inside VM. The most cost i have seen is > inside idle path. > This patch introduces a new mechanism to poll for a while before > entering idle state. If schedule is needed during poll, then we > don't need to goes through the heavy overhead path. > > Here is the data i get when running benchmark contextswitch > (https://github.com/tsuna/contextswitch) > before patch: > 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw) > after patch: > 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw) > > > Yang Zhang (2): > x86/idle: add halt poll for halt idle > x86/idle: use dynamic halt poll > > Documentation/sysctl/kernel.txt | 24 ++++++++++ > arch/x86/include/asm/processor.h | 6 +++ > arch/x86/kernel/apic/apic.c | 6 +++ > arch/x86/kernel/apic/vector.c | 1 + > arch/x86/kernel/cpu/mcheck/mce_amd.c | 2 + > arch/x86/kernel/cpu/mcheck/therm_throt.c | 2 + > arch/x86/kernel/cpu/mcheck/threshold.c | 2 + > arch/x86/kernel/irq.c | 5 ++ > arch/x86/kernel/irq_work.c | 2 + > arch/x86/kernel/process.c | 80 ++++++++++++++++++++++++++++++++ > arch/x86/kernel/smp.c | 6 +++ > include/linux/kernel.h | 5 ++ > kernel/sched/idle.c | 3 ++ > kernel/sysctl.c | 23 +++++++++ > 14 files changed, 167 insertions(+) > -- Yang Alibaba Cloud Computing
[toc] | [prev] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2017-06-22 14:00 +0200 |
| Message-ID | <tV8Wf-4b0-29@gated-at.bofh.it> |
| In reply to | #1672552 |
2017-06-22 19:22 GMT+08:00 root <yang.zhang.wz@gmail.com>: > From: Yang Zhang <yang.zhang.wz@gmail.com> > > Some latency-intensive workload will see obviously performance > drop when running inside VM. The main reason is that the overhead > is amplified when running inside VM. The most cost i have seen is > inside idle path. > This patch introduces a new mechanism to poll for a while before > entering idle state. If schedule is needed during poll, then we > don't need to goes through the heavy overhead path. > > Here is the data i get when running benchmark contextswitch > (https://github.com/tsuna/contextswitch) > before patch: > 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw) > after patch: > 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw) If you test this after disabling the adaptive halt-polling in kvm? What's the performance data of w/ this patchset and w/o the adaptive halt-polling in kvm, and w/o this patchset and w/ the adaptive halt-polling in kvm? In addition, both linux and windows guests can get benefit as we have already done this in kvm. Regards, Wanpeng Li > Yang Zhang (2): > x86/idle: add halt poll for halt idle > x86/idle: use dynamic halt poll > > Documentation/sysctl/kernel.txt | 24 ++++++++++ > arch/x86/include/asm/processor.h | 6 +++ > arch/x86/kernel/apic/apic.c | 6 +++ > arch/x86/kernel/apic/vector.c | 1 + > arch/x86/kernel/cpu/mcheck/mce_amd.c | 2 + > arch/x86/kernel/cpu/mcheck/therm_throt.c | 2 + > arch/x86/kernel/cpu/mcheck/threshold.c | 2 + > arch/x86/kernel/irq.c | 5 ++ > arch/x86/kernel/irq_work.c | 2 + > arch/x86/kernel/process.c | 80 ++++++++++++++++++++++++++++++++ > arch/x86/kernel/smp.c | 6 +++ > include/linux/kernel.h | 5 ++ > kernel/sched/idle.c | 3 ++ > kernel/sysctl.c | 23 +++++++++ > 14 files changed, 167 insertions(+) > > -- > 1.8.3.1 >
[toc] | [prev] | [next] | [standalone]
| From | Yang Zhang <yang.zhang.wz@gmail.com> |
|---|---|
| Date | 2017-06-23 06:10 +0200 |
| Message-ID | <tVo4W-5HP-11@gated-at.bofh.it> |
| In reply to | #1672565 |
On 2017/6/22 19:50, Wanpeng Li wrote: > 2017-06-22 19:22 GMT+08:00 root <yang.zhang.wz@gmail.com>: >> From: Yang Zhang <yang.zhang.wz@gmail.com> >> >> Some latency-intensive workload will see obviously performance >> drop when running inside VM. The main reason is that the overhead >> is amplified when running inside VM. The most cost i have seen is >> inside idle path. >> This patch introduces a new mechanism to poll for a while before >> entering idle state. If schedule is needed during poll, then we >> don't need to goes through the heavy overhead path. >> >> Here is the data i get when running benchmark contextswitch >> (https://github.com/tsuna/contextswitch) >> before patch: >> 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw) >> after patch: >> 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw) > > If you test this after disabling the adaptive halt-polling in kvm? > What's the performance data of w/ this patchset and w/o the adaptive > halt-polling in kvm, and w/o this patchset and w/ the adaptive > halt-polling in kvm? In addition, both linux and windows guests can > get benefit as we have already done this in kvm. I will provide more data in next version. But it doesn't conflict with current halt polling inside kvm. This is just another enhancement. > > Regards, > Wanpeng Li > >> Yang Zhang (2): >> x86/idle: add halt poll for halt idle >> x86/idle: use dynamic halt poll >> >> Documentation/sysctl/kernel.txt | 24 ++++++++++ >> arch/x86/include/asm/processor.h | 6 +++ >> arch/x86/kernel/apic/apic.c | 6 +++ >> arch/x86/kernel/apic/vector.c | 1 + >> arch/x86/kernel/cpu/mcheck/mce_amd.c | 2 + >> arch/x86/kernel/cpu/mcheck/therm_throt.c | 2 + >> arch/x86/kernel/cpu/mcheck/threshold.c | 2 + >> arch/x86/kernel/irq.c | 5 ++ >> arch/x86/kernel/irq_work.c | 2 + >> arch/x86/kernel/process.c | 80 ++++++++++++++++++++++++++++++++ >> arch/x86/kernel/smp.c | 6 +++ >> include/linux/kernel.h | 5 ++ >> kernel/sched/idle.c | 3 ++ >> kernel/sysctl.c | 23 +++++++++ >> 14 files changed, 167 insertions(+) >> >> -- >> 1.8.3.1 >> -- Yang Alibaba Cloud Computing
[toc] | [prev] | [next] | [standalone]
| From | Wanpeng Li <kernellwp@gmail.com> |
|---|---|
| Date | 2017-06-23 06:40 +0200 |
| Message-ID | <tVoxX-5VY-13@gated-at.bofh.it> |
| In reply to | #1673216 |
2017-06-23 12:08 GMT+08:00 Yang Zhang <yang.zhang.wz@gmail.com>: > On 2017/6/22 19:50, Wanpeng Li wrote: >> >> 2017-06-22 19:22 GMT+08:00 root <yang.zhang.wz@gmail.com>: >>> >>> From: Yang Zhang <yang.zhang.wz@gmail.com> >>> >>> Some latency-intensive workload will see obviously performance >>> drop when running inside VM. The main reason is that the overhead >>> is amplified when running inside VM. The most cost i have seen is >>> inside idle path. >>> This patch introduces a new mechanism to poll for a while before >>> entering idle state. If schedule is needed during poll, then we >>> don't need to goes through the heavy overhead path. >>> >>> Here is the data i get when running benchmark contextswitch >>> (https://github.com/tsuna/contextswitch) >>> before patch: >>> 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw) >>> after patch: >>> 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw) >> >> >> If you test this after disabling the adaptive halt-polling in kvm? >> What's the performance data of w/ this patchset and w/o the adaptive >> halt-polling in kvm, and w/o this patchset and w/ the adaptive >> halt-polling in kvm? In addition, both linux and windows guests can >> get benefit as we have already done this in kvm. > > > I will provide more data in next version. But it doesn't conflict with Another case I can think of is w/ both this patchset and the adaptive halt-polling in kvm. > current halt polling inside kvm. This is just another enhancement. I didn't look close to the patchset, however, maybe there is another poll in the kvm part again sometimes if you fails the poll in the guest. In addition, the adaptive halt-polling in kvm has performance penalty when the pCPU is heavily overcommitted though there is a single_task_running() in my testing, it is hard to accurately aware whether there are other tasks waiting on the pCPU in the guest which will make it worser. Depending on vcpu_is_preempted() or steal time maybe not accurately or directly. So I'm not sure how much sense it makes by adaptive halt-polling in both guest and kvm. I prefer to just keep adaptive halt-polling in kvm(then both linux/windows or other guests can get benefit) and avoid to churn the core x86 path. Regards, Wanpeng Li
[toc] | [prev] | [next] | [standalone]
| From | Yang Zhang <yang.zhang.wz@gmail.com> |
|---|---|
| Date | 2017-06-23 08:50 +0200 |
| Message-ID | <tVqzM-78A-13@gated-at.bofh.it> |
| In reply to | #1673237 |
On 2017/6/23 12:35, Wanpeng Li wrote: > 2017-06-23 12:08 GMT+08:00 Yang Zhang <yang.zhang.wz@gmail.com>: >> On 2017/6/22 19:50, Wanpeng Li wrote: >>> >>> 2017-06-22 19:22 GMT+08:00 root <yang.zhang.wz@gmail.com>: >>>> >>>> From: Yang Zhang <yang.zhang.wz@gmail.com> >>>> >>>> Some latency-intensive workload will see obviously performance >>>> drop when running inside VM. The main reason is that the overhead >>>> is amplified when running inside VM. The most cost i have seen is >>>> inside idle path. >>>> This patch introduces a new mechanism to poll for a while before >>>> entering idle state. If schedule is needed during poll, then we >>>> don't need to goes through the heavy overhead path. >>>> >>>> Here is the data i get when running benchmark contextswitch >>>> (https://github.com/tsuna/contextswitch) >>>> before patch: >>>> 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw) >>>> after patch: >>>> 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw) >>> >>> >>> If you test this after disabling the adaptive halt-polling in kvm? >>> What's the performance data of w/ this patchset and w/o the adaptive >>> halt-polling in kvm, and w/o this patchset and w/ the adaptive >>> halt-polling in kvm? In addition, both linux and windows guests can >>> get benefit as we have already done this in kvm. >> >> >> I will provide more data in next version. But it doesn't conflict with > > Another case I can think of is w/ both this patchset and the adaptive > halt-polling in kvm. > >> current halt polling inside kvm. This is just another enhancement. > > I didn't look close to the patchset, however, maybe there is another > poll in the kvm part again sometimes if you fails the poll in the > guest. In addition, the adaptive halt-polling in kvm has performance > penalty when the pCPU is heavily overcommitted though there is a > single_task_running() in my testing, it is hard to accurately aware > whether there are other tasks waiting on the pCPU in the guest which > will make it worser. Depending on vcpu_is_preempted() or steal time > maybe not accurately or directly. > > So I'm not sure how much sense it makes by adaptive halt-polling in > both guest and kvm. I prefer to just keep adaptive halt-polling in > kvm(then both linux/windows or other guests can get benefit) and avoid > to churn the core x86 path. This mechanism is not specific to KVM. It is a kernel feature which can benefit guest when running inside X86 virtualization environment. The guest includes KVM,Xen,VMWARE,Hyper-v. Administrator can control KVM to use adaptive halt poll but he cannot control the user to use halt polling inside guest. Lots of user set idle=poll inside guest to improve performance which occupy more CPU cycles. This mechanism is a enhancement to it not to KVM halt polling. > > Regards, > Wanpeng Li > -- Yang Alibaba Cloud Computing
[toc] | [prev] | [next] | [standalone]
| From | Radim Krčmář <rkrcmar@redhat.com> |
|---|---|
| Date | 2017-06-27 16:10 +0200 |
| Message-ID | <tWZlL-1Bo-7@gated-at.bofh.it> |
| In reply to | #1673297 |
2017-06-23 14:49+0800, Yang Zhang: > On 2017/6/23 12:35, Wanpeng Li wrote: > > 2017-06-23 12:08 GMT+08:00 Yang Zhang <yang.zhang.wz@gmail.com>: > > > On 2017/6/22 19:50, Wanpeng Li wrote: > > > > > > > > 2017-06-22 19:22 GMT+08:00 root <yang.zhang.wz@gmail.com>: > > > > > > > > > > From: Yang Zhang <yang.zhang.wz@gmail.com> > > > > > > > > > > Some latency-intensive workload will see obviously performance > > > > > drop when running inside VM. The main reason is that the overhead > > > > > is amplified when running inside VM. The most cost i have seen is > > > > > inside idle path. > > > > > This patch introduces a new mechanism to poll for a while before > > > > > entering idle state. If schedule is needed during poll, then we > > > > > don't need to goes through the heavy overhead path. > > > > > > > > > > Here is the data i get when running benchmark contextswitch > > > > > (https://github.com/tsuna/contextswitch) > > > > > before patch: > > > > > 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw) > > > > > after patch: > > > > > 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw) > > > > > > > > > > > > If you test this after disabling the adaptive halt-polling in kvm? > > > > What's the performance data of w/ this patchset and w/o the adaptive > > > > halt-polling in kvm, and w/o this patchset and w/ the adaptive > > > > halt-polling in kvm? In addition, both linux and windows guests can > > > > get benefit as we have already done this in kvm. > > > > > > > > > I will provide more data in next version. But it doesn't conflict with > > > > Another case I can think of is w/ both this patchset and the adaptive > > halt-polling in kvm. > > > > > current halt polling inside kvm. This is just another enhancement. > > > > I didn't look close to the patchset, however, maybe there is another > > poll in the kvm part again sometimes if you fails the poll in the > > guest. In addition, the adaptive halt-polling in kvm has performance > > penalty when the pCPU is heavily overcommitted though there is a > > single_task_running() in my testing, it is hard to accurately aware > > whether there are other tasks waiting on the pCPU in the guest which > > will make it worser. Depending on vcpu_is_preempted() or steal time > > maybe not accurately or directly. > > > > So I'm not sure how much sense it makes by adaptive halt-polling in > > both guest and kvm. I prefer to just keep adaptive halt-polling in > > kvm(then both linux/windows or other guests can get benefit) and avoid > > to churn the core x86 path. > > This mechanism is not specific to KVM. It is a kernel feature which can > benefit guest when running inside X86 virtualization environment. The guest > includes KVM,Xen,VMWARE,Hyper-v. Administrator can control KVM to use > adaptive halt poll but he cannot control the user to use halt polling inside > guest. Lots of user set idle=poll inside guest to improve performance which > occupy more CPU cycles. This mechanism is a enhancement to it not to KVM > halt polling. Users of idle=poll shouln't overcommit, so the goal seems to be energy savings without crippling the guest performance too much ... Wouldn't switching to idle=mwait work as well? Thanks.
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web