Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1672552 > unrolled thread

[PATCH 0/2] x86/idle: add halt poll support

Started byroot <yang.zhang.wz@gmail.com>
First post2017-06-22 13:30 +0200
Last post2017-06-27 16:10 +0200
Articles 6 on this page of 26 — 7 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH 0/2] x86/idle: add halt poll support root <yang.zhang.wz@gmail.com> - 2017-06-22 13:30 +0200
    [PATCH 1/2] x86/idle: add halt poll for halt idle root <yang.zhang.wz@gmail.com> - 2017-06-22 13:30 +0200
      Re: [PATCH 1/2] x86/idle: add halt poll for halt idle Thomas Gleixner <tglx@linutronix.de> - 2017-06-22 16:30 +0200
        Re: [PATCH 1/2] x86/idle: add halt poll for halt idle Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-23 06:10 +0200
    [PATCH 2/2] x86/idle: use dynamic halt poll root <yang.zhang.wz@gmail.com> - 2017-06-22 13:30 +0200
      Re: [PATCH 2/2] x86/idle: use dynamic halt poll Paolo Bonzini <pbonzini@redhat.com> - 2017-06-22 14:00 +0200
        Re: [PATCH 2/2] x86/idle: use dynamic halt poll Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-23 06:00 +0200
          Re: [PATCH 2/2] x86/idle: use dynamic halt poll Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-27 13:30 +0200
            Re: [PATCH 2/2] x86/idle: use dynamic halt poll Paolo Bonzini <pbonzini@redhat.com> - 2017-06-27 14:10 +0200
              Re: [PATCH 2/2] x86/idle: use dynamic halt poll Wanpeng Li <kernellwp@gmail.com> - 2017-06-27 14:30 +0200
                Re: [PATCH 2/2] x86/idle: use dynamic halt poll Paolo Bonzini <pbonzini@redhat.com> - 2017-06-27 14:30 +0200
                  Re: [PATCH 2/2] x86/idle: use dynamic halt poll Radim Krčmář <rkrcmar@redhat.com> - 2017-06-27 15:50 +0200
                    Re: [PATCH 2/2] x86/idle: use dynamic halt poll Paolo Bonzini <pbonzini@redhat.com> - 2017-06-27 16:00 +0200
                      Re: [PATCH 2/2] x86/idle: use dynamic halt poll Paolo Bonzini <pbonzini@redhat.com> - 2017-06-27 16:30 +0200
                      Re: [PATCH 2/2] x86/idle: use dynamic halt poll Radim Krčmář <rkrcmar@redhat.com> - 2017-06-27 16:30 +0200
                        Re: [PATCH 2/2] x86/idle: use dynamic halt poll Yang Zhang <yang.zhang.wz@gmail.com> - 2017-07-03 11:30 +0200
                          Re: [PATCH 2/2] x86/idle: use dynamic halt poll Thomas Gleixner <tglx@linutronix.de> - 2017-07-03 12:10 +0200
      Re: [PATCH 2/2] x86/idle: use dynamic halt poll Thomas Gleixner <tglx@linutronix.de> - 2017-06-22 16:40 +0200
        Re: [PATCH 2/2] x86/idle: use dynamic halt poll Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-23 06:10 +0200
      Re: [PATCH 2/2] x86/idle: use dynamic halt poll kbuild test robot <lkp@intel.com> - 2017-06-23 00:50 +0200
    Re: [PATCH 0/2] x86/idle: add halt poll support Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-22 13:40 +0200
    Re: [PATCH 0/2] x86/idle: add halt poll support Wanpeng Li <kernellwp@gmail.com> - 2017-06-22 14:00 +0200
      Re: [PATCH 0/2] x86/idle: add halt poll support Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-23 06:10 +0200
        Re: [PATCH 0/2] x86/idle: add halt poll support Wanpeng Li <kernellwp@gmail.com> - 2017-06-23 06:40 +0200
          Re: [PATCH 0/2] x86/idle: add halt poll support Yang Zhang <yang.zhang.wz@gmail.com> - 2017-06-23 08:50 +0200
            Re: [PATCH 0/2] x86/idle: add halt poll support Radim Krčmář <rkrcmar@redhat.com> - 2017-06-27 16:10 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1672556

FromYang Zhang <yang.zhang.wz@gmail.com>
Date2017-06-22 13:40 +0200
Message-ID<tV8CS-42M-21@gated-at.bofh.it>
In reply to#1672552
On 2017/6/22 19:22, root wrote:
> From: Yang Zhang <yang.zhang.wz@gmail.com>

Sorry to use wrong username to send patch because i am using a new
machine which don't setup the git config well.

>
> Some latency-intensive workload will see obviously performance
> drop when running inside VM. The main reason is that the overhead
> is amplified when running inside VM. The most cost i have seen is
> inside idle path.
> This patch introduces a new mechanism to poll for a while before
> entering idle state. If schedule is needed during poll, then we
> don't need to goes through the heavy overhead path.
>
> Here is the data i get when running benchmark contextswitch
> (https://github.com/tsuna/contextswitch)
> before patch:
> 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw)
> after patch:
> 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw)
>
>
> Yang Zhang (2):
>   x86/idle: add halt poll for halt idle
>   x86/idle: use dynamic halt poll
>
>  Documentation/sysctl/kernel.txt          | 24 ++++++++++
>  arch/x86/include/asm/processor.h         |  6 +++
>  arch/x86/kernel/apic/apic.c              |  6 +++
>  arch/x86/kernel/apic/vector.c            |  1 +
>  arch/x86/kernel/cpu/mcheck/mce_amd.c     |  2 +
>  arch/x86/kernel/cpu/mcheck/therm_throt.c |  2 +
>  arch/x86/kernel/cpu/mcheck/threshold.c   |  2 +
>  arch/x86/kernel/irq.c                    |  5 ++
>  arch/x86/kernel/irq_work.c               |  2 +
>  arch/x86/kernel/process.c                | 80 ++++++++++++++++++++++++++++++++
>  arch/x86/kernel/smp.c                    |  6 +++
>  include/linux/kernel.h                   |  5 ++
>  kernel/sched/idle.c                      |  3 ++
>  kernel/sysctl.c                          | 23 +++++++++
>  14 files changed, 167 insertions(+)
>


-- 
Yang
Alibaba Cloud Computing

[toc] | [prev] | [next] | [standalone]


#1672565

FromWanpeng Li <kernellwp@gmail.com>
Date2017-06-22 14:00 +0200
Message-ID<tV8Wf-4b0-29@gated-at.bofh.it>
In reply to#1672552
2017-06-22 19:22 GMT+08:00 root <yang.zhang.wz@gmail.com>:
> From: Yang Zhang <yang.zhang.wz@gmail.com>
>
> Some latency-intensive workload will see obviously performance
> drop when running inside VM. The main reason is that the overhead
> is amplified when running inside VM. The most cost i have seen is
> inside idle path.
> This patch introduces a new mechanism to poll for a while before
> entering idle state. If schedule is needed during poll, then we
> don't need to goes through the heavy overhead path.
>
> Here is the data i get when running benchmark contextswitch
> (https://github.com/tsuna/contextswitch)
> before patch:
> 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw)
> after patch:
> 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw)

If you test this after disabling the adaptive halt-polling in kvm?
What's the performance data of w/ this patchset and w/o the adaptive
halt-polling in kvm, and w/o this patchset and w/ the adaptive
halt-polling in kvm? In addition, both linux and windows guests can
get benefit as we have already done this in kvm.

Regards,
Wanpeng Li

> Yang Zhang (2):
>   x86/idle: add halt poll for halt idle
>   x86/idle: use dynamic halt poll
>
>  Documentation/sysctl/kernel.txt          | 24 ++++++++++
>  arch/x86/include/asm/processor.h         |  6 +++
>  arch/x86/kernel/apic/apic.c              |  6 +++
>  arch/x86/kernel/apic/vector.c            |  1 +
>  arch/x86/kernel/cpu/mcheck/mce_amd.c     |  2 +
>  arch/x86/kernel/cpu/mcheck/therm_throt.c |  2 +
>  arch/x86/kernel/cpu/mcheck/threshold.c   |  2 +
>  arch/x86/kernel/irq.c                    |  5 ++
>  arch/x86/kernel/irq_work.c               |  2 +
>  arch/x86/kernel/process.c                | 80 ++++++++++++++++++++++++++++++++
>  arch/x86/kernel/smp.c                    |  6 +++
>  include/linux/kernel.h                   |  5 ++
>  kernel/sched/idle.c                      |  3 ++
>  kernel/sysctl.c                          | 23 +++++++++
>  14 files changed, 167 insertions(+)
>
> --
> 1.8.3.1
>

[toc] | [prev] | [next] | [standalone]


#1673216

FromYang Zhang <yang.zhang.wz@gmail.com>
Date2017-06-23 06:10 +0200
Message-ID<tVo4W-5HP-11@gated-at.bofh.it>
In reply to#1672565
On 2017/6/22 19:50, Wanpeng Li wrote:
> 2017-06-22 19:22 GMT+08:00 root <yang.zhang.wz@gmail.com>:
>> From: Yang Zhang <yang.zhang.wz@gmail.com>
>>
>> Some latency-intensive workload will see obviously performance
>> drop when running inside VM. The main reason is that the overhead
>> is amplified when running inside VM. The most cost i have seen is
>> inside idle path.
>> This patch introduces a new mechanism to poll for a while before
>> entering idle state. If schedule is needed during poll, then we
>> don't need to goes through the heavy overhead path.
>>
>> Here is the data i get when running benchmark contextswitch
>> (https://github.com/tsuna/contextswitch)
>> before patch:
>> 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw)
>> after patch:
>> 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw)
>
> If you test this after disabling the adaptive halt-polling in kvm?
> What's the performance data of w/ this patchset and w/o the adaptive
> halt-polling in kvm, and w/o this patchset and w/ the adaptive
> halt-polling in kvm? In addition, both linux and windows guests can
> get benefit as we have already done this in kvm.

I will provide more data in next version. But it doesn't conflict with 
current halt polling inside kvm. This is just another enhancement.

>
> Regards,
> Wanpeng Li
>
>> Yang Zhang (2):
>>   x86/idle: add halt poll for halt idle
>>   x86/idle: use dynamic halt poll
>>
>>  Documentation/sysctl/kernel.txt          | 24 ++++++++++
>>  arch/x86/include/asm/processor.h         |  6 +++
>>  arch/x86/kernel/apic/apic.c              |  6 +++
>>  arch/x86/kernel/apic/vector.c            |  1 +
>>  arch/x86/kernel/cpu/mcheck/mce_amd.c     |  2 +
>>  arch/x86/kernel/cpu/mcheck/therm_throt.c |  2 +
>>  arch/x86/kernel/cpu/mcheck/threshold.c   |  2 +
>>  arch/x86/kernel/irq.c                    |  5 ++
>>  arch/x86/kernel/irq_work.c               |  2 +
>>  arch/x86/kernel/process.c                | 80 ++++++++++++++++++++++++++++++++
>>  arch/x86/kernel/smp.c                    |  6 +++
>>  include/linux/kernel.h                   |  5 ++
>>  kernel/sched/idle.c                      |  3 ++
>>  kernel/sysctl.c                          | 23 +++++++++
>>  14 files changed, 167 insertions(+)
>>
>> --
>> 1.8.3.1
>>


-- 
Yang
Alibaba Cloud Computing

[toc] | [prev] | [next] | [standalone]


#1673237

FromWanpeng Li <kernellwp@gmail.com>
Date2017-06-23 06:40 +0200
Message-ID<tVoxX-5VY-13@gated-at.bofh.it>
In reply to#1673216
2017-06-23 12:08 GMT+08:00 Yang Zhang <yang.zhang.wz@gmail.com>:
> On 2017/6/22 19:50, Wanpeng Li wrote:
>>
>> 2017-06-22 19:22 GMT+08:00 root <yang.zhang.wz@gmail.com>:
>>>
>>> From: Yang Zhang <yang.zhang.wz@gmail.com>
>>>
>>> Some latency-intensive workload will see obviously performance
>>> drop when running inside VM. The main reason is that the overhead
>>> is amplified when running inside VM. The most cost i have seen is
>>> inside idle path.
>>> This patch introduces a new mechanism to poll for a while before
>>> entering idle state. If schedule is needed during poll, then we
>>> don't need to goes through the heavy overhead path.
>>>
>>> Here is the data i get when running benchmark contextswitch
>>> (https://github.com/tsuna/contextswitch)
>>> before patch:
>>> 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw)
>>> after patch:
>>> 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw)
>>
>>
>> If you test this after disabling the adaptive halt-polling in kvm?
>> What's the performance data of w/ this patchset and w/o the adaptive
>> halt-polling in kvm, and w/o this patchset and w/ the adaptive
>> halt-polling in kvm? In addition, both linux and windows guests can
>> get benefit as we have already done this in kvm.
>
>
> I will provide more data in next version. But it doesn't conflict with

Another case I can think of is w/ both this patchset and the adaptive
halt-polling in kvm.

> current halt polling inside kvm. This is just another enhancement.

I didn't look close to the patchset, however, maybe there is another
poll in the kvm part again sometimes if you fails the poll in the
guest. In addition, the adaptive halt-polling in kvm has performance
penalty when the pCPU is heavily overcommitted though there is a
single_task_running() in my testing, it is hard to accurately aware
whether there are other tasks waiting on the pCPU in the guest which
will make it worser. Depending on vcpu_is_preempted() or steal time
maybe not accurately or directly.

So I'm not sure how much sense it makes by adaptive halt-polling in
both guest and kvm. I prefer to just keep adaptive halt-polling in
kvm(then both linux/windows or other guests can get benefit) and avoid
to churn the core x86 path.

Regards,
Wanpeng Li

[toc] | [prev] | [next] | [standalone]


#1673297

FromYang Zhang <yang.zhang.wz@gmail.com>
Date2017-06-23 08:50 +0200
Message-ID<tVqzM-78A-13@gated-at.bofh.it>
In reply to#1673237
On 2017/6/23 12:35, Wanpeng Li wrote:
> 2017-06-23 12:08 GMT+08:00 Yang Zhang <yang.zhang.wz@gmail.com>:
>> On 2017/6/22 19:50, Wanpeng Li wrote:
>>>
>>> 2017-06-22 19:22 GMT+08:00 root <yang.zhang.wz@gmail.com>:
>>>>
>>>> From: Yang Zhang <yang.zhang.wz@gmail.com>
>>>>
>>>> Some latency-intensive workload will see obviously performance
>>>> drop when running inside VM. The main reason is that the overhead
>>>> is amplified when running inside VM. The most cost i have seen is
>>>> inside idle path.
>>>> This patch introduces a new mechanism to poll for a while before
>>>> entering idle state. If schedule is needed during poll, then we
>>>> don't need to goes through the heavy overhead path.
>>>>
>>>> Here is the data i get when running benchmark contextswitch
>>>> (https://github.com/tsuna/contextswitch)
>>>> before patch:
>>>> 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw)
>>>> after patch:
>>>> 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw)
>>>
>>>
>>> If you test this after disabling the adaptive halt-polling in kvm?
>>> What's the performance data of w/ this patchset and w/o the adaptive
>>> halt-polling in kvm, and w/o this patchset and w/ the adaptive
>>> halt-polling in kvm? In addition, both linux and windows guests can
>>> get benefit as we have already done this in kvm.
>>
>>
>> I will provide more data in next version. But it doesn't conflict with
>
> Another case I can think of is w/ both this patchset and the adaptive
> halt-polling in kvm.
>
>> current halt polling inside kvm. This is just another enhancement.
>
> I didn't look close to the patchset, however, maybe there is another
> poll in the kvm part again sometimes if you fails the poll in the
> guest. In addition, the adaptive halt-polling in kvm has performance
> penalty when the pCPU is heavily overcommitted though there is a
> single_task_running() in my testing, it is hard to accurately aware
> whether there are other tasks waiting on the pCPU in the guest which
> will make it worser. Depending on vcpu_is_preempted() or steal time
> maybe not accurately or directly.
>
> So I'm not sure how much sense it makes by adaptive halt-polling in
> both guest and kvm. I prefer to just keep adaptive halt-polling in
> kvm(then both linux/windows or other guests can get benefit) and avoid
> to churn the core x86 path.

This mechanism is not specific to KVM. It is a kernel feature which can 
benefit guest when running inside X86 virtualization environment. The 
guest includes KVM,Xen,VMWARE,Hyper-v. Administrator can control KVM to 
use adaptive halt poll but he cannot control the user to use halt 
polling inside guest. Lots of user set idle=poll inside guest to improve 
performance which occupy more CPU cycles. This mechanism is a 
enhancement to it not to KVM halt polling.

>
> Regards,
> Wanpeng Li
>


-- 
Yang
Alibaba Cloud Computing

[toc] | [prev] | [next] | [standalone]


#1675731

FromRadim Krčmář <rkrcmar@redhat.com>
Date2017-06-27 16:10 +0200
Message-ID<tWZlL-1Bo-7@gated-at.bofh.it>
In reply to#1673297
2017-06-23 14:49+0800, Yang Zhang:
> On 2017/6/23 12:35, Wanpeng Li wrote:
> > 2017-06-23 12:08 GMT+08:00 Yang Zhang <yang.zhang.wz@gmail.com>:
> > > On 2017/6/22 19:50, Wanpeng Li wrote:
> > > > 
> > > > 2017-06-22 19:22 GMT+08:00 root <yang.zhang.wz@gmail.com>:
> > > > > 
> > > > > From: Yang Zhang <yang.zhang.wz@gmail.com>
> > > > > 
> > > > > Some latency-intensive workload will see obviously performance
> > > > > drop when running inside VM. The main reason is that the overhead
> > > > > is amplified when running inside VM. The most cost i have seen is
> > > > > inside idle path.
> > > > > This patch introduces a new mechanism to poll for a while before
> > > > > entering idle state. If schedule is needed during poll, then we
> > > > > don't need to goes through the heavy overhead path.
> > > > > 
> > > > > Here is the data i get when running benchmark contextswitch
> > > > > (https://github.com/tsuna/contextswitch)
> > > > > before patch:
> > > > > 2000000 process context switches in 4822613801ns (2411.3ns/ctxsw)
> > > > > after patch:
> > > > > 2000000 process context switches in 3584098241ns (1792.0ns/ctxsw)
> > > > 
> > > > 
> > > > If you test this after disabling the adaptive halt-polling in kvm?
> > > > What's the performance data of w/ this patchset and w/o the adaptive
> > > > halt-polling in kvm, and w/o this patchset and w/ the adaptive
> > > > halt-polling in kvm? In addition, both linux and windows guests can
> > > > get benefit as we have already done this in kvm.
> > > 
> > > 
> > > I will provide more data in next version. But it doesn't conflict with
> > 
> > Another case I can think of is w/ both this patchset and the adaptive
> > halt-polling in kvm.
> > 
> > > current halt polling inside kvm. This is just another enhancement.
> > 
> > I didn't look close to the patchset, however, maybe there is another
> > poll in the kvm part again sometimes if you fails the poll in the
> > guest. In addition, the adaptive halt-polling in kvm has performance
> > penalty when the pCPU is heavily overcommitted though there is a
> > single_task_running() in my testing, it is hard to accurately aware
> > whether there are other tasks waiting on the pCPU in the guest which
> > will make it worser. Depending on vcpu_is_preempted() or steal time
> > maybe not accurately or directly.
> > 
> > So I'm not sure how much sense it makes by adaptive halt-polling in
> > both guest and kvm. I prefer to just keep adaptive halt-polling in
> > kvm(then both linux/windows or other guests can get benefit) and avoid
> > to churn the core x86 path.
> 
> This mechanism is not specific to KVM. It is a kernel feature which can
> benefit guest when running inside X86 virtualization environment. The guest
> includes KVM,Xen,VMWARE,Hyper-v. Administrator can control KVM to use
> adaptive halt poll but he cannot control the user to use halt polling inside
> guest. Lots of user set idle=poll inside guest to improve performance which
> occupy more CPU cycles. This mechanism is a enhancement to it not to KVM
> halt polling.

Users of idle=poll shouln't overcommit, so the goal seems to be energy
savings without crippling the guest performance too much ...

Wouldn't switching to idle=mwait work as well?

Thanks.

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web