Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1731952 > unrolled thread
| Started by | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| First post | 2017-09-13 23:50 +0200 |
| Last post | 2017-09-19 11:20 +0200 |
| Articles | 4 — 4 participants |
Back to article view | Back to linux.kernel
[patch 00/52] x86: Rework the vector management Thomas Gleixner <tglx@linutronix.de> - 2017-09-13 23:50 +0200
Re: [patch 00/52] x86: Rework the vector management Juergen Gross <jgross@suse.com> - 2017-09-14 13:30 +0200
Re: [patch 00/52] x86: Rework the vector management Paolo Bonzini <pbonzini@redhat.com> - 2017-09-20 12:30 +0200
Re: [patch 00/52] x86: Rework the vector management Yu Chen <yu.c.chen@intel.com> - 2017-09-19 11:20 +0200
| From | Thomas Gleixner <tglx@linutronix.de> |
|---|---|
| Date | 2017-09-13 23:50 +0200 |
| Subject | [patch 00/52] x86: Rework the vector management |
| Message-ID | <upny2-4iM-3@gated-at.bofh.it> |
Sorry for the large CC list, but this is a major surgery.
The vector management in x86 including the surrounding code is a
conglomorate of ancient bits and pieces which have been subject to
'modernization' and featuritis over the years. The most obscure parts are
the vector allocation mechanics, the cleanup vector handling and the cpu
hotplug machinery. Replacing these pieces of art was on my todo list for a
long time.
Recent attempts to 'solve' CPU offline / hibernation issues which are
partially caused by the current vector management implementation made me
look for real. Further information in this thread:
http://lkml.kernel.org/r/cover.1504235838.git.yu.c.chen@intel.com
Aside of drivers allocating gazillion of interrupts, there are quite some
things which can be addressed in the x86 vector management and in the core
code.
- Multi CPU affinities:
A dubious property which is not available on all machines and causes
major complexity both in the allocator and the cleanup/hotplug
management. See:
http://lkml.kernel.org/r/alpine.DEB.2.20.1709071045440.1827@nanos
- Priority level spreading:
An obscure and undocumented property which I think is sufficiently
argued to be not required in:
http://lkml.kernel.org/r/alpine.DEB.2.20.1709071045440.1827@nanos
- Allocation of vectors when interrupt descriptors are allocated.
This is a historical implementation detail, which is not really
required when the vector allocation is delayed up to the point when
request_irq() is invoked. This might make request_irq() fail, when the
vector space is exhausted, but drivers should handle request_irq()
fails anyway.
The upside of changing this is that the active vector space becomes
smaller especially on hibernation/cpu offline when drivers shut down
queue interrupts of outgoing CPUs.
Some of this is already addressed with the managed interrupt facility,
but that was bolted on top of the existing vector management because
proper integration was not possible at that point. I take the blame
for this, but the tradeoff of not doing it would have been more
broken driver boiler plate code all over the place. So I went for the
lesser of two evils.
- Allocation of vectors on the wrong place
Even for managed interrupts the vector allocation at descriptor
allocation happens on the wrong place and gets fixed after the fact
with a call to set_affinity(). In case of not remapped interrupts
this results in at least one interrupt on the wrong CPU before it is
migrated to the desired target.
- Lack of instrumentation
All of this is a black box which allows no insight into the actual
vector usage.
The series addresses these points and converts the x86 vector management to
a bitmap based allocator which provides proper reservation management for
'managed interrupts' and best effort reservation for regular interrupts.
The latter allows overcommitment, which 'fixes' some of hotplug/hibernation
problems in a clean way. It can't fix all of them depending on the driver
involved.
This rework is no excuse for driver writers to do exhaustive vector
allocations instead of utilizing the managed interrupt infrastructure, but
it addresses long standing issues in this code with the side effect of
mitigating some of the driver oddities. The proper solution for multi queue
management are 'managed interrupts' which has been proven in the block-mq
work as they solve issues which are worked around in other drivers in
creative ways with lots of copied code and often enough broken attempts to
handle interrupt affinity and CPU hotplug problems.
The new bitmap allocator and the x86 vector management code are
instrumented with tracepoints and the irq domain debugfs files allow deep
insight into the vector allocation and reservations.
The patches work on machines with and without interrupt remapping and
inside of KVM guests of various flavours, though I have no idea what I
broke on the way with other hypervisors, posted interrupts etc. So I kindly
ask for your support in testing and review.
The series applies on top of Linus tree and is available as git branch:
git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git WIP.x86/apic
Note, that this branch is Linus tree plus scheduler and x86 fixes which I
required to do proper testing. They have outstanding pull requests and
might be merged already when you read this.
Thanks,
tglx
---
arch/x86/include/asm/x2apic.h | 49 -
b/arch/x86/Kconfig | 1
b/arch/x86/include/asm/apic.h | 255 +-----
b/arch/x86/include/asm/desc.h | 2
b/arch/x86/include/asm/hw_irq.h | 6
b/arch/x86/include/asm/io_apic.h | 2
b/arch/x86/include/asm/irq.h | 4
b/arch/x86/include/asm/irq_vectors.h | 8
b/arch/x86/include/asm/irqdomain.h | 5
b/arch/x86/include/asm/kvm_host.h | 2
b/arch/x86/include/asm/trace/irq_vectors.h | 244 ++++++
b/arch/x86/kernel/apic/Makefile | 2
b/arch/x86/kernel/apic/apic.c | 38 -
b/arch/x86/kernel/apic/apic_common.c | 46 +
b/arch/x86/kernel/apic/apic_flat_64.c | 10
b/arch/x86/kernel/apic/apic_noop.c | 25
b/arch/x86/kernel/apic/apic_numachip.c | 12
b/arch/x86/kernel/apic/bigsmp_32.c | 8
b/arch/x86/kernel/apic/htirq.c | 5
b/arch/x86/kernel/apic/io_apic.c | 94 --
b/arch/x86/kernel/apic/msi.c | 5
b/arch/x86/kernel/apic/probe_32.c | 29
b/arch/x86/kernel/apic/vector.c | 1090 +++++++++++++++++------------
b/arch/x86/kernel/apic/x2apic.h | 9
b/arch/x86/kernel/apic/x2apic_cluster.c | 196 +----
b/arch/x86/kernel/apic/x2apic_phys.c | 44 +
b/arch/x86/kernel/apic/x2apic_uv_x.c | 17
b/arch/x86/kernel/i8259.c | 1
b/arch/x86/kernel/idt.c | 12
b/arch/x86/kernel/irq.c | 101 --
b/arch/x86/kernel/irqinit.c | 1
b/arch/x86/kernel/setup.c | 12
b/arch/x86/kernel/smpboot.c | 14
b/arch/x86/kernel/traps.c | 2
b/arch/x86/kernel/vsmp_64.c | 19
b/arch/x86/platform/uv/uv_irq.c | 5
b/arch/x86/xen/apic.c | 6
b/drivers/gpio/gpio-xgene-sb.c | 7
b/drivers/iommu/amd_iommu.c | 44 -
b/drivers/iommu/intel_irq_remapping.c | 43 -
b/drivers/irqchip/irq-gic-v3-its.c | 5
b/drivers/pinctrl/stm32/pinctrl-stm32.c | 5
b/include/linux/irq.h | 22
b/include/linux/irqdesc.h | 1
b/include/linux/irqdomain.h | 14
b/include/linux/msi.h | 5
b/include/trace/events/irq_matrix.h | 201 +++++
b/kernel/irq/Kconfig | 3
b/kernel/irq/Makefile | 1
b/kernel/irq/autoprobe.c | 2
b/kernel/irq/chip.c | 37
b/kernel/irq/debugfs.c | 12
b/kernel/irq/internals.h | 19
b/kernel/irq/irqdesc.c | 3
b/kernel/irq/irqdomain.c | 43 -
b/kernel/irq/manage.c | 18
b/kernel/irq/matrix.c | 443 +++++++++++
b/kernel/irq/msi.c | 32
58 files changed, 2133 insertions(+), 1208 deletions(-)
[toc] | [next] | [standalone]
| From | Juergen Gross <jgross@suse.com> |
|---|---|
| Date | 2017-09-14 13:30 +0200 |
| Message-ID | <upAvf-4nM-11@gated-at.bofh.it> |
| In reply to | #1731952 |
On 13/09/17 23:29, Thomas Gleixner wrote: > Sorry for the large CC list, but this is a major surgery. > > The vector management in x86 including the surrounding code is a > conglomorate of ancient bits and pieces which have been subject to > 'modernization' and featuritis over the years. The most obscure parts are > the vector allocation mechanics, the cleanup vector handling and the cpu > hotplug machinery. Replacing these pieces of art was on my todo list for a > long time. > > Recent attempts to 'solve' CPU offline / hibernation issues which are > partially caused by the current vector management implementation made me > look for real. Further information in this thread: > > http://lkml.kernel.org/r/cover.1504235838.git.yu.c.chen@intel.com > > Aside of drivers allocating gazillion of interrupts, there are quite some > things which can be addressed in the x86 vector management and in the core > code. > > - Multi CPU affinities: > > A dubious property which is not available on all machines and causes > major complexity both in the allocator and the cleanup/hotplug > management. See: > > http://lkml.kernel.org/r/alpine.DEB.2.20.1709071045440.1827@nanos > > - Priority level spreading: > > An obscure and undocumented property which I think is sufficiently > argued to be not required in: > > http://lkml.kernel.org/r/alpine.DEB.2.20.1709071045440.1827@nanos > > - Allocation of vectors when interrupt descriptors are allocated. > > This is a historical implementation detail, which is not really > required when the vector allocation is delayed up to the point when > request_irq() is invoked. This might make request_irq() fail, when the > vector space is exhausted, but drivers should handle request_irq() > fails anyway. > > The upside of changing this is that the active vector space becomes > smaller especially on hibernation/cpu offline when drivers shut down > queue interrupts of outgoing CPUs. > > Some of this is already addressed with the managed interrupt facility, > but that was bolted on top of the existing vector management because > proper integration was not possible at that point. I take the blame > for this, but the tradeoff of not doing it would have been more > broken driver boiler plate code all over the place. So I went for the > lesser of two evils. > > - Allocation of vectors on the wrong place > > Even for managed interrupts the vector allocation at descriptor > allocation happens on the wrong place and gets fixed after the fact > with a call to set_affinity(). In case of not remapped interrupts > this results in at least one interrupt on the wrong CPU before it is > migrated to the desired target. > > - Lack of instrumentation > > All of this is a black box which allows no insight into the actual > vector usage. > > The series addresses these points and converts the x86 vector management to > a bitmap based allocator which provides proper reservation management for > 'managed interrupts' and best effort reservation for regular interrupts. > The latter allows overcommitment, which 'fixes' some of hotplug/hibernation > problems in a clean way. It can't fix all of them depending on the driver > involved. > > This rework is no excuse for driver writers to do exhaustive vector > allocations instead of utilizing the managed interrupt infrastructure, but > it addresses long standing issues in this code with the side effect of > mitigating some of the driver oddities. The proper solution for multi queue > management are 'managed interrupts' which has been proven in the block-mq > work as they solve issues which are worked around in other drivers in > creative ways with lots of copied code and often enough broken attempts to > handle interrupt affinity and CPU hotplug problems. > > The new bitmap allocator and the x86 vector management code are > instrumented with tracepoints and the irq domain debugfs files allow deep > insight into the vector allocation and reservations. > > The patches work on machines with and without interrupt remapping and > inside of KVM guests of various flavours, though I have no idea what I > broke on the way with other hypervisors, posted interrupts etc. So I kindly > ask for your support in testing and review. > > The series applies on top of Linus tree and is available as git branch: > > git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git WIP.x86/apic > > Note, that this branch is Linus tree plus scheduler and x86 fixes which I > required to do proper testing. They have outstanding pull requests and > might be merged already when you read this. > > Thanks, > > tglx > --- > arch/x86/include/asm/x2apic.h | 49 - > b/arch/x86/Kconfig | 1 > b/arch/x86/include/asm/apic.h | 255 +----- > b/arch/x86/include/asm/desc.h | 2 > b/arch/x86/include/asm/hw_irq.h | 6 > b/arch/x86/include/asm/io_apic.h | 2 > b/arch/x86/include/asm/irq.h | 4 > b/arch/x86/include/asm/irq_vectors.h | 8 > b/arch/x86/include/asm/irqdomain.h | 5 > b/arch/x86/include/asm/kvm_host.h | 2 > b/arch/x86/include/asm/trace/irq_vectors.h | 244 ++++++ > b/arch/x86/kernel/apic/Makefile | 2 > b/arch/x86/kernel/apic/apic.c | 38 - > b/arch/x86/kernel/apic/apic_common.c | 46 + > b/arch/x86/kernel/apic/apic_flat_64.c | 10 > b/arch/x86/kernel/apic/apic_noop.c | 25 > b/arch/x86/kernel/apic/apic_numachip.c | 12 > b/arch/x86/kernel/apic/bigsmp_32.c | 8 > b/arch/x86/kernel/apic/htirq.c | 5 > b/arch/x86/kernel/apic/io_apic.c | 94 -- > b/arch/x86/kernel/apic/msi.c | 5 > b/arch/x86/kernel/apic/probe_32.c | 29 > b/arch/x86/kernel/apic/vector.c | 1090 +++++++++++++++++------------ > b/arch/x86/kernel/apic/x2apic.h | 9 > b/arch/x86/kernel/apic/x2apic_cluster.c | 196 +---- > b/arch/x86/kernel/apic/x2apic_phys.c | 44 + > b/arch/x86/kernel/apic/x2apic_uv_x.c | 17 > b/arch/x86/kernel/i8259.c | 1 > b/arch/x86/kernel/idt.c | 12 > b/arch/x86/kernel/irq.c | 101 -- > b/arch/x86/kernel/irqinit.c | 1 > b/arch/x86/kernel/setup.c | 12 > b/arch/x86/kernel/smpboot.c | 14 > b/arch/x86/kernel/traps.c | 2 > b/arch/x86/kernel/vsmp_64.c | 19 > b/arch/x86/platform/uv/uv_irq.c | 5 > b/arch/x86/xen/apic.c | 6 > b/drivers/gpio/gpio-xgene-sb.c | 7 > b/drivers/iommu/amd_iommu.c | 44 - > b/drivers/iommu/intel_irq_remapping.c | 43 - > b/drivers/irqchip/irq-gic-v3-its.c | 5 > b/drivers/pinctrl/stm32/pinctrl-stm32.c | 5 > b/include/linux/irq.h | 22 > b/include/linux/irqdesc.h | 1 > b/include/linux/irqdomain.h | 14 > b/include/linux/msi.h | 5 > b/include/trace/events/irq_matrix.h | 201 +++++ > b/kernel/irq/Kconfig | 3 > b/kernel/irq/Makefile | 1 > b/kernel/irq/autoprobe.c | 2 > b/kernel/irq/chip.c | 37 > b/kernel/irq/debugfs.c | 12 > b/kernel/irq/internals.h | 19 > b/kernel/irq/irqdesc.c | 3 > b/kernel/irq/irqdomain.c | 43 - > b/kernel/irq/manage.c | 18 > b/kernel/irq/matrix.c | 443 +++++++++++ > b/kernel/irq/msi.c | 32 > 58 files changed, 2133 insertions(+), 1208 deletions(-) Complete series tested with paravirt + xen enabled 64 bit kernel: bare metal boot okay boot as Xen dom0 okay boot as Xen pv-domU okay boot as Xen HVM-domU with PV-drivers okay Vcpu onlining/offlining in pv-domU okay So you can add my: Tested-by: Juergen Gross <jgross@suse.com> Acked-by: Juergen Gross <jgross@suse.com> Juergen
[toc] | [prev] | [next] | [standalone]
| From | Paolo Bonzini <pbonzini@redhat.com> |
|---|---|
| Date | 2017-09-20 12:30 +0200 |
| Message-ID | <urKqu-1H9-13@gated-at.bofh.it> |
| In reply to | #1732223 |
On 14/09/2017 13:21, Juergen Gross wrote: > Complete series tested with paravirt + xen enabled 64 bit kernel: > > bare metal boot okay > boot as Xen dom0 okay > boot as Xen pv-domU okay > boot as Xen HVM-domU with PV-drivers okay > Vcpu onlining/offlining in pv-domU okay Intel has now tested posted interrupts with no regression. Thanks, Paolo
[toc] | [prev] | [next] | [standalone]
| From | Yu Chen <yu.c.chen@intel.com> |
|---|---|
| Date | 2017-09-19 11:20 +0200 |
| Message-ID | <urmRb-2k0-1@gated-at.bofh.it> |
| In reply to | #1731952 |
On Wed, Sep 13, 2017 at 11:29:02PM +0200, Thomas Gleixner wrote:
> Sorry for the large CC list, but this is a major surgery.
>
> The vector management in x86 including the surrounding code is a
> conglomorate of ancient bits and pieces which have been subject to
> 'modernization' and featuritis over the years. The most obscure parts are
> the vector allocation mechanics, the cleanup vector handling and the cpu
> hotplug machinery. Replacing these pieces of art was on my todo list for a
> long time.
>
> Recent attempts to 'solve' CPU offline / hibernation issues which are
> partially caused by the current vector management implementation made me
> look for real. Further information in this thread:
>
> http://lkml.kernel.org/r/cover.1504235838.git.yu.c.chen@intel.com
>
> Aside of drivers allocating gazillion of interrupts, there are quite some
> things which can be addressed in the x86 vector management and in the core
> code.
>
> - Multi CPU affinities:
>
> A dubious property which is not available on all machines and causes
> major complexity both in the allocator and the cleanup/hotplug
> management. See:
>
> http://lkml.kernel.org/r/alpine.DEB.2.20.1709071045440.1827@nanos
>
> - Priority level spreading:
>
> An obscure and undocumented property which I think is sufficiently
> argued to be not required in:
>
> http://lkml.kernel.org/r/alpine.DEB.2.20.1709071045440.1827@nanos
>
> - Allocation of vectors when interrupt descriptors are allocated.
>
> This is a historical implementation detail, which is not really
> required when the vector allocation is delayed up to the point when
> request_irq() is invoked. This might make request_irq() fail, when the
> vector space is exhausted, but drivers should handle request_irq()
> fails anyway.
>
> The upside of changing this is that the active vector space becomes
> smaller especially on hibernation/cpu offline when drivers shut down
> queue interrupts of outgoing CPUs.
>
> Some of this is already addressed with the managed interrupt facility,
> but that was bolted on top of the existing vector management because
> proper integration was not possible at that point. I take the blame
> for this, but the tradeoff of not doing it would have been more
> broken driver boiler plate code all over the place. So I went for the
> lesser of two evils.
>
> - Allocation of vectors on the wrong place
>
> Even for managed interrupts the vector allocation at descriptor
> allocation happens on the wrong place and gets fixed after the fact
> with a call to set_affinity(). In case of not remapped interrupts
> this results in at least one interrupt on the wrong CPU before it is
> migrated to the desired target.
>
> - Lack of instrumentation
>
> All of this is a black box which allows no insight into the actual
> vector usage.
>
> The series addresses these points and converts the x86 vector management to
> a bitmap based allocator which provides proper reservation management for
> 'managed interrupts' and best effort reservation for regular interrupts.
> The latter allows overcommitment, which 'fixes' some of hotplug/hibernation
> problems in a clean way. It can't fix all of them depending on the driver
> involved.
>
> This rework is no excuse for driver writers to do exhaustive vector
> allocations instead of utilizing the managed interrupt infrastructure, but
> it addresses long standing issues in this code with the side effect of
> mitigating some of the driver oddities. The proper solution for multi queue
> management are 'managed interrupts' which has been proven in the block-mq
> work as they solve issues which are worked around in other drivers in
> creative ways with lots of copied code and often enough broken attempts to
> handle interrupt affinity and CPU hotplug problems.
>
> The new bitmap allocator and the x86 vector management code are
> instrumented with tracepoints and the irq domain debugfs files allow deep
> insight into the vector allocation and reservations.
>
> The patches work on machines with and without interrupt remapping and
> inside of KVM guests of various flavours, though I have no idea what I
> broke on the way with other hypervisors, posted interrupts etc. So I kindly
> ask for your support in testing and review.
>
> The series applies on top of Linus tree and is available as git branch:
>
> git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git WIP.x86/apic
>
> Note, that this branch is Linus tree plus scheduler and x86 fixes which I
> required to do proper testing. They have outstanding pull requests and
> might be merged already when you read this.
>
> Thanks,
>
> tglx
> ---
Tested on top of:
commit e1b476ae32fcfa59fc6752b4b01988e759269dc3
Author: Thomas Gleixner <tglx@linutronix.de>
Date: Thu Sep 14 09:53:10 2017 +0200
x86/vector: Exclude IRQ0 from reservation mode
from branch WIP.x86/apic, on a platform with 16 cores,
bootup okay, cpu[1-31] offline/online okay.
Before offline:
name: VECTOR
size: 0
mapped: 484
flags: 0x00000041
Online bitmaps: 32
Global available: 6419
Global reserved: 407
Total allocated: 77
System: 41: 0-19,32,50,128,238-255
| CPU | avl | man | act | vectors
0 126 0 77 33-49,51-110
1 203 0 0
2 203 0 0
3 203 0 0
4 203 0 0
5 203 0 0
6 203 0 0
7 203 0 0
8 203 0 0
9 203 0 0
10 203 0 0
11 203 0 0
12 203 0 0
13 203 0 0
14 203 0 0
15 203 0 0
16 203 0 0
17 203 0 0
18 203 0 0
19 203 0 0
20 203 0 0
21 203 0 0
22 203 0 0
23 203 0 0
24 203 0 0
25 203 0 0
26 203 0 0
27 203 0 0
28 203 0 0
29 203 0 0
30 203 0 0
31 203 0 0
After offline:
name: VECTOR
size: 0
mapped: 484
flags: 0x00000041
Online bitmaps: 1
Global available: 126
Global reserved: 407
Total allocated: 77
System: 41: 0-19,32,50,128,238-255
| CPU | avl | man | act | vectors
0 126 0 77 33-49,51-110
Thanks,
Yu
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web