Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1514870 > unrolled thread
| Started by | Eric Auger <eric.auger@redhat.com> |
|---|---|
| First post | 2016-11-03 22:50 +0100 |
| Last post | 2016-11-23 21:20 +0100 |
| Articles | 20 on this page of 47 — 9 participants |
Back to article view | Back to linux.kernel
[RFC 0/8] KVM PCIe/MSI passthrough on ARM/ARM64 (Alt II) Eric Auger <eric.auger@redhat.com> - 2016-11-03 22:50 +0100
[RFC 2/8] iommu/iova: fix __alloc_and_insert_iova_range Eric Auger <eric.auger@redhat.com> - 2016-11-03 22:50 +0100
[RFC 4/8] iommu: Add a list of iommu_reserved_region in iommu_domain Eric Auger <eric.auger@redhat.com> - 2016-11-03 22:50 +0100
[RFC 5/8] vfio/type1: Introduce RESV_IOVA_RANGE capability Eric Auger <eric.auger@redhat.com> - 2016-11-03 22:50 +0100
[RFC 3/8] iommu/dma: Allow MSI-only cookies Eric Auger <eric.auger@redhat.com> - 2016-11-03 22:50 +0100
[RFC 7/8] iommu/vt-d: Implement add_reserved_regions callback Eric Auger <eric.auger@redhat.com> - 2016-11-03 22:50 +0100
[RFC 1/8] vfio: fix vfio_info_cap_add/shift Eric Auger <eric.auger@redhat.com> - 2016-11-03 22:50 +0100
[RFC 6/8] iommu: Handle the list of reserved regions Eric Auger <eric.auger@redhat.com> - 2016-11-03 22:50 +0100
Re: [RFC 0/8] KVM PCIe/MSI passthrough on ARM/ARM64 (Alt II) Alex Williamson <alex.williamson@redhat.com> - 2016-11-04 05:10 +0100
Summary of LPC guest MSI discussion in Santa Fe (was: Re: [RFC 0/8] KVM PCIe/MSI passthrough on ARM/ARM64 (Alt II)) Will Deacon <will.deacon@arm.com> - 2016-11-08 03:50 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Auger Eric <eric.auger@redhat.com> - 2016-11-08 15:30 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Will Deacon <will.deacon@arm.com> - 2016-11-08 19:00 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Don Dutile <ddutile@redhat.com> - 2016-11-08 20:10 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Will Deacon <will.deacon@arm.com> - 2016-11-08 20:20 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Auger Eric <eric.auger@redhat.com> - 2016-11-09 08:50 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Don Dutile <ddutile@redhat.com> - 2016-11-08 17:10 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe (was: Re: [RFC 0/8] KVM PCIe/MSI passthrough on ARM/ARM64 (Alt II)) Christoffer Dall <christoffer.dall@linaro.org> - 2016-11-08 21:30 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe (was: Re: [RFC 0/8] KVM PCIe/MSI passthrough on ARM/ARM64 (Alt II)) Alex Williamson <alex.williamson@redhat.com> - 2016-11-09 00:40 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Don Dutile <ddutile@redhat.com> - 2016-11-09 04:00 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Will Deacon <will.deacon@arm.com> - 2016-11-09 18:10 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Don Dutile <ddutile@redhat.com> - 2016-11-09 20:00 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Christoffer Dall <christoffer.dall@linaro.org> - 2016-11-09 20:30 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Alex Williamson <alex.williamson@redhat.com> - 2016-11-09 21:10 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Joerg Roedel <joro@8bytes.org> - 2016-11-10 15:50 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Alex Williamson <alex.williamson@redhat.com> - 2016-11-10 18:10 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Will Deacon <will.deacon@arm.com> - 2016-11-09 21:40 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Alex Williamson <alex.williamson@redhat.com> - 2016-11-09 23:20 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Will Deacon <will.deacon@arm.com> - 2016-11-09 23:30 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Alex Williamson <alex.williamson@redhat.com> - 2016-11-10 00:30 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Will Deacon <will.deacon@arm.com> - 2016-11-10 00:40 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Alex Williamson <alex.williamson@redhat.com> - 2016-11-10 01:10 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Auger Eric <eric.auger@redhat.com> - 2016-11-10 01:20 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Alex Williamson <alex.williamson@redhat.com> - 2016-11-10 02:00 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Will Deacon <will.deacon@arm.com> - 2016-11-10 03:10 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Auger Eric <eric.auger@redhat.com> - 2016-11-10 12:20 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Alex Williamson <alex.williamson@redhat.com> - 2016-11-10 18:50 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Joerg Roedel <joro@8bytes.org> - 2016-11-11 12:20 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Alex Williamson <alex.williamson@redhat.com> - 2016-11-11 17:00 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Alex Williamson <alex.williamson@redhat.com> - 2016-11-11 17:10 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Joerg Roedel <joro@8bytes.org> - 2016-11-14 16:20 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Don Dutile <ddutile@redhat.com> - 2016-11-11 17:30 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Don Dutile <ddutile@redhat.com> - 2016-11-11 17:10 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Joerg Roedel <joro@8bytes.org> - 2016-11-10 16:00 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Robin Murphy <robin.murphy@arm.com> - 2016-11-09 21:20 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Joerg Roedel <joro@8bytes.org> - 2016-11-10 16:30 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Jon Masters <jcm@jonmasters.org> - 2016-11-21 07:00 +0100
Re: Summary of LPC guest MSI discussion in Santa Fe Don Dutile <ddutile@redhat.com> - 2016-11-23 21:20 +0100
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Don Dutile <ddutile@redhat.com> |
|---|---|
| Date | 2016-11-09 20:00 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBGgh-2DN-1@gated-at.bofh.it> |
| In reply to | #1518317 |
On 11/09/2016 12:03 PM, Will Deacon wrote: > On Tue, Nov 08, 2016 at 09:52:33PM -0500, Don Dutile wrote: >> On 11/08/2016 06:35 PM, Alex Williamson wrote: >>> On Tue, 8 Nov 2016 21:29:22 +0100 >>> Christoffer Dall <christoffer.dall@linaro.org> wrote: >>>> Is my understanding correct, that you need to tell userspace about the >>>> location of the doorbell (in the IOVA space) in case (2), because even >>>> though the configuration of the device is handled by the (host) kernel >>>> through trapping of the BARs, we have to avoid the VFIO user programming >>>> the device to create other DMA transactions to this particular address, >>>> since that will obviously conflict and either not produce the desired >>>> DMA transactions or result in unintended weird interrupts? > > Yes, that's the crux of the issue. > >>> Correct, if the MSI doorbell IOVA range overlaps RAM in the VM, then >>> it's potentially a DMA target and we'll get bogus data on DMA read from >>> the device, and lose data and potentially trigger spurious interrupts on >>> DMA write from the device. Thanks, >>> >> That's b/c the MSI doorbells are not positioned *above* the SMMU, i.e., >> they address match before the SMMU checks are done. if >> all DMA addrs had to go through SMMU first, then the DMA access could >> be ignored/rejected. > > That's actually not true :( The SMMU can't generally distinguish between MSI > writes and DMA writes, so it would just see a write transaction to the > doorbell address, regardless of how it was generated by the endpoint. > > Will > So, we have real systems where MSI doorbells are placed at the same IOVA that could have memory for a guest, but not at the same IOVA as memory on real hw ? How are memory holes passed to SMMU so it doesn't have this issue for bare-metal (assign an IOVA that overlaps an MSI doorbell address)?
[toc] | [prev] | [next] | [standalone]
| From | Christoffer Dall <christoffer.dall@linaro.org> |
|---|---|
| Date | 2016-11-09 20:30 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBGJj-34V-1@gated-at.bofh.it> |
| In reply to | #1518411 |
On Wed, Nov 09, 2016 at 01:59:07PM -0500, Don Dutile wrote: > On 11/09/2016 12:03 PM, Will Deacon wrote: > >On Tue, Nov 08, 2016 at 09:52:33PM -0500, Don Dutile wrote: > >>On 11/08/2016 06:35 PM, Alex Williamson wrote: > >>>On Tue, 8 Nov 2016 21:29:22 +0100 > >>>Christoffer Dall <christoffer.dall@linaro.org> wrote: > >>>>Is my understanding correct, that you need to tell userspace about the > >>>>location of the doorbell (in the IOVA space) in case (2), because even > >>>>though the configuration of the device is handled by the (host) kernel > >>>>through trapping of the BARs, we have to avoid the VFIO user programming > >>>>the device to create other DMA transactions to this particular address, > >>>>since that will obviously conflict and either not produce the desired > >>>>DMA transactions or result in unintended weird interrupts? > > > >Yes, that's the crux of the issue. > > > >>>Correct, if the MSI doorbell IOVA range overlaps RAM in the VM, then > >>>it's potentially a DMA target and we'll get bogus data on DMA read from > >>>the device, and lose data and potentially trigger spurious interrupts on > >>>DMA write from the device. Thanks, > >>> > >>That's b/c the MSI doorbells are not positioned *above* the SMMU, i.e., > >>they address match before the SMMU checks are done. if > >>all DMA addrs had to go through SMMU first, then the DMA access could > >>be ignored/rejected. > > > >That's actually not true :( The SMMU can't generally distinguish between MSI > >writes and DMA writes, so it would just see a write transaction to the > >doorbell address, regardless of how it was generated by the endpoint. > > > >Will > > > So, we have real systems where MSI doorbells are placed at the same IOVA > that could have memory for a guest I don't think this is a property of a hardware system. THe problem is userspace not knowing where in the IOVA space the kernel is going to place the doorbell, so you can end up (basically by chance) that some IPA range of guest memory overlaps with the IOVA space for the doorbell. >, but not at the same IOVA as memory on real hw ? On real hardware without an IOMMU the system designer would have to separate the IOVA and RAM in the physical address space. With an IOMMU, the SMMU driver just makes sure to allocate separate regions in the IOVA space. The challenge, as I understand it, happens with the VM, because the VM doesn't allocate the IOVA for the MSI doorbell itself, but the host kernel does this, independently from the attributes (e.g. memory map) of the VM. Because the IOVA is a single resource, but with two independent entities allocating chunks of it (the host kernel for the MSI doorbell IOVA, and the VFIO user for other DMA operations), you have to provide some coordination between those to entities to avoid conflicts. In the case of KVM, the two entities are the host kernel and the VFIO user (QEMU/the VM), and the host kernel informs the VFIO user to never attempt to use the doorbell IOVA already reserved by the host kernel for DMA. One way to do that is to ensure that the IPA space of the VFIO user corresponding to the doorbell IOVA is simply not valid, ie. the reserved regions that avoid for example QEMU to allocate RAM there. (I suppose it's technically possible to get around this issue by letting QEMU place RAM wherever it wants but tell the guest to never use a particular subset of its RAM for DMA, because that would conflict with the doorbell IOVA or be seen as p2p transactions. But I think we all probably agree that it's a disgusting idea.) > How are memory holes passed to SMMU so it doesn't have this issue for bare-metal > (assign an IOVA that overlaps an MSI doorbell address)? > As I understand it, the SMMU driver manages the whole IOVA space when VFIO is *not* involved, so it simply allocates non-overlapping regions. The problem occurs when you have two independent entities essentially attempting to mange the same resource (and the problem is exacerbated by the VM potentially allocating slots in the IOVA space which may have other limitations it doesn't know about, for example the p2p regions, because the VM doesn't know anything about the topology of the underlying physical system). Christoffer
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-09 21:10 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBHm1-3y6-11@gated-at.bofh.it> |
| In reply to | #1518428 |
On Wed, 9 Nov 2016 20:23:03 +0100 Christoffer Dall <christoffer.dall@linaro.org> wrote: > On Wed, Nov 09, 2016 at 01:59:07PM -0500, Don Dutile wrote: > > On 11/09/2016 12:03 PM, Will Deacon wrote: > > >On Tue, Nov 08, 2016 at 09:52:33PM -0500, Don Dutile wrote: > > >>On 11/08/2016 06:35 PM, Alex Williamson wrote: > > >>>On Tue, 8 Nov 2016 21:29:22 +0100 > > >>>Christoffer Dall <christoffer.dall@linaro.org> wrote: > > >>>>Is my understanding correct, that you need to tell userspace about the > > >>>>location of the doorbell (in the IOVA space) in case (2), because even > > >>>>though the configuration of the device is handled by the (host) kernel > > >>>>through trapping of the BARs, we have to avoid the VFIO user programming > > >>>>the device to create other DMA transactions to this particular address, > > >>>>since that will obviously conflict and either not produce the desired > > >>>>DMA transactions or result in unintended weird interrupts? > > > > > >Yes, that's the crux of the issue. > > > > > >>>Correct, if the MSI doorbell IOVA range overlaps RAM in the VM, then > > >>>it's potentially a DMA target and we'll get bogus data on DMA read from > > >>>the device, and lose data and potentially trigger spurious interrupts on > > >>>DMA write from the device. Thanks, > > >>> > > >>That's b/c the MSI doorbells are not positioned *above* the SMMU, i.e., > > >>they address match before the SMMU checks are done. if > > >>all DMA addrs had to go through SMMU first, then the DMA access could > > >>be ignored/rejected. > > > > > >That's actually not true :( The SMMU can't generally distinguish between MSI > > >writes and DMA writes, so it would just see a write transaction to the > > >doorbell address, regardless of how it was generated by the endpoint. > > > > > >Will > > > > > So, we have real systems where MSI doorbells are placed at the same IOVA > > that could have memory for a guest > > I don't think this is a property of a hardware system. THe problem is > userspace not knowing where in the IOVA space the kernel is going to > place the doorbell, so you can end up (basically by chance) that some > IPA range of guest memory overlaps with the IOVA space for the doorbell. > > > >, but not at the same IOVA as memory on real hw ? > > On real hardware without an IOMMU the system designer would have to > separate the IOVA and RAM in the physical address space. With an IOMMU, > the SMMU driver just makes sure to allocate separate regions in the IOVA > space. > > The challenge, as I understand it, happens with the VM, because the VM > doesn't allocate the IOVA for the MSI doorbell itself, but the host > kernel does this, independently from the attributes (e.g. memory map) of > the VM. > > Because the IOVA is a single resource, but with two independent entities > allocating chunks of it (the host kernel for the MSI doorbell IOVA, and > the VFIO user for other DMA operations), you have to provide some > coordination between those to entities to avoid conflicts. In the case > of KVM, the two entities are the host kernel and the VFIO user (QEMU/the > VM), and the host kernel informs the VFIO user to never attempt to use > the doorbell IOVA already reserved by the host kernel for DMA. > > One way to do that is to ensure that the IPA space of the VFIO user > corresponding to the doorbell IOVA is simply not valid, ie. the reserved > regions that avoid for example QEMU to allocate RAM there. > > (I suppose it's technically possible to get around this issue by letting > QEMU place RAM wherever it wants but tell the guest to never use a > particular subset of its RAM for DMA, because that would conflict with > the doorbell IOVA or be seen as p2p transactions. But I think we all > probably agree that it's a disgusting idea.) Well, it's not like QEMU or libvirt stumbling through sysfs to figure out where holes could be in order to instantiate a VM with matching holes, just in case someone might decide to hot-add a device into the VM, at some point, and hopefully they don't migrate the VM to another host with a different layout first, is all that much less disgusting or foolproof. It's just that in order to dynamically remove a page as a possible DMA target we require a paravirt channel, such as a balloon driver that's able to pluck a specific page. In some ways it's actually less disgusting, but it puts some prerequisites on enlightening the guest OS. Thanks, Alex
[toc] | [prev] | [next] | [standalone]
| From | Joerg Roedel <joro@8bytes.org> |
|---|---|
| Date | 2016-11-10 15:50 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBYPT-7uF-5@gated-at.bofh.it> |
| In reply to | #1518445 |
On Wed, Nov 09, 2016 at 01:01:14PM -0700, Alex Williamson wrote: > Well, it's not like QEMU or libvirt stumbling through sysfs to figure > out where holes could be in order to instantiate a VM with matching > holes, just in case someone might decide to hot-add a device into the > VM, at some point, and hopefully they don't migrate the VM to another > host with a different layout first, is all that much less disgusting or > foolproof. It's just that in order to dynamically remove a page as a > possible DMA target we require a paravirt channel, such as a balloon > driver that's able to pluck a specific page. In some ways it's > actually less disgusting, but it puts some prerequisites on > enlightening the guest OS. Thanks, I think it is much simpler if libvirt/qemu just go through all potentially assignable devices on a system and pre-exclude any addresses from guest RAM beforehand, rather than doing something like this with paravirt/ballooning when a device is hot-added. There is no guarantee that you can take a page away from a linux-guest. Joerg
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-10 18:10 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sC11n-Cz-7@gated-at.bofh.it> |
| In reply to | #1519021 |
On Thu, 10 Nov 2016 15:40:07 +0100 Joerg Roedel <joro@8bytes.org> wrote: > On Wed, Nov 09, 2016 at 01:01:14PM -0700, Alex Williamson wrote: > > Well, it's not like QEMU or libvirt stumbling through sysfs to figure > > out where holes could be in order to instantiate a VM with matching > > holes, just in case someone might decide to hot-add a device into the > > VM, at some point, and hopefully they don't migrate the VM to another > > host with a different layout first, is all that much less disgusting or > > foolproof. It's just that in order to dynamically remove a page as a > > possible DMA target we require a paravirt channel, such as a balloon > > driver that's able to pluck a specific page. In some ways it's > > actually less disgusting, but it puts some prerequisites on > > enlightening the guest OS. Thanks, > > I think it is much simpler if libvirt/qemu just go through all > potentially assignable devices on a system and pre-exclude any addresses > from guest RAM beforehand, rather than doing something like this with > paravirt/ballooning when a device is hot-added. There is no guarantee > that you can take a page away from a linux-guest. Right, I'm not advocating a paravirt ballooning approach, just pointing out that they're both terrible, just in different ways. Thanks, Alex
[toc] | [prev] | [next] | [standalone]
| From | Will Deacon <will.deacon@arm.com> |
|---|---|
| Date | 2016-11-09 21:40 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBHP4-3Kz-31@gated-at.bofh.it> |
| In reply to | #1518428 |
On Wed, Nov 09, 2016 at 08:23:03PM +0100, Christoffer Dall wrote: > On Wed, Nov 09, 2016 at 01:59:07PM -0500, Don Dutile wrote: > > On 11/09/2016 12:03 PM, Will Deacon wrote: > > >On Tue, Nov 08, 2016 at 09:52:33PM -0500, Don Dutile wrote: > > >>On 11/08/2016 06:35 PM, Alex Williamson wrote: > > >>>Correct, if the MSI doorbell IOVA range overlaps RAM in the VM, then > > >>>it's potentially a DMA target and we'll get bogus data on DMA read from > > >>>the device, and lose data and potentially trigger spurious interrupts on > > >>>DMA write from the device. Thanks, > > >>> > > >>That's b/c the MSI doorbells are not positioned *above* the SMMU, i.e., > > >>they address match before the SMMU checks are done. if > > >>all DMA addrs had to go through SMMU first, then the DMA access could > > >>be ignored/rejected. > > > > > >That's actually not true :( The SMMU can't generally distinguish between MSI > > >writes and DMA writes, so it would just see a write transaction to the > > >doorbell address, regardless of how it was generated by the endpoint. > > > > > So, we have real systems where MSI doorbells are placed at the same IOVA > > that could have memory for a guest > > I don't think this is a property of a hardware system. THe problem is > userspace not knowing where in the IOVA space the kernel is going to > place the doorbell, so you can end up (basically by chance) that some > IPA range of guest memory overlaps with the IOVA space for the doorbell. I think the case that Don has in mind is where the host is using the SMMU for DMA mapping. In that case, yes, the IOVAs assigned by things like dma_map_single mustn't collide with any fixed MSI mapping. We currently take care to avoid PCI windows, but nobody has added the code for the fixed MSI mappings yet (I think we should put the onus on the people with the broken systems for that). Depending on how the firmware describes the fixed MSI address, either the irqchip driver can take care of it in compose_msi_msg, or we could do something in the iommu_dma_map_msi_msg path to ensure that the fixed region is preallocated in the msi_page_list. I'm less fussed about this issue because there's not a user ABI involved, so it can all be fixed later. > >, but not at the same IOVA as memory on real hw ? > > On real hardware without an IOMMU the system designer would have to > separate the IOVA and RAM in the physical address space. With an IOMMU, > the SMMU driver just makes sure to allocate separate regions in the IOVA > space. > > The challenge, as I understand it, happens with the VM, because the VM > doesn't allocate the IOVA for the MSI doorbell itself, but the host > kernel does this, independently from the attributes (e.g. memory map) of > the VM. > > Because the IOVA is a single resource, but with two independent entities > allocating chunks of it (the host kernel for the MSI doorbell IOVA, and > the VFIO user for other DMA operations), you have to provide some > coordination between those to entities to avoid conflicts. In the case > of KVM, the two entities are the host kernel and the VFIO user (QEMU/the > VM), and the host kernel informs the VFIO user to never attempt to use > the doorbell IOVA already reserved by the host kernel for DMA. > > One way to do that is to ensure that the IPA space of the VFIO user > corresponding to the doorbell IOVA is simply not valid, ie. the reserved > regions that avoid for example QEMU to allocate RAM there. > > (I suppose it's technically possible to get around this issue by letting > QEMU place RAM wherever it wants but tell the guest to never use a > particular subset of its RAM for DMA, because that would conflict with > the doorbell IOVA or be seen as p2p transactions. But I think we all > probably agree that it's a disgusting idea.) Disgusting, yes, but Ben's idea of hotplugging on the host controller with firmware tables describing the reserved regions is something that we could do in the distant future. In the meantime, I don't think that VFIO should explicitly reject overlapping mappings if userspace asks for them. Will
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-09 23:20 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBJnP-50l-1@gated-at.bofh.it> |
| In reply to | #1518466 |
On Wed, 9 Nov 2016 20:31:45 +0000 Will Deacon <will.deacon@arm.com> wrote: > On Wed, Nov 09, 2016 at 08:23:03PM +0100, Christoffer Dall wrote: > > > > (I suppose it's technically possible to get around this issue by letting > > QEMU place RAM wherever it wants but tell the guest to never use a > > particular subset of its RAM for DMA, because that would conflict with > > the doorbell IOVA or be seen as p2p transactions. But I think we all > > probably agree that it's a disgusting idea.) > > Disgusting, yes, but Ben's idea of hotplugging on the host controller with > firmware tables describing the reserved regions is something that we could > do in the distant future. In the meantime, I don't think that VFIO should > explicitly reject overlapping mappings if userspace asks for them. I'm confused by the last sentence here, rejecting user mappings that overlap reserved ranges, such as MSI doorbell pages, is exactly how we'd reject hot-adding a device when we meet such a conflict. If we don't reject such a mapping, we're knowingly creating a situation that potentially leads to data loss. Minimally, QEMU would need to know about the reserved region, map around it through VFIO, and take responsibility (somehow) for making sure that region is never used for DMA. Thanks, Alex
[toc] | [prev] | [next] | [standalone]
| From | Will Deacon <will.deacon@arm.com> |
|---|---|
| Date | 2016-11-09 23:30 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBJxw-56q-23@gated-at.bofh.it> |
| In reply to | #1518519 |
On Wed, Nov 09, 2016 at 03:17:09PM -0700, Alex Williamson wrote: > On Wed, 9 Nov 2016 20:31:45 +0000 > Will Deacon <will.deacon@arm.com> wrote: > > On Wed, Nov 09, 2016 at 08:23:03PM +0100, Christoffer Dall wrote: > > > > > > (I suppose it's technically possible to get around this issue by letting > > > QEMU place RAM wherever it wants but tell the guest to never use a > > > particular subset of its RAM for DMA, because that would conflict with > > > the doorbell IOVA or be seen as p2p transactions. But I think we all > > > probably agree that it's a disgusting idea.) > > > > Disgusting, yes, but Ben's idea of hotplugging on the host controller with > > firmware tables describing the reserved regions is something that we could > > do in the distant future. In the meantime, I don't think that VFIO should > > explicitly reject overlapping mappings if userspace asks for them. > > I'm confused by the last sentence here, rejecting user mappings that > overlap reserved ranges, such as MSI doorbell pages, is exactly how > we'd reject hot-adding a device when we meet such a conflict. If we > don't reject such a mapping, we're knowingly creating a situation that > potentially leads to data loss. Minimally, QEMU would need to know > about the reserved region, map around it through VFIO, and take > responsibility (somehow) for making sure that region is never used for > DMA. Thanks, Yes, but my point is that it should be up to QEMU to abort the hotplug, not the host kernel, since there may be ways in which a guest can tolerate the overlapping region (e.g. by avoiding that range of memory for DMA). Will
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-10 00:30 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBKtz-5LS-13@gated-at.bofh.it> |
| In reply to | #1518529 |
On Wed, 9 Nov 2016 22:25:22 +0000 Will Deacon <will.deacon@arm.com> wrote: > On Wed, Nov 09, 2016 at 03:17:09PM -0700, Alex Williamson wrote: > > On Wed, 9 Nov 2016 20:31:45 +0000 > > Will Deacon <will.deacon@arm.com> wrote: > > > On Wed, Nov 09, 2016 at 08:23:03PM +0100, Christoffer Dall wrote: > > > > > > > > (I suppose it's technically possible to get around this issue by letting > > > > QEMU place RAM wherever it wants but tell the guest to never use a > > > > particular subset of its RAM for DMA, because that would conflict with > > > > the doorbell IOVA or be seen as p2p transactions. But I think we all > > > > probably agree that it's a disgusting idea.) > > > > > > Disgusting, yes, but Ben's idea of hotplugging on the host controller with > > > firmware tables describing the reserved regions is something that we could > > > do in the distant future. In the meantime, I don't think that VFIO should > > > explicitly reject overlapping mappings if userspace asks for them. > > > > I'm confused by the last sentence here, rejecting user mappings that > > overlap reserved ranges, such as MSI doorbell pages, is exactly how > > we'd reject hot-adding a device when we meet such a conflict. If we > > don't reject such a mapping, we're knowingly creating a situation that > > potentially leads to data loss. Minimally, QEMU would need to know > > about the reserved region, map around it through VFIO, and take > > responsibility (somehow) for making sure that region is never used for > > DMA. Thanks, > > Yes, but my point is that it should be up to QEMU to abort the hotplug, not > the host kernel, since there may be ways in which a guest can tolerate the > overlapping region (e.g. by avoiding that range of memory for DMA). The VFIO_IOMMU_MAP_DMA ioctl is a contract, the user ask to map a range of IOVAs to a range of virtual addresses for a given device. If VFIO cannot reasonably fulfill that contract, it must fail. It's up to QEMU how to manage the hotplug and what memory regions it asks VFIO to map for a device, but VFIO must reject mappings that it (or the SMMU by virtue of using the IOMMU API) know to overlap reserved ranges. So I still disagree with the referenced statement. Thanks, Alex
[toc] | [prev] | [next] | [standalone]
| From | Will Deacon <will.deacon@arm.com> |
|---|---|
| Date | 2016-11-10 00:40 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBKDf-5OO-25@gated-at.bofh.it> |
| In reply to | #1518563 |
On Wed, Nov 09, 2016 at 04:24:58PM -0700, Alex Williamson wrote: > On Wed, 9 Nov 2016 22:25:22 +0000 > Will Deacon <will.deacon@arm.com> wrote: > > > On Wed, Nov 09, 2016 at 03:17:09PM -0700, Alex Williamson wrote: > > > On Wed, 9 Nov 2016 20:31:45 +0000 > > > Will Deacon <will.deacon@arm.com> wrote: > > > > On Wed, Nov 09, 2016 at 08:23:03PM +0100, Christoffer Dall wrote: > > > > > > > > > > (I suppose it's technically possible to get around this issue by letting > > > > > QEMU place RAM wherever it wants but tell the guest to never use a > > > > > particular subset of its RAM for DMA, because that would conflict with > > > > > the doorbell IOVA or be seen as p2p transactions. But I think we all > > > > > probably agree that it's a disgusting idea.) > > > > > > > > Disgusting, yes, but Ben's idea of hotplugging on the host controller with > > > > firmware tables describing the reserved regions is something that we could > > > > do in the distant future. In the meantime, I don't think that VFIO should > > > > explicitly reject overlapping mappings if userspace asks for them. > > > > > > I'm confused by the last sentence here, rejecting user mappings that > > > overlap reserved ranges, such as MSI doorbell pages, is exactly how > > > we'd reject hot-adding a device when we meet such a conflict. If we > > > don't reject such a mapping, we're knowingly creating a situation that > > > potentially leads to data loss. Minimally, QEMU would need to know > > > about the reserved region, map around it through VFIO, and take > > > responsibility (somehow) for making sure that region is never used for > > > DMA. Thanks, > > > > Yes, but my point is that it should be up to QEMU to abort the hotplug, not > > the host kernel, since there may be ways in which a guest can tolerate the > > overlapping region (e.g. by avoiding that range of memory for DMA). > > The VFIO_IOMMU_MAP_DMA ioctl is a contract, the user ask to map a range > of IOVAs to a range of virtual addresses for a given device. If VFIO > cannot reasonably fulfill that contract, it must fail. It's up to QEMU > how to manage the hotplug and what memory regions it asks VFIO to map > for a device, but VFIO must reject mappings that it (or the SMMU by > virtue of using the IOMMU API) know to overlap reserved ranges. So I > still disagree with the referenced statement. Thanks, I think that's a pity. Not only does it mean that both QEMU and the kernel have more work to do (the former has to carve up its mapping requests, whilst the latter has to check that it is indeed doing this), but it also precludes the use of hugepage mappings on the IOMMU because of reserved regions. For example, a 4k hole someplace may mean we can't put down 1GB table entries for the guest memory in the SMMU. All this seems to do is add complexity and decrease performance. For what? QEMU has to go read the reserved regions from someplace anyway. It's also the way that VFIO works *today* on arm64 wrt reserved regions, it just has no way to identify those holes at present. Will
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-10 01:10 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBL6i-6fT-19@gated-at.bofh.it> |
| In reply to | #1518567 |
On Wed, 9 Nov 2016 23:38:50 +0000 Will Deacon <will.deacon@arm.com> wrote: > On Wed, Nov 09, 2016 at 04:24:58PM -0700, Alex Williamson wrote: > > On Wed, 9 Nov 2016 22:25:22 +0000 > > Will Deacon <will.deacon@arm.com> wrote: > > > > > On Wed, Nov 09, 2016 at 03:17:09PM -0700, Alex Williamson wrote: > > > > On Wed, 9 Nov 2016 20:31:45 +0000 > > > > Will Deacon <will.deacon@arm.com> wrote: > > > > > On Wed, Nov 09, 2016 at 08:23:03PM +0100, Christoffer Dall wrote: > > > > > > > > > > > > (I suppose it's technically possible to get around this issue by letting > > > > > > QEMU place RAM wherever it wants but tell the guest to never use a > > > > > > particular subset of its RAM for DMA, because that would conflict with > > > > > > the doorbell IOVA or be seen as p2p transactions. But I think we all > > > > > > probably agree that it's a disgusting idea.) > > > > > > > > > > Disgusting, yes, but Ben's idea of hotplugging on the host controller with > > > > > firmware tables describing the reserved regions is something that we could > > > > > do in the distant future. In the meantime, I don't think that VFIO should > > > > > explicitly reject overlapping mappings if userspace asks for them. > > > > > > > > I'm confused by the last sentence here, rejecting user mappings that > > > > overlap reserved ranges, such as MSI doorbell pages, is exactly how > > > > we'd reject hot-adding a device when we meet such a conflict. If we > > > > don't reject such a mapping, we're knowingly creating a situation that > > > > potentially leads to data loss. Minimally, QEMU would need to know > > > > about the reserved region, map around it through VFIO, and take > > > > responsibility (somehow) for making sure that region is never used for > > > > DMA. Thanks, > > > > > > Yes, but my point is that it should be up to QEMU to abort the hotplug, not > > > the host kernel, since there may be ways in which a guest can tolerate the > > > overlapping region (e.g. by avoiding that range of memory for DMA). > > > > The VFIO_IOMMU_MAP_DMA ioctl is a contract, the user ask to map a range > > of IOVAs to a range of virtual addresses for a given device. If VFIO > > cannot reasonably fulfill that contract, it must fail. It's up to QEMU > > how to manage the hotplug and what memory regions it asks VFIO to map > > for a device, but VFIO must reject mappings that it (or the SMMU by > > virtue of using the IOMMU API) know to overlap reserved ranges. So I > > still disagree with the referenced statement. Thanks, > > I think that's a pity. Not only does it mean that both QEMU and the kernel > have more work to do (the former has to carve up its mapping requests, > whilst the latter has to check that it is indeed doing this), but it also > precludes the use of hugepage mappings on the IOMMU because of reserved > regions. For example, a 4k hole someplace may mean we can't put down 1GB > table entries for the guest memory in the SMMU. > > All this seems to do is add complexity and decrease performance. For what? > QEMU has to go read the reserved regions from someplace anyway. It's also > the way that VFIO works *today* on arm64 wrt reserved regions, it just has > no way to identify those holes at present. Sure, that sucks, but how is the alternative even an option? The user asked to map something, we can't, if we allow that to happen now it's a bug. Put the MSI doorbells somewhere that this won't be an issue. If the platform has it fixed somewhere that this is an issue, don't use that platform. The correctness of the interface is more important than catering to a poorly designed system layout IMO. Thanks, Alex
[toc] | [prev] | [next] | [standalone]
| From | Auger Eric <eric.auger@redhat.com> |
|---|---|
| Date | 2016-11-10 01:20 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBLfY-6pf-15@gated-at.bofh.it> |
| In reply to | #1518591 |
Hi, On 10/11/2016 00:59, Alex Williamson wrote: > On Wed, 9 Nov 2016 23:38:50 +0000 > Will Deacon <will.deacon@arm.com> wrote: > >> On Wed, Nov 09, 2016 at 04:24:58PM -0700, Alex Williamson wrote: >>> On Wed, 9 Nov 2016 22:25:22 +0000 >>> Will Deacon <will.deacon@arm.com> wrote: >>> >>>> On Wed, Nov 09, 2016 at 03:17:09PM -0700, Alex Williamson wrote: >>>>> On Wed, 9 Nov 2016 20:31:45 +0000 >>>>> Will Deacon <will.deacon@arm.com> wrote: >>>>>> On Wed, Nov 09, 2016 at 08:23:03PM +0100, Christoffer Dall wrote: >>>>>>> >>>>>>> (I suppose it's technically possible to get around this issue by letting >>>>>>> QEMU place RAM wherever it wants but tell the guest to never use a >>>>>>> particular subset of its RAM for DMA, because that would conflict with >>>>>>> the doorbell IOVA or be seen as p2p transactions. But I think we all >>>>>>> probably agree that it's a disgusting idea.) >>>>>> >>>>>> Disgusting, yes, but Ben's idea of hotplugging on the host controller with >>>>>> firmware tables describing the reserved regions is something that we could >>>>>> do in the distant future. In the meantime, I don't think that VFIO should >>>>>> explicitly reject overlapping mappings if userspace asks for them. >>>>> >>>>> I'm confused by the last sentence here, rejecting user mappings that >>>>> overlap reserved ranges, such as MSI doorbell pages, is exactly how >>>>> we'd reject hot-adding a device when we meet such a conflict. If we >>>>> don't reject such a mapping, we're knowingly creating a situation that >>>>> potentially leads to data loss. Minimally, QEMU would need to know >>>>> about the reserved region, map around it through VFIO, and take >>>>> responsibility (somehow) for making sure that region is never used for >>>>> DMA. Thanks, >>>> >>>> Yes, but my point is that it should be up to QEMU to abort the hotplug, not >>>> the host kernel, since there may be ways in which a guest can tolerate the >>>> overlapping region (e.g. by avoiding that range of memory for DMA). >>> >>> The VFIO_IOMMU_MAP_DMA ioctl is a contract, the user ask to map a range >>> of IOVAs to a range of virtual addresses for a given device. If VFIO >>> cannot reasonably fulfill that contract, it must fail. It's up to QEMU >>> how to manage the hotplug and what memory regions it asks VFIO to map >>> for a device, but VFIO must reject mappings that it (or the SMMU by >>> virtue of using the IOMMU API) know to overlap reserved ranges. So I >>> still disagree with the referenced statement. Thanks, >> >> I think that's a pity. Not only does it mean that both QEMU and the kernel >> have more work to do (the former has to carve up its mapping requests, >> whilst the latter has to check that it is indeed doing this), but it also >> precludes the use of hugepage mappings on the IOMMU because of reserved >> regions. For example, a 4k hole someplace may mean we can't put down 1GB >> table entries for the guest memory in the SMMU. >> >> All this seems to do is add complexity and decrease performance. For what? >> QEMU has to go read the reserved regions from someplace anyway. It's also >> the way that VFIO works *today* on arm64 wrt reserved regions, it just has >> no way to identify those holes at present. > > Sure, that sucks, but how is the alternative even an option? The user > asked to map something, we can't, if we allow that to happen now it's a > bug. Put the MSI doorbells somewhere that this won't be an issue. If > the platform has it fixed somewhere that this is an issue, don't use > that platform. The correctness of the interface is more important than > catering to a poorly designed system layout IMO. Thanks, Besides above problematic, I started to prototype the sysfs API. A first issue I face is the reserved regions become global to the iommu instead of characterizing the iommu_domain, ie. the "reserved_regions" attribute file sits below an iommu instance (~ /sys/class/iommu/dmar0/intel-iommu/reserved_regions || /sys/class/iommu/arm-smmu0/arm-smmu/reserved_regions). MSI reserved window can be considered global to the IOMMU. However PCIe host bridge P2P regions rather are per iommu-domain. Do you confirm the attribute file should contain both global reserved regions and all per iommu_domain reserved regions? Thoughts? Thanks Eric > > Alex > > _______________________________________________ > linux-arm-kernel mailing list > linux-arm-kernel@lists.infradead.org > http://lists.infradead.org/mailman/listinfo/linux-arm-kernel >
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-10 02:00 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBLSF-6Eq-1@gated-at.bofh.it> |
| In reply to | #1518596 |
On Thu, 10 Nov 2016 01:14:42 +0100 Auger Eric <eric.auger@redhat.com> wrote: > Hi, > > On 10/11/2016 00:59, Alex Williamson wrote: > > On Wed, 9 Nov 2016 23:38:50 +0000 > > Will Deacon <will.deacon@arm.com> wrote: > > > >> On Wed, Nov 09, 2016 at 04:24:58PM -0700, Alex Williamson wrote: > >>> On Wed, 9 Nov 2016 22:25:22 +0000 > >>> Will Deacon <will.deacon@arm.com> wrote: > >>> > >>>> On Wed, Nov 09, 2016 at 03:17:09PM -0700, Alex Williamson wrote: > >>>>> On Wed, 9 Nov 2016 20:31:45 +0000 > >>>>> Will Deacon <will.deacon@arm.com> wrote: > >>>>>> On Wed, Nov 09, 2016 at 08:23:03PM +0100, Christoffer Dall wrote: > >>>>>>> > >>>>>>> (I suppose it's technically possible to get around this issue by letting > >>>>>>> QEMU place RAM wherever it wants but tell the guest to never use a > >>>>>>> particular subset of its RAM for DMA, because that would conflict with > >>>>>>> the doorbell IOVA or be seen as p2p transactions. But I think we all > >>>>>>> probably agree that it's a disgusting idea.) > >>>>>> > >>>>>> Disgusting, yes, but Ben's idea of hotplugging on the host controller with > >>>>>> firmware tables describing the reserved regions is something that we could > >>>>>> do in the distant future. In the meantime, I don't think that VFIO should > >>>>>> explicitly reject overlapping mappings if userspace asks for them. > >>>>> > >>>>> I'm confused by the last sentence here, rejecting user mappings that > >>>>> overlap reserved ranges, such as MSI doorbell pages, is exactly how > >>>>> we'd reject hot-adding a device when we meet such a conflict. If we > >>>>> don't reject such a mapping, we're knowingly creating a situation that > >>>>> potentially leads to data loss. Minimally, QEMU would need to know > >>>>> about the reserved region, map around it through VFIO, and take > >>>>> responsibility (somehow) for making sure that region is never used for > >>>>> DMA. Thanks, > >>>> > >>>> Yes, but my point is that it should be up to QEMU to abort the hotplug, not > >>>> the host kernel, since there may be ways in which a guest can tolerate the > >>>> overlapping region (e.g. by avoiding that range of memory for DMA). > >>> > >>> The VFIO_IOMMU_MAP_DMA ioctl is a contract, the user ask to map a range > >>> of IOVAs to a range of virtual addresses for a given device. If VFIO > >>> cannot reasonably fulfill that contract, it must fail. It's up to QEMU > >>> how to manage the hotplug and what memory regions it asks VFIO to map > >>> for a device, but VFIO must reject mappings that it (or the SMMU by > >>> virtue of using the IOMMU API) know to overlap reserved ranges. So I > >>> still disagree with the referenced statement. Thanks, > >> > >> I think that's a pity. Not only does it mean that both QEMU and the kernel > >> have more work to do (the former has to carve up its mapping requests, > >> whilst the latter has to check that it is indeed doing this), but it also > >> precludes the use of hugepage mappings on the IOMMU because of reserved > >> regions. For example, a 4k hole someplace may mean we can't put down 1GB > >> table entries for the guest memory in the SMMU. > >> > >> All this seems to do is add complexity and decrease performance. For what? > >> QEMU has to go read the reserved regions from someplace anyway. It's also > >> the way that VFIO works *today* on arm64 wrt reserved regions, it just has > >> no way to identify those holes at present. > > > > Sure, that sucks, but how is the alternative even an option? The user > > asked to map something, we can't, if we allow that to happen now it's a > > bug. Put the MSI doorbells somewhere that this won't be an issue. If > > the platform has it fixed somewhere that this is an issue, don't use > > that platform. The correctness of the interface is more important than > > catering to a poorly designed system layout IMO. Thanks, > > Besides above problematic, I started to prototype the sysfs API. A first > issue I face is the reserved regions become global to the iommu instead > of characterizing the iommu_domain, ie. the "reserved_regions" attribute > file sits below an iommu instance (~ > /sys/class/iommu/dmar0/intel-iommu/reserved_regions || > /sys/class/iommu/arm-smmu0/arm-smmu/reserved_regions). > > MSI reserved window can be considered global to the IOMMU. However PCIe > host bridge P2P regions rather are per iommu-domain. > > Do you confirm the attribute file should contain both global reserved > regions and all per iommu_domain reserved regions? > > Thoughts? I don't think we have any business describing IOVA addresses consumed by peer devices in an IOMMU sysfs file. If it's a separate device it should be exposed by examining the rest of the topology. Regions consumed by PCI endpoints and interconnects are already exposed in sysfs. In fact, is this perhaps a more accurate model for these MSI controllers too? Perhaps they should be exposed in the bus topology somewhere as consuming the IOVA range. If DMA to an IOVA is consumed by an intermediate device before it hits the IOMMU vs not being translated as specified by the user at the IOMMU, I'm less inclined to call that something VFIO should reject. However, instantiating a VM with holes to account for every potential peer device seems like it borders on insanity. Thanks, Alex
[toc] | [prev] | [next] | [standalone]
| From | Will Deacon <will.deacon@arm.com> |
|---|---|
| Date | 2016-11-10 03:10 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBMYp-7B7-3@gated-at.bofh.it> |
| In reply to | #1518606 |
On Wed, Nov 09, 2016 at 05:55:17PM -0700, Alex Williamson wrote:
> On Thu, 10 Nov 2016 01:14:42 +0100
> Auger Eric <eric.auger@redhat.com> wrote:
> > On 10/11/2016 00:59, Alex Williamson wrote:
> > > On Wed, 9 Nov 2016 23:38:50 +0000
> > > Will Deacon <will.deacon@arm.com> wrote:
> > >> On Wed, Nov 09, 2016 at 04:24:58PM -0700, Alex Williamson wrote:
> > >>> The VFIO_IOMMU_MAP_DMA ioctl is a contract, the user ask to map a range
> > >>> of IOVAs to a range of virtual addresses for a given device. If VFIO
> > >>> cannot reasonably fulfill that contract, it must fail. It's up to QEMU
> > >>> how to manage the hotplug and what memory regions it asks VFIO to map
> > >>> for a device, but VFIO must reject mappings that it (or the SMMU by
> > >>> virtue of using the IOMMU API) know to overlap reserved ranges. So I
> > >>> still disagree with the referenced statement. Thanks,
> > >>
> > >> I think that's a pity. Not only does it mean that both QEMU and the kernel
> > >> have more work to do (the former has to carve up its mapping requests,
> > >> whilst the latter has to check that it is indeed doing this), but it also
> > >> precludes the use of hugepage mappings on the IOMMU because of reserved
> > >> regions. For example, a 4k hole someplace may mean we can't put down 1GB
> > >> table entries for the guest memory in the SMMU.
> > >>
> > >> All this seems to do is add complexity and decrease performance. For what?
> > >> QEMU has to go read the reserved regions from someplace anyway. It's also
> > >> the way that VFIO works *today* on arm64 wrt reserved regions, it just has
> > >> no way to identify those holes at present.
> > >
> > > Sure, that sucks, but how is the alternative even an option? The user
> > > asked to map something, we can't, if we allow that to happen now it's a
> > > bug. Put the MSI doorbells somewhere that this won't be an issue. If
> > > the platform has it fixed somewhere that this is an issue, don't use
> > > that platform. The correctness of the interface is more important than
> > > catering to a poorly designed system layout IMO. Thanks,
> >
> > Besides above problematic, I started to prototype the sysfs API. A first
> > issue I face is the reserved regions become global to the iommu instead
> > of characterizing the iommu_domain, ie. the "reserved_regions" attribute
> > file sits below an iommu instance (~
> > /sys/class/iommu/dmar0/intel-iommu/reserved_regions ||
> > /sys/class/iommu/arm-smmu0/arm-smmu/reserved_regions).
> >
> > MSI reserved window can be considered global to the IOMMU. However PCIe
> > host bridge P2P regions rather are per iommu-domain.
I don't think we can treat them as per-domain, given that we want to
enumerate this stuff before we've decided to do a hotplug (and therefore
don't have a domain).
> >
> > Do you confirm the attribute file should contain both global reserved
> > regions and all per iommu_domain reserved regions?
> >
> > Thoughts?
>
> I don't think we have any business describing IOVA addresses consumed
> by peer devices in an IOMMU sysfs file. If it's a separate device it
> should be exposed by examining the rest of the topology. Regions
> consumed by PCI endpoints and interconnects are already exposed in
> sysfs. In fact, is this perhaps a more accurate model for these MSI
> controllers too? Perhaps they should be exposed in the bus topology
> somewhere as consuming the IOVA range. If DMA to an IOVA is consumed
> by an intermediate device before it hits the IOMMU vs not being
> translated as specified by the user at the IOMMU, I'm less inclined to
> call that something VFIO should reject.
Oh, so perhaps we've been talking past each other. In all of these cases,
the SMMU can translate the access if it makes it that far. The issue is
that not all accesses do make it that far, because they may be "consumed"
by another device, such as an MSI doorbell or another endpoint. In other
words, I don't envisage a scenario where e.g. some address range just
bypasses the SMMU straight to memory. I realise now that that's not clear
from the slides I presented.
> However, instantiating a VM
> with holes to account for every potential peer device seems like it
> borders on insanity. Thanks,
Ok, so rather than having a list of reserved regions under the iommu node,
you're proposing that each region is attributed to the device that "owns"
(consumes) it? I think that can work, but we need to make sure that:
(a) The topology is walkable from userspace (where do you start?)
(b) It also works for platform (non-PCI) devices, that lack much in the
way of bus hierarchy
(c) It doesn't require Linux to have a driver bound to a device in order
for the ranges consumed by that device to be advertised (again,
more of an issue for non-PCI).
How is this currently advertised for PCI? I'd really like to use the same
scheme irrespective of the bus type.
Will
[toc] | [prev] | [next] | [standalone]
| From | Auger Eric <eric.auger@redhat.com> |
|---|---|
| Date | 2016-11-10 12:20 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sBVyG-5qG-33@gated-at.bofh.it> |
| In reply to | #1518645 |
Hi Will, Alex, On 10/11/2016 03:01, Will Deacon wrote: > On Wed, Nov 09, 2016 at 05:55:17PM -0700, Alex Williamson wrote: >> On Thu, 10 Nov 2016 01:14:42 +0100 >> Auger Eric <eric.auger@redhat.com> wrote: >>> On 10/11/2016 00:59, Alex Williamson wrote: >>>> On Wed, 9 Nov 2016 23:38:50 +0000 >>>> Will Deacon <will.deacon@arm.com> wrote: >>>>> On Wed, Nov 09, 2016 at 04:24:58PM -0700, Alex Williamson wrote: >>>>>> The VFIO_IOMMU_MAP_DMA ioctl is a contract, the user ask to map a range >>>>>> of IOVAs to a range of virtual addresses for a given device. If VFIO >>>>>> cannot reasonably fulfill that contract, it must fail. It's up to QEMU >>>>>> how to manage the hotplug and what memory regions it asks VFIO to map >>>>>> for a device, but VFIO must reject mappings that it (or the SMMU by >>>>>> virtue of using the IOMMU API) know to overlap reserved ranges. So I >>>>>> still disagree with the referenced statement. Thanks, >>>>> >>>>> I think that's a pity. Not only does it mean that both QEMU and the kernel >>>>> have more work to do (the former has to carve up its mapping requests, >>>>> whilst the latter has to check that it is indeed doing this), but it also >>>>> precludes the use of hugepage mappings on the IOMMU because of reserved >>>>> regions. For example, a 4k hole someplace may mean we can't put down 1GB >>>>> table entries for the guest memory in the SMMU. >>>>> >>>>> All this seems to do is add complexity and decrease performance. For what? >>>>> QEMU has to go read the reserved regions from someplace anyway. It's also >>>>> the way that VFIO works *today* on arm64 wrt reserved regions, it just has >>>>> no way to identify those holes at present. >>>> >>>> Sure, that sucks, but how is the alternative even an option? The user >>>> asked to map something, we can't, if we allow that to happen now it's a >>>> bug. Put the MSI doorbells somewhere that this won't be an issue. If >>>> the platform has it fixed somewhere that this is an issue, don't use >>>> that platform. The correctness of the interface is more important than >>>> catering to a poorly designed system layout IMO. Thanks, >>> >>> Besides above problematic, I started to prototype the sysfs API. A first >>> issue I face is the reserved regions become global to the iommu instead >>> of characterizing the iommu_domain, ie. the "reserved_regions" attribute >>> file sits below an iommu instance (~ >>> /sys/class/iommu/dmar0/intel-iommu/reserved_regions || >>> /sys/class/iommu/arm-smmu0/arm-smmu/reserved_regions). >>> >>> MSI reserved window can be considered global to the IOMMU. However PCIe >>> host bridge P2P regions rather are per iommu-domain. > > I don't think we can treat them as per-domain, given that we want to > enumerate this stuff before we've decided to do a hotplug (and therefore > don't have a domain). That's the issue indeed. We need to wait for the PCIe device to be connected to the iommu. Only on the VFIO SET_IOMMU we get the comprehensive list of P2P regions that can impact IOVA mapping for this iommu. This removes any advantage of sysfs API over previous VFIO capability chain API for P2P IOVA range enumeration at early stage. > >>> >>> Do you confirm the attribute file should contain both global reserved >>> regions and all per iommu_domain reserved regions? >>> >>> Thoughts? >> >> I don't think we have any business describing IOVA addresses consumed >> by peer devices in an IOMMU sysfs file. If it's a separate device it >> should be exposed by examining the rest of the topology. Regions >> consumed by PCI endpoints and interconnects are already exposed in >> sysfs. In fact, is this perhaps a more accurate model for these MSI >> controllers too? Perhaps they should be exposed in the bus topology >> somewhere as consuming the IOVA range. Currently on x86 the P2P regions are not checked when allowing passthrough. Aren't we more papist that the pope? As Don mentioned, shouldn't we simply consider that a platform that does not support proper ACS is not candidate for safe passthrough, like Juno? At least we can state the feature also is missing on x86 and it would be nice to report the risk to the userspace and urge him to opt-in. To me taking into account those P2P still is controversial and induce the bulk of the complexity. Considering the migration use case discussed at LPC while only handling the MSI problem looks much easier. host can choose an MSI base that is QEMU mach-virt friendly, ie. non RAM region. Problem is to satisfy all potential uses though. When migrating, mach-virt still is being used so there should not be any collision. Am I missing some migration weird use cases here? Of course if we take into consideration new host PCIe P2P regions this becomes completely different. We still have the good old v14 where the user space chose where MSI IOVA's are put without any risk of collision ;-) >> If DMA to an IOVA is consumed >> by an intermediate device before it hits the IOMMU vs not being >> translated as specified by the user at the IOMMU, I'm less inclined to >> call that something VFIO should reject. > > Oh, so perhaps we've been talking past each other. In all of these cases, > the SMMU can translate the access if it makes it that far. The issue is > that not all accesses do make it that far, because they may be "consumed" > by another device, such as an MSI doorbell or another endpoint. In other > words, I don't envisage a scenario where e.g. some address range just > bypasses the SMMU straight to memory. I realise now that that's not clear > from the slides I presented. > >> However, instantiating a VM >> with holes to account for every potential peer device seems like it >> borders on insanity. Thanks, > > Ok, so rather than having a list of reserved regions under the iommu node, > you're proposing that each region is attributed to the device that "owns" > (consumes) it? I think that can work, but we need to make sure that: > > (a) The topology is walkable from userspace (where do you start?) > > (b) It also works for platform (non-PCI) devices, that lack much in the > way of bus hierarchy > > (c) It doesn't require Linux to have a driver bound to a device in order > for the ranges consumed by that device to be advertised (again, > more of an issue for non-PCI). Looks a long way ... Thanks Eric > > How is this currently advertised for PCI? I'd really like to use the same > scheme irrespective of the bus type. > > Will > > _______________________________________________ > linux-arm-kernel mailing list > linux-arm-kernel@lists.infradead.org > http://lists.infradead.org/mailman/listinfo/linux-arm-kernel >
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-10 18:50 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sC1E5-QH-3@gated-at.bofh.it> |
| In reply to | #1518876 |
On Thu, 10 Nov 2016 12:14:40 +0100 Auger Eric <eric.auger@redhat.com> wrote: > Hi Will, Alex, > > On 10/11/2016 03:01, Will Deacon wrote: > > On Wed, Nov 09, 2016 at 05:55:17PM -0700, Alex Williamson wrote: > >> On Thu, 10 Nov 2016 01:14:42 +0100 > >> Auger Eric <eric.auger@redhat.com> wrote: > >>> On 10/11/2016 00:59, Alex Williamson wrote: > >>>> On Wed, 9 Nov 2016 23:38:50 +0000 > >>>> Will Deacon <will.deacon@arm.com> wrote: > >>>>> On Wed, Nov 09, 2016 at 04:24:58PM -0700, Alex Williamson wrote: > >>>>>> The VFIO_IOMMU_MAP_DMA ioctl is a contract, the user ask to map a range > >>>>>> of IOVAs to a range of virtual addresses for a given device. If VFIO > >>>>>> cannot reasonably fulfill that contract, it must fail. It's up to QEMU > >>>>>> how to manage the hotplug and what memory regions it asks VFIO to map > >>>>>> for a device, but VFIO must reject mappings that it (or the SMMU by > >>>>>> virtue of using the IOMMU API) know to overlap reserved ranges. So I > >>>>>> still disagree with the referenced statement. Thanks, > >>>>> > >>>>> I think that's a pity. Not only does it mean that both QEMU and the kernel > >>>>> have more work to do (the former has to carve up its mapping requests, > >>>>> whilst the latter has to check that it is indeed doing this), but it also > >>>>> precludes the use of hugepage mappings on the IOMMU because of reserved > >>>>> regions. For example, a 4k hole someplace may mean we can't put down 1GB > >>>>> table entries for the guest memory in the SMMU. > >>>>> > >>>>> All this seems to do is add complexity and decrease performance. For what? > >>>>> QEMU has to go read the reserved regions from someplace anyway. It's also > >>>>> the way that VFIO works *today* on arm64 wrt reserved regions, it just has > >>>>> no way to identify those holes at present. > >>>> > >>>> Sure, that sucks, but how is the alternative even an option? The user > >>>> asked to map something, we can't, if we allow that to happen now it's a > >>>> bug. Put the MSI doorbells somewhere that this won't be an issue. If > >>>> the platform has it fixed somewhere that this is an issue, don't use > >>>> that platform. The correctness of the interface is more important than > >>>> catering to a poorly designed system layout IMO. Thanks, > >>> > >>> Besides above problematic, I started to prototype the sysfs API. A first > >>> issue I face is the reserved regions become global to the iommu instead > >>> of characterizing the iommu_domain, ie. the "reserved_regions" attribute > >>> file sits below an iommu instance (~ > >>> /sys/class/iommu/dmar0/intel-iommu/reserved_regions || > >>> /sys/class/iommu/arm-smmu0/arm-smmu/reserved_regions). > >>> > >>> MSI reserved window can be considered global to the IOMMU. However PCIe > >>> host bridge P2P regions rather are per iommu-domain. > > > > I don't think we can treat them as per-domain, given that we want to > > enumerate this stuff before we've decided to do a hotplug (and therefore > > don't have a domain). > That's the issue indeed. We need to wait for the PCIe device to be > connected to the iommu. Only on the VFIO SET_IOMMU we get the > comprehensive list of P2P regions that can impact IOVA mapping for this > iommu. This removes any advantage of sysfs API over previous VFIO > capability chain API for P2P IOVA range enumeration at early stage. For use through vfio we know that an iommu_domain is minimally composed of an iommu_group and we can find all the p2p resources of that group referencing /proc/iomem, at least for PCI-based groups. This is the part that I don't think any sort of iommu sysfs attributes should be duplicating. > >>> Do you confirm the attribute file should contain both global reserved > >>> regions and all per iommu_domain reserved regions? > >>> > >>> Thoughts? > >> > >> I don't think we have any business describing IOVA addresses consumed > >> by peer devices in an IOMMU sysfs file. If it's a separate device it > >> should be exposed by examining the rest of the topology. Regions > >> consumed by PCI endpoints and interconnects are already exposed in > >> sysfs. In fact, is this perhaps a more accurate model for these MSI > >> controllers too? Perhaps they should be exposed in the bus topology > >> somewhere as consuming the IOVA range. > Currently on x86 the P2P regions are not checked when allowing > passthrough. Aren't we more papist that the pope? As Don mentioned, > shouldn't we simply consider that a platform that does not support > proper ACS is not candidate for safe passthrough, like Juno? There are two sides here, there's the kernel side vfio and there's how QEMU makes use of vfio. On the kernel side, we create iommu groups as the set of devices we consider isolated, that doesn't necessarily mean that there isn't p2p within the group, in fact that potential often determines the composition of the group. It's the user's problem how to deal with that potential. When I talk about the contract with userspace, I consider that to be at the iommu mapping, ie. for transactions that actually make it to the iommu. In the case of x86, we know that DMA mappings overlapping the MSI doorbells won't be translated correctly, it's not a valid mapping for that range, and therefore the iommu driver backing the IOMMU API should describe that reserved range and reject mappings to it. For devices downstream of the IOMMU, whether they be p2p or MSI controllers consuming fixed IOVA space, I consider these to be problems beyond the scope of the IOMMU API, and maybe that's where we've been going wrong all along. Users like QEMU can currently discover potential p2p conflicts by looking at the composition of an iommu group and taking into account the host PCI resources of each device. We don't currently do this, though we probably should. The reason we typically don't run into problems with this is that (again) x86 has a fairly standard memory layout. Potential p2p regions are typically in an MMIO hole in the host that sufficiently matches an MMIO hole in the guest. So we don't often have VM RAM, which could be a DMA target, matching those p2p addresses. We also hope that any serious device assignment users have singleton iommu groups, ie. the IO subsystem is designed to support proper, fine grained isolation. > At least we can state the feature also is missing on x86 and it would be > nice to report the risk to the userspace and urge him to opt-in. Sure, but the information is already there, it's "just" a matter of QEMU taking it into account, which has some implications that VMs with any potential of doing device assignment need to be instantiated with address maps compatible with the host system, which is not an easy feat for something often considered the ugly step-child of virtualization. > To me taking into account those P2P still is controversial and induce > the bulk of the complexity. Considering the migration use case discussed > at LPC while only handling the MSI problem looks much easier. > host can choose an MSI base that is QEMU mach-virt friendly, ie. non RAM > region. Problem is to satisfy all potential uses though. When migrating, > mach-virt still is being used so there should not be any collision. Am I > missing some migration weird use cases here? Of course if we take into > consideration new host PCIe P2P regions this becomes completely different. Yep, x86 having a standard MSI range is a nice happenstance, so long as we're running an x86 VM, we don't worry about that being a DMA target. Running non-x86 VMs on x86 hosts hits this problem, but is several orders of magnitude lower priority. > We still have the good old v14 where the user space chose where MSI > IOVA's are put without any risk of collision ;-) > > >> If DMA to an IOVA is consumed > >> by an intermediate device before it hits the IOMMU vs not being > >> translated as specified by the user at the IOMMU, I'm less inclined to > >> call that something VFIO should reject. > > > > Oh, so perhaps we've been talking past each other. In all of these cases, > > the SMMU can translate the access if it makes it that far. The issue is > > that not all accesses do make it that far, because they may be "consumed" > > by another device, such as an MSI doorbell or another endpoint. In other > > words, I don't envisage a scenario where e.g. some address range just > > bypasses the SMMU straight to memory. I realise now that that's not clear > > from the slides I presented. As above, so long as a transaction that does make it to the iommu is translated as prescribed by the user, I have no basis for rejecting a user requested translation. Downstream MSI controllers consuming IOVA space is no different than the existing p2p problem that vfio considers a userspace issue. > >> However, instantiating a VM > >> with holes to account for every potential peer device seems like it > >> borders on insanity. Thanks, > > > > Ok, so rather than having a list of reserved regions under the iommu node, > > you're proposing that each region is attributed to the device that "owns" > > (consumes) it? I think that can work, but we need to make sure that: > > > > (a) The topology is walkable from userspace (where do you start?) For PCI devices userspace can examine the topology of the iommu group and exclude MMIO ranges of peer devices based on the BARs, which are exposed in various places, pci-sysfs as well as /proc/iomem. For non-PCI or MSI controllers... ??? > > (b) It also works for platform (non-PCI) devices, that lack much in the > > way of bus hierarchy No idea here, without a userspace visible topology the user is in the dark as to what devices potentially sit between them and the iommu. > > (c) It doesn't require Linux to have a driver bound to a device in order > > for the ranges consumed by that device to be advertised (again, > > more of an issue for non-PCI). Right, PCI has this problem solved, be more like PCI ;) > > How is this currently advertised for PCI? I'd really like to use the same > > scheme irrespective of the bus type. For all devices within an IOMMU group, /proc/iomem might be the solution, but I don't know how the MSI controller works. If the MSI controller belongs to the group, then maybe it's a matter of creating a struct device for it that consumes resources and making it show up in both the iommu group and /proc/iomem. An MSI controller shared between groups, which already sounds like a bit of a violation of iommu groups, would need to be discoverable some other way. Thanks, Alex
[toc] | [prev] | [next] | [standalone]
| From | Joerg Roedel <joro@8bytes.org> |
|---|---|
| Date | 2016-11-11 12:20 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sCi2d-3Bo-1@gated-at.bofh.it> |
| In reply to | #1519223 |
On Thu, Nov 10, 2016 at 10:46:01AM -0700, Alex Williamson wrote: > In the case of x86, we know that DMA mappings overlapping the MSI > doorbells won't be translated correctly, it's not a valid mapping for > that range, and therefore the iommu driver backing the IOMMU API > should describe that reserved range and reject mappings to it. The drivers actually allow mappings to the MSI region via the IOMMU-API, and I think it should stay this way also for other reserved ranges. Address space management is done by the IOMMU-API user already (and has to be done there nowadays), be it a DMA-API implementation which just reserves these regions in its address space allocator or be it VFIO with QEMU, which don't map RAM there anyway. So there is no point of checking this again in the IOMMU drivers and we can keep that out of the mapping/unmapping fast-path. > For PCI devices userspace can examine the topology of the iommu group > and exclude MMIO ranges of peer devices based on the BARs, which are > exposed in various places, pci-sysfs as well as /proc/iomem. For > non-PCI or MSI controllers... ??? Right, the hardware resources can be examined. But maybe this can be extended to also cover RMRR ranges? Then we would be able to assign devices with RMRR mappings to guests. Joerg
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-11 17:00 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sCmpc-6c5-27@gated-at.bofh.it> |
| In reply to | #1519698 |
On Fri, 11 Nov 2016 12:19:44 +0100 Joerg Roedel <joro@8bytes.org> wrote: > On Thu, Nov 10, 2016 at 10:46:01AM -0700, Alex Williamson wrote: > > In the case of x86, we know that DMA mappings overlapping the MSI > > doorbells won't be translated correctly, it's not a valid mapping for > > that range, and therefore the iommu driver backing the IOMMU API > > should describe that reserved range and reject mappings to it. > > The drivers actually allow mappings to the MSI region via the IOMMU-API, > and I think it should stay this way also for other reserved ranges. > Address space management is done by the IOMMU-API user already (and has > to be done there nowadays), be it a DMA-API implementation which just > reserves these regions in its address space allocator or be it VFIO with > QEMU, which don't map RAM there anyway. So there is no point of checking > this again in the IOMMU drivers and we can keep that out of the > mapping/unmapping fast-path. It's really just a happenstance that we don't map RAM over the x86 MSI range though. That property really can't be guaranteed once we mix architectures, such as running an aarch64 VM on x86 host via TCG. AIUI, the MSI range is actually handled differently than other DMA ranges, so a iommu_map() overlapping a range that the iommu cannot map should fail just like an attempt to map beyond the address width of the iommu. > > For PCI devices userspace can examine the topology of the iommu group > > and exclude MMIO ranges of peer devices based on the BARs, which are > > exposed in various places, pci-sysfs as well as /proc/iomem. For > > non-PCI or MSI controllers... ??? > > Right, the hardware resources can be examined. But maybe this can be > extended to also cover RMRR ranges? Then we would be able to assign > devices with RMRR mappings to guests. RMRRs are special in a different way, the VT-d spec requires that the OS honor RMRRs, the user has no responsibility (and currently no visibility) to make that same arrangement. In order to potentially protect the physical host platform, the iommu drivers should prevent a user from remapping RMRRS. Maybe there needs to be a different interface used by untrusted users vs in-kernel drivers, but I think the kernel really needs to be defensive in the case of user mappings, which is where the IOMMU API is rooted. Thanks, Alex
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-11 17:10 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sCmyR-6y4-5@gated-at.bofh.it> |
| In reply to | #1519874 |
On Fri, 11 Nov 2016 08:50:56 -0700 Alex Williamson <alex.williamson@redhat.com> wrote: > On Fri, 11 Nov 2016 12:19:44 +0100 > Joerg Roedel <joro@8bytes.org> wrote: > > > On Thu, Nov 10, 2016 at 10:46:01AM -0700, Alex Williamson wrote: > > > In the case of x86, we know that DMA mappings overlapping the MSI > > > doorbells won't be translated correctly, it's not a valid mapping for > > > that range, and therefore the iommu driver backing the IOMMU API > > > should describe that reserved range and reject mappings to it. > > > > The drivers actually allow mappings to the MSI region via the IOMMU-API, > > and I think it should stay this way also for other reserved ranges. > > Address space management is done by the IOMMU-API user already (and has > > to be done there nowadays), be it a DMA-API implementation which just > > reserves these regions in its address space allocator or be it VFIO with > > QEMU, which don't map RAM there anyway. So there is no point of checking > > this again in the IOMMU drivers and we can keep that out of the > > mapping/unmapping fast-path. > > It's really just a happenstance that we don't map RAM over the x86 MSI > range though. That property really can't be guaranteed once we mix > architectures, such as running an aarch64 VM on x86 host via TCG. > AIUI, the MSI range is actually handled differently than other DMA > ranges, so a iommu_map() overlapping a range that the iommu cannot map > should fail just like an attempt to map beyond the address width of the > iommu. (clarification, this is x86 specific, the MSI controller - interrupt remapper - is embedded in the iommu AIUI, so the iommu is actually not able to provide DMA translation for this range. In architectures where the MSI controller is separate from the iommu, I agree that the iommu has no responsibility to fault mapping of iova ranges in the shadow of an external MSI controller) > > > For PCI devices userspace can examine the topology of the iommu group > > > and exclude MMIO ranges of peer devices based on the BARs, which are > > > exposed in various places, pci-sysfs as well as /proc/iomem. For > > > non-PCI or MSI controllers... ??? > > > > Right, the hardware resources can be examined. But maybe this can be > > extended to also cover RMRR ranges? Then we would be able to assign > > devices with RMRR mappings to guests. > > RMRRs are special in a different way, the VT-d spec requires that the > OS honor RMRRs, the user has no responsibility (and currently no > visibility) to make that same arrangement. In order to potentially > protect the physical host platform, the iommu drivers should prevent a > user from remapping RMRRS. Maybe there needs to be a different > interface used by untrusted users vs in-kernel drivers, but I think the > kernel really needs to be defensive in the case of user mappings, which > is where the IOMMU API is rooted. Thanks, > > Alex > _______________________________________________ > iommu mailing list > iommu@lists.linux-foundation.org > https://lists.linuxfoundation.org/mailman/listinfo/iommu
[toc] | [prev] | [next] | [standalone]
| From | Joerg Roedel <joro@8bytes.org> |
|---|---|
| Date | 2016-11-14 16:20 +0100 |
| Subject | Re: Summary of LPC guest MSI discussion in Santa Fe |
| Message-ID | <sDrd7-eg-17@gated-at.bofh.it> |
| In reply to | #1519879 |
On Fri, Nov 11, 2016 at 09:05:43AM -0700, Alex Williamson wrote: > On Fri, 11 Nov 2016 08:50:56 -0700 > Alex Williamson <alex.williamson@redhat.com> wrote: > > > > It's really just a happenstance that we don't map RAM over the x86 MSI > > range though. That property really can't be guaranteed once we mix > > architectures, such as running an aarch64 VM on x86 host via TCG. > > AIUI, the MSI range is actually handled differently than other DMA > > ranges, so a iommu_map() overlapping a range that the iommu cannot map > > should fail just like an attempt to map beyond the address width of the > > iommu. > > (clarification, this is x86 specific, the MSI controller - interrupt > remapper - is embedded in the iommu AIUI, so the iommu is actually not > able to provide DMA translation for this range. Right, on x86 the MSI range can be covered by page-tables, but those are ignored by the IOMMU hardware. But what I am trying to say is, that checking for these ranges happens already on a higher level (in the dma-api implementations by marking these regions as allocted) so that there is no need to check for them again in the iommu_map/unmap path. Joerg
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web