Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1272447 > unrolled thread

Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel

Started byAlex Williamson <alex.williamson@redhat.com>
First post2015-11-18 19:20 +0100
Last post2015-11-24 15:20 +0100
Articles 12 on this page of 32 — 9 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-18 19:20 +0100
    RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-19 05:10 +0100
      Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-19 08:30 +0100
        Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Stefano Stabellini <stefano.stabellini@eu.citrix.com> - 2015-11-19 16:40 +0100
          Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-19 17:00 +0100
            Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-20 04:00 +0100
              Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-20 05:30 +0100
                Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-20 07:00 +0100
                  RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-20 07:10 +0100
                  Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-20 17:50 +0100
                    Re: [Qemu-devel] [Intel-gfx] [Announcement] 2015-Q3 release of XenGT  - a Mediated Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-23 06:00 +0100
          Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Paolo Bonzini <pbonzini@redhat.com> - 2015-11-19 17:00 +0100
            Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Stefano Stabellini <stefano.stabellini@eu.citrix.com> - 2015-11-19 17:20 +0100
      Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Gerd Hoffmann <kraxel@redhat.com> - 2015-11-19 09:50 +0100
        Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Paolo Bonzini <pbonzini@redhat.com> - 2015-11-19 12:10 +0100
          Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-20 03:50 +0100
        RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-20 07:20 +0100
          Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Gerd Hoffmann <kraxel@redhat.com> - 2015-11-20 09:30 +0100
            RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-20 09:40 +0100
              Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Zhiyuan Lv <zhiyuan.lv@intel.com> - 2015-11-20 10:00 +0100
      Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-19 21:10 +0100
        RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-20 08:20 +0100
          Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-20 18:10 +0100
        RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-20 09:20 +0100
          Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-20 18:30 +0100
            Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-23 06:10 +0100
        Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Daniel Vetter <daniel@ffwll.ch> - 2015-11-24 12:20 +0100
          Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Chris Wilson <chris@chris-wilson.co.uk> - 2015-11-24 13:00 +0100
          Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Gerd Hoffmann <kraxel@redhat.com> - 2015-11-24 13:40 +0100
            Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Daniel Vetter <daniel@ffwll.ch> - 2015-11-24 14:40 +0100
              Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a  Mediated Graphics Passthrough Solution from Intel Gerd Hoffmann <kraxel@redhat.com> - 2015-11-24 15:20 +0100
                Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated  Graphics Passthrough Solution from Intel Daniel Vetter <daniel@ffwll.ch> - 2015-11-24 15:20 +0100

Page 2 of 2 — ← Prev page 1 [2]


#1273447

FromAlex Williamson <alex.williamson@redhat.com>
Date2015-11-19 21:10 +0100
Message-ID<qwDGN-5Ex-15@gated-at.bofh.it>
In reply to#1272827
Hi Kevin,

On Thu, 2015-11-19 at 04:06 +0000, Tian, Kevin wrote:
> > From: Alex Williamson [mailto:alex.williamson@redhat.com]
> > Sent: Thursday, November 19, 2015 2:12 AM
> > 
> > [cc +qemu-devel, +paolo, +gerd]
> > 
> > On Tue, 2015-10-27 at 17:25 +0800, Jike Song wrote:
> > > Hi all,
> > >
> > > We are pleased to announce another update of Intel GVT-g for Xen.
> > >
> > > Intel GVT-g is a full GPU virtualization solution with mediated
> > > pass-through, starting from 4th generation Intel Core(TM) processors
> > > with Intel Graphics processors. A virtual GPU instance is maintained
> > > for each VM, with part of performance critical resources directly
> > > assigned. The capability of running native graphics driver inside a
> > > VM, without hypervisor intervention in performance critical paths,
> > > achieves a good balance among performance, feature, and sharing
> > > capability. Xen is currently supported on Intel Processor Graphics
> > > (a.k.a. XenGT); and the core logic can be easily ported to other
> > > hypervisors.
> > >
> > >
> > > Repositories
> > >
> > >      Kernel: https://github.com/01org/igvtg-kernel (2015q3-3.18.0 branch)
> > >      Xen: https://github.com/01org/igvtg-xen (2015q3-4.5 branch)
> > >      Qemu: https://github.com/01org/igvtg-qemu (xengt_public2015q3 branch)
> > >
> > >
> > > This update consists of:
> > >
> > >      - XenGT is now merged with KVMGT in unified repositories(kernel and qemu), but
> > currently
> > >        different branches for qemu.  XenGT and KVMGT share same iGVT-g core logic.
> > 
> > Hi!
> > 
> > At redhat we've been thinking about how to support vGPUs from multiple
> > vendors in a common way within QEMU.  We want to enable code sharing
> > between vendors and give new vendors an easy path to add their own
> > support.  We also have the complication that not all vGPU vendors are as
> > open source friendly as Intel, so being able to abstract the device
> > mediation and access outside of QEMU is a big advantage.
> > 
> > The proposal I'd like to make is that a vGPU, whether it is from Intel
> > or another vendor, is predominantly a PCI(e) device.  We have an
> > interface in QEMU already for exposing arbitrary PCI devices, vfio-pci.
> > Currently vfio-pci uses the VFIO API to interact with "physical" devices
> > and system IOMMUs.  I highlight /physical/ there because some of these
> > physical devices are SR-IOV VFs, which is somewhat of a fuzzy concept,
> > somewhere between fixed hardware and a virtual device implemented in
> > software.  That software just happens to be running on the physical
> > endpoint.
> 
> Agree. 
> 
> One clarification for rest discussion, is that we're talking about GVT-g vGPU 
> here which is a pure software GPU virtualization technique. GVT-d (note 
> some use in the text) refers to passing through the whole GPU or a specific 
> VF. GVT-d already falls into existing VFIO APIs nicely (though some on-going
> effort to remove Intel specific platform stickness from gfx driver). :-)
> 
> > 
> > vGPUs are similar, with the virtual device created at a different point,
> > host software.  They also rely on different IOMMU constructs, making use
> > of the MMU capabilities of the GPU (GTTs and such), but really having
> > similar requirements.
> 
> One important difference between system IOMMU and GPU-MMU here.
> System IOMMU is very much about translation from a DMA target
> (IOVA on native, or GPA in virtualization case) to HPA. However GPU
> internal MMUs is to translate from Graphics Memory Address (GMA)
> to DMA target (HPA if system IOMMU is disabled, or IOVA/GPA if system
> IOMMU is enabled). GMA is an internal addr space within GPU, not 
> exposed to Qemu and fully managed by GVT-g device model. Since it's 
> not a standard PCI defined resource, we don't need abstract this capability
> in VFIO interface.
> 
> > 
> > The proposal is therefore that GPU vendors can expose vGPUs to
> > userspace, and thus to QEMU, using the VFIO API.  For instance, vfio
> > supports modular bus drivers and IOMMU drivers.  An intel-vfio-gvt-d
> > module (or extension of i915) can register as a vfio bus driver, create
> > a struct device per vGPU, create an IOMMU group for that device, and
> > register that device with the vfio-core.  Since we don't rely on the
> > system IOMMU for GVT-d vGPU assignment, another vGPU vendor driver (or
> > extension of the same module) can register a "type1" compliant IOMMU
> > driver into vfio-core.  From the perspective of QEMU then, all of the
> > existing vfio-pci code is re-used, QEMU remains largely unaware of any
> > specifics of the vGPU being assigned, and the only necessary change so
> > far is how QEMU traverses sysfs to find the device and thus the IOMMU
> > group leading to the vfio group.
> 
> GVT-g requires to pin guest memory and query GPA->HPA information,
> upon which shadow GTTs will be updated accordingly from (GMA->GPA)
> to (GMA->HPA). So yes, here a dummy or simple "type1" compliant IOMMU 
> can be introduced just for this requirement.
> 
> However there's one tricky point which I'm not sure whether overall
> VFIO concept will be violated. GVT-g doesn't require system IOMMU
> to function, however host system may enable system IOMMU just for 
> hardening purpose. This means two-level translations existing (GMA->
> IOVA->HPA), so the dummy IOMMU driver has to request system IOMMU 
> driver to allocate IOVA for VMs and then setup IOVA->HPA mapping
> in IOMMU page table. In this case, multiple VM's translations are 
> multiplexed in one IOMMU page table.
> 
> We might need create some group/sub-group or parent/child concepts
> among those IOMMUs for thorough permission control.

My thought here is that this is all abstracted through the vGPU IOMMU
and device vfio backends.  It's the GPU driver itself, or some vfio
extension of that driver, mediating access to the device and deciding
when to configure GPU MMU mappings.  That driver has access to the GPA
to HVA translations thanks to the type1 complaint IOMMU it implements
and can pin pages as needed to create GPA to HPA mappings.  That should
give it all the pieces it needs to fully setup mappings for the vGPU.
Whether or not there's a system IOMMU is simply an exercise for that
driver.  It needs to do a DMA mapping operation through the system IOMMU
the same for a vGPU as if it was doing it for itself, because they are
in fact one in the same.  The GMA to IOVA mapping seems like an internal
detail.  I assume the IOVA is some sort of GPA, and the GMA is managed
through mediation of the device.


> > There are a few areas where we know we'll need to extend the VFIO API to
> > make this work, but it seems like they can all be done generically.  One
> > is that PCI BARs are described through the VFIO API as regions and each
> > region has a single flag describing whether mmap (ie. direct mapping) of
> > that region is possible.  We expect that vGPUs likely need finer
> > granularity, enabling some areas within a BAR to be trapped and fowarded
> > as a read or write access for the vGPU-vfio-device module to emulate,
> > while other regions, like framebuffers or texture regions, are directly
> > mapped.  I have prototype code to enable this already.
> 
> Yes in GVT-g one BAR resource might be partitioned among multiple vGPUs.
> If VFIO can support such partial resource assignment, it'd be great. Similar
> parent/child concept might also be required here, so any resource enumerated 
> on a vGPU shouldn't break limitations enforced on the physical device.

To be clear, I'm talking about partitioning of the BAR exposed to the
guest.  Partitioning of the physical BAR would be managed by the vGPU
vfio device driver.  For instance when the guest mmap's a section of the
virtual BAR, the vGPU device driver would map that to a portion of the
physical device BAR.

> One unique requirement for GVT-g here, though, is that vGPU device model
> need to know guest BAR configuration for proper emulation (e.g. register
> IO emulation handler to KVM). Similar is about guest MSI vector for virtual 
> interrupt injection. Not sure how this can be fit into common VFIO model. 
> Does VFIO allow vendor specific extension today?

As a vfio device driver all config accesses and interrupt configuration
would be forwarded to you, so I don't see this being a problem.

> > 
> > Another area is that we really don't want to proliferate each vGPU
> > needing a new IOMMU type within vfio.  The existing type1 IOMMU provides
> > potentially the most simple mapping and unmapping interface possible.
> > We'd therefore need to allow multiple "type1" IOMMU drivers for vfio,
> > making type1 be more of an interface specification rather than a single
> > implementation.  This is a trivial change to make within vfio and one
> > that I believe is compatible with the existing API.  Note that
> > implementing a type1-compliant vfio IOMMU does not imply pinning an
> > mapping every registered page.  A vGPU, with mediated device access, may
> > use this only to track the current HVA to GPA mappings for a VM.  Only
> > when a DMA is enabled for the vGPU instance is that HVA pinned and an
> > HPA to GPA translation programmed into the GPU MMU.
> > 
> > Another area of extension is how to expose a framebuffer to QEMU for
> > seamless integration into a SPICE/VNC channel.  For this I believe we
> > could use a new region, much like we've done to expose VGA access
> > through a vfio device file descriptor.  An area within this new
> > framebuffer region could be directly mappable in QEMU while a
> > non-mappable page, at a standard location with standardized format,
> > provides a description of framebuffer and potentially even a
> > communication channel to synchronize framebuffer captures.  This would
> > be new code for QEMU, but something we could share among all vGPU
> > implementations.
> 
> Now GVT-g already provides an interface to decode framebuffer information,
> w/ an assumption that the framebuffer will be further composited into 
> OpenGL APIs. So the format is defined according to OpenGL definition.
> Does that meet SPICE requirement?
> 
> Another thing to be added. Framebuffers are frequently switched in
> reality. So either Qemu needs to poll or a notification mechanism is required.
> And since it's dynamic, having framebuffer page directly exposed in the
> new region might be tricky. We can just expose framebuffer information
> (including base, format, etc.) and let Qemu to map separately out of VFIO
> interface.

Sure, we'll need to work out that interface, but it's also possible that
the framebuffer region is simply remapped to another area of the device
(ie. multiple interfaces mapping the same thing) by the vfio device
driver.  Whether it's easier to do that or make the framebuffer region
reference another region is something we'll need to see.

> And... this works fine with vGPU model since software knows all the
> detail about framebuffer. However in pass-through case, who do you expect
> to provide that information? Is it OK to introduce vGPU specific APIs in
> VFIO?

Yes, vGPU may have additional features, like a framebuffer area, that
aren't present or optional for direct assignment.  Obviously we support
direct assignment of GPUs for some vendors already without this feature.

> > Another obvious area to be standardized would be how to discover,
> > create, and destroy vGPU instances.  SR-IOV has a standard mechanism to
> > create VFs in sysfs and I would propose that vGPU vendors try to
> > standardize on similar interfaces to enable libvirt to easily discover
> > the vGPU capabilities of a given GPU and manage the lifecycle of a vGPU
> > instance.
> 
> Now there is no standard. We expose vGPU life-cycle mgmt. APIs through
> sysfs (under i915 node), which is very Intel specific. In reality different
> vendors have quite different capabilities for their own vGPUs, so not sure
> how standard we can define such a mechanism. But this code should be
> minor to be maintained in libvirt.

Every difference is a barrier.  I imagine we can come up with some basic
interfaces that everyone could use, even if they don't allow fine tuning
every detail specific to a vendor.

> > This is obviously a lot to digest, but I'd certainly be interested in
> > hearing feedback on this proposal as well as try to clarify anything
> > I've left out or misrepresented above.  Another benefit to this
> > mechanism is that direct GPU assignment and vGPU assignment use the same
> > code within QEMU and same API to the kernel, which should make debugging
> > and code support between the two easier.  I'd really like to start a
> > discussion around this proposal, and of course the first open source
> > implementation of this sort of model will really help to drive the
> > direction it takes.  Thanks!
> > 
> 
> Thanks for starting this discussion. Intel will definitely work with 
> community on this work. Based on earlier comments, I'm not sure
> whether we can exactly same code for direct GPU assignment and
> vGPU assignment, since even we extend VFIO some interfaces might
> be vGPU specific. Does this way still achieve your end goal?

The backends will certainly be different for vGPU vs direct assignment,
but hopefully the QEMU code is almost entirely reused, modulo some
features like framebuffers that are likely only to be seen on vGPU.
Thanks,

Alex

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1273752

From"Tian, Kevin" <kevin.tian@intel.com>
Date2015-11-20 08:20 +0100
Message-ID<qwO9b-44h-7@gated-at.bofh.it>
In reply to#1273447
PiBGcm9tOiBBbGV4IFdpbGxpYW1zb24gW21haWx0bzphbGV4LndpbGxpYW1zb25AcmVkaGF0LmNv
bV0NCj4gU2VudDogRnJpZGF5LCBOb3ZlbWJlciAyMCwgMjAxNSA0OjAzIEFNDQo+IA0KPiA+ID4N
Cj4gPiA+IFRoZSBwcm9wb3NhbCBpcyB0aGVyZWZvcmUgdGhhdCBHUFUgdmVuZG9ycyBjYW4gZXhw
b3NlIHZHUFVzIHRvDQo+ID4gPiB1c2Vyc3BhY2UsIGFuZCB0aHVzIHRvIFFFTVUsIHVzaW5nIHRo
ZSBWRklPIEFQSS4gIEZvciBpbnN0YW5jZSwgdmZpbw0KPiA+ID4gc3VwcG9ydHMgbW9kdWxhciBi
dXMgZHJpdmVycyBhbmQgSU9NTVUgZHJpdmVycy4gIEFuIGludGVsLXZmaW8tZ3Z0LWQNCj4gPiA+
IG1vZHVsZSAob3IgZXh0ZW5zaW9uIG9mIGk5MTUpIGNhbiByZWdpc3RlciBhcyBhIHZmaW8gYnVz
IGRyaXZlciwgY3JlYXRlDQo+ID4gPiBhIHN0cnVjdCBkZXZpY2UgcGVyIHZHUFUsIGNyZWF0ZSBh
biBJT01NVSBncm91cCBmb3IgdGhhdCBkZXZpY2UsIGFuZA0KPiA+ID4gcmVnaXN0ZXIgdGhhdCBk
ZXZpY2Ugd2l0aCB0aGUgdmZpby1jb3JlLiAgU2luY2Ugd2UgZG9uJ3QgcmVseSBvbiB0aGUNCj4g
PiA+IHN5c3RlbSBJT01NVSBmb3IgR1ZULWQgdkdQVSBhc3NpZ25tZW50LCBhbm90aGVyIHZHUFUg
dmVuZG9yIGRyaXZlciAob3INCj4gPiA+IGV4dGVuc2lvbiBvZiB0aGUgc2FtZSBtb2R1bGUpIGNh
biByZWdpc3RlciBhICJ0eXBlMSIgY29tcGxpYW50IElPTU1VDQo+ID4gPiBkcml2ZXIgaW50byB2
ZmlvLWNvcmUuICBGcm9tIHRoZSBwZXJzcGVjdGl2ZSBvZiBRRU1VIHRoZW4sIGFsbCBvZiB0aGUN
Cj4gPiA+IGV4aXN0aW5nIHZmaW8tcGNpIGNvZGUgaXMgcmUtdXNlZCwgUUVNVSByZW1haW5zIGxh
cmdlbHkgdW5hd2FyZSBvZiBhbnkNCj4gPiA+IHNwZWNpZmljcyBvZiB0aGUgdkdQVSBiZWluZyBh
c3NpZ25lZCwgYW5kIHRoZSBvbmx5IG5lY2Vzc2FyeSBjaGFuZ2Ugc28NCj4gPiA+IGZhciBpcyBo
b3cgUUVNVSB0cmF2ZXJzZXMgc3lzZnMgdG8gZmluZCB0aGUgZGV2aWNlIGFuZCB0aHVzIHRoZSBJ
T01NVQ0KPiA+ID4gZ3JvdXAgbGVhZGluZyB0byB0aGUgdmZpbyBncm91cC4NCj4gPg0KPiA+IEdW
VC1nIHJlcXVpcmVzIHRvIHBpbiBndWVzdCBtZW1vcnkgYW5kIHF1ZXJ5IEdQQS0+SFBBIGluZm9y
bWF0aW9uLA0KPiA+IHVwb24gd2hpY2ggc2hhZG93IEdUVHMgd2lsbCBiZSB1cGRhdGVkIGFjY29y
ZGluZ2x5IGZyb20gKEdNQS0+R1BBKQ0KPiA+IHRvIChHTUEtPkhQQSkuIFNvIHllcywgaGVyZSBh
IGR1bW15IG9yIHNpbXBsZSAidHlwZTEiIGNvbXBsaWFudCBJT01NVQ0KPiA+IGNhbiBiZSBpbnRy
b2R1Y2VkIGp1c3QgZm9yIHRoaXMgcmVxdWlyZW1lbnQuDQo+ID4NCj4gPiBIb3dldmVyIHRoZXJl
J3Mgb25lIHRyaWNreSBwb2ludCB3aGljaCBJJ20gbm90IHN1cmUgd2hldGhlciBvdmVyYWxsDQo+
ID4gVkZJTyBjb25jZXB0IHdpbGwgYmUgdmlvbGF0ZWQuIEdWVC1nIGRvZXNuJ3QgcmVxdWlyZSBz
eXN0ZW0gSU9NTVUNCj4gPiB0byBmdW5jdGlvbiwgaG93ZXZlciBob3N0IHN5c3RlbSBtYXkgZW5h
YmxlIHN5c3RlbSBJT01NVSBqdXN0IGZvcg0KPiA+IGhhcmRlbmluZyBwdXJwb3NlLiBUaGlzIG1l
YW5zIHR3by1sZXZlbCB0cmFuc2xhdGlvbnMgZXhpc3RpbmcgKEdNQS0+DQo+ID4gSU9WQS0+SFBB
KSwgc28gdGhlIGR1bW15IElPTU1VIGRyaXZlciBoYXMgdG8gcmVxdWVzdCBzeXN0ZW0gSU9NTVUN
Cj4gPiBkcml2ZXIgdG8gYWxsb2NhdGUgSU9WQSBmb3IgVk1zIGFuZCB0aGVuIHNldHVwIElPVkEt
PkhQQSBtYXBwaW5nDQo+ID4gaW4gSU9NTVUgcGFnZSB0YWJsZS4gSW4gdGhpcyBjYXNlLCBtdWx0
aXBsZSBWTSdzIHRyYW5zbGF0aW9ucyBhcmUNCj4gPiBtdWx0aXBsZXhlZCBpbiBvbmUgSU9NTVUg
cGFnZSB0YWJsZS4NCj4gPg0KPiA+IFdlIG1pZ2h0IG5lZWQgY3JlYXRlIHNvbWUgZ3JvdXAvc3Vi
LWdyb3VwIG9yIHBhcmVudC9jaGlsZCBjb25jZXB0cw0KPiA+IGFtb25nIHRob3NlIElPTU1VcyBm
b3IgdGhvcm91Z2ggcGVybWlzc2lvbiBjb250cm9sLg0KPiANCj4gTXkgdGhvdWdodCBoZXJlIGlz
IHRoYXQgdGhpcyBpcyBhbGwgYWJzdHJhY3RlZCB0aHJvdWdoIHRoZSB2R1BVIElPTU1VDQo+IGFu
ZCBkZXZpY2UgdmZpbyBiYWNrZW5kcy4gIEl0J3MgdGhlIEdQVSBkcml2ZXIgaXRzZWxmLCBvciBz
b21lIHZmaW8NCj4gZXh0ZW5zaW9uIG9mIHRoYXQgZHJpdmVyLCBtZWRpYXRpbmcgYWNjZXNzIHRv
IHRoZSBkZXZpY2UgYW5kIGRlY2lkaW5nDQo+IHdoZW4gdG8gY29uZmlndXJlIEdQVSBNTVUgbWFw
cGluZ3MuICBUaGF0IGRyaXZlciBoYXMgYWNjZXNzIHRvIHRoZSBHUEENCj4gdG8gSFZBIHRyYW5z
bGF0aW9ucyB0aGFua3MgdG8gdGhlIHR5cGUxIGNvbXBsYWludCBJT01NVSBpdCBpbXBsZW1lbnRz
DQo+IGFuZCBjYW4gcGluIHBhZ2VzIGFzIG5lZWRlZCB0byBjcmVhdGUgR1BBIHRvIEhQQSBtYXBw
aW5ncy4gIFRoYXQgc2hvdWxkDQo+IGdpdmUgaXQgYWxsIHRoZSBwaWVjZXMgaXQgbmVlZHMgdG8g
ZnVsbHkgc2V0dXAgbWFwcGluZ3MgZm9yIHRoZSB2R1BVLg0KPiBXaGV0aGVyIG9yIG5vdCB0aGVy
ZSdzIGEgc3lzdGVtIElPTU1VIGlzIHNpbXBseSBhbiBleGVyY2lzZSBmb3IgdGhhdA0KPiBkcml2
ZXIuICBJdCBuZWVkcyB0byBkbyBhIERNQSBtYXBwaW5nIG9wZXJhdGlvbiB0aHJvdWdoIHRoZSBz
eXN0ZW0gSU9NTVUNCj4gdGhlIHNhbWUgZm9yIGEgdkdQVSBhcyBpZiBpdCB3YXMgZG9pbmcgaXQg
Zm9yIGl0c2VsZiwgYmVjYXVzZSB0aGV5IGFyZQ0KPiBpbiBmYWN0IG9uZSBpbiB0aGUgc2FtZS4g
IFRoZSBHTUEgdG8gSU9WQSBtYXBwaW5nIHNlZW1zIGxpa2UgYW4gaW50ZXJuYWwNCj4gZGV0YWls
LiAgSSBhc3N1bWUgdGhlIElPVkEgaXMgc29tZSBzb3J0IG9mIEdQQSwgYW5kIHRoZSBHTUEgaXMg
bWFuYWdlZA0KPiB0aHJvdWdoIG1lZGlhdGlvbiBvZiB0aGUgZGV2aWNlLg0KDQpTb3JyeSBJJ20g
bm90IGZhbWlsaWFyIHdpdGggVkZJTyBpbnRlcm5hbC4gTXkgb3JpZ2luYWwgd29ycnkgaXMgdGhh
dCBzeXN0ZW0gDQpJT01NVSBmb3IgR1BVIG1heSBiZSBhbHJlYWR5IGNsYWltZWQgYnkgYW5vdGhl
ciB2ZmlvIGRyaXZlciAoZS5nLiBob3N0IGtlcm5lbA0Kd2FudHMgdG8gaGFyZGVuIGdmeCBkcml2
ZXIgZnJvbSByZXN0IHN1Yi1zeXN0ZW1zLCByZWdhcmRsZXNzIG9mIHdoZXRoZXIgdkdQVSANCmlz
IGNyZWF0ZWQgb3Igbm90KS4gSW4gdGhhdCBjYXNlIHZHUFUgSU9NTVUgZHJpdmVyIHNob3VsZG4n
dCBtYW5hZ2Ugc3lzdGVtDQpJT01NVSBkaXJlY3RseS4NCg0KYnR3LCBjdXJpb3VzIHRvZGF5IGhv
dyBWRklPIGNvb3JkaW5hdGVzIHdpdGggc3lzdGVtIElPTU1VIGRyaXZlciByZWdhcmRpbmcNCnRv
IHdoZXRoZXIgYSBJT01NVSBpcyB1c2VkIHRvIGNvbnRyb2wgZGV2aWNlIGFzc2lnbm1lbnQsIG9y
IHVzZWQgZm9yIGtlcm5lbCANCmhhcmRlbmluZy4gU29tZWhvdyB0d28gYXJlIGNvbmZsaWN0aW5n
IHNpbmNlIGRpZmZlcmVudCBhZGRyZXNzIHNwYWNlcyBhcmUNCmNvbmNlcm5lZCAoR1BBIHZzLiBJ
T1ZBKS4uLg0KDQo+IA0KPiANCj4gPiA+IFRoZXJlIGFyZSBhIGZldyBhcmVhcyB3aGVyZSB3ZSBr
bm93IHdlJ2xsIG5lZWQgdG8gZXh0ZW5kIHRoZSBWRklPIEFQSSB0bw0KPiA+ID4gbWFrZSB0aGlz
IHdvcmssIGJ1dCBpdCBzZWVtcyBsaWtlIHRoZXkgY2FuIGFsbCBiZSBkb25lIGdlbmVyaWNhbGx5
LiAgT25lDQo+ID4gPiBpcyB0aGF0IFBDSSBCQVJzIGFyZSBkZXNjcmliZWQgdGhyb3VnaCB0aGUg
VkZJTyBBUEkgYXMgcmVnaW9ucyBhbmQgZWFjaA0KPiA+ID4gcmVnaW9uIGhhcyBhIHNpbmdsZSBm
bGFnIGRlc2NyaWJpbmcgd2hldGhlciBtbWFwIChpZS4gZGlyZWN0IG1hcHBpbmcpIG9mDQo+ID4g
PiB0aGF0IHJlZ2lvbiBpcyBwb3NzaWJsZS4gIFdlIGV4cGVjdCB0aGF0IHZHUFVzIGxpa2VseSBu
ZWVkIGZpbmVyDQo+ID4gPiBncmFudWxhcml0eSwgZW5hYmxpbmcgc29tZSBhcmVhcyB3aXRoaW4g
YSBCQVIgdG8gYmUgdHJhcHBlZCBhbmQgZm93YXJkZWQNCj4gPiA+IGFzIGEgcmVhZCBvciB3cml0
ZSBhY2Nlc3MgZm9yIHRoZSB2R1BVLXZmaW8tZGV2aWNlIG1vZHVsZSB0byBlbXVsYXRlLA0KPiA+
ID4gd2hpbGUgb3RoZXIgcmVnaW9ucywgbGlrZSBmcmFtZWJ1ZmZlcnMgb3IgdGV4dHVyZSByZWdp
b25zLCBhcmUgZGlyZWN0bHkNCj4gPiA+IG1hcHBlZC4gIEkgaGF2ZSBwcm90b3R5cGUgY29kZSB0
byBlbmFibGUgdGhpcyBhbHJlYWR5Lg0KPiA+DQo+ID4gWWVzIGluIEdWVC1nIG9uZSBCQVIgcmVz
b3VyY2UgbWlnaHQgYmUgcGFydGl0aW9uZWQgYW1vbmcgbXVsdGlwbGUgdkdQVXMuDQo+ID4gSWYg
VkZJTyBjYW4gc3VwcG9ydCBzdWNoIHBhcnRpYWwgcmVzb3VyY2UgYXNzaWdubWVudCwgaXQnZCBi
ZSBncmVhdC4gU2ltaWxhcg0KPiA+IHBhcmVudC9jaGlsZCBjb25jZXB0IG1pZ2h0IGFsc28gYmUg
cmVxdWlyZWQgaGVyZSwgc28gYW55IHJlc291cmNlIGVudW1lcmF0ZWQNCj4gPiBvbiBhIHZHUFUg
c2hvdWxkbid0IGJyZWFrIGxpbWl0YXRpb25zIGVuZm9yY2VkIG9uIHRoZSBwaHlzaWNhbCBkZXZp
Y2UuDQo+IA0KPiBUbyBiZSBjbGVhciwgSSdtIHRhbGtpbmcgYWJvdXQgcGFydGl0aW9uaW5nIG9m
IHRoZSBCQVIgZXhwb3NlZCB0byB0aGUNCj4gZ3Vlc3QuICBQYXJ0aXRpb25pbmcgb2YgdGhlIHBo
eXNpY2FsIEJBUiB3b3VsZCBiZSBtYW5hZ2VkIGJ5IHRoZSB2R1BVDQo+IHZmaW8gZGV2aWNlIGRy
aXZlci4gIEZvciBpbnN0YW5jZSB3aGVuIHRoZSBndWVzdCBtbWFwJ3MgYSBzZWN0aW9uIG9mIHRo
ZQ0KPiB2aXJ0dWFsIEJBUiwgdGhlIHZHUFUgZGV2aWNlIGRyaXZlciB3b3VsZCBtYXAgdGhhdCB0
byBhIHBvcnRpb24gb2YgdGhlDQo+IHBoeXNpY2FsIGRldmljZSBCQVIuDQo+IA0KPiA+IE9uZSB1
bmlxdWUgcmVxdWlyZW1lbnQgZm9yIEdWVC1nIGhlcmUsIHRob3VnaCwgaXMgdGhhdCB2R1BVIGRl
dmljZSBtb2RlbA0KPiA+IG5lZWQgdG8ga25vdyBndWVzdCBCQVIgY29uZmlndXJhdGlvbiBmb3Ig
cHJvcGVyIGVtdWxhdGlvbiAoZS5nLiByZWdpc3Rlcg0KPiA+IElPIGVtdWxhdGlvbiBoYW5kbGVy
IHRvIEtWTSkuIFNpbWlsYXIgaXMgYWJvdXQgZ3Vlc3QgTVNJIHZlY3RvciBmb3IgdmlydHVhbA0K
PiA+IGludGVycnVwdCBpbmplY3Rpb24uIE5vdCBzdXJlIGhvdyB0aGlzIGNhbiBiZSBmaXQgaW50
byBjb21tb24gVkZJTyBtb2RlbC4NCj4gPiBEb2VzIFZGSU8gYWxsb3cgdmVuZG9yIHNwZWNpZmlj
IGV4dGVuc2lvbiB0b2RheT8NCj4gDQo+IEFzIGEgdmZpbyBkZXZpY2UgZHJpdmVyIGFsbCBjb25m
aWcgYWNjZXNzZXMgYW5kIGludGVycnVwdCBjb25maWd1cmF0aW9uDQo+IHdvdWxkIGJlIGZvcndh
cmRlZCB0byB5b3UsIHNvIEkgZG9uJ3Qgc2VlIHRoaXMgYmVpbmcgYSBwcm9ibGVtLg0KDQpTdXJl
LCBuaWNlIHRvIGtub3cgdGhhdC4NCg0KVGhhbmtzDQpLZXZpbg0K
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1274274

FromAlex Williamson <alex.williamson@redhat.com>
Date2015-11-20 18:10 +0100
Message-ID<qwXm9-1Fh-11@gated-at.bofh.it>
In reply to#1273752
On Fri, 2015-11-20 at 07:09 +0000, Tian, Kevin wrote:
> > From: Alex Williamson [mailto:alex.williamson@redhat.com]
> > Sent: Friday, November 20, 2015 4:03 AM
> > 
> > > >
> > > > The proposal is therefore that GPU vendors can expose vGPUs to
> > > > userspace, and thus to QEMU, using the VFIO API.  For instance, vfio
> > > > supports modular bus drivers and IOMMU drivers.  An intel-vfio-gvt-d
> > > > module (or extension of i915) can register as a vfio bus driver, create
> > > > a struct device per vGPU, create an IOMMU group for that device, and
> > > > register that device with the vfio-core.  Since we don't rely on the
> > > > system IOMMU for GVT-d vGPU assignment, another vGPU vendor driver (or
> > > > extension of the same module) can register a "type1" compliant IOMMU
> > > > driver into vfio-core.  From the perspective of QEMU then, all of the
> > > > existing vfio-pci code is re-used, QEMU remains largely unaware of any
> > > > specifics of the vGPU being assigned, and the only necessary change so
> > > > far is how QEMU traverses sysfs to find the device and thus the IOMMU
> > > > group leading to the vfio group.
> > >
> > > GVT-g requires to pin guest memory and query GPA->HPA information,
> > > upon which shadow GTTs will be updated accordingly from (GMA->GPA)
> > > to (GMA->HPA). So yes, here a dummy or simple "type1" compliant IOMMU
> > > can be introduced just for this requirement.
> > >
> > > However there's one tricky point which I'm not sure whether overall
> > > VFIO concept will be violated. GVT-g doesn't require system IOMMU
> > > to function, however host system may enable system IOMMU just for
> > > hardening purpose. This means two-level translations existing (GMA->
> > > IOVA->HPA), so the dummy IOMMU driver has to request system IOMMU
> > > driver to allocate IOVA for VMs and then setup IOVA->HPA mapping
> > > in IOMMU page table. In this case, multiple VM's translations are
> > > multiplexed in one IOMMU page table.
> > >
> > > We might need create some group/sub-group or parent/child concepts
> > > among those IOMMUs for thorough permission control.
> > 
> > My thought here is that this is all abstracted through the vGPU IOMMU
> > and device vfio backends.  It's the GPU driver itself, or some vfio
> > extension of that driver, mediating access to the device and deciding
> > when to configure GPU MMU mappings.  That driver has access to the GPA
> > to HVA translations thanks to the type1 complaint IOMMU it implements
> > and can pin pages as needed to create GPA to HPA mappings.  That should
> > give it all the pieces it needs to fully setup mappings for the vGPU.
> > Whether or not there's a system IOMMU is simply an exercise for that
> > driver.  It needs to do a DMA mapping operation through the system IOMMU
> > the same for a vGPU as if it was doing it for itself, because they are
> > in fact one in the same.  The GMA to IOVA mapping seems like an internal
> > detail.  I assume the IOVA is some sort of GPA, and the GMA is managed
> > through mediation of the device.
> 
> Sorry I'm not familiar with VFIO internal. My original worry is that system 
> IOMMU for GPU may be already claimed by another vfio driver (e.g. host kernel
> wants to harden gfx driver from rest sub-systems, regardless of whether vGPU 
> is created or not). In that case vGPU IOMMU driver shouldn't manage system
> IOMMU directly.

There are different APIs for the IOMMU depending on how it's being use.
If the IOMMU is being used for inter-device isolation in the host, then
the DMA API (ex. dma_map_page) transparently makes use of the IOMMU.
When we're doing device assignment, we make use of the IOMMU API which
allows more explicit control (ex. iommu_domain_alloc,
iommu_attach_device, iommu_map, etc).  A vGPU is not an SR-IOV VF, it
doesn't have a unique requester ID that allows the IOMMU to
differentiate one vGPU from another, or vGPU from GPU.  All mappings for
vGPUs need to occur for the GPU.  It's therefore the responsibility of
the GPU driver, or this vfio extension of that driver, that needs to
perform the IOMMU mapping for the vGPU.

My expectation is therefore that once the GMA to IOVA mapping is
configured in the GPU MMU, the IOVA to HPA needs to be programmed, as if
the GPU driver was performing the setup itself, which it is.  Before the
device mediation that triggered the mapping setup is complete, the GPU
MMU and the system IOMMU (if preset) should be configured to enable that
DMA.  The GPU MMU provides the isolation of the vGPU, the system IOMMU
enable the DMA to occur.

> btw, curious today how VFIO coordinates with system IOMMU driver regarding
> to whether a IOMMU is used to control device assignment, or used for kernel 
> hardening. Somehow two are conflicting since different address spaces are
> concerned (GPA vs. IOVA)...

When devices unbind from native host drivers, any previous IOMMU
mappings and domains are removed.  These are typically created via the
DMA API above.  The initialization operations of the VFIO API (creating
containers, attaching groups to containers, and setting the IOMMU model
for a container) work through the IOMMU API to create a new domain and
isolate devices within it.  The type1 VFIO IOMMU interface is then
effectively a passthrough to the iommu_map() and iommu_unmap()
interfaces of the IOMMU API, modulo page pinning, accounting and
tracking.  When a VFIO instance is destroyed, the devices are detached
from the IOMMU domain, the devices are unbound from vfio and re-bound to
host drivers and the DMA API can reclaim the devices for host isolation.
Thanks,

Alex

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1273796

From"Tian, Kevin" <kevin.tian@intel.com>
Date2015-11-20 09:20 +0100
Message-ID<qwP5f-4EO-5@gated-at.bofh.it>
In reply to#1273447
PiBGcm9tOiBUaWFuLCBLZXZpbg0KPiBTZW50OiBGcmlkYXksIE5vdmVtYmVyIDIwLCAyMDE1IDM6
MTAgUE0NCg0KPiA+ID4gPg0KPiA+ID4gPiBUaGUgcHJvcG9zYWwgaXMgdGhlcmVmb3JlIHRoYXQg
R1BVIHZlbmRvcnMgY2FuIGV4cG9zZSB2R1BVcyB0bw0KPiA+ID4gPiB1c2Vyc3BhY2UsIGFuZCB0
aHVzIHRvIFFFTVUsIHVzaW5nIHRoZSBWRklPIEFQSS4gIEZvciBpbnN0YW5jZSwgdmZpbw0KPiA+
ID4gPiBzdXBwb3J0cyBtb2R1bGFyIGJ1cyBkcml2ZXJzIGFuZCBJT01NVSBkcml2ZXJzLiAgQW4g
aW50ZWwtdmZpby1ndnQtZA0KPiA+ID4gPiBtb2R1bGUgKG9yIGV4dGVuc2lvbiBvZiBpOTE1KSBj
YW4gcmVnaXN0ZXIgYXMgYSB2ZmlvIGJ1cyBkcml2ZXIsIGNyZWF0ZQ0KPiA+ID4gPiBhIHN0cnVj
dCBkZXZpY2UgcGVyIHZHUFUsIGNyZWF0ZSBhbiBJT01NVSBncm91cCBmb3IgdGhhdCBkZXZpY2Us
IGFuZA0KPiA+ID4gPiByZWdpc3RlciB0aGF0IGRldmljZSB3aXRoIHRoZSB2ZmlvLWNvcmUuICBT
aW5jZSB3ZSBkb24ndCByZWx5IG9uIHRoZQ0KPiA+ID4gPiBzeXN0ZW0gSU9NTVUgZm9yIEdWVC1k
IHZHUFUgYXNzaWdubWVudCwgYW5vdGhlciB2R1BVIHZlbmRvciBkcml2ZXIgKG9yDQo+ID4gPiA+
IGV4dGVuc2lvbiBvZiB0aGUgc2FtZSBtb2R1bGUpIGNhbiByZWdpc3RlciBhICJ0eXBlMSIgY29t
cGxpYW50IElPTU1VDQo+ID4gPiA+IGRyaXZlciBpbnRvIHZmaW8tY29yZS4gIEZyb20gdGhlIHBl
cnNwZWN0aXZlIG9mIFFFTVUgdGhlbiwgYWxsIG9mIHRoZQ0KPiA+ID4gPiBleGlzdGluZyB2Zmlv
LXBjaSBjb2RlIGlzIHJlLXVzZWQsIFFFTVUgcmVtYWlucyBsYXJnZWx5IHVuYXdhcmUgb2YgYW55
DQo+ID4gPiA+IHNwZWNpZmljcyBvZiB0aGUgdkdQVSBiZWluZyBhc3NpZ25lZCwgYW5kIHRoZSBv
bmx5IG5lY2Vzc2FyeSBjaGFuZ2Ugc28NCj4gPiA+ID4gZmFyIGlzIGhvdyBRRU1VIHRyYXZlcnNl
cyBzeXNmcyB0byBmaW5kIHRoZSBkZXZpY2UgYW5kIHRodXMgdGhlIElPTU1VDQo+ID4gPiA+IGdy
b3VwIGxlYWRpbmcgdG8gdGhlIHZmaW8gZ3JvdXAuDQo+ID4gPg0KPiA+ID4gR1ZULWcgcmVxdWly
ZXMgdG8gcGluIGd1ZXN0IG1lbW9yeSBhbmQgcXVlcnkgR1BBLT5IUEEgaW5mb3JtYXRpb24sDQo+
ID4gPiB1cG9uIHdoaWNoIHNoYWRvdyBHVFRzIHdpbGwgYmUgdXBkYXRlZCBhY2NvcmRpbmdseSBm
cm9tIChHTUEtPkdQQSkNCj4gPiA+IHRvIChHTUEtPkhQQSkuIFNvIHllcywgaGVyZSBhIGR1bW15
IG9yIHNpbXBsZSAidHlwZTEiIGNvbXBsaWFudCBJT01NVQ0KPiA+ID4gY2FuIGJlIGludHJvZHVj
ZWQganVzdCBmb3IgdGhpcyByZXF1aXJlbWVudC4NCj4gPiA+DQo+ID4gPiBIb3dldmVyIHRoZXJl
J3Mgb25lIHRyaWNreSBwb2ludCB3aGljaCBJJ20gbm90IHN1cmUgd2hldGhlciBvdmVyYWxsDQo+
ID4gPiBWRklPIGNvbmNlcHQgd2lsbCBiZSB2aW9sYXRlZC4gR1ZULWcgZG9lc24ndCByZXF1aXJl
IHN5c3RlbSBJT01NVQ0KPiA+ID4gdG8gZnVuY3Rpb24sIGhvd2V2ZXIgaG9zdCBzeXN0ZW0gbWF5
IGVuYWJsZSBzeXN0ZW0gSU9NTVUganVzdCBmb3INCj4gPiA+IGhhcmRlbmluZyBwdXJwb3NlLiBU
aGlzIG1lYW5zIHR3by1sZXZlbCB0cmFuc2xhdGlvbnMgZXhpc3RpbmcgKEdNQS0+DQo+ID4gPiBJ
T1ZBLT5IUEEpLCBzbyB0aGUgZHVtbXkgSU9NTVUgZHJpdmVyIGhhcyB0byByZXF1ZXN0IHN5c3Rl
bSBJT01NVQ0KPiA+ID4gZHJpdmVyIHRvIGFsbG9jYXRlIElPVkEgZm9yIFZNcyBhbmQgdGhlbiBz
ZXR1cCBJT1ZBLT5IUEEgbWFwcGluZw0KPiA+ID4gaW4gSU9NTVUgcGFnZSB0YWJsZS4gSW4gdGhp
cyBjYXNlLCBtdWx0aXBsZSBWTSdzIHRyYW5zbGF0aW9ucyBhcmUNCj4gPiA+IG11bHRpcGxleGVk
IGluIG9uZSBJT01NVSBwYWdlIHRhYmxlLg0KPiA+ID4NCj4gPiA+IFdlIG1pZ2h0IG5lZWQgY3Jl
YXRlIHNvbWUgZ3JvdXAvc3ViLWdyb3VwIG9yIHBhcmVudC9jaGlsZCBjb25jZXB0cw0KPiA+ID4g
YW1vbmcgdGhvc2UgSU9NTVVzIGZvciB0aG9yb3VnaCBwZXJtaXNzaW9uIGNvbnRyb2wuDQo+ID4N
Cj4gPiBNeSB0aG91Z2h0IGhlcmUgaXMgdGhhdCB0aGlzIGlzIGFsbCBhYnN0cmFjdGVkIHRocm91
Z2ggdGhlIHZHUFUgSU9NTVUNCj4gPiBhbmQgZGV2aWNlIHZmaW8gYmFja2VuZHMuICBJdCdzIHRo
ZSBHUFUgZHJpdmVyIGl0c2VsZiwgb3Igc29tZSB2ZmlvDQo+ID4gZXh0ZW5zaW9uIG9mIHRoYXQg
ZHJpdmVyLCBtZWRpYXRpbmcgYWNjZXNzIHRvIHRoZSBkZXZpY2UgYW5kIGRlY2lkaW5nDQo+ID4g
d2hlbiB0byBjb25maWd1cmUgR1BVIE1NVSBtYXBwaW5ncy4gIFRoYXQgZHJpdmVyIGhhcyBhY2Nl
c3MgdG8gdGhlIEdQQQ0KPiA+IHRvIEhWQSB0cmFuc2xhdGlvbnMgdGhhbmtzIHRvIHRoZSB0eXBl
MSBjb21wbGFpbnQgSU9NTVUgaXQgaW1wbGVtZW50cw0KPiA+IGFuZCBjYW4gcGluIHBhZ2VzIGFz
IG5lZWRlZCB0byBjcmVhdGUgR1BBIHRvIEhQQSBtYXBwaW5ncy4gIFRoYXQgc2hvdWxkDQo+ID4g
Z2l2ZSBpdCBhbGwgdGhlIHBpZWNlcyBpdCBuZWVkcyB0byBmdWxseSBzZXR1cCBtYXBwaW5ncyBm
b3IgdGhlIHZHUFUuDQo+ID4gV2hldGhlciBvciBub3QgdGhlcmUncyBhIHN5c3RlbSBJT01NVSBp
cyBzaW1wbHkgYW4gZXhlcmNpc2UgZm9yIHRoYXQNCj4gPiBkcml2ZXIuICBJdCBuZWVkcyB0byBk
byBhIERNQSBtYXBwaW5nIG9wZXJhdGlvbiB0aHJvdWdoIHRoZSBzeXN0ZW0gSU9NTVUNCj4gPiB0
aGUgc2FtZSBmb3IgYSB2R1BVIGFzIGlmIGl0IHdhcyBkb2luZyBpdCBmb3IgaXRzZWxmLCBiZWNh
dXNlIHRoZXkgYXJlDQo+ID4gaW4gZmFjdCBvbmUgaW4gdGhlIHNhbWUuICBUaGUgR01BIHRvIElP
VkEgbWFwcGluZyBzZWVtcyBsaWtlIGFuIGludGVybmFsDQo+ID4gZGV0YWlsLiAgSSBhc3N1bWUg
dGhlIElPVkEgaXMgc29tZSBzb3J0IG9mIEdQQSwgYW5kIHRoZSBHTUEgaXMgbWFuYWdlZA0KPiA+
IHRocm91Z2ggbWVkaWF0aW9uIG9mIHRoZSBkZXZpY2UuDQo+IA0KPiBTb3JyeSBJJ20gbm90IGZh
bWlsaWFyIHdpdGggVkZJTyBpbnRlcm5hbC4gTXkgb3JpZ2luYWwgd29ycnkgaXMgdGhhdCBzeXN0
ZW0NCj4gSU9NTVUgZm9yIEdQVSBtYXkgYmUgYWxyZWFkeSBjbGFpbWVkIGJ5IGFub3RoZXIgdmZp
byBkcml2ZXIgKGUuZy4gaG9zdCBrZXJuZWwNCj4gd2FudHMgdG8gaGFyZGVuIGdmeCBkcml2ZXIg
ZnJvbSByZXN0IHN1Yi1zeXN0ZW1zLCByZWdhcmRsZXNzIG9mIHdoZXRoZXIgdkdQVQ0KPiBpcyBj
cmVhdGVkIG9yIG5vdCkuIEluIHRoYXQgY2FzZSB2R1BVIElPTU1VIGRyaXZlciBzaG91bGRuJ3Qg
bWFuYWdlIHN5c3RlbQ0KPiBJT01NVSBkaXJlY3RseS4NCj4gDQo+IGJ0dywgY3VyaW91cyB0b2Rh
eSBob3cgVkZJTyBjb29yZGluYXRlcyB3aXRoIHN5c3RlbSBJT01NVSBkcml2ZXIgcmVnYXJkaW5n
DQo+IHRvIHdoZXRoZXIgYSBJT01NVSBpcyB1c2VkIHRvIGNvbnRyb2wgZGV2aWNlIGFzc2lnbm1l
bnQsIG9yIHVzZWQgZm9yIGtlcm5lbA0KPiBoYXJkZW5pbmcuIFNvbWVob3cgdHdvIGFyZSBjb25m
bGljdGluZyBzaW5jZSBkaWZmZXJlbnQgYWRkcmVzcyBzcGFjZXMgYXJlDQo+IGNvbmNlcm5lZCAo
R1BBIHZzLiBJT1ZBKS4uLg0KPiANCg0KSGVyZSBpcyBhIG1vcmUgY29uY3JldGUgZXhhbXBsZToN
Cg0KS1ZNR1QgZG9lc24ndCByZXF1aXJlIElPTU1VLiBBbGwgRE1BIHRhcmdldHMgYXJlIGFscmVh
ZHkgcmVwbGFjZWQgd2l0aCANCkhQQSB0aHJ1IHNoYWRvdyBHVFQuIFNvIERNQSByZXF1ZXN0cyBm
cm9tIEdQVSBhbGwgY29udGFpbiBIUEFzLg0KDQpXaGVuIElPTU1VIGlzIGVuYWJsZWQsIG9uZSBz
aW1wbGUgYXBwcm9hY2ggaXMgdG8gaGF2ZSB2R1BVIElPTU1VDQpkcml2ZXIgY29uZmlndXJlIHN5
c3RlbSBJT01NVSB3aXRoIGlkZW50aXR5IG1hcHBpbmcgKEhQQS0+SFBBKS4gV2UgDQpjYW4ndCB1
c2UgKEdQQS0+SFBBKSBzaW5jZSBHUEFzIGZyb20gbXVsdGlwbGUgVk1zIGFyZSBjb25mbGljdGlu
Zy4gDQoNCkhvd2V2ZXIsIHdlIHN0aWxsIGhhdmUgaG9zdCBnZnggZHJpdmVyIHJ1bm5pbmcuIFdo
ZW4gSU9NTVUgaXMgZW5hYmxlZCwgDQpkbWFfYWxsb2NfKioqIHdpbGwgcmV0dXJuIElPVkEgKGRy
dmVycy9pb21tdS9pb3ZhLmMpIGluIGhvc3QgZ2Z4IGRyaXZlciwNCndoaWNoIHdpbGwgaGF2ZSBJ
T1ZBLT5IUEEgcHJvZ3JhbW1lZCB0byBzeXN0ZW0gSU9NTVUuDQoNCk9uZSBJT01NVSBkZXZpY2Ug
ZW50cnkgY2FuIG9ubHkgdHJhbnNsYXRlIG9uZSBhZGRyZXNzIHNwYWNlLCBzbyBoZXJlDQpjb21l
cyBhIGNvbmZsaWN0IChIUEEtPkhQQSB2cy4gSU9WQS0+SFBBKS4gVG8gc29sdmUgdGhpcywgdkdQ
VSBJT01NVQ0KZHJpdmVyIG5lZWRzIHRvIGFsbG9jYXRlIElPVkEgZnJvbSBpb3ZhLmMgZm9yIGVh
Y2ggVk0gdy8gdkdQVSBhc3NpZ25lZCwNCmFuZCB0aGVuIEtWTUdUIHdpbGwgcHJvZ3JhbSBJT1ZB
IGluIHNoYWRvdyBHVFQgYWNjb3JkaW5nbHkuIEl0IGFkZHMNCm9uZSBhZGRpdGlvbmFsIG1hcHBp
bmcgbGF5ZXIgKEdQQS0+SU9WQS0+SFBBKS4gSW4gdGhpcyB3YXkgdHdvIA0KcmVxdWlyZW1lbnRz
IGNhbiBiZSB1bmlmaWVkIHRvZ2V0aGVyIHNpbmNlIG9ubHkgSU9WQS0+SFBBIG1hcHBpbmcgDQpu
ZWVkcyB0byBiZSBidWlsdC4NCg0KU28gdW5saWtlIGV4aXN0aW5nIHR5cGUxIElPTU1VIGRyaXZl
ciB3aGljaCBjb250cm9scyBJT01NVSBhbG9uZSwgdkdQVSANCklPTU1VIGRyaXZlciBuZWVkcyB0
byBjb29wZXJhdGUgd2l0aCBvdGhlciBhZ2VudCAoaW92YS5jIGhlcmUpIHRvDQpjby1tYW5hZ2Ug
c3lzdGVtIElPTU1VLiBUaGlzIG1heSBub3QgaW1wYWN0IGV4aXN0aW5nIFZGSU8gZnJhbWV3b3Jr
Lg0KSnVzdCB3YW50IHRvIGhpZ2hsaWdodCBhZGRpdGlvbmFsIHdvcmsgaGVyZSB3aGVuIGltcGxl
bWVudGluZyB0aGUgdkdQVQ0KSU9NTVUgZHJpdmVyLg0KDQpUaGFua3MNCktldmluDQogDQoNCg0K
VGhhbmtzDQpLZXZpbg0K
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1274303

FromAlex Williamson <alex.williamson@redhat.com>
Date2015-11-20 18:30 +0100
Message-ID<qwXFw-1Np-1@gated-at.bofh.it>
In reply to#1273796
On Fri, 2015-11-20 at 08:10 +0000, Tian, Kevin wrote:
> > From: Tian, Kevin
> > Sent: Friday, November 20, 2015 3:10 PM
> 
> > > > >
> > > > > The proposal is therefore that GPU vendors can expose vGPUs to
> > > > > userspace, and thus to QEMU, using the VFIO API.  For instance, vfio
> > > > > supports modular bus drivers and IOMMU drivers.  An intel-vfio-gvt-d
> > > > > module (or extension of i915) can register as a vfio bus driver, create
> > > > > a struct device per vGPU, create an IOMMU group for that device, and
> > > > > register that device with the vfio-core.  Since we don't rely on the
> > > > > system IOMMU for GVT-d vGPU assignment, another vGPU vendor driver (or
> > > > > extension of the same module) can register a "type1" compliant IOMMU
> > > > > driver into vfio-core.  From the perspective of QEMU then, all of the
> > > > > existing vfio-pci code is re-used, QEMU remains largely unaware of any
> > > > > specifics of the vGPU being assigned, and the only necessary change so
> > > > > far is how QEMU traverses sysfs to find the device and thus the IOMMU
> > > > > group leading to the vfio group.
> > > >
> > > > GVT-g requires to pin guest memory and query GPA->HPA information,
> > > > upon which shadow GTTs will be updated accordingly from (GMA->GPA)
> > > > to (GMA->HPA). So yes, here a dummy or simple "type1" compliant IOMMU
> > > > can be introduced just for this requirement.
> > > >
> > > > However there's one tricky point which I'm not sure whether overall
> > > > VFIO concept will be violated. GVT-g doesn't require system IOMMU
> > > > to function, however host system may enable system IOMMU just for
> > > > hardening purpose. This means two-level translations existing (GMA->
> > > > IOVA->HPA), so the dummy IOMMU driver has to request system IOMMU
> > > > driver to allocate IOVA for VMs and then setup IOVA->HPA mapping
> > > > in IOMMU page table. In this case, multiple VM's translations are
> > > > multiplexed in one IOMMU page table.
> > > >
> > > > We might need create some group/sub-group or parent/child concepts
> > > > among those IOMMUs for thorough permission control.
> > >
> > > My thought here is that this is all abstracted through the vGPU IOMMU
> > > and device vfio backends.  It's the GPU driver itself, or some vfio
> > > extension of that driver, mediating access to the device and deciding
> > > when to configure GPU MMU mappings.  That driver has access to the GPA
> > > to HVA translations thanks to the type1 complaint IOMMU it implements
> > > and can pin pages as needed to create GPA to HPA mappings.  That should
> > > give it all the pieces it needs to fully setup mappings for the vGPU.
> > > Whether or not there's a system IOMMU is simply an exercise for that
> > > driver.  It needs to do a DMA mapping operation through the system IOMMU
> > > the same for a vGPU as if it was doing it for itself, because they are
> > > in fact one in the same.  The GMA to IOVA mapping seems like an internal
> > > detail.  I assume the IOVA is some sort of GPA, and the GMA is managed
> > > through mediation of the device.
> > 
> > Sorry I'm not familiar with VFIO internal. My original worry is that system
> > IOMMU for GPU may be already claimed by another vfio driver (e.g. host kernel
> > wants to harden gfx driver from rest sub-systems, regardless of whether vGPU
> > is created or not). In that case vGPU IOMMU driver shouldn't manage system
> > IOMMU directly.
> > 
> > btw, curious today how VFIO coordinates with system IOMMU driver regarding
> > to whether a IOMMU is used to control device assignment, or used for kernel
> > hardening. Somehow two are conflicting since different address spaces are
> > concerned (GPA vs. IOVA)...
> > 
> 
> Here is a more concrete example:
> 
> KVMGT doesn't require IOMMU. All DMA targets are already replaced with 
> HPA thru shadow GTT. So DMA requests from GPU all contain HPAs.
> 
> When IOMMU is enabled, one simple approach is to have vGPU IOMMU
> driver configure system IOMMU with identity mapping (HPA->HPA). We 
> can't use (GPA->HPA) since GPAs from multiple VMs are conflicting. 
> 
> However, we still have host gfx driver running. When IOMMU is enabled, 
> dma_alloc_*** will return IOVA (drvers/iommu/iova.c) in host gfx driver,
> which will have IOVA->HPA programmed to system IOMMU.
> 
> One IOMMU device entry can only translate one address space, so here
> comes a conflict (HPA->HPA vs. IOVA->HPA). To solve this, vGPU IOMMU
> driver needs to allocate IOVA from iova.c for each VM w/ vGPU assigned,
> and then KVMGT will program IOVA in shadow GTT accordingly. It adds
> one additional mapping layer (GPA->IOVA->HPA). In this way two 
> requirements can be unified together since only IOVA->HPA mapping 
> needs to be built.
> 
> So unlike existing type1 IOMMU driver which controls IOMMU alone, vGPU 
> IOMMU driver needs to cooperate with other agent (iova.c here) to
> co-manage system IOMMU. This may not impact existing VFIO framework.
> Just want to highlight additional work here when implementing the vGPU
> IOMMU driver.

Right, so the existing i915 driver needs to use the DMA API and calls
like dma_map_page() to enable translations through the IOMMU.  With
dma_map_page(), the caller provides a page address (~HPA) and is
returned an IOVA.  So unfortunately you don't get to take the shortcut
of having an identity mapping through the IOMMU unless you want to
convert i915 entirely to using the IOMMU API, because we also can't have
the conflict that an HPA could overlap an IOVA for a previously mapped
page.

The double translation, once through the GPU MMU and once through the
system IOMMU is going to happen regardless of whether we can identity
map through the IOMMU.  The only solution to this would be for the GPU
to participate in ATS and provide pre-translated transactions from the
GPU.  All of this is internal to the i915 driver (or vfio extension of
that driver) and needs to be done regardless of what sort of interface
we're using to expose the vGPU to QEMU.  It just seems like VFIO
provides a convenient way of doing this since you'll have ready access
to the HVA-GPA mappings for the user.

I think the key points though are:

      * the VFIO type1 IOMMU stores GPA to HVA translations
      * get_user_pages() on the HVA will pin the page and give you a
        page
      * dma_map_page() receives that page, programs the system IOMMU and
        provides an IOVA
      * the GPU MMU can then be programmed with the GPA to IOVA
        translations

Thanks,
Alex

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1275049 — Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel

FromJike Song <jike.song@intel.com>
Date2015-11-23 06:10 +0100
SubjectRe: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel
Message-ID<qxRy2-5Kz-1@gated-at.bofh.it>
In reply to#1274303
On 11/21/2015 01:25 AM, Alex Williamson wrote:
> On Fri, 2015-11-20 at 08:10 +0000, Tian, Kevin wrote:
>>
>> Here is a more concrete example:
>>
>> KVMGT doesn't require IOMMU. All DMA targets are already replaced with
>> HPA thru shadow GTT. So DMA requests from GPU all contain HPAs.
>>
>> When IOMMU is enabled, one simple approach is to have vGPU IOMMU
>> driver configure system IOMMU with identity mapping (HPA->HPA). We
>> can't use (GPA->HPA) since GPAs from multiple VMs are conflicting.
>>
>> However, we still have host gfx driver running. When IOMMU is enabled,
>> dma_alloc_*** will return IOVA (drvers/iommu/iova.c) in host gfx driver,
>> which will have IOVA->HPA programmed to system IOMMU.
>>
>> One IOMMU device entry can only translate one address space, so here
>> comes a conflict (HPA->HPA vs. IOVA->HPA). To solve this, vGPU IOMMU
>> driver needs to allocate IOVA from iova.c for each VM w/ vGPU assigned,
>> and then KVMGT will program IOVA in shadow GTT accordingly. It adds
>> one additional mapping layer (GPA->IOVA->HPA). In this way two
>> requirements can be unified together since only IOVA->HPA mapping
>> needs to be built.
>>
>> So unlike existing type1 IOMMU driver which controls IOMMU alone, vGPU
>> IOMMU driver needs to cooperate with other agent (iova.c here) to
>> co-manage system IOMMU. This may not impact existing VFIO framework.
>> Just want to highlight additional work here when implementing the vGPU
>> IOMMU driver.
>
> Right, so the existing i915 driver needs to use the DMA API and calls
> like dma_map_page() to enable translations through the IOMMU.  With
> dma_map_page(), the caller provides a page address (~HPA) and is
> returned an IOVA.  So unfortunately you don't get to take the shortcut
> of having an identity mapping through the IOMMU unless you want to
> convert i915 entirely to using the IOMMU API, because we also can't have
> the conflict that an HPA could overlap an IOVA for a previously mapped
> page.
>
> The double translation, once through the GPU MMU and once through the
> system IOMMU is going to happen regardless of whether we can identity
> map through the IOMMU.  The only solution to this would be for the GPU
> to participate in ATS and provide pre-translated transactions from the
> GPU.  All of this is internal to the i915 driver (or vfio extension of
> that driver) and needs to be done regardless of what sort of interface
> we're using to expose the vGPU to QEMU.  It just seems like VFIO
> provides a convenient way of doing this since you'll have ready access
> to the HVA-GPA mappings for the user.
>
> I think the key points though are:
>
>        * the VFIO type1 IOMMU stores GPA to HVA translations
>        * get_user_pages() on the HVA will pin the page and give you a
>          page
>        * dma_map_page() receives that page, programs the system IOMMU and
>          provides an IOVA
>        * the GPU MMU can then be programmed with the GPA to IOVA
>          translations

Thanks for such a nice example! I'll do my home work and get back to you
shortly :)

>
> Thanks,
> Alex
>

--
Thanks,
Jike
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1276337 — Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel

FromDaniel Vetter <daniel@ffwll.ch>
Date2015-11-24 12:20 +0100
SubjectRe: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel
Message-ID<qyjND-7oT-9@gated-at.bofh.it>
In reply to#1273447
On Thu, Nov 19, 2015 at 01:02:36PM -0700, Alex Williamson wrote:
> On Thu, 2015-11-19 at 04:06 +0000, Tian, Kevin wrote:
> > > From: Alex Williamson [mailto:alex.williamson@redhat.com]
> > > Sent: Thursday, November 19, 2015 2:12 AM
> > > 
> > > [cc +qemu-devel, +paolo, +gerd]
> > > 
> > > Another area of extension is how to expose a framebuffer to QEMU for
> > > seamless integration into a SPICE/VNC channel.  For this I believe we
> > > could use a new region, much like we've done to expose VGA access
> > > through a vfio device file descriptor.  An area within this new
> > > framebuffer region could be directly mappable in QEMU while a
> > > non-mappable page, at a standard location with standardized format,
> > > provides a description of framebuffer and potentially even a
> > > communication channel to synchronize framebuffer captures.  This would
> > > be new code for QEMU, but something we could share among all vGPU
> > > implementations.
> > 
> > Now GVT-g already provides an interface to decode framebuffer information,
> > w/ an assumption that the framebuffer will be further composited into 
> > OpenGL APIs. So the format is defined according to OpenGL definition.
> > Does that meet SPICE requirement?
> > 
> > Another thing to be added. Framebuffers are frequently switched in
> > reality. So either Qemu needs to poll or a notification mechanism is required.
> > And since it's dynamic, having framebuffer page directly exposed in the
> > new region might be tricky. We can just expose framebuffer information
> > (including base, format, etc.) and let Qemu to map separately out of VFIO
> > interface.
> 
> Sure, we'll need to work out that interface, but it's also possible that
> the framebuffer region is simply remapped to another area of the device
> (ie. multiple interfaces mapping the same thing) by the vfio device
> driver.  Whether it's easier to do that or make the framebuffer region
> reference another region is something we'll need to see.
> 
> > And... this works fine with vGPU model since software knows all the
> > detail about framebuffer. However in pass-through case, who do you expect
> > to provide that information? Is it OK to introduce vGPU specific APIs in
> > VFIO?
> 
> Yes, vGPU may have additional features, like a framebuffer area, that
> aren't present or optional for direct assignment.  Obviously we support
> direct assignment of GPUs for some vendors already without this feature.

For exposing framebuffers for spice/vnc I highly recommend against
anything that looks like a bar/fixed mmio range mapping. First this means
the kernel driver needs to internally fake remapping, which isn't fun.
Second we can't get at the memory in an easy fashion for hw-accelerated
compositing.

My recoomendation is to build the actual memory access for underlying
framebuffers on top of dma-buf, so that it can be vacuumed up by e.g. the
host gpu driver again for rendering. For userspace the generic part would
simply be an invalidate-fb signal, with the new dma-buf supplied.

Upsides:
- You can composit stuff with the gpu.
- VRAM and other kinds of resources (even stuff not visible in pci bars)
  can be represented.

Downside: Tracking mapping changes on the guest side won't be any easier.
This is mostly a problem for integrated gpus, since discrete ones usually
require contiguous vram for scanout. I think saying "don't do that" is a
valid option though, i.e. we're assuming that page mappings for a in-use
scanout range never changes on the guest side. That is true for at least
all the current linux drivers.
-Daniel
-- 
Daniel Vetter
Software Engineer, Intel Corporation
http://blog.ffwll.ch
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1276359 — Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel

FromChris Wilson <chris@chris-wilson.co.uk>
Date2015-11-24 13:00 +0100
SubjectRe: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel
Message-ID<qykqm-7Cj-7@gated-at.bofh.it>
In reply to#1276337
On Tue, Nov 24, 2015 at 12:19:18PM +0100, Daniel Vetter wrote:
> Downside: Tracking mapping changes on the guest side won't be any easier.
> This is mostly a problem for integrated gpus, since discrete ones usually
> require contiguous vram for scanout. I think saying "don't do that" is a
> valid option though, i.e. we're assuming that page mappings for a in-use
> scanout range never changes on the guest side. That is true for at least
> all the current linux drivers.

Apart from we already suffer limitations of fixed mappings and have patches
that want to change the page mapping of active scanouts.
-Chris

-- 
Chris Wilson, Intel Open Source Technology Centre
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1276390

FromGerd Hoffmann <kraxel@redhat.com>
Date2015-11-24 13:40 +0100
Message-ID<qyl34-86O-13@gated-at.bofh.it>
In reply to#1276337
  Hi,

> > Yes, vGPU may have additional features, like a framebuffer area, that
> > aren't present or optional for direct assignment.  Obviously we support
> > direct assignment of GPUs for some vendors already without this feature.
> 
> For exposing framebuffers for spice/vnc I highly recommend against
> anything that looks like a bar/fixed mmio range mapping. First this means
> the kernel driver needs to internally fake remapping, which isn't fun.

Sure.  I don't think we should remap here.  More below.

> My recoomendation is to build the actual memory access for underlying
> framebuffers on top of dma-buf, so that it can be vacuumed up by e.g. the
> host gpu driver again for rendering.

We want that too ;)

Some more background:

OpenGL support in qemu is still young and emerging, and we are actually
building on dma-bufs here.  There are a bunch of different ways how
guest display output is handled.  At the end of the day it boils down to
only two fundamental cases though:

  (a) Where qemu doesn't need access to the guest framebuffer
      - qemu directly renders via opengl (works today with virtio-gpu
        and will be in the qemu 2.5 release)
      - qemu passed on the dma-buf to spice client for local display
        (experimental code exists).
      - qemu feeds the guest display into gpu-assisted video encoder
        to send a stream over the network (no code yet).

  (b) Where qemu must read the guest framebuffer.
      - qemu's builtin vnc server.
      - qemu writing screenshots to file.
      - (non-opengl legacy code paths for local display, will
         hopefully disappear long-term though ...)

So, the question is how to support (b) best.  Even with OpenGL support
in qemu improving over time I don't expect this going away completely
anytime soon.

I think it makes sense to have a special vfio region for that.  I don't
think remapping makes sense there.  It doesn't need to be "live", it
doesn't need support high refresh rates.  Placing a copy of the guest
framebuffer there on request (and convert from tiled to linear while
being at it) is perfectly fine.  qemu has a adaptive update rate and
will stop doing frequent update requests when the vnc client
disconnects, so there will be nothing to do if nobody wants actually see
the guest display.

Possible alternative approach would be to import a dma-buf, then use
glReadPixels().  I suspect when doing the copy in the kernel the driver
could ask just the gpu to blit the guest framebuffer.  Don't know gfx
hardware good enough to be sure though, comments are welcome.

cheers,
  Gerd


--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1276458 — Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel

FromDaniel Vetter <daniel@ffwll.ch>
Date2015-11-24 14:40 +0100
SubjectRe: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel
Message-ID<qylZ9-g8-29@gated-at.bofh.it>
In reply to#1276390
On Tue, Nov 24, 2015 at 01:38:55PM +0100, Gerd Hoffmann wrote:
>   Hi,
> 
> > > Yes, vGPU may have additional features, like a framebuffer area, that
> > > aren't present or optional for direct assignment.  Obviously we support
> > > direct assignment of GPUs for some vendors already without this feature.
> > 
> > For exposing framebuffers for spice/vnc I highly recommend against
> > anything that looks like a bar/fixed mmio range mapping. First this means
> > the kernel driver needs to internally fake remapping, which isn't fun.
> 
> Sure.  I don't think we should remap here.  More below.
> 
> > My recoomendation is to build the actual memory access for underlying
> > framebuffers on top of dma-buf, so that it can be vacuumed up by e.g. the
> > host gpu driver again for rendering.
> 
> We want that too ;)
> 
> Some more background:
> 
> OpenGL support in qemu is still young and emerging, and we are actually
> building on dma-bufs here.  There are a bunch of different ways how
> guest display output is handled.  At the end of the day it boils down to
> only two fundamental cases though:
> 
>   (a) Where qemu doesn't need access to the guest framebuffer
>       - qemu directly renders via opengl (works today with virtio-gpu
>         and will be in the qemu 2.5 release)
>       - qemu passed on the dma-buf to spice client for local display
>         (experimental code exists).
>       - qemu feeds the guest display into gpu-assisted video encoder
>         to send a stream over the network (no code yet).
> 
>   (b) Where qemu must read the guest framebuffer.
>       - qemu's builtin vnc server.
>       - qemu writing screenshots to file.
>       - (non-opengl legacy code paths for local display, will
>          hopefully disappear long-term though ...)
> 
> So, the question is how to support (b) best.  Even with OpenGL support
> in qemu improving over time I don't expect this going away completely
> anytime soon.
> 
> I think it makes sense to have a special vfio region for that.  I don't
> think remapping makes sense there.  It doesn't need to be "live", it
> doesn't need support high refresh rates.  Placing a copy of the guest
> framebuffer there on request (and convert from tiled to linear while
> being at it) is perfectly fine.  qemu has a adaptive update rate and
> will stop doing frequent update requests when the vnc client
> disconnects, so there will be nothing to do if nobody wants actually see
> the guest display.
> 
> Possible alternative approach would be to import a dma-buf, then use
> glReadPixels().  I suspect when doing the copy in the kernel the driver
> could ask just the gpu to blit the guest framebuffer.  Don't know gfx
> hardware good enough to be sure though, comments are welcome.

Generally the kernel can't do gpu blts since the required massive state
setup is only in the userspace side of the GL driver stack. But
glReadPixels can do tricks for detiling, and if you use pixel buffer
objects or something similar it'll even be amortized reasonably.

But there's some work to add generic mmap support to dma-bufs, and for
really simple case (where we don't have a gl driver to handle the dma-buf
specially) for untiled framebuffers that would be all we need?
-Daniel
-- 
Daniel Vetter
Software Engineer, Intel Corporation
http://blog.ffwll.ch
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1276509

FromGerd Hoffmann <kraxel@redhat.com>
Date2015-11-24 15:20 +0100
Message-ID<qymBP-JT-3@gated-at.bofh.it>
In reply to#1276458
  Hi,

> But there's some work to add generic mmap support to dma-bufs, and for
> really simple case (where we don't have a gl driver to handle the dma-buf
> specially) for untiled framebuffers that would be all we need?

Not requiring gl is certainly a bonus, people might want build qemu
without opengl support to reduce the attach surface and/or package
dependency chain.

And, yes, requirements for the non-gl rendering path are pretty low.
qemu needs something it can mmap, and which it can ask pixman to handle.
Preferred format is PIXMAN_x8r8g8b8 (qemu uses that internally in alot
of places so this avoids conversions).

Current plan is to have a special vfio region (not visible to the guest)
where the framebuffer lives, with one or two pages at the end for meta
data (format and size).  Status field is there too and will be used by
qemu to request updates and the kernel to signal update completion.
Guess I should write that down as vfio rfc patch ...

I don't think it makes sense to have fields to notify qemu about which
framebuffer regions have been updated, I'd expect with full-screen
composing we have these days this information isn't available anyway.
Maybe a flag telling whenever there have been updates or not, so qemu
can skip update processing in case we have the screensaver showing a
black screen all day long.

cheers,
  Gerd


--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1276512 — Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel

FromDaniel Vetter <daniel@ffwll.ch>
Date2015-11-24 15:20 +0100
SubjectRe: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel
Message-ID<qymBQ-JT-21@gated-at.bofh.it>
In reply to#1276509
On Tue, Nov 24, 2015 at 03:12:31PM +0100, Gerd Hoffmann wrote:
>   Hi,
> 
> > But there's some work to add generic mmap support to dma-bufs, and for
> > really simple case (where we don't have a gl driver to handle the dma-buf
> > specially) for untiled framebuffers that would be all we need?
> 
> Not requiring gl is certainly a bonus, people might want build qemu
> without opengl support to reduce the attach surface and/or package
> dependency chain.
> 
> And, yes, requirements for the non-gl rendering path are pretty low.
> qemu needs something it can mmap, and which it can ask pixman to handle.
> Preferred format is PIXMAN_x8r8g8b8 (qemu uses that internally in alot
> of places so this avoids conversions).
> 
> Current plan is to have a special vfio region (not visible to the guest)
> where the framebuffer lives, with one or two pages at the end for meta
> data (format and size).  Status field is there too and will be used by
> qemu to request updates and the kernel to signal update completion.
> Guess I should write that down as vfio rfc patch ...
> 
> I don't think it makes sense to have fields to notify qemu about which
> framebuffer regions have been updated, I'd expect with full-screen
> composing we have these days this information isn't available anyway.
> Maybe a flag telling whenever there have been updates or not, so qemu
> can skip update processing in case we have the screensaver showing a
> black screen all day long.

GL, wayland, X, EGL and soonish Android's surface flinger (hwc already has
it afaik) all track damage. There's plans to add the same to the atomic
kms api too. But if you do damage tracking you really don't want to
support (maybe allow for perf reasons if the guest is stupid) frontbuffer
rendering, which means you need buffer handles + damage, and not a static
region.
-Daniel
-- 
Daniel Vetter
Software Engineer, Intel Corporation
http://blog.ffwll.ch
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web