Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1272447 > unrolled thread
| Started by | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| First post | 2015-11-18 19:20 +0100 |
| Last post | 2015-11-24 15:20 +0100 |
| Articles | 12 on this page of 32 — 9 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-18 19:20 +0100
RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-19 05:10 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-19 08:30 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Stefano Stabellini <stefano.stabellini@eu.citrix.com> - 2015-11-19 16:40 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-19 17:00 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-20 04:00 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-20 05:30 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-20 07:00 +0100
RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-20 07:10 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-20 17:50 +0100
Re: [Qemu-devel] [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-23 06:00 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Paolo Bonzini <pbonzini@redhat.com> - 2015-11-19 17:00 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Stefano Stabellini <stefano.stabellini@eu.citrix.com> - 2015-11-19 17:20 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Gerd Hoffmann <kraxel@redhat.com> - 2015-11-19 09:50 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Paolo Bonzini <pbonzini@redhat.com> - 2015-11-19 12:10 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-20 03:50 +0100
RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-20 07:20 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Gerd Hoffmann <kraxel@redhat.com> - 2015-11-20 09:30 +0100
RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-20 09:40 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Zhiyuan Lv <zhiyuan.lv@intel.com> - 2015-11-20 10:00 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-19 21:10 +0100
RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-20 08:20 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-20 18:10 +0100
RE: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel "Tian, Kevin" <kevin.tian@intel.com> - 2015-11-20 09:20 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Alex Williamson <alex.williamson@redhat.com> - 2015-11-20 18:30 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Jike Song <jike.song@intel.com> - 2015-11-23 06:10 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Daniel Vetter <daniel@ffwll.ch> - 2015-11-24 12:20 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Chris Wilson <chris@chris-wilson.co.uk> - 2015-11-24 13:00 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Gerd Hoffmann <kraxel@redhat.com> - 2015-11-24 13:40 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Daniel Vetter <daniel@ffwll.ch> - 2015-11-24 14:40 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Gerd Hoffmann <kraxel@redhat.com> - 2015-11-24 15:20 +0100
Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel Daniel Vetter <daniel@ffwll.ch> - 2015-11-24 15:20 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2015-11-19 21:10 +0100 |
| Message-ID | <qwDGN-5Ex-15@gated-at.bofh.it> |
| In reply to | #1272827 |
Hi Kevin, On Thu, 2015-11-19 at 04:06 +0000, Tian, Kevin wrote: > > From: Alex Williamson [mailto:alex.williamson@redhat.com] > > Sent: Thursday, November 19, 2015 2:12 AM > > > > [cc +qemu-devel, +paolo, +gerd] > > > > On Tue, 2015-10-27 at 17:25 +0800, Jike Song wrote: > > > Hi all, > > > > > > We are pleased to announce another update of Intel GVT-g for Xen. > > > > > > Intel GVT-g is a full GPU virtualization solution with mediated > > > pass-through, starting from 4th generation Intel Core(TM) processors > > > with Intel Graphics processors. A virtual GPU instance is maintained > > > for each VM, with part of performance critical resources directly > > > assigned. The capability of running native graphics driver inside a > > > VM, without hypervisor intervention in performance critical paths, > > > achieves a good balance among performance, feature, and sharing > > > capability. Xen is currently supported on Intel Processor Graphics > > > (a.k.a. XenGT); and the core logic can be easily ported to other > > > hypervisors. > > > > > > > > > Repositories > > > > > > Kernel: https://github.com/01org/igvtg-kernel (2015q3-3.18.0 branch) > > > Xen: https://github.com/01org/igvtg-xen (2015q3-4.5 branch) > > > Qemu: https://github.com/01org/igvtg-qemu (xengt_public2015q3 branch) > > > > > > > > > This update consists of: > > > > > > - XenGT is now merged with KVMGT in unified repositories(kernel and qemu), but > > currently > > > different branches for qemu. XenGT and KVMGT share same iGVT-g core logic. > > > > Hi! > > > > At redhat we've been thinking about how to support vGPUs from multiple > > vendors in a common way within QEMU. We want to enable code sharing > > between vendors and give new vendors an easy path to add their own > > support. We also have the complication that not all vGPU vendors are as > > open source friendly as Intel, so being able to abstract the device > > mediation and access outside of QEMU is a big advantage. > > > > The proposal I'd like to make is that a vGPU, whether it is from Intel > > or another vendor, is predominantly a PCI(e) device. We have an > > interface in QEMU already for exposing arbitrary PCI devices, vfio-pci. > > Currently vfio-pci uses the VFIO API to interact with "physical" devices > > and system IOMMUs. I highlight /physical/ there because some of these > > physical devices are SR-IOV VFs, which is somewhat of a fuzzy concept, > > somewhere between fixed hardware and a virtual device implemented in > > software. That software just happens to be running on the physical > > endpoint. > > Agree. > > One clarification for rest discussion, is that we're talking about GVT-g vGPU > here which is a pure software GPU virtualization technique. GVT-d (note > some use in the text) refers to passing through the whole GPU or a specific > VF. GVT-d already falls into existing VFIO APIs nicely (though some on-going > effort to remove Intel specific platform stickness from gfx driver). :-) > > > > > vGPUs are similar, with the virtual device created at a different point, > > host software. They also rely on different IOMMU constructs, making use > > of the MMU capabilities of the GPU (GTTs and such), but really having > > similar requirements. > > One important difference between system IOMMU and GPU-MMU here. > System IOMMU is very much about translation from a DMA target > (IOVA on native, or GPA in virtualization case) to HPA. However GPU > internal MMUs is to translate from Graphics Memory Address (GMA) > to DMA target (HPA if system IOMMU is disabled, or IOVA/GPA if system > IOMMU is enabled). GMA is an internal addr space within GPU, not > exposed to Qemu and fully managed by GVT-g device model. Since it's > not a standard PCI defined resource, we don't need abstract this capability > in VFIO interface. > > > > > The proposal is therefore that GPU vendors can expose vGPUs to > > userspace, and thus to QEMU, using the VFIO API. For instance, vfio > > supports modular bus drivers and IOMMU drivers. An intel-vfio-gvt-d > > module (or extension of i915) can register as a vfio bus driver, create > > a struct device per vGPU, create an IOMMU group for that device, and > > register that device with the vfio-core. Since we don't rely on the > > system IOMMU for GVT-d vGPU assignment, another vGPU vendor driver (or > > extension of the same module) can register a "type1" compliant IOMMU > > driver into vfio-core. From the perspective of QEMU then, all of the > > existing vfio-pci code is re-used, QEMU remains largely unaware of any > > specifics of the vGPU being assigned, and the only necessary change so > > far is how QEMU traverses sysfs to find the device and thus the IOMMU > > group leading to the vfio group. > > GVT-g requires to pin guest memory and query GPA->HPA information, > upon which shadow GTTs will be updated accordingly from (GMA->GPA) > to (GMA->HPA). So yes, here a dummy or simple "type1" compliant IOMMU > can be introduced just for this requirement. > > However there's one tricky point which I'm not sure whether overall > VFIO concept will be violated. GVT-g doesn't require system IOMMU > to function, however host system may enable system IOMMU just for > hardening purpose. This means two-level translations existing (GMA-> > IOVA->HPA), so the dummy IOMMU driver has to request system IOMMU > driver to allocate IOVA for VMs and then setup IOVA->HPA mapping > in IOMMU page table. In this case, multiple VM's translations are > multiplexed in one IOMMU page table. > > We might need create some group/sub-group or parent/child concepts > among those IOMMUs for thorough permission control. My thought here is that this is all abstracted through the vGPU IOMMU and device vfio backends. It's the GPU driver itself, or some vfio extension of that driver, mediating access to the device and deciding when to configure GPU MMU mappings. That driver has access to the GPA to HVA translations thanks to the type1 complaint IOMMU it implements and can pin pages as needed to create GPA to HPA mappings. That should give it all the pieces it needs to fully setup mappings for the vGPU. Whether or not there's a system IOMMU is simply an exercise for that driver. It needs to do a DMA mapping operation through the system IOMMU the same for a vGPU as if it was doing it for itself, because they are in fact one in the same. The GMA to IOVA mapping seems like an internal detail. I assume the IOVA is some sort of GPA, and the GMA is managed through mediation of the device. > > There are a few areas where we know we'll need to extend the VFIO API to > > make this work, but it seems like they can all be done generically. One > > is that PCI BARs are described through the VFIO API as regions and each > > region has a single flag describing whether mmap (ie. direct mapping) of > > that region is possible. We expect that vGPUs likely need finer > > granularity, enabling some areas within a BAR to be trapped and fowarded > > as a read or write access for the vGPU-vfio-device module to emulate, > > while other regions, like framebuffers or texture regions, are directly > > mapped. I have prototype code to enable this already. > > Yes in GVT-g one BAR resource might be partitioned among multiple vGPUs. > If VFIO can support such partial resource assignment, it'd be great. Similar > parent/child concept might also be required here, so any resource enumerated > on a vGPU shouldn't break limitations enforced on the physical device. To be clear, I'm talking about partitioning of the BAR exposed to the guest. Partitioning of the physical BAR would be managed by the vGPU vfio device driver. For instance when the guest mmap's a section of the virtual BAR, the vGPU device driver would map that to a portion of the physical device BAR. > One unique requirement for GVT-g here, though, is that vGPU device model > need to know guest BAR configuration for proper emulation (e.g. register > IO emulation handler to KVM). Similar is about guest MSI vector for virtual > interrupt injection. Not sure how this can be fit into common VFIO model. > Does VFIO allow vendor specific extension today? As a vfio device driver all config accesses and interrupt configuration would be forwarded to you, so I don't see this being a problem. > > > > Another area is that we really don't want to proliferate each vGPU > > needing a new IOMMU type within vfio. The existing type1 IOMMU provides > > potentially the most simple mapping and unmapping interface possible. > > We'd therefore need to allow multiple "type1" IOMMU drivers for vfio, > > making type1 be more of an interface specification rather than a single > > implementation. This is a trivial change to make within vfio and one > > that I believe is compatible with the existing API. Note that > > implementing a type1-compliant vfio IOMMU does not imply pinning an > > mapping every registered page. A vGPU, with mediated device access, may > > use this only to track the current HVA to GPA mappings for a VM. Only > > when a DMA is enabled for the vGPU instance is that HVA pinned and an > > HPA to GPA translation programmed into the GPU MMU. > > > > Another area of extension is how to expose a framebuffer to QEMU for > > seamless integration into a SPICE/VNC channel. For this I believe we > > could use a new region, much like we've done to expose VGA access > > through a vfio device file descriptor. An area within this new > > framebuffer region could be directly mappable in QEMU while a > > non-mappable page, at a standard location with standardized format, > > provides a description of framebuffer and potentially even a > > communication channel to synchronize framebuffer captures. This would > > be new code for QEMU, but something we could share among all vGPU > > implementations. > > Now GVT-g already provides an interface to decode framebuffer information, > w/ an assumption that the framebuffer will be further composited into > OpenGL APIs. So the format is defined according to OpenGL definition. > Does that meet SPICE requirement? > > Another thing to be added. Framebuffers are frequently switched in > reality. So either Qemu needs to poll or a notification mechanism is required. > And since it's dynamic, having framebuffer page directly exposed in the > new region might be tricky. We can just expose framebuffer information > (including base, format, etc.) and let Qemu to map separately out of VFIO > interface. Sure, we'll need to work out that interface, but it's also possible that the framebuffer region is simply remapped to another area of the device (ie. multiple interfaces mapping the same thing) by the vfio device driver. Whether it's easier to do that or make the framebuffer region reference another region is something we'll need to see. > And... this works fine with vGPU model since software knows all the > detail about framebuffer. However in pass-through case, who do you expect > to provide that information? Is it OK to introduce vGPU specific APIs in > VFIO? Yes, vGPU may have additional features, like a framebuffer area, that aren't present or optional for direct assignment. Obviously we support direct assignment of GPUs for some vendors already without this feature. > > Another obvious area to be standardized would be how to discover, > > create, and destroy vGPU instances. SR-IOV has a standard mechanism to > > create VFs in sysfs and I would propose that vGPU vendors try to > > standardize on similar interfaces to enable libvirt to easily discover > > the vGPU capabilities of a given GPU and manage the lifecycle of a vGPU > > instance. > > Now there is no standard. We expose vGPU life-cycle mgmt. APIs through > sysfs (under i915 node), which is very Intel specific. In reality different > vendors have quite different capabilities for their own vGPUs, so not sure > how standard we can define such a mechanism. But this code should be > minor to be maintained in libvirt. Every difference is a barrier. I imagine we can come up with some basic interfaces that everyone could use, even if they don't allow fine tuning every detail specific to a vendor. > > This is obviously a lot to digest, but I'd certainly be interested in > > hearing feedback on this proposal as well as try to clarify anything > > I've left out or misrepresented above. Another benefit to this > > mechanism is that direct GPU assignment and vGPU assignment use the same > > code within QEMU and same API to the kernel, which should make debugging > > and code support between the two easier. I'd really like to start a > > discussion around this proposal, and of course the first open source > > implementation of this sort of model will really help to drive the > > direction it takes. Thanks! > > > > Thanks for starting this discussion. Intel will definitely work with > community on this work. Based on earlier comments, I'm not sure > whether we can exactly same code for direct GPU assignment and > vGPU assignment, since even we extend VFIO some interfaces might > be vGPU specific. Does this way still achieve your end goal? The backends will certainly be different for vGPU vs direct assignment, but hopefully the QEMU code is almost entirely reused, modulo some features like framebuffers that are likely only to be seen on vGPU. Thanks, Alex -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Tian, Kevin" <kevin.tian@intel.com> |
|---|---|
| Date | 2015-11-20 08:20 +0100 |
| Message-ID | <qwO9b-44h-7@gated-at.bofh.it> |
| In reply to | #1273447 |
PiBGcm9tOiBBbGV4IFdpbGxpYW1zb24gW21haWx0bzphbGV4LndpbGxpYW1zb25AcmVkaGF0LmNv bV0NCj4gU2VudDogRnJpZGF5LCBOb3ZlbWJlciAyMCwgMjAxNSA0OjAzIEFNDQo+IA0KPiA+ID4N Cj4gPiA+IFRoZSBwcm9wb3NhbCBpcyB0aGVyZWZvcmUgdGhhdCBHUFUgdmVuZG9ycyBjYW4gZXhw b3NlIHZHUFVzIHRvDQo+ID4gPiB1c2Vyc3BhY2UsIGFuZCB0aHVzIHRvIFFFTVUsIHVzaW5nIHRo ZSBWRklPIEFQSS4gIEZvciBpbnN0YW5jZSwgdmZpbw0KPiA+ID4gc3VwcG9ydHMgbW9kdWxhciBi dXMgZHJpdmVycyBhbmQgSU9NTVUgZHJpdmVycy4gIEFuIGludGVsLXZmaW8tZ3Z0LWQNCj4gPiA+ IG1vZHVsZSAob3IgZXh0ZW5zaW9uIG9mIGk5MTUpIGNhbiByZWdpc3RlciBhcyBhIHZmaW8gYnVz IGRyaXZlciwgY3JlYXRlDQo+ID4gPiBhIHN0cnVjdCBkZXZpY2UgcGVyIHZHUFUsIGNyZWF0ZSBh biBJT01NVSBncm91cCBmb3IgdGhhdCBkZXZpY2UsIGFuZA0KPiA+ID4gcmVnaXN0ZXIgdGhhdCBk ZXZpY2Ugd2l0aCB0aGUgdmZpby1jb3JlLiAgU2luY2Ugd2UgZG9uJ3QgcmVseSBvbiB0aGUNCj4g PiA+IHN5c3RlbSBJT01NVSBmb3IgR1ZULWQgdkdQVSBhc3NpZ25tZW50LCBhbm90aGVyIHZHUFUg dmVuZG9yIGRyaXZlciAob3INCj4gPiA+IGV4dGVuc2lvbiBvZiB0aGUgc2FtZSBtb2R1bGUpIGNh biByZWdpc3RlciBhICJ0eXBlMSIgY29tcGxpYW50IElPTU1VDQo+ID4gPiBkcml2ZXIgaW50byB2 ZmlvLWNvcmUuICBGcm9tIHRoZSBwZXJzcGVjdGl2ZSBvZiBRRU1VIHRoZW4sIGFsbCBvZiB0aGUN Cj4gPiA+IGV4aXN0aW5nIHZmaW8tcGNpIGNvZGUgaXMgcmUtdXNlZCwgUUVNVSByZW1haW5zIGxh cmdlbHkgdW5hd2FyZSBvZiBhbnkNCj4gPiA+IHNwZWNpZmljcyBvZiB0aGUgdkdQVSBiZWluZyBh c3NpZ25lZCwgYW5kIHRoZSBvbmx5IG5lY2Vzc2FyeSBjaGFuZ2Ugc28NCj4gPiA+IGZhciBpcyBo b3cgUUVNVSB0cmF2ZXJzZXMgc3lzZnMgdG8gZmluZCB0aGUgZGV2aWNlIGFuZCB0aHVzIHRoZSBJ T01NVQ0KPiA+ID4gZ3JvdXAgbGVhZGluZyB0byB0aGUgdmZpbyBncm91cC4NCj4gPg0KPiA+IEdW VC1nIHJlcXVpcmVzIHRvIHBpbiBndWVzdCBtZW1vcnkgYW5kIHF1ZXJ5IEdQQS0+SFBBIGluZm9y bWF0aW9uLA0KPiA+IHVwb24gd2hpY2ggc2hhZG93IEdUVHMgd2lsbCBiZSB1cGRhdGVkIGFjY29y ZGluZ2x5IGZyb20gKEdNQS0+R1BBKQ0KPiA+IHRvIChHTUEtPkhQQSkuIFNvIHllcywgaGVyZSBh IGR1bW15IG9yIHNpbXBsZSAidHlwZTEiIGNvbXBsaWFudCBJT01NVQ0KPiA+IGNhbiBiZSBpbnRy b2R1Y2VkIGp1c3QgZm9yIHRoaXMgcmVxdWlyZW1lbnQuDQo+ID4NCj4gPiBIb3dldmVyIHRoZXJl J3Mgb25lIHRyaWNreSBwb2ludCB3aGljaCBJJ20gbm90IHN1cmUgd2hldGhlciBvdmVyYWxsDQo+ ID4gVkZJTyBjb25jZXB0IHdpbGwgYmUgdmlvbGF0ZWQuIEdWVC1nIGRvZXNuJ3QgcmVxdWlyZSBz eXN0ZW0gSU9NTVUNCj4gPiB0byBmdW5jdGlvbiwgaG93ZXZlciBob3N0IHN5c3RlbSBtYXkgZW5h YmxlIHN5c3RlbSBJT01NVSBqdXN0IGZvcg0KPiA+IGhhcmRlbmluZyBwdXJwb3NlLiBUaGlzIG1l YW5zIHR3by1sZXZlbCB0cmFuc2xhdGlvbnMgZXhpc3RpbmcgKEdNQS0+DQo+ID4gSU9WQS0+SFBB KSwgc28gdGhlIGR1bW15IElPTU1VIGRyaXZlciBoYXMgdG8gcmVxdWVzdCBzeXN0ZW0gSU9NTVUN Cj4gPiBkcml2ZXIgdG8gYWxsb2NhdGUgSU9WQSBmb3IgVk1zIGFuZCB0aGVuIHNldHVwIElPVkEt PkhQQSBtYXBwaW5nDQo+ID4gaW4gSU9NTVUgcGFnZSB0YWJsZS4gSW4gdGhpcyBjYXNlLCBtdWx0 aXBsZSBWTSdzIHRyYW5zbGF0aW9ucyBhcmUNCj4gPiBtdWx0aXBsZXhlZCBpbiBvbmUgSU9NTVUg cGFnZSB0YWJsZS4NCj4gPg0KPiA+IFdlIG1pZ2h0IG5lZWQgY3JlYXRlIHNvbWUgZ3JvdXAvc3Vi LWdyb3VwIG9yIHBhcmVudC9jaGlsZCBjb25jZXB0cw0KPiA+IGFtb25nIHRob3NlIElPTU1VcyBm b3IgdGhvcm91Z2ggcGVybWlzc2lvbiBjb250cm9sLg0KPiANCj4gTXkgdGhvdWdodCBoZXJlIGlz IHRoYXQgdGhpcyBpcyBhbGwgYWJzdHJhY3RlZCB0aHJvdWdoIHRoZSB2R1BVIElPTU1VDQo+IGFu ZCBkZXZpY2UgdmZpbyBiYWNrZW5kcy4gIEl0J3MgdGhlIEdQVSBkcml2ZXIgaXRzZWxmLCBvciBz b21lIHZmaW8NCj4gZXh0ZW5zaW9uIG9mIHRoYXQgZHJpdmVyLCBtZWRpYXRpbmcgYWNjZXNzIHRv IHRoZSBkZXZpY2UgYW5kIGRlY2lkaW5nDQo+IHdoZW4gdG8gY29uZmlndXJlIEdQVSBNTVUgbWFw cGluZ3MuICBUaGF0IGRyaXZlciBoYXMgYWNjZXNzIHRvIHRoZSBHUEENCj4gdG8gSFZBIHRyYW5z bGF0aW9ucyB0aGFua3MgdG8gdGhlIHR5cGUxIGNvbXBsYWludCBJT01NVSBpdCBpbXBsZW1lbnRz DQo+IGFuZCBjYW4gcGluIHBhZ2VzIGFzIG5lZWRlZCB0byBjcmVhdGUgR1BBIHRvIEhQQSBtYXBw aW5ncy4gIFRoYXQgc2hvdWxkDQo+IGdpdmUgaXQgYWxsIHRoZSBwaWVjZXMgaXQgbmVlZHMgdG8g ZnVsbHkgc2V0dXAgbWFwcGluZ3MgZm9yIHRoZSB2R1BVLg0KPiBXaGV0aGVyIG9yIG5vdCB0aGVy ZSdzIGEgc3lzdGVtIElPTU1VIGlzIHNpbXBseSBhbiBleGVyY2lzZSBmb3IgdGhhdA0KPiBkcml2 ZXIuICBJdCBuZWVkcyB0byBkbyBhIERNQSBtYXBwaW5nIG9wZXJhdGlvbiB0aHJvdWdoIHRoZSBz eXN0ZW0gSU9NTVUNCj4gdGhlIHNhbWUgZm9yIGEgdkdQVSBhcyBpZiBpdCB3YXMgZG9pbmcgaXQg Zm9yIGl0c2VsZiwgYmVjYXVzZSB0aGV5IGFyZQ0KPiBpbiBmYWN0IG9uZSBpbiB0aGUgc2FtZS4g IFRoZSBHTUEgdG8gSU9WQSBtYXBwaW5nIHNlZW1zIGxpa2UgYW4gaW50ZXJuYWwNCj4gZGV0YWls LiAgSSBhc3N1bWUgdGhlIElPVkEgaXMgc29tZSBzb3J0IG9mIEdQQSwgYW5kIHRoZSBHTUEgaXMg bWFuYWdlZA0KPiB0aHJvdWdoIG1lZGlhdGlvbiBvZiB0aGUgZGV2aWNlLg0KDQpTb3JyeSBJJ20g bm90IGZhbWlsaWFyIHdpdGggVkZJTyBpbnRlcm5hbC4gTXkgb3JpZ2luYWwgd29ycnkgaXMgdGhh dCBzeXN0ZW0gDQpJT01NVSBmb3IgR1BVIG1heSBiZSBhbHJlYWR5IGNsYWltZWQgYnkgYW5vdGhl ciB2ZmlvIGRyaXZlciAoZS5nLiBob3N0IGtlcm5lbA0Kd2FudHMgdG8gaGFyZGVuIGdmeCBkcml2 ZXIgZnJvbSByZXN0IHN1Yi1zeXN0ZW1zLCByZWdhcmRsZXNzIG9mIHdoZXRoZXIgdkdQVSANCmlz IGNyZWF0ZWQgb3Igbm90KS4gSW4gdGhhdCBjYXNlIHZHUFUgSU9NTVUgZHJpdmVyIHNob3VsZG4n dCBtYW5hZ2Ugc3lzdGVtDQpJT01NVSBkaXJlY3RseS4NCg0KYnR3LCBjdXJpb3VzIHRvZGF5IGhv dyBWRklPIGNvb3JkaW5hdGVzIHdpdGggc3lzdGVtIElPTU1VIGRyaXZlciByZWdhcmRpbmcNCnRv IHdoZXRoZXIgYSBJT01NVSBpcyB1c2VkIHRvIGNvbnRyb2wgZGV2aWNlIGFzc2lnbm1lbnQsIG9y IHVzZWQgZm9yIGtlcm5lbCANCmhhcmRlbmluZy4gU29tZWhvdyB0d28gYXJlIGNvbmZsaWN0aW5n IHNpbmNlIGRpZmZlcmVudCBhZGRyZXNzIHNwYWNlcyBhcmUNCmNvbmNlcm5lZCAoR1BBIHZzLiBJ T1ZBKS4uLg0KDQo+IA0KPiANCj4gPiA+IFRoZXJlIGFyZSBhIGZldyBhcmVhcyB3aGVyZSB3ZSBr bm93IHdlJ2xsIG5lZWQgdG8gZXh0ZW5kIHRoZSBWRklPIEFQSSB0bw0KPiA+ID4gbWFrZSB0aGlz IHdvcmssIGJ1dCBpdCBzZWVtcyBsaWtlIHRoZXkgY2FuIGFsbCBiZSBkb25lIGdlbmVyaWNhbGx5 LiAgT25lDQo+ID4gPiBpcyB0aGF0IFBDSSBCQVJzIGFyZSBkZXNjcmliZWQgdGhyb3VnaCB0aGUg VkZJTyBBUEkgYXMgcmVnaW9ucyBhbmQgZWFjaA0KPiA+ID4gcmVnaW9uIGhhcyBhIHNpbmdsZSBm bGFnIGRlc2NyaWJpbmcgd2hldGhlciBtbWFwIChpZS4gZGlyZWN0IG1hcHBpbmcpIG9mDQo+ID4g PiB0aGF0IHJlZ2lvbiBpcyBwb3NzaWJsZS4gIFdlIGV4cGVjdCB0aGF0IHZHUFVzIGxpa2VseSBu ZWVkIGZpbmVyDQo+ID4gPiBncmFudWxhcml0eSwgZW5hYmxpbmcgc29tZSBhcmVhcyB3aXRoaW4g YSBCQVIgdG8gYmUgdHJhcHBlZCBhbmQgZm93YXJkZWQNCj4gPiA+IGFzIGEgcmVhZCBvciB3cml0 ZSBhY2Nlc3MgZm9yIHRoZSB2R1BVLXZmaW8tZGV2aWNlIG1vZHVsZSB0byBlbXVsYXRlLA0KPiA+ ID4gd2hpbGUgb3RoZXIgcmVnaW9ucywgbGlrZSBmcmFtZWJ1ZmZlcnMgb3IgdGV4dHVyZSByZWdp b25zLCBhcmUgZGlyZWN0bHkNCj4gPiA+IG1hcHBlZC4gIEkgaGF2ZSBwcm90b3R5cGUgY29kZSB0 byBlbmFibGUgdGhpcyBhbHJlYWR5Lg0KPiA+DQo+ID4gWWVzIGluIEdWVC1nIG9uZSBCQVIgcmVz b3VyY2UgbWlnaHQgYmUgcGFydGl0aW9uZWQgYW1vbmcgbXVsdGlwbGUgdkdQVXMuDQo+ID4gSWYg VkZJTyBjYW4gc3VwcG9ydCBzdWNoIHBhcnRpYWwgcmVzb3VyY2UgYXNzaWdubWVudCwgaXQnZCBi ZSBncmVhdC4gU2ltaWxhcg0KPiA+IHBhcmVudC9jaGlsZCBjb25jZXB0IG1pZ2h0IGFsc28gYmUg cmVxdWlyZWQgaGVyZSwgc28gYW55IHJlc291cmNlIGVudW1lcmF0ZWQNCj4gPiBvbiBhIHZHUFUg c2hvdWxkbid0IGJyZWFrIGxpbWl0YXRpb25zIGVuZm9yY2VkIG9uIHRoZSBwaHlzaWNhbCBkZXZp Y2UuDQo+IA0KPiBUbyBiZSBjbGVhciwgSSdtIHRhbGtpbmcgYWJvdXQgcGFydGl0aW9uaW5nIG9m IHRoZSBCQVIgZXhwb3NlZCB0byB0aGUNCj4gZ3Vlc3QuICBQYXJ0aXRpb25pbmcgb2YgdGhlIHBo eXNpY2FsIEJBUiB3b3VsZCBiZSBtYW5hZ2VkIGJ5IHRoZSB2R1BVDQo+IHZmaW8gZGV2aWNlIGRy aXZlci4gIEZvciBpbnN0YW5jZSB3aGVuIHRoZSBndWVzdCBtbWFwJ3MgYSBzZWN0aW9uIG9mIHRo ZQ0KPiB2aXJ0dWFsIEJBUiwgdGhlIHZHUFUgZGV2aWNlIGRyaXZlciB3b3VsZCBtYXAgdGhhdCB0 byBhIHBvcnRpb24gb2YgdGhlDQo+IHBoeXNpY2FsIGRldmljZSBCQVIuDQo+IA0KPiA+IE9uZSB1 bmlxdWUgcmVxdWlyZW1lbnQgZm9yIEdWVC1nIGhlcmUsIHRob3VnaCwgaXMgdGhhdCB2R1BVIGRl dmljZSBtb2RlbA0KPiA+IG5lZWQgdG8ga25vdyBndWVzdCBCQVIgY29uZmlndXJhdGlvbiBmb3Ig cHJvcGVyIGVtdWxhdGlvbiAoZS5nLiByZWdpc3Rlcg0KPiA+IElPIGVtdWxhdGlvbiBoYW5kbGVy IHRvIEtWTSkuIFNpbWlsYXIgaXMgYWJvdXQgZ3Vlc3QgTVNJIHZlY3RvciBmb3IgdmlydHVhbA0K PiA+IGludGVycnVwdCBpbmplY3Rpb24uIE5vdCBzdXJlIGhvdyB0aGlzIGNhbiBiZSBmaXQgaW50 byBjb21tb24gVkZJTyBtb2RlbC4NCj4gPiBEb2VzIFZGSU8gYWxsb3cgdmVuZG9yIHNwZWNpZmlj IGV4dGVuc2lvbiB0b2RheT8NCj4gDQo+IEFzIGEgdmZpbyBkZXZpY2UgZHJpdmVyIGFsbCBjb25m aWcgYWNjZXNzZXMgYW5kIGludGVycnVwdCBjb25maWd1cmF0aW9uDQo+IHdvdWxkIGJlIGZvcndh cmRlZCB0byB5b3UsIHNvIEkgZG9uJ3Qgc2VlIHRoaXMgYmVpbmcgYSBwcm9ibGVtLg0KDQpTdXJl LCBuaWNlIHRvIGtub3cgdGhhdC4NCg0KVGhhbmtzDQpLZXZpbg0K -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2015-11-20 18:10 +0100 |
| Message-ID | <qwXm9-1Fh-11@gated-at.bofh.it> |
| In reply to | #1273752 |
On Fri, 2015-11-20 at 07:09 +0000, Tian, Kevin wrote: > > From: Alex Williamson [mailto:alex.williamson@redhat.com] > > Sent: Friday, November 20, 2015 4:03 AM > > > > > > > > > > The proposal is therefore that GPU vendors can expose vGPUs to > > > > userspace, and thus to QEMU, using the VFIO API. For instance, vfio > > > > supports modular bus drivers and IOMMU drivers. An intel-vfio-gvt-d > > > > module (or extension of i915) can register as a vfio bus driver, create > > > > a struct device per vGPU, create an IOMMU group for that device, and > > > > register that device with the vfio-core. Since we don't rely on the > > > > system IOMMU for GVT-d vGPU assignment, another vGPU vendor driver (or > > > > extension of the same module) can register a "type1" compliant IOMMU > > > > driver into vfio-core. From the perspective of QEMU then, all of the > > > > existing vfio-pci code is re-used, QEMU remains largely unaware of any > > > > specifics of the vGPU being assigned, and the only necessary change so > > > > far is how QEMU traverses sysfs to find the device and thus the IOMMU > > > > group leading to the vfio group. > > > > > > GVT-g requires to pin guest memory and query GPA->HPA information, > > > upon which shadow GTTs will be updated accordingly from (GMA->GPA) > > > to (GMA->HPA). So yes, here a dummy or simple "type1" compliant IOMMU > > > can be introduced just for this requirement. > > > > > > However there's one tricky point which I'm not sure whether overall > > > VFIO concept will be violated. GVT-g doesn't require system IOMMU > > > to function, however host system may enable system IOMMU just for > > > hardening purpose. This means two-level translations existing (GMA-> > > > IOVA->HPA), so the dummy IOMMU driver has to request system IOMMU > > > driver to allocate IOVA for VMs and then setup IOVA->HPA mapping > > > in IOMMU page table. In this case, multiple VM's translations are > > > multiplexed in one IOMMU page table. > > > > > > We might need create some group/sub-group or parent/child concepts > > > among those IOMMUs for thorough permission control. > > > > My thought here is that this is all abstracted through the vGPU IOMMU > > and device vfio backends. It's the GPU driver itself, or some vfio > > extension of that driver, mediating access to the device and deciding > > when to configure GPU MMU mappings. That driver has access to the GPA > > to HVA translations thanks to the type1 complaint IOMMU it implements > > and can pin pages as needed to create GPA to HPA mappings. That should > > give it all the pieces it needs to fully setup mappings for the vGPU. > > Whether or not there's a system IOMMU is simply an exercise for that > > driver. It needs to do a DMA mapping operation through the system IOMMU > > the same for a vGPU as if it was doing it for itself, because they are > > in fact one in the same. The GMA to IOVA mapping seems like an internal > > detail. I assume the IOVA is some sort of GPA, and the GMA is managed > > through mediation of the device. > > Sorry I'm not familiar with VFIO internal. My original worry is that system > IOMMU for GPU may be already claimed by another vfio driver (e.g. host kernel > wants to harden gfx driver from rest sub-systems, regardless of whether vGPU > is created or not). In that case vGPU IOMMU driver shouldn't manage system > IOMMU directly. There are different APIs for the IOMMU depending on how it's being use. If the IOMMU is being used for inter-device isolation in the host, then the DMA API (ex. dma_map_page) transparently makes use of the IOMMU. When we're doing device assignment, we make use of the IOMMU API which allows more explicit control (ex. iommu_domain_alloc, iommu_attach_device, iommu_map, etc). A vGPU is not an SR-IOV VF, it doesn't have a unique requester ID that allows the IOMMU to differentiate one vGPU from another, or vGPU from GPU. All mappings for vGPUs need to occur for the GPU. It's therefore the responsibility of the GPU driver, or this vfio extension of that driver, that needs to perform the IOMMU mapping for the vGPU. My expectation is therefore that once the GMA to IOVA mapping is configured in the GPU MMU, the IOVA to HPA needs to be programmed, as if the GPU driver was performing the setup itself, which it is. Before the device mediation that triggered the mapping setup is complete, the GPU MMU and the system IOMMU (if preset) should be configured to enable that DMA. The GPU MMU provides the isolation of the vGPU, the system IOMMU enable the DMA to occur. > btw, curious today how VFIO coordinates with system IOMMU driver regarding > to whether a IOMMU is used to control device assignment, or used for kernel > hardening. Somehow two are conflicting since different address spaces are > concerned (GPA vs. IOVA)... When devices unbind from native host drivers, any previous IOMMU mappings and domains are removed. These are typically created via the DMA API above. The initialization operations of the VFIO API (creating containers, attaching groups to containers, and setting the IOMMU model for a container) work through the IOMMU API to create a new domain and isolate devices within it. The type1 VFIO IOMMU interface is then effectively a passthrough to the iommu_map() and iommu_unmap() interfaces of the IOMMU API, modulo page pinning, accounting and tracking. When a VFIO instance is destroyed, the devices are detached from the IOMMU domain, the devices are unbound from vfio and re-bound to host drivers and the DMA API can reclaim the devices for host isolation. Thanks, Alex -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Tian, Kevin" <kevin.tian@intel.com> |
|---|---|
| Date | 2015-11-20 09:20 +0100 |
| Message-ID | <qwP5f-4EO-5@gated-at.bofh.it> |
| In reply to | #1273447 |
PiBGcm9tOiBUaWFuLCBLZXZpbg0KPiBTZW50OiBGcmlkYXksIE5vdmVtYmVyIDIwLCAyMDE1IDM6 MTAgUE0NCg0KPiA+ID4gPg0KPiA+ID4gPiBUaGUgcHJvcG9zYWwgaXMgdGhlcmVmb3JlIHRoYXQg R1BVIHZlbmRvcnMgY2FuIGV4cG9zZSB2R1BVcyB0bw0KPiA+ID4gPiB1c2Vyc3BhY2UsIGFuZCB0 aHVzIHRvIFFFTVUsIHVzaW5nIHRoZSBWRklPIEFQSS4gIEZvciBpbnN0YW5jZSwgdmZpbw0KPiA+ ID4gPiBzdXBwb3J0cyBtb2R1bGFyIGJ1cyBkcml2ZXJzIGFuZCBJT01NVSBkcml2ZXJzLiAgQW4g aW50ZWwtdmZpby1ndnQtZA0KPiA+ID4gPiBtb2R1bGUgKG9yIGV4dGVuc2lvbiBvZiBpOTE1KSBj YW4gcmVnaXN0ZXIgYXMgYSB2ZmlvIGJ1cyBkcml2ZXIsIGNyZWF0ZQ0KPiA+ID4gPiBhIHN0cnVj dCBkZXZpY2UgcGVyIHZHUFUsIGNyZWF0ZSBhbiBJT01NVSBncm91cCBmb3IgdGhhdCBkZXZpY2Us IGFuZA0KPiA+ID4gPiByZWdpc3RlciB0aGF0IGRldmljZSB3aXRoIHRoZSB2ZmlvLWNvcmUuICBT aW5jZSB3ZSBkb24ndCByZWx5IG9uIHRoZQ0KPiA+ID4gPiBzeXN0ZW0gSU9NTVUgZm9yIEdWVC1k IHZHUFUgYXNzaWdubWVudCwgYW5vdGhlciB2R1BVIHZlbmRvciBkcml2ZXIgKG9yDQo+ID4gPiA+ IGV4dGVuc2lvbiBvZiB0aGUgc2FtZSBtb2R1bGUpIGNhbiByZWdpc3RlciBhICJ0eXBlMSIgY29t cGxpYW50IElPTU1VDQo+ID4gPiA+IGRyaXZlciBpbnRvIHZmaW8tY29yZS4gIEZyb20gdGhlIHBl cnNwZWN0aXZlIG9mIFFFTVUgdGhlbiwgYWxsIG9mIHRoZQ0KPiA+ID4gPiBleGlzdGluZyB2Zmlv LXBjaSBjb2RlIGlzIHJlLXVzZWQsIFFFTVUgcmVtYWlucyBsYXJnZWx5IHVuYXdhcmUgb2YgYW55 DQo+ID4gPiA+IHNwZWNpZmljcyBvZiB0aGUgdkdQVSBiZWluZyBhc3NpZ25lZCwgYW5kIHRoZSBv bmx5IG5lY2Vzc2FyeSBjaGFuZ2Ugc28NCj4gPiA+ID4gZmFyIGlzIGhvdyBRRU1VIHRyYXZlcnNl cyBzeXNmcyB0byBmaW5kIHRoZSBkZXZpY2UgYW5kIHRodXMgdGhlIElPTU1VDQo+ID4gPiA+IGdy b3VwIGxlYWRpbmcgdG8gdGhlIHZmaW8gZ3JvdXAuDQo+ID4gPg0KPiA+ID4gR1ZULWcgcmVxdWly ZXMgdG8gcGluIGd1ZXN0IG1lbW9yeSBhbmQgcXVlcnkgR1BBLT5IUEEgaW5mb3JtYXRpb24sDQo+ ID4gPiB1cG9uIHdoaWNoIHNoYWRvdyBHVFRzIHdpbGwgYmUgdXBkYXRlZCBhY2NvcmRpbmdseSBm cm9tIChHTUEtPkdQQSkNCj4gPiA+IHRvIChHTUEtPkhQQSkuIFNvIHllcywgaGVyZSBhIGR1bW15 IG9yIHNpbXBsZSAidHlwZTEiIGNvbXBsaWFudCBJT01NVQ0KPiA+ID4gY2FuIGJlIGludHJvZHVj ZWQganVzdCBmb3IgdGhpcyByZXF1aXJlbWVudC4NCj4gPiA+DQo+ID4gPiBIb3dldmVyIHRoZXJl J3Mgb25lIHRyaWNreSBwb2ludCB3aGljaCBJJ20gbm90IHN1cmUgd2hldGhlciBvdmVyYWxsDQo+ ID4gPiBWRklPIGNvbmNlcHQgd2lsbCBiZSB2aW9sYXRlZC4gR1ZULWcgZG9lc24ndCByZXF1aXJl IHN5c3RlbSBJT01NVQ0KPiA+ID4gdG8gZnVuY3Rpb24sIGhvd2V2ZXIgaG9zdCBzeXN0ZW0gbWF5 IGVuYWJsZSBzeXN0ZW0gSU9NTVUganVzdCBmb3INCj4gPiA+IGhhcmRlbmluZyBwdXJwb3NlLiBU aGlzIG1lYW5zIHR3by1sZXZlbCB0cmFuc2xhdGlvbnMgZXhpc3RpbmcgKEdNQS0+DQo+ID4gPiBJ T1ZBLT5IUEEpLCBzbyB0aGUgZHVtbXkgSU9NTVUgZHJpdmVyIGhhcyB0byByZXF1ZXN0IHN5c3Rl bSBJT01NVQ0KPiA+ID4gZHJpdmVyIHRvIGFsbG9jYXRlIElPVkEgZm9yIFZNcyBhbmQgdGhlbiBz ZXR1cCBJT1ZBLT5IUEEgbWFwcGluZw0KPiA+ID4gaW4gSU9NTVUgcGFnZSB0YWJsZS4gSW4gdGhp cyBjYXNlLCBtdWx0aXBsZSBWTSdzIHRyYW5zbGF0aW9ucyBhcmUNCj4gPiA+IG11bHRpcGxleGVk IGluIG9uZSBJT01NVSBwYWdlIHRhYmxlLg0KPiA+ID4NCj4gPiA+IFdlIG1pZ2h0IG5lZWQgY3Jl YXRlIHNvbWUgZ3JvdXAvc3ViLWdyb3VwIG9yIHBhcmVudC9jaGlsZCBjb25jZXB0cw0KPiA+ID4g YW1vbmcgdGhvc2UgSU9NTVVzIGZvciB0aG9yb3VnaCBwZXJtaXNzaW9uIGNvbnRyb2wuDQo+ID4N Cj4gPiBNeSB0aG91Z2h0IGhlcmUgaXMgdGhhdCB0aGlzIGlzIGFsbCBhYnN0cmFjdGVkIHRocm91 Z2ggdGhlIHZHUFUgSU9NTVUNCj4gPiBhbmQgZGV2aWNlIHZmaW8gYmFja2VuZHMuICBJdCdzIHRo ZSBHUFUgZHJpdmVyIGl0c2VsZiwgb3Igc29tZSB2ZmlvDQo+ID4gZXh0ZW5zaW9uIG9mIHRoYXQg ZHJpdmVyLCBtZWRpYXRpbmcgYWNjZXNzIHRvIHRoZSBkZXZpY2UgYW5kIGRlY2lkaW5nDQo+ID4g d2hlbiB0byBjb25maWd1cmUgR1BVIE1NVSBtYXBwaW5ncy4gIFRoYXQgZHJpdmVyIGhhcyBhY2Nl c3MgdG8gdGhlIEdQQQ0KPiA+IHRvIEhWQSB0cmFuc2xhdGlvbnMgdGhhbmtzIHRvIHRoZSB0eXBl MSBjb21wbGFpbnQgSU9NTVUgaXQgaW1wbGVtZW50cw0KPiA+IGFuZCBjYW4gcGluIHBhZ2VzIGFz IG5lZWRlZCB0byBjcmVhdGUgR1BBIHRvIEhQQSBtYXBwaW5ncy4gIFRoYXQgc2hvdWxkDQo+ID4g Z2l2ZSBpdCBhbGwgdGhlIHBpZWNlcyBpdCBuZWVkcyB0byBmdWxseSBzZXR1cCBtYXBwaW5ncyBm b3IgdGhlIHZHUFUuDQo+ID4gV2hldGhlciBvciBub3QgdGhlcmUncyBhIHN5c3RlbSBJT01NVSBp cyBzaW1wbHkgYW4gZXhlcmNpc2UgZm9yIHRoYXQNCj4gPiBkcml2ZXIuICBJdCBuZWVkcyB0byBk byBhIERNQSBtYXBwaW5nIG9wZXJhdGlvbiB0aHJvdWdoIHRoZSBzeXN0ZW0gSU9NTVUNCj4gPiB0 aGUgc2FtZSBmb3IgYSB2R1BVIGFzIGlmIGl0IHdhcyBkb2luZyBpdCBmb3IgaXRzZWxmLCBiZWNh dXNlIHRoZXkgYXJlDQo+ID4gaW4gZmFjdCBvbmUgaW4gdGhlIHNhbWUuICBUaGUgR01BIHRvIElP VkEgbWFwcGluZyBzZWVtcyBsaWtlIGFuIGludGVybmFsDQo+ID4gZGV0YWlsLiAgSSBhc3N1bWUg dGhlIElPVkEgaXMgc29tZSBzb3J0IG9mIEdQQSwgYW5kIHRoZSBHTUEgaXMgbWFuYWdlZA0KPiA+ IHRocm91Z2ggbWVkaWF0aW9uIG9mIHRoZSBkZXZpY2UuDQo+IA0KPiBTb3JyeSBJJ20gbm90IGZh bWlsaWFyIHdpdGggVkZJTyBpbnRlcm5hbC4gTXkgb3JpZ2luYWwgd29ycnkgaXMgdGhhdCBzeXN0 ZW0NCj4gSU9NTVUgZm9yIEdQVSBtYXkgYmUgYWxyZWFkeSBjbGFpbWVkIGJ5IGFub3RoZXIgdmZp byBkcml2ZXIgKGUuZy4gaG9zdCBrZXJuZWwNCj4gd2FudHMgdG8gaGFyZGVuIGdmeCBkcml2ZXIg ZnJvbSByZXN0IHN1Yi1zeXN0ZW1zLCByZWdhcmRsZXNzIG9mIHdoZXRoZXIgdkdQVQ0KPiBpcyBj cmVhdGVkIG9yIG5vdCkuIEluIHRoYXQgY2FzZSB2R1BVIElPTU1VIGRyaXZlciBzaG91bGRuJ3Qg bWFuYWdlIHN5c3RlbQ0KPiBJT01NVSBkaXJlY3RseS4NCj4gDQo+IGJ0dywgY3VyaW91cyB0b2Rh eSBob3cgVkZJTyBjb29yZGluYXRlcyB3aXRoIHN5c3RlbSBJT01NVSBkcml2ZXIgcmVnYXJkaW5n DQo+IHRvIHdoZXRoZXIgYSBJT01NVSBpcyB1c2VkIHRvIGNvbnRyb2wgZGV2aWNlIGFzc2lnbm1l bnQsIG9yIHVzZWQgZm9yIGtlcm5lbA0KPiBoYXJkZW5pbmcuIFNvbWVob3cgdHdvIGFyZSBjb25m bGljdGluZyBzaW5jZSBkaWZmZXJlbnQgYWRkcmVzcyBzcGFjZXMgYXJlDQo+IGNvbmNlcm5lZCAo R1BBIHZzLiBJT1ZBKS4uLg0KPiANCg0KSGVyZSBpcyBhIG1vcmUgY29uY3JldGUgZXhhbXBsZToN Cg0KS1ZNR1QgZG9lc24ndCByZXF1aXJlIElPTU1VLiBBbGwgRE1BIHRhcmdldHMgYXJlIGFscmVh ZHkgcmVwbGFjZWQgd2l0aCANCkhQQSB0aHJ1IHNoYWRvdyBHVFQuIFNvIERNQSByZXF1ZXN0cyBm cm9tIEdQVSBhbGwgY29udGFpbiBIUEFzLg0KDQpXaGVuIElPTU1VIGlzIGVuYWJsZWQsIG9uZSBz aW1wbGUgYXBwcm9hY2ggaXMgdG8gaGF2ZSB2R1BVIElPTU1VDQpkcml2ZXIgY29uZmlndXJlIHN5 c3RlbSBJT01NVSB3aXRoIGlkZW50aXR5IG1hcHBpbmcgKEhQQS0+SFBBKS4gV2UgDQpjYW4ndCB1 c2UgKEdQQS0+SFBBKSBzaW5jZSBHUEFzIGZyb20gbXVsdGlwbGUgVk1zIGFyZSBjb25mbGljdGlu Zy4gDQoNCkhvd2V2ZXIsIHdlIHN0aWxsIGhhdmUgaG9zdCBnZnggZHJpdmVyIHJ1bm5pbmcuIFdo ZW4gSU9NTVUgaXMgZW5hYmxlZCwgDQpkbWFfYWxsb2NfKioqIHdpbGwgcmV0dXJuIElPVkEgKGRy dmVycy9pb21tdS9pb3ZhLmMpIGluIGhvc3QgZ2Z4IGRyaXZlciwNCndoaWNoIHdpbGwgaGF2ZSBJ T1ZBLT5IUEEgcHJvZ3JhbW1lZCB0byBzeXN0ZW0gSU9NTVUuDQoNCk9uZSBJT01NVSBkZXZpY2Ug ZW50cnkgY2FuIG9ubHkgdHJhbnNsYXRlIG9uZSBhZGRyZXNzIHNwYWNlLCBzbyBoZXJlDQpjb21l cyBhIGNvbmZsaWN0IChIUEEtPkhQQSB2cy4gSU9WQS0+SFBBKS4gVG8gc29sdmUgdGhpcywgdkdQ VSBJT01NVQ0KZHJpdmVyIG5lZWRzIHRvIGFsbG9jYXRlIElPVkEgZnJvbSBpb3ZhLmMgZm9yIGVh Y2ggVk0gdy8gdkdQVSBhc3NpZ25lZCwNCmFuZCB0aGVuIEtWTUdUIHdpbGwgcHJvZ3JhbSBJT1ZB IGluIHNoYWRvdyBHVFQgYWNjb3JkaW5nbHkuIEl0IGFkZHMNCm9uZSBhZGRpdGlvbmFsIG1hcHBp bmcgbGF5ZXIgKEdQQS0+SU9WQS0+SFBBKS4gSW4gdGhpcyB3YXkgdHdvIA0KcmVxdWlyZW1lbnRz IGNhbiBiZSB1bmlmaWVkIHRvZ2V0aGVyIHNpbmNlIG9ubHkgSU9WQS0+SFBBIG1hcHBpbmcgDQpu ZWVkcyB0byBiZSBidWlsdC4NCg0KU28gdW5saWtlIGV4aXN0aW5nIHR5cGUxIElPTU1VIGRyaXZl ciB3aGljaCBjb250cm9scyBJT01NVSBhbG9uZSwgdkdQVSANCklPTU1VIGRyaXZlciBuZWVkcyB0 byBjb29wZXJhdGUgd2l0aCBvdGhlciBhZ2VudCAoaW92YS5jIGhlcmUpIHRvDQpjby1tYW5hZ2Ug c3lzdGVtIElPTU1VLiBUaGlzIG1heSBub3QgaW1wYWN0IGV4aXN0aW5nIFZGSU8gZnJhbWV3b3Jr Lg0KSnVzdCB3YW50IHRvIGhpZ2hsaWdodCBhZGRpdGlvbmFsIHdvcmsgaGVyZSB3aGVuIGltcGxl bWVudGluZyB0aGUgdkdQVQ0KSU9NTVUgZHJpdmVyLg0KDQpUaGFua3MNCktldmluDQogDQoNCg0K VGhhbmtzDQpLZXZpbg0K -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2015-11-20 18:30 +0100 |
| Message-ID | <qwXFw-1Np-1@gated-at.bofh.it> |
| In reply to | #1273796 |
On Fri, 2015-11-20 at 08:10 +0000, Tian, Kevin wrote:
> > From: Tian, Kevin
> > Sent: Friday, November 20, 2015 3:10 PM
>
> > > > >
> > > > > The proposal is therefore that GPU vendors can expose vGPUs to
> > > > > userspace, and thus to QEMU, using the VFIO API. For instance, vfio
> > > > > supports modular bus drivers and IOMMU drivers. An intel-vfio-gvt-d
> > > > > module (or extension of i915) can register as a vfio bus driver, create
> > > > > a struct device per vGPU, create an IOMMU group for that device, and
> > > > > register that device with the vfio-core. Since we don't rely on the
> > > > > system IOMMU for GVT-d vGPU assignment, another vGPU vendor driver (or
> > > > > extension of the same module) can register a "type1" compliant IOMMU
> > > > > driver into vfio-core. From the perspective of QEMU then, all of the
> > > > > existing vfio-pci code is re-used, QEMU remains largely unaware of any
> > > > > specifics of the vGPU being assigned, and the only necessary change so
> > > > > far is how QEMU traverses sysfs to find the device and thus the IOMMU
> > > > > group leading to the vfio group.
> > > >
> > > > GVT-g requires to pin guest memory and query GPA->HPA information,
> > > > upon which shadow GTTs will be updated accordingly from (GMA->GPA)
> > > > to (GMA->HPA). So yes, here a dummy or simple "type1" compliant IOMMU
> > > > can be introduced just for this requirement.
> > > >
> > > > However there's one tricky point which I'm not sure whether overall
> > > > VFIO concept will be violated. GVT-g doesn't require system IOMMU
> > > > to function, however host system may enable system IOMMU just for
> > > > hardening purpose. This means two-level translations existing (GMA->
> > > > IOVA->HPA), so the dummy IOMMU driver has to request system IOMMU
> > > > driver to allocate IOVA for VMs and then setup IOVA->HPA mapping
> > > > in IOMMU page table. In this case, multiple VM's translations are
> > > > multiplexed in one IOMMU page table.
> > > >
> > > > We might need create some group/sub-group or parent/child concepts
> > > > among those IOMMUs for thorough permission control.
> > >
> > > My thought here is that this is all abstracted through the vGPU IOMMU
> > > and device vfio backends. It's the GPU driver itself, or some vfio
> > > extension of that driver, mediating access to the device and deciding
> > > when to configure GPU MMU mappings. That driver has access to the GPA
> > > to HVA translations thanks to the type1 complaint IOMMU it implements
> > > and can pin pages as needed to create GPA to HPA mappings. That should
> > > give it all the pieces it needs to fully setup mappings for the vGPU.
> > > Whether or not there's a system IOMMU is simply an exercise for that
> > > driver. It needs to do a DMA mapping operation through the system IOMMU
> > > the same for a vGPU as if it was doing it for itself, because they are
> > > in fact one in the same. The GMA to IOVA mapping seems like an internal
> > > detail. I assume the IOVA is some sort of GPA, and the GMA is managed
> > > through mediation of the device.
> >
> > Sorry I'm not familiar with VFIO internal. My original worry is that system
> > IOMMU for GPU may be already claimed by another vfio driver (e.g. host kernel
> > wants to harden gfx driver from rest sub-systems, regardless of whether vGPU
> > is created or not). In that case vGPU IOMMU driver shouldn't manage system
> > IOMMU directly.
> >
> > btw, curious today how VFIO coordinates with system IOMMU driver regarding
> > to whether a IOMMU is used to control device assignment, or used for kernel
> > hardening. Somehow two are conflicting since different address spaces are
> > concerned (GPA vs. IOVA)...
> >
>
> Here is a more concrete example:
>
> KVMGT doesn't require IOMMU. All DMA targets are already replaced with
> HPA thru shadow GTT. So DMA requests from GPU all contain HPAs.
>
> When IOMMU is enabled, one simple approach is to have vGPU IOMMU
> driver configure system IOMMU with identity mapping (HPA->HPA). We
> can't use (GPA->HPA) since GPAs from multiple VMs are conflicting.
>
> However, we still have host gfx driver running. When IOMMU is enabled,
> dma_alloc_*** will return IOVA (drvers/iommu/iova.c) in host gfx driver,
> which will have IOVA->HPA programmed to system IOMMU.
>
> One IOMMU device entry can only translate one address space, so here
> comes a conflict (HPA->HPA vs. IOVA->HPA). To solve this, vGPU IOMMU
> driver needs to allocate IOVA from iova.c for each VM w/ vGPU assigned,
> and then KVMGT will program IOVA in shadow GTT accordingly. It adds
> one additional mapping layer (GPA->IOVA->HPA). In this way two
> requirements can be unified together since only IOVA->HPA mapping
> needs to be built.
>
> So unlike existing type1 IOMMU driver which controls IOMMU alone, vGPU
> IOMMU driver needs to cooperate with other agent (iova.c here) to
> co-manage system IOMMU. This may not impact existing VFIO framework.
> Just want to highlight additional work here when implementing the vGPU
> IOMMU driver.
Right, so the existing i915 driver needs to use the DMA API and calls
like dma_map_page() to enable translations through the IOMMU. With
dma_map_page(), the caller provides a page address (~HPA) and is
returned an IOVA. So unfortunately you don't get to take the shortcut
of having an identity mapping through the IOMMU unless you want to
convert i915 entirely to using the IOMMU API, because we also can't have
the conflict that an HPA could overlap an IOVA for a previously mapped
page.
The double translation, once through the GPU MMU and once through the
system IOMMU is going to happen regardless of whether we can identity
map through the IOMMU. The only solution to this would be for the GPU
to participate in ATS and provide pre-translated transactions from the
GPU. All of this is internal to the i915 driver (or vfio extension of
that driver) and needs to be done regardless of what sort of interface
we're using to expose the vGPU to QEMU. It just seems like VFIO
provides a convenient way of doing this since you'll have ready access
to the HVA-GPA mappings for the user.
I think the key points though are:
* the VFIO type1 IOMMU stores GPA to HVA translations
* get_user_pages() on the HVA will pin the page and give you a
page
* dma_map_page() receives that page, programs the system IOMMU and
provides an IOVA
* the GPU MMU can then be programmed with the GPA to IOVA
translations
Thanks,
Alex
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Jike Song <jike.song@intel.com> |
|---|---|
| Date | 2015-11-23 06:10 +0100 |
| Subject | Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel |
| Message-ID | <qxRy2-5Kz-1@gated-at.bofh.it> |
| In reply to | #1274303 |
On 11/21/2015 01:25 AM, Alex Williamson wrote: > On Fri, 2015-11-20 at 08:10 +0000, Tian, Kevin wrote: >> >> Here is a more concrete example: >> >> KVMGT doesn't require IOMMU. All DMA targets are already replaced with >> HPA thru shadow GTT. So DMA requests from GPU all contain HPAs. >> >> When IOMMU is enabled, one simple approach is to have vGPU IOMMU >> driver configure system IOMMU with identity mapping (HPA->HPA). We >> can't use (GPA->HPA) since GPAs from multiple VMs are conflicting. >> >> However, we still have host gfx driver running. When IOMMU is enabled, >> dma_alloc_*** will return IOVA (drvers/iommu/iova.c) in host gfx driver, >> which will have IOVA->HPA programmed to system IOMMU. >> >> One IOMMU device entry can only translate one address space, so here >> comes a conflict (HPA->HPA vs. IOVA->HPA). To solve this, vGPU IOMMU >> driver needs to allocate IOVA from iova.c for each VM w/ vGPU assigned, >> and then KVMGT will program IOVA in shadow GTT accordingly. It adds >> one additional mapping layer (GPA->IOVA->HPA). In this way two >> requirements can be unified together since only IOVA->HPA mapping >> needs to be built. >> >> So unlike existing type1 IOMMU driver which controls IOMMU alone, vGPU >> IOMMU driver needs to cooperate with other agent (iova.c here) to >> co-manage system IOMMU. This may not impact existing VFIO framework. >> Just want to highlight additional work here when implementing the vGPU >> IOMMU driver. > > Right, so the existing i915 driver needs to use the DMA API and calls > like dma_map_page() to enable translations through the IOMMU. With > dma_map_page(), the caller provides a page address (~HPA) and is > returned an IOVA. So unfortunately you don't get to take the shortcut > of having an identity mapping through the IOMMU unless you want to > convert i915 entirely to using the IOMMU API, because we also can't have > the conflict that an HPA could overlap an IOVA for a previously mapped > page. > > The double translation, once through the GPU MMU and once through the > system IOMMU is going to happen regardless of whether we can identity > map through the IOMMU. The only solution to this would be for the GPU > to participate in ATS and provide pre-translated transactions from the > GPU. All of this is internal to the i915 driver (or vfio extension of > that driver) and needs to be done regardless of what sort of interface > we're using to expose the vGPU to QEMU. It just seems like VFIO > provides a convenient way of doing this since you'll have ready access > to the HVA-GPA mappings for the user. > > I think the key points though are: > > * the VFIO type1 IOMMU stores GPA to HVA translations > * get_user_pages() on the HVA will pin the page and give you a > page > * dma_map_page() receives that page, programs the system IOMMU and > provides an IOVA > * the GPU MMU can then be programmed with the GPA to IOVA > translations Thanks for such a nice example! I'll do my home work and get back to you shortly :) > > Thanks, > Alex > -- Thanks, Jike -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Daniel Vetter <daniel@ffwll.ch> |
|---|---|
| Date | 2015-11-24 12:20 +0100 |
| Subject | Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel |
| Message-ID | <qyjND-7oT-9@gated-at.bofh.it> |
| In reply to | #1273447 |
On Thu, Nov 19, 2015 at 01:02:36PM -0700, Alex Williamson wrote: > On Thu, 2015-11-19 at 04:06 +0000, Tian, Kevin wrote: > > > From: Alex Williamson [mailto:alex.williamson@redhat.com] > > > Sent: Thursday, November 19, 2015 2:12 AM > > > > > > [cc +qemu-devel, +paolo, +gerd] > > > > > > Another area of extension is how to expose a framebuffer to QEMU for > > > seamless integration into a SPICE/VNC channel. For this I believe we > > > could use a new region, much like we've done to expose VGA access > > > through a vfio device file descriptor. An area within this new > > > framebuffer region could be directly mappable in QEMU while a > > > non-mappable page, at a standard location with standardized format, > > > provides a description of framebuffer and potentially even a > > > communication channel to synchronize framebuffer captures. This would > > > be new code for QEMU, but something we could share among all vGPU > > > implementations. > > > > Now GVT-g already provides an interface to decode framebuffer information, > > w/ an assumption that the framebuffer will be further composited into > > OpenGL APIs. So the format is defined according to OpenGL definition. > > Does that meet SPICE requirement? > > > > Another thing to be added. Framebuffers are frequently switched in > > reality. So either Qemu needs to poll or a notification mechanism is required. > > And since it's dynamic, having framebuffer page directly exposed in the > > new region might be tricky. We can just expose framebuffer information > > (including base, format, etc.) and let Qemu to map separately out of VFIO > > interface. > > Sure, we'll need to work out that interface, but it's also possible that > the framebuffer region is simply remapped to another area of the device > (ie. multiple interfaces mapping the same thing) by the vfio device > driver. Whether it's easier to do that or make the framebuffer region > reference another region is something we'll need to see. > > > And... this works fine with vGPU model since software knows all the > > detail about framebuffer. However in pass-through case, who do you expect > > to provide that information? Is it OK to introduce vGPU specific APIs in > > VFIO? > > Yes, vGPU may have additional features, like a framebuffer area, that > aren't present or optional for direct assignment. Obviously we support > direct assignment of GPUs for some vendors already without this feature. For exposing framebuffers for spice/vnc I highly recommend against anything that looks like a bar/fixed mmio range mapping. First this means the kernel driver needs to internally fake remapping, which isn't fun. Second we can't get at the memory in an easy fashion for hw-accelerated compositing. My recoomendation is to build the actual memory access for underlying framebuffers on top of dma-buf, so that it can be vacuumed up by e.g. the host gpu driver again for rendering. For userspace the generic part would simply be an invalidate-fb signal, with the new dma-buf supplied. Upsides: - You can composit stuff with the gpu. - VRAM and other kinds of resources (even stuff not visible in pci bars) can be represented. Downside: Tracking mapping changes on the guest side won't be any easier. This is mostly a problem for integrated gpus, since discrete ones usually require contiguous vram for scanout. I think saying "don't do that" is a valid option though, i.e. we're assuming that page mappings for a in-use scanout range never changes on the guest side. That is true for at least all the current linux drivers. -Daniel -- Daniel Vetter Software Engineer, Intel Corporation http://blog.ffwll.ch -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Chris Wilson <chris@chris-wilson.co.uk> |
|---|---|
| Date | 2015-11-24 13:00 +0100 |
| Subject | Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel |
| Message-ID | <qykqm-7Cj-7@gated-at.bofh.it> |
| In reply to | #1276337 |
On Tue, Nov 24, 2015 at 12:19:18PM +0100, Daniel Vetter wrote: > Downside: Tracking mapping changes on the guest side won't be any easier. > This is mostly a problem for integrated gpus, since discrete ones usually > require contiguous vram for scanout. I think saying "don't do that" is a > valid option though, i.e. we're assuming that page mappings for a in-use > scanout range never changes on the guest side. That is true for at least > all the current linux drivers. Apart from we already suffer limitations of fixed mappings and have patches that want to change the page mapping of active scanouts. -Chris -- Chris Wilson, Intel Open Source Technology Centre -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Gerd Hoffmann <kraxel@redhat.com> |
|---|---|
| Date | 2015-11-24 13:40 +0100 |
| Message-ID | <qyl34-86O-13@gated-at.bofh.it> |
| In reply to | #1276337 |
Hi,
> > Yes, vGPU may have additional features, like a framebuffer area, that
> > aren't present or optional for direct assignment. Obviously we support
> > direct assignment of GPUs for some vendors already without this feature.
>
> For exposing framebuffers for spice/vnc I highly recommend against
> anything that looks like a bar/fixed mmio range mapping. First this means
> the kernel driver needs to internally fake remapping, which isn't fun.
Sure. I don't think we should remap here. More below.
> My recoomendation is to build the actual memory access for underlying
> framebuffers on top of dma-buf, so that it can be vacuumed up by e.g. the
> host gpu driver again for rendering.
We want that too ;)
Some more background:
OpenGL support in qemu is still young and emerging, and we are actually
building on dma-bufs here. There are a bunch of different ways how
guest display output is handled. At the end of the day it boils down to
only two fundamental cases though:
(a) Where qemu doesn't need access to the guest framebuffer
- qemu directly renders via opengl (works today with virtio-gpu
and will be in the qemu 2.5 release)
- qemu passed on the dma-buf to spice client for local display
(experimental code exists).
- qemu feeds the guest display into gpu-assisted video encoder
to send a stream over the network (no code yet).
(b) Where qemu must read the guest framebuffer.
- qemu's builtin vnc server.
- qemu writing screenshots to file.
- (non-opengl legacy code paths for local display, will
hopefully disappear long-term though ...)
So, the question is how to support (b) best. Even with OpenGL support
in qemu improving over time I don't expect this going away completely
anytime soon.
I think it makes sense to have a special vfio region for that. I don't
think remapping makes sense there. It doesn't need to be "live", it
doesn't need support high refresh rates. Placing a copy of the guest
framebuffer there on request (and convert from tiled to linear while
being at it) is perfectly fine. qemu has a adaptive update rate and
will stop doing frequent update requests when the vnc client
disconnects, so there will be nothing to do if nobody wants actually see
the guest display.
Possible alternative approach would be to import a dma-buf, then use
glReadPixels(). I suspect when doing the copy in the kernel the driver
could ask just the gpu to blit the guest framebuffer. Don't know gfx
hardware good enough to be sure though, comments are welcome.
cheers,
Gerd
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Daniel Vetter <daniel@ffwll.ch> |
|---|---|
| Date | 2015-11-24 14:40 +0100 |
| Subject | Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel |
| Message-ID | <qylZ9-g8-29@gated-at.bofh.it> |
| In reply to | #1276390 |
On Tue, Nov 24, 2015 at 01:38:55PM +0100, Gerd Hoffmann wrote: > Hi, > > > > Yes, vGPU may have additional features, like a framebuffer area, that > > > aren't present or optional for direct assignment. Obviously we support > > > direct assignment of GPUs for some vendors already without this feature. > > > > For exposing framebuffers for spice/vnc I highly recommend against > > anything that looks like a bar/fixed mmio range mapping. First this means > > the kernel driver needs to internally fake remapping, which isn't fun. > > Sure. I don't think we should remap here. More below. > > > My recoomendation is to build the actual memory access for underlying > > framebuffers on top of dma-buf, so that it can be vacuumed up by e.g. the > > host gpu driver again for rendering. > > We want that too ;) > > Some more background: > > OpenGL support in qemu is still young and emerging, and we are actually > building on dma-bufs here. There are a bunch of different ways how > guest display output is handled. At the end of the day it boils down to > only two fundamental cases though: > > (a) Where qemu doesn't need access to the guest framebuffer > - qemu directly renders via opengl (works today with virtio-gpu > and will be in the qemu 2.5 release) > - qemu passed on the dma-buf to spice client for local display > (experimental code exists). > - qemu feeds the guest display into gpu-assisted video encoder > to send a stream over the network (no code yet). > > (b) Where qemu must read the guest framebuffer. > - qemu's builtin vnc server. > - qemu writing screenshots to file. > - (non-opengl legacy code paths for local display, will > hopefully disappear long-term though ...) > > So, the question is how to support (b) best. Even with OpenGL support > in qemu improving over time I don't expect this going away completely > anytime soon. > > I think it makes sense to have a special vfio region for that. I don't > think remapping makes sense there. It doesn't need to be "live", it > doesn't need support high refresh rates. Placing a copy of the guest > framebuffer there on request (and convert from tiled to linear while > being at it) is perfectly fine. qemu has a adaptive update rate and > will stop doing frequent update requests when the vnc client > disconnects, so there will be nothing to do if nobody wants actually see > the guest display. > > Possible alternative approach would be to import a dma-buf, then use > glReadPixels(). I suspect when doing the copy in the kernel the driver > could ask just the gpu to blit the guest framebuffer. Don't know gfx > hardware good enough to be sure though, comments are welcome. Generally the kernel can't do gpu blts since the required massive state setup is only in the userspace side of the GL driver stack. But glReadPixels can do tricks for detiling, and if you use pixel buffer objects or something similar it'll even be amortized reasonably. But there's some work to add generic mmap support to dma-bufs, and for really simple case (where we don't have a gl driver to handle the dma-buf specially) for untiled framebuffers that would be all we need? -Daniel -- Daniel Vetter Software Engineer, Intel Corporation http://blog.ffwll.ch -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Gerd Hoffmann <kraxel@redhat.com> |
|---|---|
| Date | 2015-11-24 15:20 +0100 |
| Message-ID | <qymBP-JT-3@gated-at.bofh.it> |
| In reply to | #1276458 |
Hi, > But there's some work to add generic mmap support to dma-bufs, and for > really simple case (where we don't have a gl driver to handle the dma-buf > specially) for untiled framebuffers that would be all we need? Not requiring gl is certainly a bonus, people might want build qemu without opengl support to reduce the attach surface and/or package dependency chain. And, yes, requirements for the non-gl rendering path are pretty low. qemu needs something it can mmap, and which it can ask pixman to handle. Preferred format is PIXMAN_x8r8g8b8 (qemu uses that internally in alot of places so this avoids conversions). Current plan is to have a special vfio region (not visible to the guest) where the framebuffer lives, with one or two pages at the end for meta data (format and size). Status field is there too and will be used by qemu to request updates and the kernel to signal update completion. Guess I should write that down as vfio rfc patch ... I don't think it makes sense to have fields to notify qemu about which framebuffer regions have been updated, I'd expect with full-screen composing we have these days this information isn't available anyway. Maybe a flag telling whenever there have been updates or not, so qemu can skip update processing in case we have the screensaver showing a black screen all day long. cheers, Gerd -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Daniel Vetter <daniel@ffwll.ch> |
|---|---|
| Date | 2015-11-24 15:20 +0100 |
| Subject | Re: [Intel-gfx] [Announcement] 2015-Q3 release of XenGT - a Mediated Graphics Passthrough Solution from Intel |
| Message-ID | <qymBQ-JT-21@gated-at.bofh.it> |
| In reply to | #1276509 |
On Tue, Nov 24, 2015 at 03:12:31PM +0100, Gerd Hoffmann wrote: > Hi, > > > But there's some work to add generic mmap support to dma-bufs, and for > > really simple case (where we don't have a gl driver to handle the dma-buf > > specially) for untiled framebuffers that would be all we need? > > Not requiring gl is certainly a bonus, people might want build qemu > without opengl support to reduce the attach surface and/or package > dependency chain. > > And, yes, requirements for the non-gl rendering path are pretty low. > qemu needs something it can mmap, and which it can ask pixman to handle. > Preferred format is PIXMAN_x8r8g8b8 (qemu uses that internally in alot > of places so this avoids conversions). > > Current plan is to have a special vfio region (not visible to the guest) > where the framebuffer lives, with one or two pages at the end for meta > data (format and size). Status field is there too and will be used by > qemu to request updates and the kernel to signal update completion. > Guess I should write that down as vfio rfc patch ... > > I don't think it makes sense to have fields to notify qemu about which > framebuffer regions have been updated, I'd expect with full-screen > composing we have these days this information isn't available anyway. > Maybe a flag telling whenever there have been updates or not, so qemu > can skip update processing in case we have the screensaver showing a > black screen all day long. GL, wayland, X, EGL and soonish Android's surface flinger (hwc already has it afaik) all track damage. There's plans to add the same to the atomic kms api too. But if you do damage tracking you really don't want to support (maybe allow for perf reasons if the guest is stupid) frontbuffer rendering, which means you need buffer handles + damage, and not a static region. -Daniel -- Daniel Vetter Software Engineer, Intel Corporation http://blog.ffwll.ch -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web