Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1434540 > unrolled thread

[PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed

Started byPaolo Bonzini <pbonzini@redhat.com>
First post2016-06-30 15:10 +0200
Last post2016-07-06 18:00 +0200
Articles 20 on this page of 43 — 4 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-06-30 15:10 +0200
    [PATCH 1/2] KVM: MMU: prepare to support mapping of VM_IO and VM_PFNMAP frames Paolo Bonzini <pbonzini@redhat.com> - 2016-06-30 15:10 +0200
    Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-01 00:10 +0200
    Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 08:50 +0200
      Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-04 09:10 +0200
        Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 09:50 +0200
          Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-04 09:50 +0200
            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 10:10 +0200
              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-04 10:20 +0200
                Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 10:30 +0200
                  Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-04 10:50 +0200
          Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-04 10:00 +0200
            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-04 17:40 +0200
              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-05 03:30 +0200
                Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 03:40 +0200
                  Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-05 06:10 +0200
                    Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 07:20 +0200
                      Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-05 08:40 +0200
                        Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 09:40 +0200
                          Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-05 11:10 +0200
                            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 17:10 +0200
                              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-06 04:30 +0200
                                Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-06 06:10 +0200
            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 10:30 +0200
              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 12:30 +0200
              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-04 10:50 +0200
                Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 10:50 +0200
                  Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 11:00 +0200
                  Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-04 11:20 +0200
      Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-04 09:40 +0200
        Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 09:50 +0200
    Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 07:50 +0200
      Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-05 14:20 +0200
        Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 16:10 +0200
        Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-06 04:10 +0200
          Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-06 04:20 +0200
            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-06 04:40 +0200
              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-06 05:00 +0200
                Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-06 06:10 +0200
                  Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-06 13:50 +0200
                    Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-07 04:50 +0200
          Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-06 08:10 +0200
            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Alex Williamson <alex.williamson@redhat.com> - 2016-07-06 18:00 +0200

Page 2 of 3 — ← Prev page 1 [2] 3  Next page →


#1437112

FromNeo Jia <cjia@nvidia.com>
Date2016-07-05 17:10 +0200
Message-ID<rRA93-4vb-13@gated-at.bofh.it>
In reply to#1436861
On Tue, Jul 05, 2016 at 05:02:46PM +0800, Xiao Guangrong wrote:
> 
> >
> >It is physically contiguous but it is done during the runtime, physically contiguous doesn't mean
> >static partition at boot time. And only during runtime, the proper HW resource will be requested therefore
> >the right portion of MMIO region will be granted by the mediated device driver on the host.
> 
> Okay. This is your implantation design rather than the hardware limitation, right?

I don't think it matters here. We are talking about framework so it should
provide the flexibility for different driver vendor.

> 
> For example, if the instance require 512M memory (the size can be specified by QEMU
> command line), it can tell its requirement to the mediated device driver via create()
> interface, then the driver can allocate then memory for this instance before it is running.

BAR != your device memory

We don't set the BAR size via QEMU command line, BAR size is extracted by QEMU
from config space provided by vendor driver.

> 
> Theoretically, the hardware is able to do memory management as this style, but for some
> reasons you choose allocating memory in the runtime. right? If my understanding is right,
> could you please tell us what benefit you want to get from this running-allocation style?

Your understanding is incorrect.

> 
> >
> >Then the req_size and pgoff will both come from the mediated device driver based on his internal book
> >keeping of the hw resource allocation, which is only available during runtime. And such book keeping
> >can be built part of para-virtualization scheme between guest and host device driver.
> >
> 
> I am talking the parameters you passed to validate_map_request(). req_size is calculated like this:
> 
> +       offset   = virtaddr - vma->vm_start;
> +       phyaddr  = (vma->vm_pgoff << PAGE_SHIFT) + offset;
> +       pgoff    = phyaddr >> PAGE_SHIFT;
> 
> All these info is from vma which is available in mmmap().
> 
> pgoff is got from:
> +       pg_prot  = vma->vm_page_prot;
> that is also available in mmap().

This is kept there in case the validate_map_request() is not provided by vendor
driver then by default assume 1:1 mapping. So if validate_map_request() is not
provided, fault handler should not fail.

> 
> >None of such information is available at VFIO mmap() time. For example, several VMs
> >are sharing the same physical device to provide mediated access. All VMs will
> >call the VFIO mmap() on their virtual BAR as part of QEMU vfio/pci initialization
> >process, at that moment, we definitely can't mmap the entire physical MMIO
> >into both VM blindly for obvious reason.
> >
> 
> mmap() carries @length information, so you only need to allocate the specified size
> (corresponding to @length) of memory for them.

Again, you still look at this as a static partition at QEMU configuration time
where the guest mmio will be mapped as a whole at some offset of the physical
mmio region. (You still can do that like I said above by not providing
validate_map_request() in your vendor driver.)

But this is not the framework we are defining here.

The framework we have here is to provide the driver vendor flexibility to decide
the guest mmio and physical mmio mapping on page basis, and such information is
available during runtime.

How such information gets communicated between guest and host driver is up to
driver vendor.

Thanks,
Neo

[toc] | [prev] | [next] | [standalone]


#1437370

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-06 04:30 +0200
Message-ID<rRKL7-3c5-1@gated-at.bofh.it>
In reply to#1437112

On 07/05/2016 11:07 PM, Neo Jia wrote:
> On Tue, Jul 05, 2016 at 05:02:46PM +0800, Xiao Guangrong wrote:
>>
>>>
>>> It is physically contiguous but it is done during the runtime, physically contiguous doesn't mean
>>> static partition at boot time. And only during runtime, the proper HW resource will be requested therefore
>>> the right portion of MMIO region will be granted by the mediated device driver on the host.
>>
>> Okay. This is your implantation design rather than the hardware limitation, right?
>
> I don't think it matters here. We are talking about framework so it should
> provide the flexibility for different driver vendor.

It really matters. It is the reason why we design the framework like this and
we need to make sure whether we have a better design to fill the requirements.

>
>>
>> For example, if the instance require 512M memory (the size can be specified by QEMU
>> command line), it can tell its requirement to the mediated device driver via create()
>> interface, then the driver can allocate then memory for this instance before it is running.
>
> BAR != your device memory
>
> We don't set the BAR size via QEMU command line, BAR size is extracted by QEMU
> from config space provided by vendor driver.
>

Anyway, the BAR size have a way to configure, e.g, specify the size as a parameter when
you create a mdev via sysfs.

>>
>> Theoretically, the hardware is able to do memory management as this style, but for some
>> reasons you choose allocating memory in the runtime. right? If my understanding is right,
>> could you please tell us what benefit you want to get from this running-allocation style?
>
> Your understanding is incorrect.

Then WHY?

>
>>
>>>
>>> Then the req_size and pgoff will both come from the mediated device driver based on his internal book
>>> keeping of the hw resource allocation, which is only available during runtime. And such book keeping
>>> can be built part of para-virtualization scheme between guest and host device driver.
>>>
>>
>> I am talking the parameters you passed to validate_map_request(). req_size is calculated like this:
>>
>> +       offset   = virtaddr - vma->vm_start;
>> +       phyaddr  = (vma->vm_pgoff << PAGE_SHIFT) + offset;
>> +       pgoff    = phyaddr >> PAGE_SHIFT;
>>
>> All these info is from vma which is available in mmmap().
>>
>> pgoff is got from:
>> +       pg_prot  = vma->vm_page_prot;
>> that is also available in mmap().
>
> This is kept there in case the validate_map_request() is not provided by vendor
> driver then by default assume 1:1 mapping. So if validate_map_request() is not
> provided, fault handler should not fail.

THESE are the parameters you passed to validate_map_request(), and these info is
available in mmap(), it really does not matter if you move validate_map_request()
to mmap(). That's what i want to say.

>
>>
>>> None of such information is available at VFIO mmap() time. For example, several VMs
>>> are sharing the same physical device to provide mediated access. All VMs will
>>> call the VFIO mmap() on their virtual BAR as part of QEMU vfio/pci initialization
>>> process, at that moment, we definitely can't mmap the entire physical MMIO
>>> into both VM blindly for obvious reason.
>>>
>>
>> mmap() carries @length information, so you only need to allocate the specified size
>> (corresponding to @length) of memory for them.
>
> Again, you still look at this as a static partition at QEMU configuration time
> where the guest mmio will be mapped as a whole at some offset of the physical
> mmio region. (You still can do that like I said above by not providing
> validate_map_request() in your vendor driver.)
>

Then you can move validate_map_request() to here to achieve custom allocation-policy.

> But this is not the framework we are defining here.
>
> The framework we have here is to provide the driver vendor flexibility to decide
> the guest mmio and physical mmio mapping on page basis, and such information is
> available during runtime.
>
> How such information gets communicated between guest and host driver is up to
> driver vendor.

The problems is the sequence of the way "provide the driver vendor
flexibility to decide the guest mmio and physical mmio mapping on page basis"
and mmap().

We should provide such allocation info first then do mmap(). You current design,
do mmap() -> communication telling such info -> use such info when fault happens,
is really BAD, because you can not control the time when memory fault will happen.
The guest may access this memory before the communication you mentioned above,
and another reason is that KVM MMU can prefetch memory at any time.

[toc] | [prev] | [next] | [standalone]


#1437397

FromNeo Jia <cjia@nvidia.com>
Date2016-07-06 06:10 +0200
Message-ID<rRMjT-4hL-1@gated-at.bofh.it>
In reply to#1437370
On Wed, Jul 06, 2016 at 10:22:59AM +0800, Xiao Guangrong wrote:
> 
> 
> On 07/05/2016 11:07 PM, Neo Jia wrote:
> >This is kept there in case the validate_map_request() is not provided by vendor
> >driver then by default assume 1:1 mapping. So if validate_map_request() is not
> >provided, fault handler should not fail.
> 
> THESE are the parameters you passed to validate_map_request(), and these info is
> available in mmap(), it really does not matter if you move validate_map_request()
> to mmap(). That's what i want to say.

Let me answer this at the end of my response.

> 
> >
> >>
> >>>None of such information is available at VFIO mmap() time. For example, several VMs
> >>>are sharing the same physical device to provide mediated access. All VMs will
> >>>call the VFIO mmap() on their virtual BAR as part of QEMU vfio/pci initialization
> >>>process, at that moment, we definitely can't mmap the entire physical MMIO
> >>>into both VM blindly for obvious reason.
> >>>
> >>
> >>mmap() carries @length information, so you only need to allocate the specified size
> >>(corresponding to @length) of memory for them.
> >
> >Again, you still look at this as a static partition at QEMU configuration time
> >where the guest mmio will be mapped as a whole at some offset of the physical
> >mmio region. (You still can do that like I said above by not providing
> >validate_map_request() in your vendor driver.)
> >
> 
> Then you can move validate_map_request() to here to achieve custom allocation-policy.
> 
> >But this is not the framework we are defining here.
> >
> >The framework we have here is to provide the driver vendor flexibility to decide
> >the guest mmio and physical mmio mapping on page basis, and such information is
> >available during runtime.
> >
> >How such information gets communicated between guest and host driver is up to
> >driver vendor.
> 
> The problems is the sequence of the way "provide the driver vendor
> flexibility to decide the guest mmio and physical mmio mapping on page basis"
> and mmap().
> 
> We should provide such allocation info first then do mmap(). You current design,
> do mmap() -> communication telling such info -> use such info when fault happens,
> is really BAD, because you can not control the time when memory fault will happen.
> The guest may access this memory before the communication you mentioned above,
> and another reason is that KVM MMU can prefetch memory at any time.

Like I have said before if your implementation doesn't need such flexibility,
you can still do a static mapping at VFIO mmap() time, then your mediated driver
doesn't have to provide validate_map_request, also the fault handler will not be
called. 

Let me address your questions below.

1. Information available at VFIO mmap() time?

So you are saying that the @req_size and &pgoff are both available in the time
when people are calling VFIO mmap() when guest OS is not even running, right?

The answer is No, the only thing are available at VFIO mmap are the following:

1) guest MMIO size

2) host physical MMIO size

3) guest MMIO starting address

4) host MMIO starting address

But none of above are the @req_size and @pgoff that we are talking about at the
validate_map_request time.

Our host MMIO is representing the means to access GPU HW resource. Those GPU HW
resources are allocated dynamically at runtime. So we have no visibility of the
@pgoff and @req_size that is covering some specific type of GPU HW resource at
VFIO mmap time. Also we don't even know if such resource will be required for a
particular VM or not.

For example, VM1 will need to launch a lot of graphics workload than VM2. So the
end result is that VM1 will gets a lot of resource A allocated than VM2 to
support his graphics workload. And to access resource A, the host mmio region
will be allocated as well, say [pfn_a size_a], the VM2 is [pfn_b, size_b].

Clearly, such region can be destroyed and reallocated through mediated driver
lifetime. This is why we need to have a fault handler there to map the proper
pages into guest after validation in the runtime.

I hope above response can address your question why we can't provide such
allocation info at VFIO mmap() time.

2. Guest might access mmio region at any time ...

Guest with a mediated GPU inside can definitely access his BARs at any time. If
guest is accessing some his BAR region that is not previously allocated, then
such access will be denied and with current scheme VM will crash to prevent
malicious access from the guest. This is another reason we choose to keep the
guest MMIO mediated.

3. KVM MMU can prefetch memory at any time.

You are talking about the KVM MMU prefetch the guest mmio region which is marked
as prefetchable right?

On baremetal, the prefetch is basic a cache line fill, where the range needs to
be marked as cachable for the CPU, then it issues a read to anywhere in the
cache line.

Is KVM MMU prefetch the same as baremetal? If so, it is at least *not at any
time" right?

And the prefetch will only happen when there is a valid guest CPU mapping to its
guest mmio region. Then, it goes back to issue (2), if the CPU mapping setup is
done via proper driver and get validated by the mediated device driver, such
prefetch will work as expected. If not, then such prefech is no different than
the malicious or unsupported access from the guest, VM will crash.

Thanks,
Neo

[toc] | [prev] | [next] | [standalone]


#1436274

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-04 10:30 +0200
Message-ID<rR7qq-3hN-19@gated-at.bofh.it>
In reply to#1436239

On 07/04/2016 03:53 PM, Neo Jia wrote:
> On Mon, Jul 04, 2016 at 03:37:35PM +0800, Xiao Guangrong wrote:
>>
>>
>> On 07/04/2016 03:03 PM, Neo Jia wrote:
>>> On Mon, Jul 04, 2016 at 02:39:22PM +0800, Xiao Guangrong wrote:
>>>>
>>>>
>>>> On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
>>>>> The vGPU folks would like to trap the first access to a BAR by setting
>>>>> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
>>>>> then can use remap_pfn_range to place some non-reserved pages in the VMA.
>>>>
>>>> Why does it require fetching the pfn when the fault is triggered rather
>>>> than when mmap() is called?
>>>
>>> Hi Guangrong,
>>>
>>> as such mapping information between virtual mmio to physical mmio is only available
>>> at runtime.
>>
>> Sorry, i do not know what the different between mmap() and the time VM actually
>> accesses the memory for your case. Could you please more detail?
>
> Hi Guangrong,
>
> Sure. The mmap() gets called by qemu or any VFIO API userspace consumer when
> setting up the virtual mmio, at that moment nobody has any knowledge about how
> the physical mmio gets virtualized.
>
> When the vm (or application if we don't want to limit ourselves to vmm term)
> starts, the virtual and physical mmio gets mapped by mpci kernel module with the
> help from vendor supplied mediated host driver according to the hw resource
> assigned to this vm / application.

Thanks for your expiation.

It sounds like a strategy of resource allocation, you delay the allocation until VM really
accesses it, right?

[toc] | [prev] | [next] | [standalone]


#1436286

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-04 12:30 +0200
Message-ID<rR9iy-4q7-7@gated-at.bofh.it>
In reply to#1436274

On 07/04/2016 05:16 PM, Neo Jia wrote:
> On Mon, Jul 04, 2016 at 04:45:05PM +0800, Xiao Guangrong wrote:
>>
>>
>> On 07/04/2016 04:41 PM, Neo Jia wrote:
>>> On Mon, Jul 04, 2016 at 04:19:20PM +0800, Xiao Guangrong wrote:
>>>>
>>>>
>>>> On 07/04/2016 03:53 PM, Neo Jia wrote:
>>>>> On Mon, Jul 04, 2016 at 03:37:35PM +0800, Xiao Guangrong wrote:
>>>>>>
>>>>>>
>>>>>> On 07/04/2016 03:03 PM, Neo Jia wrote:
>>>>>>> On Mon, Jul 04, 2016 at 02:39:22PM +0800, Xiao Guangrong wrote:
>>>>>>>>
>>>>>>>>
>>>>>>>> On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
>>>>>>>>> The vGPU folks would like to trap the first access to a BAR by setting
>>>>>>>>> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
>>>>>>>>> then can use remap_pfn_range to place some non-reserved pages in the VMA.
>>>>>>>>
>>>>>>>> Why does it require fetching the pfn when the fault is triggered rather
>>>>>>>> than when mmap() is called?
>>>>>>>
>>>>>>> Hi Guangrong,
>>>>>>>
>>>>>>> as such mapping information between virtual mmio to physical mmio is only available
>>>>>>> at runtime.
>>>>>>
>>>>>> Sorry, i do not know what the different between mmap() and the time VM actually
>>>>>> accesses the memory for your case. Could you please more detail?
>>>>>
>>>>> Hi Guangrong,
>>>>>
>>>>> Sure. The mmap() gets called by qemu or any VFIO API userspace consumer when
>>>>> setting up the virtual mmio, at that moment nobody has any knowledge about how
>>>>> the physical mmio gets virtualized.
>>>>>
>>>>> When the vm (or application if we don't want to limit ourselves to vmm term)
>>>>> starts, the virtual and physical mmio gets mapped by mpci kernel module with the
>>>>> help from vendor supplied mediated host driver according to the hw resource
>>>>> assigned to this vm / application.
>>>>
>>>> Thanks for your expiation.
>>>>
>>>> It sounds like a strategy of resource allocation, you delay the allocation until VM really
>>>> accesses it, right?
>>>
>>> Yes, that is where the fault handler inside mpci code comes to the picture.
>>
>>
>> I am not sure this strategy is good. The instance is successfully created, and it is started
>> successful, but the VM is crashed due to the resource of that instance is not enough. That sounds
>> unreasonable.
>
>
> Sorry, I think I misread the "allocation" as "mapping". We only delay the
> cpu mapping, not the allocation.

So how to understand your statement:
"at that moment nobody has any knowledge about how the physical mmio gets virtualized"

The resource, physical MMIO region, has been allocated, why we do not know the physical
address mapped to the VM?

[toc] | [prev] | [next] | [standalone]


#1436319

FromNeo Jia <cjia@nvidia.com>
Date2016-07-04 10:50 +0200
Message-ID<rR7JM-3ot-21@gated-at.bofh.it>
In reply to#1436274
On Mon, Jul 04, 2016 at 04:19:20PM +0800, Xiao Guangrong wrote:
> 
> 
> On 07/04/2016 03:53 PM, Neo Jia wrote:
> >On Mon, Jul 04, 2016 at 03:37:35PM +0800, Xiao Guangrong wrote:
> >>
> >>
> >>On 07/04/2016 03:03 PM, Neo Jia wrote:
> >>>On Mon, Jul 04, 2016 at 02:39:22PM +0800, Xiao Guangrong wrote:
> >>>>
> >>>>
> >>>>On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
> >>>>>The vGPU folks would like to trap the first access to a BAR by setting
> >>>>>vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> >>>>>then can use remap_pfn_range to place some non-reserved pages in the VMA.
> >>>>
> >>>>Why does it require fetching the pfn when the fault is triggered rather
> >>>>than when mmap() is called?
> >>>
> >>>Hi Guangrong,
> >>>
> >>>as such mapping information between virtual mmio to physical mmio is only available
> >>>at runtime.
> >>
> >>Sorry, i do not know what the different between mmap() and the time VM actually
> >>accesses the memory for your case. Could you please more detail?
> >
> >Hi Guangrong,
> >
> >Sure. The mmap() gets called by qemu or any VFIO API userspace consumer when
> >setting up the virtual mmio, at that moment nobody has any knowledge about how
> >the physical mmio gets virtualized.
> >
> >When the vm (or application if we don't want to limit ourselves to vmm term)
> >starts, the virtual and physical mmio gets mapped by mpci kernel module with the
> >help from vendor supplied mediated host driver according to the hw resource
> >assigned to this vm / application.
> 
> Thanks for your expiation.
> 
> It sounds like a strategy of resource allocation, you delay the allocation until VM really
> accesses it, right?

Yes, that is where the fault handler inside mpci code comes to the picture.

Thanks,
Neo

[toc] | [prev] | [next] | [standalone]


#1436322

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-04 10:50 +0200
Message-ID<rR7JM-3ot-25@gated-at.bofh.it>
In reply to#1436319

On 07/04/2016 04:41 PM, Neo Jia wrote:
> On Mon, Jul 04, 2016 at 04:19:20PM +0800, Xiao Guangrong wrote:
>>
>>
>> On 07/04/2016 03:53 PM, Neo Jia wrote:
>>> On Mon, Jul 04, 2016 at 03:37:35PM +0800, Xiao Guangrong wrote:
>>>>
>>>>
>>>> On 07/04/2016 03:03 PM, Neo Jia wrote:
>>>>> On Mon, Jul 04, 2016 at 02:39:22PM +0800, Xiao Guangrong wrote:
>>>>>>
>>>>>>
>>>>>> On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
>>>>>>> The vGPU folks would like to trap the first access to a BAR by setting
>>>>>>> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
>>>>>>> then can use remap_pfn_range to place some non-reserved pages in the VMA.
>>>>>>
>>>>>> Why does it require fetching the pfn when the fault is triggered rather
>>>>>> than when mmap() is called?
>>>>>
>>>>> Hi Guangrong,
>>>>>
>>>>> as such mapping information between virtual mmio to physical mmio is only available
>>>>> at runtime.
>>>>
>>>> Sorry, i do not know what the different between mmap() and the time VM actually
>>>> accesses the memory for your case. Could you please more detail?
>>>
>>> Hi Guangrong,
>>>
>>> Sure. The mmap() gets called by qemu or any VFIO API userspace consumer when
>>> setting up the virtual mmio, at that moment nobody has any knowledge about how
>>> the physical mmio gets virtualized.
>>>
>>> When the vm (or application if we don't want to limit ourselves to vmm term)
>>> starts, the virtual and physical mmio gets mapped by mpci kernel module with the
>>> help from vendor supplied mediated host driver according to the hw resource
>>> assigned to this vm / application.
>>
>> Thanks for your expiation.
>>
>> It sounds like a strategy of resource allocation, you delay the allocation until VM really
>> accesses it, right?
>
> Yes, that is where the fault handler inside mpci code comes to the picture.


I am not sure this strategy is good. The instance is successfully created, and it is started
successful, but the VM is crashed due to the resource of that instance is not enough. That sounds
unreasonable.

[toc] | [prev] | [next] | [standalone]


#1436339

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-04 11:00 +0200
Message-ID<rR7Tr-3rR-1@gated-at.bofh.it>
In reply to#1436322

On 07/04/2016 04:45 PM, Xiao Guangrong wrote:
>
>
> On 07/04/2016 04:41 PM, Neo Jia wrote:
>> On Mon, Jul 04, 2016 at 04:19:20PM +0800, Xiao Guangrong wrote:
>>>
>>>
>>> On 07/04/2016 03:53 PM, Neo Jia wrote:
>>>> On Mon, Jul 04, 2016 at 03:37:35PM +0800, Xiao Guangrong wrote:
>>>>>
>>>>>
>>>>> On 07/04/2016 03:03 PM, Neo Jia wrote:
>>>>>> On Mon, Jul 04, 2016 at 02:39:22PM +0800, Xiao Guangrong wrote:
>>>>>>>
>>>>>>>
>>>>>>> On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
>>>>>>>> The vGPU folks would like to trap the first access to a BAR by setting
>>>>>>>> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
>>>>>>>> then can use remap_pfn_range to place some non-reserved pages in the VMA.
>>>>>>>
>>>>>>> Why does it require fetching the pfn when the fault is triggered rather
>>>>>>> than when mmap() is called?
>>>>>>
>>>>>> Hi Guangrong,
>>>>>>
>>>>>> as such mapping information between virtual mmio to physical mmio is only available
>>>>>> at runtime.
>>>>>
>>>>> Sorry, i do not know what the different between mmap() and the time VM actually
>>>>> accesses the memory for your case. Could you please more detail?
>>>>
>>>> Hi Guangrong,
>>>>
>>>> Sure. The mmap() gets called by qemu or any VFIO API userspace consumer when
>>>> setting up the virtual mmio, at that moment nobody has any knowledge about how
>>>> the physical mmio gets virtualized.
>>>>
>>>> When the vm (or application if we don't want to limit ourselves to vmm term)
>>>> starts, the virtual and physical mmio gets mapped by mpci kernel module with the
>>>> help from vendor supplied mediated host driver according to the hw resource
>>>> assigned to this vm / application.
>>>
>>> Thanks for your expiation.
>>>
>>> It sounds like a strategy of resource allocation, you delay the allocation until VM really
>>> accesses it, right?
>>
>> Yes, that is where the fault handler inside mpci code comes to the picture.
>
>
> I am not sure this strategy is good. The instance is successfully created, and it is started
> successful, but the VM is crashed due to the resource of that instance is not enough. That sounds
> unreasonable.
>

Especially, you can not squeeze this kind of memory to balance the usage between all VMs. Does
this strategy still make sense?

[toc] | [prev] | [next] | [standalone]


#1436371

FromNeo Jia <cjia@nvidia.com>
Date2016-07-04 11:20 +0200
Message-ID<rR8cO-3NA-27@gated-at.bofh.it>
In reply to#1436322
On Mon, Jul 04, 2016 at 04:45:05PM +0800, Xiao Guangrong wrote:
> 
> 
> On 07/04/2016 04:41 PM, Neo Jia wrote:
> >On Mon, Jul 04, 2016 at 04:19:20PM +0800, Xiao Guangrong wrote:
> >>
> >>
> >>On 07/04/2016 03:53 PM, Neo Jia wrote:
> >>>On Mon, Jul 04, 2016 at 03:37:35PM +0800, Xiao Guangrong wrote:
> >>>>
> >>>>
> >>>>On 07/04/2016 03:03 PM, Neo Jia wrote:
> >>>>>On Mon, Jul 04, 2016 at 02:39:22PM +0800, Xiao Guangrong wrote:
> >>>>>>
> >>>>>>
> >>>>>>On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
> >>>>>>>The vGPU folks would like to trap the first access to a BAR by setting
> >>>>>>>vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> >>>>>>>then can use remap_pfn_range to place some non-reserved pages in the VMA.
> >>>>>>
> >>>>>>Why does it require fetching the pfn when the fault is triggered rather
> >>>>>>than when mmap() is called?
> >>>>>
> >>>>>Hi Guangrong,
> >>>>>
> >>>>>as such mapping information between virtual mmio to physical mmio is only available
> >>>>>at runtime.
> >>>>
> >>>>Sorry, i do not know what the different between mmap() and the time VM actually
> >>>>accesses the memory for your case. Could you please more detail?
> >>>
> >>>Hi Guangrong,
> >>>
> >>>Sure. The mmap() gets called by qemu or any VFIO API userspace consumer when
> >>>setting up the virtual mmio, at that moment nobody has any knowledge about how
> >>>the physical mmio gets virtualized.
> >>>
> >>>When the vm (or application if we don't want to limit ourselves to vmm term)
> >>>starts, the virtual and physical mmio gets mapped by mpci kernel module with the
> >>>help from vendor supplied mediated host driver according to the hw resource
> >>>assigned to this vm / application.
> >>
> >>Thanks for your expiation.
> >>
> >>It sounds like a strategy of resource allocation, you delay the allocation until VM really
> >>accesses it, right?
> >
> >Yes, that is where the fault handler inside mpci code comes to the picture.
> 
> 
> I am not sure this strategy is good. The instance is successfully created, and it is started
> successful, but the VM is crashed due to the resource of that instance is not enough. That sounds
> unreasonable.


Sorry, I think I misread the "allocation" as "mapping". We only delay the
cpu mapping, not the allocation.

Thanks,
Neo

> 
> 
> 

[toc] | [prev] | [next] | [standalone]


#1436178

FromPaolo Bonzini <pbonzini@redhat.com>
Date2016-07-04 09:40 +0200
Message-ID<rR6E2-2LX-27@gated-at.bofh.it>
In reply to#1436126

On 04/07/2016 08:39, Xiao Guangrong wrote:
> Why the memory mapped by this mmap() is not a portion of MMIO from
> underlayer physical device? If it is a valid system memory, is this
> interface
> really needed to implemented in vfio? (you at least need to set VM_MIXEDMAP
> if it mixed system memory with MMIO)

The KVM code does not care if VM_MIXEDMAP is set or not, it works in
either case.

Paolo

> IIUC, the kernel assumes that VM_PFNMAP is a continuous memory, e.g, like
> current KVM and vaddr_get_pfn() in vfio, but it seems nvdia's patchset
> breaks this semantic as ops->validate_map_request() can adjust the physical
> address arbitrarily. (again, the name 'validate' should be changed to match
> the thing as it is really doing)
> 
> 
> -- 
> To unsubscribe from this list: send the line "unsubscribe kvm" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
> 

[toc] | [prev] | [next] | [standalone]


#1436232

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-04 09:50 +0200
Message-ID<rR6NI-2Pw-23@gated-at.bofh.it>
In reply to#1436178

On 07/04/2016 03:38 PM, Paolo Bonzini wrote:
>
>
> On 04/07/2016 08:39, Xiao Guangrong wrote:
>> Why the memory mapped by this mmap() is not a portion of MMIO from
>> underlayer physical device? If it is a valid system memory, is this
>> interface
>> really needed to implemented in vfio? (you at least need to set VM_MIXEDMAP
>> if it mixed system memory with MMIO)
>
> The KVM code does not care if VM_MIXEDMAP is set or not, it works in
> either case.

Yes, it is. I mean nvdia's vfio patchset should use VM_MIXEDMAP if the memory
is mixed. :)

[toc] | [prev] | [next] | [standalone]


#1436771

FromNeo Jia <cjia@nvidia.com>
Date2016-07-05 07:50 +0200
Message-ID<rRrp7-76C-3@gated-at.bofh.it>
In reply to#1434540
On Thu, Jun 30, 2016 at 03:01:49PM +0200, Paolo Bonzini wrote:
> The vGPU folks would like to trap the first access to a BAR by setting
> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> then can use remap_pfn_range to place some non-reserved pages in the VMA.
> 
> KVM lacks support for this kind of non-linear VM_PFNMAP mapping, and these
> patches should fix this.

Hi Paolo,

I have tested your patches with the mediated passthru patchset that is being
reviewed in KVM and QEMU mailing list.

The fault handler gets called successfully and the previously mapped memory gets
unmmaped correctly via unmap_mapping_range.

Thanks,
Neo

> 
> Thanks,
> 
> Paolo
> 
> Paolo Bonzini (2):
>   KVM: MMU: prepare to support mapping of VM_IO and VM_PFNMAP frames
>   KVM: MMU: try to fix up page faults before giving up
> 
>  mm/gup.c            |  1 +
>  virt/kvm/kvm_main.c | 55 ++++++++++++++++++++++++++++++++++++++++++++++++-----
>  2 files changed, 51 insertions(+), 5 deletions(-)
> 
> -- 
> 1.8.3.1
> 

[toc] | [prev] | [next] | [standalone]


#1436958

FromPaolo Bonzini <pbonzini@redhat.com>
Date2016-07-05 14:20 +0200
Message-ID<rRxuy-2G4-19@gated-at.bofh.it>
In reply to#1436771

On 05/07/2016 07:41, Neo Jia wrote:
> On Thu, Jun 30, 2016 at 03:01:49PM +0200, Paolo Bonzini wrote:
>> The vGPU folks would like to trap the first access to a BAR by setting
>> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
>> then can use remap_pfn_range to place some non-reserved pages in the VMA.
>>
>> KVM lacks support for this kind of non-linear VM_PFNMAP mapping, and these
>> patches should fix this.
> 
> Hi Paolo,
> 
> I have tested your patches with the mediated passthru patchset that is being
> reviewed in KVM and QEMU mailing list.
> 
> The fault handler gets called successfully and the previously mapped memory gets
> unmmaped correctly via unmap_mapping_range.

Great, then I'll include them in 4.8.

Paolo

[toc] | [prev] | [next] | [standalone]


#1437043

FromNeo Jia <cjia@nvidia.com>
Date2016-07-05 16:10 +0200
Message-ID<rRzcZ-3Ok-1@gated-at.bofh.it>
In reply to#1436958
On Tue, Jul 05, 2016 at 02:18:28PM +0200, Paolo Bonzini wrote:
> 
> 
> On 05/07/2016 07:41, Neo Jia wrote:
> > On Thu, Jun 30, 2016 at 03:01:49PM +0200, Paolo Bonzini wrote:
> >> The vGPU folks would like to trap the first access to a BAR by setting
> >> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> >> then can use remap_pfn_range to place some non-reserved pages in the VMA.
> >>
> >> KVM lacks support for this kind of non-linear VM_PFNMAP mapping, and these
> >> patches should fix this.
> > 
> > Hi Paolo,
> > 
> > I have tested your patches with the mediated passthru patchset that is being
> > reviewed in KVM and QEMU mailing list.
> > 
> > The fault handler gets called successfully and the previously mapped memory gets
> > unmmaped correctly via unmap_mapping_range.
> 
> Great, then I'll include them in 4.8.

Thanks!

> 
> Paolo
> 

[toc] | [prev] | [next] | [standalone]


#1437364

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-06 04:10 +0200
Message-ID<rRKrL-2YQ-7@gated-at.bofh.it>
In reply to#1436958

On 07/05/2016 08:18 PM, Paolo Bonzini wrote:
>
>
> On 05/07/2016 07:41, Neo Jia wrote:
>> On Thu, Jun 30, 2016 at 03:01:49PM +0200, Paolo Bonzini wrote:
>>> The vGPU folks would like to trap the first access to a BAR by setting
>>> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
>>> then can use remap_pfn_range to place some non-reserved pages in the VMA.
>>>
>>> KVM lacks support for this kind of non-linear VM_PFNMAP mapping, and these
>>> patches should fix this.
>>
>> Hi Paolo,
>>
>> I have tested your patches with the mediated passthru patchset that is being
>> reviewed in KVM and QEMU mailing list.
>>
>> The fault handler gets called successfully and the previously mapped memory gets
>> unmmaped correctly via unmap_mapping_range.
>
> Great, then I'll include them in 4.8.

Code is okay, but i still suspect if this implementation, fetch mmio pages in fault
handler, is needed. We'd better include these patches after the design of vfio
framework is decided.

[toc] | [prev] | [next] | [standalone]


#1437365

FromNeo Jia <cjia@nvidia.com>
Date2016-07-06 04:20 +0200
Message-ID<rRKBr-38O-1@gated-at.bofh.it>
In reply to#1437364
On Wed, Jul 06, 2016 at 10:00:46AM +0800, Xiao Guangrong wrote:
> 
> 
> On 07/05/2016 08:18 PM, Paolo Bonzini wrote:
> >
> >
> >On 05/07/2016 07:41, Neo Jia wrote:
> >>On Thu, Jun 30, 2016 at 03:01:49PM +0200, Paolo Bonzini wrote:
> >>>The vGPU folks would like to trap the first access to a BAR by setting
> >>>vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> >>>then can use remap_pfn_range to place some non-reserved pages in the VMA.
> >>>
> >>>KVM lacks support for this kind of non-linear VM_PFNMAP mapping, and these
> >>>patches should fix this.
> >>
> >>Hi Paolo,
> >>
> >>I have tested your patches with the mediated passthru patchset that is being
> >>reviewed in KVM and QEMU mailing list.
> >>
> >>The fault handler gets called successfully and the previously mapped memory gets
> >>unmmaped correctly via unmap_mapping_range.
> >
> >Great, then I'll include them in 4.8.
> 
> Code is okay, but i still suspect if this implementation, fetch mmio pages in fault
> handler, is needed. We'd better include these patches after the design of vfio
> framework is decided.

Hi Guangrong,

I disagree. The design of VFIO framework has been actively discussed in the KVM
and QEMU mailing for a while and the fault handler is agreed upon to provide the
flexibility for different driver vendors' implementation. With that said, I am
still open to discuss with you and anybody else about this framework as the goal
is to allow multiple vendor to plugin into this framework to support their
mediated device virtualization scheme, such as Intel, IBM and us.

May I ask you what the exact issue you have with this interface for Intel to support 
your own GPU virtualization? 

Thanks,
Neo

> 

[toc] | [prev] | [next] | [standalone]


#1437372

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-06 04:40 +0200
Message-ID<rRKUN-3fr-1@gated-at.bofh.it>
In reply to#1437365

On 07/06/2016 10:18 AM, Neo Jia wrote:
> On Wed, Jul 06, 2016 at 10:00:46AM +0800, Xiao Guangrong wrote:
>>
>>
>> On 07/05/2016 08:18 PM, Paolo Bonzini wrote:
>>>
>>>
>>> On 05/07/2016 07:41, Neo Jia wrote:
>>>> On Thu, Jun 30, 2016 at 03:01:49PM +0200, Paolo Bonzini wrote:
>>>>> The vGPU folks would like to trap the first access to a BAR by setting
>>>>> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
>>>>> then can use remap_pfn_range to place some non-reserved pages in the VMA.
>>>>>
>>>>> KVM lacks support for this kind of non-linear VM_PFNMAP mapping, and these
>>>>> patches should fix this.
>>>>
>>>> Hi Paolo,
>>>>
>>>> I have tested your patches with the mediated passthru patchset that is being
>>>> reviewed in KVM and QEMU mailing list.
>>>>
>>>> The fault handler gets called successfully and the previously mapped memory gets
>>>> unmmaped correctly via unmap_mapping_range.
>>>
>>> Great, then I'll include them in 4.8.
>>
>> Code is okay, but i still suspect if this implementation, fetch mmio pages in fault
>> handler, is needed. We'd better include these patches after the design of vfio
>> framework is decided.
>
> Hi Guangrong,
>
> I disagree. The design of VFIO framework has been actively discussed in the KVM
> and QEMU mailing for a while and the fault handler is agreed upon to provide the
> flexibility for different driver vendors' implementation. With that said, I am
> still open to discuss with you and anybody else about this framework as the goal
> is to allow multiple vendor to plugin into this framework to support their
> mediated device virtualization scheme, such as Intel, IBM and us.

The discussion is still going on. And current vfio patchset we reviewed is still
problematic.

>
> May I ask you what the exact issue you have with this interface for Intel to support
> your own GPU virtualization?

Intel's vGPU can work with this framework. We really appreciate your / nvidia's
contribution.

i didn’t mean to offend you, i just want to make sure if this complexity is really
needed and inspect if this framework is safe enough and think it over if we have
a better implementation.

[toc] | [prev] | [next] | [standalone]


#1437387

FromNeo Jia <cjia@nvidia.com>
Date2016-07-06 05:00 +0200
Message-ID<rRLe9-3nf-11@gated-at.bofh.it>
In reply to#1437372
On Wed, Jul 06, 2016 at 10:35:18AM +0800, Xiao Guangrong wrote:
> 
> 
> On 07/06/2016 10:18 AM, Neo Jia wrote:
> >On Wed, Jul 06, 2016 at 10:00:46AM +0800, Xiao Guangrong wrote:
> >>
> >>
> >>On 07/05/2016 08:18 PM, Paolo Bonzini wrote:
> >>>
> >>>
> >>>On 05/07/2016 07:41, Neo Jia wrote:
> >>>>On Thu, Jun 30, 2016 at 03:01:49PM +0200, Paolo Bonzini wrote:
> >>>>>The vGPU folks would like to trap the first access to a BAR by setting
> >>>>>vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> >>>>>then can use remap_pfn_range to place some non-reserved pages in the VMA.
> >>>>>
> >>>>>KVM lacks support for this kind of non-linear VM_PFNMAP mapping, and these
> >>>>>patches should fix this.
> >>>>
> >>>>Hi Paolo,
> >>>>
> >>>>I have tested your patches with the mediated passthru patchset that is being
> >>>>reviewed in KVM and QEMU mailing list.
> >>>>
> >>>>The fault handler gets called successfully and the previously mapped memory gets
> >>>>unmmaped correctly via unmap_mapping_range.
> >>>
> >>>Great, then I'll include them in 4.8.
> >>
> >>Code is okay, but i still suspect if this implementation, fetch mmio pages in fault
> >>handler, is needed. We'd better include these patches after the design of vfio
> >>framework is decided.
> >
> >Hi Guangrong,
> >
> >I disagree. The design of VFIO framework has been actively discussed in the KVM
> >and QEMU mailing for a while and the fault handler is agreed upon to provide the
> >flexibility for different driver vendors' implementation. With that said, I am
> >still open to discuss with you and anybody else about this framework as the goal
> >is to allow multiple vendor to plugin into this framework to support their
> >mediated device virtualization scheme, such as Intel, IBM and us.
> 
> The discussion is still going on. And current vfio patchset we reviewed is still
> problematic.

My point is the fault handler part has been discussed already, with that said I
am always open to any constructive suggestions to make things better and
maintainable. (Appreciate your code review on the VFIO thread, I think we still
own you another response, will do that.)

> 
> >
> >May I ask you what the exact issue you have with this interface for Intel to support
> >your own GPU virtualization?
> 
> Intel's vGPU can work with this framework. We really appreciate your / nvidia's
> contribution.

Then, I don't think we should embargo Paolo's patch.

> 
> i didn’t mean to offend you, i just want to make sure if this complexity is really
> needed and inspect if this framework is safe enough and think it over if we have
> a better implementation.

Not at all. :-)

Suggestions are always welcome, I just want to know the exact issues you have
with the code so I can have a better response to address that with proper 
information.

Thanks,
Neo

[toc] | [prev] | [next] | [standalone]


#1437399

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-06 06:10 +0200
Message-ID<rRMjT-4hL-3@gated-at.bofh.it>
In reply to#1437387

On 07/06/2016 10:57 AM, Neo Jia wrote:
> On Wed, Jul 06, 2016 at 10:35:18AM +0800, Xiao Guangrong wrote:
>>
>>
>> On 07/06/2016 10:18 AM, Neo Jia wrote:
>>> On Wed, Jul 06, 2016 at 10:00:46AM +0800, Xiao Guangrong wrote:
>>>>
>>>>
>>>> On 07/05/2016 08:18 PM, Paolo Bonzini wrote:
>>>>>
>>>>>
>>>>> On 05/07/2016 07:41, Neo Jia wrote:
>>>>>> On Thu, Jun 30, 2016 at 03:01:49PM +0200, Paolo Bonzini wrote:
>>>>>>> The vGPU folks would like to trap the first access to a BAR by setting
>>>>>>> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
>>>>>>> then can use remap_pfn_range to place some non-reserved pages in the VMA.
>>>>>>>
>>>>>>> KVM lacks support for this kind of non-linear VM_PFNMAP mapping, and these
>>>>>>> patches should fix this.
>>>>>>
>>>>>> Hi Paolo,
>>>>>>
>>>>>> I have tested your patches with the mediated passthru patchset that is being
>>>>>> reviewed in KVM and QEMU mailing list.
>>>>>>
>>>>>> The fault handler gets called successfully and the previously mapped memory gets
>>>>>> unmmaped correctly via unmap_mapping_range.
>>>>>
>>>>> Great, then I'll include them in 4.8.
>>>>
>>>> Code is okay, but i still suspect if this implementation, fetch mmio pages in fault
>>>> handler, is needed. We'd better include these patches after the design of vfio
>>>> framework is decided.
>>>
>>> Hi Guangrong,
>>>
>>> I disagree. The design of VFIO framework has been actively discussed in the KVM
>>> and QEMU mailing for a while and the fault handler is agreed upon to provide the
>>> flexibility for different driver vendors' implementation. With that said, I am
>>> still open to discuss with you and anybody else about this framework as the goal
>>> is to allow multiple vendor to plugin into this framework to support their
>>> mediated device virtualization scheme, such as Intel, IBM and us.
>>
>> The discussion is still going on. And current vfio patchset we reviewed is still
>> problematic.
>
> My point is the fault handler part has been discussed already, with that said I
> am always open to any constructive suggestions to make things better and
> maintainable. (Appreciate your code review on the VFIO thread, I think we still
> own you another response, will do that.)
>

It always can be changed especially the vfio patchset is not in a good shape.

>>
>>>
>>> May I ask you what the exact issue you have with this interface for Intel to support
>>> your own GPU virtualization?
>>
>> Intel's vGPU can work with this framework. We really appreciate your / nvidia's
>> contribution.
>
> Then, I don't think we should embargo Paolo's patch.

This patchset is specific for the framework design, i.e, mapping memory when fault
happens rather than mmap(), and this design is exact what we are discussing for
nearly two days.

[toc] | [prev] | [next] | [standalone]


#1437659

FromPaolo Bonzini <pbonzini@redhat.com>
Date2016-07-06 13:50 +0200
Message-ID<rRTv4-fc-27@gated-at.bofh.it>
In reply to#1437399

On 06/07/2016 06:02, Xiao Guangrong wrote:
>>>
>>>>
>>>> May I ask you what the exact issue you have with this interface for
>>>> Intel to support
>>>> your own GPU virtualization?
>>>
>>> Intel's vGPU can work with this framework. We really appreciate your
>>> / nvidia's
>>> contribution.
>>
>> Then, I don't think we should embargo Paolo's patch.
> 
> This patchset is specific for the framework design, i.e, mapping memory when
> fault happens rather than mmap(), and this design is exact what we are
> discussing for nearly two days.

I disagree, this patch fixes a bug because what Neo is doing is legal.
It may not be the design that will be committed, but the bug they found
in KVM is real.

Paolo

[toc] | [prev] | [next] | [standalone]


Page 2 of 3 — ← Prev page 1 [2] 3  Next page →

Back to top | Article view | linux.kernel


csiph-web