Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1434540 > unrolled thread

[PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed

Started byPaolo Bonzini <pbonzini@redhat.com>
First post2016-06-30 15:10 +0200
Last post2016-07-06 18:00 +0200
Articles 20 on this page of 43 — 4 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-06-30 15:10 +0200
    [PATCH 1/2] KVM: MMU: prepare to support mapping of VM_IO and VM_PFNMAP frames Paolo Bonzini <pbonzini@redhat.com> - 2016-06-30 15:10 +0200
    Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-01 00:10 +0200
    Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 08:50 +0200
      Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-04 09:10 +0200
        Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 09:50 +0200
          Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-04 09:50 +0200
            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 10:10 +0200
              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-04 10:20 +0200
                Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 10:30 +0200
                  Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-04 10:50 +0200
          Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-04 10:00 +0200
            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-04 17:40 +0200
              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-05 03:30 +0200
                Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 03:40 +0200
                  Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-05 06:10 +0200
                    Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 07:20 +0200
                      Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-05 08:40 +0200
                        Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 09:40 +0200
                          Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-05 11:10 +0200
                            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 17:10 +0200
                              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-06 04:30 +0200
                                Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-06 06:10 +0200
            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 10:30 +0200
              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 12:30 +0200
              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-04 10:50 +0200
                Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 10:50 +0200
                  Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 11:00 +0200
                  Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-04 11:20 +0200
      Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-04 09:40 +0200
        Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-04 09:50 +0200
    Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 07:50 +0200
      Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-05 14:20 +0200
        Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-05 16:10 +0200
        Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-06 04:10 +0200
          Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-06 04:20 +0200
            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-06 04:40 +0200
              Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Neo Jia <cjia@nvidia.com> - 2016-07-06 05:00 +0200
                Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-06 06:10 +0200
                  Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-06 13:50 +0200
                    Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-07-07 04:50 +0200
          Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Paolo Bonzini <pbonzini@redhat.com> - 2016-07-06 08:10 +0200
            Re: [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed Alex Williamson <alex.williamson@redhat.com> - 2016-07-06 18:00 +0200

Page 1 of 3  [1] 2 3  Next page →


#1434540 — [PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed

FromPaolo Bonzini <pbonzini@redhat.com>
Date2016-06-30 15:10 +0200
Subject[PATCH 0/2] KVM: MMU: support VMAs that got remap_pfn_range-ed
Message-ID<rPJTb-1y9-7@gated-at.bofh.it>
The vGPU folks would like to trap the first access to a BAR by setting
vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
then can use remap_pfn_range to place some non-reserved pages in the VMA.

KVM lacks support for this kind of non-linear VM_PFNMAP mapping, and these
patches should fix this.

Thanks,

Paolo

Paolo Bonzini (2):
  KVM: MMU: prepare to support mapping of VM_IO and VM_PFNMAP frames
  KVM: MMU: try to fix up page faults before giving up

 mm/gup.c            |  1 +
 virt/kvm/kvm_main.c | 55 ++++++++++++++++++++++++++++++++++++++++++++++++-----
 2 files changed, 51 insertions(+), 5 deletions(-)

-- 
1.8.3.1

[toc] | [next] | [standalone]


#1434541 — [PATCH 1/2] KVM: MMU: prepare to support mapping of VM_IO and VM_PFNMAP frames

FromPaolo Bonzini <pbonzini@redhat.com>
Date2016-06-30 15:10 +0200
Subject[PATCH 1/2] KVM: MMU: prepare to support mapping of VM_IO and VM_PFNMAP frames
Message-ID<rPJTc-1y9-23@gated-at.bofh.it>
In reply to#1434540
Handle VM_IO like VM_PFNMAP, as is common in the rest of Linux; extract
the formula to convert hva->pfn into a new function, which will soon
gain more capabilities.

Cc: Xiao Guangrong <guangrong.xiao@linux.intel.com>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Radim Krčmář <rkrcmar@redhat.com>
Signed-off-by: Paolo Bonzini <pbonzini@redhat.com>
---
 virt/kvm/kvm_main.c | 20 +++++++++++++++-----
 1 file changed, 15 insertions(+), 5 deletions(-)

diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c
index ef54b4c31792..5aae59e00bef 100644
--- a/virt/kvm/kvm_main.c
+++ b/virt/kvm/kvm_main.c
@@ -1442,6 +1442,16 @@ static bool vma_is_valid(struct vm_area_struct *vma, bool write_fault)
 	return true;
 }
 
+static int hva_to_pfn_remapped(struct vm_area_struct *vma,
+			       unsigned long addr, bool *async,
+			       bool write_fault, kvm_pfn_t *p_pfn)
+{
+	*p_pfn = ((addr - vma->vm_start) >> PAGE_SHIFT) +
+		vma->vm_pgoff;
+	BUG_ON(!kvm_is_reserved_pfn(*p_pfn));
+	return 0;
+}
+
 /*
  * Pin guest page in memory and return its pfn.
  * @addr: host virtual address which maps memory to the guest
@@ -1461,7 +1471,7 @@ static kvm_pfn_t hva_to_pfn(unsigned long addr, bool atomic, bool *async,
 {
 	struct vm_area_struct *vma;
 	kvm_pfn_t pfn = 0;
-	int npages;
+	int npages, r;
 
 	/* we can do it either atomically or asynchronously, not both */
 	BUG_ON(atomic && async);
@@ -1487,10 +1497,10 @@ static kvm_pfn_t hva_to_pfn(unsigned long addr, bool atomic, bool *async,
 
 	if (vma == NULL)
 		pfn = KVM_PFN_ERR_FAULT;
-	else if ((vma->vm_flags & VM_PFNMAP)) {
-		pfn = ((addr - vma->vm_start) >> PAGE_SHIFT) +
-			vma->vm_pgoff;
-		BUG_ON(!kvm_is_reserved_pfn(pfn));
+	else if (vma->vm_flags & (VM_IO | VM_PFNMAP)) {
+		r = hva_to_pfn_remapped(vma, addr, async, write_fault, &pfn);
+		if (r < 0)
+			pfn = KVM_PFN_ERR_FAULT;
 	} else {
 		if (async && vma_is_valid(vma, write_fault))
 			*async = true;
-- 
1.8.3.1

[toc] | [prev] | [next] | [standalone]


#1434872

FromNeo Jia <cjia@nvidia.com>
Date2016-07-01 00:10 +0200
Message-ID<rPSjM-6FA-11@gated-at.bofh.it>
In reply to#1434540
On Thu, Jun 30, 2016 at 03:01:49PM +0200, Paolo Bonzini wrote:
> The vGPU folks would like to trap the first access to a BAR by setting
> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> then can use remap_pfn_range to place some non-reserved pages in the VMA.

Hi Paolo,

Thanks for the quick patches, I am in the middle of verifying them and will
report back asap.

Thanks,
Neo

> 
> KVM lacks support for this kind of non-linear VM_PFNMAP mapping, and these
> patches should fix this.
> 
> Thanks,
> 
> Paolo
> 
> Paolo Bonzini (2):
>   KVM: MMU: prepare to support mapping of VM_IO and VM_PFNMAP frames
>   KVM: MMU: try to fix up page faults before giving up
> 
>  mm/gup.c            |  1 +
>  virt/kvm/kvm_main.c | 55 ++++++++++++++++++++++++++++++++++++++++++++++++-----
>  2 files changed, 51 insertions(+), 5 deletions(-)
> 
> -- 
> 1.8.3.1
> 

[toc] | [prev] | [next] | [standalone]


#1436126

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-04 08:50 +0200
Message-ID<rR5RE-2gx-7@gated-at.bofh.it>
In reply to#1434540

On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
> The vGPU folks would like to trap the first access to a BAR by setting
> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> then can use remap_pfn_range to place some non-reserved pages in the VMA.

Why does it require fetching the pfn when the fault is triggered rather
than when mmap() is called?

Why the memory mapped by this mmap() is not a portion of MMIO from
underlayer physical device? If it is a valid system memory, is this interface
really needed to implemented in vfio? (you at least need to set VM_MIXEDMAP
if it mixed system memory with MMIO)

IIUC, the kernel assumes that VM_PFNMAP is a continuous memory, e.g, like
current KVM and vaddr_get_pfn() in vfio, but it seems nvdia's patchset
breaks this semantic as ops->validate_map_request() can adjust the physical
address arbitrarily. (again, the name 'validate' should be changed to match
the thing as it is really doing)

[toc] | [prev] | [next] | [standalone]


#1436137

FromNeo Jia <cjia@nvidia.com>
Date2016-07-04 09:10 +0200
Message-ID<rR6aZ-2Cq-29@gated-at.bofh.it>
In reply to#1436126
On Mon, Jul 04, 2016 at 02:39:22PM +0800, Xiao Guangrong wrote:
> 
> 
> On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
> >The vGPU folks would like to trap the first access to a BAR by setting
> >vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> >then can use remap_pfn_range to place some non-reserved pages in the VMA.
> 
> Why does it require fetching the pfn when the fault is triggered rather
> than when mmap() is called?

Hi Guangrong,

as such mapping information between virtual mmio to physical mmio is only available 
at runtime.

> 
> Why the memory mapped by this mmap() is not a portion of MMIO from
> underlayer physical device? If it is a valid system memory, is this interface
> really needed to implemented in vfio? (you at least need to set VM_MIXEDMAP
> if it mixed system memory with MMIO)
> 

It actually is a portion of the physical mmio which is set by vfio mmap.

> IIUC, the kernel assumes that VM_PFNMAP is a continuous memory, e.g, like
> current KVM and vaddr_get_pfn() in vfio, but it seems nvdia's patchset
> breaks this semantic as ops->validate_map_request() can adjust the physical
> address arbitrarily. (again, the name 'validate' should be changed to match
> the thing as it is really doing)

The vgpu api will allow you to adjust the target mmio address and the size via 
validate_map_request, but it is still physical contiguous as <start_pfn, size>.

Thanks,
Neo

> 
> 

[toc] | [prev] | [next] | [standalone]


#1436206

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-04 09:50 +0200
Message-ID<rR6NH-2Pw-1@gated-at.bofh.it>
In reply to#1436137

On 07/04/2016 03:03 PM, Neo Jia wrote:
> On Mon, Jul 04, 2016 at 02:39:22PM +0800, Xiao Guangrong wrote:
>>
>>
>> On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
>>> The vGPU folks would like to trap the first access to a BAR by setting
>>> vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
>>> then can use remap_pfn_range to place some non-reserved pages in the VMA.
>>
>> Why does it require fetching the pfn when the fault is triggered rather
>> than when mmap() is called?
>
> Hi Guangrong,
>
> as such mapping information between virtual mmio to physical mmio is only available
> at runtime.

Sorry, i do not know what the different between mmap() and the time VM actually
accesses the memory for your case. Could you please more detail?

>
>>
>> Why the memory mapped by this mmap() is not a portion of MMIO from
>> underlayer physical device? If it is a valid system memory, is this interface
>> really needed to implemented in vfio? (you at least need to set VM_MIXEDMAP
>> if it mixed system memory with MMIO)
>>
>
> It actually is a portion of the physical mmio which is set by vfio mmap.

So i do not think we need to care its refcount, i,e, we can consider it as reserved_pfn,
Paolo?

>
>> IIUC, the kernel assumes that VM_PFNMAP is a continuous memory, e.g, like
>> current KVM and vaddr_get_pfn() in vfio, but it seems nvdia's patchset
>> breaks this semantic as ops->validate_map_request() can adjust the physical
>> address arbitrarily. (again, the name 'validate' should be changed to match
>> the thing as it is really doing)
>
> The vgpu api will allow you to adjust the target mmio address and the size via
> validate_map_request, but it is still physical contiguous as <start_pfn, size>.

Okay, the interface confused us, maybe this interface need to be cooked to reflect
to this fact.

Thanks!

[toc] | [prev] | [next] | [standalone]


#1436209

FromPaolo Bonzini <pbonzini@redhat.com>
Date2016-07-04 09:50 +0200
Message-ID<rR6NH-2Pw-19@gated-at.bofh.it>
In reply to#1436206

On 04/07/2016 09:37, Xiao Guangrong wrote:
>>>
>>
>> It actually is a portion of the physical mmio which is set by vfio mmap.
> 
> So i do not think we need to care its refcount, i,e, we can consider it
> as reserved_pfn,
> Paolo?

nVidia provided me (offlist) with a simple patch that modified VFIO to
exhibit the problem, and it didn't use reserved PFNs.  This is why the
commit message for the patch is not entirely accurate.

But apart from this, it's much more obvious to consider the refcount.
The x86 MMU code doesn't care if the page is reserved or not;
mmu_set_spte does a kvm_release_pfn_clean, hence it makes sense for
hva_to_pfn_remapped to try doing a get_page (via kvm_get_pfn) after
invoking the fault handler, just like the get_user_pages family of
function does.

Paolo

[toc] | [prev] | [next] | [standalone]


#1436250

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-04 10:10 +0200
Message-ID<rR773-3bg-7@gated-at.bofh.it>
In reply to#1436209

On 07/04/2016 03:48 PM, Paolo Bonzini wrote:
>
>
> On 04/07/2016 09:37, Xiao Guangrong wrote:
>>>>
>>>
>>> It actually is a portion of the physical mmio which is set by vfio mmap.
>>
>> So i do not think we need to care its refcount, i,e, we can consider it
>> as reserved_pfn,
>> Paolo?
>
> nVidia provided me (offlist) with a simple patch that modified VFIO to
> exhibit the problem, and it didn't use reserved PFNs.  This is why the
> commit message for the patch is not entirely accurate.
>

It's clear now.

> But apart from this, it's much more obvious to consider the refcount.
> The x86 MMU code doesn't care if the page is reserved or not;
> mmu_set_spte does a kvm_release_pfn_clean, hence it makes sense for
> hva_to_pfn_remapped to try doing a get_page (via kvm_get_pfn) after
> invoking the fault handler, just like the get_user_pages family of
> function does.

Well,  it's little strange as you always try to get refcont
for a PFNMAP region without MIXEDMAP which indicates all the memory
in this region is no 'struct page' backend.

But it works as kvm_{get, release}_* have already been aware of
reserved_pfn, so i am okay with it......

[toc] | [prev] | [next] | [standalone]


#1436253

FromPaolo Bonzini <pbonzini@redhat.com>
Date2016-07-04 10:20 +0200
Message-ID<rR7gJ-3eC-1@gated-at.bofh.it>
In reply to#1436250

On 04/07/2016 09:59, Xiao Guangrong wrote:
> 
>> But apart from this, it's much more obvious to consider the refcount.
>> The x86 MMU code doesn't care if the page is reserved or not;
>> mmu_set_spte does a kvm_release_pfn_clean, hence it makes sense for
>> hva_to_pfn_remapped to try doing a get_page (via kvm_get_pfn) after
>> invoking the fault handler, just like the get_user_pages family of
>> function does.
> 
> Well,  it's little strange as you always try to get refcont
> for a PFNMAP region without MIXEDMAP which indicates all the memory
> in this region is no 'struct page' backend.

Fair enough, I can modify the comment.

	/*
	 * In case the VMA has VM_MIXEDMAP set, whoever called remap_pfn_range
	 * is also going to call e.g. unmap_mapping_range before the underlying
	 * non-reserved pages are freed, which will then call our MMU notifier.
	 * We still have to get a reference here to the page, because the callers
	 * of *hva_to_pfn* and *gfn_to_pfn* ultimately end up doing a
	 * kvm_release_pfn_clean on the returned pfn.  If the pfn is
	 * reserved, the kvm_get_pfn/kvm_release_pfn_clean pair will simply
	 * do nothing.
	 */

Paolo

> But it works as kvm_{get, release}_* have already been aware of
> reserved_pfn, so i am okay with it......

[toc] | [prev] | [next] | [standalone]


#1436270

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-04 10:30 +0200
Message-ID<rR7qp-3hN-11@gated-at.bofh.it>
In reply to#1436253

On 07/04/2016 04:14 PM, Paolo Bonzini wrote:
>
>
> On 04/07/2016 09:59, Xiao Guangrong wrote:
>>
>>> But apart from this, it's much more obvious to consider the refcount.
>>> The x86 MMU code doesn't care if the page is reserved or not;
>>> mmu_set_spte does a kvm_release_pfn_clean, hence it makes sense for
>>> hva_to_pfn_remapped to try doing a get_page (via kvm_get_pfn) after
>>> invoking the fault handler, just like the get_user_pages family of
>>> function does.
>>
>> Well,  it's little strange as you always try to get refcont
>> for a PFNMAP region without MIXEDMAP which indicates all the memory
>> in this region is no 'struct page' backend.
>
> Fair enough, I can modify the comment.
>
> 	/*
> 	 * In case the VMA has VM_MIXEDMAP set, whoever called remap_pfn_range
> 	 * is also going to call e.g. unmap_mapping_range before the underlying
> 	 * non-reserved pages are freed, which will then call our MMU notifier.
> 	 * We still have to get a reference here to the page, because the callers
> 	 * of *hva_to_pfn* and *gfn_to_pfn* ultimately end up doing a
> 	 * kvm_release_pfn_clean on the returned pfn.  If the pfn is
> 	 * reserved, the kvm_get_pfn/kvm_release_pfn_clean pair will simply
> 	 * do nothing.
> 	 */
>

Excellent. I like it. :)

[toc] | [prev] | [next] | [standalone]


#1436310

FromPaolo Bonzini <pbonzini@redhat.com>
Date2016-07-04 10:50 +0200
Message-ID<rR7JM-3ot-15@gated-at.bofh.it>
In reply to#1436270

On 04/07/2016 10:21, Xiao Guangrong wrote:
>>
>>     /*
>>      * In case the VMA has VM_MIXEDMAP set, whoever called
>> remap_pfn_range
>>      * is also going to call e.g. unmap_mapping_range before the
>> underlying
>>      * non-reserved pages are freed, which will then call our MMU
>> notifier.
>>      * We still have to get a reference here to the page, because the
>> callers
>>      * of *hva_to_pfn* and *gfn_to_pfn* ultimately end up doing a
>>      * kvm_release_pfn_clean on the returned pfn.  If the pfn is
>>      * reserved, the kvm_get_pfn/kvm_release_pfn_clean pair will simply
>>      * do nothing.
>>      */
>>
> 
> Excellent. I like it. :)

So is it Reviewed-by Guangrong? :)

Paolo

[toc] | [prev] | [next] | [standalone]


#1436239

FromNeo Jia <cjia@nvidia.com>
Date2016-07-04 10:00 +0200
Message-ID<rR6Xo-2SK-19@gated-at.bofh.it>
In reply to#1436206
On Mon, Jul 04, 2016 at 03:37:35PM +0800, Xiao Guangrong wrote:
> 
> 
> On 07/04/2016 03:03 PM, Neo Jia wrote:
> >On Mon, Jul 04, 2016 at 02:39:22PM +0800, Xiao Guangrong wrote:
> >>
> >>
> >>On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
> >>>The vGPU folks would like to trap the first access to a BAR by setting
> >>>vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> >>>then can use remap_pfn_range to place some non-reserved pages in the VMA.
> >>
> >>Why does it require fetching the pfn when the fault is triggered rather
> >>than when mmap() is called?
> >
> >Hi Guangrong,
> >
> >as such mapping information between virtual mmio to physical mmio is only available
> >at runtime.
> 
> Sorry, i do not know what the different between mmap() and the time VM actually
> accesses the memory for your case. Could you please more detail?

Hi Guangrong,

Sure. The mmap() gets called by qemu or any VFIO API userspace consumer when
setting up the virtual mmio, at that moment nobody has any knowledge about how
the physical mmio gets virtualized.

When the vm (or application if we don't want to limit ourselves to vmm term)
starts, the virtual and physical mmio gets mapped by mpci kernel module with the
help from vendor supplied mediated host driver according to the hw resource
assigned to this vm / application.

> 
> >
> >>
> >>Why the memory mapped by this mmap() is not a portion of MMIO from
> >>underlayer physical device? If it is a valid system memory, is this interface
> >>really needed to implemented in vfio? (you at least need to set VM_MIXEDMAP
> >>if it mixed system memory with MMIO)
> >>
> >
> >It actually is a portion of the physical mmio which is set by vfio mmap.
> 
> So i do not think we need to care its refcount, i,e, we can consider it as reserved_pfn,
> Paolo?
> 
> >
> >>IIUC, the kernel assumes that VM_PFNMAP is a continuous memory, e.g, like
> >>current KVM and vaddr_get_pfn() in vfio, but it seems nvdia's patchset
> >>breaks this semantic as ops->validate_map_request() can adjust the physical
> >>address arbitrarily. (again, the name 'validate' should be changed to match
> >>the thing as it is really doing)
> >
> >The vgpu api will allow you to adjust the target mmio address and the size via
> >validate_map_request, but it is still physical contiguous as <start_pfn, size>.
> 
> Okay, the interface confused us, maybe this interface need to be cooked to reflect
> to this fact.

Sure. We can address this in the RFC mediated device thread.

Thanks,
Neo

> 
> Thanks!
> 

[toc] | [prev] | [next] | [standalone]


#1436267

FromNeo Jia <cjia@nvidia.com>
Date2016-07-04 17:40 +0200
Message-ID<rRe8x-7mI-9@gated-at.bofh.it>
In reply to#1436239
On Mon, Jul 04, 2016 at 06:16:46PM +0800, Xiao Guangrong wrote:
> 
> 
> On 07/04/2016 05:16 PM, Neo Jia wrote:
> >On Mon, Jul 04, 2016 at 04:45:05PM +0800, Xiao Guangrong wrote:
> >>
> >>
> >>On 07/04/2016 04:41 PM, Neo Jia wrote:
> >>>On Mon, Jul 04, 2016 at 04:19:20PM +0800, Xiao Guangrong wrote:
> >>>>
> >>>>
> >>>>On 07/04/2016 03:53 PM, Neo Jia wrote:
> >>>>>On Mon, Jul 04, 2016 at 03:37:35PM +0800, Xiao Guangrong wrote:
> >>>>>>
> >>>>>>
> >>>>>>On 07/04/2016 03:03 PM, Neo Jia wrote:
> >>>>>>>On Mon, Jul 04, 2016 at 02:39:22PM +0800, Xiao Guangrong wrote:
> >>>>>>>>
> >>>>>>>>
> >>>>>>>>On 06/30/2016 09:01 PM, Paolo Bonzini wrote:
> >>>>>>>>>The vGPU folks would like to trap the first access to a BAR by setting
> >>>>>>>>>vm_ops on the VMAs produced by mmap-ing a VFIO device.  The fault handler
> >>>>>>>>>then can use remap_pfn_range to place some non-reserved pages in the VMA.
> >>>>>>>>
> >>>>>>>>Why does it require fetching the pfn when the fault is triggered rather
> >>>>>>>>than when mmap() is called?
> >>>>>>>
> >>>>>>>Hi Guangrong,
> >>>>>>>
> >>>>>>>as such mapping information between virtual mmio to physical mmio is only available
> >>>>>>>at runtime.
> >>>>>>
> >>>>>>Sorry, i do not know what the different between mmap() and the time VM actually
> >>>>>>accesses the memory for your case. Could you please more detail?
> >>>>>
> >>>>>Hi Guangrong,
> >>>>>
> >>>>>Sure. The mmap() gets called by qemu or any VFIO API userspace consumer when
> >>>>>setting up the virtual mmio, at that moment nobody has any knowledge about how
> >>>>>the physical mmio gets virtualized.
> >>>>>
> >>>>>When the vm (or application if we don't want to limit ourselves to vmm term)
> >>>>>starts, the virtual and physical mmio gets mapped by mpci kernel module with the
> >>>>>help from vendor supplied mediated host driver according to the hw resource
> >>>>>assigned to this vm / application.
> >>>>
> >>>>Thanks for your expiation.
> >>>>
> >>>>It sounds like a strategy of resource allocation, you delay the allocation until VM really
> >>>>accesses it, right?
> >>>
> >>>Yes, that is where the fault handler inside mpci code comes to the picture.
> >>
> >>
> >>I am not sure this strategy is good. The instance is successfully created, and it is started
> >>successful, but the VM is crashed due to the resource of that instance is not enough. That sounds
> >>unreasonable.
> >
> >
> >Sorry, I think I misread the "allocation" as "mapping". We only delay the
> >cpu mapping, not the allocation.
> 
> So how to understand your statement:
> "at that moment nobody has any knowledge about how the physical mmio gets virtualized"
> 
> The resource, physical MMIO region, has been allocated, why we do not know the physical
> address mapped to the VM?
> 

From a device driver point of view, the physical mmio region never gets allocated until 
the corresponding resource is requested by clients and granted by the mediated device driver. 

The resource here is the internal hw resource.

"at that moment" == vfio client triggers mmap() call.

Thanks,
Neo

> 
> --
> To unsubscribe from this list: send the line "unsubscribe kvm" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html

[toc] | [prev] | [next] | [standalone]


#1436716

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-05 03:30 +0200
Message-ID<rRnlw-4AT-23@gated-at.bofh.it>
In reply to#1436267

On 07/04/2016 11:33 PM, Neo Jia wrote:

>>>
>>> Sorry, I think I misread the "allocation" as "mapping". We only delay the
>>> cpu mapping, not the allocation.
>>
>> So how to understand your statement:
>> "at that moment nobody has any knowledge about how the physical mmio gets virtualized"
>>
>> The resource, physical MMIO region, has been allocated, why we do not know the physical
>> address mapped to the VM?
>>
>
>>From a device driver point of view, the physical mmio region never gets allocated until
> the corresponding resource is requested by clients and granted by the mediated device driver.

Hmm... but you told me that you did not delay the allocation. :(

So it returns to my original question: why not allocate the physical mmio region in mmap()?

[toc] | [prev] | [next] | [standalone]


#1436721

FromNeo Jia <cjia@nvidia.com>
Date2016-07-05 03:40 +0200
Message-ID<rRnvb-4E3-5@gated-at.bofh.it>
In reply to#1436716
On Tue, Jul 05, 2016 at 09:19:40AM +0800, Xiao Guangrong wrote:
> 
> 
> On 07/04/2016 11:33 PM, Neo Jia wrote:
> 
> >>>
> >>>Sorry, I think I misread the "allocation" as "mapping". We only delay the
> >>>cpu mapping, not the allocation.
> >>
> >>So how to understand your statement:
> >>"at that moment nobody has any knowledge about how the physical mmio gets virtualized"
> >>
> >>The resource, physical MMIO region, has been allocated, why we do not know the physical
> >>address mapped to the VM?
> >>
> >
> >>From a device driver point of view, the physical mmio region never gets allocated until
> >the corresponding resource is requested by clients and granted by the mediated device driver.
> 
> Hmm... but you told me that you did not delay the allocation. :(

Hi Guangrong,

The allocation here is the allocation of device resource, and the only way to
access that kind of device resource is via a mmio region of some pages there.

For example, if VM needs resource A, and the only way to access resource A is
via some kind of device memory at mmio address X.

So, we never defer the allocation request during runtime, we just setup the
CPU mapping later when it actually gets accessed.

> 
> So it returns to my original question: why not allocate the physical mmio region in mmap()?
> 

Without running anything inside the VM, how do you know how the hw resource gets
allocated, therefore no knowledge of the use of mmio region.

Thanks,
Neo

> 
> 
> 
> --
> To unsubscribe from this list: send the line "unsubscribe kvm" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html

[toc] | [prev] | [next] | [standalone]


#1436747

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-05 06:10 +0200
Message-ID<rRpQl-6gQ-9@gated-at.bofh.it>
In reply to#1436721

On 07/05/2016 09:35 AM, Neo Jia wrote:
> On Tue, Jul 05, 2016 at 09:19:40AM +0800, Xiao Guangrong wrote:
>>
>>
>> On 07/04/2016 11:33 PM, Neo Jia wrote:
>>
>>>>>
>>>>> Sorry, I think I misread the "allocation" as "mapping". We only delay the
>>>>> cpu mapping, not the allocation.
>>>>
>>>> So how to understand your statement:
>>>> "at that moment nobody has any knowledge about how the physical mmio gets virtualized"
>>>>
>>>> The resource, physical MMIO region, has been allocated, why we do not know the physical
>>>> address mapped to the VM?
>>>>
>>>
>>> >From a device driver point of view, the physical mmio region never gets allocated until
>>> the corresponding resource is requested by clients and granted by the mediated device driver.
>>
>> Hmm... but you told me that you did not delay the allocation. :(
>
> Hi Guangrong,
>
> The allocation here is the allocation of device resource, and the only way to
> access that kind of device resource is via a mmio region of some pages there.
>
> For example, if VM needs resource A, and the only way to access resource A is
> via some kind of device memory at mmio address X.
>
> So, we never defer the allocation request during runtime, we just setup the
> CPU mapping later when it actually gets accessed.
>
>>
>> So it returns to my original question: why not allocate the physical mmio region in mmap()?
>>
>
> Without running anything inside the VM, how do you know how the hw resource gets
> allocated, therefore no knowledge of the use of mmio region.

The allocation and mapping can be two independent processes:
- the first process is just allocation. The MMIO region is allocated from physical
   hardware and this region is mapped into _QEMU's_ arbitrary virtual address by mmap().
   At this time, VM can not actually use this resource.

- the second process is mapping. When VM enable this region, e.g, it enables the
   PCI BAR, then QEMU maps its virtual address returned by mmap() to VM's physical
   memory. After that, VM can access this region.

The second process is completed handled in userspace, that means, the mediated
device driver needn't care how the resource is mapped into VM.

This is how QEMU/VFIO currently works, could you please tell me the special points
of your solution comparing with current QEMU/VFIO and why current model can not fit
your requirement? So that we can better understand your scenario?

[toc] | [prev] | [next] | [standalone]


#1436755

FromNeo Jia <cjia@nvidia.com>
Date2016-07-05 07:20 +0200
Message-ID<rRqW5-6Xe-3@gated-at.bofh.it>
In reply to#1436747
On Tue, Jul 05, 2016 at 12:02:42PM +0800, Xiao Guangrong wrote:
> 
> 
> On 07/05/2016 09:35 AM, Neo Jia wrote:
> >On Tue, Jul 05, 2016 at 09:19:40AM +0800, Xiao Guangrong wrote:
> >>
> >>
> >>On 07/04/2016 11:33 PM, Neo Jia wrote:
> >>
> >>>>>
> >>>>>Sorry, I think I misread the "allocation" as "mapping". We only delay the
> >>>>>cpu mapping, not the allocation.
> >>>>
> >>>>So how to understand your statement:
> >>>>"at that moment nobody has any knowledge about how the physical mmio gets virtualized"
> >>>>
> >>>>The resource, physical MMIO region, has been allocated, why we do not know the physical
> >>>>address mapped to the VM?
> >>>>
> >>>
> >>>>From a device driver point of view, the physical mmio region never gets allocated until
> >>>the corresponding resource is requested by clients and granted by the mediated device driver.
> >>
> >>Hmm... but you told me that you did not delay the allocation. :(
> >
> >Hi Guangrong,
> >
> >The allocation here is the allocation of device resource, and the only way to
> >access that kind of device resource is via a mmio region of some pages there.
> >
> >For example, if VM needs resource A, and the only way to access resource A is
> >via some kind of device memory at mmio address X.
> >
> >So, we never defer the allocation request during runtime, we just setup the
> >CPU mapping later when it actually gets accessed.
> >
> >>
> >>So it returns to my original question: why not allocate the physical mmio region in mmap()?
> >>
> >
> >Without running anything inside the VM, how do you know how the hw resource gets
> >allocated, therefore no knowledge of the use of mmio region.
> 
> The allocation and mapping can be two independent processes:
> - the first process is just allocation. The MMIO region is allocated from physical
>   hardware and this region is mapped into _QEMU's_ arbitrary virtual address by mmap().
>   At this time, VM can not actually use this resource.
> 
> - the second process is mapping. When VM enable this region, e.g, it enables the
>   PCI BAR, then QEMU maps its virtual address returned by mmap() to VM's physical
>   memory. After that, VM can access this region.
> 
> The second process is completed handled in userspace, that means, the mediated
> device driver needn't care how the resource is mapped into VM.

In your example, you are still picturing it as VFIO direct assign, but the solution we are 
talking here is mediated passthru via VFIO framework to virtualize DMA devices without SR-IOV.

(Just for completeness, if you really want to use a device in above example as
VFIO passthru, the second step is not completely handled in userspace, it is actually the guest
driver who will allocate and setup the proper hw resource which will later ready
for you to access via some mmio pages.)

> 
> This is how QEMU/VFIO currently works, could you please tell me the special points
> of your solution comparing with current QEMU/VFIO and why current model can not fit
> your requirement? So that we can better understand your scenario?

The scenario I am describing here is mediated passthru case, but what you are
describing here (more or less) is VFIO direct assigned case. It is different in several
areas, but major difference related to this topic here is:

1) In VFIO direct assigned case, the device (and its resource) is completely owned by the VM
therefore its mmio region can be mapped directly into the VM during the VFIO mmap() call as
there is no resource sharing among VMs and there is no mediated device driver on
the host to manage such resource, so it is completely owned by the guest.

2) In mediated passthru case, multiple VMs are sharing the same physical device, so how
the HW resource gets allocated is completely decided by the guest and host device driver of 
the virtualized DMA device, here is the GPU, same as the MMIO pages used to access those Hw resource.

Thanks,
Neo

[toc] | [prev] | [next] | [standalone]


#1436795

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-05 08:40 +0200
Message-ID<rRsbv-7C5-13@gated-at.bofh.it>
In reply to#1436755

On 07/05/2016 01:16 PM, Neo Jia wrote:
> On Tue, Jul 05, 2016 at 12:02:42PM +0800, Xiao Guangrong wrote:
>>
>>
>> On 07/05/2016 09:35 AM, Neo Jia wrote:
>>> On Tue, Jul 05, 2016 at 09:19:40AM +0800, Xiao Guangrong wrote:
>>>>
>>>>
>>>> On 07/04/2016 11:33 PM, Neo Jia wrote:
>>>>
>>>>>>>
>>>>>>> Sorry, I think I misread the "allocation" as "mapping". We only delay the
>>>>>>> cpu mapping, not the allocation.
>>>>>>
>>>>>> So how to understand your statement:
>>>>>> "at that moment nobody has any knowledge about how the physical mmio gets virtualized"
>>>>>>
>>>>>> The resource, physical MMIO region, has been allocated, why we do not know the physical
>>>>>> address mapped to the VM?
>>>>>>
>>>>>
>>>>> >From a device driver point of view, the physical mmio region never gets allocated until
>>>>> the corresponding resource is requested by clients and granted by the mediated device driver.
>>>>
>>>> Hmm... but you told me that you did not delay the allocation. :(
>>>
>>> Hi Guangrong,
>>>
>>> The allocation here is the allocation of device resource, and the only way to
>>> access that kind of device resource is via a mmio region of some pages there.
>>>
>>> For example, if VM needs resource A, and the only way to access resource A is
>>> via some kind of device memory at mmio address X.
>>>
>>> So, we never defer the allocation request during runtime, we just setup the
>>> CPU mapping later when it actually gets accessed.
>>>
>>>>
>>>> So it returns to my original question: why not allocate the physical mmio region in mmap()?
>>>>
>>>
>>> Without running anything inside the VM, how do you know how the hw resource gets
>>> allocated, therefore no knowledge of the use of mmio region.
>>
>> The allocation and mapping can be two independent processes:
>> - the first process is just allocation. The MMIO region is allocated from physical
>>    hardware and this region is mapped into _QEMU's_ arbitrary virtual address by mmap().
>>    At this time, VM can not actually use this resource.
>>
>> - the second process is mapping. When VM enable this region, e.g, it enables the
>>    PCI BAR, then QEMU maps its virtual address returned by mmap() to VM's physical
>>    memory. After that, VM can access this region.
>>
>> The second process is completed handled in userspace, that means, the mediated
>> device driver needn't care how the resource is mapped into VM.
>
> In your example, you are still picturing it as VFIO direct assign, but the solution we are
> talking here is mediated passthru via VFIO framework to virtualize DMA devices without SR-IOV.
>

Please see my comments below.

> (Just for completeness, if you really want to use a device in above example as
> VFIO passthru, the second step is not completely handled in userspace, it is actually the guest
> driver who will allocate and setup the proper hw resource which will later ready
> for you to access via some mmio pages.)

Hmm... i always treat the VM as userspace.

>
>>
>> This is how QEMU/VFIO currently works, could you please tell me the special points
>> of your solution comparing with current QEMU/VFIO and why current model can not fit
>> your requirement? So that we can better understand your scenario?
>
> The scenario I am describing here is mediated passthru case, but what you are
> describing here (more or less) is VFIO direct assigned case. It is different in several
> areas, but major difference related to this topic here is:
>
> 1) In VFIO direct assigned case, the device (and its resource) is completely owned by the VM
> therefore its mmio region can be mapped directly into the VM during the VFIO mmap() call as
> there is no resource sharing among VMs and there is no mediated device driver on
> the host to manage such resource, so it is completely owned by the guest.

I understand this difference, However, as you told to me that the MMIO region allocated for the
VM is continuous, so i assume the portion of physical MMIO region is completely owned by guest.
The only difference i can see is mediated device driver need to allocate that region.

>
> 2) In mediated passthru case, multiple VMs are sharing the same physical device, so how
> the HW resource gets allocated is completely decided by the guest and host device driver of
> the virtualized DMA device, here is the GPU, same as the MMIO pages used to access those Hw resource.

I can not see what guest's affair is here, look at your code, you cooked the fault handler like
this:

+               ret = parent->ops->validate_map_request(mdev, virtaddr,
+                                                        &pgoff, &req_size,
+                                                        &pg_prot);

Please tell me what information is got from guest? All these info can be found at the time of
mmap().

[toc] | [prev] | [next] | [standalone]


#1436817

FromNeo Jia <cjia@nvidia.com>
Date2016-07-05 09:40 +0200
Message-ID<rRt7A-8au-11@gated-at.bofh.it>
In reply to#1436795
On Tue, Jul 05, 2016 at 02:26:46PM +0800, Xiao Guangrong wrote:
> 
> 
> On 07/05/2016 01:16 PM, Neo Jia wrote:
> >On Tue, Jul 05, 2016 at 12:02:42PM +0800, Xiao Guangrong wrote:
> >>
> >>
> >>On 07/05/2016 09:35 AM, Neo Jia wrote:
> >>>On Tue, Jul 05, 2016 at 09:19:40AM +0800, Xiao Guangrong wrote:
> >>>>
> >>>>
> >>>>On 07/04/2016 11:33 PM, Neo Jia wrote:
> >>>>
> >>>>>>>
> >>>>>>>Sorry, I think I misread the "allocation" as "mapping". We only delay the
> >>>>>>>cpu mapping, not the allocation.
> >>>>>>
> >>>>>>So how to understand your statement:
> >>>>>>"at that moment nobody has any knowledge about how the physical mmio gets virtualized"
> >>>>>>
> >>>>>>The resource, physical MMIO region, has been allocated, why we do not know the physical
> >>>>>>address mapped to the VM?
> >>>>>>
> >>>>>
> >>>>>>From a device driver point of view, the physical mmio region never gets allocated until
> >>>>>the corresponding resource is requested by clients and granted by the mediated device driver.
> >>>>
> >>>>Hmm... but you told me that you did not delay the allocation. :(
> >>>
> >>>Hi Guangrong,
> >>>
> >>>The allocation here is the allocation of device resource, and the only way to
> >>>access that kind of device resource is via a mmio region of some pages there.
> >>>
> >>>For example, if VM needs resource A, and the only way to access resource A is
> >>>via some kind of device memory at mmio address X.
> >>>
> >>>So, we never defer the allocation request during runtime, we just setup the
> >>>CPU mapping later when it actually gets accessed.
> >>>
> >>>>
> >>>>So it returns to my original question: why not allocate the physical mmio region in mmap()?
> >>>>
> >>>
> >>>Without running anything inside the VM, how do you know how the hw resource gets
> >>>allocated, therefore no knowledge of the use of mmio region.
> >>
> >>The allocation and mapping can be two independent processes:
> >>- the first process is just allocation. The MMIO region is allocated from physical
> >>   hardware and this region is mapped into _QEMU's_ arbitrary virtual address by mmap().
> >>   At this time, VM can not actually use this resource.
> >>
> >>- the second process is mapping. When VM enable this region, e.g, it enables the
> >>   PCI BAR, then QEMU maps its virtual address returned by mmap() to VM's physical
> >>   memory. After that, VM can access this region.
> >>
> >>The second process is completed handled in userspace, that means, the mediated
> >>device driver needn't care how the resource is mapped into VM.
> >
> >In your example, you are still picturing it as VFIO direct assign, but the solution we are
> >talking here is mediated passthru via VFIO framework to virtualize DMA devices without SR-IOV.
> >
> 
> Please see my comments below.
> 
> >(Just for completeness, if you really want to use a device in above example as
> >VFIO passthru, the second step is not completely handled in userspace, it is actually the guest
> >driver who will allocate and setup the proper hw resource which will later ready
> >for you to access via some mmio pages.)
> 
> Hmm... i always treat the VM as userspace.

It is OK to treat VM as userspace, but I think it is better to put out details
so we are always on the same page.

> 
> >
> >>
> >>This is how QEMU/VFIO currently works, could you please tell me the special points
> >>of your solution comparing with current QEMU/VFIO and why current model can not fit
> >>your requirement? So that we can better understand your scenario?
> >
> >The scenario I am describing here is mediated passthru case, but what you are
> >describing here (more or less) is VFIO direct assigned case. It is different in several
> >areas, but major difference related to this topic here is:
> >
> >1) In VFIO direct assigned case, the device (and its resource) is completely owned by the VM
> >therefore its mmio region can be mapped directly into the VM during the VFIO mmap() call as
> >there is no resource sharing among VMs and there is no mediated device driver on
> >the host to manage such resource, so it is completely owned by the guest.
> 
> I understand this difference, However, as you told to me that the MMIO region allocated for the
> VM is continuous, so i assume the portion of physical MMIO region is completely owned by guest.
> The only difference i can see is mediated device driver need to allocate that region.

It is physically contiguous but it is done during the runtime, physically contiguous doesn't mean 
static partition at boot time. And only during runtime, the proper HW resource will be requested therefore 
the right portion of MMIO region will be granted by the mediated device driver on the host.

Also, the physically contiguous doesn't mean the guest and host mmio is 1:1
always. You can have a 8GB host physical mmio while the guest will only have
256MB.

> 
> >
> >2) In mediated passthru case, multiple VMs are sharing the same physical device, so how
> >the HW resource gets allocated is completely decided by the guest and host device driver of
> >the virtualized DMA device, here is the GPU, same as the MMIO pages used to access those Hw resource.
> 
> I can not see what guest's affair is here, look at your code, you cooked the fault handler like
> this:

You shouldn't as that depends on how different devices are getting
para-virtualized by their own implementations.

> 
> +               ret = parent->ops->validate_map_request(mdev, virtaddr,
> +                                                        &pgoff, &req_size,
> +                                                        &pg_prot);
> 
> Please tell me what information is got from guest? All these info can be found at the time of
> mmap().

The virtaddr is the guest mmio address that triggers this fault, which will be
used for the mediated device driver to locate the resource that he has previously allocated.

Then the req_size and pgoff will both come from the mediated device driver based on his internal book 
keeping of the hw resource allocation, which is only available during runtime. And such book keeping 
can be built part of para-virtualization scheme between guest and host device driver.

None of such information is available at VFIO mmap() time. For example, several VMs
are sharing the same physical device to provide mediated access. All VMs will
call the VFIO mmap() on their virtual BAR as part of QEMU vfio/pci initialization 
process, at that moment, we definitely can't mmap the entire physical MMIO
into both VM blindly for obvious reason. 

And the pgoff will be different for different VMs as they will not have access
to others hw resource for the same reason.

Thanks,
Neo

[toc] | [prev] | [next] | [standalone]


#1436861

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-07-05 11:10 +0200
Message-ID<rRuwG-MJ-29@gated-at.bofh.it>
In reply to#1436817

On 07/05/2016 03:30 PM, Neo Jia wrote:

>>
>>> (Just for completeness, if you really want to use a device in above example as
>>> VFIO passthru, the second step is not completely handled in userspace, it is actually the guest
>>> driver who will allocate and setup the proper hw resource which will later ready
>>> for you to access via some mmio pages.)
>>
>> Hmm... i always treat the VM as userspace.
>
> It is OK to treat VM as userspace, but I think it is better to put out details
> so we are always on the same page.
>

Okay. I should pay more attention on it when i discuss with the driver people. :)

>>
>>>
>>>>
>>>> This is how QEMU/VFIO currently works, could you please tell me the special points
>>>> of your solution comparing with current QEMU/VFIO and why current model can not fit
>>>> your requirement? So that we can better understand your scenario?
>>>
>>> The scenario I am describing here is mediated passthru case, but what you are
>>> describing here (more or less) is VFIO direct assigned case. It is different in several
>>> areas, but major difference related to this topic here is:
>>>
>>> 1) In VFIO direct assigned case, the device (and its resource) is completely owned by the VM
>>> therefore its mmio region can be mapped directly into the VM during the VFIO mmap() call as
>>> there is no resource sharing among VMs and there is no mediated device driver on
>>> the host to manage such resource, so it is completely owned by the guest.
>>
>> I understand this difference, However, as you told to me that the MMIO region allocated for the
>> VM is continuous, so i assume the portion of physical MMIO region is completely owned by guest.
>> The only difference i can see is mediated device driver need to allocate that region.
>
> It is physically contiguous but it is done during the runtime, physically contiguous doesn't mean
> static partition at boot time. And only during runtime, the proper HW resource will be requested therefore
> the right portion of MMIO region will be granted by the mediated device driver on the host.

Okay. This is your implantation design rather than the hardware limitation, right?

For example, if the instance require 512M memory (the size can be specified by QEMU
command line), it can tell its requirement to the mediated device driver via create()
interface, then the driver can allocate then memory for this instance before it is running.

Theoretically, the hardware is able to do memory management as this style, but for some
reasons you choose allocating memory in the runtime. right? If my understanding is right,
could you please tell us what benefit you want to get from this running-allocation style?

>
> Also, the physically contiguous doesn't mean the guest and host mmio is 1:1
> always. You can have a 8GB host physical mmio while the guest will only have
> 256MB.

Thanks for your patience, it is clearer to me and at least i am able to try to guess the
whole picture. :)

>
>>
>>>
>>> 2) In mediated passthru case, multiple VMs are sharing the same physical device, so how
>>> the HW resource gets allocated is completely decided by the guest and host device driver of
>>> the virtualized DMA device, here is the GPU, same as the MMIO pages used to access those Hw resource.
>>
>> I can not see what guest's affair is here, look at your code, you cooked the fault handler like
>> this:
>
> You shouldn't as that depends on how different devices are getting
> para-virtualized by their own implementations.
>

PV method. It is interesting. More comments below.

>>
>> +               ret = parent->ops->validate_map_request(mdev, virtaddr,
>> +                                                        &pgoff, &req_size,
>> +                                                        &pg_prot);
>>
>> Please tell me what information is got from guest? All these info can be found at the time of
>> mmap().
>
> The virtaddr is the guest mmio address that triggers this fault, which will be
> used for the mediated device driver to locate the resource that he has previously allocated.

The virtaddr is not the guest mmio address, it is the virtual address of QEMU. vfio is not
able to figure out the guest mmio address as the mapping is handled in userspace as we
discussed above.

And we can get the virtaddr from [vma->start, vma->end) when we do mmap().

>
> Then the req_size and pgoff will both come from the mediated device driver based on his internal book
> keeping of the hw resource allocation, which is only available during runtime. And such book keeping
> can be built part of para-virtualization scheme between guest and host device driver.
>

I am talking the parameters you passed to validate_map_request(). req_size is calculated like this:

+       offset   = virtaddr - vma->vm_start;
+       phyaddr  = (vma->vm_pgoff << PAGE_SHIFT) + offset;
+       pgoff    = phyaddr >> PAGE_SHIFT;

All these info is from vma which is available in mmmap().

pgoff is got from:
+       pg_prot  = vma->vm_page_prot;
that is also available in mmap().

> None of such information is available at VFIO mmap() time. For example, several VMs
> are sharing the same physical device to provide mediated access. All VMs will
> call the VFIO mmap() on their virtual BAR as part of QEMU vfio/pci initialization
> process, at that moment, we definitely can't mmap the entire physical MMIO
> into both VM blindly for obvious reason.
>

mmap() carries @length information, so you only need to allocate the specified size
(corresponding to @length) of memory for them.

Now i guess there has some operations, e.g, PV operations, between mmap() and memory fault,
these operations tell the mediated device driver how to allocate memory for this instance.
Right?

[toc] | [prev] | [next] | [standalone]


Page 1 of 3  [1] 2 3  Next page →

Back to top | Article view | linux.kernel


csiph-web