Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1522815 > unrolled thread
| Started by | Kirti Wankhede <kwankhede@nvidia.com> |
|---|---|
| First post | 2016-11-15 16:30 +0100 |
| Last post | 2016-11-16 16:30 +0100 |
| Articles | 10 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
[PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP Kirti Wankhede <kwankhede@nvidia.com> - 2016-11-15 16:30 +0100
Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP Alex Williamson <alex.williamson@redhat.com> - 2016-11-15 23:30 +0100
Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP Kirti Wankhede <kwankhede@nvidia.com> - 2016-11-16 03:50 +0100
Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP Alex Williamson <alex.williamson@redhat.com> - 2016-11-16 04:20 +0100
Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP Alex Williamson <alex.williamson@redhat.com> - 2016-11-16 04:30 +0100
Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP Kirti Wankhede <kwankhede@nvidia.com> - 2016-11-16 04:50 +0100
Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP Alex Williamson <alex.williamson@redhat.com> - 2016-11-16 05:00 +0100
Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP Kirti Wankhede <kwankhede@nvidia.com> - 2016-11-16 05:20 +0100
Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP Alex Williamson <alex.williamson@redhat.com> - 2016-11-16 05:40 +0100
Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP Kirti Wankhede <kwankhede@nvidia.com> - 2016-11-16 16:30 +0100
| From | Kirti Wankhede <kwankhede@nvidia.com> |
|---|---|
| Date | 2016-11-15 16:30 +0100 |
| Subject | [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP |
| Message-ID | <sDNQl-6Go-3@gated-at.bofh.it> |
Added blocking notifier to IOMMU TYPE1 driver to notify vendor drivers
about DMA_UNMAP.
Exported two APIs vfio_register_notifier() and vfio_unregister_notifier().
Notifier should be registered, if external user wants to use
vfio_pin_pages()/vfio_unpin_pages() APIs to pin/unpin pages.
Vendor driver should use VFIO_IOMMU_NOTIFY_DMA_UNMAP action to invalidate
mappings.
Signed-off-by: Kirti Wankhede <kwankhede@nvidia.com>
Signed-off-by: Neo Jia <cjia@nvidia.com>
Change-Id: I5910d0024d6be87f3e8d3e0ca0eaeaaa0b17f271
---
drivers/vfio/vfio.c | 73 +++++++++++++++++++++++++++++++++++++++++
drivers/vfio/vfio_iommu_type1.c | 63 +++++++++++++++++++++++++++++------
include/linux/vfio.h | 11 +++++++
3 files changed, 137 insertions(+), 10 deletions(-)
diff --git a/drivers/vfio/vfio.c b/drivers/vfio/vfio.c
index 3bf8a01bf67b..fa121d983991 100644
--- a/drivers/vfio/vfio.c
+++ b/drivers/vfio/vfio.c
@@ -1902,6 +1902,79 @@ err_unpin_pages:
}
EXPORT_SYMBOL(vfio_unpin_pages);
+int vfio_register_notifier(struct device *dev, struct notifier_block *nb)
+{
+ struct vfio_container *container;
+ struct vfio_group *group;
+ struct vfio_iommu_driver *driver;
+ ssize_t ret;
+
+ if (!dev || !nb)
+ return -EINVAL;
+
+ group = vfio_group_get_from_dev(dev);
+ if (IS_ERR(group))
+ return PTR_ERR(group);
+
+ ret = vfio_group_add_container_user(group);
+ if (ret)
+ goto err_register_nb;
+
+ container = group->container;
+ down_read(&container->group_lock);
+
+ driver = container->iommu_driver;
+ if (likely(driver && driver->ops->register_notifier))
+ ret = driver->ops->register_notifier(container->iommu_data, nb);
+ else
+ ret = -ENOTTY;
+
+ up_read(&container->group_lock);
+ vfio_group_try_dissolve_container(group);
+
+err_register_nb:
+ vfio_group_put(group);
+ return ret;
+}
+EXPORT_SYMBOL(vfio_register_notifier);
+
+int vfio_unregister_notifier(struct device *dev, struct notifier_block *nb)
+{
+ struct vfio_container *container;
+ struct vfio_group *group;
+ struct vfio_iommu_driver *driver;
+ ssize_t ret;
+
+ if (!dev || !nb)
+ return -EINVAL;
+
+ group = vfio_group_get_from_dev(dev);
+ if (IS_ERR(group))
+ return PTR_ERR(group);
+
+ ret = vfio_group_add_container_user(group);
+ if (ret)
+ goto err_unregister_nb;
+
+ container = group->container;
+ down_read(&container->group_lock);
+
+ driver = container->iommu_driver;
+ if (likely(driver && driver->ops->unregister_notifier))
+ ret = driver->ops->unregister_notifier(container->iommu_data,
+ nb);
+ else
+ ret = -ENOTTY;
+
+ up_read(&container->group_lock);
+ vfio_group_try_dissolve_container(group);
+
+err_unregister_nb:
+ vfio_group_put(group);
+ return ret;
+}
+EXPORT_SYMBOL(vfio_unregister_notifier);
+
/**
* Module/class support
*/
diff --git a/drivers/vfio/vfio_iommu_type1.c b/drivers/vfio/vfio_iommu_type1.c
index 0de7c20f66b1..c45a4822784e 100644
--- a/drivers/vfio/vfio_iommu_type1.c
+++ b/drivers/vfio/vfio_iommu_type1.c
@@ -38,6 +38,7 @@
#include <linux/workqueue.h>
#include <linux/pid_namespace.h>
#include <linux/mdev.h>
+#include <linux/notifier.h>
#define DRIVER_VERSION "0.2"
#define DRIVER_AUTHOR "Alex Williamson <alex.williamson@redhat.com>"
@@ -60,6 +61,7 @@ struct vfio_iommu {
struct vfio_domain *external_domain; /* domain for external user */
struct mutex lock;
struct rb_root dma_list;
+ struct blocking_notifier_head notifier;
bool v2;
bool nesting;
};
@@ -571,7 +573,8 @@ static int vfio_iommu_type1_pin_pages(void *iommu_data,
mutex_lock(&iommu->lock);
- if (!iommu->external_domain) {
+ /* Fail if notifier list is empty */
+ if ((!iommu->external_domain) || (!iommu->notifier.head)) {
ret = -EINVAL;
goto pin_done;
}
@@ -854,7 +857,28 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
*/
if (dma->task->mm != current->mm)
break;
+
unmapped += dma->size;
+
+ if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
+ struct vfio_iommu_type1_dma_unmap nb_unmap;
+
+ nb_unmap.iova = dma->iova;
+ nb_unmap.size = dma->size;
+
+ /*
+ * Notifier callback would call vfio_unpin_pages() which
+ * would acquire iommu->lock. Release lock here and
+ * reacquire it again.
+ */
+ mutex_unlock(&iommu->lock);
+ blocking_notifier_call_chain(&iommu->notifier,
+ VFIO_IOMMU_NOTIFY_DMA_UNMAP,
+ &nb_unmap);
+ mutex_lock(&iommu->lock);
+ if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
+ break;
+ }
vfio_remove_dma(iommu, dma);
}
@@ -1439,6 +1463,7 @@ static void *vfio_iommu_type1_open(unsigned long arg)
INIT_LIST_HEAD(&iommu->domain_list);
iommu->dma_list = RB_ROOT;
mutex_init(&iommu->lock);
+ BLOCKING_INIT_NOTIFIER_HEAD(&iommu->notifier);
return iommu;
}
@@ -1574,16 +1599,34 @@ static long vfio_iommu_type1_ioctl(void *iommu_data,
return -ENOTTY;
}
+static int vfio_iommu_type1_register_notifier(void *iommu_data,
+ struct notifier_block *nb)
+{
+ struct vfio_iommu *iommu = iommu_data;
+
+ return blocking_notifier_chain_register(&iommu->notifier, nb);
+}
+
+static int vfio_iommu_type1_unregister_notifier(void *iommu_data,
+ struct notifier_block *nb)
+{
+ struct vfio_iommu *iommu = iommu_data;
+
+ return blocking_notifier_chain_unregister(&iommu->notifier, nb);
+}
+
static const struct vfio_iommu_driver_ops vfio_iommu_driver_ops_type1 = {
- .name = "vfio-iommu-type1",
- .owner = THIS_MODULE,
- .open = vfio_iommu_type1_open,
- .release = vfio_iommu_type1_release,
- .ioctl = vfio_iommu_type1_ioctl,
- .attach_group = vfio_iommu_type1_attach_group,
- .detach_group = vfio_iommu_type1_detach_group,
- .pin_pages = vfio_iommu_type1_pin_pages,
- .unpin_pages = vfio_iommu_type1_unpin_pages,
+ .name = "vfio-iommu-type1",
+ .owner = THIS_MODULE,
+ .open = vfio_iommu_type1_open,
+ .release = vfio_iommu_type1_release,
+ .ioctl = vfio_iommu_type1_ioctl,
+ .attach_group = vfio_iommu_type1_attach_group,
+ .detach_group = vfio_iommu_type1_detach_group,
+ .pin_pages = vfio_iommu_type1_pin_pages,
+ .unpin_pages = vfio_iommu_type1_unpin_pages,
+ .register_notifier = vfio_iommu_type1_register_notifier,
+ .unregister_notifier = vfio_iommu_type1_unregister_notifier,
};
static int __init vfio_iommu_type1_init(void)
diff --git a/include/linux/vfio.h b/include/linux/vfio.h
index 420cdc928786..997442398c09 100644
--- a/include/linux/vfio.h
+++ b/include/linux/vfio.h
@@ -80,6 +80,10 @@ struct vfio_iommu_driver_ops {
unsigned long *phys_pfn);
int (*unpin_pages)(void *iommu_data,
unsigned long *user_pfn, int npage);
+ int (*register_notifier)(void *iommu_data,
+ struct notifier_block *nb);
+ int (*unregister_notifier)(void *iommu_data,
+ struct notifier_block *nb);
};
extern int vfio_register_iommu_driver(const struct vfio_iommu_driver_ops *ops);
@@ -139,6 +143,13 @@ extern int vfio_pin_pages(struct device *dev, unsigned long *user_pfn,
extern int vfio_unpin_pages(struct device *dev, unsigned long *user_pfn,
int npage);
+#define VFIO_IOMMU_NOTIFY_DMA_UNMAP 1
+
+extern int vfio_register_notifier(struct device *dev,
+ struct notifier_block *nb);
+
+extern int vfio_unregister_notifier(struct device *dev,
+ struct notifier_block *nb);
/*
* IRQfd - generic
*/
--
2.7.0
[toc] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-15 23:30 +0100 |
| Subject | Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP |
| Message-ID | <sDUoN-2tp-15@gated-at.bofh.it> |
| In reply to | #1522815 |
On Tue, 15 Nov 2016 20:59:54 +0530
Kirti Wankhede <kwankhede@nvidia.com> wrote:
> Added blocking notifier to IOMMU TYPE1 driver to notify vendor drivers
> about DMA_UNMAP.
> Exported two APIs vfio_register_notifier() and vfio_unregister_notifier().
> Notifier should be registered, if external user wants to use
> vfio_pin_pages()/vfio_unpin_pages() APIs to pin/unpin pages.
> Vendor driver should use VFIO_IOMMU_NOTIFY_DMA_UNMAP action to invalidate
> mappings.
>
> Signed-off-by: Kirti Wankhede <kwankhede@nvidia.com>
> Signed-off-by: Neo Jia <cjia@nvidia.com>
> Change-Id: I5910d0024d6be87f3e8d3e0ca0eaeaaa0b17f271
> ---
> drivers/vfio/vfio.c | 73 +++++++++++++++++++++++++++++++++++++++++
> drivers/vfio/vfio_iommu_type1.c | 63 +++++++++++++++++++++++++++++------
> include/linux/vfio.h | 11 +++++++
> 3 files changed, 137 insertions(+), 10 deletions(-)
>
> diff --git a/drivers/vfio/vfio.c b/drivers/vfio/vfio.c
> index 3bf8a01bf67b..fa121d983991 100644
> --- a/drivers/vfio/vfio.c
> +++ b/drivers/vfio/vfio.c
> @@ -1902,6 +1902,79 @@ err_unpin_pages:
> }
> EXPORT_SYMBOL(vfio_unpin_pages);
>
> +int vfio_register_notifier(struct device *dev, struct notifier_block *nb)
> +{
> + struct vfio_container *container;
> + struct vfio_group *group;
> + struct vfio_iommu_driver *driver;
> + ssize_t ret;
> +
> + if (!dev || !nb)
> + return -EINVAL;
> +
> + group = vfio_group_get_from_dev(dev);
> + if (IS_ERR(group))
> + return PTR_ERR(group);
> +
> + ret = vfio_group_add_container_user(group);
> + if (ret)
> + goto err_register_nb;
> +
> + container = group->container;
> + down_read(&container->group_lock);
> +
> + driver = container->iommu_driver;
> + if (likely(driver && driver->ops->register_notifier))
> + ret = driver->ops->register_notifier(container->iommu_data, nb);
> + else
> + ret = -ENOTTY;
> +
> + up_read(&container->group_lock);
> + vfio_group_try_dissolve_container(group);
> +
> +err_register_nb:
> + vfio_group_put(group);
> + return ret;
> +}
> +EXPORT_SYMBOL(vfio_register_notifier);
> +
> +int vfio_unregister_notifier(struct device *dev, struct notifier_block *nb)
> +{
> + struct vfio_container *container;
> + struct vfio_group *group;
> + struct vfio_iommu_driver *driver;
> + ssize_t ret;
> +
> + if (!dev || !nb)
> + return -EINVAL;
> +
> + group = vfio_group_get_from_dev(dev);
> + if (IS_ERR(group))
> + return PTR_ERR(group);
> +
> + ret = vfio_group_add_container_user(group);
> + if (ret)
> + goto err_unregister_nb;
> +
> + container = group->container;
> + down_read(&container->group_lock);
> +
> + driver = container->iommu_driver;
> + if (likely(driver && driver->ops->unregister_notifier))
> + ret = driver->ops->unregister_notifier(container->iommu_data,
> + nb);
> + else
> + ret = -ENOTTY;
> +
> + up_read(&container->group_lock);
> + vfio_group_try_dissolve_container(group);
> +
> +err_unregister_nb:
> + vfio_group_put(group);
> + return ret;
> +}
> +EXPORT_SYMBOL(vfio_unregister_notifier);
> +
> /**
> * Module/class support
> */
> diff --git a/drivers/vfio/vfio_iommu_type1.c b/drivers/vfio/vfio_iommu_type1.c
> index 0de7c20f66b1..c45a4822784e 100644
> --- a/drivers/vfio/vfio_iommu_type1.c
> +++ b/drivers/vfio/vfio_iommu_type1.c
> @@ -38,6 +38,7 @@
> #include <linux/workqueue.h>
> #include <linux/pid_namespace.h>
> #include <linux/mdev.h>
> +#include <linux/notifier.h>
>
> #define DRIVER_VERSION "0.2"
> #define DRIVER_AUTHOR "Alex Williamson <alex.williamson@redhat.com>"
> @@ -60,6 +61,7 @@ struct vfio_iommu {
> struct vfio_domain *external_domain; /* domain for external user */
> struct mutex lock;
> struct rb_root dma_list;
> + struct blocking_notifier_head notifier;
> bool v2;
> bool nesting;
> };
> @@ -571,7 +573,8 @@ static int vfio_iommu_type1_pin_pages(void *iommu_data,
>
> mutex_lock(&iommu->lock);
>
> - if (!iommu->external_domain) {
> + /* Fail if notifier list is empty */
> + if ((!iommu->external_domain) || (!iommu->notifier.head)) {
> ret = -EINVAL;
> goto pin_done;
> }
> @@ -854,7 +857,28 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
> */
> if (dma->task->mm != current->mm)
> break;
> +
> unmapped += dma->size;
> +
> + if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
> + struct vfio_iommu_type1_dma_unmap nb_unmap;
> +
> + nb_unmap.iova = dma->iova;
> + nb_unmap.size = dma->size;
> +
> + /*
> + * Notifier callback would call vfio_unpin_pages() which
> + * would acquire iommu->lock. Release lock here and
> + * reacquire it again.
> + */
> + mutex_unlock(&iommu->lock);
> + blocking_notifier_call_chain(&iommu->notifier,
> + VFIO_IOMMU_NOTIFY_DMA_UNMAP,
> + &nb_unmap);
> + mutex_lock(&iommu->lock);
> + if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
> + break;
> + }
Why exactly do we need to notify per vfio_dma rather than per unmap
request? If we do the latter we can send the notify first, limiting us
to races where a page is pinned between the notify and the locking,
whereas here, even our dma pointer is suspect once we re-acquire the
lock, we don't technically know if another unmap could have removed
that already. Perhaps something like this (untested):
diff --git a/drivers/vfio/vfio_iommu_type1.c b/drivers/vfio/vfio_iommu_type1.c
index ee9a680..8504501 100644
--- a/drivers/vfio/vfio_iommu_type1.c
+++ b/drivers/vfio/vfio_iommu_type1.c
@@ -785,6 +785,8 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
struct vfio_dma *dma;
size_t unmapped = 0;
int ret = 0;
+ struct vfio_iommu_type1_dma_unmap nb_unmap = { .iova = unmap->iova,
+ .size = unmap->size };
mask = ((uint64_t)1 << __ffs(vfio_pgsize_bitmap(iommu))) - 1;
@@ -795,6 +797,14 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
WARN_ON(mask & PAGE_MASK);
+ /*
+ * Notify anyone (mdev vendor drivers) to invalidate and unmap
+ * iovas within the range we're about to unmap. Vendor drivers MUST
+ * unpin pages in response to an invalidation.
+ */
+ blocking_notifier_call_chain(&iommu->notifier,
+ VFIO_IOMMU_NOTIFY_DMA_UNMAP, &nb_unmap);
+
mutex_lock(&iommu->lock);
/*
@@ -853,25 +863,8 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
unmapped += dma->size;
- if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
- struct vfio_iommu_type1_dma_unmap nb_unmap;
+ WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list));
- nb_unmap.iova = dma->iova;
- nb_unmap.size = dma->size;
-
- /*
- * Notifier callback would call vfio_unpin_pages() which
- * would acquire iommu->lock. Release lock here and
- * reacquire it again.
- */
- mutex_unlock(&iommu->lock);
- blocking_notifier_call_chain(&iommu->notifier,
- VFIO_IOMMU_NOTIFY_DMA_UNMAP,
- &nb_unmap);
- mutex_lock(&iommu->lock);
- if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
- break;
- }
vfio_remove_dma(iommu, dma);
}
> vfio_remove_dma(iommu, dma);
> }
>
> @@ -1439,6 +1463,7 @@ static void *vfio_iommu_type1_open(unsigned long arg)
> INIT_LIST_HEAD(&iommu->domain_list);
> iommu->dma_list = RB_ROOT;
> mutex_init(&iommu->lock);
> + BLOCKING_INIT_NOTIFIER_HEAD(&iommu->notifier);
>
> return iommu;
> }
> @@ -1574,16 +1599,34 @@ static long vfio_iommu_type1_ioctl(void *iommu_data,
> return -ENOTTY;
> }
>
> +static int vfio_iommu_type1_register_notifier(void *iommu_data,
> + struct notifier_block *nb)
> +{
> + struct vfio_iommu *iommu = iommu_data;
> +
> + return blocking_notifier_chain_register(&iommu->notifier, nb);
> +}
> +
> +static int vfio_iommu_type1_unregister_notifier(void *iommu_data,
> + struct notifier_block *nb)
> +{
> + struct vfio_iommu *iommu = iommu_data;
> +
> + return blocking_notifier_chain_unregister(&iommu->notifier, nb);
> +}
> +
> static const struct vfio_iommu_driver_ops vfio_iommu_driver_ops_type1 = {
> - .name = "vfio-iommu-type1",
> - .owner = THIS_MODULE,
> - .open = vfio_iommu_type1_open,
> - .release = vfio_iommu_type1_release,
> - .ioctl = vfio_iommu_type1_ioctl,
> - .attach_group = vfio_iommu_type1_attach_group,
> - .detach_group = vfio_iommu_type1_detach_group,
> - .pin_pages = vfio_iommu_type1_pin_pages,
> - .unpin_pages = vfio_iommu_type1_unpin_pages,
> + .name = "vfio-iommu-type1",
> + .owner = THIS_MODULE,
> + .open = vfio_iommu_type1_open,
> + .release = vfio_iommu_type1_release,
> + .ioctl = vfio_iommu_type1_ioctl,
> + .attach_group = vfio_iommu_type1_attach_group,
> + .detach_group = vfio_iommu_type1_detach_group,
> + .pin_pages = vfio_iommu_type1_pin_pages,
> + .unpin_pages = vfio_iommu_type1_unpin_pages,
> + .register_notifier = vfio_iommu_type1_register_notifier,
> + .unregister_notifier = vfio_iommu_type1_unregister_notifier,
> };
>
> static int __init vfio_iommu_type1_init(void)
> diff --git a/include/linux/vfio.h b/include/linux/vfio.h
> index 420cdc928786..997442398c09 100644
> --- a/include/linux/vfio.h
> +++ b/include/linux/vfio.h
> @@ -80,6 +80,10 @@ struct vfio_iommu_driver_ops {
> unsigned long *phys_pfn);
> int (*unpin_pages)(void *iommu_data,
> unsigned long *user_pfn, int npage);
> + int (*register_notifier)(void *iommu_data,
> + struct notifier_block *nb);
> + int (*unregister_notifier)(void *iommu_data,
> + struct notifier_block *nb);
> };
>
> extern int vfio_register_iommu_driver(const struct vfio_iommu_driver_ops *ops);
> @@ -139,6 +143,13 @@ extern int vfio_pin_pages(struct device *dev, unsigned long *user_pfn,
> extern int vfio_unpin_pages(struct device *dev, unsigned long *user_pfn,
> int npage);
>
> +#define VFIO_IOMMU_NOTIFY_DMA_UNMAP 1
> +
> +extern int vfio_register_notifier(struct device *dev,
> + struct notifier_block *nb);
> +
> +extern int vfio_unregister_notifier(struct device *dev,
> + struct notifier_block *nb);
> /*
> * IRQfd - generic
> */
[toc] | [prev] | [next] | [standalone]
| From | Kirti Wankhede <kwankhede@nvidia.com> |
|---|---|
| Date | 2016-11-16 03:50 +0100 |
| Subject | Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP |
| Message-ID | <sDYsq-4Ud-1@gated-at.bofh.it> |
| In reply to | #1523100 |
On 11/16/2016 3:49 AM, Alex Williamson wrote:
> On Tue, 15 Nov 2016 20:59:54 +0530
> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>
...
>> @@ -854,7 +857,28 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
>> */
>> if (dma->task->mm != current->mm)
>> break;
>> +
>> unmapped += dma->size;
>> +
>> + if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
>> + struct vfio_iommu_type1_dma_unmap nb_unmap;
>> +
>> + nb_unmap.iova = dma->iova;
>> + nb_unmap.size = dma->size;
>> +
>> + /*
>> + * Notifier callback would call vfio_unpin_pages() which
>> + * would acquire iommu->lock. Release lock here and
>> + * reacquire it again.
>> + */
>> + mutex_unlock(&iommu->lock);
>> + blocking_notifier_call_chain(&iommu->notifier,
>> + VFIO_IOMMU_NOTIFY_DMA_UNMAP,
>> + &nb_unmap);
>> + mutex_lock(&iommu->lock);
>> + if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
>> + break;
>> + }
>
>
> Why exactly do we need to notify per vfio_dma rather than per unmap
> request? If we do the latter we can send the notify first, limiting us
> to races where a page is pinned between the notify and the locking,
> whereas here, even our dma pointer is suspect once we re-acquire the
> lock, we don't technically know if another unmap could have removed
> that already. Perhaps something like this (untested):
>
There are checks to validate unmap request, like v2 check and who is
calling unmap and is it allowed for that task to unmap. Before these
checks its not sure that unmap region range which asked for would be
unmapped all. Notify call should be at the place where its sure that the
range provided to notify call is definitely going to be removed. My
change do that.
Thanks,
Kirti
> diff --git a/drivers/vfio/vfio_iommu_type1.c b/drivers/vfio/vfio_iommu_type1.c
> index ee9a680..8504501 100644
> --- a/drivers/vfio/vfio_iommu_type1.c
> +++ b/drivers/vfio/vfio_iommu_type1.c
> @@ -785,6 +785,8 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
> struct vfio_dma *dma;
> size_t unmapped = 0;
> int ret = 0;
> + struct vfio_iommu_type1_dma_unmap nb_unmap = { .iova = unmap->iova,
> + .size = unmap->size };
>
> mask = ((uint64_t)1 << __ffs(vfio_pgsize_bitmap(iommu))) - 1;
>
> @@ -795,6 +797,14 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
>
> WARN_ON(mask & PAGE_MASK);
>
> + /*
> + * Notify anyone (mdev vendor drivers) to invalidate and unmap
> + * iovas within the range we're about to unmap. Vendor drivers MUST
> + * unpin pages in response to an invalidation.
> + */
> + blocking_notifier_call_chain(&iommu->notifier,
> + VFIO_IOMMU_NOTIFY_DMA_UNMAP, &nb_unmap);
> +
> mutex_lock(&iommu->lock);
>
> /*
> @@ -853,25 +863,8 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
>
> unmapped += dma->size;
>
> - if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
> - struct vfio_iommu_type1_dma_unmap nb_unmap;
> + WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list));
>
> - nb_unmap.iova = dma->iova;
> - nb_unmap.size = dma->size;
> -
> - /*
> - * Notifier callback would call vfio_unpin_pages() which
> - * would acquire iommu->lock. Release lock here and
> - * reacquire it again.
> - */
> - mutex_unlock(&iommu->lock);
> - blocking_notifier_call_chain(&iommu->notifier,
> - VFIO_IOMMU_NOTIFY_DMA_UNMAP,
> - &nb_unmap);
> - mutex_lock(&iommu->lock);
> - if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
> - break;
> - }
> vfio_remove_dma(iommu, dma);
> }
>
>
>
>> vfio_remove_dma(iommu, dma);
>> }
>>
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-16 04:20 +0100 |
| Subject | Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP |
| Message-ID | <sDYVs-5pO-29@gated-at.bofh.it> |
| In reply to | #1523188 |
On Wed, 16 Nov 2016 08:16:15 +0530
Kirti Wankhede <kwankhede@nvidia.com> wrote:
> On 11/16/2016 3:49 AM, Alex Williamson wrote:
> > On Tue, 15 Nov 2016 20:59:54 +0530
> > Kirti Wankhede <kwankhede@nvidia.com> wrote:
> >
> ...
>
> >> @@ -854,7 +857,28 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
> >> */
> >> if (dma->task->mm != current->mm)
> >> break;
> >> +
> >> unmapped += dma->size;
> >> +
> >> + if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
> >> + struct vfio_iommu_type1_dma_unmap nb_unmap;
> >> +
> >> + nb_unmap.iova = dma->iova;
> >> + nb_unmap.size = dma->size;
> >> +
> >> + /*
> >> + * Notifier callback would call vfio_unpin_pages() which
> >> + * would acquire iommu->lock. Release lock here and
> >> + * reacquire it again.
> >> + */
> >> + mutex_unlock(&iommu->lock);
> >> + blocking_notifier_call_chain(&iommu->notifier,
> >> + VFIO_IOMMU_NOTIFY_DMA_UNMAP,
> >> + &nb_unmap);
> >> + mutex_lock(&iommu->lock);
> >> + if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
> >> + break;
> >> + }
> >
> >
> > Why exactly do we need to notify per vfio_dma rather than per unmap
> > request? If we do the latter we can send the notify first, limiting us
> > to races where a page is pinned between the notify and the locking,
> > whereas here, even our dma pointer is suspect once we re-acquire the
> > lock, we don't technically know if another unmap could have removed
> > that already. Perhaps something like this (untested):
> >
>
> There are checks to validate unmap request, like v2 check and who is
> calling unmap and is it allowed for that task to unmap. Before these
> checks its not sure that unmap region range which asked for would be
> unmapped all. Notify call should be at the place where its sure that the
> range provided to notify call is definitely going to be removed. My
> change do that.
Ok, but that does solve the problem. What about this (untested):
diff --git a/drivers/vfio/vfio_iommu_type1.c b/drivers/vfio/vfio_iommu_type1.c
index ee9a680..50cafdf 100644
--- a/drivers/vfio/vfio_iommu_type1.c
+++ b/drivers/vfio/vfio_iommu_type1.c
@@ -782,9 +782,9 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
struct vfio_iommu_type1_dma_unmap *unmap)
{
uint64_t mask;
- struct vfio_dma *dma;
+ struct vfio_dma *dma, *dma_last = NULL;
size_t unmapped = 0;
- int ret = 0;
+ int ret = 0, retries;
mask = ((uint64_t)1 << __ffs(vfio_pgsize_bitmap(iommu))) - 1;
@@ -794,7 +794,7 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
return -EINVAL;
WARN_ON(mask & PAGE_MASK);
-
+again:
mutex_lock(&iommu->lock);
/*
@@ -851,11 +851,16 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
if (dma->task->mm != current->mm)
break;
- unmapped += dma->size;
-
- if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
+ if (!RB_EMPTY_ROOT(&dma->pfn_list)) {
struct vfio_iommu_type1_dma_unmap nb_unmap;
+ if (dma_last == dma) {
+ BUG_ON(++retries > 10);
+ } else {
+ dma_last = dma;
+ retries = 0;
+ }
+
nb_unmap.iova = dma->iova;
nb_unmap.size = dma->size;
@@ -868,11 +873,11 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
blocking_notifier_call_chain(&iommu->notifier,
VFIO_IOMMU_NOTIFY_DMA_UNMAP,
&nb_unmap);
- mutex_lock(&iommu->lock);
- if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
- break;
+ goto again:
}
+ unmapped += dma->size;
vfio_remove_dma(iommu, dma);
+
}
unlock:
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-16 04:30 +0100 |
| Subject | Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP |
| Message-ID | <sDZ58-5t5-9@gated-at.bofh.it> |
| In reply to | #1523207 |
On Tue, 15 Nov 2016 20:16:12 -0700
Alex Williamson <alex.williamson@redhat.com> wrote:
> On Wed, 16 Nov 2016 08:16:15 +0530
> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>
> > On 11/16/2016 3:49 AM, Alex Williamson wrote:
> > > On Tue, 15 Nov 2016 20:59:54 +0530
> > > Kirti Wankhede <kwankhede@nvidia.com> wrote:
> > >
> > ...
> >
> > >> @@ -854,7 +857,28 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
> > >> */
> > >> if (dma->task->mm != current->mm)
> > >> break;
> > >> +
> > >> unmapped += dma->size;
> > >> +
> > >> + if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
> > >> + struct vfio_iommu_type1_dma_unmap nb_unmap;
> > >> +
> > >> + nb_unmap.iova = dma->iova;
> > >> + nb_unmap.size = dma->size;
> > >> +
> > >> + /*
> > >> + * Notifier callback would call vfio_unpin_pages() which
> > >> + * would acquire iommu->lock. Release lock here and
> > >> + * reacquire it again.
> > >> + */
> > >> + mutex_unlock(&iommu->lock);
> > >> + blocking_notifier_call_chain(&iommu->notifier,
> > >> + VFIO_IOMMU_NOTIFY_DMA_UNMAP,
> > >> + &nb_unmap);
> > >> + mutex_lock(&iommu->lock);
> > >> + if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
> > >> + break;
> > >> + }
> > >
> > >
> > > Why exactly do we need to notify per vfio_dma rather than per unmap
> > > request? If we do the latter we can send the notify first, limiting us
> > > to races where a page is pinned between the notify and the locking,
> > > whereas here, even our dma pointer is suspect once we re-acquire the
> > > lock, we don't technically know if another unmap could have removed
> > > that already. Perhaps something like this (untested):
> > >
> >
> > There are checks to validate unmap request, like v2 check and who is
> > calling unmap and is it allowed for that task to unmap. Before these
> > checks its not sure that unmap region range which asked for would be
> > unmapped all. Notify call should be at the place where its sure that the
> > range provided to notify call is definitely going to be removed. My
> > change do that.
>
> Ok, but that does solve the problem. What about this (untested):
s/does/does not/
BTW, I like how the retries here fill the gap in my previous proposal
where we could still race re-pinning. We've given it an honest shot or
someone is not participating if we've retried 10 times. I don't
understand why the test for iommu->external_domain was there, clearly
if the list is not empty, we need to notify. Thanks,
Alex
> diff --git a/drivers/vfio/vfio_iommu_type1.c b/drivers/vfio/vfio_iommu_type1.c
> index ee9a680..50cafdf 100644
> --- a/drivers/vfio/vfio_iommu_type1.c
> +++ b/drivers/vfio/vfio_iommu_type1.c
> @@ -782,9 +782,9 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
> struct vfio_iommu_type1_dma_unmap *unmap)
> {
> uint64_t mask;
> - struct vfio_dma *dma;
> + struct vfio_dma *dma, *dma_last = NULL;
> size_t unmapped = 0;
> - int ret = 0;
> + int ret = 0, retries;
>
> mask = ((uint64_t)1 << __ffs(vfio_pgsize_bitmap(iommu))) - 1;
>
> @@ -794,7 +794,7 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
> return -EINVAL;
>
> WARN_ON(mask & PAGE_MASK);
> -
> +again:
> mutex_lock(&iommu->lock);
>
> /*
> @@ -851,11 +851,16 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
> if (dma->task->mm != current->mm)
> break;
>
> - unmapped += dma->size;
> -
> - if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
> + if (!RB_EMPTY_ROOT(&dma->pfn_list)) {
> struct vfio_iommu_type1_dma_unmap nb_unmap;
>
> + if (dma_last == dma) {
> + BUG_ON(++retries > 10);
> + } else {
> + dma_last = dma;
> + retries = 0;
> + }
> +
> nb_unmap.iova = dma->iova;
> nb_unmap.size = dma->size;
>
> @@ -868,11 +873,11 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
> blocking_notifier_call_chain(&iommu->notifier,
> VFIO_IOMMU_NOTIFY_DMA_UNMAP,
> &nb_unmap);
> - mutex_lock(&iommu->lock);
> - if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
> - break;
> + goto again:
> }
> + unmapped += dma->size;
> vfio_remove_dma(iommu, dma);
> +
> }
>
> unlock:
[toc] | [prev] | [next] | [standalone]
| From | Kirti Wankhede <kwankhede@nvidia.com> |
|---|---|
| Date | 2016-11-16 04:50 +0100 |
| Subject | Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP |
| Message-ID | <sDZot-5zy-9@gated-at.bofh.it> |
| In reply to | #1523213 |
On 11/16/2016 8:55 AM, Alex Williamson wrote:
> On Tue, 15 Nov 2016 20:16:12 -0700
> Alex Williamson <alex.williamson@redhat.com> wrote:
>
>> On Wed, 16 Nov 2016 08:16:15 +0530
>> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>>
>>> On 11/16/2016 3:49 AM, Alex Williamson wrote:
>>>> On Tue, 15 Nov 2016 20:59:54 +0530
>>>> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>>>>
>>> ...
>>>
>>>>> @@ -854,7 +857,28 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
>>>>> */
>>>>> if (dma->task->mm != current->mm)
>>>>> break;
>>>>> +
>>>>> unmapped += dma->size;
>>>>> +
>>>>> + if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
>>>>> + struct vfio_iommu_type1_dma_unmap nb_unmap;
>>>>> +
>>>>> + nb_unmap.iova = dma->iova;
>>>>> + nb_unmap.size = dma->size;
>>>>> +
>>>>> + /*
>>>>> + * Notifier callback would call vfio_unpin_pages() which
>>>>> + * would acquire iommu->lock. Release lock here and
>>>>> + * reacquire it again.
>>>>> + */
>>>>> + mutex_unlock(&iommu->lock);
>>>>> + blocking_notifier_call_chain(&iommu->notifier,
>>>>> + VFIO_IOMMU_NOTIFY_DMA_UNMAP,
>>>>> + &nb_unmap);
>>>>> + mutex_lock(&iommu->lock);
>>>>> + if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
>>>>> + break;
>>>>> + }
>>>>
>>>>
>>>> Why exactly do we need to notify per vfio_dma rather than per unmap
>>>> request? If we do the latter we can send the notify first, limiting us
>>>> to races where a page is pinned between the notify and the locking,
>>>> whereas here, even our dma pointer is suspect once we re-acquire the
>>>> lock, we don't technically know if another unmap could have removed
>>>> that already. Perhaps something like this (untested):
>>>>
>>>
>>> There are checks to validate unmap request, like v2 check and who is
>>> calling unmap and is it allowed for that task to unmap. Before these
>>> checks its not sure that unmap region range which asked for would be
>>> unmapped all. Notify call should be at the place where its sure that the
>>> range provided to notify call is definitely going to be removed. My
>>> change do that.
>>
>> Ok, but that does solve the problem. What about this (untested):
>
> s/does/does not/
>
> BTW, I like how the retries here fill the gap in my previous proposal
> where we could still race re-pinning. We've given it an honest shot or
> someone is not participating if we've retried 10 times. I don't
> understand why the test for iommu->external_domain was there, clearly
> if the list is not empty, we need to notify. Thanks,
>
Ok. Retry is good to give a chance to unpin all. But is it really
required to use BUG_ON() that would panic the host. I think WARN_ON
should be fine and then when container is closed or when the last group
is removed from the container, vfio_iommu_type1_release() is called and
we have a chance to unpin it all.
Thanks,
Kirti
> Alex
>
>> diff --git a/drivers/vfio/vfio_iommu_type1.c b/drivers/vfio/vfio_iommu_type1.c
>> index ee9a680..50cafdf 100644
>> --- a/drivers/vfio/vfio_iommu_type1.c
>> +++ b/drivers/vfio/vfio_iommu_type1.c
>> @@ -782,9 +782,9 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
>> struct vfio_iommu_type1_dma_unmap *unmap)
>> {
>> uint64_t mask;
>> - struct vfio_dma *dma;
>> + struct vfio_dma *dma, *dma_last = NULL;
>> size_t unmapped = 0;
>> - int ret = 0;
>> + int ret = 0, retries;
>>
>> mask = ((uint64_t)1 << __ffs(vfio_pgsize_bitmap(iommu))) - 1;
>>
>> @@ -794,7 +794,7 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
>> return -EINVAL;
>>
>> WARN_ON(mask & PAGE_MASK);
>> -
>> +again:
>> mutex_lock(&iommu->lock);
>>
>> /*
>> @@ -851,11 +851,16 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
>> if (dma->task->mm != current->mm)
>> break;
>>
>> - unmapped += dma->size;
>> -
>> - if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
>> + if (!RB_EMPTY_ROOT(&dma->pfn_list)) {
>> struct vfio_iommu_type1_dma_unmap nb_unmap;
>>
>> + if (dma_last == dma) {
>> + BUG_ON(++retries > 10);
>> + } else {
>> + dma_last = dma;
>> + retries = 0;
>> + }
>> +
>> nb_unmap.iova = dma->iova;
>> nb_unmap.size = dma->size;
>>
>> @@ -868,11 +873,11 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
>> blocking_notifier_call_chain(&iommu->notifier,
>> VFIO_IOMMU_NOTIFY_DMA_UNMAP,
>> &nb_unmap);
>> - mutex_lock(&iommu->lock);
>> - if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
>> - break;
>> + goto again:
>> }
>> + unmapped += dma->size;
>> vfio_remove_dma(iommu, dma);
>> +
>> }
>>
>> unlock:
>
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-16 05:00 +0100 |
| Subject | Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP |
| Message-ID | <sDZy9-5CR-5@gated-at.bofh.it> |
| In reply to | #1523217 |
On Wed, 16 Nov 2016 09:13:37 +0530
Kirti Wankhede <kwankhede@nvidia.com> wrote:
> On 11/16/2016 8:55 AM, Alex Williamson wrote:
> > On Tue, 15 Nov 2016 20:16:12 -0700
> > Alex Williamson <alex.williamson@redhat.com> wrote:
> >
> >> On Wed, 16 Nov 2016 08:16:15 +0530
> >> Kirti Wankhede <kwankhede@nvidia.com> wrote:
> >>
> >>> On 11/16/2016 3:49 AM, Alex Williamson wrote:
> >>>> On Tue, 15 Nov 2016 20:59:54 +0530
> >>>> Kirti Wankhede <kwankhede@nvidia.com> wrote:
> >>>>
> >>> ...
> >>>
> >>>>> @@ -854,7 +857,28 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
> >>>>> */
> >>>>> if (dma->task->mm != current->mm)
> >>>>> break;
> >>>>> +
> >>>>> unmapped += dma->size;
> >>>>> +
> >>>>> + if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
> >>>>> + struct vfio_iommu_type1_dma_unmap nb_unmap;
> >>>>> +
> >>>>> + nb_unmap.iova = dma->iova;
> >>>>> + nb_unmap.size = dma->size;
> >>>>> +
> >>>>> + /*
> >>>>> + * Notifier callback would call vfio_unpin_pages() which
> >>>>> + * would acquire iommu->lock. Release lock here and
> >>>>> + * reacquire it again.
> >>>>> + */
> >>>>> + mutex_unlock(&iommu->lock);
> >>>>> + blocking_notifier_call_chain(&iommu->notifier,
> >>>>> + VFIO_IOMMU_NOTIFY_DMA_UNMAP,
> >>>>> + &nb_unmap);
> >>>>> + mutex_lock(&iommu->lock);
> >>>>> + if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
> >>>>> + break;
> >>>>> + }
> >>>>
> >>>>
> >>>> Why exactly do we need to notify per vfio_dma rather than per unmap
> >>>> request? If we do the latter we can send the notify first, limiting us
> >>>> to races where a page is pinned between the notify and the locking,
> >>>> whereas here, even our dma pointer is suspect once we re-acquire the
> >>>> lock, we don't technically know if another unmap could have removed
> >>>> that already. Perhaps something like this (untested):
> >>>>
> >>>
> >>> There are checks to validate unmap request, like v2 check and who is
> >>> calling unmap and is it allowed for that task to unmap. Before these
> >>> checks its not sure that unmap region range which asked for would be
> >>> unmapped all. Notify call should be at the place where its sure that the
> >>> range provided to notify call is definitely going to be removed. My
> >>> change do that.
> >>
> >> Ok, but that does solve the problem. What about this (untested):
> >
> > s/does/does not/
> >
> > BTW, I like how the retries here fill the gap in my previous proposal
> > where we could still race re-pinning. We've given it an honest shot or
> > someone is not participating if we've retried 10 times. I don't
> > understand why the test for iommu->external_domain was there, clearly
> > if the list is not empty, we need to notify. Thanks,
> >
>
> Ok. Retry is good to give a chance to unpin all. But is it really
> required to use BUG_ON() that would panic the host. I think WARN_ON
> should be fine and then when container is closed or when the last group
> is removed from the container, vfio_iommu_type1_release() is called and
> we have a chance to unpin it all.
See my comments on patch 10/22, we need to be vigilant that the vendor
driver is participating. I don't think we should be cleaning up after
the vendor driver on release, if we need to do that, it implies we
already have problems in multi-mdev containers since we'll be left with
pfn_list entries that no longer have an owner. Thanks,
Alex
[toc] | [prev] | [next] | [standalone]
| From | Kirti Wankhede <kwankhede@nvidia.com> |
|---|---|
| Date | 2016-11-16 05:20 +0100 |
| Subject | Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP |
| Message-ID | <sDZRv-62y-9@gated-at.bofh.it> |
| In reply to | #1523219 |
On 11/16/2016 9:28 AM, Alex Williamson wrote:
> On Wed, 16 Nov 2016 09:13:37 +0530
> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>
>> On 11/16/2016 8:55 AM, Alex Williamson wrote:
>>> On Tue, 15 Nov 2016 20:16:12 -0700
>>> Alex Williamson <alex.williamson@redhat.com> wrote:
>>>
>>>> On Wed, 16 Nov 2016 08:16:15 +0530
>>>> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>>>>
>>>>> On 11/16/2016 3:49 AM, Alex Williamson wrote:
>>>>>> On Tue, 15 Nov 2016 20:59:54 +0530
>>>>>> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>>>>>>
>>>>> ...
>>>>>
>>>>>>> @@ -854,7 +857,28 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
>>>>>>> */
>>>>>>> if (dma->task->mm != current->mm)
>>>>>>> break;
>>>>>>> +
>>>>>>> unmapped += dma->size;
>>>>>>> +
>>>>>>> + if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
>>>>>>> + struct vfio_iommu_type1_dma_unmap nb_unmap;
>>>>>>> +
>>>>>>> + nb_unmap.iova = dma->iova;
>>>>>>> + nb_unmap.size = dma->size;
>>>>>>> +
>>>>>>> + /*
>>>>>>> + * Notifier callback would call vfio_unpin_pages() which
>>>>>>> + * would acquire iommu->lock. Release lock here and
>>>>>>> + * reacquire it again.
>>>>>>> + */
>>>>>>> + mutex_unlock(&iommu->lock);
>>>>>>> + blocking_notifier_call_chain(&iommu->notifier,
>>>>>>> + VFIO_IOMMU_NOTIFY_DMA_UNMAP,
>>>>>>> + &nb_unmap);
>>>>>>> + mutex_lock(&iommu->lock);
>>>>>>> + if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
>>>>>>> + break;
>>>>>>> + }
>>>>>>
>>>>>>
>>>>>> Why exactly do we need to notify per vfio_dma rather than per unmap
>>>>>> request? If we do the latter we can send the notify first, limiting us
>>>>>> to races where a page is pinned between the notify and the locking,
>>>>>> whereas here, even our dma pointer is suspect once we re-acquire the
>>>>>> lock, we don't technically know if another unmap could have removed
>>>>>> that already. Perhaps something like this (untested):
>>>>>>
>>>>>
>>>>> There are checks to validate unmap request, like v2 check and who is
>>>>> calling unmap and is it allowed for that task to unmap. Before these
>>>>> checks its not sure that unmap region range which asked for would be
>>>>> unmapped all. Notify call should be at the place where its sure that the
>>>>> range provided to notify call is definitely going to be removed. My
>>>>> change do that.
>>>>
>>>> Ok, but that does solve the problem. What about this (untested):
>>>
>>> s/does/does not/
>>>
>>> BTW, I like how the retries here fill the gap in my previous proposal
>>> where we could still race re-pinning. We've given it an honest shot or
>>> someone is not participating if we've retried 10 times. I don't
>>> understand why the test for iommu->external_domain was there, clearly
>>> if the list is not empty, we need to notify. Thanks,
>>>
>>
>> Ok. Retry is good to give a chance to unpin all. But is it really
>> required to use BUG_ON() that would panic the host. I think WARN_ON
>> should be fine and then when container is closed or when the last group
>> is removed from the container, vfio_iommu_type1_release() is called and
>> we have a chance to unpin it all.
>
> See my comments on patch 10/22, we need to be vigilant that the vendor
> driver is participating. I don't think we should be cleaning up after
> the vendor driver on release, if we need to do that, it implies we
> already have problems in multi-mdev containers since we'll be left with
> pfn_list entries that no longer have an owner. Thanks,
>
If any vendor driver doesn't clean its pinned pages and there are
entries in pfn_list with no owner, that would be indicated by WARN_ON,
which should be fixed by that vendor driver. I still feel it shouldn't
cause host panic.
When such warning is seen with multiple mdev devices in container, it is
easy to isolate and find which vendor driver is not cleaning their
stuff, same warning would be seen with single mdev device in a
container. To isolate and find which vendor driver is culprit check with
one mdev device at a time.
Finally, we have a chance to clean all residue from
vfio_iommu_type1_release() so that vfio_iommu_type1 module doesn't leave
any leaks.
Thanks,
Kirti
[toc] | [prev] | [next] | [standalone]
| From | Alex Williamson <alex.williamson@redhat.com> |
|---|---|
| Date | 2016-11-16 05:40 +0100 |
| Subject | Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP |
| Message-ID | <sE0aR-6d1-9@gated-at.bofh.it> |
| In reply to | #1523222 |
On Wed, 16 Nov 2016 09:46:20 +0530
Kirti Wankhede <kwankhede@nvidia.com> wrote:
> On 11/16/2016 9:28 AM, Alex Williamson wrote:
> > On Wed, 16 Nov 2016 09:13:37 +0530
> > Kirti Wankhede <kwankhede@nvidia.com> wrote:
> >
> >> On 11/16/2016 8:55 AM, Alex Williamson wrote:
> >>> On Tue, 15 Nov 2016 20:16:12 -0700
> >>> Alex Williamson <alex.williamson@redhat.com> wrote:
> >>>
> >>>> On Wed, 16 Nov 2016 08:16:15 +0530
> >>>> Kirti Wankhede <kwankhede@nvidia.com> wrote:
> >>>>
> >>>>> On 11/16/2016 3:49 AM, Alex Williamson wrote:
> >>>>>> On Tue, 15 Nov 2016 20:59:54 +0530
> >>>>>> Kirti Wankhede <kwankhede@nvidia.com> wrote:
> >>>>>>
> >>>>> ...
> >>>>>
> >>>>>>> @@ -854,7 +857,28 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
> >>>>>>> */
> >>>>>>> if (dma->task->mm != current->mm)
> >>>>>>> break;
> >>>>>>> +
> >>>>>>> unmapped += dma->size;
> >>>>>>> +
> >>>>>>> + if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
> >>>>>>> + struct vfio_iommu_type1_dma_unmap nb_unmap;
> >>>>>>> +
> >>>>>>> + nb_unmap.iova = dma->iova;
> >>>>>>> + nb_unmap.size = dma->size;
> >>>>>>> +
> >>>>>>> + /*
> >>>>>>> + * Notifier callback would call vfio_unpin_pages() which
> >>>>>>> + * would acquire iommu->lock. Release lock here and
> >>>>>>> + * reacquire it again.
> >>>>>>> + */
> >>>>>>> + mutex_unlock(&iommu->lock);
> >>>>>>> + blocking_notifier_call_chain(&iommu->notifier,
> >>>>>>> + VFIO_IOMMU_NOTIFY_DMA_UNMAP,
> >>>>>>> + &nb_unmap);
> >>>>>>> + mutex_lock(&iommu->lock);
> >>>>>>> + if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
> >>>>>>> + break;
> >>>>>>> + }
> >>>>>>
> >>>>>>
> >>>>>> Why exactly do we need to notify per vfio_dma rather than per unmap
> >>>>>> request? If we do the latter we can send the notify first, limiting us
> >>>>>> to races where a page is pinned between the notify and the locking,
> >>>>>> whereas here, even our dma pointer is suspect once we re-acquire the
> >>>>>> lock, we don't technically know if another unmap could have removed
> >>>>>> that already. Perhaps something like this (untested):
> >>>>>>
> >>>>>
> >>>>> There are checks to validate unmap request, like v2 check and who is
> >>>>> calling unmap and is it allowed for that task to unmap. Before these
> >>>>> checks its not sure that unmap region range which asked for would be
> >>>>> unmapped all. Notify call should be at the place where its sure that the
> >>>>> range provided to notify call is definitely going to be removed. My
> >>>>> change do that.
> >>>>
> >>>> Ok, but that does solve the problem. What about this (untested):
> >>>
> >>> s/does/does not/
> >>>
> >>> BTW, I like how the retries here fill the gap in my previous proposal
> >>> where we could still race re-pinning. We've given it an honest shot or
> >>> someone is not participating if we've retried 10 times. I don't
> >>> understand why the test for iommu->external_domain was there, clearly
> >>> if the list is not empty, we need to notify. Thanks,
> >>>
> >>
> >> Ok. Retry is good to give a chance to unpin all. But is it really
> >> required to use BUG_ON() that would panic the host. I think WARN_ON
> >> should be fine and then when container is closed or when the last group
> >> is removed from the container, vfio_iommu_type1_release() is called and
> >> we have a chance to unpin it all.
> >
> > See my comments on patch 10/22, we need to be vigilant that the vendor
> > driver is participating. I don't think we should be cleaning up after
> > the vendor driver on release, if we need to do that, it implies we
> > already have problems in multi-mdev containers since we'll be left with
> > pfn_list entries that no longer have an owner. Thanks,
> >
>
> If any vendor driver doesn't clean its pinned pages and there are
> entries in pfn_list with no owner, that would be indicated by WARN_ON,
> which should be fixed by that vendor driver. I still feel it shouldn't
> cause host panic.
> When such warning is seen with multiple mdev devices in container, it is
> easy to isolate and find which vendor driver is not cleaning their
> stuff, same warning would be seen with single mdev device in a
> container. To isolate and find which vendor driver is culprit check with
> one mdev device at a time.
> Finally, we have a chance to clean all residue from
> vfio_iommu_type1_release() so that vfio_iommu_type1 module doesn't leave
> any leaks.
How can we claim that we've resolved anything by unpinning the
residue? In fact, is it actually safe to unpin any residue left by the
vendor driver or does it imply that we're promoting a simple memory
leak to a security issue because we can't verify whether the vendor
driver has disabled access to that pfn, which may not reference a user
page after we unpin it. That, in addition to the fact that I don't
need to figure out how to break from the loop with a BUG_ON, is why I
chose that rather than a WARN_ON. The release path could probably be a
WARN_ON since the user no longer has access to the device, so we have a
consistency error with the vendor driver, but we're probably not
promoting it further by unpinning the pages. Thanks,
Alex
[toc] | [prev] | [next] | [standalone]
| From | Kirti Wankhede <kwankhede@nvidia.com> |
|---|---|
| Date | 2016-11-16 16:30 +0100 |
| Subject | Re: [PATCH v13 11/22] vfio iommu: Add blocking notifier to notify DMA_UNMAP |
| Message-ID | <sEajV-4Eo-61@gated-at.bofh.it> |
| In reply to | #1523227 |
On 11/16/2016 10:06 AM, Alex Williamson wrote:
> On Wed, 16 Nov 2016 09:46:20 +0530
> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>
>> On 11/16/2016 9:28 AM, Alex Williamson wrote:
>>> On Wed, 16 Nov 2016 09:13:37 +0530
>>> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>>>
>>>> On 11/16/2016 8:55 AM, Alex Williamson wrote:
>>>>> On Tue, 15 Nov 2016 20:16:12 -0700
>>>>> Alex Williamson <alex.williamson@redhat.com> wrote:
>>>>>
>>>>>> On Wed, 16 Nov 2016 08:16:15 +0530
>>>>>> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>>>>>>
>>>>>>> On 11/16/2016 3:49 AM, Alex Williamson wrote:
>>>>>>>> On Tue, 15 Nov 2016 20:59:54 +0530
>>>>>>>> Kirti Wankhede <kwankhede@nvidia.com> wrote:
>>>>>>>>
>>>>>>> ...
>>>>>>>
>>>>>>>>> @@ -854,7 +857,28 @@ static int vfio_dma_do_unmap(struct vfio_iommu *iommu,
>>>>>>>>> */
>>>>>>>>> if (dma->task->mm != current->mm)
>>>>>>>>> break;
>>>>>>>>> +
>>>>>>>>> unmapped += dma->size;
>>>>>>>>> +
>>>>>>>>> + if (iommu->external_domain && !RB_EMPTY_ROOT(&dma->pfn_list)) {
>>>>>>>>> + struct vfio_iommu_type1_dma_unmap nb_unmap;
>>>>>>>>> +
>>>>>>>>> + nb_unmap.iova = dma->iova;
>>>>>>>>> + nb_unmap.size = dma->size;
>>>>>>>>> +
>>>>>>>>> + /*
>>>>>>>>> + * Notifier callback would call vfio_unpin_pages() which
>>>>>>>>> + * would acquire iommu->lock. Release lock here and
>>>>>>>>> + * reacquire it again.
>>>>>>>>> + */
>>>>>>>>> + mutex_unlock(&iommu->lock);
>>>>>>>>> + blocking_notifier_call_chain(&iommu->notifier,
>>>>>>>>> + VFIO_IOMMU_NOTIFY_DMA_UNMAP,
>>>>>>>>> + &nb_unmap);
>>>>>>>>> + mutex_lock(&iommu->lock);
>>>>>>>>> + if (WARN_ON(!RB_EMPTY_ROOT(&dma->pfn_list)))
>>>>>>>>> + break;
>>>>>>>>> + }
>>>>>>>>
>>>>>>>>
>>>>>>>> Why exactly do we need to notify per vfio_dma rather than per unmap
>>>>>>>> request? If we do the latter we can send the notify first, limiting us
>>>>>>>> to races where a page is pinned between the notify and the locking,
>>>>>>>> whereas here, even our dma pointer is suspect once we re-acquire the
>>>>>>>> lock, we don't technically know if another unmap could have removed
>>>>>>>> that already. Perhaps something like this (untested):
>>>>>>>>
>>>>>>>
>>>>>>> There are checks to validate unmap request, like v2 check and who is
>>>>>>> calling unmap and is it allowed for that task to unmap. Before these
>>>>>>> checks its not sure that unmap region range which asked for would be
>>>>>>> unmapped all. Notify call should be at the place where its sure that the
>>>>>>> range provided to notify call is definitely going to be removed. My
>>>>>>> change do that.
>>>>>>
>>>>>> Ok, but that does solve the problem. What about this (untested):
>>>>>
>>>>> s/does/does not/
>>>>>
>>>>> BTW, I like how the retries here fill the gap in my previous proposal
>>>>> where we could still race re-pinning. We've given it an honest shot or
>>>>> someone is not participating if we've retried 10 times. I don't
>>>>> understand why the test for iommu->external_domain was there, clearly
>>>>> if the list is not empty, we need to notify. Thanks,
>>>>>
>>>>
>>>> Ok. Retry is good to give a chance to unpin all. But is it really
>>>> required to use BUG_ON() that would panic the host. I think WARN_ON
>>>> should be fine and then when container is closed or when the last group
>>>> is removed from the container, vfio_iommu_type1_release() is called and
>>>> we have a chance to unpin it all.
>>>
>>> See my comments on patch 10/22, we need to be vigilant that the vendor
>>> driver is participating. I don't think we should be cleaning up after
>>> the vendor driver on release, if we need to do that, it implies we
>>> already have problems in multi-mdev containers since we'll be left with
>>> pfn_list entries that no longer have an owner. Thanks,
>>>
>>
>> If any vendor driver doesn't clean its pinned pages and there are
>> entries in pfn_list with no owner, that would be indicated by WARN_ON,
>> which should be fixed by that vendor driver. I still feel it shouldn't
>> cause host panic.
>> When such warning is seen with multiple mdev devices in container, it is
>> easy to isolate and find which vendor driver is not cleaning their
>> stuff, same warning would be seen with single mdev device in a
>> container. To isolate and find which vendor driver is culprit check with
>> one mdev device at a time.
>> Finally, we have a chance to clean all residue from
>> vfio_iommu_type1_release() so that vfio_iommu_type1 module doesn't leave
>> any leaks.
>
> How can we claim that we've resolved anything by unpinning the
> residue? In fact, is it actually safe to unpin any residue left by the
> vendor driver or does it imply that we're promoting a simple memory
> leak to a security issue because we can't verify whether the vendor
> driver has disabled access to that pfn, which may not reference a user
> page after we unpin it. That, in addition to the fact that I don't
> need to figure out how to break from the loop with a BUG_ON, is why I
> chose that rather than a WARN_ON. The release path could probably be a
> WARN_ON since the user no longer has access to the device, so we have a
> consistency error with the vendor driver, but we're probably not
> promoting it further by unpinning the pages. Thanks,
>
Ok. Agree with the security concern you mentioned.
Changing to BUG_ON as you suggested in vfio_dma_do_unmap() and replacing
'unpinning remaining pages on detach_group and release' with 'WARN_ON'
if there are unpinned pages.
Thanks,
Kirti
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web