Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1498046 > unrolled thread
| Started by | Haozhong Zhang <haozhong.zhang@intel.com> |
|---|---|
| First post | 2016-10-10 02:40 +0200 |
| Last post | 2016-10-12 09:30 +0200 |
| Articles | 20 on this page of 30 — 6 participants |
Back to article view | Back to linux.kernel
[RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Haozhong Zhang <haozhong.zhang@intel.com> - 2016-10-10 02:40 +0200
Re: [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Dan Williams <dan.j.williams@intel.com> - 2016-10-10 05:50 +0200
Re: [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Haozhong Zhang <haozhong.zhang@intel.com> - 2016-10-10 08:40 +0200
Re: [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Dan Williams <dan.j.williams@intel.com> - 2016-10-10 18:30 +0200
Re: [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Haozhong Zhang <haozhong.zhang@intel.com> - 2016-10-11 09:20 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Andrew Cooper <andrew.cooper3@citrix.com> - 2016-10-10 18:50 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Haozhong Zhang <haozhong.zhang@intel.com> - 2016-10-11 08:00 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-11 20:50 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-11 21:00 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Andrew Cooper <andrew.cooper3@citrix.com> - 2016-10-11 21:00 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen "Jan Beulich" <jbeulich@suse.com> - 2016-10-11 15:10 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Dan Williams <dan.j.williams@intel.com> - 2016-10-11 18:00 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-11 19:00 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Dan Williams <dan.j.williams@intel.com> - 2016-10-11 20:00 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Andrew Cooper <andrew.cooper3@citrix.com> - 2016-10-11 20:30 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-11 20:50 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-11 21:50 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-11 20:40 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Dan Williams <dan.j.williams@intel.com> - 2016-10-11 21:30 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> - 2016-10-11 22:00 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Andrew Cooper <andrew.cooper3@citrix.com> - 2016-10-11 22:20 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Dan Williams <dan.j.williams@intel.com> - 2016-10-11 22:20 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Haozhong Zhang <haozhong.zhang@intel.com> - 2016-10-12 12:40 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen "Jan Beulich" <JBeulich@suse.com> - 2016-10-12 13:40 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Haozhong Zhang <haozhong.zhang@intel.com> - 2016-10-12 17:10 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen "Jan Beulich" <JBeulich@suse.com> - 2016-10-12 17:40 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Dan Williams <dan.j.williams@intel.com> - 2016-10-12 17:50 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen "Jan Beulich" <JBeulich@suse.com> - 2016-10-12 18:30 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen Dan Williams <dan.j.williams@intel.com> - 2016-10-12 18:30 +0200
Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen "Jan Beulich" <JBeulich@suse.com> - 2016-10-12 09:30 +0200
Page 1 of 2 [1] 2 Next page →
| From | Haozhong Zhang <haozhong.zhang@intel.com> |
|---|---|
| Date | 2016-10-10 02:40 +0200 |
| Subject | [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <sqwNj-2FC-3@gated-at.bofh.it> |
Overview ======== This RFC kernel patch series along with corresponding patch series of Xen, QEMU and ndctl implements Xen vNVDIMM, which can map the host NVDIMM devices to Xen HVM domU as vNVDIMM devices. Xen hypervisor does not include an NVDIMM driver, so it needs the assistance from the driver in Dom0 Linux kernel to manage NVDIMM devices. We currently only supports NVDIMM devices in pmem mode. Design and Implementation ========================= The complete design can be found at https://lists.xenproject.org/archives/html/xen-devel/2016-07/msg01921.html. All patch series can be found at Xen: https://github.com/hzzhan9/xen.git nvdimm-rfc-v1 QEMU: https://github.com/hzzhan9/qemu.git xen-nvdimm-rfc-v1 Linux kernel: https://github.com/hzzhan9/nvdimm.git xen-nvdimm-rfc-v1 ndctl: https://github.com/hzzhan9/ndctl.git pfn-xen-rfc-v1 Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks: 1) Reserve an area on NVDIMM devices for Xen hypervisor to place memory management data structures, i.e. frame table and M2P table. 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen hypervisor. For 1), Patch 1 implements a new mode PFN_MODE_XEN to pfn devices, which make the reservation for Xen hypervisor. For 2), Patch 2 uses a new Xen hypercall to report the address information of pfn devices in PFN_MODE_XEN. How to test =========== Please refer to the cover letter of Xen patch series "[RFC XEN PATCH 00/16] Add vNVDIMM support to HVM domains". Haozhong Zhang (2): nvdimm: add PFN_MODE_XEN to pfn device for Xen usage xen, nvdimm: report pfn devices in PFN_MODE_XEN to Xen hypervisor drivers/nvdimm/namespace_devs.c | 2 ++ drivers/nvdimm/nd.h | 7 +++++ drivers/nvdimm/pfn_devs.c | 37 +++++++++++++++++++++--- drivers/nvdimm/pmem.c | 61 ++++++++++++++++++++++++++++++++++++++-- drivers/xen/Makefile | 2 +- drivers/xen/pmem.c | 53 ++++++++++++++++++++++++++++++++++ include/linux/pfn_t.h | 2 ++ include/xen/interface/platform.h | 13 +++++++++ include/xen/pmem.h | 32 +++++++++++++++++++++ 9 files changed, 201 insertions(+), 8 deletions(-) create mode 100644 drivers/xen/pmem.c create mode 100644 include/xen/pmem.h -- 2.10.1
[toc] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-10-10 05:50 +0200 |
| Message-ID | <sqzLb-4wh-5@gated-at.bofh.it> |
| In reply to | #1498046 |
On Sun, Oct 9, 2016 at 5:35 PM, Haozhong Zhang <haozhong.zhang@intel.com> wrote: > Overview > ======== > This RFC kernel patch series along with corresponding patch series of > Xen, QEMU and ndctl implements Xen vNVDIMM, which can map the host > NVDIMM devices to Xen HVM domU as vNVDIMM devices. > > Xen hypervisor does not include an NVDIMM driver, so it needs the > assistance from the driver in Dom0 Linux kernel to manage NVDIMM > devices. We currently only supports NVDIMM devices in pmem mode. > > Design and Implementation > ========================= > The complete design can be found at > https://lists.xenproject.org/archives/html/xen-devel/2016-07/msg01921.html. The KVM enabling for persistent memory does not need this support from the kernel, and as far as I can see neither does Xen. If the hypervisor needs to reserve some space it can simply trim the amount that it hands to the guest. The usage of fiemap and the sysfs resource for the pmem device, as mentioned in the design document, does not seem to comprehend that file block allocations may be discontiguous and may change over time depending on the file.
[toc] | [prev] | [next] | [standalone]
| From | Haozhong Zhang <haozhong.zhang@intel.com> |
|---|---|
| Date | 2016-10-10 08:40 +0200 |
| Message-ID | <sqCpH-6ic-9@gated-at.bofh.it> |
| In reply to | #1498079 |
On 10/09/16 20:45, Dan Williams wrote: > On Sun, Oct 9, 2016 at 5:35 PM, Haozhong Zhang <haozhong.zhang@intel.com> wrote: > > Overview > > ======== > > This RFC kernel patch series along with corresponding patch series of > > Xen, QEMU and ndctl implements Xen vNVDIMM, which can map the host > > NVDIMM devices to Xen HVM domU as vNVDIMM devices. > > > > Xen hypervisor does not include an NVDIMM driver, so it needs the > > assistance from the driver in Dom0 Linux kernel to manage NVDIMM > > devices. We currently only supports NVDIMM devices in pmem mode. > > > > Design and Implementation > > ========================= > > The complete design can be found at > > https://lists.xenproject.org/archives/html/xen-devel/2016-07/msg01921.html. > > The KVM enabling for persistent memory does not need this support from > the kernel, and as far as I can see neither does Xen. If the > hypervisor needs to reserve some space it can simply trim the amount > that it hands to the guest. > Xen does not have the NVDIMM driver, so it cannot operate on NVDIMM devices by itself. Instead it relies on the driver in Dom0 Linux to probe NVDIMM and make the reservation. > The usage of fiemap and the sysfs resource for the pmem device, as > mentioned in the design document, does not seem to comprehend that > file block allocations may be discontiguous and may change over time > depending on the file. True. I may need to find a way to notify Xen of the underlying changes, so that Xen can then adjust the address mapping. Thanks, Haozhong
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-10-10 18:30 +0200 |
| Message-ID | <sqLCG-3rZ-13@gated-at.bofh.it> |
| In reply to | #1498128 |
On Sun, Oct 9, 2016 at 11:32 PM, Haozhong Zhang <haozhong.zhang@intel.com> wrote: > On 10/09/16 20:45, Dan Williams wrote: >> On Sun, Oct 9, 2016 at 5:35 PM, Haozhong Zhang <haozhong.zhang@intel.com> wrote: >> > Overview >> > ======== >> > This RFC kernel patch series along with corresponding patch series of >> > Xen, QEMU and ndctl implements Xen vNVDIMM, which can map the host >> > NVDIMM devices to Xen HVM domU as vNVDIMM devices. >> > >> > Xen hypervisor does not include an NVDIMM driver, so it needs the >> > assistance from the driver in Dom0 Linux kernel to manage NVDIMM >> > devices. We currently only supports NVDIMM devices in pmem mode. >> > >> > Design and Implementation >> > ========================= >> > The complete design can be found at >> > https://lists.xenproject.org/archives/html/xen-devel/2016-07/msg01921.html. >> >> The KVM enabling for persistent memory does not need this support from >> the kernel, and as far as I can see neither does Xen. If the >> hypervisor needs to reserve some space it can simply trim the amount >> that it hands to the guest. >> > > Xen does not have the NVDIMM driver, so it cannot operate on NVDIMM > devices by itself. Instead it relies on the driver in Dom0 Linux to > probe NVDIMM and make the reservation. I'm missing something because the design document talks about mmap'ing files on a DAX filesystem. So, I'm assuming it is similar to the KVM NVDIMM virtualization case where an mmap range in dom0 is translated into a guest physical range. The suggestion is to reserve some memory out of that mapping rather than introduce a new info block / reservation type to the sub-system.
[toc] | [prev] | [next] | [standalone]
| From | Haozhong Zhang <haozhong.zhang@intel.com> |
|---|---|
| Date | 2016-10-11 09:20 +0200 |
| Message-ID | <sqZvX-3En-7@gated-at.bofh.it> |
| In reply to | #1498407 |
On 10/10/16 09:24, Dan Williams wrote: > On Sun, Oct 9, 2016 at 11:32 PM, Haozhong Zhang > <haozhong.zhang@intel.com> wrote: > > On 10/09/16 20:45, Dan Williams wrote: > >> On Sun, Oct 9, 2016 at 5:35 PM, Haozhong Zhang <haozhong.zhang@intel.com> wrote: > >> > Overview > >> > ======== > >> > This RFC kernel patch series along with corresponding patch series of > >> > Xen, QEMU and ndctl implements Xen vNVDIMM, which can map the host > >> > NVDIMM devices to Xen HVM domU as vNVDIMM devices. > >> > > >> > Xen hypervisor does not include an NVDIMM driver, so it needs the > >> > assistance from the driver in Dom0 Linux kernel to manage NVDIMM > >> > devices. We currently only supports NVDIMM devices in pmem mode. > >> > > >> > Design and Implementation > >> > ========================= > >> > The complete design can be found at > >> > https://lists.xenproject.org/archives/html/xen-devel/2016-07/msg01921.html. > >> > >> The KVM enabling for persistent memory does not need this support from > >> the kernel, and as far as I can see neither does Xen. If the > >> hypervisor needs to reserve some space it can simply trim the amount > >> that it hands to the guest. > >> > > > > Xen does not have the NVDIMM driver, so it cannot operate on NVDIMM > > devices by itself. Instead it relies on the driver in Dom0 Linux to > > probe NVDIMM and make the reservation. > > I'm missing something because the design document talks about mmap'ing > files on a DAX filesystem. So, I'm assuming it is similar to the KVM > NVDIMM virtualization case where an mmap range in dom0 is translated > into a guest physical range. The suggestion is to reserve some memory > out of that mapping rather than introduce a new info block / > reservation type to the sub-system. Just like struct page to linux, Xen hypervisor uses a struct page_info for its memory management. We are facing the same problem as linux kernel: where we store those structs for pmem, and decided to put them on a reserved area on pmem, similar to what pfn device in kernel does. Reserving at the moment of mmap and out of what is mapped does not work. It's a bootstrap problem: Xen needs the information of those pages, which are stored in struct page_info, at the moment of mapping. That is, page_info structs for pmem pages should be prepared before those pages are actually used. However, as the ongoing discussion in another thread with Andrew Cooper, if Xen hypervisor turns to treat pmem pages as MMIO, then the reservation may not be needed. Let's see what conclusion will be reached there. Thanks, Haozhong
[toc] | [prev] | [next] | [standalone]
| From | Andrew Cooper <andrew.cooper3@citrix.com> |
|---|---|
| Date | 2016-10-10 18:50 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <sqLW2-3yz-37@gated-at.bofh.it> |
| In reply to | #1498046 |
On 10/10/16 01:35, Haozhong Zhang wrote: > Overview > ======== > This RFC kernel patch series along with corresponding patch series of > Xen, QEMU and ndctl implements Xen vNVDIMM, which can map the host > NVDIMM devices to Xen HVM domU as vNVDIMM devices. > > Xen hypervisor does not include an NVDIMM driver, so it needs the > assistance from the driver in Dom0 Linux kernel to manage NVDIMM > devices. We currently only supports NVDIMM devices in pmem mode. > > Design and Implementation > ========================= > The complete design can be found at > https://lists.xenproject.org/archives/html/xen-devel/2016-07/msg01921.html. > > All patch series can be found at > Xen: https://github.com/hzzhan9/xen.git nvdimm-rfc-v1 > QEMU: https://github.com/hzzhan9/qemu.git xen-nvdimm-rfc-v1 > Linux kernel: https://github.com/hzzhan9/nvdimm.git xen-nvdimm-rfc-v1 > ndctl: https://github.com/hzzhan9/ndctl.git pfn-xen-rfc-v1 > > Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks: > 1) Reserve an area on NVDIMM devices for Xen hypervisor to place > memory management data structures, i.e. frame table and M2P table. > 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen > hypervisor. Please can we take a step back here before diving down a rabbit hole. How do pblk/pmem regions appear in the E820 map at boot? At the very least, I would expect at least a large reserved region. Is the MFN information (SPA in your terminology, so far as I can tell) available in any static APCI tables, or are they only available as a result of executing AML methods? If the MFN information is only available via AML, then point 2) is needed, although the reporting back to Xen should be restricted to a xen component, rather than polluting the main device driver. However, I can't see any justification for 1). Dom0 should not be involved in Xen's management of its own frame table and m2p. The mfns making up the pmem/pblk regions should be treated just like any other MMIO regions, and be handed wholesale to dom0 by default. ~Andrew
[toc] | [prev] | [next] | [standalone]
| From | Haozhong Zhang <haozhong.zhang@intel.com> |
|---|---|
| Date | 2016-10-11 08:00 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <sqYgx-2Hs-7@gated-at.bofh.it> |
| In reply to | #1498425 |
On 10/10/16 17:43, Andrew Cooper wrote: > On 10/10/16 01:35, Haozhong Zhang wrote: > > Overview > > ======== > > This RFC kernel patch series along with corresponding patch series of > > Xen, QEMU and ndctl implements Xen vNVDIMM, which can map the host > > NVDIMM devices to Xen HVM domU as vNVDIMM devices. > > > > Xen hypervisor does not include an NVDIMM driver, so it needs the > > assistance from the driver in Dom0 Linux kernel to manage NVDIMM > > devices. We currently only supports NVDIMM devices in pmem mode. > > > > Design and Implementation > > ========================= > > The complete design can be found at > > https://lists.xenproject.org/archives/html/xen-devel/2016-07/msg01921.html. > > > > All patch series can be found at > > Xen: https://github.com/hzzhan9/xen.git nvdimm-rfc-v1 > > QEMU: https://github.com/hzzhan9/qemu.git xen-nvdimm-rfc-v1 > > Linux kernel: https://github.com/hzzhan9/nvdimm.git xen-nvdimm-rfc-v1 > > ndctl: https://github.com/hzzhan9/ndctl.git pfn-xen-rfc-v1 > > > > Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks: > > 1) Reserve an area on NVDIMM devices for Xen hypervisor to place > > memory management data structures, i.e. frame table and M2P table. > > 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen > > hypervisor. > > Please can we take a step back here before diving down a rabbit hole. > > > How do pblk/pmem regions appear in the E820 map at boot? At the very > least, I would expect at least a large reserved region. ACPI specification does not require them to appear in E820, though it defines E820 type-7 for persistent memory. > > Is the MFN information (SPA in your terminology, so far as I can tell) > available in any static APCI tables, or are they only available as a > result of executing AML methods? > For NVDIMM devices already plugged at power on, their MFN information can be got from NFIT table. However, MFN information for hotplugged NVDIMM devices should be got via AML _FIT method, so point 2) is needed. > > If the MFN information is only available via AML, then point 2) is > needed, although the reporting back to Xen should be restricted to a xen > component, rather than polluting the main device driver. > > However, I can't see any justification for 1). Dom0 should not be > involved in Xen's management of its own frame table and m2p. The mfns > making up the pmem/pblk regions should be treated just like any other > MMIO regions, and be handed wholesale to dom0 by default. > Do you mean to treat them as mmio pages of type p2m_mmio_direct and map them to guest by map_mmio_regions()? Thanks, Haozhong
[toc] | [prev] | [next] | [standalone]
| From | Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> |
|---|---|
| Date | 2016-10-11 20:50 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <srahH-1Pi-11@gated-at.bofh.it> |
| In reply to | #1498620 |
On Tue, Oct 11, 2016 at 07:37:09PM +0100, Andrew Cooper wrote: > On 11/10/16 06:52, Haozhong Zhang wrote: > > On 10/10/16 17:43, Andrew Cooper wrote: > >> On 10/10/16 01:35, Haozhong Zhang wrote: > >>> Overview > >>> ======== > >>> This RFC kernel patch series along with corresponding patch series of > >>> Xen, QEMU and ndctl implements Xen vNVDIMM, which can map the host > >>> NVDIMM devices to Xen HVM domU as vNVDIMM devices. > >>> > >>> Xen hypervisor does not include an NVDIMM driver, so it needs the > >>> assistance from the driver in Dom0 Linux kernel to manage NVDIMM > >>> devices. We currently only supports NVDIMM devices in pmem mode. > >>> > >>> Design and Implementation > >>> ========================= > >>> The complete design can be found at > >>> https://lists.xenproject.org/archives/html/xen-devel/2016-07/msg01921.html. > >>> > >>> All patch series can be found at > >>> Xen: https://github.com/hzzhan9/xen.git nvdimm-rfc-v1 > >>> QEMU: https://github.com/hzzhan9/qemu.git xen-nvdimm-rfc-v1 > >>> Linux kernel: https://github.com/hzzhan9/nvdimm.git xen-nvdimm-rfc-v1 > >>> ndctl: https://github.com/hzzhan9/ndctl.git pfn-xen-rfc-v1 > >>> > >>> Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks: > >>> 1) Reserve an area on NVDIMM devices for Xen hypervisor to place > >>> memory management data structures, i.e. frame table and M2P table. > >>> 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen > >>> hypervisor. > >> Please can we take a step back here before diving down a rabbit hole. > >> > >> > >> How do pblk/pmem regions appear in the E820 map at boot? At the very > >> least, I would expect at least a large reserved region. > > ACPI specification does not require them to appear in E820, though > > it defines E820 type-7 for persistent memory. > > Ok, so we might get some E820 type-7 ranges, or some holes. > > > > >> Is the MFN information (SPA in your terminology, so far as I can tell) > >> available in any static APCI tables, or are they only available as a > >> result of executing AML methods? > >> > > For NVDIMM devices already plugged at power on, their MFN information > > can be got from NFIT table. However, MFN information for hotplugged > > NVDIMM devices should be got via AML _FIT method, so point 2) is needed. > > How does NVDIMM hotplug compare to RAM hotplug? Are the hotplug regions > described at boot and marked as initially not present, or do you only > know the hotplugged SPA at the point that it is hotplugged? > > I certainly agree that there needs to be a propagation of the hotplug > notification from OSPM to Xen, which will involve some glue in the Xen > subsystem in Linux, but I would expect that this would be similar to the > existing plain RAM hotplug mechanism. > > > > >> If the MFN information is only available via AML, then point 2) is > >> needed, although the reporting back to Xen should be restricted to a xen > >> component, rather than polluting the main device driver. > >> > >> However, I can't see any justification for 1). Dom0 should not be > >> involved in Xen's management of its own frame table and m2p. The mfns > >> making up the pmem/pblk regions should be treated just like any other > >> MMIO regions, and be handed wholesale to dom0 by default. > >> > > Do you mean to treat them as mmio pages of type p2m_mmio_direct and > > map them to guest by map_mmio_regions()? > > I don't see any reason why it shouldn't be treated like this. Xen > shouldn't be treating it as anything other than an opaque block of MFNs. > > The concept of trying to map a DAX file into the guest physical address > space of a VM is indeed new and doesn't fit into Xen's current model, > but all that fixing this requires is a new privileged mapping hypercall > which takes a source domid and gfn scatter list, and a destination domid > and scatter list. (I see from a quick look at your Xen series that your > XENMEM_populate_pmemmap looks roughly like this) That can be quite big. Say you want to map an DAX file that has size of 1TB and the this GFN scatter list has 1073741824 entries? How do you envision handling this in Xen and populating the P2M entries with this information? > > ~Andrew > _______________________________________________ > Linux-nvdimm mailing list > Linux-nvdimm@lists.01.org > https://lists.01.org/mailman/listinfo/linux-nvdimm
[toc] | [prev] | [next] | [standalone]
| From | Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> |
|---|---|
| Date | 2016-10-11 21:00 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <srarn-1SH-5@gated-at.bofh.it> |
| In reply to | #1498620 |
On Tue, Oct 11, 2016 at 07:37:09PM +0100, Andrew Cooper wrote: > On 11/10/16 06:52, Haozhong Zhang wrote: > > On 10/10/16 17:43, Andrew Cooper wrote: > >> On 10/10/16 01:35, Haozhong Zhang wrote: > >>> Overview > >>> ======== > >>> This RFC kernel patch series along with corresponding patch series of > >>> Xen, QEMU and ndctl implements Xen vNVDIMM, which can map the host > >>> NVDIMM devices to Xen HVM domU as vNVDIMM devices. > >>> > >>> Xen hypervisor does not include an NVDIMM driver, so it needs the > >>> assistance from the driver in Dom0 Linux kernel to manage NVDIMM > >>> devices. We currently only supports NVDIMM devices in pmem mode. > >>> > >>> Design and Implementation > >>> ========================= > >>> The complete design can be found at > >>> https://lists.xenproject.org/archives/html/xen-devel/2016-07/msg01921.html. > >>> > >>> All patch series can be found at > >>> Xen: https://github.com/hzzhan9/xen.git nvdimm-rfc-v1 > >>> QEMU: https://github.com/hzzhan9/qemu.git xen-nvdimm-rfc-v1 > >>> Linux kernel: https://github.com/hzzhan9/nvdimm.git xen-nvdimm-rfc-v1 > >>> ndctl: https://github.com/hzzhan9/ndctl.git pfn-xen-rfc-v1 > >>> > >>> Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks: > >>> 1) Reserve an area on NVDIMM devices for Xen hypervisor to place > >>> memory management data structures, i.e. frame table and M2P table. > >>> 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen > >>> hypervisor. > >> Please can we take a step back here before diving down a rabbit hole. > >> > >> > >> How do pblk/pmem regions appear in the E820 map at boot? At the very > >> least, I would expect at least a large reserved region. > > ACPI specification does not require them to appear in E820, though > > it defines E820 type-7 for persistent memory. > > Ok, so we might get some E820 type-7 ranges, or some holes. > > > > >> Is the MFN information (SPA in your terminology, so far as I can tell) > >> available in any static APCI tables, or are they only available as a > >> result of executing AML methods? > >> > > For NVDIMM devices already plugged at power on, their MFN information > > can be got from NFIT table. However, MFN information for hotplugged > > NVDIMM devices should be got via AML _FIT method, so point 2) is needed. > > How does NVDIMM hotplug compare to RAM hotplug? Are the hotplug regions > described at boot and marked as initially not present, or do you only > know the hotplugged SPA at the point that it is hotplugged? The latter. You have no idea of the size until you get an ACPI hotplug. The ACPI hotplug contains the NFIT MADT table so based on that you can populate the machine. > > I certainly agree that there needs to be a propagation of the hotplug > notification from OSPM to Xen, which will involve some glue in the Xen > subsystem in Linux, but I would expect that this would be similar to the > existing plain RAM hotplug mechanism. I am actually not sure how ACPI RAM hotplug mechanism is suppose to work in practice. I thought that the regions (E820) are marked as reserved and the 'RAM' slots nicely in there.
[toc] | [prev] | [next] | [standalone]
| From | Andrew Cooper <andrew.cooper3@citrix.com> |
|---|---|
| Date | 2016-10-11 21:00 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <srahH-1Pi-13@gated-at.bofh.it> |
| In reply to | #1498620 |
On 11/10/16 06:52, Haozhong Zhang wrote: > On 10/10/16 17:43, Andrew Cooper wrote: >> On 10/10/16 01:35, Haozhong Zhang wrote: >>> Overview >>> ======== >>> This RFC kernel patch series along with corresponding patch series of >>> Xen, QEMU and ndctl implements Xen vNVDIMM, which can map the host >>> NVDIMM devices to Xen HVM domU as vNVDIMM devices. >>> >>> Xen hypervisor does not include an NVDIMM driver, so it needs the >>> assistance from the driver in Dom0 Linux kernel to manage NVDIMM >>> devices. We currently only supports NVDIMM devices in pmem mode. >>> >>> Design and Implementation >>> ========================= >>> The complete design can be found at >>> https://lists.xenproject.org/archives/html/xen-devel/2016-07/msg01921.html. >>> >>> All patch series can be found at >>> Xen: https://github.com/hzzhan9/xen.git nvdimm-rfc-v1 >>> QEMU: https://github.com/hzzhan9/qemu.git xen-nvdimm-rfc-v1 >>> Linux kernel: https://github.com/hzzhan9/nvdimm.git xen-nvdimm-rfc-v1 >>> ndctl: https://github.com/hzzhan9/ndctl.git pfn-xen-rfc-v1 >>> >>> Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks: >>> 1) Reserve an area on NVDIMM devices for Xen hypervisor to place >>> memory management data structures, i.e. frame table and M2P table. >>> 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen >>> hypervisor. >> Please can we take a step back here before diving down a rabbit hole. >> >> >> How do pblk/pmem regions appear in the E820 map at boot? At the very >> least, I would expect at least a large reserved region. > ACPI specification does not require them to appear in E820, though > it defines E820 type-7 for persistent memory. Ok, so we might get some E820 type-7 ranges, or some holes. > >> Is the MFN information (SPA in your terminology, so far as I can tell) >> available in any static APCI tables, or are they only available as a >> result of executing AML methods? >> > For NVDIMM devices already plugged at power on, their MFN information > can be got from NFIT table. However, MFN information for hotplugged > NVDIMM devices should be got via AML _FIT method, so point 2) is needed. How does NVDIMM hotplug compare to RAM hotplug? Are the hotplug regions described at boot and marked as initially not present, or do you only know the hotplugged SPA at the point that it is hotplugged? I certainly agree that there needs to be a propagation of the hotplug notification from OSPM to Xen, which will involve some glue in the Xen subsystem in Linux, but I would expect that this would be similar to the existing plain RAM hotplug mechanism. > >> If the MFN information is only available via AML, then point 2) is >> needed, although the reporting back to Xen should be restricted to a xen >> component, rather than polluting the main device driver. >> >> However, I can't see any justification for 1). Dom0 should not be >> involved in Xen's management of its own frame table and m2p. The mfns >> making up the pmem/pblk regions should be treated just like any other >> MMIO regions, and be handed wholesale to dom0 by default. >> > Do you mean to treat them as mmio pages of type p2m_mmio_direct and > map them to guest by map_mmio_regions()? I don't see any reason why it shouldn't be treated like this. Xen shouldn't be treating it as anything other than an opaque block of MFNs. The concept of trying to map a DAX file into the guest physical address space of a VM is indeed new and doesn't fit into Xen's current model, but all that fixing this requires is a new privileged mapping hypercall which takes a source domid and gfn scatter list, and a destination domid and scatter list. (I see from a quick look at your Xen series that your XENMEM_populate_pmemmap looks roughly like this) ~Andrew
[toc] | [prev] | [next] | [standalone]
| From | "Jan Beulich" <jbeulich@suse.com> |
|---|---|
| Date | 2016-10-11 15:10 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <sr4YG-77t-15@gated-at.bofh.it> |
| In reply to | #1498425 |
>>> Andrew Cooper <andrew.cooper3@citrix.com> 10/10/16 6:44 PM >>> >On 10/10/16 01:35, Haozhong Zhang wrote: >> Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks: >> 1) Reserve an area on NVDIMM devices for Xen hypervisor to place >> memory management data structures, i.e. frame table and M2P table. >> 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen >> hypervisor. > >However, I can't see any justification for 1). Dom0 should not be >involved in Xen's management of its own frame table and m2p. The mfns >making up the pmem/pblk regions should be treated just like any other >MMIO regions, and be handed wholesale to dom0 by default. That precludes the use as RAM extension, and I thought earlier rounds of discussion had got everyone in agreement that at least for the pmem case we will need some control data in Xen. Jan
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-10-11 18:00 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <sr7Dc-84-1@gated-at.bofh.it> |
| In reply to | #1498871 |
On Tue, Oct 11, 2016 at 6:08 AM, Jan Beulich <jbeulich@suse.com> wrote: >>>> Andrew Cooper <andrew.cooper3@citrix.com> 10/10/16 6:44 PM >>> >>On 10/10/16 01:35, Haozhong Zhang wrote: >>> Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks: >>> 1) Reserve an area on NVDIMM devices for Xen hypervisor to place >>> memory management data structures, i.e. frame table and M2P table. >>> 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen >>> hypervisor. >> >>However, I can't see any justification for 1). Dom0 should not be >>involved in Xen's management of its own frame table and m2p. The mfns >>making up the pmem/pblk regions should be treated just like any other >>MMIO regions, and be handed wholesale to dom0 by default. > > That precludes the use as RAM extension, and I thought earlier rounds of > discussion had got everyone in agreement that at least for the pmem case > we will need some control data in Xen. The missing piece for me is why this reservation for control data needs to be done in the libnvdimm core? I would expect that any dax capable file could be mapped and made available to a guest. This includes /dev/ramX devices that are dax capable, but are external to the libnvdimm sub-system.
[toc] | [prev] | [next] | [standalone]
| From | Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> |
|---|---|
| Date | 2016-10-11 19:00 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <sr8zf-H0-17@gated-at.bofh.it> |
| In reply to | #1498992 |
On Tue, Oct 11, 2016 at 08:53:33AM -0700, Dan Williams wrote: > On Tue, Oct 11, 2016 at 6:08 AM, Jan Beulich <jbeulich@suse.com> wrote: > >>>> Andrew Cooper <andrew.cooper3@citrix.com> 10/10/16 6:44 PM >>> > >>On 10/10/16 01:35, Haozhong Zhang wrote: > >>> Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks: > >>> 1) Reserve an area on NVDIMM devices for Xen hypervisor to place > >>> memory management data structures, i.e. frame table and M2P table. > >>> 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen > >>> hypervisor. > >> > >>However, I can't see any justification for 1). Dom0 should not be > >>involved in Xen's management of its own frame table and m2p. The mfns > >>making up the pmem/pblk regions should be treated just like any other > >>MMIO regions, and be handed wholesale to dom0 by default. > > > > That precludes the use as RAM extension, and I thought earlier rounds of > > discussion had got everyone in agreement that at least for the pmem case > > we will need some control data in Xen. > > The missing piece for me is why this reservation for control data > needs to be done in the libnvdimm core? I would expect that any dax Isn't it done this way with Linux? That is say if the machine has 4GB of RAM and the NVDIMM is in TB range. You want to put the 'struct page' for the NVDIMM ranges somewhere. That place can be in regions on the NVDIMM that ndctl can reserve. > capable file could be mapped and made available to a guest. This > includes /dev/ramX devices that are dax capable, but are external to > the libnvdimm sub-system. This is more of just keeping track of the ranges if say the DAX file is extremely fragmented and requires a lot of 'struct pages' to keep track of when stiching up the VMA. > > _______________________________________________ > Xen-devel mailing list > Xen-devel@lists.xen.org > https://lists.xen.org/xen-devel
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-10-11 20:00 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <sr9vj-1fU-1@gated-at.bofh.it> |
| In reply to | #1499017 |
On Tue, Oct 11, 2016 at 9:58 AM, Konrad Rzeszutek Wilk
<konrad.wilk@oracle.com> wrote:
> On Tue, Oct 11, 2016 at 08:53:33AM -0700, Dan Williams wrote:
>> On Tue, Oct 11, 2016 at 6:08 AM, Jan Beulich <jbeulich@suse.com> wrote:
>> >>>> Andrew Cooper <andrew.cooper3@citrix.com> 10/10/16 6:44 PM >>>
>> >>On 10/10/16 01:35, Haozhong Zhang wrote:
>> >>> Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks:
>> >>> 1) Reserve an area on NVDIMM devices for Xen hypervisor to place
>> >>> memory management data structures, i.e. frame table and M2P table.
>> >>> 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen
>> >>> hypervisor.
>> >>
>> >>However, I can't see any justification for 1). Dom0 should not be
>> >>involved in Xen's management of its own frame table and m2p. The mfns
>> >>making up the pmem/pblk regions should be treated just like any other
>> >>MMIO regions, and be handed wholesale to dom0 by default.
>> >
>> > That precludes the use as RAM extension, and I thought earlier rounds of
>> > discussion had got everyone in agreement that at least for the pmem case
>> > we will need some control data in Xen.
>>
>> The missing piece for me is why this reservation for control data
>> needs to be done in the libnvdimm core? I would expect that any dax
>
> Isn't it done this way with Linux? That is say if the machine has
> 4GB of RAM and the NVDIMM is in TB range. You want to put the 'struct page'
> for the NVDIMM ranges somewhere. That place can be in regions on the
> NVDIMM that ndctl can reserve.
Yes.
>> capable file could be mapped and made available to a guest. This
>> includes /dev/ramX devices that are dax capable, but are external to
>> the libnvdimm sub-system.
>
> This is more of just keeping track of the ranges if say the DAX file is
> extremely fragmented and requires a lot of 'struct pages' to keep track of
> when stiching up the VMA.
Right, but why does the libnvdimm core need to know about this
specific Xen reservation? For example, if Xen wants some in-kernel
driver to own a pmem region and place its own metadata on the device I
would recommend something like:
bdev = blkdev_get_by_path("/dev/pmemX", FMODE_EXCL...);
bdev_direct_access(bdev, ...);
...in other words, I don't think we want libnvdimm to grow new device
types for every possible in-kernel user, Xen, MD, DM, etc. Instead,
just claim the resulting device.
[toc] | [prev] | [next] | [standalone]
| From | Andrew Cooper <andrew.cooper3@citrix.com> |
|---|---|
| Date | 2016-10-11 20:30 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <sr9Ym-1IC-27@gated-at.bofh.it> |
| In reply to | #1499085 |
On 11/10/16 18:51, Dan Williams wrote:
> On Tue, Oct 11, 2016 at 9:58 AM, Konrad Rzeszutek Wilk
> <konrad.wilk@oracle.com> wrote:
>> On Tue, Oct 11, 2016 at 08:53:33AM -0700, Dan Williams wrote:
>>> On Tue, Oct 11, 2016 at 6:08 AM, Jan Beulich <jbeulich@suse.com> wrote:
>>>>>>> Andrew Cooper <andrew.cooper3@citrix.com> 10/10/16 6:44 PM >>>
>>>>> On 10/10/16 01:35, Haozhong Zhang wrote:
>>>>>> Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks:
>>>>>> 1) Reserve an area on NVDIMM devices for Xen hypervisor to place
>>>>>> memory management data structures, i.e. frame table and M2P table.
>>>>>> 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen
>>>>>> hypervisor.
>>>>> However, I can't see any justification for 1). Dom0 should not be
>>>>> involved in Xen's management of its own frame table and m2p. The mfns
>>>>> making up the pmem/pblk regions should be treated just like any other
>>>>> MMIO regions, and be handed wholesale to dom0 by default.
>>>> That precludes the use as RAM extension, and I thought earlier rounds of
>>>> discussion had got everyone in agreement that at least for the pmem case
>>>> we will need some control data in Xen.
>>> The missing piece for me is why this reservation for control data
>>> needs to be done in the libnvdimm core? I would expect that any dax
>> Isn't it done this way with Linux? That is say if the machine has
>> 4GB of RAM and the NVDIMM is in TB range. You want to put the 'struct page'
>> for the NVDIMM ranges somewhere. That place can be in regions on the
>> NVDIMM that ndctl can reserve.
> Yes.
I do not see any sensible usecase for Xen to use NVDIMMs as plain RAM;
NVDIMMs are far more valuable for higher level management in dom0.
I certainly think that such a usecase should be out-of-scope for initial
Xen/NVDIMM support, even if only to reduce the complexity to start with.
A repeated complain I have of large feature submissions like this is
that, by trying to solve all potential usecases at one, end up being
overly complicated to develop, understand and review.
>
>>> capable file could be mapped and made available to a guest. This
>>> includes /dev/ramX devices that are dax capable, but are external to
>>> the libnvdimm sub-system.
>> This is more of just keeping track of the ranges if say the DAX file is
>> extremely fragmented and requires a lot of 'struct pages' to keep track of
>> when stiching up the VMA.
> Right, but why does the libnvdimm core need to know about this
> specific Xen reservation? For example, if Xen wants some in-kernel
> driver to own a pmem region and place its own metadata on the device I
> would recommend something like:
>
> bdev = blkdev_get_by_path("/dev/pmemX", FMODE_EXCL...);
> bdev_direct_access(bdev, ...);
>
> ...in other words, I don't think we want libnvdimm to grow new device
> types for every possible in-kernel user, Xen, MD, DM, etc. Instead,
> just claim the resulting device.
I completely agree.
Whatever ends up happening between Xen and dom0, there should be no
modifications like this to the nvdimm driver. I will go so far as to
say that there shouldn't be any modifications to the nvdimm driver
(other than perhaps new query hooks so the Xen subsystem in Linux can
query information to then pass up to Xen, if the existing queryability
is insufficient).
~Andrew
[toc] | [prev] | [next] | [standalone]
| From | Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> |
|---|---|
| Date | 2016-10-11 20:50 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <srahI-1Pi-15@gated-at.bofh.it> |
| In reply to | #1499135 |
On Tue, Oct 11, 2016 at 07:15:42PM +0100, Andrew Cooper wrote:
> On 11/10/16 18:51, Dan Williams wrote:
> > On Tue, Oct 11, 2016 at 9:58 AM, Konrad Rzeszutek Wilk
> > <konrad.wilk@oracle.com> wrote:
> >> On Tue, Oct 11, 2016 at 08:53:33AM -0700, Dan Williams wrote:
> >>> On Tue, Oct 11, 2016 at 6:08 AM, Jan Beulich <jbeulich@suse.com> wrote:
> >>>>>>> Andrew Cooper <andrew.cooper3@citrix.com> 10/10/16 6:44 PM >>>
> >>>>> On 10/10/16 01:35, Haozhong Zhang wrote:
> >>>>>> Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks:
> >>>>>> 1) Reserve an area on NVDIMM devices for Xen hypervisor to place
> >>>>>> memory management data structures, i.e. frame table and M2P table.
> >>>>>> 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen
> >>>>>> hypervisor.
> >>>>> However, I can't see any justification for 1). Dom0 should not be
> >>>>> involved in Xen's management of its own frame table and m2p. The mfns
> >>>>> making up the pmem/pblk regions should be treated just like any other
> >>>>> MMIO regions, and be handed wholesale to dom0 by default.
> >>>> That precludes the use as RAM extension, and I thought earlier rounds of
> >>>> discussion had got everyone in agreement that at least for the pmem case
> >>>> we will need some control data in Xen.
> >>> The missing piece for me is why this reservation for control data
> >>> needs to be done in the libnvdimm core? I would expect that any dax
> >> Isn't it done this way with Linux? That is say if the machine has
> >> 4GB of RAM and the NVDIMM is in TB range. You want to put the 'struct page'
> >> for the NVDIMM ranges somewhere. That place can be in regions on the
> >> NVDIMM that ndctl can reserve.
> > Yes.
>
> I do not see any sensible usecase for Xen to use NVDIMMs as plain RAM;
I just gave you one. This is the 'usecase' that Linux has to deal with
now that the core kernel folks have pointed out that they don't want
'struct page' for the MMIO regions. This mechanism came about this and
finding a place _somewhere_ to deal with having to have 'struct page'
for the SPA ranges of the NVDIMM.
> NVDIMMs are far more valuable for higher level management in dom0.
Andrew, why are you providing input to this so late?
Haozhong provided an nice design document outlining the problem and
the solution he suggested.
>
> I certainly think that such a usecase should be out-of-scope for initial
> Xen/NVDIMM support, even if only to reduce the complexity to start with.
>
> A repeated complain I have of large feature submissions like this is
> that, by trying to solve all potential usecases at one, end up being
> overly complicated to develop, understand and review.
On the other hand - if you don't take these complicated issues from the
start, then you may have to redesign and redevelop this after the first
version which has been set in stone and committed.
>
> >
> >>> capable file could be mapped and made available to a guest. This
> >>> includes /dev/ramX devices that are dax capable, but are external to
> >>> the libnvdimm sub-system.
> >> This is more of just keeping track of the ranges if say the DAX file is
> >> extremely fragmented and requires a lot of 'struct pages' to keep track of
> >> when stiching up the VMA.
> > Right, but why does the libnvdimm core need to know about this
> > specific Xen reservation? For example, if Xen wants some in-kernel
> > driver to own a pmem region and place its own metadata on the device I
> > would recommend something like:
> >
> > bdev = blkdev_get_by_path("/dev/pmemX", FMODE_EXCL...);
> > bdev_direct_access(bdev, ...);
> >
> > ...in other words, I don't think we want libnvdimm to grow new device
> > types for every possible in-kernel user, Xen, MD, DM, etc. Instead,
> > just claim the resulting device.
>
> I completely agree.
>
> Whatever ends up happening between Xen and dom0, there should be no
> modifications like this to the nvdimm driver. I will go so far as to
> say that there shouldn't be any modifications to the nvdimm driver
> (other than perhaps new query hooks so the Xen subsystem in Linux can
> query information to then pass up to Xen, if the existing queryability
> is insufficient).
Haozhong and Jan had been chatting about this in terms of how to keep
track of a guest having non-contingous SPAs of NVDIMM stiched to a guest.
The initial idea was to treat it as MMIO, but of course if you have 1
page ranges over say 1TB you end up consuming tons of memory to keep
track of this (the same way Linux would if you wanted to mmap an file
from DAX fs).
Other solutions were an bitmap, but that can also be cumbersome to deal
with. In the end the suggestion that was proposed was the one that Linux
choose - stash the 'struct page' in the NVDIMM.
[toc] | [prev] | [next] | [standalone]
| From | Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> |
|---|---|
| Date | 2016-10-11 21:50 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <srbdL-2oc-5@gated-at.bofh.it> |
| In reply to | #1499141 |
> Andrew, why are you providing input to this so late? First of sorry for this outburst. It was quite uncalled for and quite unprofessional. You of all people have so much on your plate that I am astonished that you are able to operate with some many pokers in the fire. Again, my sincere apology.
[toc] | [prev] | [next] | [standalone]
| From | Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> |
|---|---|
| Date | 2016-10-11 20:40 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <sra82-1LS-11@gated-at.bofh.it> |
| In reply to | #1499085 |
On Tue, Oct 11, 2016 at 10:51:19AM -0700, Dan Williams wrote:
> 260sn3756f-1
> (version=TLSv1.2 cipher=ECDHE-RSA-AES128-GCM-SHA256 bits=128 verify=NOT)
> for <konrad.wilk@oracle.com>; Tue, 11 Oct 2016 17:51:21 +0000
> Received: by mail-oi0-f43.google.com with SMTP id d132so32700570oib.2
> for <konrad.wilk@oracle.com>; Tue, 11 Oct 2016 10:51:20 -0700 (PDT)
> DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
> d=intel-com.20150623.gappssmtp.com; s=20150623;
> h=mime-version:in-reply-to:references:from:date:message-id:subject:to
> :cc;
> bh=vXHG8Ke0lr+jk8ivMDq3ZpmmHHjC205aTSytpqjXFgo=;
> b=CiKg4tJf1DGU2x/pSCYU7Jx79oCXMSIApwY2zJjO9Lny3erPxUyjNhszNyQkceYK1A
> Gzuw05eETGT/k0UWamFdN/ZXF3PucSXIXqrVtTS9kLQBlKPTWQJvndSRqZ6lPb36mlSA
> BrkdOREz5O/V7p/iGYhnxZU9eyfVY1ekgeMvTKP3su9Ye4Nk6GJYMEb5HSTCm1Ckmoq5
> T4Rlw6gcnbHCLx27vcghySG4YXcQ4r2qSPcSmAysve77sYCPYlM9XRVpzfPBTmINKGUo
> 9w7MgVs5KG0dG60j1fJNjXoY0WSoP3uI67e69afqjAChzVndGDgMXjOzGrQ6+KQF088Q
> JeiQ==
> X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
> d=1e100.net; s=20130820;
> h=x-gm-message-state:mime-version:in-reply-to:references:from:date
> :message-id:subject:to:cc;
> bh=vXHG8Ke0lr+jk8ivMDq3ZpmmHHjC205aTSytpqjXFgo=;
> b=SmizBvFSmUHAy/WKfbD4m+QVSajIfcD9SQW7hwqmiwUtrACa2PxQWyx0dHe6DOqVVx
> jYHSxbbMiz105BMwxfv2pZlAl+phFkj8APxpL2XF36SIsq5u9+evlqBUuzGcpVJ+tXyI
> 0xO0qfyspvNwLwJnkZ2bOxO9FM5cRhGGIAQ2uJCVIixLTPstJgkFL3taQ6bfr/epJGoF
> VbYrGRu0nxGTWEqk14q0YBt2uiDLWu6WiF8izG/fnyM39wzS0ZsO31hco3jpBWiq7X5N
> Ehn8ePiR9iYfowHhT3s2PefnrirD0zlJAamVqnbTNQS93PT26dWpm/vc8HVYiMLj+Fq8
> s2rw==
> X-Gm-Message-State: AA6/9RlGCiscMzjRlXRLSGCPLACOp/VdD9I/y/dQ+vytyQN0tniPrwPxFp4VQtNbW/PYF1zzfyAX+iUOa+dgEsrg
> X-Received: by 10.202.84.69 with SMTP id i66mr3504473oib.93.1476208279931;
> Tue, 11 Oct 2016 10:51:19 -0700 (PDT)
> MIME-Version: 1.0
> Received: by 10.157.39.201 with HTTP; Tue, 11 Oct 2016 10:51:19 -0700 (PDT)
> In-Reply-To: <20161011165811.GO19349@localhost.localdomain>
> References: <20161010003523.4423-1-haozhong.zhang@intel.com>
> <dde78bbd-4739-98a1-4b69-2c2dff0a9d71@citrix.com> <57FCF26A02000078000F15E0@prv-mh.provo.novell.com>
> <CAPcyv4gD3JTq93ET0SAYOxyPO9c0RTkPKSQKqZLo_8KPn53TiA@mail.gmail.com> <20161011165811.GO19349@localhost.localdomain>
> From: Dan Williams <dan.j.williams@intel.com>
> Date: Tue, 11 Oct 2016 10:51:19 -0700
> Message-ID: <CAPcyv4jX_xzj=gz=tfoNMx0qFtyeKwqttzCE5GrOi6Kz5anhiQ@mail.gmail.com>
> Subject: Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen
> To: Konrad Rzeszutek Wilk <konrad.wilk@oracle.com>
> Cc: Jan Beulich <jbeulich@suse.com>, Juergen Gross <JGross@suse.com>,
> Haozhong Zhang <haozhong.zhang@intel.com>,
> Xiao Guangrong <guangrong.xiao@linux.intel.com>,
> Arnd Bergmann <arnd@arndb.de>,
> "linux-nvdimm@lists.01.org" <linux-nvdimm@ml01.01.org>,
> Boris Ostrovsky <boris.ostrovsky@oracle.com>,
> andrew.cooper3@citrix.com,
> "linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
> Stefano Stabellini <stefano@aporeto.com>,
> David Vrabel <david.vrabel@citrix.com>,
> Johannes Thumshirn <jthumshirn@suse.de>,
> xen-devel@lists.xenproject.org,
> Andrew Morton <akpm@linux-foundation.org>,
> Ross Zwisler <ross.zwisler@linux.intel.com>
> Content-Type: text/plain; charset=UTF-8
> X-Source-IP: 209.85.218.43
> X-ServerName: mail-oi0-f43.google.com
> X-Proofpoint-SPF-Result: pass
> X-Proofpoint-SPF-Record: v=spf1 mx:intel.com include:_spf.google.com -all
> X-Proofpoint-Virus-Version: vendor=nai engine=5800 definitions=8315 signatures=670727
> X-Proofpoint-Spam-Details: rule=notspam policy=default score=0 spamscore=0 suspectscore=1
> malwarescore=0 phishscore=0 adultscore=0 bulkscore=0 classifier=spam
> adjust=0 reason=mlx scancount=1 engine=8.0.1-1609300000
> definitions=main-1610110304
> X-Spam: Clean
>
> On Tue, Oct 11, 2016 at 9:58 AM, Konrad Rzeszutek Wilk
> <konrad.wilk@oracle.com> wrote:
> > On Tue, Oct 11, 2016 at 08:53:33AM -0700, Dan Williams wrote:
> >> On Tue, Oct 11, 2016 at 6:08 AM, Jan Beulich <jbeulich@suse.com> wrote:
> >> >>>> Andrew Cooper <andrew.cooper3@citrix.com> 10/10/16 6:44 PM >>>
> >> >>On 10/10/16 01:35, Haozhong Zhang wrote:
> >> >>> Xen hypervisor needs assistance from Dom0 Linux kernel for following tasks:
> >> >>> 1) Reserve an area on NVDIMM devices for Xen hypervisor to place
> >> >>> memory management data structures, i.e. frame table and M2P table.
> >> >>> 2) Report SPA ranges of NVDIMM devices and the reserved area to Xen
> >> >>> hypervisor.
> >> >>
> >> >>However, I can't see any justification for 1). Dom0 should not be
> >> >>involved in Xen's management of its own frame table and m2p. The mfns
> >> >>making up the pmem/pblk regions should be treated just like any other
> >> >>MMIO regions, and be handed wholesale to dom0 by default.
> >> >
> >> > That precludes the use as RAM extension, and I thought earlier rounds of
> >> > discussion had got everyone in agreement that at least for the pmem case
> >> > we will need some control data in Xen.
> >>
> >> The missing piece for me is why this reservation for control data
> >> needs to be done in the libnvdimm core? I would expect that any dax
> >
> > Isn't it done this way with Linux? That is say if the machine has
> > 4GB of RAM and the NVDIMM is in TB range. You want to put the 'struct page'
> > for the NVDIMM ranges somewhere. That place can be in regions on the
> > NVDIMM that ndctl can reserve.
>
> Yes.
>
> >> capable file could be mapped and made available to a guest. This
> >> includes /dev/ramX devices that are dax capable, but are external to
> >> the libnvdimm sub-system.
> >
> > This is more of just keeping track of the ranges if say the DAX file is
> > extremely fragmented and requires a lot of 'struct pages' to keep track of
> > when stiching up the VMA.
>
> Right, but why does the libnvdimm core need to know about this
> specific Xen reservation? For example, if Xen wants some in-kernel
Let me turn this around - why does the libnvdimm core need to know about
Linux specific parts? Shouldn't this be OS agnostic, so that FreeBSD
for example can also poke a hole in this and fill it with its
OS-management meta-data?
> driver to own a pmem region and place its own metadata on the device I
> would recommend something like:
>
> bdev = blkdev_get_by_path("/dev/pmemX", FMODE_EXCL...);
> bdev_direct_access(bdev, ...);
>
> ...in other words, I don't think we want libnvdimm to grow new device
> types for every possible in-kernel user, Xen, MD, DM, etc. Instead,
> just claim the resulting device.
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-10-11 21:30 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <sraUp-2hG-1@gated-at.bofh.it> |
| In reply to | #1499136 |
On Tue, Oct 11, 2016 at 11:33 AM, Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> wrote: > On Tue, Oct 11, 2016 at 10:51:19AM -0700, Dan Williams wrote: [..] >> Right, but why does the libnvdimm core need to know about this >> specific Xen reservation? For example, if Xen wants some in-kernel > > Let me turn this around - why does the libnvdimm core need to know about > Linux specific parts? Shouldn't this be OS agnostic, so that FreeBSD > for example can also poke a hole in this and fill it with its > OS-management meta-data? Specifically the core needs to know so that it can answer the Linux specific question of whether the pfn returned by ->direct_access() has a corresponding struct page or not. It's tied to the lifetime of the device and the usage of the reservation needs to be coordinated against the references of those pages. If FreeBSD decides it needs to reserve "struct page" capacity at the start of the device, I would hope that it reuses the same on-device info block that Linux is using and not create a new "FreeBSD-mode" device type. To be honest I do not yet understand what metadata Xen wants to store in the device, but it seems the producer and consumer of that metadata is Xen itself and not the wider Linux kernel as is the case with struct page. Can you fill me in on what problem Xen solves with this reservation?
[toc] | [prev] | [next] | [standalone]
| From | Konrad Rzeszutek Wilk <konrad.wilk@oracle.com> |
|---|---|
| Date | 2016-10-11 22:00 +0200 |
| Subject | Re: [Xen-devel] [RFC KERNEL PATCH 0/2] Add Dom0 NVDIMM support for Xen |
| Message-ID | <srbnr-2rH-23@gated-at.bofh.it> |
| In reply to | #1499163 |
On Tue, Oct 11, 2016 at 12:28:56PM -0700, Dan Williams wrote: > On Tue, Oct 11, 2016 at 11:33 AM, Konrad Rzeszutek Wilk > <konrad.wilk@oracle.com> wrote: > > On Tue, Oct 11, 2016 at 10:51:19AM -0700, Dan Williams wrote: > [..] > >> Right, but why does the libnvdimm core need to know about this > >> specific Xen reservation? For example, if Xen wants some in-kernel > > > > Let me turn this around - why does the libnvdimm core need to know about > > Linux specific parts? Shouldn't this be OS agnostic, so that FreeBSD > > for example can also poke a hole in this and fill it with its > > OS-management meta-data? > > Specifically the core needs to know so that it can answer the Linux > specific question of whether the pfn returned by ->direct_access() has > a corresponding struct page or not. It's tied to the lifetime of the > device and the usage of the reservation needs to be coordinated > against the references of those pages. If FreeBSD decides it needs to > reserve "struct page" capacity at the start of the device, I would > hope that it reuses the same on-device info block that Linux is using > and not create a new "FreeBSD-mode" device type. The issue here (as I understand, I may be missing something new) is that the size of this special namespace may be different. That is the 'struct page' on FreeBSD could be 256 bytes while on Linux it is 64 bytes (numbers pulled out of the sky). Hence one would have to expand or such to re-use this. > > To be honest I do not yet understand what metadata Xen wants to store > in the device, but it seems the producer and consumer of that metadata > is Xen itself and not the wider Linux kernel as is the case with > struct page. Can you fill me in on what problem Xen solves with this Exactly! > reservation? The same as Linux - its variant of 'struct page'. Which I think is smaller than the Linux one, but perhaps it is not?
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web