Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1171983 > unrolled thread

[PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices

Started byDan Williams <dan.j.williams@intel.com>
First post2015-06-25 11:50 +0200
Last post2015-06-26 03:30 +0200
Articles 8 — 2 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  [PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices Dan Williams <dan.j.williams@intel.com> - 2015-06-25 11:50 +0200
    Re: [PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices Dan Williams <dan.j.williams@intel.com> - 2015-06-25 23:40 +0200
      Re: [PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices Toshi Kani <toshi.kani@hp.com> - 2015-06-26 00:00 +0200
        Re: [PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices Dan Williams <dan.j.williams@intel.com> - 2015-06-26 00:10 +0200
          Re: [PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices Toshi Kani <toshi.kani@hp.com> - 2015-06-26 00:20 +0200
            Re: [PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices Dan Williams <dan.j.williams@intel.com> - 2015-06-26 00:40 +0200
              Re: [PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices Toshi Kani <toshi.kani@hp.com> - 2015-06-26 01:00 +0200
                Re: [PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices Toshi Kani <toshi.kani@hp.com> - 2015-06-26 03:30 +0200

#1171983 — [PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices

FromDan Williams <dan.j.williams@intel.com>
Date2015-06-25 11:50 +0200
Subject[PATCH v2 15/17] libnvdimm: Set numa_node to NVDIMM devices
Message-ID<pFbXc-4YO-13@gated-at.bofh.it>
From: Toshi Kani <toshi.kani@hp.com>

ACPI NFIT table has System Physical Address Range Structure entries that
describe a proximity ID of each range when ACPI_NFIT_PROXIMITY_VALID is
set in the flags.

Change acpi_nfit_register_region() to map a proximity ID to its node ID,
and set it to a new numa_node field of nd_region_desc, which is then
conveyed to the nd_region device.

The device core arranges for btt and namespace devices to inherit their
node from their parent region.

Signed-off-by: Toshi Kani <toshi.kani@hp.com>
[djbw: move set_dev_node() from region 'probe' to 'create']
Signed-off-by: Dan Williams <dan.j.williams@intel.com>
---
 drivers/acpi/nfit.c          |    6 ++++++
 drivers/nvdimm/region_devs.c |    4 +++-
 include/linux/libnvdimm.h    |    1 +
 3 files changed, 10 insertions(+), 1 deletion(-)

diff --git a/drivers/acpi/nfit.c b/drivers/acpi/nfit.c
index 1f6f1b1a54f4..d96c8fe974dd 100644
--- a/drivers/acpi/nfit.c
+++ b/drivers/acpi/nfit.c
@@ -1392,6 +1392,12 @@ static int acpi_nfit_register_region(struct acpi_nfit_desc *acpi_desc,
 	ndr_desc->res = &res;
 	ndr_desc->provider_data = nfit_spa;
 	ndr_desc->attr_groups = acpi_nfit_region_attribute_groups;
+	if (spa->flags & ACPI_NFIT_PROXIMITY_VALID)
+		ndr_desc->numa_node = acpi_map_pxm_to_online_node(
+						spa->proximity_domain);
+	else
+		ndr_desc->numa_node = NUMA_NO_NODE;
+
 	list_for_each_entry(nfit_memdev, &acpi_desc->memdevs, list) {
 		struct acpi_nfit_memory_map *memdev = nfit_memdev->memdev;
 		struct nd_mapping *nd_mapping;
diff --git a/drivers/nvdimm/region_devs.c b/drivers/nvdimm/region_devs.c
index 8f8c7ea485f1..b2ec5045f5a5 100644
--- a/drivers/nvdimm/region_devs.c
+++ b/drivers/nvdimm/region_devs.c
@@ -738,13 +738,15 @@ static struct nd_region *nd_region_create(struct nvdimm_bus *nvdimm_bus,
 	nd_region->ro = ro;
 	ida_init(&nd_region->ns_ida);
 	dev = &nd_region->dev;
+	device_initialize(dev);
 	dev_set_name(dev, "region%d", nd_region->id);
 	dev->parent = &nvdimm_bus->dev;
 	dev->type = dev_type;
 	dev->groups = ndr_desc->attr_groups;
+	set_dev_node(dev, ndr_desc->numa_node);
 	nd_region->ndr_size = resource_size(ndr_desc->res);
 	nd_region->ndr_start = ndr_desc->res->start;
-	nd_device_register(dev);
+	__nd_device_register(dev);
 
 	return nd_region;
 
diff --git a/include/linux/libnvdimm.h b/include/linux/libnvdimm.h
index dc799a29ed1a..30b3deaafd51 100644
--- a/include/linux/libnvdimm.h
+++ b/include/linux/libnvdimm.h
@@ -89,6 +89,7 @@ struct nd_region_desc {
 	struct nd_interleave_set *nd_set;
 	void *provider_data;
 	int num_lanes;
+	int numa_node;
 };
 
 struct nvdimm_bus;

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1172462

FromDan Williams <dan.j.williams@intel.com>
Date2015-06-25 23:40 +0200
Message-ID<pFn2i-42w-31@gated-at.bofh.it>
In reply to#1171983
On Thu, Jun 25, 2015 at 11:34 AM, Williams, Dan J
<dan.j.williams@intel.com> wrote:
> On Thu, 2015-06-25 at 11:45 -0600, Toshi Kani wrote:
>> On Thu, 2015-06-25 at 05:37 -0400, Dan Williams wrote:
>> > From: Toshi Kani <toshi.kani@hp.com>
>> >
>> > ACPI NFIT table has System Physical Address Range Structure entries that
>> > describe a proximity ID of each range when ACPI_NFIT_PROXIMITY_VALID is
>> > set in the flags.
>> >
>> > Change acpi_nfit_register_region() to map a proximity ID to its node ID,
>> > and set it to a new numa_node field of nd_region_desc, which is then
>> > conveyed to the nd_region device.
>> >
>> > The device core arranges for btt and namespace devices to inherit their
>> > node from their parent region.
>> >
>> > Signed-off-by: Toshi Kani <toshi.kani@hp.com>
>> > [djbw: move set_dev_node() from region 'probe' to 'create']
>>
>> Sorry, I failed to mention other issue, which led me call set_dev_node()
>> in probe.  nd_async_device_register() calls device_add(), which does:
>>
>>         /* use parent numa_node */
>>         if (parent)
>>                 set_dev_node(dev, dev_to_node(parent));
>>
>> and overwrites numa_node to -1.  Since region's parent is ndbusN, we
>> cannot set numa_node to the parent.  So, I had to set it in probe.
>
> In general, I still don't like leaving it up to ->probe() which is
> within its rights to fail and not set the node.  How about the following
> that moves it to the bus uevent code?  Should get triggered before probe
> so the numa_node is valid before userspace is ever notified about the
> device.
>
> device_add() does:
>
>         kobject_uevent(&dev->kobj, KOBJ_ADD);
>         bus_probe_device(dev);
>
> ...so I think we're good, agree?  I also added a missing init of
> ndr_desc.numa_node in arch/x86/kernel/pmem.c, see below.

This looks good in a quick manual test.  It's interesting/illustrative
that I inadvertently broke the one bit of the libnvdimm sysfs
interface that did not have unit test coverage.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1172485

FromToshi Kani <toshi.kani@hp.com>
Date2015-06-26 00:00 +0200
Message-ID<pFnlD-4p2-7@gated-at.bofh.it>
In reply to#1172462
On Thu, 2015-06-25 at 14:31 -0700, Dan Williams wrote:
> On Thu, Jun 25, 2015 at 11:34 AM, Williams, Dan J
> <dan.j.williams@intel.com> wrote:
> > On Thu, 2015-06-25 at 11:45 -0600, Toshi Kani wrote:
> >> On Thu, 2015-06-25 at 05:37 -0400, Dan Williams wrote:
> >> > From: Toshi Kani <toshi.kani@hp.com>
> >> >
> >> > ACPI NFIT table has System Physical Address Range Structure entries that
> >> > describe a proximity ID of each range when ACPI_NFIT_PROXIMITY_VALID is
> >> > set in the flags.
> >> >
> >> > Change acpi_nfit_register_region() to map a proximity ID to its node ID,
> >> > and set it to a new numa_node field of nd_region_desc, which is then
> >> > conveyed to the nd_region device.
> >> >
> >> > The device core arranges for btt and namespace devices to inherit their
> >> > node from their parent region.
> >> >
> >> > Signed-off-by: Toshi Kani <toshi.kani@hp.com>
> >> > [djbw: move set_dev_node() from region 'probe' to 'create']
> >>
> >> Sorry, I failed to mention other issue, which led me call set_dev_node()
> >> in probe.  nd_async_device_register() calls device_add(), which does:
> >>
> >>         /* use parent numa_node */
> >>         if (parent)
> >>                 set_dev_node(dev, dev_to_node(parent));
> >>
> >> and overwrites numa_node to -1.  Since region's parent is ndbusN, we
> >> cannot set numa_node to the parent.  So, I had to set it in probe.
> >
> > In general, I still don't like leaving it up to ->probe() which is
> > within its rights to fail and not set the node.  How about the following
> > that moves it to the bus uevent code?  Should get triggered before probe
> > so the numa_node is valid before userspace is ever notified about the
> > device.
> >
> > device_add() does:
> >
> >         kobject_uevent(&dev->kobj, KOBJ_ADD);
> >         bus_probe_device(dev);
> >
> > ...so I think we're good, agree?  I also added a missing init of
> > ndr_desc.numa_node in arch/x86/kernel/pmem.c, see below.
> 
> This looks good in a quick manual test.  It's interesting/illustrative
> that I inadvertently broke the one bit of the libnvdimm sysfs
> interface that did not have unit test coverage.

Sorry I had some interrupt.  Yes, this works fine for region &
namespace.  I'd like to check with you for btt since the attach logic
has changed in v2.

Previously, as described in patch 16/17, bttN bound to pmem had a valid
numa_node value, and seeding btt0 had -1.

  /sys/bus/nd/devices
  |-- btt0/numa_node:-1
  |-- btt1/numa_node:0

In this version, there are unbound (seeding?) btt0-3 for every region
(there are 4 regions) and btt4 & 5 bound to pmem0 & 3 on my system.

btt0/numa_node:0
btt1/numa_node:0
btt2/numa_node:1
btt3/numa_node:1
btt4/numa_node:0
btt5/numa_node:1

btt0
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt0
btt1
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region1/btt1
btt2
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region2/btt2
btt3
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt3
btt4
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt4
btt5
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt5

And unbound bttNs attach to different regions across a reboot.

btt0/numa_node:0
btt1/numa_node:1
btt2/numa_node:1
btt3/numa_node:0
btt4/numa_node:0
btt5/numa_node:1

btt0
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt0
btt1
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt1
btt2
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region2/btt2
btt3
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region1/btt3
btt4
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt4
btt5
-> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt5

Is this how you'd expect btt to work in this version?  (I have not
looked at the btt changes yet)

Thanks,
-Toshi

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1172490

FromDan Williams <dan.j.williams@intel.com>
Date2015-06-26 00:10 +0200
Message-ID<pFnvk-4Pp-17@gated-at.bofh.it>
In reply to#1172485
On Thu, Jun 25, 2015 at 2:51 PM, Toshi Kani <toshi.kani@hp.com> wrote:
> On Thu, 2015-06-25 at 14:31 -0700, Dan Williams wrote:
>> On Thu, Jun 25, 2015 at 11:34 AM, Williams, Dan J
>> <dan.j.williams@intel.com> wrote:
>> > On Thu, 2015-06-25 at 11:45 -0600, Toshi Kani wrote:
>> >> On Thu, 2015-06-25 at 05:37 -0400, Dan Williams wrote:
>> >> > From: Toshi Kani <toshi.kani@hp.com>
>> >> >
>> >> > ACPI NFIT table has System Physical Address Range Structure entries that
>> >> > describe a proximity ID of each range when ACPI_NFIT_PROXIMITY_VALID is
>> >> > set in the flags.
>> >> >
>> >> > Change acpi_nfit_register_region() to map a proximity ID to its node ID,
>> >> > and set it to a new numa_node field of nd_region_desc, which is then
>> >> > conveyed to the nd_region device.
>> >> >
>> >> > The device core arranges for btt and namespace devices to inherit their
>> >> > node from their parent region.
>> >> >
>> >> > Signed-off-by: Toshi Kani <toshi.kani@hp.com>
>> >> > [djbw: move set_dev_node() from region 'probe' to 'create']
>> >>
>> >> Sorry, I failed to mention other issue, which led me call set_dev_node()
>> >> in probe.  nd_async_device_register() calls device_add(), which does:
>> >>
>> >>         /* use parent numa_node */
>> >>         if (parent)
>> >>                 set_dev_node(dev, dev_to_node(parent));
>> >>
>> >> and overwrites numa_node to -1.  Since region's parent is ndbusN, we
>> >> cannot set numa_node to the parent.  So, I had to set it in probe.
>> >
>> > In general, I still don't like leaving it up to ->probe() which is
>> > within its rights to fail and not set the node.  How about the following
>> > that moves it to the bus uevent code?  Should get triggered before probe
>> > so the numa_node is valid before userspace is ever notified about the
>> > device.
>> >
>> > device_add() does:
>> >
>> >         kobject_uevent(&dev->kobj, KOBJ_ADD);
>> >         bus_probe_device(dev);
>> >
>> > ...so I think we're good, agree?  I also added a missing init of
>> > ndr_desc.numa_node in arch/x86/kernel/pmem.c, see below.
>>
>> This looks good in a quick manual test.  It's interesting/illustrative
>> that I inadvertently broke the one bit of the libnvdimm sysfs
>> interface that did not have unit test coverage.
>
> Sorry I had some interrupt.  Yes, this works fine for region &
> namespace.  I'd like to check with you for btt since the attach logic
> has changed in v2.
>
> Previously, as described in patch 16/17, bttN bound to pmem had a valid
> numa_node value, and seeding btt0 had -1.
>
>   /sys/bus/nd/devices
>   |-- btt0/numa_node:-1
>   |-- btt1/numa_node:0
>
> In this version, there are unbound (seeding?) btt0-3 for every region
> (there are 4 regions) and btt4 & 5 bound to pmem0 & 3 on my system.
>
> btt0/numa_node:0
> btt1/numa_node:0
> btt2/numa_node:1
> btt3/numa_node:1
> btt4/numa_node:0
> btt5/numa_node:1
>
> btt0
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt0
> btt1
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region1/btt1
> btt2
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region2/btt2
> btt3
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt3
> btt4
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt4
> btt5
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt5
>
> And unbound bttNs attach to different regions across a reboot.
>
> btt0/numa_node:0
> btt1/numa_node:1
> btt2/numa_node:1
> btt3/numa_node:0
> btt4/numa_node:0
> btt5/numa_node:1
>
> btt0
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt0
> btt1
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt1
> btt2
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region2/btt2
> btt3
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region1/btt3
> btt4
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt4
> btt5
> -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt5
>
> Is this how you'd expect btt to work in this version?  (I have not
> looked at the btt changes yet)

Yes, this looks fine.

As requested by Christoph, in the latest version BTTs are child
devices of regions rather than busses.  They automatically inherit the
numa_node of the parent region.  In your dump above the numa_nodes are
not changing from boot-to-boot, instead the BTTs are registered
asynchronously so get different ids from boot-to-boot.  Userspace
should not care what the btt id is and the same naming trick we use to
give block devices static names would not work for BTTs.  The child
block device of the BTT will still have the static name as we
discussed earlier (/dev/pmemXs or /dev/ndblkX.Ys) because the scan
order of those is deterministic.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1172493

FromToshi Kani <toshi.kani@hp.com>
Date2015-06-26 00:20 +0200
Message-ID<pFnEZ-50r-3@gated-at.bofh.it>
In reply to#1172490
On Thu, 2015-06-25 at 15:00 -0700, Dan Williams wrote:
> On Thu, Jun 25, 2015 at 2:51 PM, Toshi Kani <toshi.kani@hp.com> wrote:
> > On Thu, 2015-06-25 at 14:31 -0700, Dan Williams wrote:
> >> On Thu, Jun 25, 2015 at 11:34 AM, Williams, Dan J
> >> <dan.j.williams@intel.com> wrote:
> >> > On Thu, 2015-06-25 at 11:45 -0600, Toshi Kani wrote:
> >> >> On Thu, 2015-06-25 at 05:37 -0400, Dan Williams wrote:
> >> >> > From: Toshi Kani <toshi.kani@hp.com>
> >> >> >
> >> >> > ACPI NFIT table has System Physical Address Range Structure entries that
> >> >> > describe a proximity ID of each range when ACPI_NFIT_PROXIMITY_VALID is
> >> >> > set in the flags.
> >> >> >
> >> >> > Change acpi_nfit_register_region() to map a proximity ID to its node ID,
> >> >> > and set it to a new numa_node field of nd_region_desc, which is then
> >> >> > conveyed to the nd_region device.
> >> >> >
> >> >> > The device core arranges for btt and namespace devices to inherit their
> >> >> > node from their parent region.
> >> >> >
> >> >> > Signed-off-by: Toshi Kani <toshi.kani@hp.com>
> >> >> > [djbw: move set_dev_node() from region 'probe' to 'create']
> >> >>
> >> >> Sorry, I failed to mention other issue, which led me call set_dev_node()
> >> >> in probe.  nd_async_device_register() calls device_add(), which does:
> >> >>
> >> >>         /* use parent numa_node */
> >> >>         if (parent)
> >> >>                 set_dev_node(dev, dev_to_node(parent));
> >> >>
> >> >> and overwrites numa_node to -1.  Since region's parent is ndbusN, we
> >> >> cannot set numa_node to the parent.  So, I had to set it in probe.
> >> >
> >> > In general, I still don't like leaving it up to ->probe() which is
> >> > within its rights to fail and not set the node.  How about the following
> >> > that moves it to the bus uevent code?  Should get triggered before probe
> >> > so the numa_node is valid before userspace is ever notified about the
> >> > device.
> >> >
> >> > device_add() does:
> >> >
> >> >         kobject_uevent(&dev->kobj, KOBJ_ADD);
> >> >         bus_probe_device(dev);
> >> >
> >> > ...so I think we're good, agree?  I also added a missing init of
> >> > ndr_desc.numa_node in arch/x86/kernel/pmem.c, see below.
> >>
> >> This looks good in a quick manual test.  It's interesting/illustrative
> >> that I inadvertently broke the one bit of the libnvdimm sysfs
> >> interface that did not have unit test coverage.
> >
> > Sorry I had some interrupt.  Yes, this works fine for region &
> > namespace.  I'd like to check with you for btt since the attach logic
> > has changed in v2.
> >
> > Previously, as described in patch 16/17, bttN bound to pmem had a valid
> > numa_node value, and seeding btt0 had -1.
> >
> >   /sys/bus/nd/devices
> >   |-- btt0/numa_node:-1
> >   |-- btt1/numa_node:0
> >
> > In this version, there are unbound (seeding?) btt0-3 for every region
> > (there are 4 regions) and btt4 & 5 bound to pmem0 & 3 on my system.
> >
> > btt0/numa_node:0
> > btt1/numa_node:0
> > btt2/numa_node:1
> > btt3/numa_node:1
> > btt4/numa_node:0
> > btt5/numa_node:1
> >
> > btt0
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt0
> > btt1
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region1/btt1
> > btt2
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region2/btt2
> > btt3
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt3
> > btt4
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt4
> > btt5
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt5
> >
> > And unbound bttNs attach to different regions across a reboot.
> >
> > btt0/numa_node:0
> > btt1/numa_node:1
> > btt2/numa_node:1
> > btt3/numa_node:0
> > btt4/numa_node:0
> > btt5/numa_node:1
> >
> > btt0
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt0
> > btt1
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt1
> > btt2
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region2/btt2
> > btt3
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region1/btt3
> > btt4
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region0/btt4
> > btt5
> > -> ../../../devices/LNXSYSTM:00/LNXSYBUS:00/ACPI0012:00/ndbus0/region3/btt5
> >
> > Is this how you'd expect btt to work in this version?  (I have not
> > looked at the btt changes yet)
> 
> Yes, this looks fine.
> 
> As requested by Christoph, in the latest version BTTs are child
> devices of regions rather than busses.  They automatically inherit the
> numa_node of the parent region.  In your dump above the numa_nodes are
> not changing from boot-to-boot, instead the BTTs are registered
> asynchronously so get different ids from boot-to-boot.  Userspace
> should not care what the btt id is and the same naming trick we use to
> give block devices static names would not work for BTTs.  The child
> block device of the BTT will still have the static name as we
> discussed earlier (/dev/pmemXs or /dev/ndblkX.Ys) because the scan
> order of those is deterministic.

Yes, I see no problem with bound BTTs and their device files.  So, how
do we bind BTT with this new version?

Thanks,
-Toshi  



--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1172501

FromDan Williams <dan.j.williams@intel.com>
Date2015-06-26 00:40 +0200
Message-ID<pFnYm-5mM-19@gated-at.bofh.it>
In reply to#1172493
On Thu, Jun 25, 2015 at 3:11 PM, Toshi Kani <toshi.kani@hp.com> wrote:
> On Thu, 2015-06-25 at 15:00 -0700, Dan Williams wrote:
> Yes, I see no problem with bound BTTs and their device files.  So, how
> do we bind BTT with this new version?
>

# cd /sys/bus/nd/devices
# uuidgen > btt6/uuid
# echo 4096 > btt6/sector_size
# echo namespace6.0 > btt6/namespace
# echo namespace6.0 > ../drivers/nd_pmem/unbind
# echo btt6 > ../drivers/nd_pmem/bind

After reboot, when the system sees namespace6.0 again it will notice
the btt instance and attach bttX instead.  The net effect is that now
you'll only ever have /dev/pmem6 or /dev/pmem6s, never both at the
same time that was a side effect of the stacking approach.

I'll post the patch that updates libndctl and the unit tests shortly
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1172507

FromToshi Kani <toshi.kani@hp.com>
Date2015-06-26 01:00 +0200
Message-ID<pFohI-5Jn-13@gated-at.bofh.it>
In reply to#1172501
On Thu, 2015-06-25 at 15:34 -0700, Dan Williams wrote:
> On Thu, Jun 25, 2015 at 3:11 PM, Toshi Kani <toshi.kani@hp.com> wrote:
> > On Thu, 2015-06-25 at 15:00 -0700, Dan Williams wrote:
> > Yes, I see no problem with bound BTTs and their device files.  So, how
> > do we bind BTT with this new version?
> >
> 
> # cd /sys/bus/nd/devices
> # uuidgen > btt6/uuid
> # echo 4096 > btt6/sector_size
> # echo namespace6.0 > btt6/namespace
> # echo namespace6.0 > ../drivers/nd_pmem/unbind
> # echo btt6 > ../drivers/nd_pmem/bind
> 
> After reboot, when the system sees namespace6.0 again it will notice
> the btt instance and attach bttX instead.  The net effect is that now
> you'll only ever have /dev/pmem6 or /dev/pmem6s, never both at the
> same time that was a side effect of the stacking approach.
> 
> I'll post the patch that updates libndctl and the unit tests shortly

Maybe I am missing something, but I am getting errors on my system.  (I
used btt0 since there is no btt6.)

# cat bind.sh
set -x
cd /sys/bus/nd/devices
uuidgen > btt0/uuid
echo 4096 > btt0/sector_size
echo namespace0.0 > btt0/namespace
echo namespace0.0 > ../drivers/nd_pmem/unbind
echo btt0 > ../drivers/nd_pmem/bind

# sh bind.sh
+ cd /sys/bus/nd/devices
+ uuidgen
+ echo 4096
+ echo namespace0.0
bind.sh: line 6: echo: write error: Device or resource busy
+ echo namespace0.0
bind.sh: line 7: echo: write error: No such device
+ echo btt0
bind.sh: line 8: echo: write error: No such device

# dmesg
 :
[12513.839162] nd btt0: uuid_store: result: 0 wrote:
b32cd195-9aae-4c54-a5ac-49adb50a8a98
[12513.880286] nd btt0: sector_size_store: result: 0 wrote: 4096
[12513.909494] nd btt0: namespace0.0 already claimed
[12513.933364] nd btt0: namespace_store: result: -16 wrote: namespace0.0
[12513.966808]  ndbus0: nd_pmem.probe(btt0) = -19

Thanks,
-Toshi

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1172540

FromToshi Kani <toshi.kani@hp.com>
Date2015-06-26 03:30 +0200
Message-ID<pFqCT-JB-5@gated-at.bofh.it>
In reply to#1172507
On Thu, 2015-06-25 at 18:08 -0700, Dan Williams wrote:
> On Thu, Jun 25, 2015 at 5:55 PM, Toshi Kani <toshi.kani@hp.com> wrote:
> > On Thu, 2015-06-25 at 23:42 +0000, Williams, Dan J wrote:
> >> On Thu, 2015-06-25 at 16:55 -0600, Toshi Kani wrote:
 :
> >
> > Now, how do I unbind BTT?  I did the following as a guess, but BTT got
> > reattached again before I have a chance to delete the metadata, which I
> > need /dev/pmemN.
> >
> > NUM=1
> > cd /sys/bus/nd/devices
> > echo ""  > btt${NUM}.1/namespace
> > echo btt${NUM}.1 > ../drivers/nd_pmem/unbind
> 
> echo 1 > namespace${NUM}.0/force_raw
> 
> > echo namespace${NUM}.0 > ../drivers/nd_pmem/bind
> >

Cool!  Yes, it works with the force_raw.  (I changed the ordering a
bit).

NUM=1
cd /sys/bus/nd/devices
echo btt${NUM}.1 > ../drivers/nd_pmem/unbind
echo "" > btt${NUM}.1/namespace
echo 1 > namespace${NUM}.0/force_raw
echo namespace${NUM}.0 > ../drivers/nd_pmem/bind

Thanks,
-Toshi


--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web