Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1421388 > unrolled thread

[PATCH 0/6] Support DAX for device-mapper dm-linear devices

Started byToshi Kani <toshi.kani@hpe.com>
First post2016-06-14 00:40 +0200
Last post2016-06-14 18:00 +0200
Articles 20 on this page of 40 — 5 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH 0/6] Support DAX for device-mapper dm-linear devices Toshi Kani <toshi.kani@hpe.com> - 2016-06-14 00:40 +0200
    [PATCH 1/6] genhd: Add GENHD_FL_DAX to gendisk flags Toshi Kani <toshi.kani@hpe.com> - 2016-06-14 00:40 +0200
    [PATCH 4/6] dm-linear: Add linear_direct_access() Toshi Kani <toshi.kani@hpe.com> - 2016-06-14 00:40 +0200
    [PATCH 6/6] dm: Enable DAX support for mapper device Toshi Kani <toshi.kani@hpe.com> - 2016-06-14 00:40 +0200
    [PATCH 3/6] dm: Add dm_blk_direct_access() for mapped device Toshi Kani <toshi.kani@hpe.com> - 2016-06-14 00:40 +0200
    [PATCH 2/6] block: Check GENHD_FL_DAX for DAX capability Toshi Kani <toshi.kani@hpe.com> - 2016-06-14 00:40 +0200
    [PATCH 5/6] dm, dm-linear: Add dax_supported to dm_target Toshi Kani <toshi.kani@hpe.com> - 2016-06-14 00:40 +0200
    Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-14 01:00 +0200
      Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-20 20:30 +0200
        Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-20 21:50 +0200
          Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-20 23:50 +0200
            Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-21 00:10 +0200
              Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-21 00:40 +0200
                Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-21 15:50 +0200
                  Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-21 18:00 +0200
                  Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-21 18:10 +0200
                    Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Dan Williams <dan.j.williams@intel.com> - 2016-06-21 18:30 +0200
                      Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Dan Williams <dan.j.williams@intel.com> - 2016-06-21 18:50 +0200
                        Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-21 19:20 +0200
                      Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-21 19:00 +0200
                    Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-21 20:30 +0200
                      Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-22 19:50 +0200
                        Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Dan Williams <dan.j.williams@intel.com> - 2016-06-22 21:20 +0200
                          Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-22 22:20 +0200
                            Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-23 00:40 +0200
                              Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-23 01:20 +0200
            Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-21 00:50 +0200
        Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-20 23:10 +0200
    Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Dan Williams <dan.j.williams@intel.com> - 2016-06-14 01:20 +0200
      Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-14 02:00 +0200
        Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Dan Williams <dan.j.williams@intel.com> - 2016-06-14 02:10 +0200
          Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Dan Williams <dan.j.williams@intel.com> - 2016-06-14 09:40 +0200
        Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Jeff Moyer <jmoyer@redhat.com> - 2016-06-14 16:00 +0200
          Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-14 17:50 +0200
            Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-14 20:10 +0200
            Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Jeff Moyer <jmoyer@redhat.com> - 2016-06-14 22:20 +0200
              Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-15 03:50 +0200
                Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Dan Williams <dan.j.williams@intel.com> - 2016-06-15 04:10 +0200
                  Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices Mike Snitzer <snitzer@redhat.com> - 2016-06-15 04:40 +0200
          Re: [PATCH 0/6] Support DAX for device-mapper dm-linear devices "Kani, Toshimitsu" <toshi.kani@hpe.com> - 2016-06-14 18:00 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1428050

FromMike Snitzer <snitzer@redhat.com>
Date2016-06-21 20:30 +0200
Message-ID<rMyAW-Xa-15@gated-at.bofh.it>
In reply to#1427909
On Tue, Jun 21 2016 at 11:44am -0400,
Kani, Toshimitsu <toshi.kani@hpe.com> wrote:

> On Tue, 2016-06-21 at 09:41 -0400, Mike Snitzer wrote:
> > On Mon, Jun 20 2016 at  6:22pm -0400,
> > Mike Snitzer <snitzer@redhat.com> wrote:
> > > 
> > > On Mon, Jun 20 2016 at  5:28pm -0400,
> > > Kani, Toshimitsu <toshi.kani@hpe.com> wrote:
> > > 
>  :
> > > Looks good, I folded it in and tested it to work.  Pushed to my 'wip'
> > > branch.
> > > 
> > > No longer seeing any corruption in my test that was using partitions to
> > > span pmem devices with a dm-linear device.
> > > 
> > > Jens, any chance you'd be open to picking up the first 2 patches in this
> > > series?  Or would you like to see them folded or something different?
> >
> > I'm now wondering if we'd be better off setting a new QUEUE_FLAG_DAX
> > rather than establish GENHD_FL_DAX on the genhd?
> > 
> > It'd be quite a bit easier to allow upper layers (e.g. XFS and ext4) to
> > check for a queue flag.
> 
> I think GENHD_FL_DAX is more appropriate since DAX does not use a request
> queue, except for protecting the underlining device being disabled while
> direct_access() is called (b2e0d1625e19).  

The devices in question have a request_queue.  All bio-based device have
a request_queue.

I don't have a big problem with GENHD_FL_DAX.  Just wanted to point out
that such block device capabilities are generally advertised in terms of
a QUEUE_FLAG.
 
> About protecting direct_access, this patch assumes that the underlining
> device cannot be disabled until dtr() is called.  Is this correct?  If not,
> I will need to call dax_map_atomic().

One of the big design considerations for DM that a DM device can be
suspended (with or without flush) and any new IO will be blocked until
the DM device is resumed.

So ideally DM should be able to have the same capability even if using
DAX.

But that is different than what commit b2e0d1625e19 is addressing.  For
DM, I wouldn't think you'd need the extra protections that
dax_map_atomic() is providing given that the underlying block device
lifetime is managed via DM core's dm_get_device/dm_put_device (see also:
dm.c:open_table_device/close_table_device).

[toc] | [prev] | [next] | [standalone]


#1429025

From"Kani, Toshimitsu" <toshi.kani@hpe.com>
Date2016-06-22 19:50 +0200
Message-ID<rMUrL-6vZ-1@gated-at.bofh.it>
In reply to#1428050
On Tue, 2016-06-21 at 14:17 -0400, Mike Snitzer wrote:
> On Tue, Jun 21 2016 at 11:44am -0400,
> Kani, Toshimitsu <toshi.kani@hpe.com> wrote:
> > 
> > On Tue, 2016-06-21 at 09:41 -0400, Mike Snitzer wrote:
> > > 
> > > On Mon, Jun 20 2016 at  6:22pm -0400,
> > > Mike Snitzer <snitzer@redhat.com> wrote:
 :
> > > I'm now wondering if we'd be better off setting a new QUEUE_FLAG_DAX
> > > rather than establish GENHD_FL_DAX on the genhd?
> > > 
> > > It'd be quite a bit easier to allow upper layers (e.g. XFS and ext4) to
> > > check for a queue flag.
> >
> > I think GENHD_FL_DAX is more appropriate since DAX does not use a request
> > queue, except for protecting the underlining device being disabled while
> > direct_access() is called (b2e0d1625e19).  
>
> The devices in question have a request_queue.  All bio-based device have
> a request_queue.

DAX-capable devices have two operation modes, bio-based and DAX.  I agree that
bio-based operation is associated with a request queue, and its capabilities
should be set to it.  DAX, on the other hand, is rather independent from a
request queue.

> I don't have a big problem with GENHD_FL_DAX.  Just wanted to point out
> that such block device capabilities are generally advertised in terms of
> a QUEUE_FLAG.

I do not have a strong opinion, but feel a bit odd to associate DAX to a
request queue. 
 
> > About protecting direct_access, this patch assumes that the underlining
> > device cannot be disabled until dtr() is called.  Is this correct?  If
> > not, I will need to call dax_map_atomic().
>
> One of the big design considerations for DM that a DM device can be
> suspended (with or without flush) and any new IO will be blocked until
> the DM device is resumed.
> 
> So ideally DM should be able to have the same capability even if using
> DAX.

Supporting suspend for DAX is challenging since it allows user applications to
access a device directly.  Once a device range is mmap'd, there is no kernel
intervention to access the range, unless we invalidate user mappings.  This
isn't done today even after a driver is unbind'd from a device.

> But that is different than what commit b2e0d1625e19 is addressing.  For
> DM, I wouldn't think you'd need the extra protections that
> dax_map_atomic() is providing given that the underlying block device
> lifetime is managed via DM core's dm_get_device/dm_put_device (see also:
> dm.c:open_table_device/close_table_device).

I thought so as well.  But I realized that there is (almost) nothing that can
prevent the unbind operation.  It cannot fail, either.  This unbind proceeds
even when a device is in-use.  In case of a pmem device, it is only protected
by pmem_release_queue(), which is called when a pmem device is being deleted
and calls blk_cleanup_queue() to serialize a critical section between
blk_queue_enter() and blk_queue_exit() per b2e0d1625e19.  This prevents from a
kernel DTLB fault, but does not prevent a device disappeared while in-use.

Protecting DM's underlining device with blk_queue_enter() (or something
similar) requires more thoughts...  blk_queue_enter() to a DM device cannot be
redirected to its underlining device.  So, this is TBD for now.  But I do not
think this is a blocker issue since doing unbind to a underlining device is
quite harmful no matter what we do - even if it is protected with
blk_queue_enter().

Thanks,
-Toshi

[toc] | [prev] | [next] | [standalone]


#1429062

FromDan Williams <dan.j.williams@intel.com>
Date2016-06-22 21:20 +0200
Message-ID<rMVQS-7xn-7@gated-at.bofh.it>
In reply to#1429025
On Wed, Jun 22, 2016 at 10:44 AM, Kani, Toshimitsu <toshi.kani@hpe.com> wrote:
> On Tue, 2016-06-21 at 14:17 -0400, Mike Snitzer wrote:
>> On Tue, Jun 21 2016 at 11:44am -0400,
>> Kani, Toshimitsu <toshi.kani@hpe.com> wrote:
>> >
>> > On Tue, 2016-06-21 at 09:41 -0400, Mike Snitzer wrote:
>> > >
>> > > On Mon, Jun 20 2016 at  6:22pm -0400,
>> > > Mike Snitzer <snitzer@redhat.com> wrote:
>  :
>> > > I'm now wondering if we'd be better off setting a new QUEUE_FLAG_DAX
>> > > rather than establish GENHD_FL_DAX on the genhd?
>> > >
>> > > It'd be quite a bit easier to allow upper layers (e.g. XFS and ext4) to
>> > > check for a queue flag.
>> >
>> > I think GENHD_FL_DAX is more appropriate since DAX does not use a request
>> > queue, except for protecting the underlining device being disabled while
>> > direct_access() is called (b2e0d1625e19).
>>
>> The devices in question have a request_queue.  All bio-based device have
>> a request_queue.
>
> DAX-capable devices have two operation modes, bio-based and DAX.  I agree that
> bio-based operation is associated with a request queue, and its capabilities
> should be set to it.  DAX, on the other hand, is rather independent from a
> request queue.
>
>> I don't have a big problem with GENHD_FL_DAX.  Just wanted to point out
>> that such block device capabilities are generally advertised in terms of
>> a QUEUE_FLAG.
>
> I do not have a strong opinion, but feel a bit odd to associate DAX to a
> request queue.

Given that we do not support dax to a raw block device [1] it seems a
gendisk flag is more misleading than request_queue flag that specifies
what requests can be made of the device.

[1]: acc93d30d7d4 Revert "block: enable dax for raw block devices"


>> > About protecting direct_access, this patch assumes that the underlining
>> > device cannot be disabled until dtr() is called.  Is this correct?  If
>> > not, I will need to call dax_map_atomic().
>>
>> One of the big design considerations for DM that a DM device can be
>> suspended (with or without flush) and any new IO will be blocked until
>> the DM device is resumed.
>>
>> So ideally DM should be able to have the same capability even if using
>> DAX.
>
> Supporting suspend for DAX is challenging since it allows user applications to
> access a device directly.  Once a device range is mmap'd, there is no kernel
> intervention to access the range, unless we invalidate user mappings.  This
> isn't done today even after a driver is unbind'd from a device.
>
>> But that is different than what commit b2e0d1625e19 is addressing.  For
>> DM, I wouldn't think you'd need the extra protections that
>> dax_map_atomic() is providing given that the underlying block device
>> lifetime is managed via DM core's dm_get_device/dm_put_device (see also:
>> dm.c:open_table_device/close_table_device).
>
> I thought so as well.  But I realized that there is (almost) nothing that can
> prevent the unbind operation.  It cannot fail, either.  This unbind proceeds
> even when a device is in-use.  In case of a pmem device, it is only protected
> by pmem_release_queue(), which is called when a pmem device is being deleted
> and calls blk_cleanup_queue() to serialize a critical section between
> blk_queue_enter() and blk_queue_exit() per b2e0d1625e19.  This prevents from a
> kernel DTLB fault, but does not prevent a device disappeared while in-use.
>
> Protecting DM's underlining device with blk_queue_enter() (or something
> similar) requires more thoughts...  blk_queue_enter() to a DM device cannot be
> redirected to its underlining device.  So, this is TBD for now.  But I do not
> think this is a blocker issue since doing unbind to a underlining device is
> quite harmful no matter what we do - even if it is protected with
> blk_queue_enter().

I still have the "block device removed" notification patches on my
todo list.  It's not a blocker, but there are scenarios where we can
keep accessing memory via dax of a disabled device leading to memory
corruption.  I'll bump that up in my queue now that we are looking at
additional scenarios where letting DAX mappings leak past the
reconfiguration of a block device could lead to trouble.

[toc] | [prev] | [next] | [standalone]


#1429089

From"Kani, Toshimitsu" <toshi.kani@hpe.com>
Date2016-06-22 22:20 +0200
Message-ID<rMWMW-87F-5@gated-at.bofh.it>
In reply to#1429062
On Wed, 2016-06-22 at 12:15 -0700, Dan Williams wrote:
> On Wed, Jun 22, 2016 at 10:44 AM, Kani, Toshimitsu <toshi.kani@hpe.com>
> wrote:
> > On Tue, 2016-06-21 at 14:17 -0400, Mike Snitzer wrote:
> > > 
> > > On Tue, Jun 21 2016 at 11:44am -0400,
> > > Kani, Toshimitsu <toshi.kani@hpe.com> wrote:
> > > > 
> > > > On Tue, 2016-06-21 at 09:41 -0400, Mike Snitzer wrote:
> > > > > On Mon, Jun 20 2016 at  6:22pm -0400,
> > > > > Mike Snitzer <snitzer@redhat.com> wrote:
> > > > > I'm now wondering if we'd be better off setting a new QUEUE_FLAG_DAX
> > > > > rather than establish GENHD_FL_DAX on the genhd?
> > > > > 
> > > > > It'd be quite a bit easier to allow upper layers (e.g. XFS and ext4)
> > > > > to check for a queue flag.
> > > > 
> > > > I think GENHD_FL_DAX is more appropriate since DAX does not use a
> > > > request queue, except for protecting the underlining device being
> > > > disabled while direct_access() is called (b2e0d1625e19).
> > > 
> > > The devices in question have a request_queue.  All bio-based device have
> > > a request_queue.
> >
> > DAX-capable devices have two operation modes, bio-based and DAX.  I agree
> > that bio-based operation is associated with a request queue, and its
> > capabilities should be set to it.  DAX, on the other hand, is rather
> > independent from a request queue.
> > 
> > > I don't have a big problem with GENHD_FL_DAX.  Just wanted to point out
> > > that such block device capabilities are generally advertised in terms of
> > > a QUEUE_FLAG.
> >
> > I do not have a strong opinion, but feel a bit odd to associate DAX to a
> > request queue.
>
> Given that we do not support dax to a raw block device [1] it seems a
> gendisk flag is more misleading than request_queue flag that specifies
> what requests can be made of the device.
> 
> [1]: acc93d30d7d4 Revert "block: enable dax for raw block devices"

Oh, I see.  I will change to use request_queue flag.


> > > > About protecting direct_access, this patch assumes that the
> > > > underlining device cannot be disabled until dtr() is called.  Is this
> > > > correct?  If not, I will need to call dax_map_atomic().
> > >
> > > One of the big design considerations for DM that a DM device can be
> > > suspended (with or without flush) and any new IO will be blocked until
> > > the DM device is resumed.
> > > 
> > > So ideally DM should be able to have the same capability even if using
> > > DAX.
> >
> > Supporting suspend for DAX is challenging since it allows user
> > applications to access a device directly.  Once a device range is mmap'd,
> > there is no kernel intervention to access the range, unless we invalidate
> > user mappings.  This isn't done today even after a driver is unbind'd from
> > a device.
> > 
> > > But that is different than what commit b2e0d1625e19 is addressing.  For
> > > DM, I wouldn't think you'd need the extra protections that
> > > dax_map_atomic() is providing given that the underlying block device
> > > lifetime is managed via DM core's dm_get_device/dm_put_device (see also:
> > > dm.c:open_table_device/close_table_device).
> >
> > I thought so as well.  But I realized that there is (almost) nothing that
> > can prevent the unbind operation.  It cannot fail, either.  This unbind
> > proceeds even when a device is in-use.  In case of a pmem device, it is
> > only protected by pmem_release_queue(), which is called when a pmem device
> > is being deleted and calls blk_cleanup_queue() to serialize a critical
> > section between
> > blk_queue_enter() and blk_queue_exit() per b2e0d1625e19.  This prevents
> > from a kernel DTLB fault, but does not prevent a device disappeared while
> > in-use.
> > 
> > Protecting DM's underlining device with blk_queue_enter() (or something
> > similar) requires more thoughts...  blk_queue_enter() to a DM device
> > cannot be redirected to its underlining device.  So, this is TBD for
> > now.  But I do not think this is a blocker issue since doing unbind to a
> > underlining device is quite harmful no matter what we do - even if it is
> > protected with blk_queue_enter().
>
> I still have the "block device removed" notification patches on my
> todo list.  It's not a blocker, but there are scenarios where we can
> keep accessing memory via dax of a disabled device leading to memory
> corruption.  

Right, I noticed that user applications can access mmap'd ranges on a disabled
device.

> I'll bump that up in my queue now that we are looking at
> additional scenarios where letting DAX mappings leak past the
> reconfiguration of a block device could lead to trouble.

Great.  With DM, removing a underlining device while in-use can lead to
trouble, esp. with RAID0.  Users need to remove a device from DM first...

Thanks,
-Toshi

[toc] | [prev] | [next] | [standalone]


#1429157

FromMike Snitzer <snitzer@redhat.com>
Date2016-06-23 00:40 +0200
Message-ID<rMYYq-11q-45@gated-at.bofh.it>
In reply to#1429089
On Wed, Jun 22 2016 at  4:16P -0400,
Kani, Toshimitsu <toshi.kani@hpe.com> wrote:

> On Wed, 2016-06-22 at 12:15 -0700, Dan Williams wrote:
> > On Wed, Jun 22, 2016 at 10:44 AM, Kani, Toshimitsu <toshi.kani@hpe.com>
> > wrote:
> > > On Tue, 2016-06-21 at 14:17 -0400, Mike Snitzer wrote:
> > > > 
> > > > On Tue, Jun 21 2016 at 11:44am -0400,
> > > > Kani, Toshimitsu <toshi.kani@hpe.com> wrote:
> > > > > 
> > > > > On Tue, 2016-06-21 at 09:41 -0400, Mike Snitzer wrote:
> > > > > > On Mon, Jun 20 2016 at  6:22pm -0400,
> > > > > > Mike Snitzer <snitzer@redhat.com> wrote:
> > > > > > I'm now wondering if we'd be better off setting a new QUEUE_FLAG_DAX
> > > > > > rather than establish GENHD_FL_DAX on the genhd?
> > > > > > 
> > > > > > It'd be quite a bit easier to allow upper layers (e.g. XFS and ext4)
> > > > > > to check for a queue flag.
> > > > > 
> > > > > I think GENHD_FL_DAX is more appropriate since DAX does not use a
> > > > > request queue, except for protecting the underlining device being
> > > > > disabled while direct_access() is called (b2e0d1625e19).
> > > > 
> > > > The devices in question have a request_queue.  All bio-based device have
> > > > a request_queue.
> > >
> > > DAX-capable devices have two operation modes, bio-based and DAX.  I agree
> > > that bio-based operation is associated with a request queue, and its
> > > capabilities should be set to it.  DAX, on the other hand, is rather
> > > independent from a request queue.
> > > 
> > > > I don't have a big problem with GENHD_FL_DAX.  Just wanted to point out
> > > > that such block device capabilities are generally advertised in terms of
> > > > a QUEUE_FLAG.
> > >
> > > I do not have a strong opinion, but feel a bit odd to associate DAX to a
> > > request queue.
> >
> > Given that we do not support dax to a raw block device [1] it seems a
> > gendisk flag is more misleading than request_queue flag that specifies
> > what requests can be made of the device.
> > 
> > [1]: acc93d30d7d4 Revert "block: enable dax for raw block devices"
> 
> Oh, I see.  I will change to use request_queue flag.

I implemented the block patch for this yesterday but based on your
feedback I stopped there (didn't carry the change through to the DM core
and DM linear -- can easily do so tomorrow though).

Feel free to use this as a starting point and fix/extend and add a
proper header:

From e88736ce322f248157da6c7d402e940adafffa1e Mon Sep 17 00:00:00 2001
From: Mike Snitzer <snitzer@redhat.com>
Date: Tue, 21 Jun 2016 12:23:29 -0400
Subject: [PATCH] block: add QUEUE_FLAG_DAX for devices to advertise their DAX support

Signed-off-by: Mike Snitzer <snitzer@redhat.com>
---
 drivers/block/brd.c          | 3 +++
 drivers/nvdimm/pmem.c        | 1 +
 drivers/s390/block/dcssblk.c | 1 +
 fs/block_dev.c               | 5 +++--
 include/linux/blkdev.h       | 2 ++
 5 files changed, 10 insertions(+), 2 deletions(-)

diff --git a/drivers/block/brd.c b/drivers/block/brd.c
index f5b0d6f..13eee12 100644
--- a/drivers/block/brd.c
+++ b/drivers/block/brd.c
@@ -508,6 +508,9 @@ static struct brd_device *brd_alloc(int i)
 	brd->brd_queue->limits.discard_granularity = PAGE_SIZE;
 	blk_queue_max_discard_sectors(brd->brd_queue, UINT_MAX);
 	brd->brd_queue->limits.discard_zeroes_data = 1;
+#ifdef CONFIG_BLK_DEV_RAM_DAX
+	queue_flag_set_unlocked(QUEUE_FLAG_DAX, brd->brd_queue);
+#endif
 	queue_flag_set_unlocked(QUEUE_FLAG_DISCARD, brd->brd_queue);
 
 	disk = brd->brd_disk = alloc_disk(max_part);
diff --git a/drivers/nvdimm/pmem.c b/drivers/nvdimm/pmem.c
index 608fc44..53b701b 100644
--- a/drivers/nvdimm/pmem.c
+++ b/drivers/nvdimm/pmem.c
@@ -283,6 +283,7 @@ static int pmem_attach_disk(struct device *dev,
 	blk_queue_max_hw_sectors(q, UINT_MAX);
 	blk_queue_bounce_limit(q, BLK_BOUNCE_ANY);
 	queue_flag_set_unlocked(QUEUE_FLAG_NONROT, q);
+	queue_flag_set_unlocked(QUEUE_FLAG_DAX, q);
 	q->queuedata = pmem;
 
 	disk = alloc_disk_node(0, nid);
diff --git a/drivers/s390/block/dcssblk.c b/drivers/s390/block/dcssblk.c
index bed53c4..093e9e1 100644
--- a/drivers/s390/block/dcssblk.c
+++ b/drivers/s390/block/dcssblk.c
@@ -618,6 +618,7 @@ dcssblk_add_store(struct device *dev, struct device_attribute *attr, const char
 	dev_info->gd->driverfs_dev = &dev_info->dev;
 	blk_queue_make_request(dev_info->dcssblk_queue, dcssblk_make_request);
 	blk_queue_logical_block_size(dev_info->dcssblk_queue, 4096);
+	queue_flag_set_unlocked(QUEUE_FLAG_DAX, dev_info->dcssblk_queue);
 
 	seg_byte_size = (dev_info->end - dev_info->start + 1);
 	set_capacity(dev_info->gd, seg_byte_size >> 9); // size in sectors
diff --git a/fs/block_dev.c b/fs/block_dev.c
index 71ccab1..9bcb3a9 100644
--- a/fs/block_dev.c
+++ b/fs/block_dev.c
@@ -484,6 +484,7 @@ long bdev_direct_access(struct block_device *bdev, struct blk_dax_ctl *dax)
 	sector_t sector = dax->sector;
 	long avail, size = dax->size;
 	const struct block_device_operations *ops = bdev->bd_disk->fops;
+	struct request_queue *q = bdev_get_queue(bdev);
 
 	/*
 	 * The device driver is allowed to sleep, in order to make the
@@ -493,7 +494,7 @@ long bdev_direct_access(struct block_device *bdev, struct blk_dax_ctl *dax)
 
 	if (size < 0)
 		return size;
-	if (!ops->direct_access)
+	if (!blk_queue_dax(q) || !ops->direct_access)
 		return -EOPNOTSUPP;
 	if ((sector + DIV_ROUND_UP(size, 512)) >
 					part_nr_sects_read(bdev->bd_part))
@@ -1287,7 +1288,7 @@ static int __blkdev_get(struct block_device *bdev, fmode_t mode, int for_part)
 		bdev->bd_disk = disk;
 		bdev->bd_queue = disk->queue;
 		bdev->bd_contains = bdev;
-		if (IS_ENABLED(CONFIG_BLK_DEV_DAX) && disk->fops->direct_access)
+		if (IS_ENABLED(CONFIG_BLK_DEV_DAX) && blk_queue_dax(disk->queue))
 			bdev->bd_inode->i_flags = S_DAX;
 		else
 			bdev->bd_inode->i_flags = 0;
diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h
index 9746d22..d5cb326 100644
--- a/include/linux/blkdev.h
+++ b/include/linux/blkdev.h
@@ -505,6 +505,7 @@ struct request_queue {
 #define QUEUE_FLAG_WC	       23	/* Write back caching */
 #define QUEUE_FLAG_FUA	       24	/* device supports FUA writes */
 #define QUEUE_FLAG_FLUSH_NQ    25	/* flush not queueuable */
+#define QUEUE_FLAG_DAX	       26	/* device supports DAX */
 
 #define QUEUE_FLAG_DEFAULT	((1 << QUEUE_FLAG_IO_STAT) |		\
 				 (1 << QUEUE_FLAG_STACKABLE)	|	\
@@ -594,6 +595,7 @@ static inline void queue_flag_clear(unsigned int flag, struct request_queue *q)
 #define blk_queue_discard(q)	test_bit(QUEUE_FLAG_DISCARD, &(q)->queue_flags)
 #define blk_queue_secdiscard(q)	(blk_queue_discard(q) && \
 	test_bit(QUEUE_FLAG_SECDISCARD, &(q)->queue_flags))
+#define blk_queue_dax(q)	test_bit(QUEUE_FLAG_DAX, &(q)->queue_flags)
 
 #define blk_noretry_request(rq) \
 	((rq)->cmd_flags & (REQ_FAILFAST_DEV|REQ_FAILFAST_TRANSPORT| \
-- 
2.7.4 (Apple Git-66)

[toc] | [prev] | [next] | [standalone]


#1429296

From"Kani, Toshimitsu" <toshi.kani@hpe.com>
Date2016-06-23 01:20 +0200
Message-ID<rMZB8-1wq-59@gated-at.bofh.it>
In reply to#1429157
On Wed, 2016-06-22 at 18:38 -0400, Mike Snitzer wrote:
> On Wed, Jun 22 2016 at  4:16P -0400,
> Kani, Toshimitsu <toshi.kani@hpe.com> wrote:
> > 
> > On Wed, 2016-06-22 at 12:15 -0700, Dan Williams wrote:
> > > 
> > > On Wed, Jun 22, 2016 at 10:44 AM, Kani, Toshimitsu <toshi.kani@hpe.com>
> > > wrote:
> > > > 
> > > > On Tue, 2016-06-21 at 14:17 -0400, Mike Snitzer wrote:
> > > > > 
> > > > > On Tue, Jun 21 2016 at 11:44am -0400,
> > > > > Kani, Toshimitsu <toshi.kani@hpe.com> wrote:
> > > > > > 
> > > > > > On Tue, 2016-06-21 at 09:41 -0400, Mike Snitzer wrote:
> > > > > > > 
> > > > > The devices in question have a request_queue.  All bio-based device
> > > > > have a request_queue.
> > > > 
> > > > DAX-capable devices have two operation modes, bio-based and DAX.  I
> > > > agree that bio-based operation is associated with a request queue, and
> > > > its capabilities should be set to it.  DAX, on the other hand, is
> > > > rather independent from a request queue.
> > > > 
> > > > > 
> > > > > I don't have a big problem with GENHD_FL_DAX.  Just wanted to point
> > > > > out that such block device capabilities are generally advertised in
> > > > > terms of a QUEUE_FLAG.
> > > > 
> > > > I do not have a strong opinion, but feel a bit odd to associate DAX to
> > > > a request queue.
> > >
> > > Given that we do not support dax to a raw block device [1] it seems a
> > > gendisk flag is more misleading than request_queue flag that specifies
> > > what requests can be made of the device.
> > > 
> > > [1]: acc93d30d7d4 Revert "block: enable dax for raw block devices"
> >
> > Oh, I see.  I will change to use request_queue flag.
>
> I implemented the block patch for this yesterday but based on your
> feedback I stopped there (didn't carry the change through to the DM core
> and DM linear -- can easily do so tomorrow though).
> 
> Feel free to use this as a starting point and fix/extend and add a
> proper header:

Thanks Mike! I made similar changes as well. I will take yours and finish up
the rest. :)
-Toshi

[toc] | [prev] | [next] | [standalone]


#1427115

From"Kani, Toshimitsu" <toshi.kani@hpe.com>
Date2016-06-21 00:50 +0200
Message-ID<rMfyi-5y7-29@gated-at.bofh.it>
In reply to#1427077
On Mon, 2016-06-20 at 15:52 -0400, Mike Snitzer wrote:
> On Mon, Jun 20 2016 at  3:40pm -0400,
> Mike Snitzer <snitzer@redhat.com> wrote:
>  
> > 
> > # dd if=/dev/zero of=/mnt/dax/meh bs=1024K oflag=direct
> > [11729.754671] XFS (dm-4): Metadata corruption detected at
> > xfs_agf_read_verify+0x70/0x120 [xfs], xfs_agf block 0x45a808
> > [11729.766423] XFS (dm-4): Unmount and run xfs_repair
> > [11729.771774] XFS (dm-4): First 64 bytes of corrupted metadata buffer:
> > [11729.778869] ffff8800b8038000: 00 00 00 00 00 00 00 00 00 00 00 00 00
> > 00 00 00  ................
> > [11729.788582] ffff8800b8038010: 00 00 00 00 00 00 00 00 00 00 00 00 00
> > 00 00 00  ................
> > [11729.798293] ffff8800b8038020: 00 00 00 00 00 00 00 00 00 00 00 00 00
> > 00 00 00  ................
> > [11729.808002] ffff8800b8038030: 00 00 00 00 00 00 00 00 00 00 00 00 00
> > 00 00 00  ................
> > [11729.817715] XFS (dm-4): metadata I/O error: block 0x45a808
> > ("xfs_trans_read_buf_map") error 117 numblks 8
> > 
> > When this XFS corruption occurs corruption then also manifests in lvm2's
> > metadata:
> > 
> > # vgremove pmem
> > Do you really want to remove volume group "pmem" containing 1 logical
> > volumes? [y/n]: y
> > Do you really want to remove active logical volume lv? [y/n]: y
> >   Incorrect metadata area header checksum on /dev/pmem0p1 at offset 4096
> >   WARNING: Failed to write an MDA of VG pmem.
> >   Incorrect metadata area header checksum on /dev/pmem0p2 at offset 4096
> >   WARNING: Failed to write an MDA of VG pmem.
> >   Failed to write VG pmem.
> >   Incorrect metadata area header checksum on /dev/pmem0p2 at offset 4096
> >   Incorrect metadata area header checksum on /dev/pmem0p1 at offset 4096
> > 
> > If I don't use XFS, and only issue IO directly to the /dev/pmem/lv, I
> > don't see this corruption.
> I did the same test with ext4 instead of xfs and it resulted in the same
> type of systemic corruption (lvm2 metadata corrupted too):
> 
> [12816.407147] EXT4-fs (dm-4): DAX enabled. Warning: EXPERIMENTAL, use at
> your own risk
> [12816.416123] EXT4-fs (dm-4): mounted filesystem with ordered data mode.
> Opts: dax
> [12816.766855] EXT4-fs error (device dm-4): ext4_mb_generate_buddy:758:
> group 9, block bitmap and bg descriptor inconsistent: 32768 vs 32395 free
> clusters
> [12816.782016] EXT4-fs error (device dm-4): ext4_mb_generate_buddy:758:
> group 10, block bitmap and bg descriptor inconsistent: 32768 vs 16384 free
> clusters
> [12816.797491] JBD2: Spotted dirty metadata buffer (dev = dm-4, blocknr =
> 0). There's a risk of filesystem corruption in case of system crash.
> 
> # vgremove pmem
> Do you really want to remove volume group "pmem" containing 1 logical
> volumes? [y/n]: y
> Do you really want to remove active logical volume lv? [y/n]: y
>   Incorrect metadata area header checksum on /dev/pmem0p1 at offset 4096
>   WARNING: Failed to write an MDA of VG pmem.
>   Incorrect metadata area header checksum on /dev/pmem0p2 at offset 4096
>   WARNING: Failed to write an MDA of VG pmem.
>   Failed to write VG pmem.
>   Incorrect metadata area header checksum on /dev/pmem0p2 at offset 4096
>   Incorrect metadata area header checksum on /dev/pmem0p1 at offset 4096

I will look into the issue.

Thanks for the testing!
-Toshi

[toc] | [prev] | [next] | [standalone]


#1427043

From"Kani, Toshimitsu" <toshi.kani@hpe.com>
Date2016-06-20 23:10 +0200
Message-ID<rMdmO-41g-11@gated-at.bofh.it>
In reply to#1426910
On Mon, 2016-06-20 at 14:00 -0400, Mike Snitzer wrote:
> On Mon, Jun 13 2016 at  6:57pm -0400,
> Mike Snitzer <snitzer@redhat.com> wrote:
> 
> > 
> > On Mon, Jun 13 2016 at  6:21pm -0400,
> > Toshi Kani <toshi.kani@hpe.com> wrote:
> > 
> > > 
> > > This patch-set adds DAX support to device-mapper dm-linear devices
> > > used by LVM.  It works with LVM commands as follows:
> > >  - Creation of a logical volume with all DAX capable devices (such
> > >    as pmem) sets the logical volume DAX capable as well.
> > >  - Once a logical volume is set to DAX capable, the volume may not
> > >    be extended with non-DAX capable devices.
> > > 
> > > The direct_access interface is added to dm and dm-linear to map
> > > a request to a target device.
> > > 
> > >  - Patch 1-2 introduce GENHD_FL_DAX flag to indicate DAX capability.
> > >  - Patch 3-4 add direct_access functions to dm and dm-linear.
> > >  - Patch 5-6 set GENHD_FL_DAX to dm when all targets are DAX capable.
> > > 
> > > ---
> > > Toshi Kani (6):
> > >  1/6 genhd: Add GENHD_FL_DAX to gendisk flags
> > >  2/6 block: Check GENHD_FL_DAX for DAX capability
> > >  3/6 dm: Add dm_blk_direct_access() for mapped device
> > >  4/6 dm-linear: Add linear_direct_access()
> > >  5/6 dm, dm-linear: Add dax_supported to dm_target
> > >  6/6 dm: Enable DAX support for mapper device
> > Thanks a lot for doing this.  I recently added it to my TODO so your
> > patches come at a great time.
> > 
> > I'll try to get to reviewing/testing your work by the end of this week.
>
> I rebased your patches on linux-dm.git's 'for-next' (which includes what
> I've already staged for the 4.8 merge window).  And I folded/changed
> some of the DM patches so that there are only 2 now (1 for DM core and 1
> for dm-linear).  Please see the 4 topmost commits in my 'wip' here:
> 
> http://git.kernel.org/cgit/linux/kernel/git/snitzer/linux.git/log/?h=wip
> 
> Feel free to pick these patches up to use as the basis for continued
> work or re-posting of this set.. either that or I could post them as v2
> on your behalf.
> 
> As for testing, I've verified that basic IO works to a pmem-based DM
> linear device and that mixed table types are rejected as expected.

Great! I will send additional patch, add DAX support to dm-stripe, on top of
these once I finish my testing.

Thanks,
-Toshi

[toc] | [prev] | [next] | [standalone]


#1421444

FromDan Williams <dan.j.williams@intel.com>
Date2016-06-14 01:20 +0200
Message-ID<rJJjb-4NW-7@gated-at.bofh.it>
In reply to#1421388
Thanks Toshi!

On Mon, Jun 13, 2016 at 3:21 PM, Toshi Kani <toshi.kani@hpe.com> wrote:
> This patch-set adds DAX support to device-mapper dm-linear devices
> used by LVM.  It works with LVM commands as follows:
>  - Creation of a logical volume with all DAX capable devices (such
>    as pmem) sets the logical volume DAX capable as well.
>  - Once a logical volume is set to DAX capable, the volume may not
>    be extended with non-DAX capable devices.

I don't mind this, but it seems a policy decision that the kernel does
not need to make.  A sufficiently sophisticated user could cope with
DAX being available at varying LBAs.  Would it be sufficient to move
this policy decision to userspace tooling?

> The direct_access interface is added to dm and dm-linear to map
> a request to a target device.

I had dm-linear and md-raid0 support on my list of things to look at,
did you have raid0 in your plans?

[toc] | [prev] | [next] | [standalone]


#1421456

From"Kani, Toshimitsu" <toshi.kani@hpe.com>
Date2016-06-14 02:00 +0200
Message-ID<rJJVT-554-1@gated-at.bofh.it>
In reply to#1421444
On Mon, 2016-06-13 at 16:18 -0700, Dan Williams wrote:
> Thanks Toshi!
> 
> On Mon, Jun 13, 2016 at 3:21 PM, Toshi Kani <toshi.kani@hpe.com> wrote:
> > 
> > This patch-set adds DAX support to device-mapper dm-linear devices
> > used by LVM.  It works with LVM commands as follows:
> >  - Creation of a logical volume with all DAX capable devices (such
> >    as pmem) sets the logical volume DAX capable as well.
> >  - Once a logical volume is set to DAX capable, the volume may not
> >    be extended with non-DAX capable devices.
>
> I don't mind this, but it seems a policy decision that the kernel does
> not need to make.  A sufficiently sophisticated user could cope with
> DAX being available at varying LBAs.  Would it be sufficient to move
> this policy decision to userspace tooling?

I think this is a kernel restriction.  When a block device is declared as
DAX capable, it should mean that the whole device is DAX capable.  So, I
think we need to assure the same to a mapped device.

In LVM level, a volume group may contain both DAX and non-DAX capable
devices.  There is no restriction for creating/extending a volume group.

> > The direct_access interface is added to dm and dm-linear to map
> > a request to a target device.
>
> I had dm-linear and md-raid0 support on my list of things to look at,
> did you have raid0 in your plans?

Yes, I hope to extend further and raid0 is a good candidate.   

Thanks,
-Toshi

[toc] | [prev] | [next] | [standalone]


#1421461

FromDan Williams <dan.j.williams@intel.com>
Date2016-06-14 02:10 +0200
Message-ID<rJK5z-5nH-9@gated-at.bofh.it>
In reply to#1421456
On Mon, Jun 13, 2016 at 4:59 PM, Kani, Toshimitsu <toshi.kani@hpe.com> wrote:
> On Mon, 2016-06-13 at 16:18 -0700, Dan Williams wrote:
>> Thanks Toshi!
>>
>> On Mon, Jun 13, 2016 at 3:21 PM, Toshi Kani <toshi.kani@hpe.com> wrote:
>> >
>> > This patch-set adds DAX support to device-mapper dm-linear devices
>> > used by LVM.  It works with LVM commands as follows:
>> >  - Creation of a logical volume with all DAX capable devices (such
>> >    as pmem) sets the logical volume DAX capable as well.
>> >  - Once a logical volume is set to DAX capable, the volume may not
>> >    be extended with non-DAX capable devices.
>>
>> I don't mind this, but it seems a policy decision that the kernel does
>> not need to make.  A sufficiently sophisticated user could cope with
>> DAX being available at varying LBAs.  Would it be sufficient to move
>> this policy decision to userspace tooling?
>
> I think this is a kernel restriction.  When a block device is declared as
> DAX capable, it should mean that the whole device is DAX capable.  So, I
> think we need to assure the same to a mapped device.

Hmm, but we already violate this with badblocks.  The device is DAX
capable, but certain LBAs will return an error if direct_access is
attempted.

[toc] | [prev] | [next] | [standalone]


#1421632

FromDan Williams <dan.j.williams@intel.com>
Date2016-06-14 09:40 +0200
Message-ID<rJR73-1H3-5@gated-at.bofh.it>
In reply to#1421461
On Mon, Jun 13, 2016 at 5:02 PM, Dan Williams <dan.j.williams@intel.com> wrote:
> On Mon, Jun 13, 2016 at 4:59 PM, Kani, Toshimitsu <toshi.kani@hpe.com> wrote:
>> On Mon, 2016-06-13 at 16:18 -0700, Dan Williams wrote:
>>> Thanks Toshi!
>>>
>>> On Mon, Jun 13, 2016 at 3:21 PM, Toshi Kani <toshi.kani@hpe.com> wrote:
>>> >
>>> > This patch-set adds DAX support to device-mapper dm-linear devices
>>> > used by LVM.  It works with LVM commands as follows:
>>> >  - Creation of a logical volume with all DAX capable devices (such
>>> >    as pmem) sets the logical volume DAX capable as well.
>>> >  - Once a logical volume is set to DAX capable, the volume may not
>>> >    be extended with non-DAX capable devices.
>>>
>>> I don't mind this, but it seems a policy decision that the kernel does
>>> not need to make.  A sufficiently sophisticated user could cope with
>>> DAX being available at varying LBAs.  Would it be sufficient to move
>>> this policy decision to userspace tooling?
>>
>> I think this is a kernel restriction.  When a block device is declared as
>> DAX capable, it should mean that the whole device is DAX capable.  So, I
>> think we need to assure the same to a mapped device.
>
> Hmm, but we already violate this with badblocks.  The device is DAX
> capable, but certain LBAs will return an error if direct_access is
> attempted.

Nevermind, for this to be useful we would need to fallback to regular
mmap for a portion of the linear span.  That's different than the
badblocks case.

[toc] | [prev] | [next] | [standalone]


#1421910

FromJeff Moyer <jmoyer@redhat.com>
Date2016-06-14 16:00 +0200
Message-ID<rJX2N-5z3-3@gated-at.bofh.it>
In reply to#1421456
"Kani, Toshimitsu" <toshi.kani@hpe.com> writes:

>> I had dm-linear and md-raid0 support on my list of things to look at,
>> did you have raid0 in your plans?
>
> Yes, I hope to extend further and raid0 is a good candidate.   

dm-flakey would allow more xfstests test cases to run.  I'd say that's
more important than linear or raid0.  ;-)

Also, the next step in this work is to then decide how to determine on
what numa node an LBA resides.  We had discussed this at a prior
plumbers conference, and I think the consensus was to use xattrs.
Toshi, do you also plan to do that work?

Cheers,
Jeff

[toc] | [prev] | [next] | [standalone]


#1422025

FromMike Snitzer <snitzer@redhat.com>
Date2016-06-14 17:50 +0200
Message-ID<rJYLg-6MM-9@gated-at.bofh.it>
In reply to#1421910
On Tue, Jun 14 2016 at  9:50am -0400,
Jeff Moyer <jmoyer@redhat.com> wrote:

> "Kani, Toshimitsu" <toshi.kani@hpe.com> writes:
> 
> >> I had dm-linear and md-raid0 support on my list of things to look at,
> >> did you have raid0 in your plans?
> >
> > Yes, I hope to extend further and raid0 is a good candidate.   
> 
> dm-flakey would allow more xfstests test cases to run.  I'd say that's
> more important than linear or raid0.  ;-)

Regardless of which target(s) grow DAX support the most pressing initial
concern is getting the DM device stacking correct.  And verifying that
IO that cross pmem device boundaries are being properly split by DM
core (via drivers/md/dm.c:__split_and_process_non_flush()'s call to
max_io_len).

My hope is to nail down the DM core and its dependencies in block etc.
Doing so in terms of dm-linear doesn't seem like wasted effort
considering you told me it'd be useful to have for pmem devices.
 
> Also, the next step in this work is to then decide how to determine on
> what numa node an LBA resides.  We had discussed this at a prior
> plumbers conference, and I think the consensus was to use xattrs.
> Toshi, do you also plan to do that work?

How does the associated NUMA node relate to this?  Does the
DM requests_queue need to be setup to only allocate from the NUMA node
the pmem device is attached to?  I recently added support for this to
DM.  But there will likely be some code need to propagate the NUMA node
id accordingly.

Mike

[toc] | [prev] | [next] | [standalone]


#1422159

From"Kani, Toshimitsu" <toshi.kani@hpe.com>
Date2016-06-14 20:10 +0200
Message-ID<rK0WJ-8p2-13@gated-at.bofh.it>
In reply to#1422025
On Tue, 2016-06-14 at 11:41 -0400, Mike Snitzer wrote:
> On Tue, Jun 14 2016 at  9:50am -0400,
> Jeff Moyer <jmoyer@redhat.com> wrote:
> > "Kani, Toshimitsu" <toshi.kani@hpe.com> writes:
> > > > I had dm-linear and md-raid0 support on my list of things to look
> > > > at, did you have raid0 in your plans?
> > >
> > > Yes, I hope to extend further and raid0 is a good candidate. 
> >   
> > dm-flakey would allow more xfstests test cases to run.  I'd say that's
> > more important than linear or raid0.  ;-)
>
> Regardless of which target(s) grow DAX support the most pressing initial
> concern is getting the DM device stacking correct.  And verifying that
> IO that cross pmem device boundaries are being properly split by DM
> core (via drivers/md/dm.c:__split_and_process_non_flush()'s call to
> max_io_len).

Agreed. I've briefly tested stacking and it seems working fine.  As for IO
crossing pmem device boundaries, __split_and_process_non_flush() is used
when the device is mounted without DAX option.  With DAX, this case is
handled by dm_blk_direct_access() that limits return size.  This leads the
caller to iterate (read/write) or fallback to a smaller size (mmap pfault).

> My hope is to nail down the DM core and its dependencies in block etc.
> Doing so in terms of dm-linear doesn't seem like wasted effort
> considering you told me it'd be useful to have for pmem devices.

Yes, I think dm-linear is useful as it gives more flexibility, ex. it allows
creating a large device with multiple pmem devices.

> > Also, the next step in this work is to then decide how to determine on
> > what numa node an LBA resides.  We had discussed this at a prior
> > plumbers conference, and I think the consensus was to use xattrs.
> > Toshi, do you also plan to do that work?
>
> How does the associated NUMA node relate to this?  Does the
> DM requests_queue need to be setup to only allocate from the NUMA node
> the pmem device is attached to?  I recently added support for this to
> DM.  But there will likely be some code need to propagate the NUMA node
> id accordingly.

Each pmem device has sysfs "numa_node" so that tools like numactl can be
used to bind application to run on the same locality as pmem device (since
CPU directly accesses to pmem).  This won't work well with mapped device
since it can be composed with multiple localities.  Locality info would need
to be managed file-basis as Jeff mentioned.

Thanks,
-Toshi

[toc] | [prev] | [next] | [standalone]


#1422294

FromJeff Moyer <jmoyer@redhat.com>
Date2016-06-14 22:20 +0200
Message-ID<rK2Yy-1dW-15@gated-at.bofh.it>
In reply to#1422025
Mike Snitzer <snitzer@redhat.com> writes:

> On Tue, Jun 14 2016 at  9:50am -0400,
> Jeff Moyer <jmoyer@redhat.com> wrote:
>
>> "Kani, Toshimitsu" <toshi.kani@hpe.com> writes:
>> 
>> >> I had dm-linear and md-raid0 support on my list of things to look at,
>> >> did you have raid0 in your plans?
>> >
>> > Yes, I hope to extend further and raid0 is a good candidate.   
>> 
>> dm-flakey would allow more xfstests test cases to run.  I'd say that's
>> more important than linear or raid0.  ;-)
>
> Regardless of which target(s) grow DAX support the most pressing initial
> concern is getting the DM device stacking correct.  And verifying that
> IO that cross pmem device boundaries are being properly split by DM
> core (via drivers/md/dm.c:__split_and_process_non_flush()'s call to
> max_io_len).

That was a tongue-in-cheek comment.  You're reading way too much into
it.

>> Also, the next step in this work is to then decide how to determine on
>> what numa node an LBA resides.  We had discussed this at a prior
>> plumbers conference, and I think the consensus was to use xattrs.
>> Toshi, do you also plan to do that work?
>
> How does the associated NUMA node relate to this?  Does the
> DM requests_queue need to be setup to only allocate from the NUMA node
> the pmem device is attached to?  I recently added support for this to
> DM.  But there will likely be some code need to propagate the NUMA node
> id accordingly.

I assume you mean allocate memory (the volatile kind).  That should work
the same between pmem and regular block devices, no?

What I was getting at was that applications may want to know on which
node their data resides.  Right now, it's easy to tell because a single
device cannot span numa nodes, or, if it does, it does so via an
interleave, so numa information isn't interesting.  However, once data
on a single file system can be placed on multiple different numa nodes,
applications may want to query and/or control that placement.

Here's a snippet from a blog post I never finished:

There are two essential questions that need to be answered regarding
persistent memory and NUMA: first, would an application benefit from
being able to query the NUMA locality of its data, and second, would
an application benefit from being able to specify a placement policy
for its data?  This article is an attempt to summarize the current
state of hardware and software in order to consider the above two
questions.  We begin with a short list of use cases for these
interfaces, which will frame the discussion.

First, let's consider an interface that allows an application to query
the NUMA placement of existing data.  With such information, an
application may want to perform the following actions:

- relocate application processes to the same NUMA node as their data.
  (Interfaces for moving a process are readily available.)
- specify a memory (RAM) allocation policy so that memory allocations
  come from the same NUMA node as the data.

Second, we consider an interface that allows an application to specify
a placement policy for new data.  Using this interface, an application
may:

- ensure data is stored on the same NUMA node as the one on which the
  application is running
- ensure data is stored on the same NUMA node as an I/O adapter such
  as a network card, that is a producer of data stored to NVM.
- ensure data is stored on a different NUMA node:
  - so that the data is stored on the same NUMA node as related data
  - because the data does not need the faster access afforded by local
    NUMA placement.  Presumably this is a trade-off, and other data
    will require local placement to meet the performance goals of the
    application.

Cheers,
Jeff

[toc] | [prev] | [next] | [standalone]


#1422480

FromMike Snitzer <snitzer@redhat.com>
Date2016-06-15 03:50 +0200
Message-ID<rK87T-4hO-3@gated-at.bofh.it>
In reply to#1422294
On Tue, Jun 14 2016 at  4:19pm -0400,
Jeff Moyer <jmoyer@redhat.com> wrote:

> Mike Snitzer <snitzer@redhat.com> writes:
> 
> > On Tue, Jun 14 2016 at  9:50am -0400,
> > Jeff Moyer <jmoyer@redhat.com> wrote:
> >
> >> "Kani, Toshimitsu" <toshi.kani@hpe.com> writes:
> >> 
> >> >> I had dm-linear and md-raid0 support on my list of things to look at,
> >> >> did you have raid0 in your plans?
> >> >
> >> > Yes, I hope to extend further and raid0 is a good candidate.   
> >> 
> >> dm-flakey would allow more xfstests test cases to run.  I'd say that's
> >> more important than linear or raid0.  ;-)
> >
> > Regardless of which target(s) grow DAX support the most pressing initial
> > concern is getting the DM device stacking correct.  And verifying that
> > IO that cross pmem device boundaries are being properly split by DM
> > core (via drivers/md/dm.c:__split_and_process_non_flush()'s call to
> > max_io_len).
> 
> That was a tongue-in-cheek comment.  You're reading way too much into
> it.
> 
> >> Also, the next step in this work is to then decide how to determine on
> >> what numa node an LBA resides.  We had discussed this at a prior
> >> plumbers conference, and I think the consensus was to use xattrs.
> >> Toshi, do you also plan to do that work?
> >
> > How does the associated NUMA node relate to this?  Does the
> > DM requests_queue need to be setup to only allocate from the NUMA node
> > the pmem device is attached to?  I recently added support for this to
> > DM.  But there will likely be some code need to propagate the NUMA node
> > id accordingly.
> 
> I assume you mean allocate memory (the volatile kind).  That should work
> the same between pmem and regular block devices, no?

This is the commit I made to train DM to be numa node aware:
115485e83f497fdf9b4 ("dm: add 'dm_numa_node' module parameter")

As is the DM code is focused on memory allocations.  But I think blk-mq
may use the NUMA node for via tag_set->numa_node.  But that is moot
given pmem is bio-based right?

Steps could be taken to make all threads DM creates for a a given device
get pinned to the specified NUMA node too.

[toc] | [prev] | [next] | [standalone]


#1422488

FromDan Williams <dan.j.williams@intel.com>
Date2016-06-15 04:10 +0200
Message-ID<rK8rf-4F5-11@gated-at.bofh.it>
In reply to#1422480
On Tue, Jun 14, 2016 at 6:46 PM, Mike Snitzer <snitzer@redhat.com> wrote:
> On Tue, Jun 14 2016 at  4:19pm -0400,
> Jeff Moyer <jmoyer@redhat.com> wrote:
>
>> Mike Snitzer <snitzer@redhat.com> writes:
>>
>> > On Tue, Jun 14 2016 at  9:50am -0400,
>> > Jeff Moyer <jmoyer@redhat.com> wrote:
>> >
>> >> "Kani, Toshimitsu" <toshi.kani@hpe.com> writes:
>> >>
>> >> >> I had dm-linear and md-raid0 support on my list of things to look at,
>> >> >> did you have raid0 in your plans?
>> >> >
>> >> > Yes, I hope to extend further and raid0 is a good candidate.
>> >>
>> >> dm-flakey would allow more xfstests test cases to run.  I'd say that's
>> >> more important than linear or raid0.  ;-)
>> >
>> > Regardless of which target(s) grow DAX support the most pressing initial
>> > concern is getting the DM device stacking correct.  And verifying that
>> > IO that cross pmem device boundaries are being properly split by DM
>> > core (via drivers/md/dm.c:__split_and_process_non_flush()'s call to
>> > max_io_len).
>>
>> That was a tongue-in-cheek comment.  You're reading way too much into
>> it.
>>
>> >> Also, the next step in this work is to then decide how to determine on
>> >> what numa node an LBA resides.  We had discussed this at a prior
>> >> plumbers conference, and I think the consensus was to use xattrs.
>> >> Toshi, do you also plan to do that work?
>> >
>> > How does the associated NUMA node relate to this?  Does the
>> > DM requests_queue need to be setup to only allocate from the NUMA node
>> > the pmem device is attached to?  I recently added support for this to
>> > DM.  But there will likely be some code need to propagate the NUMA node
>> > id accordingly.
>>
>> I assume you mean allocate memory (the volatile kind).  That should work
>> the same between pmem and regular block devices, no?
>
> This is the commit I made to train DM to be numa node aware:
> 115485e83f497fdf9b4 ("dm: add 'dm_numa_node' module parameter")

Hmm, but this is global for all DM device instances.

> As is the DM code is focused on memory allocations.  But I think blk-mq
> may use the NUMA node for via tag_set->numa_node.  But that is moot
> given pmem is bio-based right?

Right.

>
> Steps could be taken to make all threads DM creates for a a given device
> get pinned to the specified NUMA node too.

I think it would be useful if a DM instance inherited the numa node
from the component devices by default (assuming they're all from the
same node).  A "dev_to_node(disk_to_dev(disk))" conversion works for
pmem devices.

As far as I understand, Jeff wants to go further and have a linear
span across component devices from different nodes with an interface
to do an LBA-to-numa-node conversion.

[toc] | [prev] | [next] | [standalone]


#1422514

FromMike Snitzer <snitzer@redhat.com>
Date2016-06-15 04:40 +0200
Message-ID<rK8Ui-4WP-19@gated-at.bofh.it>
In reply to#1422488
On Tue, Jun 14 2016 at 10:07pm -0400,
Dan Williams <dan.j.williams@intel.com> wrote:

> On Tue, Jun 14, 2016 at 6:46 PM, Mike Snitzer <snitzer@redhat.com> wrote:
> > On Tue, Jun 14 2016 at  4:19pm -0400,
> > Jeff Moyer <jmoyer@redhat.com> wrote:
> >
> >> Mike Snitzer <snitzer@redhat.com> writes:
> >>
> >> > On Tue, Jun 14 2016 at  9:50am -0400,
> >> > Jeff Moyer <jmoyer@redhat.com> wrote:
> >> >
> >> >> "Kani, Toshimitsu" <toshi.kani@hpe.com> writes:
> >> >>
> >> >> >> I had dm-linear and md-raid0 support on my list of things to look at,
> >> >> >> did you have raid0 in your plans?
> >> >> >
> >> >> > Yes, I hope to extend further and raid0 is a good candidate.
> >> >>
> >> >> dm-flakey would allow more xfstests test cases to run.  I'd say that's
> >> >> more important than linear or raid0.  ;-)
> >> >
> >> > Regardless of which target(s) grow DAX support the most pressing initial
> >> > concern is getting the DM device stacking correct.  And verifying that
> >> > IO that cross pmem device boundaries are being properly split by DM
> >> > core (via drivers/md/dm.c:__split_and_process_non_flush()'s call to
> >> > max_io_len).
> >>
> >> That was a tongue-in-cheek comment.  You're reading way too much into
> >> it.
> >>
> >> >> Also, the next step in this work is to then decide how to determine on
> >> >> what numa node an LBA resides.  We had discussed this at a prior
> >> >> plumbers conference, and I think the consensus was to use xattrs.
> >> >> Toshi, do you also plan to do that work?
> >> >
> >> > How does the associated NUMA node relate to this?  Does the
> >> > DM requests_queue need to be setup to only allocate from the NUMA node
> >> > the pmem device is attached to?  I recently added support for this to
> >> > DM.  But there will likely be some code need to propagate the NUMA node
> >> > id accordingly.
> >>
> >> I assume you mean allocate memory (the volatile kind).  That should work
> >> the same between pmem and regular block devices, no?
> >
> > This is the commit I made to train DM to be numa node aware:
> > 115485e83f497fdf9b4 ("dm: add 'dm_numa_node' module parameter")
> 
> Hmm, but this is global for all DM device instances.

Right, only because I didn't have a convenient way to allow the user to
specify it on a per-device level.  But I'll defer skinning that cat for
now since in this pmem case we'd inherit from the underlying device(s)

> > As is the DM code is focused on memory allocations.  But I think blk-mq
> > may use the NUMA node for via tag_set->numa_node.  But that is moot
> > given pmem is bio-based right?
> 
> Right.
> 
> >
> > Steps could be taken to make all threads DM creates for a a given device
> > get pinned to the specified NUMA node too.
> 
> I think it would be useful if a DM instance inherited the numa node
> from the component devices by default (assuming they're all from the
> same node).  A "dev_to_node(disk_to_dev(disk))" conversion works for
> pmem devices.

OK, I can look to make that happen.
 
> As far as I understand, Jeff wants to go further and have a linear
> span across component devices from different nodes with an interface
> to do an LBA-to-numa-node conversion.

All that variability makes DM's ability to do anything sane with it
close to impossible considering memory pools, threads, etc are all
pinned during the first activation of the DM device.

[toc] | [prev] | [next] | [standalone]


#1422045

From"Kani, Toshimitsu" <toshi.kani@hpe.com>
Date2016-06-14 18:00 +0200
Message-ID<rJYUW-6QY-37@gated-at.bofh.it>
In reply to#1421910
On Tue, 2016-06-14 at 09:50 -0400, Jeff Moyer wrote:
> "Kani, Toshimitsu" <toshi.kani@hpe.com> writes:
> > > I had dm-linear and md-raid0 support on my list of things to look at,
> > > did you have raid0 in your plans?
> >
> > Yes, I hope to extend further and raid0 is a good candidate. 
>
> dm-flakey would allow more xfstests test cases to run.  I'd say that's
> more important than linear or raid0.  ;-)

That's an interesting one.  We can emulate badblocks by failing
direct_access with -EIO, but I do not think we can emulate "drop_writes" and
"corrupt_bio_byte" features with DAX...

> Also, the next step in this work is to then decide how to determine on
> what numa node an LBA resides.  We had discussed this at a prior
> plumbers conference, and I think the consensus was to use xattrs.
> Toshi, do you also plan to do that work?

No, it's not my plan at this point.

Thanks,
-Toshi

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web