Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1478773 > unrolled thread

DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

Started byDan Williams <dan.j.williams@intel.com>
First post2016-09-08 06:40 +0200
Last post2016-09-16 08:00 +0200
Articles 20 on this page of 31 — 9 participants

Back to article view | Back to linux.kernel


Contents

  DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dan Williams <dan.j.williams@intel.com> - 2016-09-08 06:40 +0200
    Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-09-09 01:00 +0200
      Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dan Williams <dan.j.williams@intel.com> - 2016-09-09 01:10 +0200
        Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-09-09 11:10 +0200
          Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dan Williams <dan.j.williams@intel.com> - 2016-09-09 17:50 +0200
            Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-09-12 08:10 +0200
          Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) "Rudoff, Andy" <andy.rudoff@intel.com> - 2016-09-12 05:50 +0200
            Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-09-12 08:40 +0200
      Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-12 03:50 +0200
        Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) "Darrick J. Wong" <darrick.wong@oracle.com> - 2016-09-15 08:00 +0200
          Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-15 08:30 +0200
      Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-12 07:30 +0200
        Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) "Oliver O'Halloran" <oohall@gmail.com> - 2016-09-12 09:30 +0200
          Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-12 10:00 +0200
            Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-12 10:10 +0200
              Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-12 17:10 +0200
                Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 03:40 +0200
                  Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dan Williams <dan.j.williams@intel.com> - 2016-09-13 06:10 +0200
                    Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 07:50 +0200
              Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-12 23:40 +0200
                Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 04:00 +0200
                  Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-13 09:20 +0200
                    Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 11:10 +0200
                  Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-14 09:40 +0200
                    Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-14 12:20 +0200
                      Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-15 04:40 +0200
                        Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-15 06:00 +0200
                          Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-15 12:40 +0200
                            Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-15 13:50 +0200
                              Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-16 00:40 +0200
                                Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in  /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-16 08:00 +0200

Page 1 of 2  [1] 2  Next page →


#1478773 — DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromDan Williams <dan.j.williams@intel.com>
Date2016-09-08 06:40 +0200
SubjectDAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<seZi1-15I-15@gated-at.bofh.it>
[ adding linux-fsdevel and linux-nvdimm ]

On Wed, Sep 7, 2016 at 8:36 PM, Xiao Guangrong
<guangrong.xiao@linux.intel.com> wrote:
[..]
> However, it is not easy to handle the case that the new VMA overlays with
> the old VMA
> already got by userspace. I think we have some choices:
> 1: One way is completely skipping the new VMA region as current kernel code
> does but i
>    do not think this is good as the later VMAs will be dropped.
>
> 2: show the un-overlayed portion of new VMA. In your case, we just show the
> region
>    (0x2000 -> 0x3000), however, it can not work well if the VMA is a new
> created
>    region with different attributions.
>
> 3: completely show the new VMA as this patch does.
>
> Which one do you prefer?
>

I don't have a preference, but perhaps this breakage and uncertainty
is a good opportunity to propose a more reliable interface for NVML to
get the information it needs?

My understanding is that it is looking for the VM_MIXEDMAP flag which
is already ambiguous for determining if DAX is enabled even if this
dynamic listing issue is fixed.  XFS has arranged for DAX to be a
per-inode capability and has an XFS-specific inode flag.  We can make
that a common inode flag, but it seems we should have a way to
interrogate the mapping itself in the case where the inode is unknown
or unavailable.  I'm thinking extensions to mincore to have flags for
DAX and possibly whether the page is part of a pte, pmd, or pud
mapping.  Just floating that idea before starting to look into the
implementation, comments or other ideas welcome...

[toc] | [next] | [standalone]


#1479573 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromRoss Zwisler <ross.zwisler@linux.intel.com>
Date2016-09-09 01:00 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sfgsz-3pz-85@gated-at.bofh.it>
In reply to#1478773
On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote:
> [ adding linux-fsdevel and linux-nvdimm ]
> 
> On Wed, Sep 7, 2016 at 8:36 PM, Xiao Guangrong
> <guangrong.xiao@linux.intel.com> wrote:
> [..]
> > However, it is not easy to handle the case that the new VMA overlays with
> > the old VMA
> > already got by userspace. I think we have some choices:
> > 1: One way is completely skipping the new VMA region as current kernel code
> > does but i
> >    do not think this is good as the later VMAs will be dropped.
> >
> > 2: show the un-overlayed portion of new VMA. In your case, we just show the
> > region
> >    (0x2000 -> 0x3000), however, it can not work well if the VMA is a new
> > created
> >    region with different attributions.
> >
> > 3: completely show the new VMA as this patch does.
> >
> > Which one do you prefer?
> >
> 
> I don't have a preference, but perhaps this breakage and uncertainty
> is a good opportunity to propose a more reliable interface for NVML to
> get the information it needs?
> 
> My understanding is that it is looking for the VM_MIXEDMAP flag which
> is already ambiguous for determining if DAX is enabled even if this
> dynamic listing issue is fixed.  XFS has arranged for DAX to be a
> per-inode capability and has an XFS-specific inode flag.  We can make
> that a common inode flag, but it seems we should have a way to
> interrogate the mapping itself in the case where the inode is unknown
> or unavailable.  I'm thinking extensions to mincore to have flags for
> DAX and possibly whether the page is part of a pte, pmd, or pud
> mapping.  Just floating that idea before starting to look into the
> implementation, comments or other ideas welcome...

I think this goes back to our previous discussion about support for the PMEM
programming model.  Really I think what NVML needs isn't a way to tell if it
is getting a DAX mapping, but whether it is getting a DAX mapping on a
filesystem that fully supports the PMEM programming model.  This of course is
defined to be a filesystem where it can do all of its flushes from userspace
safely and never call fsync/msync, and that allocations that happen in page
faults will be synchronized to media before the page fault completes.

IIUC this is what NVML needs - a way to decide "do I use fsync/msync for
everything or can I rely fully on flushes from userspace?" 

For all existing implementations, I think the answer is "you need to use
fsync/msync" because we don't yet have proper support for the PMEM programming
model.

My best idea of how to support this was a per-inode flag similar to the one
supported by XFS that says "you have a PMEM capable DAX mapping", which NVML
would then interpret to mean "you can do flushes from userspace and be fully
safe".  I think we really want this interface to be common over XFS and ext4.

If we can figure out a better way of doing this interface, say via mincore,
that's fine, but I don't think we can detangle this from the PMEM API
discussion.

[toc] | [prev] | [next] | [standalone]


#1479574

FromDan Williams <dan.j.williams@intel.com>
Date2016-09-09 01:10 +0200
Message-ID<sfgCd-3HW-23@gated-at.bofh.it>
In reply to#1479573
On Thu, Sep 8, 2016 at 3:56 PM, Ross Zwisler
<ross.zwisler@linux.intel.com> wrote:
> On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote:
>> [ adding linux-fsdevel and linux-nvdimm ]
>>
>> On Wed, Sep 7, 2016 at 8:36 PM, Xiao Guangrong
>> <guangrong.xiao@linux.intel.com> wrote:
>> [..]
>> > However, it is not easy to handle the case that the new VMA overlays with
>> > the old VMA
>> > already got by userspace. I think we have some choices:
>> > 1: One way is completely skipping the new VMA region as current kernel code
>> > does but i
>> >    do not think this is good as the later VMAs will be dropped.
>> >
>> > 2: show the un-overlayed portion of new VMA. In your case, we just show the
>> > region
>> >    (0x2000 -> 0x3000), however, it can not work well if the VMA is a new
>> > created
>> >    region with different attributions.
>> >
>> > 3: completely show the new VMA as this patch does.
>> >
>> > Which one do you prefer?
>> >
>>
>> I don't have a preference, but perhaps this breakage and uncertainty
>> is a good opportunity to propose a more reliable interface for NVML to
>> get the information it needs?
>>
>> My understanding is that it is looking for the VM_MIXEDMAP flag which
>> is already ambiguous for determining if DAX is enabled even if this
>> dynamic listing issue is fixed.  XFS has arranged for DAX to be a
>> per-inode capability and has an XFS-specific inode flag.  We can make
>> that a common inode flag, but it seems we should have a way to
>> interrogate the mapping itself in the case where the inode is unknown
>> or unavailable.  I'm thinking extensions to mincore to have flags for
>> DAX and possibly whether the page is part of a pte, pmd, or pud
>> mapping.  Just floating that idea before starting to look into the
>> implementation, comments or other ideas welcome...
>
> I think this goes back to our previous discussion about support for the PMEM
> programming model.  Really I think what NVML needs isn't a way to tell if it
> is getting a DAX mapping, but whether it is getting a DAX mapping on a
> filesystem that fully supports the PMEM programming model.  This of course is
> defined to be a filesystem where it can do all of its flushes from userspace
> safely and never call fsync/msync, and that allocations that happen in page
> faults will be synchronized to media before the page fault completes.
>
> IIUC this is what NVML needs - a way to decide "do I use fsync/msync for
> everything or can I rely fully on flushes from userspace?"
>
> For all existing implementations, I think the answer is "you need to use
> fsync/msync" because we don't yet have proper support for the PMEM programming
> model.
>
> My best idea of how to support this was a per-inode flag similar to the one
> supported by XFS that says "you have a PMEM capable DAX mapping", which NVML
> would then interpret to mean "you can do flushes from userspace and be fully
> safe".  I think we really want this interface to be common over XFS and ext4.
>
> If we can figure out a better way of doing this interface, say via mincore,
> that's fine, but I don't think we can detangle this from the PMEM API
> discussion.

Whether a persistent memory mapping requires an msync/fsync is a
filesystem specific question.  This mincore proposal is separate from
that.  Consider device-DAX for volatile memory or mincore() called on
an anonymous memory range.  In those cases persistence and filesystem
metadata are not in the picture, but it would still be useful for
userspace to know "is there page cache backing this mapping?" or "what
is the TLB geometry of this mapping?".

[toc] | [prev] | [next] | [standalone]


#1479767 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-09-09 11:10 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sfpYR-VU-19@gated-at.bofh.it>
In reply to#1479574

On 09/09/2016 07:04 AM, Dan Williams wrote:
> On Thu, Sep 8, 2016 at 3:56 PM, Ross Zwisler
> <ross.zwisler@linux.intel.com> wrote:
>> On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote:
>>> [ adding linux-fsdevel and linux-nvdimm ]
>>>
>>> On Wed, Sep 7, 2016 at 8:36 PM, Xiao Guangrong
>>> <guangrong.xiao@linux.intel.com> wrote:
>>> [..]
>>>> However, it is not easy to handle the case that the new VMA overlays with
>>>> the old VMA
>>>> already got by userspace. I think we have some choices:
>>>> 1: One way is completely skipping the new VMA region as current kernel code
>>>> does but i
>>>>    do not think this is good as the later VMAs will be dropped.
>>>>
>>>> 2: show the un-overlayed portion of new VMA. In your case, we just show the
>>>> region
>>>>    (0x2000 -> 0x3000), however, it can not work well if the VMA is a new
>>>> created
>>>>    region with different attributions.
>>>>
>>>> 3: completely show the new VMA as this patch does.
>>>>
>>>> Which one do you prefer?
>>>>
>>>
>>> I don't have a preference, but perhaps this breakage and uncertainty
>>> is a good opportunity to propose a more reliable interface for NVML to
>>> get the information it needs?
>>>
>>> My understanding is that it is looking for the VM_MIXEDMAP flag which
>>> is already ambiguous for determining if DAX is enabled even if this
>>> dynamic listing issue is fixed.  XFS has arranged for DAX to be a
>>> per-inode capability and has an XFS-specific inode flag.  We can make
>>> that a common inode flag, but it seems we should have a way to
>>> interrogate the mapping itself in the case where the inode is unknown
>>> or unavailable.  I'm thinking extensions to mincore to have flags for
>>> DAX and possibly whether the page is part of a pte, pmd, or pud
>>> mapping.  Just floating that idea before starting to look into the
>>> implementation, comments or other ideas welcome...
>>
>> I think this goes back to our previous discussion about support for the PMEM
>> programming model.  Really I think what NVML needs isn't a way to tell if it
>> is getting a DAX mapping, but whether it is getting a DAX mapping on a
>> filesystem that fully supports the PMEM programming model.  This of course is
>> defined to be a filesystem where it can do all of its flushes from userspace
>> safely and never call fsync/msync, and that allocations that happen in page
>> faults will be synchronized to media before the page fault completes.
>>
>> IIUC this is what NVML needs - a way to decide "do I use fsync/msync for
>> everything or can I rely fully on flushes from userspace?"
>>
>> For all existing implementations, I think the answer is "you need to use
>> fsync/msync" because we don't yet have proper support for the PMEM programming
>> model.
>>
>> My best idea of how to support this was a per-inode flag similar to the one
>> supported by XFS that says "you have a PMEM capable DAX mapping", which NVML
>> would then interpret to mean "you can do flushes from userspace and be fully
>> safe".  I think we really want this interface to be common over XFS and ext4.
>>
>> If we can figure out a better way of doing this interface, say via mincore,
>> that's fine, but I don't think we can detangle this from the PMEM API
>> discussion.
>
> Whether a persistent memory mapping requires an msync/fsync is a
> filesystem specific question.  This mincore proposal is separate from
> that.  Consider device-DAX for volatile memory or mincore() called on
> an anonymous memory range.  In those cases persistence and filesystem
> metadata are not in the picture, but it would still be useful for
> userspace to know "is there page cache backing this mapping?" or "what
> is the TLB geometry of this mapping?".

I got a question about msync/fsync which is beyond the topic of this thread :)

Whether msync/fsync can make data persistent depends on ADR feature on memory
controller, if it exists everything works well, otherwise, we need to have another
interface that is why 'Flush hint table' in ACPI comes in. 'Flush hint table' is
particularly useful for nvdimm virtualization if we use normal memory to emulate
nvdimm with data persistent characteristic (the data will be flushed to a
persistent storage, e.g, disk).

Does current PMEM programming model fully supports 'Flush hint table'? Is
userspace allowed to use these addresses?

Thanks!

[toc] | [prev] | [next] | [standalone]


#1480168

FromDan Williams <dan.j.williams@intel.com>
Date2016-09-09 17:50 +0200
Message-ID<sfwdY-4Pt-19@gated-at.bofh.it>
In reply to#1479767
On Fri, Sep 9, 2016 at 1:55 AM, Xiao Guangrong
<guangrong.xiao@linux.intel.com> wrote:
[..]
>>
>> Whether a persistent memory mapping requires an msync/fsync is a
>> filesystem specific question.  This mincore proposal is separate from
>> that.  Consider device-DAX for volatile memory or mincore() called on
>> an anonymous memory range.  In those cases persistence and filesystem
>> metadata are not in the picture, but it would still be useful for
>> userspace to know "is there page cache backing this mapping?" or "what
>> is the TLB geometry of this mapping?".
>
>
> I got a question about msync/fsync which is beyond the topic of this thread
> :)
>
> Whether msync/fsync can make data persistent depends on ADR feature on
> memory
> controller, if it exists everything works well, otherwise, we need to have
> another
> interface that is why 'Flush hint table' in ACPI comes in. 'Flush hint
> table' is
> particularly useful for nvdimm virtualization if we use normal memory to
> emulate
> nvdimm with data persistent characteristic (the data will be flushed to a
> persistent storage, e.g, disk).
>
> Does current PMEM programming model fully supports 'Flush hint table'? Is
> userspace allowed to use these addresses?

If you publish flush hint addresses in the virtual NFIT the guest VM
will write to them whenever a REQ_FLUSH or REQ_FUA request is sent to
the virtual /dev/pmemX device.  Yes, seems straightforward to take a
VM exit on those events and flush simulated pmem to persistent
storage.

[toc] | [prev] | [next] | [standalone]


#1480936 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-09-12 08:10 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgsBj-7Yg-3@gated-at.bofh.it>
In reply to#1480168

On 09/09/2016 11:40 PM, Dan Williams wrote:
> On Fri, Sep 9, 2016 at 1:55 AM, Xiao Guangrong
> <guangrong.xiao@linux.intel.com> wrote:
> [..]
>>>
>>> Whether a persistent memory mapping requires an msync/fsync is a
>>> filesystem specific question.  This mincore proposal is separate from
>>> that.  Consider device-DAX for volatile memory or mincore() called on
>>> an anonymous memory range.  In those cases persistence and filesystem
>>> metadata are not in the picture, but it would still be useful for
>>> userspace to know "is there page cache backing this mapping?" or "what
>>> is the TLB geometry of this mapping?".
>>
>>
>> I got a question about msync/fsync which is beyond the topic of this thread
>> :)
>>
>> Whether msync/fsync can make data persistent depends on ADR feature on
>> memory
>> controller, if it exists everything works well, otherwise, we need to have
>> another
>> interface that is why 'Flush hint table' in ACPI comes in. 'Flush hint
>> table' is
>> particularly useful for nvdimm virtualization if we use normal memory to
>> emulate
>> nvdimm with data persistent characteristic (the data will be flushed to a
>> persistent storage, e.g, disk).
>>
>> Does current PMEM programming model fully supports 'Flush hint table'? Is
>> userspace allowed to use these addresses?
>
> If you publish flush hint addresses in the virtual NFIT the guest VM
> will write to them whenever a REQ_FLUSH or REQ_FUA request is sent to
> the virtual /dev/pmemX device.  Yes, seems straightforward to take a
> VM exit on those events and flush simulated pmem to persistent
> storage.
>

Thank you, Dan!

However REQ_FLUSH or REQ_FUA is handled in kernel space, okay, after following
up the discussion in this thread, i understood that currently filesystems have
not supported the case that usespace itself make data be persistent without
kernel's involvement. So that works.

Hmm, Does device-DAX support this case (make data be persistent without
msync/fsync)? I guess no, but just want to confirm it. :)

[toc] | [prev] | [next] | [standalone]


#1480914 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

From"Rudoff, Andy" <andy.rudoff@intel.com>
Date2016-09-12 05:50 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgqpP-6rr-3@gated-at.bofh.it>
In reply to#1479767
>Whether msync/fsync can make data persistent depends on ADR feature on
>memory controller, if it exists everything works well, otherwise, we need
>to have another interface that is why 'Flush hint table' in ACPI comes
>in. 'Flush hint table' is particularly useful for nvdimm virtualization if
>we use normal memory to emulate nvdimm with data persistent characteristic
>(the data will be flushed to a persistent storage, e.g, disk).
>
>Does current PMEM programming model fully supports 'Flush hint table'? Is
>userspace allowed to use these addresses?

The Flush hint table is NOT a replacement for ADR.  To support pmem on
the x86 architecture, the platform is required to ensure that a pmem
store flushed from the CPU caches is in the persistent domain so that the
application need not take any additional steps to make it persistent.
The most common way to do this is the ADR feature.

If the above is not true, then your x86 platform does not support pmem.

Flush hints are for use by the BIOS and drivers and are not intended to
be used in user space.  Flush hints provide two things:

First, if a driver needs to write to command registers or movable windows
on a DIMM, the Flush hint (if provided in the NFIT) is required to flush
the command to the DIMM or ensure stores done through the movable window
are complete before moving it somewhere else.

Second, for the rare case where the kernel wants to flush stores to the
smallest possible failure domain (i.e. to the DIMM even though ADR will
handle flushing it from a larger domain), the flush hints provide a way
to do this.  This might be useful for things like file system journals to
help ensure the file system is consistent even in the face of ADR failure.

-andy


[toc] | [prev] | [next] | [standalone]


#1480952 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromXiao Guangrong <guangrong.xiao@linux.intel.com>
Date2016-09-12 08:40 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgt4l-88g-11@gated-at.bofh.it>
In reply to#1480914

On 09/12/2016 11:44 AM, Rudoff, Andy wrote:
>> Whether msync/fsync can make data persistent depends on ADR feature on
>> memory controller, if it exists everything works well, otherwise, we need
>> to have another interface that is why 'Flush hint table' in ACPI comes
>> in. 'Flush hint table' is particularly useful for nvdimm virtualization if
>> we use normal memory to emulate nvdimm with data persistent characteristic
>> (the data will be flushed to a persistent storage, e.g, disk).
>>
>> Does current PMEM programming model fully supports 'Flush hint table'? Is
>> userspace allowed to use these addresses?
>
> The Flush hint table is NOT a replacement for ADR.  To support pmem on
> the x86 architecture, the platform is required to ensure that a pmem
> store flushed from the CPU caches is in the persistent domain so that the
> application need not take any additional steps to make it persistent.
> The most common way to do this is the ADR feature.
>
> If the above is not true, then your x86 platform does not support pmem.

Understood.

However, virtualization is a special case as we can use normal memory
to emulate NVDIMM for the vm so that vm can bypass local file-cache,
reduce memory usage and io path, etc. Currently, this usage is useful
for lightweight virtualization, such as clean container.

Under this case, ADR is available on physical platform but it can
not help us to make data persistence for the vm. So that virtualizeing
'flush hint table' is a good way to handle it based on the acpi spec:
| software can write to any one of these Flush Hint Addresses to
| cause any preceding writes to the NVDIMM region to be flushed
| out of the intervening platform buffers 1 to the targeted NVDIMM
| (to achieve durability)

>
> Flush hints are for use by the BIOS and drivers and are not intended to
> be used in user space.  Flush hints provide two things:
>
> First, if a driver needs to write to command registers or movable windows
> on a DIMM, the Flush hint (if provided in the NFIT) is required to flush
> the command to the DIMM or ensure stores done through the movable window
> are complete before moving it somewhere else.
>
> Second, for the rare case where the kernel wants to flush stores to the
> smallest possible failure domain (i.e. to the DIMM even though ADR will
> handle flushing it from a larger domain), the flush hints provide a way
> to do this.  This might be useful for things like file system journals to
> help ensure the file system is consistent even in the face of ADR failure.

We are assuming ADR can fail, however, do we have a way to know whether
ADR works correctly? Maybe MCE can work on it?

This is necessary to support making data persistent without 'fsync/msync'
in userspace. Or do we need to unconditionally use 'flush hint address'
if it is available as current nvdimm driver does?

Thanks!

[toc] | [prev] | [next] | [standalone]


#1480884 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromDave Chinner <david@fromorbit.com>
Date2016-09-12 03:50 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgoxI-5f9-7@gated-at.bofh.it>
In reply to#1479573
On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote:
> On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote:
> > My understanding is that it is looking for the VM_MIXEDMAP flag which
> > is already ambiguous for determining if DAX is enabled even if this
> > dynamic listing issue is fixed.  XFS has arranged for DAX to be a
> > per-inode capability and has an XFS-specific inode flag.  We can make
> > that a common inode flag, but it seems we should have a way to
> > interrogate the mapping itself in the case where the inode is unknown
> > or unavailable.  I'm thinking extensions to mincore to have flags for
> > DAX and possibly whether the page is part of a pte, pmd, or pud
> > mapping.  Just floating that idea before starting to look into the
> > implementation, comments or other ideas welcome...
> 
> I think this goes back to our previous discussion about support for the PMEM
> programming model.  Really I think what NVML needs isn't a way to tell if it
> is getting a DAX mapping, but whether it is getting a DAX mapping on a
> filesystem that fully supports the PMEM programming model.  This of course is
> defined to be a filesystem where it can do all of its flushes from userspace
> safely and never call fsync/msync, and that allocations that happen in page
> faults will be synchronized to media before the page fault completes.
> 
> IIUC this is what NVML needs - a way to decide "do I use fsync/msync for
> everything or can I rely fully on flushes from userspace?" 

"need fsync/msync" is a dynamic state of an inode, not a static
property. i.e. users can do things that change an inode behind the
back of a mapping, even if they are not aware that this might
happen. As such, a filesystem can invalidate an existing mapping
at any time and userspace won't notice because it will simply fault
in a new mapping on the next access...

> For all existing implementations, I think the answer is "you need to use
> fsync/msync" because we don't yet have proper support for the PMEM programming
> model.

Yes, that is correct.

FWIW, I don't think it will ever be possible to support this ....
wonderful "PMEM programming model" from any current or future kernel
filesystem without a very specific set of restrictions on what can
be done to a file.  e.g.

	1. the file has to be fully allocated and zeroed before
	   use. Preallocation/zeroing via unwritten extents is not
	   allowed. Sparse files are not allowed. Shared extents are
	   not allowed.
	2. set the "PMEM_IMMUTABLE" inode flag - filesystem must
	   check the file is fully allocated before allowing it to
	   be set, and caller must have CAP_LINUX_IMMUTABLE.
	3. Inode metadata is now immutable, and file data can only
	   be accessed and/or modified via mmap().
	4. All non-mmap methods of inode data modification
	   will now fail with EPERM.
	5. all methods of inode metadata modification will now fail
	   with EPERM, timestamp udpdates will be ignored.
	6. PMEM_IMMUTABLE flag can only be removed if the file is
	   not currently mapped and caller has CAP_LINUX_IMMUTABLE.

A flag like this /should/ make it possible to avoid fsync/msync() on
a file for existing filesystems, but it also means that such files
have significant management issues (hence the need for
CAP_LINUX_IMMUTABLE to cover it's use).

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1483859 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

From"Darrick J. Wong" <darrick.wong@oracle.com>
Date2016-09-15 08:00 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<shxSh-2ZY-1@gated-at.bofh.it>
In reply to#1480884
On Mon, Sep 12, 2016 at 11:40:35AM +1000, Dave Chinner wrote:
> On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote:
> > On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote:
> > > My understanding is that it is looking for the VM_MIXEDMAP flag which
> > > is already ambiguous for determining if DAX is enabled even if this
> > > dynamic listing issue is fixed.  XFS has arranged for DAX to be a
> > > per-inode capability and has an XFS-specific inode flag.  We can make
> > > that a common inode flag, but it seems we should have a way to
> > > interrogate the mapping itself in the case where the inode is unknown
> > > or unavailable.  I'm thinking extensions to mincore to have flags for
> > > DAX and possibly whether the page is part of a pte, pmd, or pud
> > > mapping.  Just floating that idea before starting to look into the
> > > implementation, comments or other ideas welcome...
> > 
> > I think this goes back to our previous discussion about support for the PMEM
> > programming model.  Really I think what NVML needs isn't a way to tell if it
> > is getting a DAX mapping, but whether it is getting a DAX mapping on a
> > filesystem that fully supports the PMEM programming model.  This of course is
> > defined to be a filesystem where it can do all of its flushes from userspace
> > safely and never call fsync/msync, and that allocations that happen in page
> > faults will be synchronized to media before the page fault completes.
> > 
> > IIUC this is what NVML needs - a way to decide "do I use fsync/msync for
> > everything or can I rely fully on flushes from userspace?" 
> 
> "need fsync/msync" is a dynamic state of an inode, not a static
> property. i.e. users can do things that change an inode behind the
> back of a mapping, even if they are not aware that this might
> happen. As such, a filesystem can invalidate an existing mapping
> at any time and userspace won't notice because it will simply fault
> in a new mapping on the next access...
> 
> > For all existing implementations, I think the answer is "you need to use
> > fsync/msync" because we don't yet have proper support for the PMEM programming
> > model.
> 
> Yes, that is correct.
> 
> FWIW, I don't think it will ever be possible to support this ....
> wonderful "PMEM programming model" from any current or future kernel
> filesystem without a very specific set of restrictions on what can
> be done to a file.  e.g.
> 
> 	1. the file has to be fully allocated and zeroed before
> 	   use. Preallocation/zeroing via unwritten extents is not
> 	   allowed. Sparse files are not allowed. Shared extents are
> 	   not allowed.
> 	2. set the "PMEM_IMMUTABLE" inode flag - filesystem must
> 	   check the file is fully allocated before allowing it to
> 	   be set, and caller must have CAP_LINUX_IMMUTABLE.
> 	3. Inode metadata is now immutable, and file data can only
> 	   be accessed and/or modified via mmap().
> 	4. All non-mmap methods of inode data modification
> 	   will now fail with EPERM.
> 	5. all methods of inode metadata modification will now fail
> 	   with EPERM, timestamp udpdates will be ignored.
> 	6. PMEM_IMMUTABLE flag can only be removed if the file is
> 	   not currently mapped and caller has CAP_LINUX_IMMUTABLE.
> 
> A flag like this /should/ make it possible to avoid fsync/msync() on
> a file for existing filesystems, but it also means that such files
> have significant management issues (hence the need for
> CAP_LINUX_IMMUTABLE to cover it's use).

Hmmm... I started to ponder such a flag, but ran into some questions.
If it's PMEM_IMMUTABLE, does this mean that none of 1-6 apply if the
filesystem discovers it isn't on pmem?

I thought about just having a 'immutable metadata' flag where any
timestamp, xattr, or block mapping update just returns EPERM.  There
wouldn't be any checks as in (1); if you left a hole in the file prior
to setting the flag then you won't be filling it unless you clear the
flag.  OTOH if it merely made the metadata unchangeable then it's a
stretch to get to non-mmap data accesses also being disallowed.

Maybe the immutable metadata and mmap-only properties would only be
implied if both DAX and IMMUTABLE_META are set on a file?

Ok no more rambling until sleep. :)

--D

> 
> Cheers,
> 
> Dave.
> -- 
> Dave Chinner
> david@fromorbit.com
> --
> To unsubscribe from this list: send the line "unsubscribe linux-fsdevel" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html

[toc] | [prev] | [next] | [standalone]


#1483869 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromDave Chinner <david@fromorbit.com>
Date2016-09-15 08:30 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<shylj-3ql-15@gated-at.bofh.it>
In reply to#1483859
On Wed, Sep 14, 2016 at 10:55:03PM -0700, Darrick J. Wong wrote:
> On Mon, Sep 12, 2016 at 11:40:35AM +1000, Dave Chinner wrote:
> > On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote:
> > > On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote:
> > > > My understanding is that it is looking for the VM_MIXEDMAP flag which
> > > > is already ambiguous for determining if DAX is enabled even if this
> > > > dynamic listing issue is fixed.  XFS has arranged for DAX to be a
> > > > per-inode capability and has an XFS-specific inode flag.  We can make
> > > > that a common inode flag, but it seems we should have a way to
> > > > interrogate the mapping itself in the case where the inode is unknown
> > > > or unavailable.  I'm thinking extensions to mincore to have flags for
> > > > DAX and possibly whether the page is part of a pte, pmd, or pud
> > > > mapping.  Just floating that idea before starting to look into the
> > > > implementation, comments or other ideas welcome...
> > > 
> > > I think this goes back to our previous discussion about support for the PMEM
> > > programming model.  Really I think what NVML needs isn't a way to tell if it
> > > is getting a DAX mapping, but whether it is getting a DAX mapping on a
> > > filesystem that fully supports the PMEM programming model.  This of course is
> > > defined to be a filesystem where it can do all of its flushes from userspace
> > > safely and never call fsync/msync, and that allocations that happen in page
> > > faults will be synchronized to media before the page fault completes.
> > > 
> > > IIUC this is what NVML needs - a way to decide "do I use fsync/msync for
> > > everything or can I rely fully on flushes from userspace?" 
> > 
> > "need fsync/msync" is a dynamic state of an inode, not a static
> > property. i.e. users can do things that change an inode behind the
> > back of a mapping, even if they are not aware that this might
> > happen. As such, a filesystem can invalidate an existing mapping
> > at any time and userspace won't notice because it will simply fault
> > in a new mapping on the next access...
> > 
> > > For all existing implementations, I think the answer is "you need to use
> > > fsync/msync" because we don't yet have proper support for the PMEM programming
> > > model.
> > 
> > Yes, that is correct.
> > 
> > FWIW, I don't think it will ever be possible to support this ....
> > wonderful "PMEM programming model" from any current or future kernel
> > filesystem without a very specific set of restrictions on what can
> > be done to a file.  e.g.
> > 
> > 	1. the file has to be fully allocated and zeroed before
> > 	   use. Preallocation/zeroing via unwritten extents is not
> > 	   allowed. Sparse files are not allowed. Shared extents are
> > 	   not allowed.
> > 	2. set the "PMEM_IMMUTABLE" inode flag - filesystem must
> > 	   check the file is fully allocated before allowing it to
> > 	   be set, and caller must have CAP_LINUX_IMMUTABLE.
> > 	3. Inode metadata is now immutable, and file data can only
> > 	   be accessed and/or modified via mmap().
> > 	4. All non-mmap methods of inode data modification
> > 	   will now fail with EPERM.
> > 	5. all methods of inode metadata modification will now fail
> > 	   with EPERM, timestamp udpdates will be ignored.
> > 	6. PMEM_IMMUTABLE flag can only be removed if the file is
> > 	   not currently mapped and caller has CAP_LINUX_IMMUTABLE.
> > 
> > A flag like this /should/ make it possible to avoid fsync/msync() on
> > a file for existing filesystems, but it also means that such files
> > have significant management issues (hence the need for
> > CAP_LINUX_IMMUTABLE to cover it's use).
> 
> Hmmm... I started to ponder such a flag, but ran into some questions.
> If it's PMEM_IMMUTABLE, does this mean that none of 1-6 apply if the
> filesystem discovers it isn't on pmem?

Would only be meaningful if the FS_XFLAG_DAX/S_DAX flag is also set
on the inode and the backing store is dax capable. Hence the 'PMEM'
part of the name.

> I thought about just having a 'immutable metadata' flag where any
> timestamp, xattr, or block mapping update just returns EPERM.

And all the rest - no hard links, no perm/owner changes, no security
context changes(!), and so on. ANd it's even more complex with
filesystems that have COW metadata and pack multiple unrelated
metadata objects into single blocks - they can do all sorts of
interesting things on unrealted metadata updates... :P

You'd also have to turn off background internal filesystem mod
vectors, too, like EOF scanning, or defrag, balance, dedupe,
auto-repair, etc.  And, now that I think about it, snapshots are out
of the question too.

This gets more hairy the more I think about what our filesystems can
do these days....

> There
> wouldn't be any checks as in (1); if you left a hole in the file prior
> to setting the flag then you won't be filling it unless you clear the
> flag.

Which means writing into a hole would need to return an error, and a
write page fault into a hole would need a segv. Seems like a great
way to cause random application failures to me...

> OTOH if it merely made the metadata unchangeable then it's a
> stretch to get to non-mmap data accesses also being disallowed.

*nod*

> Maybe the immutable metadata and mmap-only properties would only be
> implied if both DAX and IMMUTABLE_META are set on a file?

I'd suggest that PMEM_IMMUTABLE could only be set on an inode that
already has the FS_XFLAG_DAX set on it (or it is being set at the
same time). And clearing the DAX flag would also remove the
PMEM_IMMUTABLE flag. Perhaps it would be better to call it
FS_XFLAG_DAX_IMMUTABLE rather than anything pmem related.

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


#1480928 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromChristoph Hellwig <hch@infradead.org>
Date2016-09-12 07:30 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgrYB-7wi-3@gated-at.bofh.it>
In reply to#1479573
On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote:
> I think this goes back to our previous discussion about support for the PMEM
> programming model.  Really I think what NVML needs isn't a way to tell if it
> is getting a DAX mapping, but whether it is getting a DAX mapping on a
> filesystem that fully supports the PMEM programming model.  This of course is
> defined to be a filesystem where it can do all of its flushes from userspace
> safely and never call fsync/msync, and that allocations that happen in page
> faults will be synchronized to media before the page fault completes.

That's a an easy way to flag:  you will never get that from a Linux
filesystem, period.

NVML folks really need to stop taking crack and dreaming this could
happen.

[toc] | [prev] | [next] | [standalone]


#1480965

From"Oliver O'Halloran" <oohall@gmail.com>
Date2016-09-12 09:30 +0200
Message-ID<sgtQJ-ct-1@gated-at.bofh.it>
In reply to#1480928
On Mon, Sep 12, 2016 at 3:27 PM, Christoph Hellwig <hch@infradead.org> wrote:
> On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote:
>> I think this goes back to our previous discussion about support for the PMEM
>> programming model.  Really I think what NVML needs isn't a way to tell if it
>> is getting a DAX mapping, but whether it is getting a DAX mapping on a
>> filesystem that fully supports the PMEM programming model.  This of course is
>> defined to be a filesystem where it can do all of its flushes from userspace
>> safely and never call fsync/msync, and that allocations that happen in page
>> faults will be synchronized to media before the page fault completes.
>
> That's a an easy way to flag:  you will never get that from a Linux
> filesystem, period.
>
> NVML folks really need to stop taking crack and dreaming this could
> happen.

Well, that's a bummer.

What are the problems here? Is this a matter of existing filesystems
being unable/unwilling to support this or is it just fundamentally
broken? The end goal is to let applications manage the persistence of
their own data without having to involve the kernel in every IOP, but
if we can't do that then what would a 90% solution look like? I think
most people would be OK with having to do an fsync() occasionally, but
not after ever write to pmem.

[toc] | [prev] | [next] | [standalone]


#1481000 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromChristoph Hellwig <hch@infradead.org>
Date2016-09-12 10:00 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgujL-mJ-3@gated-at.bofh.it>
In reply to#1480965
On Mon, Sep 12, 2016 at 05:25:15PM +1000, Oliver O'Halloran wrote:
> What are the problems here? Is this a matter of existing filesystems
> being unable/unwilling to support this or is it just fundamentally
> broken?

It's a fundamentally broken model.  See Dave's post that actually was
sent slightly earlier then mine for the list of required items, which
is fairly unrealistic.  You could probably try to architect a file
system for it, but I doubt it would gain much traction.

> The end goal is to let applications manage the persistence of
> their own data without having to involve the kernel in every IOP, but
> if we can't do that then what would a 90% solution look like? I think
> most people would be OK with having to do an fsync() occasionally, but
> not after ever write to pmem.

You need an fsync for each write that you want to persist.  This sounds
painful for now.  But I have an implementation that will allow the
atomic commit of more or less arbitrary amounts of previous writes for
XFS that I plan to land once the reflink work is in.

That way you create almost arbitrarily complex data structures in your
programs and commit them atomicly.  It's not going to fit the nvml
model, but that whole think has been complete bullshit since the
beginning anyway.

[toc] | [prev] | [next] | [standalone]


#1481021 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromNicholas Piggin <npiggin@gmail.com>
Date2016-09-12 10:10 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgutt-Fd-39@gated-at.bofh.it>
In reply to#1481000
On Mon, 12 Sep 2016 00:51:28 -0700
Christoph Hellwig <hch@infradead.org> wrote:

> On Mon, Sep 12, 2016 at 05:25:15PM +1000, Oliver O'Halloran wrote:
> > What are the problems here? Is this a matter of existing filesystems
> > being unable/unwilling to support this or is it just fundamentally
> > broken?  
> 
> It's a fundamentally broken model.  See Dave's post that actually was
> sent slightly earlier then mine for the list of required items, which
> is fairly unrealistic.  You could probably try to architect a file
> system for it, but I doubt it would gain much traction.

It's not fundamentally broken, it just doesn't fit well existing
filesystems.

Dave's post of requirements is also wrong. A filesystem does not have
to guarantee all that, it only has to guarantee that is the case for
a given block after it has a mapping and page fault returns, other
operations can be supported by invalidating mappings, etc.

[toc] | [prev] | [next] | [standalone]


#1481395 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromChristoph Hellwig <hch@infradead.org>
Date2016-09-12 17:10 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgB1U-4WE-23@gated-at.bofh.it>
In reply to#1481021
On Mon, Sep 12, 2016 at 06:05:07PM +1000, Nicholas Piggin wrote:
> It's not fundamentally broken, it just doesn't fit well existing
> filesystems.

Or the existing file system architecture for that matter.  Which makes
it a fundamentally broken model.

> Dave's post of requirements is also wrong. A filesystem does not have
> to guarantee all that, it only has to guarantee that is the case for
> a given block after it has a mapping and page fault returns, other
> operations can be supported by invalidating mappings, etc.

Which doesn't really matter if your use case is manipulating
fully mapped files.

But back to the point: if you want to use a full blown Linux or Unix
filesystem you will always have to fsync (or variants of it like msync),
period.

If you want a volume manager on stereoids that hands out large chunks
of storage memory that can't ever be moved, truncated, shared, allocated
on demand, etc - implement it in your library on top of a device file.

[toc] | [prev] | [next] | [standalone]


#1482080 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromNicholas Piggin <npiggin@gmail.com>
Date2016-09-13 03:40 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgKRz-39a-1@gated-at.bofh.it>
In reply to#1481395
On Mon, 12 Sep 2016 08:01:48 -0700
Christoph Hellwig <hch@infradead.org> wrote:

> On Mon, Sep 12, 2016 at 06:05:07PM +1000, Nicholas Piggin wrote:
> > It's not fundamentally broken, it just doesn't fit well existing
> > filesystems.  
> 
> Or the existing file system architecture for that matter.  Which makes
> it a fundamentally broken model.

Not really. A few reasonable changes can be made to improve things.
Until just now you thought it was fundamentally impossible to make a
reasonable implementation due to Dave's "constraints".

> 
> > Dave's post of requirements is also wrong. A filesystem does not have
> > to guarantee all that, it only has to guarantee that is the case for
> > a given block after it has a mapping and page fault returns, other
> > operations can be supported by invalidating mappings, etc.  
> 
> Which doesn't really matter if your use case is manipulating
> fully mapped files.

Nothing that says you have to use them fully mapped always and not
use other APIs on them.


> But back to the point: if you want to use a full blown Linux or Unix
> filesystem you will always have to fsync (or variants of it like msync),
> period.

That's circular logic. First you said that should not be done
because of your imagined constraints.

In fact, it's not unreasonable to describe some additional semantics
of the storage that is unavailable with traditional filesystems.

That said, a noop system call is on the order of 100 cycles nowadays,
so rushing to implement these APIs without seeing good numbers and
actual users ready to go seems premature. *This* is the real reason
not to implement new APIs yet.


> If you want a volume manager on stereoids that hands out large chunks
> of storage memory that can't ever be moved, truncated, shared, allocated
> on demand, etc - implement it in your library on top of a device file.

Those constraints don't exist either. I've written a filesystem
that avoids them. It isn't rocket science.

[toc] | [prev] | [next] | [standalone]


#1482131

FromDan Williams <dan.j.williams@intel.com>
Date2016-09-13 06:10 +0200
Message-ID<sgNcJ-4Tl-7@gated-at.bofh.it>
In reply to#1482080
On Mon, Sep 12, 2016 at 6:31 PM, Nicholas Piggin <npiggin@gmail.com> wrote:
> On Mon, 12 Sep 2016 08:01:48 -0700
[..]
> That said, a noop system call is on the order of 100 cycles nowadays,
> so rushing to implement these APIs without seeing good numbers and
> actual users ready to go seems premature. *This* is the real reason
> not to implement new APIs yet.

Yes, and harvesting the current crop of low hanging performance fruit
in the filesystem-DAX I/O path remains on the todo list.

In the meantime we're pursuing this mm api, mincore+ or whatever we
end up with, to allow userspace to distinguish memory address ranges
that are backed by a filesystem requiring coordination of metadata
updates + flushes for updates, vs something like device-dax that does
not.

[toc] | [prev] | [next] | [standalone]


#1482148 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromNicholas Piggin <npiggin@gmail.com>
Date2016-09-13 07:50 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgOLv-5S2-11@gated-at.bofh.it>
In reply to#1482131
On Mon, 12 Sep 2016 21:06:49 -0700
Dan Williams <dan.j.williams@intel.com> wrote:

> On Mon, Sep 12, 2016 at 6:31 PM, Nicholas Piggin <npiggin@gmail.com> wrote:
> > On Mon, 12 Sep 2016 08:01:48 -0700  
> [..]
> > That said, a noop system call is on the order of 100 cycles nowadays,
> > so rushing to implement these APIs without seeing good numbers and
> > actual users ready to go seems premature. *This* is the real reason
> > not to implement new APIs yet.  
> 
> Yes, and harvesting the current crop of low hanging performance fruit
> in the filesystem-DAX I/O path remains on the todo list.
> 
> In the meantime we're pursuing this mm api, mincore+ or whatever we
> end up with, to allow userspace to distinguish memory address ranges
> that are backed by a filesystem requiring coordination of metadata
> updates + flushes for updates, vs something like device-dax that does
> not.

Yes, that's reasonable.

Do you need page/block granularity? Do you need a way to advise/request
the fs for a particular capability? Is it enough to request and check
success? Would the capability be likely to change, and if so, how would
you notify the app asynchronously?

[toc] | [prev] | [next] | [standalone]


#1481993 — Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)

FromDave Chinner <david@fromorbit.com>
Date2016-09-12 23:40 +0200
SubjectRe: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps)
Message-ID<sgH7j-FR-21@gated-at.bofh.it>
In reply to#1481021
On Mon, Sep 12, 2016 at 06:05:07PM +1000, Nicholas Piggin wrote:
> On Mon, 12 Sep 2016 00:51:28 -0700
> Christoph Hellwig <hch@infradead.org> wrote:
> 
> > On Mon, Sep 12, 2016 at 05:25:15PM +1000, Oliver O'Halloran wrote:
> > > What are the problems here? Is this a matter of existing filesystems
> > > being unable/unwilling to support this or is it just fundamentally
> > > broken?  
> > 
> > It's a fundamentally broken model.  See Dave's post that actually was
> > sent slightly earlier then mine for the list of required items, which
> > is fairly unrealistic.  You could probably try to architect a file
> > system for it, but I doubt it would gain much traction.
> 
> It's not fundamentally broken, it just doesn't fit well existing
> filesystems.
> 
> Dave's post of requirements is also wrong. A filesystem does not have
> to guarantee all that, it only has to guarantee that is the case for
> a given block after it has a mapping and page fault returns, other
> operations can be supported by invalidating mappings, etc.

Sure, but filesystems are completely unaware of what is mapped at
any given time, or what constraints that mapping might have. Trying
to make filesystems aware of per-page mapping constraints seems like
a fairly significant layering violation based on a flawed
assumption. i.e. that operations on other parts of the file do not
affect the block that requires immutable metadata.

e.g an extent operation in some other area of the file can cause a
tip-to-root extent tree split or merge, and that moves the metadata
that points to the mapped block that we've told userspace "doesn't
need fsync".  We now need an fsync to ensure that the metadata is
consistent on disk again, even though that block has not physically
been moved. IOWs, the immutable data block updates are now not
ordered correctly w.r.t. other updates done to the file, especially
when we consider crash recovery....

All this will expose is an unfixable problem with ordering of stable
data + metadata operations and their synchronisation. As such, it
seems like nothing but a major cluster-fuck to try to do mapping
specific, per-block immutable metadata - it adds major complexity
and even more untractable problems.

Yes, we /could/ try to solve this but, quite frankly, it's far
easier to change the broken PMEM programming model assumptions than
it is to implement what you are suggesting. Or to do what Christoph
suggested and just use a wrapper around something like device
mapper to hand out chunks of unchanging, static pmem to
applications...

Cheers,

Dave.
-- 
Dave Chinner
david@fromorbit.com

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.kernel


csiph-web