Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1478773 > unrolled thread
| Started by | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| First post | 2016-09-08 06:40 +0200 |
| Last post | 2016-09-16 08:00 +0200 |
| Articles | 20 on this page of 31 — 9 participants |
Back to article view | Back to linux.kernel
DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dan Williams <dan.j.williams@intel.com> - 2016-09-08 06:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Ross Zwisler <ross.zwisler@linux.intel.com> - 2016-09-09 01:00 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dan Williams <dan.j.williams@intel.com> - 2016-09-09 01:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-09-09 11:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dan Williams <dan.j.williams@intel.com> - 2016-09-09 17:50 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-09-12 08:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) "Rudoff, Andy" <andy.rudoff@intel.com> - 2016-09-12 05:50 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Xiao Guangrong <guangrong.xiao@linux.intel.com> - 2016-09-12 08:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-12 03:50 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) "Darrick J. Wong" <darrick.wong@oracle.com> - 2016-09-15 08:00 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-15 08:30 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-12 07:30 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) "Oliver O'Halloran" <oohall@gmail.com> - 2016-09-12 09:30 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-12 10:00 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-12 10:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-12 17:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 03:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dan Williams <dan.j.williams@intel.com> - 2016-09-13 06:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 07:50 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-12 23:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 04:00 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-13 09:20 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 11:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-14 09:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-14 12:20 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-15 04:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-15 06:00 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-15 12:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-15 13:50 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-16 00:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-16 08:00 +0200
Page 1 of 2 [1] 2 Next page →
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-09-08 06:40 +0200 |
| Subject | DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <seZi1-15I-15@gated-at.bofh.it> |
[ adding linux-fsdevel and linux-nvdimm ] On Wed, Sep 7, 2016 at 8:36 PM, Xiao Guangrong <guangrong.xiao@linux.intel.com> wrote: [..] > However, it is not easy to handle the case that the new VMA overlays with > the old VMA > already got by userspace. I think we have some choices: > 1: One way is completely skipping the new VMA region as current kernel code > does but i > do not think this is good as the later VMAs will be dropped. > > 2: show the un-overlayed portion of new VMA. In your case, we just show the > region > (0x2000 -> 0x3000), however, it can not work well if the VMA is a new > created > region with different attributions. > > 3: completely show the new VMA as this patch does. > > Which one do you prefer? > I don't have a preference, but perhaps this breakage and uncertainty is a good opportunity to propose a more reliable interface for NVML to get the information it needs? My understanding is that it is looking for the VM_MIXEDMAP flag which is already ambiguous for determining if DAX is enabled even if this dynamic listing issue is fixed. XFS has arranged for DAX to be a per-inode capability and has an XFS-specific inode flag. We can make that a common inode flag, but it seems we should have a way to interrogate the mapping itself in the case where the inode is unknown or unavailable. I'm thinking extensions to mincore to have flags for DAX and possibly whether the page is part of a pte, pmd, or pud mapping. Just floating that idea before starting to look into the implementation, comments or other ideas welcome...
[toc] | [next] | [standalone]
| From | Ross Zwisler <ross.zwisler@linux.intel.com> |
|---|---|
| Date | 2016-09-09 01:00 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sfgsz-3pz-85@gated-at.bofh.it> |
| In reply to | #1478773 |
On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote: > [ adding linux-fsdevel and linux-nvdimm ] > > On Wed, Sep 7, 2016 at 8:36 PM, Xiao Guangrong > <guangrong.xiao@linux.intel.com> wrote: > [..] > > However, it is not easy to handle the case that the new VMA overlays with > > the old VMA > > already got by userspace. I think we have some choices: > > 1: One way is completely skipping the new VMA region as current kernel code > > does but i > > do not think this is good as the later VMAs will be dropped. > > > > 2: show the un-overlayed portion of new VMA. In your case, we just show the > > region > > (0x2000 -> 0x3000), however, it can not work well if the VMA is a new > > created > > region with different attributions. > > > > 3: completely show the new VMA as this patch does. > > > > Which one do you prefer? > > > > I don't have a preference, but perhaps this breakage and uncertainty > is a good opportunity to propose a more reliable interface for NVML to > get the information it needs? > > My understanding is that it is looking for the VM_MIXEDMAP flag which > is already ambiguous for determining if DAX is enabled even if this > dynamic listing issue is fixed. XFS has arranged for DAX to be a > per-inode capability and has an XFS-specific inode flag. We can make > that a common inode flag, but it seems we should have a way to > interrogate the mapping itself in the case where the inode is unknown > or unavailable. I'm thinking extensions to mincore to have flags for > DAX and possibly whether the page is part of a pte, pmd, or pud > mapping. Just floating that idea before starting to look into the > implementation, comments or other ideas welcome... I think this goes back to our previous discussion about support for the PMEM programming model. Really I think what NVML needs isn't a way to tell if it is getting a DAX mapping, but whether it is getting a DAX mapping on a filesystem that fully supports the PMEM programming model. This of course is defined to be a filesystem where it can do all of its flushes from userspace safely and never call fsync/msync, and that allocations that happen in page faults will be synchronized to media before the page fault completes. IIUC this is what NVML needs - a way to decide "do I use fsync/msync for everything or can I rely fully on flushes from userspace?" For all existing implementations, I think the answer is "you need to use fsync/msync" because we don't yet have proper support for the PMEM programming model. My best idea of how to support this was a per-inode flag similar to the one supported by XFS that says "you have a PMEM capable DAX mapping", which NVML would then interpret to mean "you can do flushes from userspace and be fully safe". I think we really want this interface to be common over XFS and ext4. If we can figure out a better way of doing this interface, say via mincore, that's fine, but I don't think we can detangle this from the PMEM API discussion.
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-09-09 01:10 +0200 |
| Message-ID | <sfgCd-3HW-23@gated-at.bofh.it> |
| In reply to | #1479573 |
On Thu, Sep 8, 2016 at 3:56 PM, Ross Zwisler <ross.zwisler@linux.intel.com> wrote: > On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote: >> [ adding linux-fsdevel and linux-nvdimm ] >> >> On Wed, Sep 7, 2016 at 8:36 PM, Xiao Guangrong >> <guangrong.xiao@linux.intel.com> wrote: >> [..] >> > However, it is not easy to handle the case that the new VMA overlays with >> > the old VMA >> > already got by userspace. I think we have some choices: >> > 1: One way is completely skipping the new VMA region as current kernel code >> > does but i >> > do not think this is good as the later VMAs will be dropped. >> > >> > 2: show the un-overlayed portion of new VMA. In your case, we just show the >> > region >> > (0x2000 -> 0x3000), however, it can not work well if the VMA is a new >> > created >> > region with different attributions. >> > >> > 3: completely show the new VMA as this patch does. >> > >> > Which one do you prefer? >> > >> >> I don't have a preference, but perhaps this breakage and uncertainty >> is a good opportunity to propose a more reliable interface for NVML to >> get the information it needs? >> >> My understanding is that it is looking for the VM_MIXEDMAP flag which >> is already ambiguous for determining if DAX is enabled even if this >> dynamic listing issue is fixed. XFS has arranged for DAX to be a >> per-inode capability and has an XFS-specific inode flag. We can make >> that a common inode flag, but it seems we should have a way to >> interrogate the mapping itself in the case where the inode is unknown >> or unavailable. I'm thinking extensions to mincore to have flags for >> DAX and possibly whether the page is part of a pte, pmd, or pud >> mapping. Just floating that idea before starting to look into the >> implementation, comments or other ideas welcome... > > I think this goes back to our previous discussion about support for the PMEM > programming model. Really I think what NVML needs isn't a way to tell if it > is getting a DAX mapping, but whether it is getting a DAX mapping on a > filesystem that fully supports the PMEM programming model. This of course is > defined to be a filesystem where it can do all of its flushes from userspace > safely and never call fsync/msync, and that allocations that happen in page > faults will be synchronized to media before the page fault completes. > > IIUC this is what NVML needs - a way to decide "do I use fsync/msync for > everything or can I rely fully on flushes from userspace?" > > For all existing implementations, I think the answer is "you need to use > fsync/msync" because we don't yet have proper support for the PMEM programming > model. > > My best idea of how to support this was a per-inode flag similar to the one > supported by XFS that says "you have a PMEM capable DAX mapping", which NVML > would then interpret to mean "you can do flushes from userspace and be fully > safe". I think we really want this interface to be common over XFS and ext4. > > If we can figure out a better way of doing this interface, say via mincore, > that's fine, but I don't think we can detangle this from the PMEM API > discussion. Whether a persistent memory mapping requires an msync/fsync is a filesystem specific question. This mincore proposal is separate from that. Consider device-DAX for volatile memory or mincore() called on an anonymous memory range. In those cases persistence and filesystem metadata are not in the picture, but it would still be useful for userspace to know "is there page cache backing this mapping?" or "what is the TLB geometry of this mapping?".
[toc] | [prev] | [next] | [standalone]
| From | Xiao Guangrong <guangrong.xiao@linux.intel.com> |
|---|---|
| Date | 2016-09-09 11:10 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sfpYR-VU-19@gated-at.bofh.it> |
| In reply to | #1479574 |
On 09/09/2016 07:04 AM, Dan Williams wrote: > On Thu, Sep 8, 2016 at 3:56 PM, Ross Zwisler > <ross.zwisler@linux.intel.com> wrote: >> On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote: >>> [ adding linux-fsdevel and linux-nvdimm ] >>> >>> On Wed, Sep 7, 2016 at 8:36 PM, Xiao Guangrong >>> <guangrong.xiao@linux.intel.com> wrote: >>> [..] >>>> However, it is not easy to handle the case that the new VMA overlays with >>>> the old VMA >>>> already got by userspace. I think we have some choices: >>>> 1: One way is completely skipping the new VMA region as current kernel code >>>> does but i >>>> do not think this is good as the later VMAs will be dropped. >>>> >>>> 2: show the un-overlayed portion of new VMA. In your case, we just show the >>>> region >>>> (0x2000 -> 0x3000), however, it can not work well if the VMA is a new >>>> created >>>> region with different attributions. >>>> >>>> 3: completely show the new VMA as this patch does. >>>> >>>> Which one do you prefer? >>>> >>> >>> I don't have a preference, but perhaps this breakage and uncertainty >>> is a good opportunity to propose a more reliable interface for NVML to >>> get the information it needs? >>> >>> My understanding is that it is looking for the VM_MIXEDMAP flag which >>> is already ambiguous for determining if DAX is enabled even if this >>> dynamic listing issue is fixed. XFS has arranged for DAX to be a >>> per-inode capability and has an XFS-specific inode flag. We can make >>> that a common inode flag, but it seems we should have a way to >>> interrogate the mapping itself in the case where the inode is unknown >>> or unavailable. I'm thinking extensions to mincore to have flags for >>> DAX and possibly whether the page is part of a pte, pmd, or pud >>> mapping. Just floating that idea before starting to look into the >>> implementation, comments or other ideas welcome... >> >> I think this goes back to our previous discussion about support for the PMEM >> programming model. Really I think what NVML needs isn't a way to tell if it >> is getting a DAX mapping, but whether it is getting a DAX mapping on a >> filesystem that fully supports the PMEM programming model. This of course is >> defined to be a filesystem where it can do all of its flushes from userspace >> safely and never call fsync/msync, and that allocations that happen in page >> faults will be synchronized to media before the page fault completes. >> >> IIUC this is what NVML needs - a way to decide "do I use fsync/msync for >> everything or can I rely fully on flushes from userspace?" >> >> For all existing implementations, I think the answer is "you need to use >> fsync/msync" because we don't yet have proper support for the PMEM programming >> model. >> >> My best idea of how to support this was a per-inode flag similar to the one >> supported by XFS that says "you have a PMEM capable DAX mapping", which NVML >> would then interpret to mean "you can do flushes from userspace and be fully >> safe". I think we really want this interface to be common over XFS and ext4. >> >> If we can figure out a better way of doing this interface, say via mincore, >> that's fine, but I don't think we can detangle this from the PMEM API >> discussion. > > Whether a persistent memory mapping requires an msync/fsync is a > filesystem specific question. This mincore proposal is separate from > that. Consider device-DAX for volatile memory or mincore() called on > an anonymous memory range. In those cases persistence and filesystem > metadata are not in the picture, but it would still be useful for > userspace to know "is there page cache backing this mapping?" or "what > is the TLB geometry of this mapping?". I got a question about msync/fsync which is beyond the topic of this thread :) Whether msync/fsync can make data persistent depends on ADR feature on memory controller, if it exists everything works well, otherwise, we need to have another interface that is why 'Flush hint table' in ACPI comes in. 'Flush hint table' is particularly useful for nvdimm virtualization if we use normal memory to emulate nvdimm with data persistent characteristic (the data will be flushed to a persistent storage, e.g, disk). Does current PMEM programming model fully supports 'Flush hint table'? Is userspace allowed to use these addresses? Thanks!
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-09-09 17:50 +0200 |
| Message-ID | <sfwdY-4Pt-19@gated-at.bofh.it> |
| In reply to | #1479767 |
On Fri, Sep 9, 2016 at 1:55 AM, Xiao Guangrong <guangrong.xiao@linux.intel.com> wrote: [..] >> >> Whether a persistent memory mapping requires an msync/fsync is a >> filesystem specific question. This mincore proposal is separate from >> that. Consider device-DAX for volatile memory or mincore() called on >> an anonymous memory range. In those cases persistence and filesystem >> metadata are not in the picture, but it would still be useful for >> userspace to know "is there page cache backing this mapping?" or "what >> is the TLB geometry of this mapping?". > > > I got a question about msync/fsync which is beyond the topic of this thread > :) > > Whether msync/fsync can make data persistent depends on ADR feature on > memory > controller, if it exists everything works well, otherwise, we need to have > another > interface that is why 'Flush hint table' in ACPI comes in. 'Flush hint > table' is > particularly useful for nvdimm virtualization if we use normal memory to > emulate > nvdimm with data persistent characteristic (the data will be flushed to a > persistent storage, e.g, disk). > > Does current PMEM programming model fully supports 'Flush hint table'? Is > userspace allowed to use these addresses? If you publish flush hint addresses in the virtual NFIT the guest VM will write to them whenever a REQ_FLUSH or REQ_FUA request is sent to the virtual /dev/pmemX device. Yes, seems straightforward to take a VM exit on those events and flush simulated pmem to persistent storage.
[toc] | [prev] | [next] | [standalone]
| From | Xiao Guangrong <guangrong.xiao@linux.intel.com> |
|---|---|
| Date | 2016-09-12 08:10 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgsBj-7Yg-3@gated-at.bofh.it> |
| In reply to | #1480168 |
On 09/09/2016 11:40 PM, Dan Williams wrote: > On Fri, Sep 9, 2016 at 1:55 AM, Xiao Guangrong > <guangrong.xiao@linux.intel.com> wrote: > [..] >>> >>> Whether a persistent memory mapping requires an msync/fsync is a >>> filesystem specific question. This mincore proposal is separate from >>> that. Consider device-DAX for volatile memory or mincore() called on >>> an anonymous memory range. In those cases persistence and filesystem >>> metadata are not in the picture, but it would still be useful for >>> userspace to know "is there page cache backing this mapping?" or "what >>> is the TLB geometry of this mapping?". >> >> >> I got a question about msync/fsync which is beyond the topic of this thread >> :) >> >> Whether msync/fsync can make data persistent depends on ADR feature on >> memory >> controller, if it exists everything works well, otherwise, we need to have >> another >> interface that is why 'Flush hint table' in ACPI comes in. 'Flush hint >> table' is >> particularly useful for nvdimm virtualization if we use normal memory to >> emulate >> nvdimm with data persistent characteristic (the data will be flushed to a >> persistent storage, e.g, disk). >> >> Does current PMEM programming model fully supports 'Flush hint table'? Is >> userspace allowed to use these addresses? > > If you publish flush hint addresses in the virtual NFIT the guest VM > will write to them whenever a REQ_FLUSH or REQ_FUA request is sent to > the virtual /dev/pmemX device. Yes, seems straightforward to take a > VM exit on those events and flush simulated pmem to persistent > storage. > Thank you, Dan! However REQ_FLUSH or REQ_FUA is handled in kernel space, okay, after following up the discussion in this thread, i understood that currently filesystems have not supported the case that usespace itself make data be persistent without kernel's involvement. So that works. Hmm, Does device-DAX support this case (make data be persistent without msync/fsync)? I guess no, but just want to confirm it. :)
[toc] | [prev] | [next] | [standalone]
| From | "Rudoff, Andy" <andy.rudoff@intel.com> |
|---|---|
| Date | 2016-09-12 05:50 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgqpP-6rr-3@gated-at.bofh.it> |
| In reply to | #1479767 |
>Whether msync/fsync can make data persistent depends on ADR feature on >memory controller, if it exists everything works well, otherwise, we need >to have another interface that is why 'Flush hint table' in ACPI comes >in. 'Flush hint table' is particularly useful for nvdimm virtualization if >we use normal memory to emulate nvdimm with data persistent characteristic >(the data will be flushed to a persistent storage, e.g, disk). > >Does current PMEM programming model fully supports 'Flush hint table'? Is >userspace allowed to use these addresses? The Flush hint table is NOT a replacement for ADR. To support pmem on the x86 architecture, the platform is required to ensure that a pmem store flushed from the CPU caches is in the persistent domain so that the application need not take any additional steps to make it persistent. The most common way to do this is the ADR feature. If the above is not true, then your x86 platform does not support pmem. Flush hints are for use by the BIOS and drivers and are not intended to be used in user space. Flush hints provide two things: First, if a driver needs to write to command registers or movable windows on a DIMM, the Flush hint (if provided in the NFIT) is required to flush the command to the DIMM or ensure stores done through the movable window are complete before moving it somewhere else. Second, for the rare case where the kernel wants to flush stores to the smallest possible failure domain (i.e. to the DIMM even though ADR will handle flushing it from a larger domain), the flush hints provide a way to do this. This might be useful for things like file system journals to help ensure the file system is consistent even in the face of ADR failure. -andy
[toc] | [prev] | [next] | [standalone]
| From | Xiao Guangrong <guangrong.xiao@linux.intel.com> |
|---|---|
| Date | 2016-09-12 08:40 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgt4l-88g-11@gated-at.bofh.it> |
| In reply to | #1480914 |
On 09/12/2016 11:44 AM, Rudoff, Andy wrote: >> Whether msync/fsync can make data persistent depends on ADR feature on >> memory controller, if it exists everything works well, otherwise, we need >> to have another interface that is why 'Flush hint table' in ACPI comes >> in. 'Flush hint table' is particularly useful for nvdimm virtualization if >> we use normal memory to emulate nvdimm with data persistent characteristic >> (the data will be flushed to a persistent storage, e.g, disk). >> >> Does current PMEM programming model fully supports 'Flush hint table'? Is >> userspace allowed to use these addresses? > > The Flush hint table is NOT a replacement for ADR. To support pmem on > the x86 architecture, the platform is required to ensure that a pmem > store flushed from the CPU caches is in the persistent domain so that the > application need not take any additional steps to make it persistent. > The most common way to do this is the ADR feature. > > If the above is not true, then your x86 platform does not support pmem. Understood. However, virtualization is a special case as we can use normal memory to emulate NVDIMM for the vm so that vm can bypass local file-cache, reduce memory usage and io path, etc. Currently, this usage is useful for lightweight virtualization, such as clean container. Under this case, ADR is available on physical platform but it can not help us to make data persistence for the vm. So that virtualizeing 'flush hint table' is a good way to handle it based on the acpi spec: | software can write to any one of these Flush Hint Addresses to | cause any preceding writes to the NVDIMM region to be flushed | out of the intervening platform buffers 1 to the targeted NVDIMM | (to achieve durability) > > Flush hints are for use by the BIOS and drivers and are not intended to > be used in user space. Flush hints provide two things: > > First, if a driver needs to write to command registers or movable windows > on a DIMM, the Flush hint (if provided in the NFIT) is required to flush > the command to the DIMM or ensure stores done through the movable window > are complete before moving it somewhere else. > > Second, for the rare case where the kernel wants to flush stores to the > smallest possible failure domain (i.e. to the DIMM even though ADR will > handle flushing it from a larger domain), the flush hints provide a way > to do this. This might be useful for things like file system journals to > help ensure the file system is consistent even in the face of ADR failure. We are assuming ADR can fail, however, do we have a way to know whether ADR works correctly? Maybe MCE can work on it? This is necessary to support making data persistent without 'fsync/msync' in userspace. Or do we need to unconditionally use 'flush hint address' if it is available as current nvdimm driver does? Thanks!
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-09-12 03:50 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgoxI-5f9-7@gated-at.bofh.it> |
| In reply to | #1479573 |
On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote: > On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote: > > My understanding is that it is looking for the VM_MIXEDMAP flag which > > is already ambiguous for determining if DAX is enabled even if this > > dynamic listing issue is fixed. XFS has arranged for DAX to be a > > per-inode capability and has an XFS-specific inode flag. We can make > > that a common inode flag, but it seems we should have a way to > > interrogate the mapping itself in the case where the inode is unknown > > or unavailable. I'm thinking extensions to mincore to have flags for > > DAX and possibly whether the page is part of a pte, pmd, or pud > > mapping. Just floating that idea before starting to look into the > > implementation, comments or other ideas welcome... > > I think this goes back to our previous discussion about support for the PMEM > programming model. Really I think what NVML needs isn't a way to tell if it > is getting a DAX mapping, but whether it is getting a DAX mapping on a > filesystem that fully supports the PMEM programming model. This of course is > defined to be a filesystem where it can do all of its flushes from userspace > safely and never call fsync/msync, and that allocations that happen in page > faults will be synchronized to media before the page fault completes. > > IIUC this is what NVML needs - a way to decide "do I use fsync/msync for > everything or can I rely fully on flushes from userspace?" "need fsync/msync" is a dynamic state of an inode, not a static property. i.e. users can do things that change an inode behind the back of a mapping, even if they are not aware that this might happen. As such, a filesystem can invalidate an existing mapping at any time and userspace won't notice because it will simply fault in a new mapping on the next access... > For all existing implementations, I think the answer is "you need to use > fsync/msync" because we don't yet have proper support for the PMEM programming > model. Yes, that is correct. FWIW, I don't think it will ever be possible to support this .... wonderful "PMEM programming model" from any current or future kernel filesystem without a very specific set of restrictions on what can be done to a file. e.g. 1. the file has to be fully allocated and zeroed before use. Preallocation/zeroing via unwritten extents is not allowed. Sparse files are not allowed. Shared extents are not allowed. 2. set the "PMEM_IMMUTABLE" inode flag - filesystem must check the file is fully allocated before allowing it to be set, and caller must have CAP_LINUX_IMMUTABLE. 3. Inode metadata is now immutable, and file data can only be accessed and/or modified via mmap(). 4. All non-mmap methods of inode data modification will now fail with EPERM. 5. all methods of inode metadata modification will now fail with EPERM, timestamp udpdates will be ignored. 6. PMEM_IMMUTABLE flag can only be removed if the file is not currently mapped and caller has CAP_LINUX_IMMUTABLE. A flag like this /should/ make it possible to avoid fsync/msync() on a file for existing filesystems, but it also means that such files have significant management issues (hence the need for CAP_LINUX_IMMUTABLE to cover it's use). Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | "Darrick J. Wong" <darrick.wong@oracle.com> |
|---|---|
| Date | 2016-09-15 08:00 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <shxSh-2ZY-1@gated-at.bofh.it> |
| In reply to | #1480884 |
On Mon, Sep 12, 2016 at 11:40:35AM +1000, Dave Chinner wrote: > On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote: > > On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote: > > > My understanding is that it is looking for the VM_MIXEDMAP flag which > > > is already ambiguous for determining if DAX is enabled even if this > > > dynamic listing issue is fixed. XFS has arranged for DAX to be a > > > per-inode capability and has an XFS-specific inode flag. We can make > > > that a common inode flag, but it seems we should have a way to > > > interrogate the mapping itself in the case where the inode is unknown > > > or unavailable. I'm thinking extensions to mincore to have flags for > > > DAX and possibly whether the page is part of a pte, pmd, or pud > > > mapping. Just floating that idea before starting to look into the > > > implementation, comments or other ideas welcome... > > > > I think this goes back to our previous discussion about support for the PMEM > > programming model. Really I think what NVML needs isn't a way to tell if it > > is getting a DAX mapping, but whether it is getting a DAX mapping on a > > filesystem that fully supports the PMEM programming model. This of course is > > defined to be a filesystem where it can do all of its flushes from userspace > > safely and never call fsync/msync, and that allocations that happen in page > > faults will be synchronized to media before the page fault completes. > > > > IIUC this is what NVML needs - a way to decide "do I use fsync/msync for > > everything or can I rely fully on flushes from userspace?" > > "need fsync/msync" is a dynamic state of an inode, not a static > property. i.e. users can do things that change an inode behind the > back of a mapping, even if they are not aware that this might > happen. As such, a filesystem can invalidate an existing mapping > at any time and userspace won't notice because it will simply fault > in a new mapping on the next access... > > > For all existing implementations, I think the answer is "you need to use > > fsync/msync" because we don't yet have proper support for the PMEM programming > > model. > > Yes, that is correct. > > FWIW, I don't think it will ever be possible to support this .... > wonderful "PMEM programming model" from any current or future kernel > filesystem without a very specific set of restrictions on what can > be done to a file. e.g. > > 1. the file has to be fully allocated and zeroed before > use. Preallocation/zeroing via unwritten extents is not > allowed. Sparse files are not allowed. Shared extents are > not allowed. > 2. set the "PMEM_IMMUTABLE" inode flag - filesystem must > check the file is fully allocated before allowing it to > be set, and caller must have CAP_LINUX_IMMUTABLE. > 3. Inode metadata is now immutable, and file data can only > be accessed and/or modified via mmap(). > 4. All non-mmap methods of inode data modification > will now fail with EPERM. > 5. all methods of inode metadata modification will now fail > with EPERM, timestamp udpdates will be ignored. > 6. PMEM_IMMUTABLE flag can only be removed if the file is > not currently mapped and caller has CAP_LINUX_IMMUTABLE. > > A flag like this /should/ make it possible to avoid fsync/msync() on > a file for existing filesystems, but it also means that such files > have significant management issues (hence the need for > CAP_LINUX_IMMUTABLE to cover it's use). Hmmm... I started to ponder such a flag, but ran into some questions. If it's PMEM_IMMUTABLE, does this mean that none of 1-6 apply if the filesystem discovers it isn't on pmem? I thought about just having a 'immutable metadata' flag where any timestamp, xattr, or block mapping update just returns EPERM. There wouldn't be any checks as in (1); if you left a hole in the file prior to setting the flag then you won't be filling it unless you clear the flag. OTOH if it merely made the metadata unchangeable then it's a stretch to get to non-mmap data accesses also being disallowed. Maybe the immutable metadata and mmap-only properties would only be implied if both DAX and IMMUTABLE_META are set on a file? Ok no more rambling until sleep. :) --D > > Cheers, > > Dave. > -- > Dave Chinner > david@fromorbit.com > -- > To unsubscribe from this list: send the line "unsubscribe linux-fsdevel" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-09-15 08:30 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <shylj-3ql-15@gated-at.bofh.it> |
| In reply to | #1483859 |
On Wed, Sep 14, 2016 at 10:55:03PM -0700, Darrick J. Wong wrote: > On Mon, Sep 12, 2016 at 11:40:35AM +1000, Dave Chinner wrote: > > On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote: > > > On Wed, Sep 07, 2016 at 09:32:36PM -0700, Dan Williams wrote: > > > > My understanding is that it is looking for the VM_MIXEDMAP flag which > > > > is already ambiguous for determining if DAX is enabled even if this > > > > dynamic listing issue is fixed. XFS has arranged for DAX to be a > > > > per-inode capability and has an XFS-specific inode flag. We can make > > > > that a common inode flag, but it seems we should have a way to > > > > interrogate the mapping itself in the case where the inode is unknown > > > > or unavailable. I'm thinking extensions to mincore to have flags for > > > > DAX and possibly whether the page is part of a pte, pmd, or pud > > > > mapping. Just floating that idea before starting to look into the > > > > implementation, comments or other ideas welcome... > > > > > > I think this goes back to our previous discussion about support for the PMEM > > > programming model. Really I think what NVML needs isn't a way to tell if it > > > is getting a DAX mapping, but whether it is getting a DAX mapping on a > > > filesystem that fully supports the PMEM programming model. This of course is > > > defined to be a filesystem where it can do all of its flushes from userspace > > > safely and never call fsync/msync, and that allocations that happen in page > > > faults will be synchronized to media before the page fault completes. > > > > > > IIUC this is what NVML needs - a way to decide "do I use fsync/msync for > > > everything or can I rely fully on flushes from userspace?" > > > > "need fsync/msync" is a dynamic state of an inode, not a static > > property. i.e. users can do things that change an inode behind the > > back of a mapping, even if they are not aware that this might > > happen. As such, a filesystem can invalidate an existing mapping > > at any time and userspace won't notice because it will simply fault > > in a new mapping on the next access... > > > > > For all existing implementations, I think the answer is "you need to use > > > fsync/msync" because we don't yet have proper support for the PMEM programming > > > model. > > > > Yes, that is correct. > > > > FWIW, I don't think it will ever be possible to support this .... > > wonderful "PMEM programming model" from any current or future kernel > > filesystem without a very specific set of restrictions on what can > > be done to a file. e.g. > > > > 1. the file has to be fully allocated and zeroed before > > use. Preallocation/zeroing via unwritten extents is not > > allowed. Sparse files are not allowed. Shared extents are > > not allowed. > > 2. set the "PMEM_IMMUTABLE" inode flag - filesystem must > > check the file is fully allocated before allowing it to > > be set, and caller must have CAP_LINUX_IMMUTABLE. > > 3. Inode metadata is now immutable, and file data can only > > be accessed and/or modified via mmap(). > > 4. All non-mmap methods of inode data modification > > will now fail with EPERM. > > 5. all methods of inode metadata modification will now fail > > with EPERM, timestamp udpdates will be ignored. > > 6. PMEM_IMMUTABLE flag can only be removed if the file is > > not currently mapped and caller has CAP_LINUX_IMMUTABLE. > > > > A flag like this /should/ make it possible to avoid fsync/msync() on > > a file for existing filesystems, but it also means that such files > > have significant management issues (hence the need for > > CAP_LINUX_IMMUTABLE to cover it's use). > > Hmmm... I started to ponder such a flag, but ran into some questions. > If it's PMEM_IMMUTABLE, does this mean that none of 1-6 apply if the > filesystem discovers it isn't on pmem? Would only be meaningful if the FS_XFLAG_DAX/S_DAX flag is also set on the inode and the backing store is dax capable. Hence the 'PMEM' part of the name. > I thought about just having a 'immutable metadata' flag where any > timestamp, xattr, or block mapping update just returns EPERM. And all the rest - no hard links, no perm/owner changes, no security context changes(!), and so on. ANd it's even more complex with filesystems that have COW metadata and pack multiple unrelated metadata objects into single blocks - they can do all sorts of interesting things on unrealted metadata updates... :P You'd also have to turn off background internal filesystem mod vectors, too, like EOF scanning, or defrag, balance, dedupe, auto-repair, etc. And, now that I think about it, snapshots are out of the question too. This gets more hairy the more I think about what our filesystems can do these days.... > There > wouldn't be any checks as in (1); if you left a hole in the file prior > to setting the flag then you won't be filling it unless you clear the > flag. Which means writing into a hole would need to return an error, and a write page fault into a hole would need a segv. Seems like a great way to cause random application failures to me... > OTOH if it merely made the metadata unchangeable then it's a > stretch to get to non-mmap data accesses also being disallowed. *nod* > Maybe the immutable metadata and mmap-only properties would only be > implied if both DAX and IMMUTABLE_META are set on a file? I'd suggest that PMEM_IMMUTABLE could only be set on an inode that already has the FS_XFLAG_DAX set on it (or it is being set at the same time). And clearing the DAX flag would also remove the PMEM_IMMUTABLE flag. Perhaps it would be better to call it FS_XFLAG_DAX_IMMUTABLE rather than anything pmem related. Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-12 07:30 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgrYB-7wi-3@gated-at.bofh.it> |
| In reply to | #1479573 |
On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote: > I think this goes back to our previous discussion about support for the PMEM > programming model. Really I think what NVML needs isn't a way to tell if it > is getting a DAX mapping, but whether it is getting a DAX mapping on a > filesystem that fully supports the PMEM programming model. This of course is > defined to be a filesystem where it can do all of its flushes from userspace > safely and never call fsync/msync, and that allocations that happen in page > faults will be synchronized to media before the page fault completes. That's a an easy way to flag: you will never get that from a Linux filesystem, period. NVML folks really need to stop taking crack and dreaming this could happen.
[toc] | [prev] | [next] | [standalone]
| From | "Oliver O'Halloran" <oohall@gmail.com> |
|---|---|
| Date | 2016-09-12 09:30 +0200 |
| Message-ID | <sgtQJ-ct-1@gated-at.bofh.it> |
| In reply to | #1480928 |
On Mon, Sep 12, 2016 at 3:27 PM, Christoph Hellwig <hch@infradead.org> wrote: > On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote: >> I think this goes back to our previous discussion about support for the PMEM >> programming model. Really I think what NVML needs isn't a way to tell if it >> is getting a DAX mapping, but whether it is getting a DAX mapping on a >> filesystem that fully supports the PMEM programming model. This of course is >> defined to be a filesystem where it can do all of its flushes from userspace >> safely and never call fsync/msync, and that allocations that happen in page >> faults will be synchronized to media before the page fault completes. > > That's a an easy way to flag: you will never get that from a Linux > filesystem, period. > > NVML folks really need to stop taking crack and dreaming this could > happen. Well, that's a bummer. What are the problems here? Is this a matter of existing filesystems being unable/unwilling to support this or is it just fundamentally broken? The end goal is to let applications manage the persistence of their own data without having to involve the kernel in every IOP, but if we can't do that then what would a 90% solution look like? I think most people would be OK with having to do an fsync() occasionally, but not after ever write to pmem.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-12 10:00 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgujL-mJ-3@gated-at.bofh.it> |
| In reply to | #1480965 |
On Mon, Sep 12, 2016 at 05:25:15PM +1000, Oliver O'Halloran wrote: > What are the problems here? Is this a matter of existing filesystems > being unable/unwilling to support this or is it just fundamentally > broken? It's a fundamentally broken model. See Dave's post that actually was sent slightly earlier then mine for the list of required items, which is fairly unrealistic. You could probably try to architect a file system for it, but I doubt it would gain much traction. > The end goal is to let applications manage the persistence of > their own data without having to involve the kernel in every IOP, but > if we can't do that then what would a 90% solution look like? I think > most people would be OK with having to do an fsync() occasionally, but > not after ever write to pmem. You need an fsync for each write that you want to persist. This sounds painful for now. But I have an implementation that will allow the atomic commit of more or less arbitrary amounts of previous writes for XFS that I plan to land once the reflink work is in. That way you create almost arbitrarily complex data structures in your programs and commit them atomicly. It's not going to fit the nvml model, but that whole think has been complete bullshit since the beginning anyway.
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-12 10:10 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgutt-Fd-39@gated-at.bofh.it> |
| In reply to | #1481000 |
On Mon, 12 Sep 2016 00:51:28 -0700 Christoph Hellwig <hch@infradead.org> wrote: > On Mon, Sep 12, 2016 at 05:25:15PM +1000, Oliver O'Halloran wrote: > > What are the problems here? Is this a matter of existing filesystems > > being unable/unwilling to support this or is it just fundamentally > > broken? > > It's a fundamentally broken model. See Dave's post that actually was > sent slightly earlier then mine for the list of required items, which > is fairly unrealistic. You could probably try to architect a file > system for it, but I doubt it would gain much traction. It's not fundamentally broken, it just doesn't fit well existing filesystems. Dave's post of requirements is also wrong. A filesystem does not have to guarantee all that, it only has to guarantee that is the case for a given block after it has a mapping and page fault returns, other operations can be supported by invalidating mappings, etc.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-12 17:10 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgB1U-4WE-23@gated-at.bofh.it> |
| In reply to | #1481021 |
On Mon, Sep 12, 2016 at 06:05:07PM +1000, Nicholas Piggin wrote: > It's not fundamentally broken, it just doesn't fit well existing > filesystems. Or the existing file system architecture for that matter. Which makes it a fundamentally broken model. > Dave's post of requirements is also wrong. A filesystem does not have > to guarantee all that, it only has to guarantee that is the case for > a given block after it has a mapping and page fault returns, other > operations can be supported by invalidating mappings, etc. Which doesn't really matter if your use case is manipulating fully mapped files. But back to the point: if you want to use a full blown Linux or Unix filesystem you will always have to fsync (or variants of it like msync), period. If you want a volume manager on stereoids that hands out large chunks of storage memory that can't ever be moved, truncated, shared, allocated on demand, etc - implement it in your library on top of a device file.
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-13 03:40 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgKRz-39a-1@gated-at.bofh.it> |
| In reply to | #1481395 |
On Mon, 12 Sep 2016 08:01:48 -0700 Christoph Hellwig <hch@infradead.org> wrote: > On Mon, Sep 12, 2016 at 06:05:07PM +1000, Nicholas Piggin wrote: > > It's not fundamentally broken, it just doesn't fit well existing > > filesystems. > > Or the existing file system architecture for that matter. Which makes > it a fundamentally broken model. Not really. A few reasonable changes can be made to improve things. Until just now you thought it was fundamentally impossible to make a reasonable implementation due to Dave's "constraints". > > > Dave's post of requirements is also wrong. A filesystem does not have > > to guarantee all that, it only has to guarantee that is the case for > > a given block after it has a mapping and page fault returns, other > > operations can be supported by invalidating mappings, etc. > > Which doesn't really matter if your use case is manipulating > fully mapped files. Nothing that says you have to use them fully mapped always and not use other APIs on them. > But back to the point: if you want to use a full blown Linux or Unix > filesystem you will always have to fsync (or variants of it like msync), > period. That's circular logic. First you said that should not be done because of your imagined constraints. In fact, it's not unreasonable to describe some additional semantics of the storage that is unavailable with traditional filesystems. That said, a noop system call is on the order of 100 cycles nowadays, so rushing to implement these APIs without seeing good numbers and actual users ready to go seems premature. *This* is the real reason not to implement new APIs yet. > If you want a volume manager on stereoids that hands out large chunks > of storage memory that can't ever be moved, truncated, shared, allocated > on demand, etc - implement it in your library on top of a device file. Those constraints don't exist either. I've written a filesystem that avoids them. It isn't rocket science.
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-09-13 06:10 +0200 |
| Message-ID | <sgNcJ-4Tl-7@gated-at.bofh.it> |
| In reply to | #1482080 |
On Mon, Sep 12, 2016 at 6:31 PM, Nicholas Piggin <npiggin@gmail.com> wrote: > On Mon, 12 Sep 2016 08:01:48 -0700 [..] > That said, a noop system call is on the order of 100 cycles nowadays, > so rushing to implement these APIs without seeing good numbers and > actual users ready to go seems premature. *This* is the real reason > not to implement new APIs yet. Yes, and harvesting the current crop of low hanging performance fruit in the filesystem-DAX I/O path remains on the todo list. In the meantime we're pursuing this mm api, mincore+ or whatever we end up with, to allow userspace to distinguish memory address ranges that are backed by a filesystem requiring coordination of metadata updates + flushes for updates, vs something like device-dax that does not.
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-13 07:50 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgOLv-5S2-11@gated-at.bofh.it> |
| In reply to | #1482131 |
On Mon, 12 Sep 2016 21:06:49 -0700 Dan Williams <dan.j.williams@intel.com> wrote: > On Mon, Sep 12, 2016 at 6:31 PM, Nicholas Piggin <npiggin@gmail.com> wrote: > > On Mon, 12 Sep 2016 08:01:48 -0700 > [..] > > That said, a noop system call is on the order of 100 cycles nowadays, > > so rushing to implement these APIs without seeing good numbers and > > actual users ready to go seems premature. *This* is the real reason > > not to implement new APIs yet. > > Yes, and harvesting the current crop of low hanging performance fruit > in the filesystem-DAX I/O path remains on the todo list. > > In the meantime we're pursuing this mm api, mincore+ or whatever we > end up with, to allow userspace to distinguish memory address ranges > that are backed by a filesystem requiring coordination of metadata > updates + flushes for updates, vs something like device-dax that does > not. Yes, that's reasonable. Do you need page/block granularity? Do you need a way to advise/request the fs for a particular capability? Is it enough to request and check success? Would the capability be likely to change, and if so, how would you notify the app asynchronously?
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-09-12 23:40 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgH7j-FR-21@gated-at.bofh.it> |
| In reply to | #1481021 |
On Mon, Sep 12, 2016 at 06:05:07PM +1000, Nicholas Piggin wrote: > On Mon, 12 Sep 2016 00:51:28 -0700 > Christoph Hellwig <hch@infradead.org> wrote: > > > On Mon, Sep 12, 2016 at 05:25:15PM +1000, Oliver O'Halloran wrote: > > > What are the problems here? Is this a matter of existing filesystems > > > being unable/unwilling to support this or is it just fundamentally > > > broken? > > > > It's a fundamentally broken model. See Dave's post that actually was > > sent slightly earlier then mine for the list of required items, which > > is fairly unrealistic. You could probably try to architect a file > > system for it, but I doubt it would gain much traction. > > It's not fundamentally broken, it just doesn't fit well existing > filesystems. > > Dave's post of requirements is also wrong. A filesystem does not have > to guarantee all that, it only has to guarantee that is the case for > a given block after it has a mapping and page fault returns, other > operations can be supported by invalidating mappings, etc. Sure, but filesystems are completely unaware of what is mapped at any given time, or what constraints that mapping might have. Trying to make filesystems aware of per-page mapping constraints seems like a fairly significant layering violation based on a flawed assumption. i.e. that operations on other parts of the file do not affect the block that requires immutable metadata. e.g an extent operation in some other area of the file can cause a tip-to-root extent tree split or merge, and that moves the metadata that points to the mapped block that we've told userspace "doesn't need fsync". We now need an fsync to ensure that the metadata is consistent on disk again, even though that block has not physically been moved. IOWs, the immutable data block updates are now not ordered correctly w.r.t. other updates done to the file, especially when we consider crash recovery.... All this will expose is an unfixable problem with ordering of stable data + metadata operations and their synchronisation. As such, it seems like nothing but a major cluster-fuck to try to do mapping specific, per-block immutable metadata - it adds major complexity and even more untractable problems. Yes, we /could/ try to solve this but, quite frankly, it's far easier to change the broken PMEM programming model assumptions than it is to implement what you are suggesting. Or to do what Christoph suggested and just use a wrapper around something like device mapper to hand out chunks of unchanging, static pmem to applications... Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web