Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1261900

Re: [PATCH v3 14/15] dax: dirty extent notification

From Dan Williams <dan.j.williams@intel.com>
Newsgroups linux.kernel
Subject Re: [PATCH v3 14/15] dax: dirty extent notification
Date 2015-11-03 22:20 +0100
Message-ID <qqR9L-87W-5@gated-at.bofh.it> (permalink)
References (2 earlier) <qqyqu-4js-15@gated-at.bofh.it> <qqC13-6GC-1@gated-at.bofh.it> <qqCDM-6X8-17@gated-at.bofh.it> <qqEcy-83d-27@gated-at.bofh.it> <qqQQp-7KR-5@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


On Tue, Nov 3, 2015 at 12:51 PM, Dave Chinner <david@fromorbit.com> wrote:
> On Mon, Nov 02, 2015 at 11:20:49PM -0800, Dan Williams wrote:
>> On Mon, Nov 2, 2015 at 9:40 PM, Dave Chinner <david@fromorbit.com> wrote:
>> > On Mon, Nov 02, 2015 at 08:56:24PM -0800, Dan Williams wrote:
>> >> No, we definitely can't do that.   I think your mental model of the
>> >> cache flushing is similar to the disk model where a small buffer is
>> >> flushed after a large streaming write.  Both Ross' patches and my
>> >> approach suffer from the same horror that the cache flushing is O(N)
>> >> currently, so we don't want to make it responsible for more data
>> >> ranges areas than is strictly necessary.
>> >
>> > I didn't see anything that was O(N) in Ross's patches. What part of
>> > the fsync algorithm that Ross proposed are you refering to here?
>>
>> We have to issue clflush per touched virtual address rather than a
>> constant number of physical ways, or a flush-all instruction.
> .....
>> > So don't tell me that tracking dirty pages in the radix tree too
>> > slow for DAX and that DAX should not be used for POSIX IO based
>> > applications - it should be as fast as buffered IO, if not faster,
>> > and if it isn't then we've screwed up real bad. And right now, we're
>> > screwing up real bad.
>>
>> Again, it's not the dirty tracking in the radix I'm worried about it's
>> looping through all the virtual addresses within those pages..
>
> So, let me summarise what I think you've just said. You are
>
> 1. fine with looping through the virtual addresses doing cache flushes
>    synchronously when doing IO despite it having significant
>    latency and performance costs.

No, like I said in the blkdev_issue_zeroout thread we need to replace
looping flushes with non-temporal stores and delayed wmb_pmem()
wherever possible.

> 2. Happy to hack a method into DAX to bypass the filesystems by
>    pushing information to the block device for it to track regions that
>    need cache flushes, then add infrastructure to the block device to
>    track those dirty regions and then walk those addresses and issue
>    cache flushes when the filesystem issues a REQ_FLUSH IO regardless
>    of whether the filesystem actually needs those cachelines flushed
>    for that specific IO?

I'm happier with a temporary driver level hack than a temporary core
kernel change.  This requirement to flush by virtual address is
something that, in my opinion, must be addressed by the platform with
a reliable global flush or by walking a small constant number of
physical-cache-ways.  I think we're getting ahead of ourselves jumping
to solving this in the core kernel while the question of how to do
efficient large flushes is still pending.

> 3. Not happy to use the generic mm/vfs level infrastructure
>    architectected specifically to provide the exact asynchronous
>    cache flushing/writeback semantics we require because it will
>    cause too many cache flushes, even though the number of cache
>    flushes will be, at worst, the same as in 2).

Correct, because if/when a platform solution arrives the need to track
dirty pfns evaporates.

> 1) will work, but as we can see it is *slow*. 3) is what Ross is
> implementing - it's a tried and tested architecture that all mm/fs
> developers understand, and his explanation of why it will work for
> pmem is pretty solid and completely platform/hardware architecture
> independent.
>
> Which leaves this question: How does 2) save us anything in terms of
> avoiding iterating virtual addresses and issuing cache flushes
> over 3)? And is it sufficient to justify hacking a bypass into DAX
> and the additional driver level complexity of having to add dirty
> region tracking, flushing and cleaning to REQ_FLUSH operations?
>

Given what we are talking about amounts to a hardware workaround I
think that kind of logic belongs in a driver.  If the cache flushing
gets fixed and we stop needing to track individual cachelines the
flush implementation will look and feel much more like existing
storage drivers.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

Back to linux.kernel | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

[PATCH v3 00/15] block, dax updates for 4.4 Dan Williams <dan.j.williams@intel.com> - 2015-11-02 05:40 +0100
  [PATCH v3 12/15] block: enable dax for raw block devices Dan Williams <dan.j.williams@intel.com> - 2015-11-02 05:40 +0100
  [PATCH v3 04/15] libnvdimm,  pmem: move request_queue allocation earlier in probe Dan Williams <dan.j.williams@intel.com> - 2015-11-02 05:40 +0100
    Re: [PATCH v3 04/15] libnvdimm, pmem: move request_queue allocation  earlier in probe Ross Zwisler <ross.zwisler@linux.intel.com> - 2015-11-03 20:20 +0100
  [PATCH v3 07/15] kvm: rename pfn_t to kvm_pfn_t Dan Williams <dan.j.williams@intel.com> - 2015-11-02 05:40 +0100
  [PATCH v3 13/15] block, dax: make dax mappings opt-in by default Dan Williams <dan.j.williams@intel.com> - 2015-11-02 05:40 +0100
    Re: [PATCH v3 13/15] block, dax: make dax mappings opt-in by default Dave Chinner <david@fromorbit.com> - 2015-11-03 01:40 +0100
      Re: [PATCH v3 13/15] block, dax: make dax mappings opt-in by default Dan Williams <dan.j.williams@intel.com> - 2015-11-03 08:40 +0100
        Re: [PATCH v3 13/15] block, dax: make dax mappings opt-in by default Dave Chinner <david@fromorbit.com> - 2015-11-03 21:30 +0100
          Re: [PATCH v3 13/15] block, dax: make dax mappings opt-in by default Dan Williams <dan.j.williams@intel.com> - 2015-11-04 00:10 +0100
            Re: [PATCH v3 13/15] block, dax: make dax mappings opt-in by default Dan Williams <dan.j.williams@intel.com> - 2015-11-04 20:30 +0100
  [PATCH v3 05/15] libnvdimm,  pmem: fix size trim in pmem_direct_access() Dan Williams <dan.j.williams@intel.com> - 2015-11-02 05:40 +0100
    Re: [PATCH v3 05/15] libnvdimm, pmem: fix size trim in  pmem_direct_access() Ross Zwisler <ross.zwisler@linux.intel.com> - 2015-11-03 20:40 +0100
      Re: [PATCH v3 05/15] libnvdimm, pmem: fix size trim in pmem_direct_access() Dan Williams <dan.j.williams@intel.com> - 2015-11-03 22:40 +0100
  [PATCH v3 14/15] dax: dirty extent notification Dan Williams <dan.j.williams@intel.com> - 2015-11-02 05:40 +0100
    Re: [PATCH v3 14/15] dax: dirty extent notification Dave Chinner <david@fromorbit.com> - 2015-11-03 02:20 +0100
      Re: [PATCH v3 14/15] dax: dirty extent notification Dan Williams <dan.j.williams@intel.com> - 2015-11-03 06:10 +0100
        Re: [PATCH v3 14/15] dax: dirty extent notification Dave Chinner <david@fromorbit.com> - 2015-11-03 06:50 +0100
          Re: [PATCH v3 14/15] dax: dirty extent notification Dan Williams <dan.j.williams@intel.com> - 2015-11-03 08:30 +0100
            Re: [PATCH v3 14/15] dax: dirty extent notification Dave Chinner <david@fromorbit.com> - 2015-11-03 22:00 +0100
              Re: [PATCH v3 14/15] dax: dirty extent notification Dan Williams <dan.j.williams@intel.com> - 2015-11-03 22:20 +0100
              Re: [PATCH v3 14/15] dax: dirty extent notification Ross Zwisler <ross.zwisler@linux.intel.com> - 2015-11-03 22:40 +0100
                Re: [PATCH v3 14/15] dax: dirty extent notification Dan Williams <dan.j.williams@intel.com> - 2015-11-03 22:50 +0100
        Re: [PATCH v3 14/15] dax: dirty extent notification Ross Zwisler <ross.zwisler@linux.intel.com> - 2015-11-03 22:20 +0100
          Re: [PATCH v3 14/15] dax: dirty extent notification Dan Williams <dan.j.williams@intel.com> - 2015-11-03 22:40 +0100
  [PATCH v3 10/15] dax,  pmem: introduce zone_device_revoke() and devm_memunmap_pages() Dan Williams <dan.j.williams@intel.com> - 2015-11-02 05:40 +0100
  [PATCH v3 06/15] um: kill pfn_t Dan Williams <dan.j.williams@intel.com> - 2015-11-02 05:40 +0100

csiph-web