Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1480928 > unrolled thread
| Started by | Christoph Hellwig <hch@infradead.org> |
|---|---|
| First post | 2016-09-12 07:30 +0200 |
| Last post | 2016-09-16 08:00 +0200 |
| Articles | 20 — 5 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-12 07:30 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) "Oliver O'Halloran" <oohall@gmail.com> - 2016-09-12 09:30 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-12 10:00 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-12 10:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-12 17:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 03:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dan Williams <dan.j.williams@intel.com> - 2016-09-13 06:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 07:50 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-12 23:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 04:00 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Christoph Hellwig <hch@infradead.org> - 2016-09-13 09:20 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-13 11:10 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-14 09:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-14 12:20 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-15 04:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-15 06:00 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-15 12:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-15 13:50 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Dave Chinner <david@fromorbit.com> - 2016-09-16 00:40 +0200
Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) Nicholas Piggin <npiggin@gmail.com> - 2016-09-16 08:00 +0200
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-12 07:30 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgrYB-7wi-3@gated-at.bofh.it> |
On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote: > I think this goes back to our previous discussion about support for the PMEM > programming model. Really I think what NVML needs isn't a way to tell if it > is getting a DAX mapping, but whether it is getting a DAX mapping on a > filesystem that fully supports the PMEM programming model. This of course is > defined to be a filesystem where it can do all of its flushes from userspace > safely and never call fsync/msync, and that allocations that happen in page > faults will be synchronized to media before the page fault completes. That's a an easy way to flag: you will never get that from a Linux filesystem, period. NVML folks really need to stop taking crack and dreaming this could happen.
[toc] | [next] | [standalone]
| From | "Oliver O'Halloran" <oohall@gmail.com> |
|---|---|
| Date | 2016-09-12 09:30 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgtQJ-ct-1@gated-at.bofh.it> |
| In reply to | #1480928 |
On Mon, Sep 12, 2016 at 3:27 PM, Christoph Hellwig <hch@infradead.org> wrote: > On Thu, Sep 08, 2016 at 04:56:36PM -0600, Ross Zwisler wrote: >> I think this goes back to our previous discussion about support for the PMEM >> programming model. Really I think what NVML needs isn't a way to tell if it >> is getting a DAX mapping, but whether it is getting a DAX mapping on a >> filesystem that fully supports the PMEM programming model. This of course is >> defined to be a filesystem where it can do all of its flushes from userspace >> safely and never call fsync/msync, and that allocations that happen in page >> faults will be synchronized to media before the page fault completes. > > That's a an easy way to flag: you will never get that from a Linux > filesystem, period. > > NVML folks really need to stop taking crack and dreaming this could > happen. Well, that's a bummer. What are the problems here? Is this a matter of existing filesystems being unable/unwilling to support this or is it just fundamentally broken? The end goal is to let applications manage the persistence of their own data without having to involve the kernel in every IOP, but if we can't do that then what would a 90% solution look like? I think most people would be OK with having to do an fsync() occasionally, but not after ever write to pmem.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-12 10:00 +0200 |
| Message-ID | <sgujL-mJ-3@gated-at.bofh.it> |
| In reply to | #1480965 |
On Mon, Sep 12, 2016 at 05:25:15PM +1000, Oliver O'Halloran wrote: > What are the problems here? Is this a matter of existing filesystems > being unable/unwilling to support this or is it just fundamentally > broken? It's a fundamentally broken model. See Dave's post that actually was sent slightly earlier then mine for the list of required items, which is fairly unrealistic. You could probably try to architect a file system for it, but I doubt it would gain much traction. > The end goal is to let applications manage the persistence of > their own data without having to involve the kernel in every IOP, but > if we can't do that then what would a 90% solution look like? I think > most people would be OK with having to do an fsync() occasionally, but > not after ever write to pmem. You need an fsync for each write that you want to persist. This sounds painful for now. But I have an implementation that will allow the atomic commit of more or less arbitrary amounts of previous writes for XFS that I plan to land once the reflink work is in. That way you create almost arbitrarily complex data structures in your programs and commit them atomicly. It's not going to fit the nvml model, but that whole think has been complete bullshit since the beginning anyway.
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-12 10:10 +0200 |
| Message-ID | <sgutt-Fd-39@gated-at.bofh.it> |
| In reply to | #1481000 |
On Mon, 12 Sep 2016 00:51:28 -0700 Christoph Hellwig <hch@infradead.org> wrote: > On Mon, Sep 12, 2016 at 05:25:15PM +1000, Oliver O'Halloran wrote: > > What are the problems here? Is this a matter of existing filesystems > > being unable/unwilling to support this or is it just fundamentally > > broken? > > It's a fundamentally broken model. See Dave's post that actually was > sent slightly earlier then mine for the list of required items, which > is fairly unrealistic. You could probably try to architect a file > system for it, but I doubt it would gain much traction. It's not fundamentally broken, it just doesn't fit well existing filesystems. Dave's post of requirements is also wrong. A filesystem does not have to guarantee all that, it only has to guarantee that is the case for a given block after it has a mapping and page fault returns, other operations can be supported by invalidating mappings, etc.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-12 17:10 +0200 |
| Message-ID | <sgB1U-4WE-23@gated-at.bofh.it> |
| In reply to | #1481021 |
On Mon, Sep 12, 2016 at 06:05:07PM +1000, Nicholas Piggin wrote: > It's not fundamentally broken, it just doesn't fit well existing > filesystems. Or the existing file system architecture for that matter. Which makes it a fundamentally broken model. > Dave's post of requirements is also wrong. A filesystem does not have > to guarantee all that, it only has to guarantee that is the case for > a given block after it has a mapping and page fault returns, other > operations can be supported by invalidating mappings, etc. Which doesn't really matter if your use case is manipulating fully mapped files. But back to the point: if you want to use a full blown Linux or Unix filesystem you will always have to fsync (or variants of it like msync), period. If you want a volume manager on stereoids that hands out large chunks of storage memory that can't ever be moved, truncated, shared, allocated on demand, etc - implement it in your library on top of a device file.
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-13 03:40 +0200 |
| Message-ID | <sgKRz-39a-1@gated-at.bofh.it> |
| In reply to | #1481395 |
On Mon, 12 Sep 2016 08:01:48 -0700 Christoph Hellwig <hch@infradead.org> wrote: > On Mon, Sep 12, 2016 at 06:05:07PM +1000, Nicholas Piggin wrote: > > It's not fundamentally broken, it just doesn't fit well existing > > filesystems. > > Or the existing file system architecture for that matter. Which makes > it a fundamentally broken model. Not really. A few reasonable changes can be made to improve things. Until just now you thought it was fundamentally impossible to make a reasonable implementation due to Dave's "constraints". > > > Dave's post of requirements is also wrong. A filesystem does not have > > to guarantee all that, it only has to guarantee that is the case for > > a given block after it has a mapping and page fault returns, other > > operations can be supported by invalidating mappings, etc. > > Which doesn't really matter if your use case is manipulating > fully mapped files. Nothing that says you have to use them fully mapped always and not use other APIs on them. > But back to the point: if you want to use a full blown Linux or Unix > filesystem you will always have to fsync (or variants of it like msync), > period. That's circular logic. First you said that should not be done because of your imagined constraints. In fact, it's not unreasonable to describe some additional semantics of the storage that is unavailable with traditional filesystems. That said, a noop system call is on the order of 100 cycles nowadays, so rushing to implement these APIs without seeing good numbers and actual users ready to go seems premature. *This* is the real reason not to implement new APIs yet. > If you want a volume manager on stereoids that hands out large chunks > of storage memory that can't ever be moved, truncated, shared, allocated > on demand, etc - implement it in your library on top of a device file. Those constraints don't exist either. I've written a filesystem that avoids them. It isn't rocket science.
[toc] | [prev] | [next] | [standalone]
| From | Dan Williams <dan.j.williams@intel.com> |
|---|---|
| Date | 2016-09-13 06:10 +0200 |
| Subject | Re: DAX mapping detection (was: Re: [PATCH] Fix region lost in /proc/self/smaps) |
| Message-ID | <sgNcJ-4Tl-7@gated-at.bofh.it> |
| In reply to | #1482080 |
On Mon, Sep 12, 2016 at 6:31 PM, Nicholas Piggin <npiggin@gmail.com> wrote: > On Mon, 12 Sep 2016 08:01:48 -0700 [..] > That said, a noop system call is on the order of 100 cycles nowadays, > so rushing to implement these APIs without seeing good numbers and > actual users ready to go seems premature. *This* is the real reason > not to implement new APIs yet. Yes, and harvesting the current crop of low hanging performance fruit in the filesystem-DAX I/O path remains on the todo list. In the meantime we're pursuing this mm api, mincore+ or whatever we end up with, to allow userspace to distinguish memory address ranges that are backed by a filesystem requiring coordination of metadata updates + flushes for updates, vs something like device-dax that does not.
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-13 07:50 +0200 |
| Message-ID | <sgOLv-5S2-11@gated-at.bofh.it> |
| In reply to | #1482131 |
On Mon, 12 Sep 2016 21:06:49 -0700 Dan Williams <dan.j.williams@intel.com> wrote: > On Mon, Sep 12, 2016 at 6:31 PM, Nicholas Piggin <npiggin@gmail.com> wrote: > > On Mon, 12 Sep 2016 08:01:48 -0700 > [..] > > That said, a noop system call is on the order of 100 cycles nowadays, > > so rushing to implement these APIs without seeing good numbers and > > actual users ready to go seems premature. *This* is the real reason > > not to implement new APIs yet. > > Yes, and harvesting the current crop of low hanging performance fruit > in the filesystem-DAX I/O path remains on the todo list. > > In the meantime we're pursuing this mm api, mincore+ or whatever we > end up with, to allow userspace to distinguish memory address ranges > that are backed by a filesystem requiring coordination of metadata > updates + flushes for updates, vs something like device-dax that does > not. Yes, that's reasonable. Do you need page/block granularity? Do you need a way to advise/request the fs for a particular capability? Is it enough to request and check success? Would the capability be likely to change, and if so, how would you notify the app asynchronously?
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-09-12 23:40 +0200 |
| Message-ID | <sgH7j-FR-21@gated-at.bofh.it> |
| In reply to | #1481021 |
On Mon, Sep 12, 2016 at 06:05:07PM +1000, Nicholas Piggin wrote: > On Mon, 12 Sep 2016 00:51:28 -0700 > Christoph Hellwig <hch@infradead.org> wrote: > > > On Mon, Sep 12, 2016 at 05:25:15PM +1000, Oliver O'Halloran wrote: > > > What are the problems here? Is this a matter of existing filesystems > > > being unable/unwilling to support this or is it just fundamentally > > > broken? > > > > It's a fundamentally broken model. See Dave's post that actually was > > sent slightly earlier then mine for the list of required items, which > > is fairly unrealistic. You could probably try to architect a file > > system for it, but I doubt it would gain much traction. > > It's not fundamentally broken, it just doesn't fit well existing > filesystems. > > Dave's post of requirements is also wrong. A filesystem does not have > to guarantee all that, it only has to guarantee that is the case for > a given block after it has a mapping and page fault returns, other > operations can be supported by invalidating mappings, etc. Sure, but filesystems are completely unaware of what is mapped at any given time, or what constraints that mapping might have. Trying to make filesystems aware of per-page mapping constraints seems like a fairly significant layering violation based on a flawed assumption. i.e. that operations on other parts of the file do not affect the block that requires immutable metadata. e.g an extent operation in some other area of the file can cause a tip-to-root extent tree split or merge, and that moves the metadata that points to the mapped block that we've told userspace "doesn't need fsync". We now need an fsync to ensure that the metadata is consistent on disk again, even though that block has not physically been moved. IOWs, the immutable data block updates are now not ordered correctly w.r.t. other updates done to the file, especially when we consider crash recovery.... All this will expose is an unfixable problem with ordering of stable data + metadata operations and their synchronisation. As such, it seems like nothing but a major cluster-fuck to try to do mapping specific, per-block immutable metadata - it adds major complexity and even more untractable problems. Yes, we /could/ try to solve this but, quite frankly, it's far easier to change the broken PMEM programming model assumptions than it is to implement what you are suggesting. Or to do what Christoph suggested and just use a wrapper around something like device mapper to hand out chunks of unchanging, static pmem to applications... Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-13 04:00 +0200 |
| Message-ID | <sgLaW-3gi-3@gated-at.bofh.it> |
| In reply to | #1481993 |
On Tue, 13 Sep 2016 07:34:36 +1000 Dave Chinner <david@fromorbit.com> wrote: > On Mon, Sep 12, 2016 at 06:05:07PM +1000, Nicholas Piggin wrote: > > On Mon, 12 Sep 2016 00:51:28 -0700 > > Christoph Hellwig <hch@infradead.org> wrote: > > > > > On Mon, Sep 12, 2016 at 05:25:15PM +1000, Oliver O'Halloran wrote: > > > > What are the problems here? Is this a matter of existing filesystems > > > > being unable/unwilling to support this or is it just fundamentally > > > > broken? > > > > > > It's a fundamentally broken model. See Dave's post that actually was > > > sent slightly earlier then mine for the list of required items, which > > > is fairly unrealistic. You could probably try to architect a file > > > system for it, but I doubt it would gain much traction. > > > > It's not fundamentally broken, it just doesn't fit well existing > > filesystems. > > > > Dave's post of requirements is also wrong. A filesystem does not have > > to guarantee all that, it only has to guarantee that is the case for > > a given block after it has a mapping and page fault returns, other > > operations can be supported by invalidating mappings, etc. > > Sure, but filesystems are completely unaware of what is mapped at > any given time, or what constraints that mapping might have. Trying > to make filesystems aware of per-page mapping constraints seems like I'm not sure what you mean. The filesystem can hand out mappings and fault them in itself. It can invalidate them. > a fairly significant layering violation based on a flawed > assumption. i.e. that operations on other parts of the file do not > affect the block that requires immutable metadata. > > e.g an extent operation in some other area of the file can cause a > tip-to-root extent tree split or merge, and that moves the metadata > that points to the mapped block that we've told userspace "doesn't > need fsync". We now need an fsync to ensure that the metadata is > consistent on disk again, even though that block has not physically > been moved. You don't, because the filesystem can invalidate existing mappings and do the right thing when they are faulted in again. That's the big^Wmedium hammer approach that can cope with most problems. But let me understand your example in the absence of that. - Application mmaps a file, faults in block 0 - FS allocates block, creates mappings, syncs metadata, sets "no fsync" flag for that block, and completes the fault. - Application writes some data to block 0, completes userspace flushes * At this point, a crash must return with above data (or newer). - Application starts writing more stuff into block 0 - Concurrently, fault in block 1 - FS starts to allocate, splits trees including mappings to block 0 * Crash Is that right? How does your filesystem lose data before the sync point? > IOWs, the immutable data block updates are now not > ordered correctly w.r.t. other updates done to the file, especially > when we consider crash recovery.... > > All this will expose is an unfixable problem with ordering of stable > data + metadata operations and their synchronisation. As such, it > seems like nothing but a major cluster-fuck to try to do mapping > specific, per-block immutable metadata - it adds major complexity > and even more untractable problems. > > Yes, we /could/ try to solve this but, quite frankly, it's far > easier to change the broken PMEM programming model assumptions than > it is to implement what you are suggesting. Or to do what Christoph > suggested and just use a wrapper around something like device > mapper to hand out chunks of unchanging, static pmem to > applications... If there is any huge complexity or unsolved problem, it is in XFS. Conceptual problem is simple.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-13 09:20 +0200 |
| Message-ID | <sgQaC-72V-23@gated-at.bofh.it> |
| In reply to | #1482084 |
On Tue, Sep 13, 2016 at 11:53:11AM +1000, Nicholas Piggin wrote: > - Application mmaps a file, faults in block 0 > - FS allocates block, creates mappings, syncs metadata, sets "no fsync" > flag for that block, and completes the fault. > - Application writes some data to block 0, completes userspace flushes > > * At this point, a crash must return with above data (or newer). > > - Application starts writing more stuff into block 0 > - Concurrently, fault in block 1 > - FS starts to allocate, splits trees including mappings to block 0 > > * Crash > > Is that right? How does your filesystem lose data before the sync > point? Witht all current file systems chances are your metadata hasn't been flushed out. You could write all metadata synchronously from the page fault handler, but that's basically asking for all kinds of deadlocks. > If there is any huge complexity or unsolved problem, it is in XFS. > Conceptual problem is simple. Good to have you back and make all the hard thing simple :)
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-13 11:10 +0200 |
| Message-ID | <sgRT3-8gk-9@gated-at.bofh.it> |
| In reply to | #1482194 |
On Tue, 13 Sep 2016 00:17:32 -0700 Christoph Hellwig <hch@infradead.org> wrote: > On Tue, Sep 13, 2016 at 11:53:11AM +1000, Nicholas Piggin wrote: > > - Application mmaps a file, faults in block 0 > > - FS allocates block, creates mappings, syncs metadata, sets "no fsync" > > flag for that block, and completes the fault. > > - Application writes some data to block 0, completes userspace flushes > > > > * At this point, a crash must return with above data (or newer). > > > > - Application starts writing more stuff into block 0 > > - Concurrently, fault in block 1 > > - FS starts to allocate, splits trees including mappings to block 0 > > > > * Crash > > > > Is that right? How does your filesystem lose data before the sync > > point? > > Witht all current file systems chances are your metadata hasn't been > flushed out. You could write all metadata synchronously from the Yes, that's a possibility. Another would be an advise call to request the capability for a given region. > page fault handler, but that's basically asking for all kinds of > deadlocks. Such as? > > If there is any huge complexity or unsolved problem, it is in XFS. > > Conceptual problem is simple. > > Good to have you back and make all the hard thing simple :) Thanks...? :) I don't mean to say it's simple to add it to any filesystem or that vfs and mm doesn't need any changes at all. If we can agree on something, no new APIs should be added without careful thought and justification and users. I only suggest not to dismiss it out of hand. Thanks, Nick
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-09-14 09:40 +0200 |
| Message-ID | <shcXw-6ki-35@gated-at.bofh.it> |
| In reply to | #1482084 |
On Tue, Sep 13, 2016 at 11:53:11AM +1000, Nicholas Piggin wrote: > On Tue, 13 Sep 2016 07:34:36 +1000 > Dave Chinner <david@fromorbit.com> wrote: > But let me understand your example in the absence of that. > > - Application mmaps a file, faults in block 0 > - FS allocates block, creates mappings, syncs metadata, sets "no fsync" > flag for that block, and completes the fault. > - Application writes some data to block 0, completes userspace flushes > > * At this point, a crash must return with above data (or newer). > > - Application starts writing more stuff into block 0 > - Concurrently, fault in block 1 > - FS starts to allocate, splits trees including mappings to block 0 > > * Crash > > Is that right? No. - app write faults block 0, fs allocates < time passes while app does stuff to block 0 mapping > - fs syncs journal, block 0 metadata now persistent < time passes while app does stuff to block 0 mapping > - app structure grows, faults block 1, fs allocates - app adds pointers to data in block 1 from block 0, does userspace pmem data sync. *crash* > How does your filesystem lose data before the sync > point? After recovery, file has a data in block 0, but no block 1 because the allocation transaction for block 1 was not flushed to the journal. Data in block 0 points to data in block 1, but block 1 does not exist. IOWs, the application has corrupt data because it never issued a data synchronisation request to the filesystem.... ---- Ok, looking back over your example, you seem to be suggesting a new page fault behaviour is required from filesystems that has not been described or explained, and that behaviour is triggered (persistently) somehow from userspace. You've also suggested filesystems store a persistent per-block "no fsync" flag in their extent map as part of the implementation. Right? Reading between the lines, I'm guessing that the "no fsync" flag has very specific update semantics, constraints and requirements. Can you outline how you expect this flag to be set and updated, how it's used consistently between different applications (e.g. cp of a file vs the app using the file), behavioural constraints it implies for page faults vs non-mmap access to the data in the block, how you'd expect filesystems to deal with things like a hole punch landing in the middle of an extent marked with "no fsync", etc? [snip] > If there is any huge complexity or unsolved problem, it is in XFS. > Conceptual problem is simple. Play nice and be constructive, please? Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-14 12:20 +0200 |
| Message-ID | <shfsm-830-13@gated-at.bofh.it> |
| In reply to | #1483006 |
On Wed, 14 Sep 2016 17:39:02 +1000 Dave Chinner <david@fromorbit.com> wrote: > On Tue, Sep 13, 2016 at 11:53:11AM +1000, Nicholas Piggin wrote: > > On Tue, 13 Sep 2016 07:34:36 +1000 > > Dave Chinner <david@fromorbit.com> wrote: > > But let me understand your example in the absence of that. > > > > - Application mmaps a file, faults in block 0 > > - FS allocates block, creates mappings, syncs metadata, sets "no fsync" > > flag for that block, and completes the fault. > > - Application writes some data to block 0, completes userspace flushes > > > > * At this point, a crash must return with above data (or newer). > > > > - Application starts writing more stuff into block 0 > > - Concurrently, fault in block 1 > > - FS starts to allocate, splits trees including mappings to block 0 > > > > * Crash > > > > Is that right? > > No. > > - app write faults block 0, fs allocates > < time passes while app does stuff to block 0 mapping > > - fs syncs journal, block 0 metadata now persistent > < time passes while app does stuff to block 0 mapping > > - app structure grows, faults block 1, fs allocates > - app adds pointers to data in block 1 from block 0, does > userspace pmem data sync. > > *crash* > > > How does your filesystem lose data before the sync > > point? > > After recovery, file has a data in block 0, but no block 1 because > the allocation transaction for block 1 was not flushed to the > journal. Data in block 0 points to data in block 1, but block 1 does > not exist. IOWs, the application has corrupt data because it never > issued a data synchronisation request to the filesystem.... > > ---- > > Ok, looking back over your example, you seem to be suggesting a new > page fault behaviour is required from filesystems that has not been > described or explained, and that behaviour is triggered > (persistently) somehow from userspace. You've also suggested > filesystems store a persistent per-block "no fsync" flag > in their extent map as part of the implementation. Right? This is what we're talking about. Of course a filesystem can't just start supporting the feature without any changes. > Reading between the lines, I'm guessing that the "no fsync" flag has > very specific update semantics, constraints and requirements. Can > you outline how you expect this flag to be set and updated, how it's > used consistently between different applications (e.g. cp of a file > vs the app using the file), behavioural constraints it implies for > page faults vs non-mmap access to the data in the block, how > you'd expect filesystems to deal with things like a hole punch > landing in the middle of an extent marked with "no fsync", etc? Well that's what's being discussed. An approach close to what I did is to allow the app request a "no sync" type of mmap. Filesystem will invalidate all such mappings before it does buffered IOs or hole punch, and will sync metadata after allocating a new block but before returning from a fault. The app could query rather than request, but I found request seemed to work better. The filesystem might be working with apps that don't use the feature for example, and doesn't want to flush just in case any one ever queried in the past. > [snip] > > > If there is any huge complexity or unsolved problem, it is in XFS. > > Conceptual problem is simple. > > Play nice and be constructive, please? So you agree that the persistent memory people who have come with some requirements and ideas for an API should not be immediately shut down with bogus handwaving. Thanks, Nick
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-09-15 04:40 +0200 |
| Message-ID | <shuKJ-Zq-1@gated-at.bofh.it> |
| In reply to | #1483153 |
On Wed, Sep 14, 2016 at 08:19:36PM +1000, Nicholas Piggin wrote: > On Wed, 14 Sep 2016 17:39:02 +1000 > Dave Chinner <david@fromorbit.com> wrote: > > Ok, looking back over your example, you seem to be suggesting a new > > page fault behaviour is required from filesystems that has not been > > described or explained, and that behaviour is triggered > > (persistently) somehow from userspace. You've also suggested > > filesystems store a persistent per-block "no fsync" flag > > in their extent map as part of the implementation. Right? > > This is what we're talking about. Of course a filesystem can't just > start supporting the feature without any changes. Sure, but one first has to describe the feature desired before all parties can discuss it. We need more than vague references and allusions from you to define the solution you are proposing. Once everyone understands what is being describing, we might be able to work out how it can be implemented in a simple, generic manner rather than require every filesystem to change their on-disk formats. IOWs, we need you to describe /details/ of semantics, behaviour and data integrity constraints that are required, not describe an implementation of something we have no knwoledge about. > > Reading between the lines, I'm guessing that the "no fsync" flag has > > very specific update semantics, constraints and requirements. Can > > you outline how you expect this flag to be set and updated, how it's > > used consistently between different applications (e.g. cp of a file > > vs the app using the file), behavioural constraints it implies for > > page faults vs non-mmap access to the data in the block, how > > you'd expect filesystems to deal with things like a hole punch > > landing in the middle of an extent marked with "no fsync", etc? > > Well that's what's being discussed. An approach close to what I did is > to allow the app request a "no sync" type of mmap. That's not an answer to the questions I asked about about the "no sync" flag you were proposing. You've redirected to the a different solution, one that .... > Filesystem will > invalidate all such mappings before it does buffered IOs or hole punch, > and will sync metadata after allocating a new block but before returning > from a fault. ... requires synchronous metadata updates from page fault context, which we already know is not a good solution. I'll quote one of Christoph's previous replies to save me the trouble: "You could write all metadata synchronously from the page fault handler, but that's basically asking for all kinds of deadlocks." So, let's redirect back to the "no sync" flag you were talking about - can you answer the questions I asked above? It would be especially important to highlight how the proposed feature would avoid requiring synchronous metadata updates in page fault contexts.... > > [snip] > > > > > If there is any huge complexity or unsolved problem, it is in XFS. > > > Conceptual problem is simple. > > > > Play nice and be constructive, please? > > So you agree that the persistent memory people who have come with some > requirements and ideas for an API should not be immediately shut down > with bogus handwaving. Pull your head in, Nick. You've been absent from the community for the last 5 years. You suddenly barge in with a massive chip on your shoulder and try to throw your weight around. You're being arrogant, obnoxious, evasive and petty. You're belittling anyone who dares to question your proclamations. You're not listening to the replies you are getting. You're baiting people to try to get an adverse reaction from them and when someone gives you the adverse reaction you were fishing for, you play the victim card. That's textbook bullying behaviour. Nick, this behaviour does not help progress the discussion in any way. It only serves to annoy the other people who are sincerely trying to understand and determine if/how we can solve the problem in some way. So, again, play nice and be constructive, please? Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-15 06:00 +0200 |
| Message-ID | <shw09-1Gg-1@gated-at.bofh.it> |
| In reply to | #1483831 |
On Thu, 15 Sep 2016 12:31:33 +1000 Dave Chinner <david@fromorbit.com> wrote: > On Wed, Sep 14, 2016 at 08:19:36PM +1000, Nicholas Piggin wrote: > > On Wed, 14 Sep 2016 17:39:02 +1000 > > Dave Chinner <david@fromorbit.com> wrote: > > > Ok, looking back over your example, you seem to be suggesting a new > > > page fault behaviour is required from filesystems that has not been > > > described or explained, and that behaviour is triggered > > > (persistently) somehow from userspace. You've also suggested > > > filesystems store a persistent per-block "no fsync" flag > > > in their extent map as part of the implementation. Right? > > > > This is what we're talking about. Of course a filesystem can't just > > start supporting the feature without any changes. > > Sure, but one first has to describe the feature desired before all The DAX people have been. They want to be able to get mappings that can be synced without doing fsync. The *exact* extent of those capabilities and what the API exactly looks like is up for discussion. > parties can discuss it. We need more than vague references and > allusions from you to define the solution you are proposing. > > Once everyone understands what is being describing, we might be able > to work out how it can be implemented in a simple, generic manner > rather than require every filesystem to change their on-disk > formats. IOWs, we need you to describe /details/ of semantics, > behaviour and data integrity constraints that are required, not > describe an implementation of something we have no knwoledge about. Well you said it was impossible already and Christoph told them they were smoking crack :) Anyway, there's a few questions. Implementation and API. Some filesystems may never cope with it. Of course it should be as generic as possible though. > > > Reading between the lines, I'm guessing that the "no fsync" flag has > > > very specific update semantics, constraints and requirements. Can > > > you outline how you expect this flag to be set and updated, how it's > > > used consistently between different applications (e.g. cp of a file > > > vs the app using the file), behavioural constraints it implies for > > > page faults vs non-mmap access to the data in the block, how > > > you'd expect filesystems to deal with things like a hole punch > > > landing in the middle of an extent marked with "no fsync", etc? > > > > Well that's what's being discussed. An approach close to what I did is > > to allow the app request a "no sync" type of mmap. > > That's not an answer to the questions I asked about about the "no > sync" flag you were proposing. You've redirected to the a different > solution, one that .... No sync flag would do the same thing exactly in terms of consistency. It would just do the no-sync sequence by default rather than being asked for it. More of an API detail than implementation. > > > Filesystem will > > invalidate all such mappings before it does buffered IOs or hole punch, > > and will sync metadata after allocating a new block but before returning > > from a fault. > > ... requires synchronous metadata updates from page fault context, > which we already know is not a good solution. I'll quote one of > Christoph's previous replies to save me the trouble: > > "You could write all metadata synchronously from the page > fault handler, but that's basically asking for all kinds of > deadlocks." > So, let's redirect back to the "no sync" flag you were talking about > - can you answer the questions I asked above? It would be especially > important to highlight how the proposed feature would avoid requiring > synchronous metadata updates in page fault contexts.... Right. So what deadlocks are you concerned about? There could be a scale of capabilities here, for different filesystems that do things differently. Some filesystems could require fsync for metadata, but allow fdatasync to be skipped. Users would need to have some knowledge of block size or do preallocation and sync. That might put more burden on libraries/applications if there are concurrent operations, but that might be something they can deal with -- fdatasync already requires some knowledge of concurrent operations (or lack thereof). > > > [snip] > > > > > > > If there is any huge complexity or unsolved problem, it is in XFS. > > > > Conceptual problem is simple. > > > > > > Play nice and be constructive, please? > > > > So you agree that the persistent memory people who have come with some > > requirements and ideas for an API should not be immediately shut down > > with bogus handwaving. > > Pull your head in, Nick. > > You've been absent from the community for the last 5 years. You > suddenly barge in with a massive chip on your shoulder and try to I'm trying to give some constructive input to the nvdimm guys. You and Christoph know a huge amount about vfs and filesystems. But sometimes you shut people down prematurely. It can be very intimidating for someone who might not know *exactly* what they are asking for or have not considered some difficult locking case in a filesystem. I'm sure it's not intentional, but that's how it can come across. That said, I don't want to derail their thread any further with this. So I apologise for my tone to you, Dave. Thanks, Nick
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-09-15 12:40 +0200 |
| Message-ID | <shCfg-5Oi-29@gated-at.bofh.it> |
| In reply to | #1483843 |
On Thu, Sep 15, 2016 at 01:49:45PM +1000, Nicholas Piggin wrote: > On Thu, 15 Sep 2016 12:31:33 +1000 > Dave Chinner <david@fromorbit.com> wrote: > > > On Wed, Sep 14, 2016 at 08:19:36PM +1000, Nicholas Piggin wrote: > > > On Wed, 14 Sep 2016 17:39:02 +1000 > > Sure, but one first has to describe the feature desired before all > > The DAX people have been. Hmmmm. the only "DAX people" I know of are kernel developers who have been working on implementing DAX in the kernel - Willy, Ross, Dan, Jan, Christoph, Kirill, myelf and a few others around the fringes. > They want to be able to get mappings > that can be synced without doing fsync. Oh, ok, the Intel Userspace PMEM library requirement. I though you had something more that this - whatever extra problem the per-block no fsync flag would solve? > The *exact* extent of > those capabilities and what the API exactly looks like is up for > discussion. Yup. > Well you said it was impossible already and Christoph told them > they were smoking crack :) I have not said that. I have said bad things about bad proposals and called the PMEM library model broken, but I most definitely have not said that solving the problem is impossible. > > That's not an answer to the questions I asked about about the "no > > sync" flag you were proposing. You've redirected to the a different > > solution, one that .... > > No sync flag would do the same thing exactly in terms of consistency. > It would just do the no-sync sequence by default rather than being > asked for it. More of an API detail than implementation. You still haven't described anything about what a per-block flag design is supposed to look like.... :/ > > > Filesystem will > > > invalidate all such mappings before it does buffered IOs or hole punch, > > > and will sync metadata after allocating a new block but before returning > > > from a fault. > > > > ... requires synchronous metadata updates from page fault context, > > which we already know is not a good solution. I'll quote one of > > Christoph's previous replies to save me the trouble: > > > > "You could write all metadata synchronously from the page > > fault handler, but that's basically asking for all kinds of > > deadlocks." > > So, let's redirect back to the "no sync" flag you were talking about > > - can you answer the questions I asked above? It would be especially > > important to highlight how the proposed feature would avoid requiring > > synchronous metadata updates in page fault contexts.... > > Right. So what deadlocks are you concerned about? It basically puts the entire journal checkpoint path under a page fault context. i.e. a whole new global locking context problem is created as this path can now be run both inside and outside the mmap_sem. Nothing ever good comes from running filesystem locking code both inside and outside the mmap_sem. FWIW, We've never executed synchronous transactions inside page faults in XFS, and I think ext4 is in the same boat - it may be even worse because of the way it does ordered data dispatch through the journal. I don't really even want to think about the level of hurt this might put btrfs or other COW/log structured filesystems under. I'm sure Christoph can reel off a bunch more issues off the top of his head.... > There could be a scale of capabilities here, for different filesystems > that do things differently. Why do we need such complexity to be defined? I'm tending towards just adding new fallocate() operation that sets up a fully allocated and zeroed file of fixed length that has immutable metadata once created. Single syscall, with well dfined semantics, and it doesn't dictate the implementation any filesystem must use. All it dictates is that the data region can be written safely on dax-enabled storage without needing fsync() to be issued. i.e. the implementation can be filesystem specific, and it is simple to implement the basic functionality and constraints in both XFS and ext4 right away, and as othe filesystems come along they can implement it in the way that best suits them. e.g. btrfs would need to add the "no COW" flag to the file as well. If someone wants to implement a per-block no-fsync flag, and do sycnhronous metadata updates in the page fault path, then they are welcome to do so. But we don't /need/ such complexity to implement the functionality that pmem programming model requires. > Some filesystems could require fsync for metadata, but allow fdatasync > to be skipped. Users would need to have some knowledge of block size > or do preallocation and sync. Not sure what you mean here - avoiding the need for using fsync() by using fsync() seems a little circular to me. :/ > That might put more burden on libraries/applications if there are > concurrent operations, but that might be something they can deal with > -- fdatasync already requires some knowledge of concurrent operations > (or lack thereof). Additional userspace complexity is something we should avoid. > You and Christoph know a huge amount about vfs and filesystems. > But sometimes you shut people down prematurely. Appearances can be deceiving. I don't shut discussions down unless my time is being wasted, and that's pretty rare. [You probably know most of what I'm about to write, but I'm not actually writing it for you.... ] > It can be very > intimidating for someone who might not know *exactly* what they > are asking for or have not considered some difficult locking case > in a filesystem. Yup, most kernel developers are aware that this is how the mailing list discussions appear from the outside. Unfortunately, too many people think they have expert knowlege when they don't (the Dunning-Kruger cognitive bias), and so they simply can't understand an expert response that points out problems 5 or 6 steps further along the logic chain and assumes the reader knows that chain intimately. They don't so they think they've been shut down. It's closer to the truth that they've suddenly been made aware of how little they know about the topic under discussion and they don't know how to react. At this point, we see a classic fight-or-flight reflex response, and then the person either runs away scared or create heat and light.... Occasionally, someone comes back with "hey, I don't quite understand that, can you explain it more/differently", then I will try to explain it as best I can and have time to do so. History has proven that most kernel developers react in this way, including Christoph. And when this happens, we often end up with a new contributor.... As Confuscius says: "Real knowledge is to know the extent of one's ignorance." [ I talk about this whole issue in more detail mid-way through this LCA 2015 presentation: https://www.youtube.com/watch?v=VpuVDfSXs-g ] This is why we tend to ask people for their problem descriptions, requirements and constraints, rather than expecting them to explain the how they think their problem needs to be solved. We do not expect anyone other than the regular kernel developers to understand deep, dark details needed to craft a workable solution. If they do, great. if they don't, then don't get upset or angry about it. Instead, while one may not like or fully understand the answer that is given, accept the answers that are given and respect that there is usually a very good reason that answer was given. > I'm sure it's not intentional, but that's how it > can come across. Yup, and there's really very little we can do about it. It's one of the consequences of having experts hang around in public.... > That said, I don't want to derail their thread any further with > this. So I apologise for my tone to you, Dave. Accepted. Let's start over, eh? Cheers, Dave. -- Dave Chinner david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-15 13:50 +0200 |
| Message-ID | <shDkZ-6rX-5@gated-at.bofh.it> |
| In reply to | #1483960 |
On Thu, 15 Sep 2016 20:32:10 +1000 Dave Chinner <david@fromorbit.com> wrote: > On Thu, Sep 15, 2016 at 01:49:45PM +1000, Nicholas Piggin wrote: > > On Thu, 15 Sep 2016 12:31:33 +1000 > > Dave Chinner <david@fromorbit.com> wrote: > > > > > On Wed, Sep 14, 2016 at 08:19:36PM +1000, Nicholas Piggin wrote: > > > > On Wed, 14 Sep 2016 17:39:02 +1000 > > > Sure, but one first has to describe the feature desired before all > > > > The DAX people have been. > > Hmmmm. the only "DAX people" I know of are kernel developers who > have been working on implementing DAX in the kernel - Willy, Ross, > Dan, Jan, Christoph, Kirill, myelf and a few others around the > fringes. > > > They want to be able to get mappings > > that can be synced without doing fsync. > > Oh, ok, the Intel Userspace PMEM library requirement. I though you > had something more that this - whatever extra problem the per-block > no fsync flag would solve? Only the PMEM really. I don't want to add more complexity than required. > > The *exact* extent of > > those capabilities and what the API exactly looks like is up for > > discussion. > > Yup. > > > Well you said it was impossible already and Christoph told them > > they were smoking crack :) > > I have not said that. I have said bad things about bad > proposals and called the PMEM library model broken, but I most > definitely have not said that solving the problem is impossible. > > > > That's not an answer to the questions I asked about about the "no > > > sync" flag you were proposing. You've redirected to the a different > > > solution, one that .... > > > > No sync flag would do the same thing exactly in terms of consistency. > > It would just do the no-sync sequence by default rather than being > > asked for it. More of an API detail than implementation. > > You still haven't described anything about what a per-block flag > design is supposed to look like.... :/ For the API, or implementation? I'm not quite sure what you mean here. For implementation it's possible to carefully ensure metadata is persistent when allocating blocks in page fault but before mapping pages. Truncate or hole punch or such things can be made to work by invalidating all such mappings and holding them off until you can cope with them again. Not necessarily for a filesystem with *all* capabilities of XFS -- I don't know -- but for a complete basic one. > > > > Filesystem will > > > > invalidate all such mappings before it does buffered IOs or hole punch, > > > > and will sync metadata after allocating a new block but before returning > > > > from a fault. > > > > > > ... requires synchronous metadata updates from page fault context, > > > which we already know is not a good solution. I'll quote one of > > > Christoph's previous replies to save me the trouble: > > > > > > "You could write all metadata synchronously from the page > > > fault handler, but that's basically asking for all kinds of > > > deadlocks." > > > So, let's redirect back to the "no sync" flag you were talking about > > > - can you answer the questions I asked above? It would be especially > > > important to highlight how the proposed feature would avoid requiring > > > synchronous metadata updates in page fault contexts.... > > > > Right. So what deadlocks are you concerned about? > > It basically puts the entire journal checkpoint path under a page > fault context. i.e. a whole new global locking context problem is Yes there are potentially some new lock orderings created if you do that, depending on what locks the filesystem does. > created as this path can now be run both inside and outside the > mmap_sem. Nothing ever good comes from running filesystem locking > code both inside and outside the mmap_sem. You mean that some cases journal checkpoint runs with mmap_sem held, and others without mmap_sem held? Not that mmap_sem is taken inside journal checkpoint. Then I don't really see why that's a problem. I mean performance could suffer a bit, but with fault retry you can almost always do the syncing outside mmap_sem in practice. Yes, I'll preemptively agree with you -- We don't want to add any such burden if it is not needed and well justified. > FWIW, We've never executed synchronous transactions inside page > faults in XFS, and I think ext4 is in the same boat - it may be even > worse because of the way it does ordered data dispatch through the > journal. I don't really even want to think about the level of hurt > this might put btrfs or other COW/log structured filesystems under. > I'm sure Christoph can reel off a bunch more issues off the top of > his head.... I asked him, we'll see what he thinks. I don't beleive there is anything fundamental about mm or fs core layers that cause deadlocks though. > > > There could be a scale of capabilities here, for different filesystems > > that do things differently. > > Why do we need such complexity to be defined? > > I'm tending towards just adding new fallocate() operation that sets > up a fully allocated and zeroed file of fixed length that has > immutable metadata once created. Single syscall, with well dfined > semantics, and it doesn't dictate the implementation any filesystem > must use. All it dictates is that the data region can be written > safely on dax-enabled storage without needing fsync() to be issued. > > i.e. the implementation can be filesystem specific, and it is simple > to implement the basic functionality and constraints in both XFS and > ext4 right away, and as othe filesystems come along they can > implement it in the way that best suits them. e.g. btrfs would need > to add the "no COW" flag to the file as well. > > If someone wants to implement a per-block no-fsync flag, and do > sycnhronous metadata updates in the page fault path, then they are > welcome to do so. But we don't /need/ such complexity to implement > the functionality that pmem programming model requires. Sure. That's about what I meant by scale of capabilities. And beyond that I would be as happy as you if that was sufficient. It raises the bar for justifying more complexity, which is always good. > > Some filesystems could require fsync for metadata, but allow fdatasync > > to be skipped. Users would need to have some knowledge of block size > > or do preallocation and sync. > > Not sure what you mean here - avoiding the need for using fsync() > by using fsync() seems a little circular to me. :/ I meant if blocks are already preallocated and metadata unchanging, basically like your above proposal. > > > That might put more burden on libraries/applications if there are > > concurrent operations, but that might be something they can deal with > > -- fdatasync already requires some knowledge of concurrent operations > > (or lack thereof). > > Additional userspace complexity is something we should avoid. > > > > You and Christoph know a huge amount about vfs and filesystems. > > But sometimes you shut people down prematurely. > > Appearances can be deceiving. I don't shut discussions down unless > my time is being wasted, and that's pretty rare. > > [You probably know most of what I'm about to write, but I'm not > actually writing it for you.... ] > > > It can be very > > intimidating for someone who might not know *exactly* what they > > are asking for or have not considered some difficult locking case > > in a filesystem. > > Yup, most kernel developers are aware that this is how the mailing > list discussions appear from the outside. [...] Point taken and I don't want to harp on about it. You guys can still be intimidating to good kernel programmers who otherwise don't know vfs and several filesystems inside out, that's all. > > That said, I don't want to derail their thread any further with > > this. So I apologise for my tone to you, Dave. > > Accepted. Let's start over, eh? That would be good. Thanks, Nick
[toc] | [prev] | [next] | [standalone]
| From | Dave Chinner <david@fromorbit.com> |
|---|---|
| Date | 2016-09-16 00:40 +0200 |
| Message-ID | <shNu1-4wD-9@gated-at.bofh.it> |
| In reply to | #1484018 |
On Thu, Sep 15, 2016 at 09:42:22PM +1000, Nicholas Piggin wrote:
> On Thu, 15 Sep 2016 20:32:10 +1000
> Dave Chinner <david@fromorbit.com> wrote:
> >
> > You still haven't described anything about what a per-block flag
> > design is supposed to look like.... :/
>
> For the API, or implementation? I'm not quite sure what you mean
> here. For implementation it's possible to carefully ensure metadata
> is persistent when allocating blocks in page fault but before
> mapping pages. Truncate or hole punch or such things can be made to
> work by invalidating all such mappings and holding them off until
> you can cope with them again. Not necessarily for a filesystem with
> *all* capabilities of XFS -- I don't know -- but for a complete basic
> one.
SO, essentially, it comes down to synchrnous metadta updates again.
but synchronous updates would be conditional on whether an extent
metadata with the "nofsync" flag asserted was updated? Where's the
nofsync flag kept? in memory at a generic layer, or in the
filesystem, potentially in an on-disk structure? How would the
application set it for a given range?
> > > > > Filesystem will
> > > > > invalidate all such mappings before it does buffered IOs or hole punch,
> > > > > and will sync metadata after allocating a new block but before returning
> > > > > from a fault.
> > > >
> > > > ... requires synchronous metadata updates from page fault context,
> > > > which we already know is not a good solution. I'll quote one of
> > > > Christoph's previous replies to save me the trouble:
> > > >
> > > > "You could write all metadata synchronously from the page
> > > > fault handler, but that's basically asking for all kinds of
> > > > deadlocks."
> > > > So, let's redirect back to the "no sync" flag you were talking about
> > > > - can you answer the questions I asked above? It would be especially
> > > > important to highlight how the proposed feature would avoid requiring
> > > > synchronous metadata updates in page fault contexts....
> > >
> > > Right. So what deadlocks are you concerned about?
> >
> > It basically puts the entire journal checkpoint path under a page
> > fault context. i.e. a whole new global locking context problem is
>
> Yes there are potentially some new lock orderings created if you
> do that, depending on what locks the filesystem does.
Well, that's the whole issue.
> > created as this path can now be run both inside and outside the
> > mmap_sem. Nothing ever good comes from running filesystem locking
> > code both inside and outside the mmap_sem.
>
> You mean that some cases journal checkpoint runs with mmap_sem
> held, and others without mmap_sem held? Not that mmap_sem is taken
> inside journal checkpoint.
Maybe not, but now we open up the potential for locks held inside
or outside mmap sem to interact with the journal locks that are now
held inside and outside mmap_sem. See below....
> Then I don't really see why that's a
> problem. I mean performance could suffer a bit, but with fault
> retry you can almost always do the syncing outside mmap_sem in
> practice.
>
> Yes, I'll preemptively agree with you -- We don't want to add any
> such burden if it is not needed and well justified.
>
> > FWIW, We've never executed synchronous transactions inside page
> > faults in XFS, and I think ext4 is in the same boat - it may be even
> > worse because of the way it does ordered data dispatch through the
> > journal. I don't really even want to think about the level of hurt
> > this might put btrfs or other COW/log structured filesystems under.
> > I'm sure Christoph can reel off a bunch more issues off the top of
> > his head....
>
> I asked him, we'll see what he thinks. I don't beleive there is
> anything fundamental about mm or fs core layers that cause deadlocks
> though.
Spent 5 minutes looking at ext4 for an example: filesystems are
allowed to take page locks during transaction commit. e.g ext4
journal commit when using the default ordered data mode:
jbd2_journal_commit_transaction
journal_submit_data_buffers()
journal_submit_inode_data_buffers
generic_writepages()
ext4_writepages()
mpage_prepare_extent_to_map()
lock_page()
i.e. if we fault on the user buffer during a write() operation and
that user buffer is a mmapped DAX file that needs to be allocated
and we have synchronous metadata updates during page faults, we
deadlock on the page lock held above the page fault context...
Cheers,
Dave.
--
Dave Chinner
david@fromorbit.com
[toc] | [prev] | [next] | [standalone]
| From | Nicholas Piggin <npiggin@gmail.com> |
|---|---|
| Date | 2016-09-16 08:00 +0200 |
| Message-ID | <shUlP-ud-3@gated-at.bofh.it> |
| In reply to | #1484597 |
On Fri, 16 Sep 2016 08:33:50 +1000 Dave Chinner <david@fromorbit.com> wrote: > On Thu, Sep 15, 2016 at 09:42:22PM +1000, Nicholas Piggin wrote: > > On Thu, 15 Sep 2016 20:32:10 +1000 > > Dave Chinner <david@fromorbit.com> wrote: > > > > > > You still haven't described anything about what a per-block flag > > > design is supposed to look like.... :/ > > > > For the API, or implementation? I'm not quite sure what you mean > > here. For implementation it's possible to carefully ensure metadata > > is persistent when allocating blocks in page fault but before > > mapping pages. Truncate or hole punch or such things can be made to > > work by invalidating all such mappings and holding them off until > > you can cope with them again. Not necessarily for a filesystem with > > *all* capabilities of XFS -- I don't know -- but for a complete basic > > one. > > SO, essentially, it comes down to synchrnous metadta updates again. Yes. I guess fundamentally you can't get away from either that or preloading at some level. (Also I don't know that there's a sane way to handle [cm]time properly, so some things like that -- this is just about block allocation / avoiding fdatasync). > but synchronous updates would be conditional on whether an extent > metadata with the "nofsync" flag asserted was updated? Where's the > nofsync flag kept? in memory at a generic layer, or in the > filesystem, potentially in an on-disk structure? How would the > application set it for a given range? I guess that comes back to the API. Whether you want it to be persistent, request based, etc. It could be derived type of storage blocks that are mapped there, stored per-inode, in-memory, or on extents on disk. I'm not advocating for a particular API and of course less complexity better. > > > > > > > Filesystem will > > > > > > invalidate all such mappings before it does buffered IOs or hole punch, > > > > > > and will sync metadata after allocating a new block but before returning > > > > > > from a fault. > > > > > > > > > > ... requires synchronous metadata updates from page fault context, > > > > > which we already know is not a good solution. I'll quote one of > > > > > Christoph's previous replies to save me the trouble: > > > > > > > > > > "You could write all metadata synchronously from the page > > > > > fault handler, but that's basically asking for all kinds of > > > > > deadlocks." > > > > > So, let's redirect back to the "no sync" flag you were talking about > > > > > - can you answer the questions I asked above? It would be especially > > > > > important to highlight how the proposed feature would avoid requiring > > > > > synchronous metadata updates in page fault contexts.... > > > > > > > > Right. So what deadlocks are you concerned about? > > > > > > It basically puts the entire journal checkpoint path under a page > > > fault context. i.e. a whole new global locking context problem is > > > > Yes there are potentially some new lock orderings created if you > > do that, depending on what locks the filesystem does. > > Well, that's the whole issue. For filesystem implementations, but perhaps not mm/vfs implemenatation AFAIKS. > > > > created as this path can now be run both inside and outside the > > > mmap_sem. Nothing ever good comes from running filesystem locking > > > code both inside and outside the mmap_sem. > > > > You mean that some cases journal checkpoint runs with mmap_sem > > held, and others without mmap_sem held? Not that mmap_sem is taken > > inside journal checkpoint. > > Maybe not, but now we open up the potential for locks held inside > or outside mmap sem to interact with the journal locks that are now > held inside and outside mmap_sem. See below.... > > > Then I don't really see why that's a > > problem. I mean performance could suffer a bit, but with fault > > retry you can almost always do the syncing outside mmap_sem in > > practice. > > > > Yes, I'll preemptively agree with you -- We don't want to add any > > such burden if it is not needed and well justified. > > > > > FWIW, We've never executed synchronous transactions inside page > > > faults in XFS, and I think ext4 is in the same boat - it may be even > > > worse because of the way it does ordered data dispatch through the > > > journal. I don't really even want to think about the level of hurt > > > this might put btrfs or other COW/log structured filesystems under. > > > I'm sure Christoph can reel off a bunch more issues off the top of > > > his head.... > > > > I asked him, we'll see what he thinks. I don't beleive there is > > anything fundamental about mm or fs core layers that cause deadlocks > > though. > > Spent 5 minutes looking at ext4 for an example: filesystems are > allowed to take page locks during transaction commit. e.g ext4 > journal commit when using the default ordered data mode: > > jbd2_journal_commit_transaction > journal_submit_data_buffers() > journal_submit_inode_data_buffers > generic_writepages() > ext4_writepages() > mpage_prepare_extent_to_map() > lock_page() > > i.e. if we fault on the user buffer during a write() operation and > that user buffer is a mmapped DAX file that needs to be allocated > and we have synchronous metadata updates during page faults, we > deadlock on the page lock held above the page fault context... Yeah, page lock is probably bigger issue for filesystems than mmap_sem. But still is filesystem implementation detail. Again, I'm not suggesting you could just switch all filesystems today to do a metadata sync with mmap sem and page lock held. Only that there aren't fundamental deadlocks enforced by the mm/vfs. Filesystems are already taking metadata page locks in the read path while holding data page lock, so there's long been some amount of nesting of page lock. It would be possible to change the page fault handler to allow a sync without holding page lock too if it came to it. But I don't want to go to far about implementation handwaving before it's even established that this would be worthwhile. Definitely the first step would be your simple preallocated per inode approach until it is shown to be insufficient.
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web