Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1697451 > unrolled thread

[PATCH v2 0/4] mm/gfs2: extend file_* API, and convert gfs2 to errseq_t error reporting

Started byJeff Layton <jlayton@kernel.org>
First post2017-07-26 20:00 +0200
Last post2017-07-26 20:00 +0200
Articles 11 — 6 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH v2 0/4] mm/gfs2: extend file_* API, and convert gfs2 to errseq_t error reporting Jeff Layton <jlayton@kernel.org> - 2017-07-26 20:00 +0200
    [PATCH v2 1/4] mm: consolidate dax / non-dax checks for writeback Jeff Layton <jlayton@kernel.org> - 2017-07-26 20:00 +0200
      Re: [PATCH v2 1/4] mm: consolidate dax / non-dax checks for writeback Jan Kara <jack@suse.cz> - 2017-07-27 10:50 +0200
    [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync Jeff Layton <jlayton@kernel.org> - 2017-07-26 20:00 +0200
      Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error  reporting for fsync Matthew Wilcox <willy@infradead.org> - 2017-07-26 21:30 +0200
        Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error  reporting for fsync Jeff Layton <jlayton@redhat.com> - 2017-07-27 00:30 +0200
          Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error  reporting for fsync Bob Peterson <rpeterso@redhat.com> - 2017-07-27 14:50 +0200
            Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error  reporting for fsync Steven Whitehouse <swhiteho@redhat.com> - 2017-07-28 14:40 +0200
              Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error  reporting for fsync Jeff Layton <jlayton@redhat.com> - 2017-07-28 14:50 +0200
                Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error  reporting for fsync Steven Whitehouse <swhiteho@redhat.com> - 2017-07-28 15:00 +0200
    [PATCH v2 3/4] fs: convert sync_file_range to use errseq_t based error-tracking Jeff Layton <jlayton@kernel.org> - 2017-07-26 20:00 +0200

#1697451 — [PATCH v2 0/4] mm/gfs2: extend file_* API, and convert gfs2 to errseq_t error reporting

FromJeff Layton <jlayton@kernel.org>
Date2017-07-26 20:00 +0200
Subject[PATCH v2 0/4] mm/gfs2: extend file_* API, and convert gfs2 to errseq_t error reporting
Message-ID<u7yLg-3p3-3@gated-at.bofh.it>
From: Jeff Layton <jlayton@redhat.com>

I sent a small patch earlier this week to make sync_file_range use
errseq_t reporting.

This set respins that patch into a patch that adds a bit more file_*
infrastructure, and then patches to make sync_file_range and fsync
on gfs2 report writeback errors properly.

There's also a small cleanup patch for mm/filemap.c to consolidate
the DAX handling checks in the existing infrastructure.

Jeff Layton (4):
  mm: consolidate dax / non-dax checks for writeback
  mm: add file_fdatawait_range and file_write_and_wait
  fs: convert sync_file_range to use errseq_t based error-tracking
  gfs2: convert to errseq_t based writeback error reporting for fsync

 fs/gfs2/file.c     |  6 +++--
 fs/sync.c          |  4 +--
 include/linux/fs.h |  7 +++++-
 mm/filemap.c       | 71 +++++++++++++++++++++++++++++++++++++++++++++++++-----
 4 files changed, 77 insertions(+), 11 deletions(-)

-- 
2.13.3

[toc] | [next] | [standalone]


#1697454 — [PATCH v2 1/4] mm: consolidate dax / non-dax checks for writeback

FromJeff Layton <jlayton@kernel.org>
Date2017-07-26 20:00 +0200
Subject[PATCH v2 1/4] mm: consolidate dax / non-dax checks for writeback
Message-ID<u7yLg-3p3-11@gated-at.bofh.it>
In reply to#1697451
From: Jeff Layton <jlayton@redhat.com>

We have this complex conditional copied to several places. Turn it into
a helper function.

Signed-off-by: Jeff Layton <jlayton@redhat.com>
---
 mm/filemap.c | 15 +++++++++------
 1 file changed, 9 insertions(+), 6 deletions(-)

diff --git a/mm/filemap.c b/mm/filemap.c
index e1cca770688f..72e46e6f0d9a 100644
--- a/mm/filemap.c
+++ b/mm/filemap.c
@@ -522,12 +522,17 @@ int filemap_fdatawait(struct address_space *mapping)
 }
 EXPORT_SYMBOL(filemap_fdatawait);
 
+static bool mapping_needs_writeback(struct address_space *mapping)
+{
+	return (!dax_mapping(mapping) && mapping->nrpages) ||
+	    (dax_mapping(mapping) && mapping->nrexceptional);
+}
+
 int filemap_write_and_wait(struct address_space *mapping)
 {
 	int err = 0;
 
-	if ((!dax_mapping(mapping) && mapping->nrpages) ||
-	    (dax_mapping(mapping) && mapping->nrexceptional)) {
+	if (mapping_needs_writeback(mapping)) {
 		err = filemap_fdatawrite(mapping);
 		/*
 		 * Even if the above returned error, the pages may be
@@ -566,8 +571,7 @@ int filemap_write_and_wait_range(struct address_space *mapping,
 {
 	int err = 0;
 
-	if ((!dax_mapping(mapping) && mapping->nrpages) ||
-	    (dax_mapping(mapping) && mapping->nrexceptional)) {
+	if (mapping_needs_writeback(mapping)) {
 		err = __filemap_fdatawrite_range(mapping, lstart, lend,
 						 WB_SYNC_ALL);
 		/* See comment of filemap_write_and_wait() */
@@ -656,8 +660,7 @@ int file_write_and_wait_range(struct file *file, loff_t lstart, loff_t lend)
 	int err = 0, err2;
 	struct address_space *mapping = file->f_mapping;
 
-	if ((!dax_mapping(mapping) && mapping->nrpages) ||
-	    (dax_mapping(mapping) && mapping->nrexceptional)) {
+	if (mapping_needs_writeback(mapping)) {
 		err = __filemap_fdatawrite_range(mapping, lstart, lend,
 						 WB_SYNC_ALL);
 		/* See comment of filemap_write_and_wait() */
-- 
2.13.3

[toc] | [prev] | [next] | [standalone]


#1697806 — Re: [PATCH v2 1/4] mm: consolidate dax / non-dax checks for writeback

FromJan Kara <jack@suse.cz>
Date2017-07-27 10:50 +0200
SubjectRe: [PATCH v2 1/4] mm: consolidate dax / non-dax checks for writeback
Message-ID<u7MEy-3Re-9@gated-at.bofh.it>
In reply to#1697454
On Wed 26-07-17 13:55:35, Jeff Layton wrote:
> From: Jeff Layton <jlayton@redhat.com>
> 
> We have this complex conditional copied to several places. Turn it into
> a helper function.
> 
> Signed-off-by: Jeff Layton <jlayton@redhat.com>

Looks good. You can add:

Reviewed-by: Jan Kara <jack@suse.cz>

								Honza

> ---
>  mm/filemap.c | 15 +++++++++------
>  1 file changed, 9 insertions(+), 6 deletions(-)
> 
> diff --git a/mm/filemap.c b/mm/filemap.c
> index e1cca770688f..72e46e6f0d9a 100644
> --- a/mm/filemap.c
> +++ b/mm/filemap.c
> @@ -522,12 +522,17 @@ int filemap_fdatawait(struct address_space *mapping)
>  }
>  EXPORT_SYMBOL(filemap_fdatawait);
>  
> +static bool mapping_needs_writeback(struct address_space *mapping)
> +{
> +	return (!dax_mapping(mapping) && mapping->nrpages) ||
> +	    (dax_mapping(mapping) && mapping->nrexceptional);
> +}
> +
>  int filemap_write_and_wait(struct address_space *mapping)
>  {
>  	int err = 0;
>  
> -	if ((!dax_mapping(mapping) && mapping->nrpages) ||
> -	    (dax_mapping(mapping) && mapping->nrexceptional)) {
> +	if (mapping_needs_writeback(mapping)) {
>  		err = filemap_fdatawrite(mapping);
>  		/*
>  		 * Even if the above returned error, the pages may be
> @@ -566,8 +571,7 @@ int filemap_write_and_wait_range(struct address_space *mapping,
>  {
>  	int err = 0;
>  
> -	if ((!dax_mapping(mapping) && mapping->nrpages) ||
> -	    (dax_mapping(mapping) && mapping->nrexceptional)) {
> +	if (mapping_needs_writeback(mapping)) {
>  		err = __filemap_fdatawrite_range(mapping, lstart, lend,
>  						 WB_SYNC_ALL);
>  		/* See comment of filemap_write_and_wait() */
> @@ -656,8 +660,7 @@ int file_write_and_wait_range(struct file *file, loff_t lstart, loff_t lend)
>  	int err = 0, err2;
>  	struct address_space *mapping = file->f_mapping;
>  
> -	if ((!dax_mapping(mapping) && mapping->nrpages) ||
> -	    (dax_mapping(mapping) && mapping->nrexceptional)) {
> +	if (mapping_needs_writeback(mapping)) {
>  		err = __filemap_fdatawrite_range(mapping, lstart, lend,
>  						 WB_SYNC_ALL);
>  		/* See comment of filemap_write_and_wait() */
> -- 
> 2.13.3
> 
-- 
Jan Kara <jack@suse.com>
SUSE Labs, CR

[toc] | [prev] | [next] | [standalone]


#1697457 — [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync

FromJeff Layton <jlayton@kernel.org>
Date2017-07-26 20:00 +0200
Subject[PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync
Message-ID<u7yLh-3p3-27@gated-at.bofh.it>
In reply to#1697451
From: Jeff Layton <jlayton@redhat.com>

This means that we need to export the new file_fdatawait_range symbol.

Also, fix a place where a writeback error might get dropped in the
gfs2_is_jdata case.

Signed-off-by: Jeff Layton <jlayton@redhat.com>
---
 fs/gfs2/file.c | 6 ++++--
 1 file changed, 4 insertions(+), 2 deletions(-)

diff --git a/fs/gfs2/file.c b/fs/gfs2/file.c
index c2062a108d19..c53ac6efd04c 100644
--- a/fs/gfs2/file.c
+++ b/fs/gfs2/file.c
@@ -668,12 +668,14 @@ static int gfs2_fsync(struct file *file, loff_t start, loff_t end,
 		if (ret)
 			return ret;
 		if (gfs2_is_jdata(ip))
-			filemap_write_and_wait(mapping);
+			ret = file_write_and_wait(file);
+		if (ret)
+			return ret;
 		gfs2_ail_flush(ip->i_gl, 1);
 	}
 
 	if (mapping->nrpages)
-		ret = filemap_fdatawait_range(mapping, start, end);
+		ret = file_fdatawait_range(file, start, end);
 
 	return ret ? ret : ret1;
 }
-- 
2.13.3

[toc] | [prev] | [next] | [standalone]


#1697507 — Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync

FromMatthew Wilcox <willy@infradead.org>
Date2017-07-26 21:30 +0200
SubjectRe: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync
Message-ID<u7Aam-4qo-11@gated-at.bofh.it>
In reply to#1697457
On Wed, Jul 26, 2017 at 01:55:38PM -0400, Jeff Layton wrote:
> @@ -668,12 +668,14 @@ static int gfs2_fsync(struct file *file, loff_t start, loff_t end,
>  		if (ret)
>  			return ret;
>  		if (gfs2_is_jdata(ip))
> -			filemap_write_and_wait(mapping);
> +			ret = file_write_and_wait(file);
> +		if (ret)
> +			return ret;
>  		gfs2_ail_flush(ip->i_gl, 1);
>  	}

Do we want to skip flushing the AIL if there was an error (possibly
previously encountered)?  I'd think we'd want to flush the AIL then report
the error, like this:

 		if (gfs2_is_jdata(ip))
-			filemap_write_and_wait(mapping);
+			ret = file_write_and_wait(file);
 		gfs2_ail_flush(ip->i_gl, 1);
+		if (ret)
+			return ret;
 	}

[toc] | [prev] | [next] | [standalone]


#1697597 — Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync

FromJeff Layton <jlayton@redhat.com>
Date2017-07-27 00:30 +0200
SubjectRe: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync
Message-ID<u7CYy-6ee-9@gated-at.bofh.it>
In reply to#1697507
On Wed, 2017-07-26 at 12:21 -0700, Matthew Wilcox wrote:
> On Wed, Jul 26, 2017 at 01:55:38PM -0400, Jeff Layton wrote:
> > @@ -668,12 +668,14 @@ static int gfs2_fsync(struct file *file, loff_t start, loff_t end,
> >  		if (ret)
> >  			return ret;
> >  		if (gfs2_is_jdata(ip))
> > -			filemap_write_and_wait(mapping);
> > +			ret = file_write_and_wait(file);
> > +		if (ret)
> > +			return ret;
> >  		gfs2_ail_flush(ip->i_gl, 1);
> >  	}
> 
> Do we want to skip flushing the AIL if there was an error (possibly
> previously encountered)?  I'd think we'd want to flush the AIL then report
> the error, like this:
> 

I wondered about that. Note that earlier in the function, we also bail
out without flushing the AIL if sync_inode_metadata fails, so I assumed
that we'd want to do the same here. 

I could definitely be wrong and am fine with changing it if so.
Discarding the error like we do today seems wrong though.

Bob, thoughts?


>  		if (gfs2_is_jdata(ip))
> -			filemap_write_and_wait(mapping);
> +			ret = file_write_and_wait(file);
>  		gfs2_ail_flush(ip->i_gl, 1);
> +		if (ret)
> +			return ret;
>  	}
-- 
Jeff Layton <jlayton@redhat.com>

[toc] | [prev] | [next] | [standalone]


#1697948 — Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync

FromBob Peterson <rpeterso@redhat.com>
Date2017-07-27 14:50 +0200
SubjectRe: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync
Message-ID<u7QoN-69t-1@gated-at.bofh.it>
In reply to#1697597
----- Original Message -----
| On Wed, 2017-07-26 at 12:21 -0700, Matthew Wilcox wrote:
| > On Wed, Jul 26, 2017 at 01:55:38PM -0400, Jeff Layton wrote:
| > > @@ -668,12 +668,14 @@ static int gfs2_fsync(struct file *file, loff_t
| > > start, loff_t end,
| > >  		if (ret)
| > >  			return ret;
| > >  		if (gfs2_is_jdata(ip))
| > > -			filemap_write_and_wait(mapping);
| > > +			ret = file_write_and_wait(file);
| > > +		if (ret)
| > > +			return ret;
| > >  		gfs2_ail_flush(ip->i_gl, 1);
| > >  	}
| > 
| > Do we want to skip flushing the AIL if there was an error (possibly
| > previously encountered)?  I'd think we'd want to flush the AIL then report
| > the error, like this:
| > 
| 
| I wondered about that. Note that earlier in the function, we also bail
| out without flushing the AIL if sync_inode_metadata fails, so I assumed
| that we'd want to do the same here.
| 
| I could definitely be wrong and am fine with changing it if so.
| Discarding the error like we do today seems wrong though.
| 
| Bob, thoughts?

Hi Jeff, Matthew,

I'm not sure there's a right or wrong answer here. I don't know what's
best from a "correctness" point of view.

I guess I'm leaning toward Jeff's original solution where we don't
call gfs2_ail_flush() on error. The main purpose of ail_flush is to
go through buffer descriptors (bds) attached to the glock and generate
revokes for them in a new transaction. If there's an error condition,
trying to go through more hoops will probably just get us into more
trouble. If the error is -ENOMEM, we don't want to allocate new memory
for the new transaction. If the error is -EIO, we probably don't
want to encourage more writing either.

So on the one hand, it might be good to get rid of the buffer descriptors
so we don't leak memory, but that's probably also done elsewhere.
I have not chased down what happens in that case, but the same thing
would happen in the existing -EIO case a few lines above.

On the other hand, we probably don't want to start a new transaction
and start adding revokes to it, and such, due to the error.

Perhaps Steve Whitehouse can weigh in?

Regards,

Bob Peterson
Red Hat File Systems

[toc] | [prev] | [next] | [standalone]


#1698731 — Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync

FromSteven Whitehouse <swhiteho@redhat.com>
Date2017-07-28 14:40 +0200
SubjectRe: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync
Message-ID<u8cIF-3Kl-5@gated-at.bofh.it>
In reply to#1697948
Hi,


On 27/07/17 13:47, Bob Peterson wrote:
> ----- Original Message -----
> | On Wed, 2017-07-26 at 12:21 -0700, Matthew Wilcox wrote:
> | > On Wed, Jul 26, 2017 at 01:55:38PM -0400, Jeff Layton wrote:
> | > > @@ -668,12 +668,14 @@ static int gfs2_fsync(struct file *file, loff_t
> | > > start, loff_t end,
> | > >  		if (ret)
> | > >  			return ret;
> | > >  		if (gfs2_is_jdata(ip))
> | > > -			filemap_write_and_wait(mapping);
> | > > +			ret = file_write_and_wait(file);
> | > > +		if (ret)
> | > > +			return ret;
> | > >  		gfs2_ail_flush(ip->i_gl, 1);
> | > >  	}
> | >
> | > Do we want to skip flushing the AIL if there was an error (possibly
> | > previously encountered)?  I'd think we'd want to flush the AIL then report
> | > the error, like this:
> | >
> |
> | I wondered about that. Note that earlier in the function, we also bail
> | out without flushing the AIL if sync_inode_metadata fails, so I assumed
> | that we'd want to do the same here.
> |
> | I could definitely be wrong and am fine with changing it if so.
> | Discarding the error like we do today seems wrong though.
> |
> | Bob, thoughts?
>
> Hi Jeff, Matthew,
>
> I'm not sure there's a right or wrong answer here. I don't know what's
> best from a "correctness" point of view.
>
> I guess I'm leaning toward Jeff's original solution where we don't
> call gfs2_ail_flush() on error. The main purpose of ail_flush is to
> go through buffer descriptors (bds) attached to the glock and generate
> revokes for them in a new transaction. If there's an error condition,
> trying to go through more hoops will probably just get us into more
> trouble. If the error is -ENOMEM, we don't want to allocate new memory
> for the new transaction. If the error is -EIO, we probably don't
> want to encourage more writing either.
>
> So on the one hand, it might be good to get rid of the buffer descriptors
> so we don't leak memory, but that's probably also done elsewhere.
> I have not chased down what happens in that case, but the same thing
> would happen in the existing -EIO case a few lines above.
>
> On the other hand, we probably don't want to start a new transaction
> and start adding revokes to it, and such, due to the error.
>
> Perhaps Steve Whitehouse can weigh in?
>
> Regards,
>
> Bob Peterson
> Red Hat File Systems

Yes, we probably do want to skip the ail flush if there is an error. We 
don't know whether the error is permanent or transient at that stage. If 
a previous stage of the fsync has failed, then there may be nothing for 
the next stage to do anyway, so it is probably not a big deal either 
way. So long as the error is reported to the caller, then we should be ok,

Steve.

[toc] | [prev] | [next] | [standalone]


#1698745 — Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync

FromJeff Layton <jlayton@redhat.com>
Date2017-07-28 14:50 +0200
SubjectRe: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync
Message-ID<u8cSl-3NQ-15@gated-at.bofh.it>
In reply to#1698731
On Fri, 2017-07-28 at 13:37 +0100, Steven Whitehouse wrote:
> Hi,
> 
> 
> On 27/07/17 13:47, Bob Peterson wrote:
> > ----- Original Message -----
> > > On Wed, 2017-07-26 at 12:21 -0700, Matthew Wilcox wrote:
> > > > On Wed, Jul 26, 2017 at 01:55:38PM -0400, Jeff Layton wrote:
> > > > > @@ -668,12 +668,14 @@ static int gfs2_fsync(struct file *file, loff_t
> > > > > start, loff_t end,
> > > > >  		if (ret)
> > > > >  			return ret;
> > > > >  		if (gfs2_is_jdata(ip))
> > > > > -			filemap_write_and_wait(mapping);
> > > > > +			ret = file_write_and_wait(file);
> > > > > +		if (ret)
> > > > > +			return ret;
> > > > >  		gfs2_ail_flush(ip->i_gl, 1);
> > > > >  	}
> > > > 
> > > > Do we want to skip flushing the AIL if there was an error (possibly
> > > > previously encountered)?  I'd think we'd want to flush the AIL then report
> > > > the error, like this:
> > > > 
> > > 
> > > I wondered about that. Note that earlier in the function, we also bail
> > > out without flushing the AIL if sync_inode_metadata fails, so I assumed
> > > that we'd want to do the same here.
> > > 
> > > I could definitely be wrong and am fine with changing it if so.
> > > Discarding the error like we do today seems wrong though.
> > > 
> > > Bob, thoughts?
> > 
> > Hi Jeff, Matthew,
> > 
> > I'm not sure there's a right or wrong answer here. I don't know what's
> > best from a "correctness" point of view.
> > 
> > I guess I'm leaning toward Jeff's original solution where we don't
> > call gfs2_ail_flush() on error. The main purpose of ail_flush is to
> > go through buffer descriptors (bds) attached to the glock and generate
> > revokes for them in a new transaction. If there's an error condition,
> > trying to go through more hoops will probably just get us into more
> > trouble. If the error is -ENOMEM, we don't want to allocate new memory
> > for the new transaction. If the error is -EIO, we probably don't
> > want to encourage more writing either.
> > 
> > So on the one hand, it might be good to get rid of the buffer descriptors
> > so we don't leak memory, but that's probably also done elsewhere.
> > I have not chased down what happens in that case, but the same thing
> > would happen in the existing -EIO case a few lines above.
> > 
> > On the other hand, we probably don't want to start a new transaction
> > and start adding revokes to it, and such, due to the error.
> > 
> > Perhaps Steve Whitehouse can weigh in?
> > 
> > Regards,
> > 
> > Bob Peterson
> > Red Hat File Systems
> 
> Yes, we probably do want to skip the ail flush if there is an error. We 
> don't know whether the error is permanent or transient at that stage. If 
> a previous stage of the fsync has failed, then there may be nothing for 
> the next stage to do anyway, so it is probably not a big deal either 
> way. So long as the error is reported to the caller, then we should be ok,
> 

Ok, cool. I'll plan to carry this patch for now as it depends on an
earlier one in the series. One more question though:

Is it correct in the gfs2_is_jdata case to ignore the range that was
passed in from the caller? ->fsync gets start and end arguments, but
this will always write back the whole range. Is that necessary in this
case?

-- 
Jeff Layton <jlayton@redhat.com>

[toc] | [prev] | [next] | [standalone]


#1698751 — Re: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync

FromSteven Whitehouse <swhiteho@redhat.com>
Date2017-07-28 15:00 +0200
SubjectRe: [PATCH v2 4/4] gfs2: convert to errseq_t based writeback error reporting for fsync
Message-ID<u8d22-3Rv-9@gated-at.bofh.it>
In reply to#1698745
Hi,


On 28/07/17 13:47, Jeff Layton wrote:
> On Fri, 2017-07-28 at 13:37 +0100, Steven Whitehouse wrote:
>> Hi,
>>
>>
>> On 27/07/17 13:47, Bob Peterson wrote:
>>> ----- Original Message -----
>>>> On Wed, 2017-07-26 at 12:21 -0700, Matthew Wilcox wrote:
>>>>> On Wed, Jul 26, 2017 at 01:55:38PM -0400, Jeff Layton wrote:
>>>>>> @@ -668,12 +668,14 @@ static int gfs2_fsync(struct file *file, loff_t
>>>>>> start, loff_t end,
>>>>>>   		if (ret)
>>>>>>   			return ret;
>>>>>>   		if (gfs2_is_jdata(ip))
>>>>>> -			filemap_write_and_wait(mapping);
>>>>>> +			ret = file_write_and_wait(file);
>>>>>> +		if (ret)
>>>>>> +			return ret;
>>>>>>   		gfs2_ail_flush(ip->i_gl, 1);
>>>>>>   	}
>>>>> Do we want to skip flushing the AIL if there was an error (possibly
>>>>> previously encountered)?  I'd think we'd want to flush the AIL then report
>>>>> the error, like this:
>>>>>
>>>> I wondered about that. Note that earlier in the function, we also bail
>>>> out without flushing the AIL if sync_inode_metadata fails, so I assumed
>>>> that we'd want to do the same here.
>>>>
>>>> I could definitely be wrong and am fine with changing it if so.
>>>> Discarding the error like we do today seems wrong though.
>>>>
>>>> Bob, thoughts?
>>> Hi Jeff, Matthew,
>>>
>>> I'm not sure there's a right or wrong answer here. I don't know what's
>>> best from a "correctness" point of view.
>>>
>>> I guess I'm leaning toward Jeff's original solution where we don't
>>> call gfs2_ail_flush() on error. The main purpose of ail_flush is to
>>> go through buffer descriptors (bds) attached to the glock and generate
>>> revokes for them in a new transaction. If there's an error condition,
>>> trying to go through more hoops will probably just get us into more
>>> trouble. If the error is -ENOMEM, we don't want to allocate new memory
>>> for the new transaction. If the error is -EIO, we probably don't
>>> want to encourage more writing either.
>>>
>>> So on the one hand, it might be good to get rid of the buffer descriptors
>>> so we don't leak memory, but that's probably also done elsewhere.
>>> I have not chased down what happens in that case, but the same thing
>>> would happen in the existing -EIO case a few lines above.
>>>
>>> On the other hand, we probably don't want to start a new transaction
>>> and start adding revokes to it, and such, due to the error.
>>>
>>> Perhaps Steve Whitehouse can weigh in?
>>>
>>> Regards,
>>>
>>> Bob Peterson
>>> Red Hat File Systems
>> Yes, we probably do want to skip the ail flush if there is an error. We
>> don't know whether the error is permanent or transient at that stage. If
>> a previous stage of the fsync has failed, then there may be nothing for
>> the next stage to do anyway, so it is probably not a big deal either
>> way. So long as the error is reported to the caller, then we should be ok,
>>
> Ok, cool. I'll plan to carry this patch for now as it depends on an
> earlier one in the series. One more question though:
>
> Is it correct in the gfs2_is_jdata case to ignore the range that was
> passed in from the caller? ->fsync gets start and end arguments, but
> this will always write back the whole range. Is that necessary in this
> case?
>
It probably doesn't matter really. We try to discourage the use of jdata 
from userspace. There are a few internal files that use it still, and it 
is there for backwards compatibility more than anything. So performance 
is generally not a problem for that. The ordered write mode is the 
important one.

So you are right that it might be better to add the range into that call 
too, but it is not likely that anybody will notice the performance 
improvement,

Steve.

[toc] | [prev] | [next] | [standalone]


#1697458 — [PATCH v2 3/4] fs: convert sync_file_range to use errseq_t based error-tracking

FromJeff Layton <jlayton@kernel.org>
Date2017-07-26 20:00 +0200
Subject[PATCH v2 3/4] fs: convert sync_file_range to use errseq_t based error-tracking
Message-ID<u7yLh-3p3-23@gated-at.bofh.it>
In reply to#1697451
From: Jeff Layton <jlayton@redhat.com>

sync_file_range doesn't call down into the filesystem directly at all.
It only kicks off writeback of pagecache pages and optionally waits
on the result.

Convert sync_file_range to use errseq_t based error tracking, under the
assumption that most users will prefer this behavior when errors occur.

Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Jeff Layton <jlayton@redhat.com>
---
 fs/sync.c | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)

diff --git a/fs/sync.c b/fs/sync.c
index 2a54c1f22035..27d6b8bbcb6a 100644
--- a/fs/sync.c
+++ b/fs/sync.c
@@ -342,7 +342,7 @@ SYSCALL_DEFINE4(sync_file_range, int, fd, loff_t, offset, loff_t, nbytes,
 
 	ret = 0;
 	if (flags & SYNC_FILE_RANGE_WAIT_BEFORE) {
-		ret = filemap_fdatawait_range(mapping, offset, endbyte);
+		ret = file_fdatawait_range(f.file, offset, endbyte);
 		if (ret < 0)
 			goto out_put;
 	}
@@ -355,7 +355,7 @@ SYSCALL_DEFINE4(sync_file_range, int, fd, loff_t, offset, loff_t, nbytes,
 	}
 
 	if (flags & SYNC_FILE_RANGE_WAIT_AFTER)
-		ret = filemap_fdatawait_range(mapping, offset, endbyte);
+		ret = file_fdatawait_range(f.file, offset, endbyte);
 
 out_put:
 	fdput(f);
-- 
2.13.3

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web