Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1420582 > unrolled thread

Re: [PATCH V2] block: correctly fallback for zeroout

Started byChristoph Hellwig <hch@infradead.org>
First post2016-06-13 10:30 +0200
Last post2016-06-15 23:30 +0200
Articles 6 — 4 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH V2] block: correctly fallback for zeroout Christoph Hellwig <hch@infradead.org> - 2016-06-13 10:30 +0200
    Re: [PATCH V2] block: correctly fallback for zeroout Mike Snitzer <snitzer@redhat.com> - 2016-06-14 20:40 +0200
      Re: [PATCH V2] block: correctly fallback for zeroout "Martin K. Petersen" <martin.petersen@oracle.com> - 2016-06-15 04:40 +0200
        Re: [PATCH V2] block: correctly fallback for zeroout Mike Snitzer <snitzer@redhat.com> - 2016-06-15 04:50 +0200
    Re: [PATCH V2] block: correctly fallback for zeroout "Martin K. Petersen" <martin.petersen@oracle.com> - 2016-06-15 04:20 +0200
      Re: [PATCH V2] block: correctly fallback for zeroout Sitsofe Wheeler <sitsofe@gmail.com> - 2016-06-15 23:30 +0200

#1420582 — Re: [PATCH V2] block: correctly fallback for zeroout

FromChristoph Hellwig <hch@infradead.org>
Date2016-06-13 10:30 +0200
SubjectRe: [PATCH V2] block: correctly fallback for zeroout
Message-ID<rJvpT-3Ie-1@gated-at.bofh.it>
On Fri, Jun 10, 2016 at 09:49:44PM -0400, Martin K. Petersen wrote:
> >> What does the extra io_err buy us? Just have this function return an
> >> error. And then in blkdev_issue_discard if you get -EOPNOTSUPP you
> >> special case it there.
> 
> Shaohua> The __blkdev_issue_discard returns -EOPNOTSUPP if disk doesn't
> Shaohua> support discard.  in that case, blkdev_issue_discard doesn't
> Shaohua> return 0. blkdev_issue_discard only returns 0 if IO error is
> Shaohua> -EOPNOTSUPP.
> 
> Oh, I see. The sanity checks are now in __blkdev_issue_discard() so
> there is no way to distinguish between -EOPNOTSUPP and the other
> -EOPNOTSUPP. *sigh*

We can move the sanity checks out.  Or even better get rid of the
stupid behavior of ignoring the late -EOPNOTSUPP in this low level
helper and instead leaving it to the caller(s) that care.  So far
the DM test suite seems to be the only one that does.

> I am OK with your patch as a stable fix but this really needs to be
> fixed up properly.

And I'd much prefer to get this right now.  It's not like this is
recently introduced behavior.

[toc] | [next] | [standalone]


#1422204

FromMike Snitzer <snitzer@redhat.com>
Date2016-06-14 20:40 +0200
Message-ID<rK1pM-88-23@gated-at.bofh.it>
In reply to#1420582
On Mon, Jun 13 2016 at  4:20am -0400,
Christoph Hellwig <hch@infradead.org> wrote:

> On Fri, Jun 10, 2016 at 09:49:44PM -0400, Martin K. Petersen wrote:
> > >> What does the extra io_err buy us? Just have this function return an
> > >> error. And then in blkdev_issue_discard if you get -EOPNOTSUPP you
> > >> special case it there.
> > 
> > Shaohua> The __blkdev_issue_discard returns -EOPNOTSUPP if disk doesn't
> > Shaohua> support discard.  in that case, blkdev_issue_discard doesn't
> > Shaohua> return 0. blkdev_issue_discard only returns 0 if IO error is
> > Shaohua> -EOPNOTSUPP.
> > 
> > Oh, I see. The sanity checks are now in __blkdev_issue_discard() so
> > there is no way to distinguish between -EOPNOTSUPP and the other
> > -EOPNOTSUPP. *sigh*
> 
> We can move the sanity checks out.  Or even better get rid of the
> stupid behavior of ignoring the late -EOPNOTSUPP in this low level
> helper and instead leaving it to the caller(s) that care.

I'm not onboard with blkdev_issue_discard() no longer masking the late
return of -EOPNOTSUPP.

I'd be fine with moving the early -EOPNOTSUPP checks and the masking of
late -EOPNOTSUPP out to blkdev_issue_discard().  But to be clear,
the masking of late -EOPNOTSUPP return is there for stacking drivers
like MD and DM.  So long as the upper level ioctl code, filesystems, etc
makes use of blkdev_issue_discard() then they'll still get the benefit
of that masking.

drivers/md/dm-thin.c is now using the new async __blkdev_issue_discard()
and it'll only ever do so to a device it knows supports discards -- BUT
it could be that the DM thin-pool's data device is itself a stacked
device that doesn't uniformly support discards throughout its entire
logical address space.  So it could issue a discard to a portion of the
stacked data device that will return -EOPNOTSUPP.. so long story short:
making this change to remove this so-called "stupid behaviour" will
require code like drivers/md/dm-thin.c:issue_discard(() to check the
return from __blkdev_issue_discard() and if it is -EOPNOTSUPP then it
should return 0.

> So far the DM test suite seems to be the only one that does.

The device-mapper-test-suite was only ever relying on
blkdev_issue_discard()'s early return of -EOPNOTSUPP.

> > I am OK with your patch as a stable fix but this really needs to be
> > fixed up properly.
> 
> And I'd much prefer to get this right now.  It's not like this is
> recently introduced behavior.

We need to sequence the fixes such that stable kernels get the zeroout
fallback fixed.  Right?  Not sure if that is a goal of shli's though..

In 4.7-rc, where you introduced __blkdev_issue_discard and I made
dm-thin.c consume it, I'm fine with seeing __blkdev_issue_discard stop
masking -EOPNOTSUPP... but at the same time that change is made
dm-thin.c would need to be fixed (in the same commit as the interface
change).  Though I'm now missing what lifting the -EOPNOTSUPP behavior
into blkdev_issue_discard() buys us... maybe purity of the new async
__blkdev_issue_discard()?

Mike

[toc] | [prev] | [next] | [standalone]


#1422515

From"Martin K. Petersen" <martin.petersen@oracle.com>
Date2016-06-15 04:40 +0200
Message-ID<rK8Ui-4WP-21@gated-at.bofh.it>
In reply to#1422204
>>>>> "Mike" == Mike Snitzer <snitzer@redhat.com> writes:

Mike,

Mike> so long story short: making this change to remove this so-called
Mike> "stupid behaviour" will require code like
Mike> drivers/md/dm-thin.c:issue_discard(() to check the return from
Mike> __blkdev_issue_discard() and if it is -EOPNOTSUPP then it should
Mike> return 0.

Yes, please.

The original -EOPNOTSUPP equals success is a remnant from the days where
discards were only a hint. And sadly that policy got encoded in the
actual interface instead of being left up to the caller.

Now the world has moved on. And reliable zeroout behavior, the SCSI
target drivers and other kernel users need an interface that tells them
exactly what happened at the bottom of the stack so they in turn can
provide a deterministic result (including partial block zeroing) to
their clients.

It's imperative that this gets fixed up. And instead of perpetuating a
weird interface that returns success on failure, let's fix DM and the
callers that actually check the return of blkdev_issue_discard() so they
do the right thing.

I really don't understand why you are objecting so much to this. It's a
trivial change that may not directly benefit DM but it helps everybody
else. And it cleans up a library call that's confusing, error prone and
goes against the very grain of how all our kernel interfaces work in
general.

-- 
Martin K. Petersen	Oracle Linux Engineering

[toc] | [prev] | [next] | [standalone]


#1422518

FromMike Snitzer <snitzer@redhat.com>
Date2016-06-15 04:50 +0200
Message-ID<rK93X-507-1@gated-at.bofh.it>
In reply to#1422515
On Tue, Jun 14 2016 at 10:30pm -0400,
Martin K. Petersen <martin.petersen@oracle.com> wrote:

> >>>>> "Mike" == Mike Snitzer <snitzer@redhat.com> writes:
> 
> Mike,
> 
> Mike> so long story short: making this change to remove this so-called
> Mike> "stupid behaviour" will require code like
> Mike> drivers/md/dm-thin.c:issue_discard(() to check the return from
> Mike> __blkdev_issue_discard() and if it is -EOPNOTSUPP then it should
> Mike> return 0.
> 
> Yes, please.
> 
> The original -EOPNOTSUPP equals success is a remnant from the days where
> discards were only a hint. And sadly that policy got encoded in the
> actual interface instead of being left up to the caller.
> 
> Now the world has moved on. And reliable zeroout behavior, the SCSI
> target drivers and other kernel users need an interface that tells them
> exactly what happened at the bottom of the stack so they in turn can
> provide a deterministic result (including partial block zeroing) to
> their clients.
> 
> It's imperative that this gets fixed up. And instead of perpetuating a
> weird interface that returns success on failure, let's fix DM and the
> callers that actually check the return of blkdev_issue_discard() so they
> do the right thing.
> 
> I really don't understand why you are objecting so much to this. It's a
> trivial change that may not directly benefit DM but it helps everybody
> else. And it cleans up a library call that's confusing, error prone and
> goes against the very grain of how all our kernel interfaces work in
> general.

I've been consistently objecting to changing the blkdev_issue_discard()
interface.  Fixing the async __blkdev_issue_discard() to offer
unfiltered return values is perfectly fine by me.

But the ship has sailed on the blkdev_issue_discard() interface.

[toc] | [prev] | [next] | [standalone]


#1422492

From"Martin K. Petersen" <martin.petersen@oracle.com>
Date2016-06-15 04:20 +0200
Message-ID<rK8AV-4PU-9@gated-at.bofh.it>
In reply to#1420582
>>>>> "Christoph" == Christoph Hellwig <hch@infradead.org> writes:

Christoph> We can move the sanity checks out.  Or even better get rid of
Christoph> the stupid behavior of ignoring the late -EOPNOTSUPP in this
Christoph> low level helper and instead leaving it to the caller(s) that
Christoph> care.

It definitely should be a caller decision whether to ignore the return
value or not.

>> I am OK with your patch as a stable fix but this really needs to be
>> fixed up properly.

Christoph> And I'd much prefer to get this right now.  It's not like
Christoph> this is recently introduced behavior.

Unfortunately there are quite a few callers of blkdev_issue_discard()
these days. Some of them ignore the return value but not all of
them. I'm concerned about causing all sorts of breakage if we suddenly
start returning errors various places in the stable trees.

-- 
Martin K. Petersen	Oracle Linux Engineering

[toc] | [prev] | [next] | [standalone]


#1423462

FromSitsofe Wheeler <sitsofe@gmail.com>
Date2016-06-15 23:30 +0200
Message-ID<rKqxQ-7NH-27@gated-at.bofh.it>
In reply to#1422492
On Tue, Jun 14, 2016 at 10:14:50PM -0400, Martin K. Petersen wrote:
> >>>>> "Christoph" == Christoph Hellwig <hch@infradead.org> writes:
> 
> Christoph> And I'd much prefer to get this right now.  It's not like
> Christoph> this is recently introduced behavior.
> 
> Unfortunately there are quite a few callers of blkdev_issue_discard()
> these days. Some of them ignore the return value but not all of
> them. I'm concerned about causing all sorts of breakage if we suddenly
> start returning errors various places in the stable trees.

This is true. We have problematic behaviour in stable kernels today so
there needs to be a "least intrusive" workaround which changes the
behaviour as little as possible for those. I would say that means
maintaining the current -EOPNOTSUPP behaviour in those kernels
regardless of what goes into master.

-- 
Sitsofe | http://sucs.org/~sits/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web