Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1480323 > unrolled thread
| Started by | Wouter Verhelst <w@uter.be> |
|---|---|
| First post | 2016-09-09 22:20 +0200 |
| Last post | 2016-09-15 14:20 +0200 |
| Articles | 20 on this page of 42 — 6 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Wouter Verhelst <w@uter.be> - 2016-09-09 22:20 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Wouter Verhelst <w@uter.be> - 2016-09-09 23:00 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Jens Axboe <axboe@kernel.dk> - 2016-09-10 01:40 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Wouter Verhelst <w@uter.be> - 2016-09-15 13:00 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 13:20 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Wouter Verhelst <w@uter.be> - 2016-09-15 13:40 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 13:50 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 13:50 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 14:00 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Wouter Verhelst <w@uter.be> - 2016-09-15 14:10 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 14:30 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 14:30 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Wouter Verhelst <w@uter.be> - 2016-09-15 14:30 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 14:10 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 13:40 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Eric Blake <eblake@redhat.com> - 2016-09-15 15:40 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Paolo Bonzini <pbonzini@redhat.com> - 2016-09-15 16:10 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 17:30 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Paolo Bonzini <pbonzini@redhat.com> - 2016-09-15 23:20 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 17:30 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 13:50 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 14:00 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 13:50 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Wouter Verhelst <w@uter.be> - 2016-09-15 14:00 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 14:10 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 14:20 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 14:20 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 14:30 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Wouter Verhelst <w@uter.be> - 2016-09-15 14:30 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 14:30 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 14:40 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 14:40 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 14:50 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 14:50 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Wouter Verhelst <w@uter.be> - 2016-09-15 15:20 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 17:20 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 18:10 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Wouter Verhelst <w@uter.be> - 2016-09-15 18:30 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Alex Bligh <alex@alex.org.uk> - 2016-09-15 18:50 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Eric Blake <eblake@redhat.com> - 2016-09-15 21:10 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 14:40 +0200
Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements Christoph Hellwig <hch@infradead.org> - 2016-09-15 14:20 +0200
Page 1 of 3 [1] 2 3 Next page →
| From | Wouter Verhelst <w@uter.be> |
|---|---|
| Date | 2016-09-09 22:20 +0200 |
| Subject | Re: [Nbd] [RESEND][PATCH 0/5] nbd improvements |
| Message-ID | <sfArf-7zf-11@gated-at.bofh.it> |
Hi Josef,
On Thu, Sep 08, 2016 at 05:12:05PM -0400, Josef Bacik wrote:
> Apologies if you are getting this a second time, it appears vger ate my last
> submission.
>
> ----------------------------------------------------------------------
>
> This is a patch series aimed at bringing NBD into 2016. The two big components
> of this series is converting nbd over to using blkmq and then allowing us to
> provide more than one connection for a nbd device. The NBD user space server
> doesn't care about how many connections it has to a particular device, so we can
> easily open multiple connections to the server and allow blkmq to handle
> multi-plexing over the different connections.
I see some practical problems with this:
- You removed the pid attribute from sysfs (unless you added it back and
I didn't notice, in which case just ignore this part). This kills
userspace in two ways:
- systemd/udev mark an NBD device as "not active" if the sysfs pid
attribute is absent. Removing that attribute causes the new nbd
systemd unit to stop working.
- nbd-client -check relies on this attribute too, which means that
even if people don't use systemd, their init scripts will still
break, and vigilant sysadmins (who check before trying to connect
something) will be surprised.
- What happens if userspace tries to connect an already-connected device
to some other server? Currently that can't happen (you get EBUSY);
with this patch, I believe it can, and data corruption would be the
result (on *two* nbd devices). Additionally, with the loss of the pid
attribute (as above) and the ensuing loss of the -check functionality,
this might actually be a somewhat likely scenario.
- What happens if one of the multiple connections drop but the others do
not?
- This all has the downside that userspace now has to predict how many
parallel connections will be necessary and/or useful. If the initial
guess was wrong, we don't have a way to correct later on.
My suggestion is to reject an additional connection unless it comes from
the same userspace process as the previous connections, and to retain
the pid attribute (since it is now guaranteed to be the same for all the
connections). That should fix the first two issues (while unfortunately
reinforcing the last one). The third would also need to have clearly
defined semantics, at the very least.
A better way, long term, would presumably be to modify the protocol to
allow multiplexing several requests in one NBD session. This would deal
with what you're trying to fix too[1], while it would not pull in all of
the above problems.
[1] after all, we have to serialize all traffic anyway, just before it
heads into the NIC.
--
< ron> I mean, the main *practical* problem with C++, is there's like a dozen
people in the world who think they really understand all of its rules,
and pretty much all of them are just lying to themselves too.
-- #debian-devel, OFTC, 2016-02-12
[toc] | [next] | [standalone]
| From | Wouter Verhelst <w@uter.be> |
|---|---|
| Date | 2016-09-09 23:00 +0200 |
| Message-ID | <sfB3Y-7LQ-13@gated-at.bofh.it> |
| In reply to | #1480323 |
On Fri, Sep 09, 2016 at 04:36:07PM -0400, Josef Bacik wrote:
> On 09/09/2016 04:02 PM, Wouter Verhelst wrote:
[...]
> > I see some practical problems with this:
> > - You removed the pid attribute from sysfs (unless you added it back and
> > I didn't notice, in which case just ignore this part). This kills
> > userspace in two ways:
> > - systemd/udev mark an NBD device as "not active" if the sysfs pid
> > attribute is absent. Removing that attribute causes the new nbd
> > systemd unit to stop working.
> > - nbd-client -check relies on this attribute too, which means that
> > even if people don't use systemd, their init scripts will still
> > break, and vigilant sysadmins (who check before trying to connect
> > something) will be surprised.
>
> Ok I can add this back, I didn't see anybody using it, but again I didn't look
> very hard.
Thank you.
> > - What happens if userspace tries to connect an already-connected device
> > to some other server? Currently that can't happen (you get EBUSY);
> > with this patch, I believe it can, and data corruption would be the
> > result (on *two* nbd devices). Additionally, with the loss of the pid
> > attribute (as above) and the ensuing loss of the -check functionality,
> > this might actually be a somewhat likely scenario.
>
> Once you do DO_IT then you'll get the EBUSY, so no problems.
Oh, okay. I missed that part.
> Now if you modify the client to connect to two different servers then yes you
> could have data corruption, but hey if you do stupid things then bad things
> happen, I'm not sure we need to explicitly keep this from happening.
Yeah, totally agree there.
> > - What happens if one of the multiple connections drop but the others do
> > not?
>
> It keeps on trucking, but the connections that break will return -EIO. That's
> not good, I'll fix it to tear down everything if that happens.
Right. Alternatively, you could perhaps make it so that the lost
connection is removed, unack'd requests on that connection are resent,
and the session moves on with one less connection (unless the lost
connection is the last one, in which case we die as before). That might
be too much work and not worth it though.
> > - This all has the downside that userspace now has to predict how many
> > parallel connections will be necessary and/or useful. If the initial
> > guess was wrong, we don't have a way to correct later on.
>
> No, it relies on the admin to specify based on their environment.
Sure, but I suppose it would be nice if things could dynamically grow
when needed, and/or that the admin could modify the number of
connections of an already-connected device. Then again, this might also
be too much work and not worth it.
[...]
> > A better way, long term, would presumably be to modify the protocol to
> > allow multiplexing several requests in one NBD session. This would deal
> > with what you're trying to fix too[1], while it would not pull in all of
> > the above problems.
> >
> > [1] after all, we have to serialize all traffic anyway, just before it
> > heads into the NIC.
>
> Yeah I considered changing the protocol to handle multiplexing different
> requests, but that runs into trouble since we can't guarantee that each discrete
> sendmsg/recvmsg is going to atomically copy our buffer in. We can accomplish
> this with KCM of course which is a road I went down for a little while, but then
> we have the issue of the actual data to send across, and KCM is limited to a
> certain buffer size (I don't remember what it was exactly). This limitation is
> fine in practice I think, but I got such good performance with multiple
> connections that I threw all that work away and went with this.
Okay, sounds like you've given that way more thought than me, and that
that's a dead end. Never mind then.
> Thanks for the review, I'll fix up these issues you've pointed out and resend,
Thanks,
--
< ron> I mean, the main *practical* problem with C++, is there's like a dozen
people in the world who think they really understand all of its rules,
and pretty much all of them are just lying to themselves too.
-- #debian-devel, OFTC, 2016-02-12
[toc] | [prev] | [next] | [standalone]
| From | Jens Axboe <axboe@kernel.dk> |
|---|---|
| Date | 2016-09-10 01:40 +0200 |
| Message-ID | <sfDyN-V8-3@gated-at.bofh.it> |
| In reply to | #1480352 |
On 09/09/2016 05:00 PM, Josef Bacik wrote: >> Right. Alternatively, you could perhaps make it so that the lost >> connection is removed, unack'd requests on that connection are resent, >> and the session moves on with one less connection (unless the lost >> connection is the last one, in which case we die as before). That might >> be too much work and not worth it though. > > Yeah I wasn't sure if we could just randomly remove hw queue's in blk mq > while the device is still up. If that is in fact easy to do then I'm in > favor of trucking along with less connections than we originally had, > otherwise I think it'll be too big of a pain. Also some users (Facebook > in this case) would rather the whole thing fail so we can figure out > what went wrong rather than suddenly going at a degraded performance. We can do that online. We do that for CPU hotplug/unplug events, and we also expose that functionality to drivers through blk_mq_update_nr_hw_queues(). So yes, should be trivial to support from nbd. -- Jens Axboe
[toc] | [prev] | [next] | [standalone]
| From | Wouter Verhelst <w@uter.be> |
|---|---|
| Date | 2016-09-15 13:00 +0200 |
| Message-ID | <shCyC-5VC-15@gated-at.bofh.it> |
| In reply to | #1480323 |
Hi,
On Fri, Sep 09, 2016 at 10:02:03PM +0200, Wouter Verhelst wrote:
> I see some practical problems with this:
[...]
One more that I didn't think about earlier:
A while back, we spent quite some time defining the semantics of the
various commands in the face of the NBD_CMD_FLUSH and NBD_CMD_FLAG_FUA
write barriers. At the time, we decided that it would be unreasonable
to expect servers to make these write barriers effective across
different connections.
Since my knowledge of kernel internals is limited, I tried finding some
documentation on this, but I guess that either it doesn't exist or I'm
looking in the wrong place; therefore, am I correct in assuming that
blk-mq knows about such semantics, and will handle them correctly (by
either sending a write barrier to all queues, or not making assumptions
about write barriers that were sent over a different queue)? If not,
this may be something that needs to be taken care of.
Thanks,
--
< ron> I mean, the main *practical* problem with C++, is there's like a dozen
people in the world who think they really understand all of its rules,
and pretty much all of them are just lying to themselves too.
-- #debian-devel, OFTC, 2016-02-12
[toc] | [prev] | [next] | [standalone]
| From | Alex Bligh <alex@alex.org.uk> |
|---|---|
| Date | 2016-09-15 13:20 +0200 |
| Message-ID | <shCRX-6hp-5@gated-at.bofh.it> |
| In reply to | #1483987 |
Wouter, Josef, (& Eric) > On 15 Sep 2016, at 11:49, Wouter Verhelst <w@uter.be> wrote: > > Hi, > > On Fri, Sep 09, 2016 at 10:02:03PM +0200, Wouter Verhelst wrote: >> I see some practical problems with this: > [...] > > One more that I didn't think about earlier: > > A while back, we spent quite some time defining the semantics of the > various commands in the face of the NBD_CMD_FLUSH and NBD_CMD_FLAG_FUA > write barriers. At the time, we decided that it would be unreasonable > to expect servers to make these write barriers effective across > different connections. Actually I wonder if there is a wider problem in that implementations might mediate access to a device by presence of an extant TCP connection, i.e. only permit one TCP connection to access a given block device at once. If you think about (for instance) a forking daemon that does writeback caching, that would be an entirely reasonable thing to do for data consistency. I also wonder whether any servers that can do caching per connection will always share a consistent cache between connections. The one I'm worried about in particular here is qemu-nbd - Eric Blake CC'd. A more general point is that with multiple queues requests may be processed in a different order even by those servers that currently process the requests in strict order, or in something similar to strict order. The server is permitted by the spec (save as mandated by NBD_CMD_FLUSH and NBD_CMD_FLAG_FUA) to process commands out of order anyway, but I suspect this has to date been little tested. Lastly I confess to lack of familiarity with the kernel side code, but how is NBD_CMD_DISCONNECT synchronised across each of the connections? Presumably you need to send it on each channel, but cannot assume the NBD connection as a whole is dead until the last tcp connection has closed? -- Alex Bligh
[toc] | [prev] | [next] | [standalone]
| From | Wouter Verhelst <w@uter.be> |
|---|---|
| Date | 2016-09-15 13:40 +0200 |
| Message-ID | <shDbk-6oz-21@gated-at.bofh.it> |
| In reply to | #1483996 |
On Thu, Sep 15, 2016 at 12:09:28PM +0100, Alex Bligh wrote:
> Wouter, Josef, (& Eric)
>
> > On 15 Sep 2016, at 11:49, Wouter Verhelst <w@uter.be> wrote:
> >
> > Hi,
> >
> > On Fri, Sep 09, 2016 at 10:02:03PM +0200, Wouter Verhelst wrote:
> >> I see some practical problems with this:
> > [...]
> >
> > One more that I didn't think about earlier:
> >
> > A while back, we spent quite some time defining the semantics of the
> > various commands in the face of the NBD_CMD_FLUSH and NBD_CMD_FLAG_FUA
> > write barriers. At the time, we decided that it would be unreasonable
> > to expect servers to make these write barriers effective across
> > different connections.
>
> Actually I wonder if there is a wider problem in that implementations
> might mediate access to a device by presence of an extant TCP connection,
> i.e. only permit one TCP connection to access a given block device at
> once. If you think about (for instance) a forking daemon that does
> writeback caching, that would be an entirely reasonable thing to do
> for data consistency.
Sure. They will have to live with the fact that clients connected to
them will run slower; I don't think that's a problem. In addition,
Josef's client implementation requires the user to explicitly ask for
multiple connections.
There are multiple contexts in which NBD can be used, and in some
performance is more important than in others. I think that is fine.
[...]
> A more general point is that with multiple queues requests
> may be processed in a different order even by those servers that
> currently process the requests in strict order, or in something
> similar to strict order. The server is permitted by the spec
> (save as mandated by NBD_CMD_FLUSH and NBD_CMD_FLAG_FUA) to
> process commands out of order anyway, but I suspect this has
> to date been little tested.
Yes, and that is why I was asking about this. If the write barriers
are expected to be shared across connections, we have a problem. If,
however, they are not, then it doesn't matter that the commands may be
processed out of order.
[...]
--
< ron> I mean, the main *practical* problem with C++, is there's like a dozen
people in the world who think they really understand all of its rules,
and pretty much all of them are just lying to themselves too.
-- #debian-devel, OFTC, 2016-02-12
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-15 13:50 +0200 |
| Message-ID | <shDkZ-6rX-3@gated-at.bofh.it> |
| In reply to | #1484015 |
On Thu, Sep 15, 2016 at 01:29:36PM +0200, Wouter Verhelst wrote: > Yes, and that is why I was asking about this. If the write barriers > are expected to be shared across connections, we have a problem. If, > however, they are not, then it doesn't matter that the commands may be > processed out of order. There is no such thing as a write barrier in the Linux kernel. We'd much prefer protocols not to introduce any pointless synchronization if we can avoid it.
[toc] | [prev] | [next] | [standalone]
| From | Alex Bligh <alex@alex.org.uk> |
|---|---|
| Date | 2016-09-15 13:50 +0200 |
| Message-ID | <shDl0-6rX-7@gated-at.bofh.it> |
| In reply to | #1484019 |
> On 15 Sep 2016, at 12:40, Christoph Hellwig <hch@infradead.org> wrote: > > On Thu, Sep 15, 2016 at 01:29:36PM +0200, Wouter Verhelst wrote: >> Yes, and that is why I was asking about this. If the write barriers >> are expected to be shared across connections, we have a problem. If, >> however, they are not, then it doesn't matter that the commands may be >> processed out of order. > > There is no such thing as a write barrier in the Linux kernel. We'd > much prefer protocols not to introduce any pointless synchronization > if we can avoid it. I suspect the issue is terminological. Essentially NBD does supports FLUSH/FUA like this: https://www.kernel.org/doc/Documentation/block/writeback_cache_control.txt IE supports the same FLUSH/FUA primitives as other block drivers (AIUI). Link to protocol (per last email) here: https://github.com/yoe/nbd/blob/master/doc/proto.md#ordering-of-messages-and-writes -- Alex Bligh
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-15 14:00 +0200 |
| Message-ID | <shDuF-6vN-25@gated-at.bofh.it> |
| In reply to | #1484020 |
On Thu, Sep 15, 2016 at 12:46:07PM +0100, Alex Bligh wrote: > Essentially NBD does supports FLUSH/FUA like this: > > https://www.kernel.org/doc/Documentation/block/writeback_cache_control.txt > > IE supports the same FLUSH/FUA primitives as other block drivers (AIUI). > > Link to protocol (per last email) here: > > https://github.com/yoe/nbd/blob/master/doc/proto.md#ordering-of-messages-and-writes Flush as defined by the Linux block layer (and supported that way in SCSI, ATA, NVMe) only requires to flush all already completed writes to non-volatile media. It does not impose any ordering unlike the nbd spec. FUA as defined by the Linux block layer (and supported that way in SCSI, ATA, NVMe) only requires the write operation the FUA bit is set on to be on non-volatile media before completing the write operation. It does not impose any ordering, which seems to match the nbd spec. Unlike the NBD spec Linux does not allow FUA to be set on anything by WRITE commands. Some other storage protocols allow a FUA bit on READ commands or other commands that write data to the device, though.
[toc] | [prev] | [next] | [standalone]
| From | Wouter Verhelst <w@uter.be> |
|---|---|
| Date | 2016-09-15 14:10 +0200 |
| Message-ID | <shDEm-6P5-27@gated-at.bofh.it> |
| In reply to | #1484034 |
On Thu, Sep 15, 2016 at 04:52:17AM -0700, Christoph Hellwig wrote:
> On Thu, Sep 15, 2016 at 12:46:07PM +0100, Alex Bligh wrote:
> > Essentially NBD does supports FLUSH/FUA like this:
> >
> > https://www.kernel.org/doc/Documentation/block/writeback_cache_control.txt
> >
> > IE supports the same FLUSH/FUA primitives as other block drivers (AIUI).
> >
> > Link to protocol (per last email) here:
> >
> > https://github.com/yoe/nbd/blob/master/doc/proto.md#ordering-of-messages-and-writes
>
> Flush as defined by the Linux block layer (and supported that way in
> SCSI, ATA, NVMe) only requires to flush all already completed writes
> to non-volatile media.
That is precisely what FLUSH in nbd does, too.
> It does not impose any ordering unlike the nbd spec.
If you read the spec differently, then that's a bug in the spec. Can you
clarify which part of it caused that confusion? We should fix it, then.
> FUA as defined by the Linux block layer (and supported that way in SCSI,
> ATA, NVMe) only requires the write operation the FUA bit is set on to be
> on non-volatile media before completing the write operation. It does
> not impose any ordering, which seems to match the nbd spec. Unlike the
> NBD spec Linux does not allow FUA to be set on anything by WRITE
> commands. Some other storage protocols allow a FUA bit on READ
> commands or other commands that write data to the device, though.
Yes. There was some discussion on that part, and we decided that setting
the flag doesn't hurt, but the spec also clarifies that using it on READ
does nothing, semantically.
The problem is that there are clients in the wild which do set it on
READ, so it's just a matter of "be liberal in what you accept".
--
< ron> I mean, the main *practical* problem with C++, is there's like a dozen
people in the world who think they really understand all of its rules,
and pretty much all of them are just lying to themselves too.
-- #debian-devel, OFTC, 2016-02-12
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-15 14:30 +0200 |
| Message-ID | <shDXH-6Zf-21@gated-at.bofh.it> |
| In reply to | #1484057 |
On Thu, Sep 15, 2016 at 02:26:31PM +0200, Wouter Verhelst wrote: > Yes. I think the kernel nbd driver should probably filter out FUA on > READ. It has no meaning in the case of nbd, and whatever expectations > the kernel may have cannot be provided for by nbd anyway. The kernel never sets FUA on reads - I just pointed out that other protocols have defined (although horrible) semantics for it.
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-15 14:30 +0200 |
| Message-ID | <shDXH-6Zf-23@gated-at.bofh.it> |
| In reply to | #1484057 |
On Thu, Sep 15, 2016 at 02:01:59PM +0200, Wouter Verhelst wrote: > Yes. There was some discussion on that part, and we decided that setting > the flag doesn't hurt, but the spec also clarifies that using it on READ > does nothing, semantically. > > > The problem is that there are clients in the wild which do set it on > READ, so it's just a matter of "be liberal in what you accept". Note that FUA on READ in SCSI and NVMe does have a meaning - it requires you to bypass any sort of cache on the target. I think it's an wrong defintion because it mandates implementation details that aren't observable by the initiator, but it's still the spec wording and nbd diverges from it. That's not nessecarily a bad thing, but a caveat to look out for.
[toc] | [prev] | [next] | [standalone]
| From | Wouter Verhelst <w@uter.be> |
|---|---|
| Date | 2016-09-15 14:30 +0200 |
| Message-ID | <shDXH-6Zf-25@gated-at.bofh.it> |
| In reply to | #1484102 |
On Thu, Sep 15, 2016 at 05:20:08AM -0700, Christoph Hellwig wrote:
> On Thu, Sep 15, 2016 at 02:01:59PM +0200, Wouter Verhelst wrote:
> > Yes. There was some discussion on that part, and we decided that setting
> > the flag doesn't hurt, but the spec also clarifies that using it on READ
> > does nothing, semantically.
> >
> >
> > The problem is that there are clients in the wild which do set it on
> > READ, so it's just a matter of "be liberal in what you accept".
>
> Note that FUA on READ in SCSI and NVMe does have a meaning - it
> requires you to bypass any sort of cache on the target. I think it's an
> wrong defintion because it mandates implementation details that aren't
> observable by the initiator, but it's still the spec wording and nbd
> diverges from it. That's not nessecarily a bad thing, but a caveat to
> look out for.
Yes. I think the kernel nbd driver should probably filter out FUA on
READ. It has no meaning in the case of nbd, and whatever expectations
the kernel may have cannot be provided for by nbd anyway.
--
< ron> I mean, the main *practical* problem with C++, is there's like a dozen
people in the world who think they really understand all of its rules,
and pretty much all of them are just lying to themselves too.
-- #debian-devel, OFTC, 2016-02-12
[toc] | [prev] | [next] | [standalone]
| From | Alex Bligh <alex@alex.org.uk> |
|---|---|
| Date | 2016-09-15 14:10 +0200 |
| Message-ID | <shDEm-6P5-39@gated-at.bofh.it> |
| In reply to | #1484034 |
> On 15 Sep 2016, at 12:52, Christoph Hellwig <hch@infradead.org> wrote: > > On Thu, Sep 15, 2016 at 12:46:07PM +0100, Alex Bligh wrote: >> Essentially NBD does supports FLUSH/FUA like this: >> >> https://www.kernel.org/doc/Documentation/block/writeback_cache_control.txt >> >> IE supports the same FLUSH/FUA primitives as other block drivers (AIUI). >> >> Link to protocol (per last email) here: >> >> https://github.com/yoe/nbd/blob/master/doc/proto.md#ordering-of-messages-and-writes > > Flush as defined by the Linux block layer (and supported that way in > SCSI, ATA, NVMe) only requires to flush all already completed writes > to non-volatile media. It does not impose any ordering unlike the > nbd spec. As maintainer of the NBD spec, I'm confused as to why you think it imposes any ordering - if you think this, clearly I need to clean up the wording. Here's what it says: > The server MAY process commands out of order, and MAY reply out of order, > except that: > > • All write commands (that includes NBD_CMD_WRITE, and NBD_CMD_TRIM) > that the server completes (i.e. replies to) prior to processing to a > NBD_CMD_FLUSH MUST be written to non-volatile storage prior to replying to that > NBD_CMD_FLUSH. This paragraph only applies if NBD_FLAG_SEND_FLUSH is set within > the transmission flags, as otherwise NBD_CMD_FLUSH will never be sent by the > client to the server. (and another bit re FUA that isn't relevant here). Here's the Linux Kernel documentation: > The REQ_PREFLUSH flag can be OR ed into the r/w flags of a bio submitted from > the filesystem and will make sure the volatile cache of the storage device > has been flushed before the actual I/O operation is started. This explicitly > guarantees that previously completed write requests are on non-volatile > storage before the flagged bio starts. In addition the REQ_PREFLUSH flag can be > set on an otherwise empty bio structure, which causes only an explicit cache > flush without any dependent I/O. It is recommend to use > the blkdev_issue_flush() helper for a pure cache flush. I believe that NBD treats NBD_CMD_FLUSH the same as a REQ_PREFLUSH and empty bio. If you don't read those two as compatible, I'd like to understand why not (i.e. what additional constraints one is applying that the other is not) as they are meant to be the same (save that NBD only has FLUSH as a command, i.e. the 'empty bio' version). I am happy to improve the docs to make it clearer. (sidenote: I am interested in the change from REQ_FLUSH to REQ_PREFLUSH, but in an empty bio it's not really relevant I think). > FUA as defined by the Linux block layer (and supported that way in SCSI, > ATA, NVMe) only requires the write operation the FUA bit is set on to be > on non-volatile media before completing the write operation. It does > not impose any ordering, which seems to match the nbd spec. Unlike the > NBD spec Linux does not allow FUA to be set on anything by WRITE > commands. Some other storage protocols allow a FUA bit on READ > commands or other commands that write data to the device, though. I think you mean "anything *but* WRITE commands". In NBD setting FUA on a command that does not write will do nothing, but FUA can be set on NBD_CMD_TRIM and has the expected effect. Interestingly the kernel docs are silent on which commands REQ_FUA can be set on. -- Alex Bligh
[toc] | [prev] | [next] | [standalone]
| From | Christoph Hellwig <hch@infradead.org> |
|---|---|
| Date | 2016-09-15 13:40 +0200 |
| Message-ID | <shDbk-6oz-29@gated-at.bofh.it> |
| In reply to | #1483996 |
On Thu, Sep 15, 2016 at 12:09:28PM +0100, Alex Bligh wrote: > A more general point is that with multiple queues requests > may be processed in a different order even by those servers that > currently process the requests in strict order, or in something > similar to strict order. The server is permitted by the spec > (save as mandated by NBD_CMD_FLUSH and NBD_CMD_FLAG_FUA) to > process commands out of order anyway, but I suspect this has > to date been little tested. The Linux kernel does not assume any synchroniztion between block I/O commands. So any sort of synchronization a protocol does is complete overkill for us.
[toc] | [prev] | [next] | [standalone]
| From | Eric Blake <eblake@redhat.com> |
|---|---|
| Date | 2016-09-15 15:40 +0200 |
| Message-ID | <shF3r-7Ca-23@gated-at.bofh.it> |
| In reply to | #1483996 |
[Multipart message — attachments visible in raw view] — view raw
On 09/15/2016 06:09 AM, Alex Bligh wrote: > > I also wonder whether any servers that can do caching per > connection will always share a consistent cache between > connections. The one I'm worried about in particular here > is qemu-nbd - Eric Blake CC'd. > I doubt that qemu-nbd would ever want to support the situation with more than one client connection writing to the same image at the same time; the implications of sorting out data consistency between multiple writers is rather complex and not worth coding into qemu. So I think qemu would probably prefer to just prohibit the multiple writer situation. And while multiple readers with no writer should be fine, I'm not even sure if multiple readers plus one writer can always be made to appear sane (if there is no coordination between the different connections, on an image where the writer changes AA to BA then flushes then changes to BB, it is still feasible that a reader could see AB (pre-flush state of the first sector, post-flush changes to the second sector, even though the writer never flushed that particular content to disk). But Paolo Bonzini (cc'd) may have more insight on qemu's NBD server and what it supports (or forbids) in the way of multiple clients to a single server. > A more general point is that with multiple queues requests > may be processed in a different order even by those servers that > currently process the requests in strict order, or in something > similar to strict order. The server is permitted by the spec > (save as mandated by NBD_CMD_FLUSH and NBD_CMD_FLAG_FUA) to > process commands out of order anyway, but I suspect this has > to date been little tested. qemu-nbd is definitely capable of serving reads and writes out-of-order to a single connection client; but that's different than the case with multiple connections. -- Eric Blake eblake redhat com +1-919-301-3266 Libvirt virtualization library http://libvirt.org
[toc] | [prev] | [next] | [standalone]
| From | Paolo Bonzini <pbonzini@redhat.com> |
|---|---|
| Date | 2016-09-15 16:10 +0200 |
| Message-ID | <shFwt-822-1@gated-at.bofh.it> |
| In reply to | #1484170 |
[Multipart message — attachments visible in raw view] — view raw
On 15/09/2016 15:34, Eric Blake wrote: > On 09/15/2016 06:09 AM, Alex Bligh wrote: >> >> I also wonder whether any servers that can do caching per >> connection will always share a consistent cache between >> connections. The one I'm worried about in particular here >> is qemu-nbd - Eric Blake CC'd. >> > > I doubt that qemu-nbd would ever want to support the situation with more > than one client connection writing to the same image at the same time; > the implications of sorting out data consistency between multiple > writers is rather complex and not worth coding into qemu. So I think > qemu would probably prefer to just prohibit the multiple writer > situation. And while multiple readers with no writer should be fine, > I'm not even sure if multiple readers plus one writer can always be made > to appear sane (if there is no coordination between the different > connections, on an image where the writer changes AA to BA then flushes > then changes to BB, it is still feasible that a reader could see AB > (pre-flush state of the first sector, post-flush changes to the second > sector, even though the writer never flushed that particular content to > disk). > > But Paolo Bonzini (cc'd) may have more insight on qemu's NBD server and > what it supports (or forbids) in the way of multiple clients to a single > server. I don't think QEMU forbids multiple clients to the single server, and guarantees consistency as long as there is no overlap between writes and reads. These are the same guarantees you have for multiple commands on a single connection. In other words, from the POV of QEMU there's no difference whether multiple commands come from one or more connections. Paolo >> A more general point is that with multiple queues requests >> may be processed in a different order even by those servers that >> currently process the requests in strict order, or in something >> similar to strict order. The server is permitted by the spec >> (save as mandated by NBD_CMD_FLUSH and NBD_CMD_FLAG_FUA) to >> process commands out of order anyway, but I suspect this has >> to date been little tested. > > qemu-nbd is definitely capable of serving reads and writes out-of-order > to a single connection client; but that's different than the case with > multiple connections. >
[toc] | [prev] | [next] | [standalone]
| From | Alex Bligh <alex@alex.org.uk> |
|---|---|
| Date | 2016-09-15 17:30 +0200 |
| Message-ID | <shGLU-hb-33@gated-at.bofh.it> |
| In reply to | #1484206 |
[Multipart message — attachments visible in raw view] — view raw
Paolo,
> On 15 Sep 2016, at 15:07, Paolo Bonzini <pbonzini@redhat.com> wrote:
>
> I don't think QEMU forbids multiple clients to the single server, and
> guarantees consistency as long as there is no overlap between writes and
> reads. These are the same guarantees you have for multiple commands on
> a single connection.
>
> In other words, from the POV of QEMU there's no difference whether
> multiple commands come from one or more connections.
This isn't really about ordering, it's about cache coherency
and persisting things to disk.
What you say is correct as far as it goes in terms of ordering. However
consider the scenario with read and writes on two channels as follows
of the same block:
Channel1 Channel2
R Block read, and cached in user space in
channel 1's cache
Reply sent
W New value written, channel 2's cache updated
channel 1's cache not
R Value returned from channel 1's cache.
In the above scenario, there is a problem if the server(s) handling the
two channels each use a read cache which is not coherent between the
two channels. An example would be a read-through cache on a server that
did fork() and shared no state between connections.
Similarly, if there is a write on channel 1 that has completed, and
the flush goes to channel 2, it may not (if state is not shared) guarantee
that the write on channel 1 (which has completed) is persisted to non-volatile
media. Obviously if the 'state' is OS block cache/buffers/whatever, it
will, but if it's (e.g.) a user-space per process write-through cache,
it won't.
I don't know whether qemu-nbd is likely to suffer from either of these.
--
Alex Bligh
[toc] | [prev] | [next] | [standalone]
| From | Paolo Bonzini <pbonzini@redhat.com> |
|---|---|
| Date | 2016-09-15 23:20 +0200 |
| Message-ID | <shMeC-3PX-9@gated-at.bofh.it> |
| In reply to | #1484322 |
[Multipart message — attachments visible in raw view] — view raw
On 15/09/2016 17:23, Alex Bligh wrote: > Paolo, > >> On 15 Sep 2016, at 15:07, Paolo Bonzini <pbonzini@redhat.com> wrote: >> >> I don't think QEMU forbids multiple clients to the single server, and >> guarantees consistency as long as there is no overlap between writes and >> reads. These are the same guarantees you have for multiple commands on >> a single connection. >> >> In other words, from the POV of QEMU there's no difference whether >> multiple commands come from one or more connections. > > This isn't really about ordering, it's about cache coherency > and persisting things to disk. > > What you say is correct as far as it goes in terms of ordering. However > consider the scenario with read and writes on two channels as follows > of the same block: > > Channel1 Channel2 > > R Block read, and cached in user space in > channel 1's cache > Reply sent > > W New value written, channel 2's cache updated > channel 1's cache not > > R Value returned from channel 1's cache. > > > In the above scenario, there is a problem if the server(s) handling the > two channels each use a read cache which is not coherent between the > two channels. An example would be a read-through cache on a server that > did fork() and shared no state between connections. qemu-nbd does not fork(), so there is no coherency issue if W has replied. However, if W hasn't replied, channel1 can get garbage. Typically the VM will be the one during writes, everyone else must be ready to handle whatever mess the VM throws at them. Paolo > Similarly, if there is a write on channel 1 that has completed, and > the flush goes to channel 2, it may not (if state is not shared) guarantee > that the write on channel 1 (which has completed) is persisted to non-volatile > media. Obviously if the 'state' is OS block cache/buffers/whatever, it > will, but if it's (e.g.) a user-space per process write-through cache, > it won't. > > I don't know whether qemu-nbd is likely to suffer from either of these. It can't happen. On the other hand, channel1 must be ready to handle garbage, it's illegal.
[toc] | [prev] | [next] | [standalone]
| From | Alex Bligh <alex@alex.org.uk> |
|---|---|
| Date | 2016-09-15 17:30 +0200 |
| Message-ID | <shGLU-hb-17@gated-at.bofh.it> |
| In reply to | #1484170 |
[Multipart message — attachments visible in raw view] — view raw
Eric, > I doubt that qemu-nbd would ever want to support the situation with more > than one client connection writing to the same image at the same time; > the implications of sorting out data consistency between multiple > writers is rather complex and not worth coding into qemu. So I think > qemu would probably prefer to just prohibit the multiple writer > situation. Yeah, I was thinking about a 'no multiple connection' flag. > And while multiple readers with no writer should be fine, > I'm not even sure if multiple readers plus one writer can always be made > to appear sane (if there is no coordination between the different > connections, on an image where the writer changes AA to BA then flushes > then changes to BB, it is still feasible that a reader could see AB > (pre-flush state of the first sector, post-flush changes to the second > sector, even though the writer never flushed that particular content to > disk). Agree -- Alex Bligh
[toc] | [prev] | [next] | [standalone]
Page 1 of 3 [1] 2 3 Next page →
Back to top | Article view | linux.kernel
csiph-web