Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1457828 > unrolled thread

[RFD] I/O scheduling in blk-mq

Started byPaolo <paolo.valente@linaro.org>
First post2016-08-08 16:20 +0200
Last post2016-08-08 22:10 +0200
Articles 3 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [RFD] I/O scheduling in blk-mq Paolo <paolo.valente@linaro.org> - 2016-08-08 16:20 +0200
    Re: [RFD] I/O scheduling in blk-mq Bart Van Assche <bvanassche@acm.org> - 2016-08-08 18:40 +0200
    Re: [RFD] I/O scheduling in blk-mq Omar Sandoval <osandov@osandov.com> - 2016-08-08 22:10 +0200

#1457828 — [RFD] I/O scheduling in blk-mq

FromPaolo <paolo.valente@linaro.org>
Date2016-08-08 16:20 +0200
Subject[RFD] I/O scheduling in blk-mq
Message-ID<s3Tzk-2sm-5@gated-at.bofh.it>
Hi Jens, Tejun, Christoph, all,
AFAIK blk-mq does not yet feature I/O schedulers. In particular, there
is no scheduler providing strong guarantees in terms of
responsiveness, latency for time-sensitive applications and bandwidth
distribution.

For this reason, I'm trying to port BFQ to blk-mq, or to develop
something simpler if even a reduced version of BFQ proves to be too
heavy (this project is supported by Linaro). If you are willing to
provide some feedback in this respect, I would like to ask for
opinions/suggestions on the following two matters, and possibly to
open a more general discussion on I/O scheduling in blk-mq.

1) My idea is to have an independent instance of BFQ, or in general of
the I/O scheduler, executed for each software queue. Then there would
be no global scheduling. The drawback of no global scheduling is that
each process cannot get more than 1/M of the total throughput of the
device, if M is the number of software queues. But, if I'm not
mistaken, it is however unfeasible to give a process more than 1/M of
the total throughput, without lowering the throughput itself. In fact,
giving a process more than 1/M of the total throughput implies serving
its software queue, say Q, more than the others.  The only way to do
it is periodically stopping the service of the other software queues
and dispatching only the requests in Q. But this would reduce
parallelism, which is the main way how blk-mq achieves a very high
throughput. Are these considerations, and, in particular, one
independent I/O scheduler per software queue, sensible?

2) To provide per-process service guarantees, an I/O scheduler must
create per-process internal queues. BFQ and CFQ use I/O contexts to
achieve this goal. Is something like that (or exactly the same)
available also in blk-mq? If so, do you have any suggestion, or link to
documentation/code on how to use what is available in blk-mq?

Thanks,
Paolo

[toc] | [next] | [standalone]


#1457901

FromBart Van Assche <bvanassche@acm.org>
Date2016-08-08 18:40 +0200
Message-ID<s3VKN-3JN-5@gated-at.bofh.it>
In reply to#1457828
On 08/08/16 07:09, Paolo wrote:
> 2) To provide per-process service guarantees, an I/O scheduler must
> create per-process internal queues. BFQ and CFQ use I/O contexts to
> achieve this goal. Is something like that (or exactly the same)
> available also in blk-mq? If so, do you have any suggestion, or link to
> documentation/code on how to use what is available in blk-mq?

Hello Paolo,

I/O contexts are, by definition, data structures that are shared by 
multiple I/O queues. blk-mq reaches high performance by keeping each 
per-CPU queue independent. This means that using I/O contexts in a 
blk-mq I/O scheduler would introduce a contention point and probably 
also a performance bottleneck. So I would appreciate it if multiqueue 
schedulers would avoid constructs similar to I/O contexts.

Thanks,

Bart.

[toc] | [prev] | [next] | [standalone]


#1458204

FromOmar Sandoval <osandov@osandov.com>
Date2016-08-08 22:10 +0200
Message-ID<s3Z22-63f-5@gated-at.bofh.it>
In reply to#1457828
On Mon, Aug 08, 2016 at 04:09:56PM +0200, Paolo wrote:
> Hi Jens, Tejun, Christoph, all,
> AFAIK blk-mq does not yet feature I/O schedulers. In particular, there
> is no scheduler providing strong guarantees in terms of
> responsiveness, latency for time-sensitive applications and bandwidth
> distribution.
> 
> For this reason, I'm trying to port BFQ to blk-mq, or to develop
> something simpler if even a reduced version of BFQ proves to be too
> heavy (this project is supported by Linaro). If you are willing to
> provide some feedback in this respect, I would like to ask for
> opinions/suggestions on the following two matters, and possibly to
> open a more general discussion on I/O scheduling in blk-mq.
> 
> 1) My idea is to have an independent instance of BFQ, or in general of
> the I/O scheduler, executed for each software queue. Then there would
> be no global scheduling. The drawback of no global scheduling is that
> each process cannot get more than 1/M of the total throughput of the
> device, if M is the number of software queues. But, if I'm not
> mistaken, it is however unfeasible to give a process more than 1/M of
> the total throughput, without lowering the throughput itself. In fact,
> giving a process more than 1/M of the total throughput implies serving
> its software queue, say Q, more than the others.  The only way to do
> it is periodically stopping the service of the other software queues
> and dispatching only the requests in Q. But this would reduce
> parallelism, which is the main way how blk-mq achieves a very high
> throughput. Are these considerations, and, in particular, one
> independent I/O scheduler per software queue, sensible?
> 
> 2) To provide per-process service guarantees, an I/O scheduler must
> create per-process internal queues. BFQ and CFQ use I/O contexts to
> achieve this goal. Is something like that (or exactly the same)
> available also in blk-mq? If so, do you have any suggestion, or link to
> documentation/code on how to use what is available in blk-mq?
> 
> Thanks,
> Paolo

Hi, Paolo,

I've been working on I/O scheduling for blk-mq with Jens for the past
few months (splitting time with other small projects), and we're making
good progress. Like you noticed, the hard part isn't really grafting a
scheduler interface onto blk-mq, it's maintaining good scalability while
providing adequate fairness.

We're working towards a scheduler more like deadline and getting the
architectural issues worked out. The goal is some sort of fairness
across all queues. The scheduler-per-software-queue model won't hold up
so well if we have a slower device with an I/O-hungry process on one CPU
and an interactive process on another CPU.

The issue I'm working through now is that on blk-mq, we only have as
many `struct request`s as the hardware has tags, so on a device with a
limited queue depth, it's really hard to do any sort of intelligent
scheduling. The solution for that is switching over to working with
`struct bio`s in the software queues instead, which abstracts away the
hardware capabilities. I have some work in progress at
https://github.com/osandov/linux/tree/blk-mq-iosched, but it's not yet
at feature-parity.

After that, I'll be back to working on the scheduling itself. The vague
idea is to amortize global scheduling decisions, but I don't have much
concrete code behind that yet.

Thanks!
-- 
Omar

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web