Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1495053 > unrolled thread

[PATCH V3 00/11] block-throttle: add .high limit

Started byShaohua Li <shli@fb.com>
First post2016-10-03 23:30 +0200
Last post2016-10-04 21:00 +0200
Articles 10 on this page of 50 — 9 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH V3 00/11] block-throttle: add .high limit Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
    [PATCH v3 10/11] block-throttle: add a simple idle detection Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
    [PATCH v3 07/11] blk-throttle: make throtl_slice tunable Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
    [PATCH v3 08/11] blk-throttle: detect completed idle cgroup Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
    [PATCH v3 06/11] blk-throttle: make sure expire time isn't too big Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
    [PATCH v3 11/11] blk-throttle: ignore idle cgroup limit Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
    [PATCH v3 09/11] block-throttle: make bandwidth change smooth Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
    [PATCH v3 03/11] block-throttle: configure bps/iops limit for cgroup in high limit Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
    [PATCH v3 01/11] block-throttle: prepare support multiple limits Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
    [PATCH v3 02/11] block-throttle: add .high interface Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
    Re: [PATCH V3 00/11] block-throttle: add .high limit Vivek Goyal <vgoyal@redhat.com> - 2016-10-04 15:30 +0200
      Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 18:00 +0200
        Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 18:30 +0200
          Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 19:10 +0200
            Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 19:50 +0200
              Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 21:00 +0200
                Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 21:10 +0200
                  Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 21:20 +0200
                    Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 21:40 +0200
                      Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 22:30 +0200
                        Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-05 14:40 +0200
                          Re: [PATCH V3 00/11] block-throttle: add .high limit Vivek Goyal <vgoyal@redhat.com> - 2016-10-05 15:20 +0200
                            Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-05 16:10 +0200
                          Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-05 17:00 +0200
                            Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-05 21:50 +0200
                              Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-05 22:10 +0200
                              Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 10:00 +0200
                                Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 15:20 +0200
                                  Re: [PATCH V3 00/11] block-throttle: add .high limit Vivek Goyal <vgoyal@redhat.com> - 2016-10-06 19:50 +0200
                                    Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 20:10 +0200
                                      Re: [PATCH V3 00/11] block-throttle: add .high limit Vivek Goyal <vgoyal@redhat.com> - 2016-10-06 20:40 +0200
                                        Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 23:00 +0200
                                    Re: [PATCH V3 00/11] block-throttle: add .high limit Mark Brown <broonie@kernel.org> - 2016-10-06 21:50 +0200
                                Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-07 00:30 +0200
                            Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-05 22:00 +0200
                              Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 09:30 +0200
                        Fwd: [PATCH V3 00/11] block-throttle: add .high limit Kyle Sanderson <kyle.leet@gmail.com> - 2016-10-09 03:20 +0200
                    Re: [PATCH V3 00/11] block-throttle: add .high limit Linus Walleij <linus.walleij@linaro.org> - 2016-10-06 10:10 +0200
                      Re: [PATCH V3 00/11] block-throttle: add .high limit Mark Brown <broonie@kernel.org> - 2016-10-06 13:10 +0200
                        Re: [PATCH V3 00/11] block-throttle: add .high limit "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-10-06 14:00 +0200
                          Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 15:00 +0200
                            Re: [PATCH V3 00/11] block-throttle: add .high limit "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-10-06 16:00 +0200
                              Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 17:10 +0200
                                Re: [PATCH V3 00/11] block-throttle: add .high limit "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-10-06 17:20 +0200
                      Re: [PATCH V3 00/11] block-throttle: add .high limit Heinz Diehl <htd+ml@fritha.org> - 2016-10-08 12:50 +0200
              Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 21:50 +0200
        Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 18:30 +0200
        Re: [PATCH V3 00/11] block-throttle: add .high limit Vivek Goyal <vgoyal@redhat.com> - 2016-10-04 20:20 +0200
          Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 21:00 +0200
            Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 21:00 +0200

Page 3 of 3 — ← Prev page 1 2 [3]


#1496635

FromPaolo Valente <paolo.valente@unimore.it>
Date2016-10-06 15:00 +0200
Message-ID<spgrf-dR-1@gated-at.bofh.it>
In reply to#1496612
> Il giorno 06 ott 2016, alle ore 13:57, Austin S. Hemmelgarn <ahferroin7@gmail.com> ha scritto:
> 
> On 2016-10-06 07:03, Mark Brown wrote:
>> On Thu, Oct 06, 2016 at 10:04:41AM +0200, Linus Walleij wrote:
>>> On Tue, Oct 4, 2016 at 9:14 PM, Tejun Heo <tj@kernel.org> wrote:
>> 
>>>> I get that bfq can be a good compromise on most desktop workloads and
>>>> behave reasonably well for some server workloads with the slice
>>>> expiration mechanism but it really isn't an IO resource partitioning
>>>> mechanism.
>> 
>>> Not just desktops, also Android phones.
>> 
>>> So why not have BFQ as a separate scheduling policy upstream,
>>> alongside CFQ, deadline and noop?
>> 
>> Right.
>> 
>>> We're already doing the per-usecase Kconfig thing for preemption.
>>> But maybe somebody already hates that and want to get rid of it,
>>> I don't know.
>> 
>> Hannes also suggested going back to making BFQ a separate scheduler
>> rather than replacing CFQ earlier, pointing out that it mitigates
>> against the risks of changing CFQ substantially at this point (which
>> seems to be the biggest issue here).
>> 
> ISTR that the original argument for this approach essentially amounted to: 'If it's so much better, why do we need both?'.
> 
> Such an argument is valid only if the new design is better in all respects (which there isn't sufficient information to decide in this case), or the negative aspects are worth the improvements (which is too workload specific to decide for something like this).

All correct, apart from the workload-specific issue, which is not very clear to me. Over the last five years I have not found a single workload for which CFQ is better than BFQ, and none has been suggested.

Anyway, leaving aside this fact, IMO the real problem here is that we are in a catch-22: "we want BFQ to replace CFQ, but, since CFQ is legacy code, then you cannot change, and thus replace, CFQ"

Thanks,
Paolo

--
Paolo Valente
Algogroup
Dipartimento di Scienze Fisiche, Informatiche e Matematiche
Via Campi 213/B
41125 Modena - Italy
http://algogroup.unimore.it/people/paolo/

[toc] | [prev] | [next] | [standalone]


#1496668

From"Austin S. Hemmelgarn" <ahferroin7@gmail.com>
Date2016-10-06 16:00 +0200
Message-ID<sphnj-Oy-9@gated-at.bofh.it>
In reply to#1496635
On 2016-10-06 08:50, Paolo Valente wrote:
>
>> Il giorno 06 ott 2016, alle ore 13:57, Austin S. Hemmelgarn <ahferroin7@gmail.com> ha scritto:
>>
>> On 2016-10-06 07:03, Mark Brown wrote:
>>> On Thu, Oct 06, 2016 at 10:04:41AM +0200, Linus Walleij wrote:
>>>> On Tue, Oct 4, 2016 at 9:14 PM, Tejun Heo <tj@kernel.org> wrote:
>>>
>>>>> I get that bfq can be a good compromise on most desktop workloads and
>>>>> behave reasonably well for some server workloads with the slice
>>>>> expiration mechanism but it really isn't an IO resource partitioning
>>>>> mechanism.
>>>
>>>> Not just desktops, also Android phones.
>>>
>>>> So why not have BFQ as a separate scheduling policy upstream,
>>>> alongside CFQ, deadline and noop?
>>>
>>> Right.
>>>
>>>> We're already doing the per-usecase Kconfig thing for preemption.
>>>> But maybe somebody already hates that and want to get rid of it,
>>>> I don't know.
>>>
>>> Hannes also suggested going back to making BFQ a separate scheduler
>>> rather than replacing CFQ earlier, pointing out that it mitigates
>>> against the risks of changing CFQ substantially at this point (which
>>> seems to be the biggest issue here).
>>>
>> ISTR that the original argument for this approach essentially amounted to: 'If it's so much better, why do we need both?'.
>>
>> Such an argument is valid only if the new design is better in all respects (which there isn't sufficient information to decide in this case), or the negative aspects are worth the improvements (which is too workload specific to decide for something like this).
>
> All correct, apart from the workload-specific issue, which is not very clear to me. Over the last five years I have not found a single workload for which CFQ is better than BFQ, and none has been suggested.
My point is that whether or not BFQ is better depends on the workload. 
You can't test for every workload, so you can't say definitively that 
BFQ is better for every workload.  At a minimum, there are workloads 
where the deadline and noop schedulers are better, but they're very 
domain specific workloads.  Based on the numbers from Shaohua, it looks 
like CFQ has better throughput than BFQ, and that will affect some 
workloads (for most, the improved fairness is worth the reduced 
throughput, but there probably are some cases where it isn't).
>
> Anyway, leaving aside this fact, IMO the real problem here is that we are in a catch-22: "we want BFQ to replace CFQ, but, since CFQ is legacy code, then you cannot change, and thus replace, CFQ"
I agree that that's part of the issue, but I also don't entirely agree 
with the reasoning on it.  Until blk-mq has proper I/O scheduling, 
people will continue to use CFQ, and based on the way things are going, 
it will be multiple months before that happens, whereas BFQ exists and 
is working now.

[toc] | [prev] | [next] | [standalone]


#1496694

FromPaolo Valente <paolo.valente@unimore.it>
Date2016-10-06 17:10 +0200
Message-ID<spit3-1Pe-17@gated-at.bofh.it>
In reply to#1496668
> Il giorno 06 ott 2016, alle ore 15:52, Austin S. Hemmelgarn <ahferroin7@gmail.com> ha scritto:
> 
> On 2016-10-06 08:50, Paolo Valente wrote:
>> 
>>> Il giorno 06 ott 2016, alle ore 13:57, Austin S. Hemmelgarn <ahferroin7@gmail.com> ha scritto:
>>> 
>>> On 2016-10-06 07:03, Mark Brown wrote:
>>>> On Thu, Oct 06, 2016 at 10:04:41AM +0200, Linus Walleij wrote:
>>>>> On Tue, Oct 4, 2016 at 9:14 PM, Tejun Heo <tj@kernel.org> wrote:
>>>> 
>>>>>> I get that bfq can be a good compromise on most desktop workloads and
>>>>>> behave reasonably well for some server workloads with the slice
>>>>>> expiration mechanism but it really isn't an IO resource partitioning
>>>>>> mechanism.
>>>> 
>>>>> Not just desktops, also Android phones.
>>>> 
>>>>> So why not have BFQ as a separate scheduling policy upstream,
>>>>> alongside CFQ, deadline and noop?
>>>> 
>>>> Right.
>>>> 
>>>>> We're already doing the per-usecase Kconfig thing for preemption.
>>>>> But maybe somebody already hates that and want to get rid of it,
>>>>> I don't know.
>>>> 
>>>> Hannes also suggested going back to making BFQ a separate scheduler
>>>> rather than replacing CFQ earlier, pointing out that it mitigates
>>>> against the risks of changing CFQ substantially at this point (which
>>>> seems to be the biggest issue here).
>>>> 
>>> ISTR that the original argument for this approach essentially amounted to: 'If it's so much better, why do we need both?'.
>>> 
>>> Such an argument is valid only if the new design is better in all respects (which there isn't sufficient information to decide in this case), or the negative aspects are worth the improvements (which is too workload specific to decide for something like this).
>> 
>> All correct, apart from the workload-specific issue, which is not very clear to me. Over the last five years I have not found a single workload for which CFQ is better than BFQ, and none has been suggested.
> My point is that whether or not BFQ is better depends on the workload. You can't test for every workload, so you can't say definitively that BFQ is better for every workload.

Yes

>  At a minimum, there are workloads where the deadline and noop schedulers are better, but they're very domain specific workloads.

Definitely

>  Based on the numbers from Shaohua, it looks like CFQ has better throughput than BFQ, and that will affect some workloads (for most, the improved fairness is worth the reduced throughput, but there probably are some cases where it isn't).

Well, no fairness as deadline and noop, but with much less throughput
than deadline and noop, doesn't sound much like the best scheduler for
those workloads.  With BFQ you have service guarantees, with noop or
deadline you have maximum throughput.

>> 
>> Anyway, leaving aside this fact, IMO the real problem here is that we are in a catch-22: "we want BFQ to replace CFQ, but, since CFQ is legacy code, then you cannot change, and thus replace, CFQ"
> I agree that that's part of the issue, but I also don't entirely agree with the reasoning on it.  Until blk-mq has proper I/O scheduling, people will continue to use CFQ, and based on the way things are going, it will be multiple months before that happens, whereas BFQ exists and is working now.

Exactly!

Thanks,
Paolo

> --
> To unsubscribe from this list: send the line "unsubscribe linux-block" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html


--
Paolo Valente
Algogroup
Dipartimento di Scienze Fisiche, Informatiche e Matematiche
Via Campi 213/B
41125 Modena - Italy
http://algogroup.unimore.it/people/paolo/

[toc] | [prev] | [next] | [standalone]


#1496701

From"Austin S. Hemmelgarn" <ahferroin7@gmail.com>
Date2016-10-06 17:20 +0200
Message-ID<spiCJ-1Up-13@gated-at.bofh.it>
In reply to#1496694
On 2016-10-06 11:05, Paolo Valente wrote:
>
>> Il giorno 06 ott 2016, alle ore 15:52, Austin S. Hemmelgarn <ahferroin7@gmail.com> ha scritto:
>>
>> On 2016-10-06 08:50, Paolo Valente wrote:
>>>
>>>> Il giorno 06 ott 2016, alle ore 13:57, Austin S. Hemmelgarn <ahferroin7@gmail.com> ha scritto:
>>>>
>>>> On 2016-10-06 07:03, Mark Brown wrote:
>>>>> On Thu, Oct 06, 2016 at 10:04:41AM +0200, Linus Walleij wrote:
>>>>>> On Tue, Oct 4, 2016 at 9:14 PM, Tejun Heo <tj@kernel.org> wrote:
>>>>>
>>>>>>> I get that bfq can be a good compromise on most desktop workloads and
>>>>>>> behave reasonably well for some server workloads with the slice
>>>>>>> expiration mechanism but it really isn't an IO resource partitioning
>>>>>>> mechanism.
>>>>>
>>>>>> Not just desktops, also Android phones.
>>>>>
>>>>>> So why not have BFQ as a separate scheduling policy upstream,
>>>>>> alongside CFQ, deadline and noop?
>>>>>
>>>>> Right.
>>>>>
>>>>>> We're already doing the per-usecase Kconfig thing for preemption.
>>>>>> But maybe somebody already hates that and want to get rid of it,
>>>>>> I don't know.
>>>>>
>>>>> Hannes also suggested going back to making BFQ a separate scheduler
>>>>> rather than replacing CFQ earlier, pointing out that it mitigates
>>>>> against the risks of changing CFQ substantially at this point (which
>>>>> seems to be the biggest issue here).
>>>>>
>>>> ISTR that the original argument for this approach essentially amounted to: 'If it's so much better, why do we need both?'.
>>>>
>>>> Such an argument is valid only if the new design is better in all respects (which there isn't sufficient information to decide in this case), or the negative aspects are worth the improvements (which is too workload specific to decide for something like this).
>>>
>>> All correct, apart from the workload-specific issue, which is not very clear to me. Over the last five years I have not found a single workload for which CFQ is better than BFQ, and none has been suggested.
>> My point is that whether or not BFQ is better depends on the workload. You can't test for every workload, so you can't say definitively that BFQ is better for every workload.
>
> Yes
>
>>  At a minimum, there are workloads where the deadline and noop schedulers are better, but they're very domain specific workloads.
>
> Definitely
>
>>  Based on the numbers from Shaohua, it looks like CFQ has better throughput than BFQ, and that will affect some workloads (for most, the improved fairness is worth the reduced throughput, but there probably are some cases where it isn't).
>
> Well, no fairness as deadline and noop, but with much less throughput
> than deadline and noop, doesn't sound much like the best scheduler for
> those workloads.  With BFQ you have service guarantees, with noop or
> deadline you have maximum throughput.
And with CFQ you have something in between, which is half of why I think 
CFQ is still worth keeping (the other half being the people who 
inevitably want to stay on CFQ).  And TBH, deadline and noop only give 
good throughput with specific workloads (and in the case of noop, it's 
usually only useful on tiny systems where the overhead of scheduling is 
greater than the time saved by doing so (like some very low power 
embedded systems), or when you have scheduling done elsewher in the 
storage stack (like in a VM)).
>
>>>
>>> Anyway, leaving aside this fact, IMO the real problem here is that we are in a catch-22: "we want BFQ to replace CFQ, but, since CFQ is legacy code, then you cannot change, and thus replace, CFQ"
>> I agree that that's part of the issue, but I also don't entirely agree with the reasoning on it.  Until blk-mq has proper I/O scheduling, people will continue to use CFQ, and based on the way things are going, it will be multiple months before that happens, whereas BFQ exists and is working now.
>
> Exactly!
>
> Thanks,
> Paolo
>

[toc] | [prev] | [next] | [standalone]


#1497705

FromHeinz Diehl <htd+ml@fritha.org>
Date2016-10-08 12:50 +0200
Message-ID<spXmx-5Yr-17@gated-at.bofh.it>
In reply to#1496204
On 06.10.2016, Linus Walleij wrote: 

> Not just desktops, also Android phones.
 
> So why not have BFQ as a separate scheduling policy upstream,
> alongside CFQ, deadline and noop?

Please, just do it! CFQ and deadline perform horribly on Android, and
the same is true for desktop systems based on SSD drives.

What e.g. this test shows is true for all of my machines:
https://www.youtube.com/watch?v=1cjZeaCXIyM

I've been using BFQ nearly right from its inception, and it improves
all of my systems noticeably, in fact in such a positive way that I
don't update to a newer kernel unless there is a BFQ patch ready for
it.

[toc] | [prev] | [next] | [standalone]


#1495574

FromPaolo Valente <paolo.valente@unimore.it>
Date2016-10-04 21:50 +0200
Message-ID<soDSV-71V-1@gated-at.bofh.it>
In reply to#1495539
> Il giorno 04 ott 2016, alle ore 20:28, Shaohua Li <shli@fb.com> ha scritto:
> 
> On Tue, Oct 04, 2016 at 07:43:48PM +0200, Paolo Valente wrote:
>> 
>>> Il giorno 04 ott 2016, alle ore 19:28, Shaohua Li <shli@fb.com> ha scritto:
>>> 
>>> On Tue, Oct 04, 2016 at 07:01:39PM +0200, Paolo Valente wrote:
>>>> 
>>>>> Il giorno 04 ott 2016, alle ore 18:27, Tejun Heo <tj@kernel.org> ha scritto:
>>>>> 
>>>>> Hello,
>>>>> 
>>>>> On Tue, Oct 04, 2016 at 06:22:28PM +0200, Paolo Valente wrote:
>>>>>> Could you please elaborate more on this point?  BFQ uses sectors
>>>>>> served to measure service, and, on the all the fast devices on which
>>>>>> we have tested it, it accurately distributes
>>>>>> bandwidth as desired, redistributes excess bandwidth with any issue,
>>>>>> and guarantees high responsiveness and low latency at application and
>>>>>> system level (e.g., ~0 drop rate in video playback, with any background
>>>>>> workload tested).
>>>>> 
>>>>> The same argument as before.  Bandwidth is a very bad measure of IO
>>>>> resources spent.  For specific use cases (like desktop or whatever),
>>>>> this can work but not generally.
>>>>> 
>>>> 
>>>> Actually, we have already discussed this point, and IMHO the arguments
>>>> that (apparently) convinced you that bandwidth is the most relevant
>>>> service guarantee for I/O in desktops and the like, prove that
>>>> bandwidth is the most important service guarantee in servers too.
>>>> 
>>>> Again, all the examples I can think of seem to confirm it:
>>>> . file hosting: a good service must guarantee reasonable read/write,
>>>> i.e., download/upload, speeds to users
>>>> . file streaming: a good service must guarantee low drop rates, and
>>>> this can be guaranteed only by guaranteeing bandwidth and latency
>>>> . web hosting: high bandwidth and low latency needed here too
>>>> . clouds: high bw and low latency needed to let, e.g., users of VMs
>>>> enjoy high responsiveness and, for example, reasonable file-copy
>>>> time
>>>> ...
>>>> 
>>>> To put in yet another way, with packet I/O in, e.g., clouds, there are
>>>> basically the same issues, and the main goal is again guaranteeing
>>>> bandwidth and low latency among nodes.
>>>> 
>>>> Could you please provide a concrete server example (assuming we still
>>>> agree about desktops), where I/O bandwidth does not matter while time
>>>> does?
>>> 
>>> I don't think IO bandwidth does not matter. The problem is bandwidth can't
>>> measure IO cost. For example, you can't say 8k IO costs 2x IO resource than 4k
>>> IO.
>>> 
>> 
>> For what goal do you need to be able to say this, once you succeeded
>> in guaranteeing bandwidth and low latency to each
>> process/client/group/node/user?
> 
> I think we are discussing if bandwidth should be used to measure IO for
> propotional IO scheduling.


Yes. But my point is upstream. It's something like this:

Can bandwidth and low latency guarantees be provided with a
sector-based proportional-share scheduler?

YOUR ANSWER: No, then we need to look for other non-trivial solutions.
Hence your arguments in this discussion.

MY ANSWER: Yes, I have already achieved this goal for years now, with
a publicly available, proportional-share scheduler.  A lot of test
results with many devices, papers discussing details, demos, and so on
are available too.

> Since bandwidth can't measure the cost and you are
> using it to do arbitration, you will either have low latency but unfair
> bandwidth, or fair bandwidth but some workloads have unexpected high latency.
> But it might be ok depending on the latency target (for example, you can set
> the latency target high, so low latency is guaranteed*) and workload
> characteristics. I think the bandwidth based proporional scheduling will only
> work for workloads disk isn't fully utilized.
> 
>>>>>> Could you please suggest me some test to show how sector-based
>>>>>> guarantees fails?
>>>>> 
>>>>> Well, mix 4k random and sequential workloads and try to distribute the
>>>>> acteual IO resources.
>>>>> 
>>>> 
>>>> 
>>>> If I'm not mistaken, we have already gone through this example too,
>>>> and I thought we agreed on what service scheme worked best, again
>>>> focusing only on desktops.  To make a long story short(er), here is a
>>>> snippet from one of our last exchanges.
>>>> 
>>>> ----------
>>>> 
>>>> On Sat, Apr 16, 2016 at 12:08:44AM +0200, Paolo Valente wrote:
>>>>> Maybe the source of confusion is the fact that a simple sector-based,
>>>>> proportional share scheduler always distributes total bandwidth
>>>>> according to weights. The catch is the additional BFQ rule: random
>>>>> workloads get only time isolation, and are charged for full budgets,
>>>>> so as to not affect the schedule of quasi-sequential workloads. So,
>>>>> the correct claim for BFQ is that it distributes total bandwidth
>>>>> according to weights (only) when all competing workloads are
>>>>> quasi-sequential. If some workloads are random, then these workloads
>>>>> are just time scheduled. This does break proportional-share bandwidth
>>>>> distribution with mixed workloads, but, much more importantly, saves
>>>>> both total throughput and individual bandwidths of quasi-sequential
>>>>> workloads.
>>>>> 
>>>>> We could then check whether I did succeed in tuning timeouts and
>>>>> budgets so as to achieve the best tradeoffs. But this is probably a
>>>>> second-order problem as of now.
>>> 
>>> I don't see why random/sequential matters for SSD. what really matters is
>>> request size and IO depth. Time scheduling is skeptical too, as workloads can
>>> dispatch all IO within almost 0 time in high queue depth disks.
>>> 
>> 
>> That's an orthogonal issue.  If what matter is, e.g., size, then it is
>> enough to replace "sequential I/O" with "large-request I/O".  In case
>> I have been too vague, here is an example: I mean that, e.g, in an I/O
>> scheduler you replace the function that computes whether a queue is
>> seeky based on request distance, with a function based on
>> request size.  And this is exactly what has been already done, for
>> example, in CFQ:
>> 
>> 	if (blk_queue_nonrot(cfqd->queue))
>> 		cfqq->seek_history |= (n_sec < CFQQ_SECT_THR_NONROT);
>> 	else
>> 		cfqq->seek_history |= (sdist > CFQQ_SEEK_THR);
> 
> CFQ is known not fair for SSD especially high queue depth SSD, so this doesn't
> mean correctness.

I'm afraid CFQ is unfair for reasons that have little or nothing to do
with the above lines of code (which I pasted just to give you an
example, sorry for creating a misunderstanding).

> And based on request size for idle detection (so let cfqq
> backlog the disk) isn't very good. iodepth 1 4k workload could be idle, but
> iodepth 128 4k workload likely isn't idle (and the workload can dispatch 128
> requests in almost 0 time in high queue depth disk).
> 

That's absolutely true.  And it is one of the most challenging issues
I have addressed in BFQ.  So far the solutions I have found proved to
work well.  But, as I said to Tejun, if you have a concrete example
for which you expect BFQ to fail, just tell me and I will try.
Maximum depth is 32 with blk devices (if I'm not missing something,
given my limited expertise), but that would be probably enough to
prove your point.

Let me add just a comment, to not be misunderstood.  I'm not
undervaluing your proposal.  I'm trying to point out that sector-based
proportional share works, and it is likely to be the best solution
exactly with devices with varying bandwidth and deep queues.  Yet I do
think that your proposal is a good and accurately-designed solution,
definitely necessary until good schedulers will be
available (of course I mean sector-based schedulers! ;) ).

Thanks,
Paolo

> Thanks,
> Shaohua


--
Paolo Valente
Algogroup
Dipartimento di Scienze Fisiche, Informatiche e Matematiche
Via Campi 213/B
41125 Modena - Italy
http://algogroup.unimore.it/people/paolo/

[toc] | [prev] | [next] | [standalone]


#1495508

FromPaolo Valente <paolo.valente@unimore.it>
Date2016-10-04 18:30 +0200
Message-ID<soALn-5bt-15@gated-at.bofh.it>
In reply to#1495488
> Il giorno 04 ott 2016, alle ore 17:56, Tejun Heo <tj@kernel.org> ha scritto:
> 
> Hello, Vivek.
> 
> On Tue, Oct 04, 2016 at 09:28:05AM -0400, Vivek Goyal wrote:
>> On Mon, Oct 03, 2016 at 02:20:19PM -0700, Shaohua Li wrote:
>>> Hi,
>>> 
>>> The background is we don't have an ioscheduler for blk-mq yet, so we can't
>>> prioritize processes/cgroups.
>> 
>> So this is an interim solution till we have ioscheduler for blk-mq?
> 
> It's a common permanent solution which applies to both !mq and mq.
> 
>>> This patch set tries to add basic arbitration
>>> between cgroups with blk-throttle. It adds a new limit io.high for
>>> blk-throttle. It's only for cgroup2.
>>> 
>>> io.max is a hard limit throttling. cgroups with a max limit never dispatch more
>>> IO than their max limit. While io.high is a best effort throttling. cgroups
>>> with high limit can run above their high limit at appropriate time.
>>> Specifically, if all cgroups reach their high limit, all cgroups can run above
>>> their high limit. If any cgroup runs under its high limit, all other cgroups
>>> will run according to their high limit.
>> 
>> Hi Shaohua,
>> 
>> I still don't understand why we should not implement a weight based
>> proportional IO mechanism and how this mechanism is better than proportional IO .
> 
> Oh, if we actually can implement proportional IO control, it'd be
> great.  The problem is that we have no way of knowing IO cost for
> highspeed ssd devices.  CFQ gets around the problem by using the
> walltime as the measure of resource usage and scheduling time slices,
> which works fine for rotating disks but horribly for highspeed ssds.
> 

Could you please elaborate more on this point?  BFQ uses sectors
served to measure service, and, on the all the fast devices on which
we have tested it, it accurately distributes
bandwidth as desired, redistributes excess bandwidth with any issue,
and guarantees high responsiveness and low latency at application and
system level (e.g., ~0 drop rate in video playback, with any background
workload tested).

Could you please suggest me some test to show how sector-based
guarantees fails?

Thanks,
Paolo

> We can get some semblance of proportional control by just counting bw
> or iops but both break down badly as a means to measure the actual
> resource consumption depending on the workload.  While limit based
> control is more tedious to configure, it doesn't misrepresent what's
> going on and is a lot less likely to produce surprising outcomes.
> 
> We *can* try to concoct something which tries to do proportional
> control for highspeed ssds but that's gonna be quite a bit of
> complexity and I'm not so sure it'd be justifiable given that we can't
> even figure out measurement of the most basic operating unit.
> 
>> Agreed that we have issues with proportional IO and we don't have good
>> solutions for these problems. But I can't see that how this mechanism
>> will overcome these problems either.
> 
> It mostly defers the burden to the one who's configuring the limits
> and expects it to know the characteristics of the device and workloads
> and configure accordingly.  It's quite a bit more tedious to use but
> should be able to cover good portion of use cases without being overly
> complicated.  I agree that it'd be nice to have a simple proportional
> control but as you said can't see a good solution for it at the
> moment.
> 
>> IIRC, biggest issue with proportional IO was that a low prio group might
>> fill up the device queue with plenty of IO requests and later when high
>> prio cgroup comes, it will still experience latencies anyway. And solution
>> to the problem probably would be to get some awareness in device about 
>> priority of request and map weights to those priority. That way higher
>> prio requests get prioritized.
> 
> Nah, the real problem is that we can't even decide what the
> proportions should be based on.  The most fundamental part is missing.
> 
>> Or run device at lower queue depth. That will improve latencies but migth
>> reduce overall throughput.
> 
> And that we can't do this (and thus basically operate close to
> scheduling time slices) for highspeed ssds.
> 
>> Or thorottle number of buffered writes (as Jens's writeback throttling)
>> patches were doing. Buffered writes seem to be biggest culprit for 
>> increased latencies and being able to control these should help.
> 
> That's a different topic.
> 
>> ioprio/weight based proportional IO mechanism is much more generic and
>> much easier to configure for any kind of storage. io.high is absolute
>> limit and makes it much harder to configure. One needs to know a lot
>> about underlying volume/device's bandwidth (which varies a lot anyway
>> based on workload).
> 
> Yeap, no disagreement there, but it still is a workable solution.
> 
>> IMHO, we seem to be trying to cater to one specific use case using
>> this mechanism. Something ioprio/weight based will be much more
>> generic and we should explore implementing that along with building
>> notion of ioprio in devices. When these two work together, we might
>> be able to see good results. Just software mechanism alone might not
>> be enough.
> 
> I don't think it's catering to specific use cases.  It is a generic
> mechanism which demands knowledge and experimentation to configure.
> It's more a way for the kernel to cop out and defer figuring out
> device characteristics to userland.  If you have a better idea, I'm
> all ears.
> 
> Thanks.
> 
> -- 
> tejun
> --
> To unsubscribe from this list: send the line "unsubscribe linux-block" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html


--
Paolo Valente
Algogroup
Dipartimento di Scienze Fisiche, Informatiche e Matematiche
Via Campi 213/B
41125 Modena - Italy
http://algogroup.unimore.it/people/paolo/

[toc] | [prev] | [next] | [standalone]


#1495547

FromVivek Goyal <vgoyal@redhat.com>
Date2016-10-04 20:20 +0200
Message-ID<soCtP-6h4-5@gated-at.bofh.it>
In reply to#1495488
On Tue, Oct 04, 2016 at 11:56:16AM -0400, Tejun Heo wrote:
> Hello, Vivek.
> 
> On Tue, Oct 04, 2016 at 09:28:05AM -0400, Vivek Goyal wrote:
> > On Mon, Oct 03, 2016 at 02:20:19PM -0700, Shaohua Li wrote:
> > > Hi,
> > > 
> > > The background is we don't have an ioscheduler for blk-mq yet, so we can't
> > > prioritize processes/cgroups.
> > 
> > So this is an interim solution till we have ioscheduler for blk-mq?
> 
> It's a common permanent solution which applies to both !mq and mq.
> 
> > > This patch set tries to add basic arbitration
> > > between cgroups with blk-throttle. It adds a new limit io.high for
> > > blk-throttle. It's only for cgroup2.
> > > 
> > > io.max is a hard limit throttling. cgroups with a max limit never dispatch more
> > > IO than their max limit. While io.high is a best effort throttling. cgroups
> > > with high limit can run above their high limit at appropriate time.
> > > Specifically, if all cgroups reach their high limit, all cgroups can run above
> > > their high limit. If any cgroup runs under its high limit, all other cgroups
> > > will run according to their high limit.
> > 
> > Hi Shaohua,
> > 
> > I still don't understand why we should not implement a weight based
> > proportional IO mechanism and how this mechanism is better than proportional IO .
> 
> Oh, if we actually can implement proportional IO control, it'd be
> great.  The problem is that we have no way of knowing IO cost for
> highspeed ssd devices.  CFQ gets around the problem by using the
> walltime as the measure of resource usage and scheduling time slices,
> which works fine for rotating disks but horribly for highspeed ssds.
> 
> We can get some semblance of proportional control by just counting bw
> or iops but both break down badly as a means to measure the actual
> resource consumption depending on the workload.  While limit based
> control is more tedious to configure, it doesn't misrepresent what's
> going on and is a lot less likely to produce surprising outcomes.
> 
> We *can* try to concoct something which tries to do proportional
> control for highspeed ssds but that's gonna be quite a bit of
> complexity and I'm not so sure it'd be justifiable given that we can't
> even figure out measurement of the most basic operating unit.

Hi Tejun,

Agreed that we don't have a good basic unit to measure IO cost. I was
thinking of measuring cost in terms of sectors as that's simple and
gets more accurate on faster devices with almost no seek penalty. And
in fact this proposal is also providing fairness in terms of bandwitdh.
One extra feature seems to be this notion of minimum bandwidth for each
cgroup and until and unless all competing groups have met their minimum,
other cgroups can't cross their limits.

(BTW, should we call io.high, io.minimum instead. To say, this is the
 minimum bandwidth group should get before others get to cross their
 minimum limit till max limit).

> 
> > Agreed that we have issues with proportional IO and we don't have good
> > solutions for these problems. But I can't see that how this mechanism
> > will overcome these problems either.
> 
> It mostly defers the burden to the one who's configuring the limits
> and expects it to know the characteristics of the device and workloads
> and configure accordingly.  It's quite a bit more tedious to use but
> should be able to cover good portion of use cases without being overly
> complicated.  I agree that it'd be nice to have a simple proportional
> control but as you said can't see a good solution for it at the
> moment.

Ok, so idea is that if we can't provide something accurate in kernel,
then expose a very low level knob, which is harder to configure but
should work in some cases where users know their devices and workload
very well. 

> 
> > IIRC, biggest issue with proportional IO was that a low prio group might
> > fill up the device queue with plenty of IO requests and later when high
> > prio cgroup comes, it will still experience latencies anyway. And solution
> > to the problem probably would be to get some awareness in device about 
> > priority of request and map weights to those priority. That way higher
> > prio requests get prioritized.
> 
> Nah, the real problem is that we can't even decide what the
> proportions should be based on.  The most fundamental part is missing.
> 
> > Or run device at lower queue depth. That will improve latencies but migth
> > reduce overall throughput.
> 
> And that we can't do this (and thus basically operate close to
> scheduling time slices) for highspeed ssds.
> 
> > Or thorottle number of buffered writes (as Jens's writeback throttling)
> > patches were doing. Buffered writes seem to be biggest culprit for 
> > increased latencies and being able to control these should help.
> 
> That's a different topic.
> 
> > ioprio/weight based proportional IO mechanism is much more generic and
> > much easier to configure for any kind of storage. io.high is absolute
> > limit and makes it much harder to configure. One needs to know a lot
> > about underlying volume/device's bandwidth (which varies a lot anyway
> > based on workload).
> 
> Yeap, no disagreement there, but it still is a workable solution.
> 
> > IMHO, we seem to be trying to cater to one specific use case using
> > this mechanism. Something ioprio/weight based will be much more
> > generic and we should explore implementing that along with building
> > notion of ioprio in devices. When these two work together, we might
> > be able to see good results. Just software mechanism alone might not
> > be enough.
> 
> I don't think it's catering to specific use cases.  It is a generic
> mechanism which demands knowledge and experimentation to configure.
> It's more a way for the kernel to cop out and defer figuring out
> device characteristics to userland.  If you have a better idea, I'm
> all ears.

I don't think I have a better idea as such. Once we had talked and you
mentioned that for faster devices we should probably do some token based
mechanism (which I believe would probably mean sector based IO
accounting). 

If a proportional IO controller based on sector as unit of measurement
not good enough and does not solve the issues real world workloads are
facing, then we can think of giving additional control in
blk-throttle to atleast get some of the use cases working.

Thanks
Vivek

[toc] | [prev] | [next] | [standalone]


#1495555

FromTejun Heo <tj@kernel.org>
Date2016-10-04 21:00 +0200
Message-ID<soD6x-6vd-3@gated-at.bofh.it>
In reply to#1495547
Hello, Vivek.

On Tue, Oct 04, 2016 at 02:12:45PM -0400, Vivek Goyal wrote:
> Agreed that we don't have a good basic unit to measure IO cost. I was
> thinking of measuring cost in terms of sectors as that's simple and
> gets more accurate on faster devices with almost no seek penalty. And

If this were true, we could simply base everything on bandwidth;
unfortunately, even highspeed ssds perform wildly differently
depending on the specifics of workloads.

> in fact this proposal is also providing fairness in terms of bandwitdh.
> One extra feature seems to be this notion of minimum bandwidth for each
> cgroup and until and unless all competing groups have met their minimum,
> other cgroups can't cross their limits.

Haven't read the patches yet but it should allow regulating in terms
of both bandwidth and iops.

> (BTW, should we call io.high, io.minimum instead. To say, this is the
>  minimum bandwidth group should get before others get to cross their
>  minimum limit till max limit).

The naming convetion is min, low, high, max but I'm not sure "min",
which means hard minimum amount (whether guaranteed or best-effort),
quite makes sense here.

> > It mostly defers the burden to the one who's configuring the limits
> > and expects it to know the characteristics of the device and workloads
> > and configure accordingly.  It's quite a bit more tedious to use but
> > should be able to cover good portion of use cases without being overly
> > complicated.  I agree that it'd be nice to have a simple proportional
> > control but as you said can't see a good solution for it at the
> > moment.
> 
> Ok, so idea is that if we can't provide something accurate in kernel,
> then expose a very low level knob, which is harder to configure but
> should work in some cases where users know their devices and workload
> very well. 

Yeah, that's the basic idea for this approach.  It'd be great if we
eventually end up with proper proportional control but having
something low level is useful anyway, so...

> > I don't think it's catering to specific use cases.  It is a generic
> > mechanism which demands knowledge and experimentation to configure.
> > It's more a way for the kernel to cop out and defer figuring out
> > device characteristics to userland.  If you have a better idea, I'm
> > all ears.
> 
> I don't think I have a better idea as such. Once we had talked and you
> mentioned that for faster devices we should probably do some token based
> mechanism (which I believe would probably mean sector based IO
> accounting). 

That's more about the implementation strategy and doesn't affect
whether we support bw, iops or combined configurations.  In terms of
implementation, I still think it'd be great to have something token
based with per-cpu batch to lower the cpu overhead on highspeed
devices but that shouldn't really affect the semantics.

Thanks.

-- 
tejun

[toc] | [prev] | [next] | [standalone]


#1495556

FromPaolo Valente <paolo.valente@unimore.it>
Date2016-10-04 21:00 +0200
Message-ID<soD6x-6vd-7@gated-at.bofh.it>
In reply to#1495555
> Il giorno 04 ott 2016, alle ore 20:50, Tejun Heo <tj@kernel.org> ha scritto:
> 
> Hello, Vivek.
> 
> On Tue, Oct 04, 2016 at 02:12:45PM -0400, Vivek Goyal wrote:
>> Agreed that we don't have a good basic unit to measure IO cost. I was
>> thinking of measuring cost in terms of sectors as that's simple and
>> gets more accurate on faster devices with almost no seek penalty. And
> 
> If this were true, we could simply base everything on bandwidth;
> unfortunately, even highspeed ssds perform wildly differently
> depending on the specifics of workloads.
> 

If you base your throttler or scheduler on time, and bandwidth varies
with workload, as you correctly point out, then the result is loss of
control on bandwidth distribution, hence unfairness and
hard-to-control (high) latency.  If you use BFQ's approach, as we
already discussed with numbers and examples, you have stable fairness
and low latency.  More precisely, given your workload, you can even
compute formally the strong service guarantees you provide.

Thanks,
Paolo

>> in fact this proposal is also providing fairness in terms of bandwitdh.
>> One extra feature seems to be this notion of minimum bandwidth for each
>> cgroup and until and unless all competing groups have met their minimum,
>> other cgroups can't cross their limits.
> 
> Haven't read the patches yet but it should allow regulating in terms
> of both bandwidth and iops.
> 
>> (BTW, should we call io.high, io.minimum instead. To say, this is the
>> minimum bandwidth group should get before others get to cross their
>> minimum limit till max limit).
> 
> The naming convetion is min, low, high, max but I'm not sure "min",
> which means hard minimum amount (whether guaranteed or best-effort),
> quite makes sense here.
> 
>>> It mostly defers the burden to the one who's configuring the limits
>>> and expects it to know the characteristics of the device and workloads
>>> and configure accordingly.  It's quite a bit more tedious to use but
>>> should be able to cover good portion of use cases without being overly
>>> complicated.  I agree that it'd be nice to have a simple proportional
>>> control but as you said can't see a good solution for it at the
>>> moment.
>> 
>> Ok, so idea is that if we can't provide something accurate in kernel,
>> then expose a very low level knob, which is harder to configure but
>> should work in some cases where users know their devices and workload
>> very well. 
> 
> Yeah, that's the basic idea for this approach.  It'd be great if we
> eventually end up with proper proportional control but having
> something low level is useful anyway, so...
> 
>>> I don't think it's catering to specific use cases.  It is a generic
>>> mechanism which demands knowledge and experimentation to configure.
>>> It's more a way for the kernel to cop out and defer figuring out
>>> device characteristics to userland.  If you have a better idea, I'm
>>> all ears.
>> 
>> I don't think I have a better idea as such. Once we had talked and you
>> mentioned that for faster devices we should probably do some token based
>> mechanism (which I believe would probably mean sector based IO
>> accounting). 
> 
> That's more about the implementation strategy and doesn't affect
> whether we support bw, iops or combined configurations.  In terms of
> implementation, I still think it'd be great to have something token
> based with per-cpu batch to lower the cpu overhead on highspeed
> devices but that shouldn't really affect the semantics.
> 
> Thanks.
> 
> -- 
> tejun
> --
> To unsubscribe from this list: send the line "unsubscribe linux-block" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html


--
Paolo Valente
Algogroup
Dipartimento di Scienze Fisiche, Informatiche e Matematiche
Via Campi 213/B
41125 Modena - Italy
http://algogroup.unimore.it/people/paolo/

[toc] | [prev] | [standalone]


Page 3 of 3 — ← Prev page 1 2 [3]

Back to top | Article view | linux.kernel


csiph-web