Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1495053 > unrolled thread
| Started by | Shaohua Li <shli@fb.com> |
|---|---|
| First post | 2016-10-03 23:30 +0200 |
| Last post | 2016-10-04 21:00 +0200 |
| Articles | 20 on this page of 50 — 9 participants |
Back to article view | Back to linux.kernel
[PATCH V3 00/11] block-throttle: add .high limit Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
[PATCH v3 10/11] block-throttle: add a simple idle detection Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
[PATCH v3 07/11] blk-throttle: make throtl_slice tunable Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
[PATCH v3 08/11] blk-throttle: detect completed idle cgroup Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
[PATCH v3 06/11] blk-throttle: make sure expire time isn't too big Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
[PATCH v3 11/11] blk-throttle: ignore idle cgroup limit Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
[PATCH v3 09/11] block-throttle: make bandwidth change smooth Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
[PATCH v3 03/11] block-throttle: configure bps/iops limit for cgroup in high limit Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
[PATCH v3 01/11] block-throttle: prepare support multiple limits Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
[PATCH v3 02/11] block-throttle: add .high interface Shaohua Li <shli@fb.com> - 2016-10-03 23:30 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Vivek Goyal <vgoyal@redhat.com> - 2016-10-04 15:30 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 18:00 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 18:30 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 19:10 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 19:50 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 21:00 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 21:10 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 21:20 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 21:40 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 22:30 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-05 14:40 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Vivek Goyal <vgoyal@redhat.com> - 2016-10-05 15:20 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-05 16:10 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-05 17:00 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-05 21:50 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-05 22:10 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 10:00 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 15:20 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Vivek Goyal <vgoyal@redhat.com> - 2016-10-06 19:50 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 20:10 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Vivek Goyal <vgoyal@redhat.com> - 2016-10-06 20:40 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 23:00 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Mark Brown <broonie@kernel.org> - 2016-10-06 21:50 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-07 00:30 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-05 22:00 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 09:30 +0200
Fwd: [PATCH V3 00/11] block-throttle: add .high limit Kyle Sanderson <kyle.leet@gmail.com> - 2016-10-09 03:20 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Linus Walleij <linus.walleij@linaro.org> - 2016-10-06 10:10 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Mark Brown <broonie@kernel.org> - 2016-10-06 13:10 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-10-06 14:00 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 15:00 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-10-06 16:00 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-06 17:10 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit "Austin S. Hemmelgarn" <ahferroin7@gmail.com> - 2016-10-06 17:20 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Heinz Diehl <htd+ml@fritha.org> - 2016-10-08 12:50 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 21:50 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 18:30 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Vivek Goyal <vgoyal@redhat.com> - 2016-10-04 20:20 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Tejun Heo <tj@kernel.org> - 2016-10-04 21:00 +0200
Re: [PATCH V3 00/11] block-throttle: add .high limit Paolo Valente <paolo.valente@unimore.it> - 2016-10-04 21:00 +0200
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-05 14:40 +0200 |
| Message-ID | <soTEm-Y7-9@gated-at.bofh.it> |
| In reply to | #1495587 |
> Il giorno 04 ott 2016, alle ore 22:27, Tejun Heo <tj@kernel.org> ha scritto: > > Hello, Paolo. > > On Tue, Oct 04, 2016 at 09:29:48PM +0200, Paolo Valente wrote: >>> Hmm... I think we already discussed this but here's a really simple >>> case. There are three unknown workloads A, B and C and we want to >>> give A certain best-effort guarantees (let's say around 80% of the >>> underlying device) whether A is sharing the device with B or C. >> >> That's the same example that you proposed me in our previous >> discussion. For this example I showed you, with many boring numbers, >> that with BFQ you get the most accurate distribution of the resource. > > Yes, it is about the same example and what I understood was that > "accurate distribution of the resources" holds as long as the > randomness is incidental (ie. due to layout on the filesystem and so > on) with the slice expiration mechanism offsetting the actually random > workloads. > For completeness, this property holds whatever the workload is, especially even if it changes. >> If you have enough stamina, I can repeat them again. To save your > > I'll go back to the thread and re-read them. > Maybe we can make this less boring, see the end of this email. >> patience, here is a very brief summary. In a concrete use case, the >> unknown workloads turn into something like this: there will be a first >> time interval during which A happens to be, say, sequential, B happens >> to be, say, random and C happens to be, say, quasi-sequential. Then >> there will be a next time interval during which their characteristics >> change, and so on. It is easy (but boring, I acknowledge it) to show >> that, for each of these time intervals BFQ provides the best possible >> service in terms of fairness, bandwidth distribution, stability and so >> on. Why? Because of the elastic bandwidth-time scheduling of BFQ >> that we already discussed, and because BFQ is naturally accurate in >> redistributing aggregate throughput proportionally, when needed. > > Yeah, that's what I remember and for workload above certain level of > randomness its time consumption is mapped to bw, right? > Exactly. >>> I get that bfq can be a good compromise on most desktop workloads and >>> behave reasonably well for some server workloads with the slice >>> expiration mechanism but it really isn't an IO resource partitioning >>> mechanism. >> >> Right. My argument is that BFQ enables you to give to each client the >> bandwidth and low-latency guarantees you want. And this IMO is way >> better than partitioning a resource and then getting unavoidable >> unfairness and high latency. > > But that statement only holds while bw is the main thing to guarantee, > no? The level of isolation that we're looking for here is fairly > strict adherence to sub/few-milliseconds in terms of high percentile > scheduling latency while within the configured bw/iops limits, not > "overall this device is being used pretty well". > Guaranteeing such a short-term latency, while guaranteeing not just bw limits, but also proportional share distribution of the bw, is the reason why we have devised BFQ years ago. Anyway, to avoid going on with trying speculations and arguments, let me retry with a practical proposal. BFQ is out there, free. Let's just test, measure and check whether we have already a solution to the problems you/we are still trying to solve in Linux. In this respect, for your generic, unpredictable scenario to make sense, there must exist at least one real system that meets the requirements of such a scenario. Or, if such a real system does not yet exist, it must be possible to emulate it. If it is impossible to achieve this last goal either, then I miss the usefulness of looking for solutions for such a scenario. That said, let's define the instance(s) of the scenario that you find most representative, and let's test BFQ on it/them. Numbers will give us the answers. For example, what about all or part of the following groups: . one cyclically doing random I/O for some second and then sequential I/O for the next seconds . one doing, say, quasi-sequential I/O in ON/OFF cycles . one starting an application cyclically . one playing back or streaming a movie For each group, we could then measure the time needed to complete each phase of I/O in each cycle, plus the responsiveness in the group starting an application, plus the frame drop in the group streaming the movie. In addition, we can measure the bandwidth/iops enjoyed by each group, plus, of course, the aggregate throughput of the whole system. In particular we could compare results with throttling, BFQ, and CFQ. Then we could write resulting numbers on the stone, and stick to them until something proves them wrong. What do you (or others) think about it? Thanks, Paolo > Thanks. > > -- > tejun > -- > To unsubscribe from this list: send the line "unsubscribe linux-block" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Vivek Goyal <vgoyal@redhat.com> |
|---|---|
| Date | 2016-10-05 15:20 +0200 |
| Message-ID | <soUh3-1sa-7@gated-at.bofh.it> |
| In reply to | #1495852 |
On Wed, Oct 05, 2016 at 02:37:00PM +0200, Paolo Valente wrote: [..] > Anyway, to avoid going on with trying speculations and arguments, let > me retry with a practical proposal. BFQ is out there, free. Let's > just test, measure and check whether we have already a solution to > the problems you/we are still trying to solve in Linux. Hi Paolo, Does BFQ implementaiton scale for fast storage devices using blk-mq interface. We will want to make sure that locking and other overhead of BFQ is very minimal so that overall throughput does not suffer. Vivek
[toc] | [prev] | [next] | [standalone]
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-05 16:10 +0200 |
| Message-ID | <soV3s-209-21@gated-at.bofh.it> |
| In reply to | #1495866 |
> Il giorno 05 ott 2016, alle ore 15:12, Vivek Goyal <vgoyal@redhat.com> ha scritto: > > On Wed, Oct 05, 2016 at 02:37:00PM +0200, Paolo Valente wrote: > > [..] >> Anyway, to avoid going on with trying speculations and arguments, let >> me retry with a practical proposal. BFQ is out there, free. Let's >> just test, measure and check whether we have already a solution to >> the problems you/we are still trying to solve in Linux. > > Hi Paolo, > > Does BFQ implementaiton scale for fast storage devices using blk-mq > interface. We will want to make sure that locking and other overhead of > BFQ is very minimal so that overall throughput does not suffer. > Of course BFQ needs to be modified to work in blk-mq. I'm rather sure its overhead will then be small enough, just because I have already collaborated to a basically equivalent port from single to multi-queue for packet scheduling (with Luigi Rizzo and others), and our prototype can make over 15 million scheduling decisions per second, and keep latency low, even with tens of concurrent clients running on a multi-core, multi-socket system. For details, here is the paper [1], plus some slides [2]. Actually, the solution in [1] is a global scheduler, which is more complex than the first blk-mq version of BFQ that I have in mind, namely, partitioned scheduling, in which there should be one independent scheduler instance per core. But this is still investigation territory. BTW, I would really appreciate help/feedback on this task [3]. Thanks, Paolo [1] http://info.iet.unipi.it/~luigi/papers/20160921-pspat.pdf [2] http://info.iet.unipi.it/~luigi/pspat/ [3] https://marc.info/?l=linux-kernel&m=147066540916339&w=2 > Vivek > -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Tejun Heo <tj@kernel.org> |
|---|---|
| Date | 2016-10-05 17:00 +0200 |
| Message-ID | <soVPQ-2hF-7@gated-at.bofh.it> |
| In reply to | #1495852 |
Hello, Paolo. On Wed, Oct 05, 2016 at 02:37:00PM +0200, Paolo Valente wrote: > In this respect, for your generic, unpredictable scenario to make > sense, there must exist at least one real system that meets the > requirements of such a scenario. Or, if such a real system does not > yet exist, it must be possible to emulate it. If it is impossible to > achieve this last goal either, then I miss the usefulness > of looking for solutions for such a scenario. > > That said, let's define the instance(s) of the scenario that you find > most representative, and let's test BFQ on it/them. Numbers will give > us the answers. For example, what about all or part of the following > groups: > . one cyclically doing random I/O for some second and then sequential I/O > for the next seconds > . one doing, say, quasi-sequential I/O in ON/OFF cycles > . one starting an application cyclically > . one playing back or streaming a movie > > For each group, we could then measure the time needed to complete each > phase of I/O in each cycle, plus the responsiveness in the group > starting an application, plus the frame drop in the group streaming > the movie. In addition, we can measure the bandwidth/iops enjoyed by > each group, plus, of course, the aggregate throughput of the whole > system. In particular we could compare results with throttling, BFQ, > and CFQ. > > Then we could write resulting numbers on the stone, and stick to them > until something proves them wrong. > > What do you (or others) think about it? That sounds great and yeah it's lame that we didn't start with that. Shaohua, would it be difficult to compare how bfq performs against blk-throttle? Thanks. -- tejun
[toc] | [prev] | [next] | [standalone]
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-05 21:50 +0200 |
| Message-ID | <sp0mu-6bE-15@gated-at.bofh.it> |
| In reply to | #1495905 |
> Il giorno 05 ott 2016, alle ore 20:30, Shaohua Li <shli@fb.com> ha scritto: > > On Wed, Oct 05, 2016 at 10:49:46AM -0400, Tejun Heo wrote: >> Hello, Paolo. >> >> On Wed, Oct 05, 2016 at 02:37:00PM +0200, Paolo Valente wrote: >>> In this respect, for your generic, unpredictable scenario to make >>> sense, there must exist at least one real system that meets the >>> requirements of such a scenario. Or, if such a real system does not >>> yet exist, it must be possible to emulate it. If it is impossible to >>> achieve this last goal either, then I miss the usefulness >>> of looking for solutions for such a scenario. >>> >>> That said, let's define the instance(s) of the scenario that you find >>> most representative, and let's test BFQ on it/them. Numbers will give >>> us the answers. For example, what about all or part of the following >>> groups: >>> . one cyclically doing random I/O for some second and then sequential I/O >>> for the next seconds >>> . one doing, say, quasi-sequential I/O in ON/OFF cycles >>> . one starting an application cyclically >>> . one playing back or streaming a movie >>> >>> For each group, we could then measure the time needed to complete each >>> phase of I/O in each cycle, plus the responsiveness in the group >>> starting an application, plus the frame drop in the group streaming >>> the movie. In addition, we can measure the bandwidth/iops enjoyed by >>> each group, plus, of course, the aggregate throughput of the whole >>> system. In particular we could compare results with throttling, BFQ, >>> and CFQ. >>> >>> Then we could write resulting numbers on the stone, and stick to them >>> until something proves them wrong. >>> >>> What do you (or others) think about it? >> >> That sounds great and yeah it's lame that we didn't start with that. >> Shaohua, would it be difficult to compare how bfq performs against >> blk-throttle? > > I had a test of BFQ. Thank you very much for testing BFQ! > I'm using BFQ found at > http://algogroup.unimore.it/people/paolo/disk_sched/sources.php. version is > 4.7.0-v8r3. That's the latest stable version. The development version [1] already contains further improvements for fairness, latency and throughput. It is however still a release candidate. [1] https://github.com/linusw/linux-bfq/tree/bfq-v8 > It's a LSI SSD, queue depth 32. I use default setting. fio script > is: > > [global] > ioengine=libaio > direct=1 > readwrite=randread > bs=4k > runtime=60 > time_based=1 > file_service_type=random:36 > overwrite=1 > thread=0 > group_reporting=1 > filename=/dev/sdb > iodepth=1 > numjobs=8 > > [groupA] > prio=2 > > [groupB] > new_group > prio=6 > > I'll change iodepth, numjobs and prio in different tests. result unit is MB/s. > > iodepth=1 numjobs=1 prio 4:4 > CFQ: 28:28 BFQ: 21:21 deadline: 29:29 > > iodepth=8 numjobs=1 prio 4:4 > CFQ: 162:162 BFQ: 102:98 deadline: 205:205 > > iodepth=1 numjobs=8 prio 4:4 > CFQ: 157:157 BFQ: 81:92 deadline: 196:197 > > iodepth=1 numjobs=1 prio 2:6 > CFQ: 26.7:27.6 BFQ: 20:6 deadline: 29:29 > > iodepth=8 numjobs=1 prio 2:6 > CFQ: 166:174 BFQ: 139:72 deadline: 202:202 > > iodepth=1 numjobs=8 prio 2:6 > CFQ: 148:150 BFQ: 90:77 deadline: 198:197 > > CFQ isn't fair at all. BFQ is very good in this side, but has poor throughput > even prio is the default value. > Throughput is lower with BFQ for two reasons. First, you certainly left the low_latency in its default state, i.e., on. As explained, e.g., here [2], low_latency mode is totally geared towards maximum responsiveness and minimum latency for soft real-time applications (e.g., video players). To achieve this goal, BFQ is willing to perform more idling, when necessary. This lowers throughput (I'll get back on this at the end of the discussion of the second reason). The second, most important reason, is that a minimum of idling is the *only* way to achieve differentiated bandwidth distribution, as you requested by setting different ioprios. I stress that this constraint is not a technological accident, but a intrinsic, logical necessity. The proof is simple, and if the following explanation is too boring or confusing, I can show it to you with any trace of sync I/O. First, to provide differentiated service, you need per-process scheduling, i.e., schedulers in which there is a separate queue associated with each process. Now, let A be the process with higher weight (ioprio), and B the process with lower weight. Both processes are sync, thus, by definition, they issue requests as follows: a few requests (probably two, or a little bit more with larger iodepth), then a little break to wait for request completion, then the next small batch and so on. For each process, the queue associated with the process (in the scheduler) is necessarily empty on the break. As a consequence, if there is no idling, then every time A reaches its break, the scheduler has only the option to switch to B (which is extremely likely to have pending requests). The service pattern of the processes then unavoidably becomes: A B A B A B ... where each letter represents a full small batch served for the process. That is, 50% of the bw for each process, and complete loss of control on the desired bandwidth distribution. So, to sum up, the reason why BFQ achieves a lower total bw is that it behaves in the only correct way to respect weights with sync I/O, i.e., it performs a little idling. If low_latency is on, then BFQ increases idling further, and this may be have caused further bw loss in your test (but this varies greatly with devices, so you can discover it only by trying). The bottom line is that if you do want to achieve differentiation with sync I/O, you have to pay a price in terms of bw, because of idling. Actually, the recent preemption mechanism that I have introduced in BFQ is proving so effective in preserving differentiation, that I'm tempted to try some almost idleness solution. A little of accuracy should however be sacrificed. Anyway, this is still work in progress. Thank you very much, Paolo [2] http://algogroup.unimore.it/people/paolo/disk_sched/description.php > Thanks, > Shaohua > -- > To unsubscribe from this list: send the line "unsubscribe linux-block" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-05 22:10 +0200 |
| Message-ID | <sp0FP-6Ex-11@gated-at.bofh.it> |
| In reply to | #1496019 |
> Il giorno 05 ott 2016, alle ore 21:47, Paolo Valente <paolo.valente@unimore.it> ha scritto: > >> >> Il giorno 05 ott 2016, alle ore 20:30, Shaohua Li <shli@fb.com> ha scritto: >> >> On Wed, Oct 05, 2016 at 10:49:46AM -0400, Tejun Heo wrote: >>> Hello, Paolo. >>> >>> On Wed, Oct 05, 2016 at 02:37:00PM +0200, Paolo Valente wrote: >>>> In this respect, for your generic, unpredictable scenario to make >>>> sense, there must exist at least one real system that meets the >>>> requirements of such a scenario. Or, if such a real system does not >>>> yet exist, it must be possible to emulate it. If it is impossible to >>>> achieve this last goal either, then I miss the usefulness >>>> of looking for solutions for such a scenario. >>>> >>>> That said, let's define the instance(s) of the scenario that you find >>>> most representative, and let's test BFQ on it/them. Numbers will give >>>> us the answers. For example, what about all or part of the following >>>> groups: >>>> . one cyclically doing random I/O for some second and then sequential I/O >>>> for the next seconds >>>> . one doing, say, quasi-sequential I/O in ON/OFF cycles >>>> . one starting an application cyclically >>>> . one playing back or streaming a movie >>>> >>>> For each group, we could then measure the time needed to complete each >>>> phase of I/O in each cycle, plus the responsiveness in the group >>>> starting an application, plus the frame drop in the group streaming >>>> the movie. In addition, we can measure the bandwidth/iops enjoyed by >>>> each group, plus, of course, the aggregate throughput of the whole >>>> system. In particular we could compare results with throttling, BFQ, >>>> and CFQ. >>>> >>>> Then we could write resulting numbers on the stone, and stick to them >>>> until something proves them wrong. >>>> >>>> What do you (or others) think about it? >>> >>> That sounds great and yeah it's lame that we didn't start with that. >>> Shaohua, would it be difficult to compare how bfq performs against >>> blk-throttle? >> >> I had a test of BFQ. > > Thank you very much for testing BFQ! > >> I'm using BFQ found at >> http://algogroup.unimore.it/people/paolo/disk_sched/sources.php. version is >> 4.7.0-v8r3. > > That's the latest stable version. The development version [1] already > contains further improvements for fairness, latency and throughput. > It is however still a release candidate. > > [1] https://github.com/linusw/linux-bfq/tree/bfq-v8 > >> It's a LSI SSD, queue depth 32. I use default setting. fio script >> is: >> >> [global] >> ioengine=libaio >> direct=1 >> readwrite=randread >> bs=4k >> runtime=60 >> time_based=1 >> file_service_type=random:36 >> overwrite=1 >> thread=0 >> group_reporting=1 >> filename=/dev/sdb >> iodepth=1 >> numjobs=8 >> >> [groupA] >> prio=2 >> >> [groupB] >> new_group >> prio=6 >> >> I'll change iodepth, numjobs and prio in different tests. result unit is MB/s. >> >> iodepth=1 numjobs=1 prio 4:4 >> CFQ: 28:28 BFQ: 21:21 deadline: 29:29 >> >> iodepth=8 numjobs=1 prio 4:4 >> CFQ: 162:162 BFQ: 102:98 deadline: 205:205 >> >> iodepth=1 numjobs=8 prio 4:4 >> CFQ: 157:157 BFQ: 81:92 deadline: 196:197 >> >> iodepth=1 numjobs=1 prio 2:6 >> CFQ: 26.7:27.6 BFQ: 20:6 deadline: 29:29 >> >> iodepth=8 numjobs=1 prio 2:6 >> CFQ: 166:174 BFQ: 139:72 deadline: 202:202 >> >> iodepth=1 numjobs=8 prio 2:6 >> CFQ: 148:150 BFQ: 90:77 deadline: 198:197 >> >> CFQ isn't fair at all. BFQ is very good in this side, but has poor throughput >> even prio is the default value. >> > > Throughput is lower with BFQ for two reasons. > > First, you certainly left the low_latency in its default state, i.e., > on. As explained, e.g., here [2], low_latency mode is totally geared > towards maximum responsiveness and minimum latency for soft real-time > applications (e.g., video players). To achieve this goal, BFQ is > willing to perform more idling, when necessary. This lowers > throughput (I'll get back on this at the end of the discussion of the > second reason). > > The second, most important reason, is that a minimum of idling is the > *only* way to achieve differentiated bandwidth distribution, as you > requested by setting different ioprios. I stress that this constraint > is not a technological accident, but a intrinsic, logical necessity. > The proof is simple, and if the following explanation is too boring or > confusing, I can show it to you with any trace of sync I/O. > > First, to provide differentiated service, you need per-process > scheduling, i.e., schedulers in which there is a separate queue > associated with each process. Now, let A be the process with higher > weight (ioprio), and B the process with lower weight. Both processes > are sync, thus, by definition, they issue requests as follows: a few > requests (probably two, or a little bit more with larger iodepth), > then a little break to wait for request completion, then the next > small batch and so on. For each process, the queue associated with > the process (in the scheduler) is necessarily empty on the break. As > a consequence, if there is no idling, then every time A reaches its > break, the scheduler has only the option to switch to B (which is > extremely likely to have pending requests). > > The service pattern of the processes then unavoidably becomes: > > A B A B A B ... > > where each letter represents a full small batch served for the > process. That is, 50% of the bw for each process, and complete loss > of control on the desired bandwidth distribution. > > So, to sum up, the reason why BFQ achieves a lower total bw is that it > behaves in the only correct way to respect weights with sync I/O, > i.e., it performs a little idling. If low_latency is on, then BFQ > increases idling further, and this may be have caused further bw loss > in your test (but this varies greatly with devices, so you can > discover it only by trying). > > The bottom line is that if you do want to achieve differentiation with > sync I/O, you have to pay a price in terms of bw, because of idling. > Actually, the recent preemption mechanism that I have introduced in > BFQ is proving so effective in preserving differentiation, that I'm > tempted to try some almost idleness solution. A little of accuracy > should however be sacrificed. Anyway, this is still work in progress. > Just for completeness, if the weights of the processes or groups are equal, then BFQ preserves service guarantees while achieving the same bw as deadline. And equal weights is the most common case according to my limited experience. Thanks, Paolo > Thank you very much, > Paolo > > [2] http://algogroup.unimore.it/people/paolo/disk_sched/description.php > >> Thanks, >> Shaohua >> -- >> To unsubscribe from this list: send the line "unsubscribe linux-block" in >> the body of a message to majordomo@vger.kernel.org >> More majordomo info at http://vger.kernel.org/majordomo-info.html > > > -- > Paolo Valente > Algogroup > Dipartimento di Scienze Fisiche, Informatiche e Matematiche > Via Campi 213/B > 41125 Modena - Italy > http://algogroup.unimore.it/people/paolo/ > > > > > > -- > To unsubscribe from this list: send the line "unsubscribe linux-block" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-06 10:00 +0200 |
| Message-ID | <spbKV-5wA-1@gated-at.bofh.it> |
| In reply to | #1496019 |
> Il giorno 05 ott 2016, alle ore 22:46, Shaohua Li <shli@fb.com> ha scritto: > > On Wed, Oct 05, 2016 at 09:47:19PM +0200, Paolo Valente wrote: >> >>> Il giorno 05 ott 2016, alle ore 20:30, Shaohua Li <shli@fb.com> ha scritto: >>> >>> On Wed, Oct 05, 2016 at 10:49:46AM -0400, Tejun Heo wrote: >>>> Hello, Paolo. >>>> >>>> On Wed, Oct 05, 2016 at 02:37:00PM +0200, Paolo Valente wrote: >>>>> In this respect, for your generic, unpredictable scenario to make >>>>> sense, there must exist at least one real system that meets the >>>>> requirements of such a scenario. Or, if such a real system does not >>>>> yet exist, it must be possible to emulate it. If it is impossible to >>>>> achieve this last goal either, then I miss the usefulness >>>>> of looking for solutions for such a scenario. >>>>> >>>>> That said, let's define the instance(s) of the scenario that you find >>>>> most representative, and let's test BFQ on it/them. Numbers will give >>>>> us the answers. For example, what about all or part of the following >>>>> groups: >>>>> . one cyclically doing random I/O for some second and then sequential I/O >>>>> for the next seconds >>>>> . one doing, say, quasi-sequential I/O in ON/OFF cycles >>>>> . one starting an application cyclically >>>>> . one playing back or streaming a movie >>>>> >>>>> For each group, we could then measure the time needed to complete each >>>>> phase of I/O in each cycle, plus the responsiveness in the group >>>>> starting an application, plus the frame drop in the group streaming >>>>> the movie. In addition, we can measure the bandwidth/iops enjoyed by >>>>> each group, plus, of course, the aggregate throughput of the whole >>>>> system. In particular we could compare results with throttling, BFQ, >>>>> and CFQ. >>>>> >>>>> Then we could write resulting numbers on the stone, and stick to them >>>>> until something proves them wrong. >>>>> >>>>> What do you (or others) think about it? >>>> >>>> That sounds great and yeah it's lame that we didn't start with that. >>>> Shaohua, would it be difficult to compare how bfq performs against >>>> blk-throttle? >>> >>> I had a test of BFQ. >> >> Thank you very much for testing BFQ! >> >>> I'm using BFQ found at >>> https://urldefense.proofpoint.com/v2/url?u=http-3A__algogroup.unimore.it_people_paolo_disk-5Fsched_sources.php&d=DQIFAg&c=5VD0RTtNlTh3ycd41b3MUw&r=i6WobKxbeG3slzHSIOxTVtYIJw7qjCE6S0spDTKL-J4&m=2pG8KEx5tRymExa_K0ddKH_YvhH3qvJxELBd1_lw0-w&s=FZKEAOu2sw95y9jZio2k012cQWoLzlBWDl0NiGPVW78&e= . version is >>> 4.7.0-v8r3. >> >> That's the latest stable version. The development version [1] already >> contains further improvements for fairness, latency and throughput. >> It is however still a release candidate. >> >> [1] https://github.com/linusw/linux-bfq/tree/bfq-v8 >> >>> It's a LSI SSD, queue depth 32. I use default setting. fio script >>> is: >>> >>> [global] >>> ioengine=libaio >>> direct=1 >>> readwrite=randread >>> bs=4k >>> runtime=60 >>> time_based=1 >>> file_service_type=random:36 >>> overwrite=1 >>> thread=0 >>> group_reporting=1 >>> filename=/dev/sdb >>> iodepth=1 >>> numjobs=8 >>> >>> [groupA] >>> prio=2 >>> >>> [groupB] >>> new_group >>> prio=6 >>> >>> I'll change iodepth, numjobs and prio in different tests. result unit is MB/s. >>> >>> iodepth=1 numjobs=1 prio 4:4 >>> CFQ: 28:28 BFQ: 21:21 deadline: 29:29 >>> >>> iodepth=8 numjobs=1 prio 4:4 >>> CFQ: 162:162 BFQ: 102:98 deadline: 205:205 >>> >>> iodepth=1 numjobs=8 prio 4:4 >>> CFQ: 157:157 BFQ: 81:92 deadline: 196:197 >>> >>> iodepth=1 numjobs=1 prio 2:6 >>> CFQ: 26.7:27.6 BFQ: 20:6 deadline: 29:29 >>> >>> iodepth=8 numjobs=1 prio 2:6 >>> CFQ: 166:174 BFQ: 139:72 deadline: 202:202 >>> >>> iodepth=1 numjobs=8 prio 2:6 >>> CFQ: 148:150 BFQ: 90:77 deadline: 198:197 >>> >>> CFQ isn't fair at all. BFQ is very good in this side, but has poor throughput >>> even prio is the default value. >>> >> >> Throughput is lower with BFQ for two reasons. >> >> First, you certainly left the low_latency in its default state, i.e., >> on. As explained, e.g., here [2], low_latency mode is totally geared >> towards maximum responsiveness and minimum latency for soft real-time >> applications (e.g., video players). To achieve this goal, BFQ is >> willing to perform more idling, when necessary. This lowers >> throughput (I'll get back on this at the end of the discussion of the >> second reason). > > changing low_latency to 0 seems not change anything, at least for the test: > iodepth=1 numjobs=1 prio 2:6 A bs 4k:64k > >> The second, most important reason, is that a minimum of idling is the >> *only* way to achieve differentiated bandwidth distribution, as you >> requested by setting different ioprios. I stress that this constraint >> is not a technological accident, but a intrinsic, logical necessity. >> The proof is simple, and if the following explanation is too boring or >> confusing, I can show it to you with any trace of sync I/O. >> >> First, to provide differentiated service, you need per-process >> scheduling, i.e., schedulers in which there is a separate queue >> associated with each process. Now, let A be the process with higher >> weight (ioprio), and B the process with lower weight. Both processes >> are sync, thus, by definition, they issue requests as follows: a few >> requests (probably two, or a little bit more with larger iodepth), >> then a little break to wait for request completion, then the next >> small batch and so on. For each process, the queue associated with >> the process (in the scheduler) is necessarily empty on the break. As >> a consequence, if there is no idling, then every time A reaches its >> break, the scheduler has only the option to switch to B (which is >> extremely likely to have pending requests). >> >> The service pattern of the processes then unavoidably becomes: >> >> A B A B A B ... >> >> where each letter represents a full small batch served for the >> process. That is, 50% of the bw for each process, and complete loss >> of control on the desired bandwidth distribution. >> >> So, to sum up, the reason why BFQ achieves a lower total bw is that it >> behaves in the only correct way to respect weights with sync I/O, >> i.e., it performs a little idling. If low_latency is on, then BFQ >> increases idling further, and this may be have caused further bw loss >> in your test (but this varies greatly with devices, so you can >> discover it only by trying). >> >> The bottom line is that if you do want to achieve differentiation with >> sync I/O, you have to pay a price in terms of bw, because of idling. >> Actually, the recent preemption mechanism that I have introduced in >> BFQ is proving so effective in preserving differentiation, that I'm >> tempted to try some almost idleness solution. A little of accuracy >> should however be sacrificed. Anyway, this is still work in progress. > > Yep, I fully understand why idle is required here. As long as workload io depth > is lower than queue io depth, idle is the only way to maintain fairness. This > is the core of CFQ, I bet the same for BFQ. Unfortunately idle disk harms > throughput too much especially for high end SSD. > Then I'm afraid I have to give you very bad news: bw limiting causes the same throughput loss. You can see it from your very tests. Here is one of your results with BFQ (one that is likely to have been affected less by the fact that you left low_latency on, or by further issues that I may have not yet addressed thoroughly): iodepth=8 numjobs=1 prio 2:6 CFQ: 166:174 BFQ: 139:72 deadline: 202:202 Here is, instead, your test with bw limitation: iodepth=8 numjobs=1 prio 2:6, group A has 50M/s limit CFQ:51:207 BFQ: 51:45 deadline: 51:216 From the first test, you see that the total bw achievable by the device is at least 404MB/s. But in the second test you get at most 267MB/s, with deadline. In this respect, the total bw achieved by BFQ in the first test is 211MB/s. So, both throttling and proportional share need to waste bw, BFQ looses about 13% more of the total bw. In return, it gives you incomparably better bw and latency guarantees, while allowing you to configure your system with zero or minimal effort. In contrast, using bw limits to properly configure a common system like, e.g., a large file server, may become a nightmare for a sysadmin. For example, if, in the simplest case, she/he configures limits for the worst-case, then per-client limits will have to be extremely low. But the system is large and dynamic, so the actual number of clients and at the actual bw consumed by each client will vary without a break, even in the short-medium term. The bw redistribution heuristic do not give any provable guarantees on accuracy of bw redistribution. The result will likely be highly varying client bandwidths, with unlucky clients unjustly limited to low limits, and then experiencing high latencies. The latter will be further emphasized by the intrinsically bursty nature of throttling. In addition, the scenario in your tests is the worst case for a proportional share solution: in a generic system, such as the the file server above, part of the workload is likely to be sequential or quasi-sequential (at least in medium-length time intervals), and this is enough to get very close to peak bw with a proportional-share scheduler. No configuration needed. With bw throttling, you must do the math very well to get peak bw all the time in a dynamic system. To sum up, I do think that bw throttling is the best possible solution (and in particular your solution is very good) till blk-mq lacks an accurate scheduler. But, where you do have such a scheduler, it makes very little sense to give users a very bad service, for fear of wasting at most 13% of the bw in a single worst case. Thanks, Paolo > Thanks, > Shaohua > -- > To unsubscribe from this list: send the line "unsubscribe linux-block" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-06 15:20 +0200 |
| Message-ID | <spgKC-A6-27@gated-at.bofh.it> |
| In reply to | #1496198 |
> Il giorno 06 ott 2016, alle ore 09:58, Paolo Valente <paolo.valente@unimore.it> ha scritto: > >> >> Il giorno 05 ott 2016, alle ore 22:46, Shaohua Li <shli@fb.com> ha scritto: >> >> On Wed, Oct 05, 2016 at 09:47:19PM +0200, Paolo Valente wrote: >>> >>>> Il giorno 05 ott 2016, alle ore 20:30, Shaohua Li <shli@fb.com> ha scritto: >>>> >>>> On Wed, Oct 05, 2016 at 10:49:46AM -0400, Tejun Heo wrote: >>>>> Hello, Paolo. >>>>> >>>>> On Wed, Oct 05, 2016 at 02:37:00PM +0200, Paolo Valente wrote: >>>>>> In this respect, for your generic, unpredictable scenario to make >>>>>> sense, there must exist at least one real system that meets the >>>>>> requirements of such a scenario. Or, if such a real system does not >>>>>> yet exist, it must be possible to emulate it. If it is impossible to >>>>>> achieve this last goal either, then I miss the usefulness >>>>>> of looking for solutions for such a scenario. >>>>>> >>>>>> That said, let's define the instance(s) of the scenario that you find >>>>>> most representative, and let's test BFQ on it/them. Numbers will give >>>>>> us the answers. For example, what about all or part of the following >>>>>> groups: >>>>>> . one cyclically doing random I/O for some second and then sequential I/O >>>>>> for the next seconds >>>>>> . one doing, say, quasi-sequential I/O in ON/OFF cycles >>>>>> . one starting an application cyclically >>>>>> . one playing back or streaming a movie >>>>>> >>>>>> For each group, we could then measure the time needed to complete each >>>>>> phase of I/O in each cycle, plus the responsiveness in the group >>>>>> starting an application, plus the frame drop in the group streaming >>>>>> the movie. In addition, we can measure the bandwidth/iops enjoyed by >>>>>> each group, plus, of course, the aggregate throughput of the whole >>>>>> system. In particular we could compare results with throttling, BFQ, >>>>>> and CFQ. >>>>>> >>>>>> Then we could write resulting numbers on the stone, and stick to them >>>>>> until something proves them wrong. >>>>>> >>>>>> What do you (or others) think about it? >>>>> >>>>> That sounds great and yeah it's lame that we didn't start with that. >>>>> Shaohua, would it be difficult to compare how bfq performs against >>>>> blk-throttle? >>>> >>>> I had a test of BFQ. >>> >>> Thank you very much for testing BFQ! >>> >>>> I'm using BFQ found at >>>> https://urldefense.proofpoint.com/v2/url?u=http-3A__algogroup.unimore.it_people_paolo_disk-5Fsched_sources.php&d=DQIFAg&c=5VD0RTtNlTh3ycd41b3MUw&r=i6WobKxbeG3slzHSIOxTVtYIJw7qjCE6S0spDTKL-J4&m=2pG8KEx5tRymExa_K0ddKH_YvhH3qvJxELBd1_lw0-w&s=FZKEAOu2sw95y9jZio2k012cQWoLzlBWDl0NiGPVW78&e= . version is >>>> 4.7.0-v8r3. >>> >>> That's the latest stable version. The development version [1] already >>> contains further improvements for fairness, latency and throughput. >>> It is however still a release candidate. >>> >>> [1] https://github.com/linusw/linux-bfq/tree/bfq-v8 >>> >>>> It's a LSI SSD, queue depth 32. I use default setting. fio script >>>> is: >>>> >>>> [global] >>>> ioengine=libaio >>>> direct=1 >>>> readwrite=randread >>>> bs=4k >>>> runtime=60 >>>> time_based=1 >>>> file_service_type=random:36 >>>> overwrite=1 >>>> thread=0 >>>> group_reporting=1 >>>> filename=/dev/sdb >>>> iodepth=1 >>>> numjobs=8 >>>> >>>> [groupA] >>>> prio=2 >>>> >>>> [groupB] >>>> new_group >>>> prio=6 >>>> >>>> I'll change iodepth, numjobs and prio in different tests. result unit is MB/s. >>>> >>>> iodepth=1 numjobs=1 prio 4:4 >>>> CFQ: 28:28 BFQ: 21:21 deadline: 29:29 >>>> >>>> iodepth=8 numjobs=1 prio 4:4 >>>> CFQ: 162:162 BFQ: 102:98 deadline: 205:205 >>>> >>>> iodepth=1 numjobs=8 prio 4:4 >>>> CFQ: 157:157 BFQ: 81:92 deadline: 196:197 >>>> >>>> iodepth=1 numjobs=1 prio 2:6 >>>> CFQ: 26.7:27.6 BFQ: 20:6 deadline: 29:29 >>>> >>>> iodepth=8 numjobs=1 prio 2:6 >>>> CFQ: 166:174 BFQ: 139:72 deadline: 202:202 >>>> >>>> iodepth=1 numjobs=8 prio 2:6 >>>> CFQ: 148:150 BFQ: 90:77 deadline: 198:197 >>>> >>>> CFQ isn't fair at all. BFQ is very good in this side, but has poor throughput >>>> even prio is the default value. >>>> >>> >>> Throughput is lower with BFQ for two reasons. >>> >>> First, you certainly left the low_latency in its default state, i.e., >>> on. As explained, e.g., here [2], low_latency mode is totally geared >>> towards maximum responsiveness and minimum latency for soft real-time >>> applications (e.g., video players). To achieve this goal, BFQ is >>> willing to perform more idling, when necessary. This lowers >>> throughput (I'll get back on this at the end of the discussion of the >>> second reason). >> >> changing low_latency to 0 seems not change anything, at least for the test: >> iodepth=1 numjobs=1 prio 2:6 A bs 4k:64k >> >>> The second, most important reason, is that a minimum of idling is the >>> *only* way to achieve differentiated bandwidth distribution, as you >>> requested by setting different ioprios. I stress that this constraint >>> is not a technological accident, but a intrinsic, logical necessity. >>> The proof is simple, and if the following explanation is too boring or >>> confusing, I can show it to you with any trace of sync I/O. >>> >>> First, to provide differentiated service, you need per-process >>> scheduling, i.e., schedulers in which there is a separate queue >>> associated with each process. Now, let A be the process with higher >>> weight (ioprio), and B the process with lower weight. Both processes >>> are sync, thus, by definition, they issue requests as follows: a few >>> requests (probably two, or a little bit more with larger iodepth), >>> then a little break to wait for request completion, then the next >>> small batch and so on. For each process, the queue associated with >>> the process (in the scheduler) is necessarily empty on the break. As >>> a consequence, if there is no idling, then every time A reaches its >>> break, the scheduler has only the option to switch to B (which is >>> extremely likely to have pending requests). >>> >>> The service pattern of the processes then unavoidably becomes: >>> >>> A B A B A B ... >>> >>> where each letter represents a full small batch served for the >>> process. That is, 50% of the bw for each process, and complete loss >>> of control on the desired bandwidth distribution. >>> >>> So, to sum up, the reason why BFQ achieves a lower total bw is that it >>> behaves in the only correct way to respect weights with sync I/O, >>> i.e., it performs a little idling. If low_latency is on, then BFQ >>> increases idling further, and this may be have caused further bw loss >>> in your test (but this varies greatly with devices, so you can >>> discover it only by trying). >>> >>> The bottom line is that if you do want to achieve differentiation with >>> sync I/O, you have to pay a price in terms of bw, because of idling. >>> Actually, the recent preemption mechanism that I have introduced in >>> BFQ is proving so effective in preserving differentiation, that I'm >>> tempted to try some almost idleness solution. A little of accuracy >>> should however be sacrificed. Anyway, this is still work in progress. >> >> Yep, I fully understand why idle is required here. As long as workload io depth >> is lower than queue io depth, idle is the only way to maintain fairness. This >> is the core of CFQ, I bet the same for BFQ. Unfortunately idle disk harms >> throughput too much especially for high end SSD. >> > > Then I'm afraid I have to give you very bad news: bw limiting causes > the same throughput loss. You can see it from your very tests. Here > is one of your results with BFQ (one that is likely to have been > affected less by the fact that you left low_latency on, or by further > issues that I may have not yet addressed thoroughly): > > iodepth=8 numjobs=1 prio 2:6 > CFQ: 166:174 BFQ: 139:72 deadline: 202:202 > > Here is, instead, your test with bw limitation: > iodepth=8 numjobs=1 prio 2:6, group A has 50M/s limit > CFQ:51:207 BFQ: 51:45 deadline: 51:216 > > From the first test, you see that the total bw achievable by the > device is at least 404MB/s. But in the second test you get at most > 267MB/s, with deadline. In this respect, the total bw achieved by BFQ > in the first test is 211MB/s. > > So, both throttling and proportional share need to waste bw, BFQ > looses about 13% more of the total bw. In return, it gives you > incomparably better bw and latency guarantees, while allowing you to > configure your system with zero or minimal effort. In contrast, using > bw limits to properly configure a common system like, e.g., a large > file server, may become a nightmare for a sysadmin. For example, if, > in the simplest case, she/he configures limits for the worst-case, > then per-client limits will have to be extremely low. But the system > is large and dynamic, so the actual number of clients and at the > actual bw consumed by each client will vary without a break, even in > the short-medium term. The bw redistribution heuristic do not give > any provable guarantees on accuracy of bw redistribution. The result > will likely be highly varying client bandwidths, with unlucky clients > unjustly limited to low limits, and then experiencing high latencies. > The latter will be further emphasized by the intrinsically bursty > nature of throttling. > > In addition, the scenario in your tests is the worst case for a > proportional share solution: in a generic system, such as the the file > server above, part of the workload is likely to be sequential or > quasi-sequential (at least in medium-length time intervals), and this > is enough to get very close to peak bw with a proportional-share > scheduler. No configuration needed. With bw throttling, you must do > the math very well to get peak bw all the time in a dynamic system. > > To sum up, I do think that bw throttling is the best possible solution > (and in particular your solution is very good) till blk-mq lacks an > accurate scheduler. But, where you do have such a scheduler, it makes > very little sense to give users a very bad service, for fear of > wasting at most 13% of the bw in a single worst case. > Shaohua, I have just realized that I have unconsciously defended a wrong argument. Although all the facts that I have reported are evidently true, I have argued as if the question was: "do we need to throw away throttling because there is proportional, or do we need to throw away proportional share because there is throttling?". This question is simply wrong, as I think consciously (sorry for my dissociated behavior :) ). The best goal to achieve is to have both a good throttling mechanism, and a good proportional share scheduler. This goal would be valid if even if there was just one important scenario for each of the two approaches. The vulnus here is that you guys are constantly, and rightly, working on solutions to achieve and consolidate reasonable QoS guarantees, but an apparently very good proportional-share scheduler has been kept off for years. If you (or others) have good arguments to support this state of affairs, then this would probably be an important point to discuss. Thanks, Paolo > Thanks, > Paolo > > >> Thanks, >> Shaohua >> -- >> To unsubscribe from this list: send the line "unsubscribe linux-block" in >> the body of a message to majordomo@vger.kernel.org >> More majordomo info at http://vger.kernel.org/majordomo-info.html > > > -- > Paolo Valente > Algogroup > Dipartimento di Scienze Fisiche, Informatiche e Matematiche > Via Campi 213/B > 41125 Modena - Italy > http://algogroup.unimore.it/people/paolo/ > > > > > > -- > To unsubscribe from this list: send the line "unsubscribe linux-block" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Vivek Goyal <vgoyal@redhat.com> |
|---|---|
| Date | 2016-10-06 19:50 +0200 |
| Message-ID | <spkXT-3Nc-9@gated-at.bofh.it> |
| In reply to | #1496649 |
On Thu, Oct 06, 2016 at 03:15:50PM +0200, Paolo Valente wrote: [..] > Shaohua, I have just realized that I have unconsciously defended a > wrong argument. Although all the facts that I have reported are > evidently true, I have argued as if the question was: "do we need to > throw away throttling because there is proportional, or do we need to > throw away proportional share because there is throttling?". This > question is simply wrong, as I think consciously (sorry for my > dissociated behavior :) ). I was wondering about the same. We need both and both should be able to work with fast devices of today using blk-mq interfaces without much overhead. > > The best goal to achieve is to have both a good throttling mechanism, > and a good proportional share scheduler. This goal would be valid if > even if there was just one important scenario for each of the two > approaches. The vulnus here is that you guys are constantly, and > rightly, working on solutions to achieve and consolidate reasonable > QoS guarantees, but an apparently very good proportional-share > scheduler has been kept off for years. If you (or others) have good > arguments to support this state of affairs, then this would probably > be an important point to discuss. Paolo, CFQ is legacy now and if we can come up with a proportional IO mechanism which works reasonably well with fast devices using blk-mq interfaces, that will be much more interesting. Vivek
[toc] | [prev] | [next] | [standalone]
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-06 20:10 +0200 |
| Message-ID | <splhf-4du-1@gated-at.bofh.it> |
| In reply to | #1496784 |
> Il giorno 06 ott 2016, alle ore 19:49, Vivek Goyal <vgoyal@redhat.com> ha scritto: > > On Thu, Oct 06, 2016 at 03:15:50PM +0200, Paolo Valente wrote: > > [..] >> Shaohua, I have just realized that I have unconsciously defended a >> wrong argument. Although all the facts that I have reported are >> evidently true, I have argued as if the question was: "do we need to >> throw away throttling because there is proportional, or do we need to >> throw away proportional share because there is throttling?". This >> question is simply wrong, as I think consciously (sorry for my >> dissociated behavior :) ). > > I was wondering about the same. We need both and both should be able > to work with fast devices of today using blk-mq interfaces without > much overhead. > >> >> The best goal to achieve is to have both a good throttling mechanism, >> and a good proportional share scheduler. This goal would be valid if >> even if there was just one important scenario for each of the two >> approaches. The vulnus here is that you guys are constantly, and >> rightly, working on solutions to achieve and consolidate reasonable >> QoS guarantees, but an apparently very good proportional-share >> scheduler has been kept off for years. If you (or others) have good >> arguments to support this state of affairs, then this would probably >> be an important point to discuss. > > Paolo, CFQ is legacy now and if we can come up with a proportional > IO mechanism which works reasonably well with fast devices using > blk-mq interfaces, that will be much more interesting. > That's absolutely true. But, why do we pretend not to know that, for (at least) hundreds of thousands of users Linux will go on giving bad responsiveness, starvation, high latency and unfairness, until blk will not be used any more (assuming that these problems will somehow disappear will blk-mq). Many of these users are fully aware of these Linux long-standing problems. We could solve these problems by just adding a scheduler that has already been adopted, and thus extensively tested, by thousands of users. And more and more people are aware of this fact too. Are we doing the right thing? Thanks, Paolo > Vivek -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Vivek Goyal <vgoyal@redhat.com> |
|---|---|
| Date | 2016-10-06 20:40 +0200 |
| Message-ID | <splKh-4z5-1@gated-at.bofh.it> |
| In reply to | #1496790 |
On Thu, Oct 06, 2016 at 08:01:42PM +0200, Paolo Valente wrote: > > > Il giorno 06 ott 2016, alle ore 19:49, Vivek Goyal <vgoyal@redhat.com> ha scritto: > > > > On Thu, Oct 06, 2016 at 03:15:50PM +0200, Paolo Valente wrote: > > > > [..] > >> Shaohua, I have just realized that I have unconsciously defended a > >> wrong argument. Although all the facts that I have reported are > >> evidently true, I have argued as if the question was: "do we need to > >> throw away throttling because there is proportional, or do we need to > >> throw away proportional share because there is throttling?". This > >> question is simply wrong, as I think consciously (sorry for my > >> dissociated behavior :) ). > > > > I was wondering about the same. We need both and both should be able > > to work with fast devices of today using blk-mq interfaces without > > much overhead. > > > >> > >> The best goal to achieve is to have both a good throttling mechanism, > >> and a good proportional share scheduler. This goal would be valid if > >> even if there was just one important scenario for each of the two > >> approaches. The vulnus here is that you guys are constantly, and > >> rightly, working on solutions to achieve and consolidate reasonable > >> QoS guarantees, but an apparently very good proportional-share > >> scheduler has been kept off for years. If you (or others) have good > >> arguments to support this state of affairs, then this would probably > >> be an important point to discuss. > > > > Paolo, CFQ is legacy now and if we can come up with a proportional > > IO mechanism which works reasonably well with fast devices using > > blk-mq interfaces, that will be much more interesting. > > > > That's absolutely true. But, why do we pretend not to know that, for > (at least) hundreds of thousands of users Linux will go on giving bad > responsiveness, starvation, high latency and unfairness, until blk > will not be used any more (assuming that these problems will somehow > disappear will blk-mq). Many of these users are fully aware of these > Linux long-standing problems. We could solve these problems by just > adding a scheduler that has already been adopted, and thus extensively > tested, by thousands of users. And more and more people are aware of > this fact too. Are we doing the right thing? Hi Paolo, People have been using CFQ for many years. I am not sure if benefits offered by BFQ over CFQ are significant enough to justify taking a completely new code and get rid of CFQ. Or are the benfits significant enough that one feels like putting time and effort into this and take chances wiht new code. At this point of time replacing CFQ with something better is not a priority for me. But if something better and stable goes upstream, I will gladly use it. Vivek
[toc] | [prev] | [next] | [standalone]
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-06 23:00 +0200 |
| Message-ID | <spnVR-63w-9@gated-at.bofh.it> |
| In reply to | #1496800 |
> Il giorno 06 ott 2016, alle ore 20:32, Vivek Goyal <vgoyal@redhat.com> ha scritto: > > On Thu, Oct 06, 2016 at 08:01:42PM +0200, Paolo Valente wrote: >> >>> Il giorno 06 ott 2016, alle ore 19:49, Vivek Goyal <vgoyal@redhat.com> ha scritto: >>> >>> On Thu, Oct 06, 2016 at 03:15:50PM +0200, Paolo Valente wrote: >>> >>> [..] >>>> Shaohua, I have just realized that I have unconsciously defended a >>>> wrong argument. Although all the facts that I have reported are >>>> evidently true, I have argued as if the question was: "do we need to >>>> throw away throttling because there is proportional, or do we need to >>>> throw away proportional share because there is throttling?". This >>>> question is simply wrong, as I think consciously (sorry for my >>>> dissociated behavior :) ). >>> >>> I was wondering about the same. We need both and both should be able >>> to work with fast devices of today using blk-mq interfaces without >>> much overhead. >>> >>>> >>>> The best goal to achieve is to have both a good throttling mechanism, >>>> and a good proportional share scheduler. This goal would be valid if >>>> even if there was just one important scenario for each of the two >>>> approaches. The vulnus here is that you guys are constantly, and >>>> rightly, working on solutions to achieve and consolidate reasonable >>>> QoS guarantees, but an apparently very good proportional-share >>>> scheduler has been kept off for years. If you (or others) have good >>>> arguments to support this state of affairs, then this would probably >>>> be an important point to discuss. >>> >>> Paolo, CFQ is legacy now and if we can come up with a proportional >>> IO mechanism which works reasonably well with fast devices using >>> blk-mq interfaces, that will be much more interesting. >>> >> >> That's absolutely true. But, why do we pretend not to know that, for >> (at least) hundreds of thousands of users Linux will go on giving bad >> responsiveness, starvation, high latency and unfairness, until blk >> will not be used any more (assuming that these problems will somehow >> disappear will blk-mq). Many of these users are fully aware of these >> Linux long-standing problems. We could solve these problems by just >> adding a scheduler that has already been adopted, and thus extensively >> tested, by thousands of users. And more and more people are aware of >> this fact too. Are we doing the right thing? > > Hi Paolo, > Hi > People have been using CFQ for many years. Yes, but allow me just to add that a lot of people have also been unhappy with CFQ for many years. > I am not sure if benefits > offered by BFQ over CFQ are significant enough to justify taking a > completely new code and get rid of CFQ. Or are the benfits significant > enough that one feels like putting time and effort into this and > take chances wiht new code. > Although I think that BFQ's benefits are relevant (but I'm a little bit an interested party :) ), I do agree that abruptly replacing the most used I/O scheduler (AFAIK) with a so different one is at least a little risky. > At this point of time replacing CFQ with something better is not a > priority for me. ok > But if something better and stable goes upstream, I > will gladly use it. > Then, in case of success, I will be glad to receive some feedback from you, and possibly use it to improve the set of ideas that we have put into BFQ. Thank you, Paolo > Vivek > -- > To unsubscribe from this list: send the line "unsubscribe linux-block" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Mark Brown <broonie@kernel.org> |
|---|---|
| Date | 2016-10-06 21:50 +0200 |
| Message-ID | <spmQ1-5lG-11@gated-at.bofh.it> |
| In reply to | #1496784 |
[Multipart message — attachments visible in raw view] — view raw
On Thu, Oct 06, 2016 at 01:49:43PM -0400, Vivek Goyal wrote: > Paolo, CFQ is legacy now and if we can come up with a proportional > IO mechanism which works reasonably well with fast devices using > blk-mq interfaces, that will be much more interesting. CFQ is not legacy yet. Until we've got scheduling in blk-mq and converted all the existing users over to blk-mq it's still going to be something that affects real, production systems that don't have any other meaningful option.
[toc] | [prev] | [next] | [standalone]
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-07 00:30 +0200 |
| Message-ID | <sppkS-7h3-1@gated-at.bofh.it> |
| In reply to | #1496198 |
> Il giorno 06 ott 2016, alle ore 21:57, Shaohua Li <shli@fb.com> ha scritto: > > On Thu, Oct 06, 2016 at 09:58:44AM +0200, Paolo Valente wrote: >> >>> Il giorno 05 ott 2016, alle ore 22:46, Shaohua Li <shli@fb.com> ha scritto: >>> >>> On Wed, Oct 05, 2016 at 09:47:19PM +0200, Paolo Valente wrote: >>>> >>>>> Il giorno 05 ott 2016, alle ore 20:30, Shaohua Li <shli@fb.com> ha scritto: >>>>> >>>>> On Wed, Oct 05, 2016 at 10:49:46AM -0400, Tejun Heo wrote: >>>>>> Hello, Paolo. >>>>>> >>>>>> On Wed, Oct 05, 2016 at 02:37:00PM +0200, Paolo Valente wrote: >>>>>>> In this respect, for your generic, unpredictable scenario to make >>>>>>> sense, there must exist at least one real system that meets the >>>>>>> requirements of such a scenario. Or, if such a real system does not >>>>>>> yet exist, it must be possible to emulate it. If it is impossible to >>>>>>> achieve this last goal either, then I miss the usefulness >>>>>>> of looking for solutions for such a scenario. >>>>>>> >>>>>>> That said, let's define the instance(s) of the scenario that you find >>>>>>> most representative, and let's test BFQ on it/them. Numbers will give >>>>>>> us the answers. For example, what about all or part of the following >>>>>>> groups: >>>>>>> . one cyclically doing random I/O for some second and then sequential I/O >>>>>>> for the next seconds >>>>>>> . one doing, say, quasi-sequential I/O in ON/OFF cycles >>>>>>> . one starting an application cyclically >>>>>>> . one playing back or streaming a movie >>>>>>> >>>>>>> For each group, we could then measure the time needed to complete each >>>>>>> phase of I/O in each cycle, plus the responsiveness in the group >>>>>>> starting an application, plus the frame drop in the group streaming >>>>>>> the movie. In addition, we can measure the bandwidth/iops enjoyed by >>>>>>> each group, plus, of course, the aggregate throughput of the whole >>>>>>> system. In particular we could compare results with throttling, BFQ, >>>>>>> and CFQ. >>>>>>> >>>>>>> Then we could write resulting numbers on the stone, and stick to them >>>>>>> until something proves them wrong. >>>>>>> >>>>>>> What do you (or others) think about it? >>>>>> >>>>>> That sounds great and yeah it's lame that we didn't start with that. >>>>>> Shaohua, would it be difficult to compare how bfq performs against >>>>>> blk-throttle? >>>>> >>>>> I had a test of BFQ. >>>> >>>> Thank you very much for testing BFQ! >>>> >>>>> I'm using BFQ found at >>>>> https://urldefense.proofpoint.com/v2/url?u=http-3A__algogroup.unimore.it_people_paolo_disk-5Fsched_sources.php&d=DQIFAg&c=5VD0RTtNlTh3ycd41b3MUw&r=i6WobKxbeG3slzHSIOxTVtYIJw7qjCE6S0spDTKL-J4&m=2pG8KEx5tRymExa_K0ddKH_YvhH3qvJxELBd1_lw0-w&s=FZKEAOu2sw95y9jZio2k012cQWoLzlBWDl0NiGPVW78&e= . version is >>>>> 4.7.0-v8r3. >>>> >>>> That's the latest stable version. The development version [1] already >>>> contains further improvements for fairness, latency and throughput. >>>> It is however still a release candidate. >>>> >>>> [1] https://github.com/linusw/linux-bfq/tree/bfq-v8 >>>> >>>>> It's a LSI SSD, queue depth 32. I use default setting. fio script >>>>> is: >>>>> >>>>> [global] >>>>> ioengine=libaio >>>>> direct=1 >>>>> readwrite=randread >>>>> bs=4k >>>>> runtime=60 >>>>> time_based=1 >>>>> file_service_type=random:36 >>>>> overwrite=1 >>>>> thread=0 >>>>> group_reporting=1 >>>>> filename=/dev/sdb >>>>> iodepth=1 >>>>> numjobs=8 >>>>> >>>>> [groupA] >>>>> prio=2 >>>>> >>>>> [groupB] >>>>> new_group >>>>> prio=6 >>>>> >>>>> I'll change iodepth, numjobs and prio in different tests. result unit is MB/s. >>>>> >>>>> iodepth=1 numjobs=1 prio 4:4 >>>>> CFQ: 28:28 BFQ: 21:21 deadline: 29:29 >>>>> >>>>> iodepth=8 numjobs=1 prio 4:4 >>>>> CFQ: 162:162 BFQ: 102:98 deadline: 205:205 >>>>> >>>>> iodepth=1 numjobs=8 prio 4:4 >>>>> CFQ: 157:157 BFQ: 81:92 deadline: 196:197 >>>>> >>>>> iodepth=1 numjobs=1 prio 2:6 >>>>> CFQ: 26.7:27.6 BFQ: 20:6 deadline: 29:29 >>>>> >>>>> iodepth=8 numjobs=1 prio 2:6 >>>>> CFQ: 166:174 BFQ: 139:72 deadline: 202:202 >>>>> >>>>> iodepth=1 numjobs=8 prio 2:6 >>>>> CFQ: 148:150 BFQ: 90:77 deadline: 198:197 >>>>> >>>>> CFQ isn't fair at all. BFQ is very good in this side, but has poor throughput >>>>> even prio is the default value. >>>>> >>>> >>>> Throughput is lower with BFQ for two reasons. >>>> >>>> First, you certainly left the low_latency in its default state, i.e., >>>> on. As explained, e.g., here [2], low_latency mode is totally geared >>>> towards maximum responsiveness and minimum latency for soft real-time >>>> applications (e.g., video players). To achieve this goal, BFQ is >>>> willing to perform more idling, when necessary. This lowers >>>> throughput (I'll get back on this at the end of the discussion of the >>>> second reason). >>> >>> changing low_latency to 0 seems not change anything, at least for the test: >>> iodepth=1 numjobs=1 prio 2:6 A bs 4k:64k >>> >>>> The second, most important reason, is that a minimum of idling is the >>>> *only* way to achieve differentiated bandwidth distribution, as you >>>> requested by setting different ioprios. I stress that this constraint >>>> is not a technological accident, but a intrinsic, logical necessity. >>>> The proof is simple, and if the following explanation is too boring or >>>> confusing, I can show it to you with any trace of sync I/O. >>>> >>>> First, to provide differentiated service, you need per-process >>>> scheduling, i.e., schedulers in which there is a separate queue >>>> associated with each process. Now, let A be the process with higher >>>> weight (ioprio), and B the process with lower weight. Both processes >>>> are sync, thus, by definition, they issue requests as follows: a few >>>> requests (probably two, or a little bit more with larger iodepth), >>>> then a little break to wait for request completion, then the next >>>> small batch and so on. For each process, the queue associated with >>>> the process (in the scheduler) is necessarily empty on the break. As >>>> a consequence, if there is no idling, then every time A reaches its >>>> break, the scheduler has only the option to switch to B (which is >>>> extremely likely to have pending requests). >>>> >>>> The service pattern of the processes then unavoidably becomes: >>>> >>>> A B A B A B ... >>>> >>>> where each letter represents a full small batch served for the >>>> process. That is, 50% of the bw for each process, and complete loss >>>> of control on the desired bandwidth distribution. >>>> >>>> So, to sum up, the reason why BFQ achieves a lower total bw is that it >>>> behaves in the only correct way to respect weights with sync I/O, >>>> i.e., it performs a little idling. If low_latency is on, then BFQ >>>> increases idling further, and this may be have caused further bw loss >>>> in your test (but this varies greatly with devices, so you can >>>> discover it only by trying). >>>> >>>> The bottom line is that if you do want to achieve differentiation with >>>> sync I/O, you have to pay a price in terms of bw, because of idling. >>>> Actually, the recent preemption mechanism that I have introduced in >>>> BFQ is proving so effective in preserving differentiation, that I'm >>>> tempted to try some almost idleness solution. A little of accuracy >>>> should however be sacrificed. Anyway, this is still work in progress. >>> >>> Yep, I fully understand why idle is required here. As long as workload io depth >>> is lower than queue io depth, idle is the only way to maintain fairness. This >>> is the core of CFQ, I bet the same for BFQ. Unfortunately idle disk harms >>> throughput too much especially for high end SSD. >>> >> >> Then I'm afraid I have to give you very bad news: bw limiting causes >> the same throughput loss. You can see it from your very tests. Here >> is one of your results with BFQ (one that is likely to have been >> affected less by the fact that you left low_latency on, or by further >> issues that I may have not yet addressed thoroughly): >> >> iodepth=8 numjobs=1 prio 2:6 >> CFQ: 166:174 BFQ: 139:72 deadline: 202:202 >> >> Here is, instead, your test with bw limitation: >> iodepth=8 numjobs=1 prio 2:6, group A has 50M/s limit >> CFQ:51:207 BFQ: 51:45 deadline: 51:216 >> >> From the first test, you see that the total bw achievable by the >> device is at least 404MB/s. But in the second test you get at most >> 267MB/s, with deadline. In this respect, the total bw achieved by BFQ >> in the first test is 211MB/s. >> >> So, both throttling and proportional share need to waste bw, BFQ >> looses about 13% more of the total bw. > > I don't think you calculate this correct. In iodepth 8 request size 4k, the > workload can only dispatch 216M/s. Even group A doesn't dispatch any IO, group > B can only dispatch 216M/s. So deadline doesn't waste any bw. CFQ wastes 216 - > 207, while BFQ wastes 216 - 45. That's the problem. > I'm sorry, I have chosen an ambiguous example for showing my point. I thought of it too superficially, because it was very simple and provided numbers on which we would have both agreed. Yes, group B doesn't lose anything with deadline, no matter what limit we impose on group A. That's why this is a case where one would never do any bw limitation. One needs bw limitation when some group is choked by other groups. For this to happen, some of the groups causing the problem must be receiving, each, less than the bw they can sustain. In this situation, the total workload is pumping the device to the maximum total bw achievable given the properties of the workload of each group (iodepth, block size, degree of randomness, ...). If one intervenes with group limits, the she/he must be quite clever in finding the right limits for each offending group, such that a total bw close to the no-limit case is still achieved. If the system is dynamic, i.e., number of groups, as well as group iodepths, block sizes and randomness may vary with time, then she/he must be very lucky too, as she/he must find an automatic limit recomputation that still achieves about maximum bw, however groups and workloads change. IMO, depending on the system, achieving such a goal may become virtually impossible. >> In return, it gives you >> incomparably better bw and latency guarantees, while allowing you to >> configure your system with zero or minimal effort. In contrast, using >> bw limits to properly configure a common system like, e.g., a large >> file server, may become a nightmare for a sysadmin. For example, if, >> in the simplest case, she/he configures limits for the worst-case, >> then per-client limits will have to be extremely low. But the system >> is large and dynamic, so the actual number of clients and at the >> actual bw consumed by each client will vary without a break, even in >> the short-medium term. The bw redistribution heuristic do not give >> any provable guarantees on accuracy of bw redistribution. The result >> will likely be highly varying client bandwidths, with unlucky clients >> unjustly limited to low limits, and then experiencing high latencies. >> The latter will be further emphasized by the intrinsically bursty >> nature of throttling. > > I don't disagree here. The bw/iops throttling is not easy to configure. It's > kind of low level configuration, people need to know the workload very well to > configure. If we have way to do proportional scheduling, everybody will cheer > up. The goal of the tests is checking if proportional scheduling is feasible, > the result isn't very optimistic so far. If you think that, in contrast, with bw limits you easily achieve about maximum total bw, then your conclusion is correct. Above I have tried to restate why this seems quite unlikely to me on moderately complex, possibly dynamic systems. > No, I don't say BFQ isn't good. It is > much more fair than CFQ. I suppose it would work well for desktop workloads. >> In addition, the scenario in your tests is the worst case for a >> proportional share solution: in a generic system, such as the the file >> server above, part of the workload is likely to be sequential or >> quasi-sequential (at least in medium-length time intervals), and this >> is enough to get very close to peak bw with a proportional-share >> scheduler. No configuration needed. With bw throttling, you must do >> the math very well to get peak bw all the time in a dynamic system. > > I'm afraid this is not true. Workloads seldomly fully utilize the bandwidth of > a highend SSD. Achieving full utilization in a complex system is certainly a rather unlikely event. The matter here is how much more you may lose when you have to provide bw or other service guarantees, to individual users or groups. In this respect, I'm unfortunately too ignorant to have an idea of the global distribution of significant workloads. For this reason, in a case like this, I try to proceed with concrete examples, to avoid hard-to-refute statements or positions. Here is why I made the file-server example. In this example, workload is not tied to be fully random. How relevant is this case? I don' know. To me, it seems relevant. We could list other examples in which the workload seems mostly sequential or quasi-sequential (video streaming, audio streaming, packet-traffic dumping, ...). In all these case, it seems to me that there may be bw to loose if the individual-bw control mechanism is not well tuned. > At least that's true here. My tests are actually pretty normal. > I didn't do weird tests at all. > I don't think your tests are weird at all! Simply they focus on one, yet very important, case. Is this the only significant case? My examples above induce me to think that it is not. Just this. Thanks, Paolo > Thanks, > Shaohua > -- > To unsubscribe from this list: send the line "unsubscribe linux-block" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-05 22:00 +0200 |
| Message-ID | <sp0w9-6iw-9@gated-at.bofh.it> |
| In reply to | #1495905 |
> Il giorno 05 ott 2016, alle ore 21:08, Shaohua Li <shli@fb.com> ha scritto: > > On Wed, Oct 05, 2016 at 11:30:53AM -0700, Shaohua Li wrote: >> On Wed, Oct 05, 2016 at 10:49:46AM -0400, Tejun Heo wrote: >>> Hello, Paolo. >>> >>> On Wed, Oct 05, 2016 at 02:37:00PM +0200, Paolo Valente wrote: >>>> In this respect, for your generic, unpredictable scenario to make >>>> sense, there must exist at least one real system that meets the >>>> requirements of such a scenario. Or, if such a real system does not >>>> yet exist, it must be possible to emulate it. If it is impossible to >>>> achieve this last goal either, then I miss the usefulness >>>> of looking for solutions for such a scenario. >>>> >>>> That said, let's define the instance(s) of the scenario that you find >>>> most representative, and let's test BFQ on it/them. Numbers will give >>>> us the answers. For example, what about all or part of the following >>>> groups: >>>> . one cyclically doing random I/O for some second and then sequential I/O >>>> for the next seconds >>>> . one doing, say, quasi-sequential I/O in ON/OFF cycles >>>> . one starting an application cyclically >>>> . one playing back or streaming a movie >>>> >>>> For each group, we could then measure the time needed to complete each >>>> phase of I/O in each cycle, plus the responsiveness in the group >>>> starting an application, plus the frame drop in the group streaming >>>> the movie. In addition, we can measure the bandwidth/iops enjoyed by >>>> each group, plus, of course, the aggregate throughput of the whole >>>> system. In particular we could compare results with throttling, BFQ, >>>> and CFQ. >>>> >>>> Then we could write resulting numbers on the stone, and stick to them >>>> until something proves them wrong. >>>> >>>> What do you (or others) think about it? >>> >>> That sounds great and yeah it's lame that we didn't start with that. >>> Shaohua, would it be difficult to compare how bfq performs against >>> blk-throttle? >> >> I had a test of BFQ. I'm using BFQ found at >> http://algogroup.unimore.it/people/paolo/disk_sched/sources.php. version is >> 4.7.0-v8r3. It's a LSI SSD, queue depth 32. I use default setting. fio script >> is: >> >> [global] >> ioengine=libaio >> direct=1 >> readwrite=randread >> bs=4k >> runtime=60 >> time_based=1 >> file_service_type=random:36 >> overwrite=1 >> thread=0 >> group_reporting=1 >> filename=/dev/sdb >> iodepth=1 >> numjobs=8 >> >> [groupA] >> prio=2 >> >> [groupB] >> new_group >> prio=6 >> >> I'll change iodepth, numjobs and prio in different tests. result unit is MB/s. >> >> iodepth=1 numjobs=1 prio 4:4 >> CFQ: 28:28 BFQ: 21:21 deadline: 29:29 >> >> iodepth=8 numjobs=1 prio 4:4 >> CFQ: 162:162 BFQ: 102:98 deadline: 205:205 >> >> iodepth=1 numjobs=8 prio 4:4 >> CFQ: 157:157 BFQ: 81:92 deadline: 196:197 >> >> iodepth=1 numjobs=1 prio 2:6 >> CFQ: 26.7:27.6 BFQ: 20:6 deadline: 29:29 >> >> iodepth=8 numjobs=1 prio 2:6 >> CFQ: 166:174 BFQ: 139:72 deadline: 202:202 >> >> iodepth=1 numjobs=8 prio 2:6 >> CFQ: 148:150 BFQ: 90:77 deadline: 198:197 > > More tests: > > iodepth=8 numjobs=1 prio 2:6, group A has 50M/s limit > CFQ:51:207 BFQ: 51:45 deadline: 51:216 > > iodepth=1 numjobs=1 prio 2:6, group A bs=4k, group B bs=64k > CFQ:25:249 BFQ: 23:42 deadline: 26:251 > A true proportional share scheduler like BFQ works under the assumption to be the only limiter of the bandwidth of its clients. And the availability of such a scheduler should apparently make bandwidth limiting useless: once you have a mechanism that allows you to give each group the desired fraction of the bandwidth, and to redistribute excess bandwidth seamlessly when needed, what do you need additional limiting for? But I'm not expert of any possible system configuration or requirement. So, if you have practical examples, I would really appreciate them. And I don't think it will be difficult to see what goes wrong in BFQ with external bw limitation, and to fix the problem. Thanks, Paolo > Thanks, > Shaohua > -- > To unsubscribe from this list: send the line "unsubscribe linux-block" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Paolo Valente <paolo.valente@unimore.it> |
|---|---|
| Date | 2016-10-06 09:30 +0200 |
| Message-ID | <spbhT-5mm-9@gated-at.bofh.it> |
| In reply to | #1496025 |
> Il giorno 05 ott 2016, alle ore 22:36, Shaohua Li <shli@fb.com> ha scritto: > > On Wed, Oct 05, 2016 at 09:57:22PM +0200, Paolo Valente wrote: >> >>> Il giorno 05 ott 2016, alle ore 21:08, Shaohua Li <shli@fb.com> ha scritto: >>> >>> On Wed, Oct 05, 2016 at 11:30:53AM -0700, Shaohua Li wrote: >>>> On Wed, Oct 05, 2016 at 10:49:46AM -0400, Tejun Heo wrote: >>>>> Hello, Paolo. >>>>> >>>>> On Wed, Oct 05, 2016 at 02:37:00PM +0200, Paolo Valente wrote: >>>>>> In this respect, for your generic, unpredictable scenario to make >>>>>> sense, there must exist at least one real system that meets the >>>>>> requirements of such a scenario. Or, if such a real system does not >>>>>> yet exist, it must be possible to emulate it. If it is impossible to >>>>>> achieve this last goal either, then I miss the usefulness >>>>>> of looking for solutions for such a scenario. >>>>>> >>>>>> That said, let's define the instance(s) of the scenario that you find >>>>>> most representative, and let's test BFQ on it/them. Numbers will give >>>>>> us the answers. For example, what about all or part of the following >>>>>> groups: >>>>>> . one cyclically doing random I/O for some second and then sequential I/O >>>>>> for the next seconds >>>>>> . one doing, say, quasi-sequential I/O in ON/OFF cycles >>>>>> . one starting an application cyclically >>>>>> . one playing back or streaming a movie >>>>>> >>>>>> For each group, we could then measure the time needed to complete each >>>>>> phase of I/O in each cycle, plus the responsiveness in the group >>>>>> starting an application, plus the frame drop in the group streaming >>>>>> the movie. In addition, we can measure the bandwidth/iops enjoyed by >>>>>> each group, plus, of course, the aggregate throughput of the whole >>>>>> system. In particular we could compare results with throttling, BFQ, >>>>>> and CFQ. >>>>>> >>>>>> Then we could write resulting numbers on the stone, and stick to them >>>>>> until something proves them wrong. >>>>>> >>>>>> What do you (or others) think about it? >>>>> >>>>> That sounds great and yeah it's lame that we didn't start with that. >>>>> Shaohua, would it be difficult to compare how bfq performs against >>>>> blk-throttle? >>>> >>>> I had a test of BFQ. I'm using BFQ found at >>>> https://urldefense.proofpoint.com/v2/url?u=http-3A__algogroup.unimore.it_people_paolo_disk-5Fsched_sources.php&d=DQIFAg&c=5VD0RTtNlTh3ycd41b3MUw&r=X13hAPkxmvBro1Ug8vcKHw&m=zB09S7v2QifXXTa6f2_r6YLjiXq3AwAi7sqO4o2UfBQ&s=oMKpjQMXfWmMwHmANB-Qnrm2EdERzz9Oef7jcLkbyFg&e= . version is >>>> 4.7.0-v8r3. It's a LSI SSD, queue depth 32. I use default setting. fio script >>>> is: >>>> >>>> [global] >>>> ioengine=libaio >>>> direct=1 >>>> readwrite=randread >>>> bs=4k >>>> runtime=60 >>>> time_based=1 >>>> file_service_type=random:36 >>>> overwrite=1 >>>> thread=0 >>>> group_reporting=1 >>>> filename=/dev/sdb >>>> iodepth=1 >>>> numjobs=8 >>>> >>>> [groupA] >>>> prio=2 >>>> >>>> [groupB] >>>> new_group >>>> prio=6 >>>> >>>> I'll change iodepth, numjobs and prio in different tests. result unit is MB/s. >>>> >>>> iodepth=1 numjobs=1 prio 4:4 >>>> CFQ: 28:28 BFQ: 21:21 deadline: 29:29 >>>> >>>> iodepth=8 numjobs=1 prio 4:4 >>>> CFQ: 162:162 BFQ: 102:98 deadline: 205:205 >>>> >>>> iodepth=1 numjobs=8 prio 4:4 >>>> CFQ: 157:157 BFQ: 81:92 deadline: 196:197 >>>> >>>> iodepth=1 numjobs=1 prio 2:6 >>>> CFQ: 26.7:27.6 BFQ: 20:6 deadline: 29:29 >>>> >>>> iodepth=8 numjobs=1 prio 2:6 >>>> CFQ: 166:174 BFQ: 139:72 deadline: 202:202 >>>> >>>> iodepth=1 numjobs=8 prio 2:6 >>>> CFQ: 148:150 BFQ: 90:77 deadline: 198:197 >>> >>> More tests: >>> >>> iodepth=8 numjobs=1 prio 2:6, group A has 50M/s limit >>> CFQ:51:207 BFQ: 51:45 deadline: 51:216 >>> >>> iodepth=1 numjobs=1 prio 2:6, group A bs=4k, group B bs=64k >>> CFQ:25:249 BFQ: 23:42 deadline: 26:251 >>> >> >> A true proportional share scheduler like BFQ works under the >> assumption to be the only limiter of the bandwidth of its clients. >> And the availability of such a scheduler should apparently make >> bandwidth limiting useless: once you have a mechanism that allows you >> to give each group the desired fraction of the bandwidth, and to >> redistribute excess bandwidth seamlessly when needed, what do you need >> additional limiting for? >> >> But I'm not expert of any possible system configuration or >> requirement. So, if you have practical examples, I would really >> appreciate them. And I don't think it will be difficult to see what >> goes wrong in BFQ with external bw limitation, and to fix the >> problem. > > I think the test emulates a very common configuration. We assign more IO > resources to high priority workload. But such workload doesn't always dispatch > enough io. That's why I set a rate limit. When this happend, we hope low > priority workload uses the disk bandwidth. That's the whole point of disk > sharing. > But that's exactly the configuration for which a proportional-share scheduler is designed: systematically and seamlessly redistribute excess bw, with no configuration needed. Or is there something else in the scenario you have in mind? Thanks, Paolo > Thanks, > Shaohua > -- > To unsubscribe from this list: send the line "unsubscribe linux-block" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html -- Paolo Valente Algogroup Dipartimento di Scienze Fisiche, Informatiche e Matematiche Via Campi 213/B 41125 Modena - Italy http://algogroup.unimore.it/people/paolo/
[toc] | [prev] | [next] | [standalone]
| From | Kyle Sanderson <kyle.leet@gmail.com> |
|---|---|
| Date | 2016-10-09 03:20 +0200 |
| Message-ID | <sqaWt-63w-5@gated-at.bofh.it> |
| In reply to | #1495587 |
Re-sending as plain-text as the Gmail Android App is still historically broken... ---------- Forwarded message ---------- From: Kyle Sanderson <kyle.leet@gmail.com> Date: Wed, Oct 5, 2016 at 7:09 AM Subject: Re: [PATCH V3 00/11] block-throttle: add .high limit To: Tejun Heo <tj@kernel.org> Cc: jmoyer@redhat.com, Paolo Valente <paolo.valente@unimore.it>, linux-kernel@vger.kernel.org, Mark Brown <broonie@kernel.org>, linux-block@vger.kernel.org, Shaohua Li <shli@fb.com>, Jens Axboe <axboe@fb.com>, Linus Walleij <linus.walleij@linaro.org>, Vivek Goyal <vgoyal@redhat.com>, Kernel-team@fb.com, Ulf Hansson <ulf.hansson@linaro.org> Obviously not to compound against this, however it has been proven for years that CFQ will lock significantly under contention, and other schedulers, such as BFQ attempt to provide fairness which is absolutely the desired outcome from using a machine. The networking space is a little wrecked in the sense there's a plethora of qdiscs that don't necessarily need to exist; but are legacy. This limitation does not exist in this realm as there are no specific tunables. There is no reason that in 2016 a user-space application can steal all of the I/O from a disk, completely locking the machine when BFQ has essentially solved this years ago. I've been a moderately happy user of BFQ for quite sometime now. There aren't tens, or hundreds of us, but thousands through the custom kernels that are spun, and the distros that helped support BFQ. How is this even a discussion when hard numbers, and trying any reproduction case easily reproduce the issues that CFQ causes. Reading this thread, and many others only grows not only my disappointment, but whenever someone launches kterm or scrot and their machine freezes, leaves a selective few individuals completely responsible for this. Help those users, help yourself, help Linux. On 4 Oct 2016 1:29 pm, "Tejun Heo" <tj@kernel.org> wrote: > > Hello, Paolo. > > On Tue, Oct 04, 2016 at 09:29:48PM +0200, Paolo Valente wrote: > > > Hmm... I think we already discussed this but here's a really simple > > > case. There are three unknown workloads A, B and C and we want to > > > give A certain best-effort guarantees (let's say around 80% of the > > > underlying device) whether A is sharing the device with B or C. > > > > That's the same example that you proposed me in our previous > > discussion. For this example I showed you, with many boring numbers, > > that with BFQ you get the most accurate distribution of the resource. > > Yes, it is about the same example and what I understood was that > "accurate distribution of the resources" holds as long as the > randomness is incidental (ie. due to layout on the filesystem and so > on) with the slice expiration mechanism offsetting the actually random > workloads. > > > If you have enough stamina, I can repeat them again. To save your > > I'll go back to the thread and re-read them. > > > patience, here is a very brief summary. In a concrete use case, the > > unknown workloads turn into something like this: there will be a first > > time interval during which A happens to be, say, sequential, B happens > > to be, say, random and C happens to be, say, quasi-sequential. Then > > there will be a next time interval during which their characteristics > > change, and so on. It is easy (but boring, I acknowledge it) to show > > that, for each of these time intervals BFQ provides the best possible > > service in terms of fairness, bandwidth distribution, stability and so > > on. Why? Because of the elastic bandwidth-time scheduling of BFQ > > that we already discussed, and because BFQ is naturally accurate in > > redistributing aggregate throughput proportionally, when needed. > > Yeah, that's what I remember and for workload above certain level of > randomness its time consumption is mapped to bw, right? > > > > I get that bfq can be a good compromise on most desktop workloads and > > > behave reasonably well for some server workloads with the slice > > > expiration mechanism but it really isn't an IO resource partitioning > > > mechanism. > > > > Right. My argument is that BFQ enables you to give to each client the > > bandwidth and low-latency guarantees you want. And this IMO is way > > better than partitioning a resource and then getting unavoidable > > unfairness and high latency. > > But that statement only holds while bw is the main thing to guarantee, > no? The level of isolation that we're looking for here is fairly > strict adherence to sub/few-milliseconds in terms of high percentile > scheduling latency while within the configured bw/iops limits, not > "overall this device is being used pretty well". > > Thanks. > > -- > tejun
[toc] | [prev] | [next] | [standalone]
| From | Linus Walleij <linus.walleij@linaro.org> |
|---|---|
| Date | 2016-10-06 10:10 +0200 |
| Message-ID | <spbUB-5Qv-17@gated-at.bofh.it> |
| In reply to | #1495565 |
On Tue, Oct 4, 2016 at 9:14 PM, Tejun Heo <tj@kernel.org> wrote:
> I get that bfq can be a good compromise on most desktop workloads and
> behave reasonably well for some server workloads with the slice
> expiration mechanism but it really isn't an IO resource partitioning
> mechanism.
Not just desktops, also Android phones.
So why not have BFQ as a separate scheduling policy upstream,
alongside CFQ, deadline and noop?
I understand the CPU scheduler people's position that they want
one scheduler for everyone's everyday loads (except RT and
SCHED_DEADLINE) and I guess that is the source of the highlander
"there can be only one" argument, but note this:
kernel/Kconfig.preempt:
config PREEMPT_NONE
bool "No Forced Preemption (Server)"
config PREEMPT_VOLUNTARY
bool "Voluntary Kernel Preemption (Desktop)"
config PREEMPT
bool "Preemptible Kernel (Low-Latency Desktop)"
We're already doing the per-usecase Kconfig thing for preemption.
But maybe somebody already hates that and want to get rid of it,
I don't know.
Yours,
Linus Walleij
[toc] | [prev] | [next] | [standalone]
| From | Mark Brown <broonie@kernel.org> |
|---|---|
| Date | 2016-10-06 13:10 +0200 |
| Message-ID | <speIO-7Nv-47@gated-at.bofh.it> |
| In reply to | #1496204 |
[Multipart message — attachments visible in raw view] — view raw
On Thu, Oct 06, 2016 at 10:04:41AM +0200, Linus Walleij wrote: > On Tue, Oct 4, 2016 at 9:14 PM, Tejun Heo <tj@kernel.org> wrote: > > I get that bfq can be a good compromise on most desktop workloads and > > behave reasonably well for some server workloads with the slice > > expiration mechanism but it really isn't an IO resource partitioning > > mechanism. > Not just desktops, also Android phones. > So why not have BFQ as a separate scheduling policy upstream, > alongside CFQ, deadline and noop? Right. > We're already doing the per-usecase Kconfig thing for preemption. > But maybe somebody already hates that and want to get rid of it, > I don't know. Hannes also suggested going back to making BFQ a separate scheduler rather than replacing CFQ earlier, pointing out that it mitigates against the risks of changing CFQ substantially at this point (which seems to be the biggest issue here).
[toc] | [prev] | [next] | [standalone]
| From | "Austin S. Hemmelgarn" <ahferroin7@gmail.com> |
|---|---|
| Date | 2016-10-06 14:00 +0200 |
| Message-ID | <spfvb-85m-3@gated-at.bofh.it> |
| In reply to | #1496594 |
On 2016-10-06 07:03, Mark Brown wrote: > On Thu, Oct 06, 2016 at 10:04:41AM +0200, Linus Walleij wrote: >> On Tue, Oct 4, 2016 at 9:14 PM, Tejun Heo <tj@kernel.org> wrote: > >>> I get that bfq can be a good compromise on most desktop workloads and >>> behave reasonably well for some server workloads with the slice >>> expiration mechanism but it really isn't an IO resource partitioning >>> mechanism. > >> Not just desktops, also Android phones. > >> So why not have BFQ as a separate scheduling policy upstream, >> alongside CFQ, deadline and noop? > > Right. > >> We're already doing the per-usecase Kconfig thing for preemption. >> But maybe somebody already hates that and want to get rid of it, >> I don't know. > > Hannes also suggested going back to making BFQ a separate scheduler > rather than replacing CFQ earlier, pointing out that it mitigates > against the risks of changing CFQ substantially at this point (which > seems to be the biggest issue here). > ISTR that the original argument for this approach essentially amounted to: 'If it's so much better, why do we need both?'. Such an argument is valid only if the new design is better in all respects (which there isn't sufficient information to decide in this case), or the negative aspects are worth the improvements (which is too workload specific to decide for something like this).
[toc] | [prev] | [next] | [standalone]
Page 2 of 3 — ← Prev page 1 [2] 3 Next page →
Back to top | Article view | linux.kernel
csiph-web