Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1340046 > unrolled thread
| Started by | Kent Overstreet <kent.overstreet@gmail.com> |
|---|---|
| First post | 2016-02-23 00:00 +0100 |
| Last post | 2016-02-23 21:50 +0100 |
| Articles | 6 — 4 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: 4.4-final: 28 bioset threads on small notebook Kent Overstreet <kent.overstreet@gmail.com> - 2016-02-23 00:00 +0100
Re: 4.4-final: 28 bioset threads on small notebook Ming Lei <ming.lei@canonical.com> - 2016-02-23 04:00 +0100
Re: 4.4-final: 28 bioset threads on small notebook Mike Snitzer <snitzer@redhat.com> - 2016-02-23 16:00 +0100
Re: 4.4-final: 28 bioset threads on small notebook Ming Lei <ming.lei@canonical.com> - 2016-02-24 03:50 +0100
Re: 4.4-final: 28 bioset threads on small notebook Kent Overstreet <kent.overstreet@gmail.com> - 2016-02-24 04:30 +0100
Re: 4.4-final: 28 bioset threads on small notebook Pavel Machek <pavel@ucw.cz> - 2016-02-23 21:50 +0100
| From | Kent Overstreet <kent.overstreet@gmail.com> |
|---|---|
| Date | 2016-02-23 00:00 +0100 |
| Subject | Re: 4.4-final: 28 bioset threads on small notebook |
| Message-ID | <r57Cq-4w9-1@gated-at.bofh.it> |
On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote: > On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote: > >>-----Original Message----- > > > > So it's almost already "per request_queue" > > Yes, that is because of the following line: > > q->bio_split = bioset_create(BIO_POOL_SIZE, 0); > > in blk_alloc_queue_node(). > > Looks like this bio_set doesn't need to be per-request_queue, and > now it is only used for fast-cloning bio for splitting, and one global > split bio_set should be enough. It does have to be per request queue for stacking block devices (which includes loopback).
[toc] | [next] | [standalone]
| From | Ming Lei <ming.lei@canonical.com> |
|---|---|
| Date | 2016-02-23 04:00 +0100 |
| Message-ID | <r5bmG-7ep-9@gated-at.bofh.it> |
| In reply to | #1340046 |
On Tue, Feb 23, 2016 at 6:58 AM, Kent Overstreet <kent.overstreet@gmail.com> wrote: > On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote: >> On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote: >> >>-----Original Message----- >> > >> > So it's almost already "per request_queue" >> >> Yes, that is because of the following line: >> >> q->bio_split = bioset_create(BIO_POOL_SIZE, 0); >> >> in blk_alloc_queue_node(). >> >> Looks like this bio_set doesn't need to be per-request_queue, and >> now it is only used for fast-cloning bio for splitting, and one global >> split bio_set should be enough. > > It does have to be per request queue for stacking block devices (which includes > loopback). In commit df2cb6daa4(block: Avoid deadlocks with bio allocation by stacking drivers), deadlock in this situation has been avoided already. Or are there other issues with global bio_set? I appreciate if you may explain it a bit if there are. Thanks, Ming Lei
[toc] | [prev] | [next] | [standalone]
| From | Mike Snitzer <snitzer@redhat.com> |
|---|---|
| Date | 2016-02-23 16:00 +0100 |
| Message-ID | <r5mBt-6Ra-27@gated-at.bofh.it> |
| In reply to | #1340193 |
On Mon, Feb 22 2016 at 9:55pm -0500, Ming Lei <ming.lei@canonical.com> wrote: > On Tue, Feb 23, 2016 at 6:58 AM, Kent Overstreet > <kent.overstreet@gmail.com> wrote: > > On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote: > >> On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote: > >> >>-----Original Message----- > >> > > >> > So it's almost already "per request_queue" > >> > >> Yes, that is because of the following line: > >> > >> q->bio_split = bioset_create(BIO_POOL_SIZE, 0); > >> > >> in blk_alloc_queue_node(). > >> > >> Looks like this bio_set doesn't need to be per-request_queue, and > >> now it is only used for fast-cloning bio for splitting, and one global > >> split bio_set should be enough. > > > > It does have to be per request queue for stacking block devices (which includes > > loopback). > > In commit df2cb6daa4(block: Avoid deadlocks with bio allocation by > stacking drivers), deadlock in this situation has been avoided already. > Or are there other issues with global bio_set? I appreciate if you may > explain it a bit if there are. Even with commit df2cb6daa4 there is still risk of deadlocks (even without low memory condition), see: https://patchwork.kernel.org/patch/7398411/ (you may recall you blocked this patch with concerns about performance, context switches, plug merging being compromised, etc.. to which I never circled back to verify your concerns) But it illustrates the type of problems that can occur when your rescue infrastructure is shared across devices (in the context of df2cb6daa4, current->bio_list contains bios from multiple devices). If a single splitting bio_set were shared across devices there would be no guarantee of forward progress with complex stacked devices (one or more devices could exhaust the reserve and starve out other devices in the stack). So keeping the bio_set per request_queue isn't prone to failure like a shared bio_set might be. Mike
[toc] | [prev] | [next] | [standalone]
| From | Ming Lei <ming.lei@canonical.com> |
|---|---|
| Date | 2016-02-24 03:50 +0100 |
| Message-ID | <r5xGx-6lj-3@gated-at.bofh.it> |
| In reply to | #1340730 |
On Tue, Feb 23, 2016 at 10:54 PM, Mike Snitzer <snitzer@redhat.com> wrote:
> On Mon, Feb 22 2016 at 9:55pm -0500,
> Ming Lei <ming.lei@canonical.com> wrote:
>
>> On Tue, Feb 23, 2016 at 6:58 AM, Kent Overstreet
>> <kent.overstreet@gmail.com> wrote:
>> > On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote:
>> >> On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote:
>> >> >>-----Original Message-----
>> >> >
>> >> > So it's almost already "per request_queue"
>> >>
>> >> Yes, that is because of the following line:
>> >>
>> >> q->bio_split = bioset_create(BIO_POOL_SIZE, 0);
>> >>
>> >> in blk_alloc_queue_node().
>> >>
>> >> Looks like this bio_set doesn't need to be per-request_queue, and
>> >> now it is only used for fast-cloning bio for splitting, and one global
>> >> split bio_set should be enough.
>> >
>> > It does have to be per request queue for stacking block devices (which includes
>> > loopback).
>>
>> In commit df2cb6daa4(block: Avoid deadlocks with bio allocation by
>> stacking drivers), deadlock in this situation has been avoided already.
>> Or are there other issues with global bio_set? I appreciate if you may
>> explain it a bit if there are.
>
> Even with commit df2cb6daa4 there is still risk of deadlocks (even
> without low memory condition), see:
> https://patchwork.kernel.org/patch/7398411/
That is definitely another problem which isn't related with low memory,
and I guess Kent means there might be deadlock risk in case of shared
bio_set.
>
> (you may recall you blocked this patch with concerns about performance,
> context switches, plug merging being compromised, etc.. to which I never
> circled back to verify your concerns)
I still remember that problem:
1) Process A
- two bio(a, b) are splitted in dm's make_request funtion
- bio(a) is submitted via generic_make_request(), so it is staged
in current->bio_list
- time t1
- before bio(b) is submitted, down_write(&s->lock) is run and
never return
2) Process B:
- just during time t1, wait completion of bio(a) by down_write(&s->lock)
Then Process A waits the lock which is acquired by B first, and the
two bio(a, b)
can't reach to driver/device at all.
Looks that current->bio_list is fragile to locks from make_request function,
and moving the lock into workqueue context should be helpful.
And I am happy to continue to discuss this issue further.
>
> But it illustrates the type of problems that can occur when your rescue
> infrastructure is shared across devices (in the context of df2cb6daa4,
> current->bio_list contains bios from multiple devices).
>
> If a single splitting bio_set were shared across devices there would be
> no guarantee of forward progress with complex stacked devices (one or
> more devices could exhaust the reserve and starve out other devices in
> the stack). So keeping the bio_set per request_queue isn't prone to
> failure like a shared bio_set might be.
Not consider the dm lock problem, from Kent's commit(df2cb6daa4) log and
the patch, looks forward progress can be guaranteed for stacked devices
with same bio_set, but better to get Kent's clarification.
If forward progress can be guaranteed, percpu mempool might avoid
easy exhausting, because it is reasonable to assume that one CPU can only
provide a certain amount of bandwidth wrt. block transfer.
Thanks
Ming
[toc] | [prev] | [next] | [standalone]
| From | Kent Overstreet <kent.overstreet@gmail.com> |
|---|---|
| Date | 2016-02-24 04:30 +0100 |
| Message-ID | <r5yjg-6Uj-1@gated-at.bofh.it> |
| In reply to | #1341261 |
On Wed, Feb 24, 2016 at 10:48:10AM +0800, Ming Lei wrote: > On Tue, Feb 23, 2016 at 10:54 PM, Mike Snitzer <snitzer@redhat.com> wrote: > > On Mon, Feb 22 2016 at 9:55pm -0500, > > Ming Lei <ming.lei@canonical.com> wrote: > > > >> On Tue, Feb 23, 2016 at 6:58 AM, Kent Overstreet > >> <kent.overstreet@gmail.com> wrote: > >> > On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote: > >> >> On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote: > >> >> >>-----Original Message----- > >> >> > > >> >> > So it's almost already "per request_queue" > >> >> > >> >> Yes, that is because of the following line: > >> >> > >> >> q->bio_split = bioset_create(BIO_POOL_SIZE, 0); > >> >> > >> >> in blk_alloc_queue_node(). > >> >> > >> >> Looks like this bio_set doesn't need to be per-request_queue, and > >> >> now it is only used for fast-cloning bio for splitting, and one global > >> >> split bio_set should be enough. > >> > > >> > It does have to be per request queue for stacking block devices (which includes > >> > loopback). > >> > >> In commit df2cb6daa4(block: Avoid deadlocks with bio allocation by > >> stacking drivers), deadlock in this situation has been avoided already. > >> Or are there other issues with global bio_set? I appreciate if you may > >> explain it a bit if there are. > > > > Even with commit df2cb6daa4 there is still risk of deadlocks (even > > without low memory condition), see: > > https://patchwork.kernel.org/patch/7398411/ > > That is definitely another problem which isn't related with low memory, > and I guess Kent means there might be deadlock risk in case of shared > bio_set. > > > > > (you may recall you blocked this patch with concerns about performance, > > context switches, plug merging being compromised, etc.. to which I never > > circled back to verify your concerns) > > I still remember that problem: > > 1) Process A > - two bio(a, b) are splitted in dm's make_request funtion > - bio(a) is submitted via generic_make_request(), so it is staged > in current->bio_list > - time t1 > - before bio(b) is submitted, down_write(&s->lock) is run and > never return > > 2) Process B: > - just during time t1, wait completion of bio(a) by down_write(&s->lock) > > Then Process A waits the lock which is acquired by B first, and the > two bio(a, b) > can't reach to driver/device at all. > > Looks that current->bio_list is fragile to locks from make_request function, > and moving the lock into workqueue context should be helpful. > > And I am happy to continue to discuss this issue further. > > > > > But it illustrates the type of problems that can occur when your rescue > > infrastructure is shared across devices (in the context of df2cb6daa4, > > current->bio_list contains bios from multiple devices). > > > > If a single splitting bio_set were shared across devices there would be > > no guarantee of forward progress with complex stacked devices (one or > > more devices could exhaust the reserve and starve out other devices in > > the stack). So keeping the bio_set per request_queue isn't prone to > > failure like a shared bio_set might be. > > Not consider the dm lock problem, from Kent's commit(df2cb6daa4) log and > the patch, looks forward progress can be guaranteed for stacked devices > with same bio_set, but better to get Kent's clarification. > > If forward progress can be guaranteed, percpu mempool might avoid > easy exhausting, because it is reasonable to assume that one CPU can only > provide a certain amount of bandwidth wrt. block transfer. Generally speaking, with potential deadlocks like this I don't bother to work out the specific scenario, it's enough to know that there's a shared resource and multiple users that depend on each other... if you've got that, you'll have a deadlock. But, if you're curious: say we've got block devices a and b, when you submit to a the bio will get passed down to b: for the bioset itself: if a bio gets split when submitted to a, then needs to be split again when it's submitted to b - you're allocating twice from the same mempool, and the first allocation can't be freed until the original bio completes. deadlock. with the rescuer threads it's more subtle, but you just need a scenario where the rescuer is required twice in a row. I'm not going to bother trying to work out the details, but it's the same principle - you can end up in a situation where you're blocked, and you need the rescuer thread to make forward progress (or you'd deadlock - that's why it exists, right?) - well, what happens if that happens twice in a row, and the second time you're running out of the rescuer thread? oops.
[toc] | [prev] | [next] | [standalone]
| From | Pavel Machek <pavel@ucw.cz> |
|---|---|
| Date | 2016-02-23 21:50 +0100 |
| Message-ID | <r5s4b-2dy-23@gated-at.bofh.it> |
| In reply to | #1340046 |
On Mon 2016-02-22 13:58:18, Kent Overstreet wrote: > On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote: > > On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote: > > >>-----Original Message----- > > > > > > So it's almost already "per request_queue" > > > > Yes, that is because of the following line: > > > > q->bio_split = bioset_create(BIO_POOL_SIZE, 0); > > > > in blk_alloc_queue_node(). > > > > Looks like this bio_set doesn't need to be per-request_queue, and > > now it is only used for fast-cloning bio for splitting, and one global > > split bio_set should be enough. > > It does have to be per request queue for stacking block devices (which includes > loopback). Could we only allocate request queues for devices that are not even opened? I have these in my system: loop0 loop2 loop4 loop6 md0 nbd1 nbd11 nbd13 nbd15 nbd3 nbd5 nbd7 nbd9 sda1 sda3 loop1 loop3 loop5 loop7 nbd0 nbd10 nbd12 nbd14 nbd2 nbd4 nbd6 nbd8 sda sda2 sda4 ...but nbd is never used, loop1+ is never used, and loop0 is only used once in a blue moon. Each process takes 8K+... Pavel -- (english) http://www.livejournal.com/~pavelmachek (cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web