Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1340046 > unrolled thread

Re: 4.4-final: 28 bioset threads on small notebook

Started byKent Overstreet <kent.overstreet@gmail.com>
First post2016-02-23 00:00 +0100
Last post2016-02-23 21:50 +0100
Articles 6 — 4 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: 4.4-final: 28 bioset threads on small notebook Kent Overstreet <kent.overstreet@gmail.com> - 2016-02-23 00:00 +0100
    Re: 4.4-final: 28 bioset threads on small notebook Ming Lei <ming.lei@canonical.com> - 2016-02-23 04:00 +0100
      Re: 4.4-final: 28 bioset threads on small notebook Mike Snitzer <snitzer@redhat.com> - 2016-02-23 16:00 +0100
        Re: 4.4-final: 28 bioset threads on small notebook Ming Lei <ming.lei@canonical.com> - 2016-02-24 03:50 +0100
          Re: 4.4-final: 28 bioset threads on small notebook Kent Overstreet <kent.overstreet@gmail.com> - 2016-02-24 04:30 +0100
    Re: 4.4-final: 28 bioset threads on small notebook Pavel Machek <pavel@ucw.cz> - 2016-02-23 21:50 +0100

#1340046 — Re: 4.4-final: 28 bioset threads on small notebook

FromKent Overstreet <kent.overstreet@gmail.com>
Date2016-02-23 00:00 +0100
SubjectRe: 4.4-final: 28 bioset threads on small notebook
Message-ID<r57Cq-4w9-1@gated-at.bofh.it>
On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote:
> On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote:
> >>-----Original Message-----
> >
> > So it's almost already "per request_queue"
> 
> Yes, that is because of the following line:
> 
> q->bio_split = bioset_create(BIO_POOL_SIZE, 0);
> 
> in blk_alloc_queue_node().
> 
> Looks like this bio_set doesn't need to be per-request_queue, and
> now it is only used for fast-cloning bio for splitting, and one global
> split bio_set should be enough.

It does have to be per request queue for stacking block devices (which includes
loopback).

[toc] | [next] | [standalone]


#1340193

FromMing Lei <ming.lei@canonical.com>
Date2016-02-23 04:00 +0100
Message-ID<r5bmG-7ep-9@gated-at.bofh.it>
In reply to#1340046
On Tue, Feb 23, 2016 at 6:58 AM, Kent Overstreet
<kent.overstreet@gmail.com> wrote:
> On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote:
>> On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote:
>> >>-----Original Message-----
>> >
>> > So it's almost already "per request_queue"
>>
>> Yes, that is because of the following line:
>>
>> q->bio_split = bioset_create(BIO_POOL_SIZE, 0);
>>
>> in blk_alloc_queue_node().
>>
>> Looks like this bio_set doesn't need to be per-request_queue, and
>> now it is only used for fast-cloning bio for splitting, and one global
>> split bio_set should be enough.
>
> It does have to be per request queue for stacking block devices (which includes
> loopback).

In commit df2cb6daa4(block: Avoid deadlocks with bio allocation by
stacking drivers), deadlock in this situation has been avoided already.
Or are there other issues with global bio_set? I appreciate if you may
explain it a bit if there are.

Thanks,
Ming Lei

[toc] | [prev] | [next] | [standalone]


#1340730

FromMike Snitzer <snitzer@redhat.com>
Date2016-02-23 16:00 +0100
Message-ID<r5mBt-6Ra-27@gated-at.bofh.it>
In reply to#1340193
On Mon, Feb 22 2016 at  9:55pm -0500,
Ming Lei <ming.lei@canonical.com> wrote:

> On Tue, Feb 23, 2016 at 6:58 AM, Kent Overstreet
> <kent.overstreet@gmail.com> wrote:
> > On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote:
> >> On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote:
> >> >>-----Original Message-----
> >> >
> >> > So it's almost already "per request_queue"
> >>
> >> Yes, that is because of the following line:
> >>
> >> q->bio_split = bioset_create(BIO_POOL_SIZE, 0);
> >>
> >> in blk_alloc_queue_node().
> >>
> >> Looks like this bio_set doesn't need to be per-request_queue, and
> >> now it is only used for fast-cloning bio for splitting, and one global
> >> split bio_set should be enough.
> >
> > It does have to be per request queue for stacking block devices (which includes
> > loopback).
> 
> In commit df2cb6daa4(block: Avoid deadlocks with bio allocation by
> stacking drivers), deadlock in this situation has been avoided already.
> Or are there other issues with global bio_set? I appreciate if you may
> explain it a bit if there are.

Even with commit df2cb6daa4 there is still risk of deadlocks (even
without low memory condition), see:
https://patchwork.kernel.org/patch/7398411/

(you may recall you blocked this patch with concerns about performance,
context switches, plug merging being compromised, etc.. to which I never
circled back to verify your concerns)

But it illustrates the type of problems that can occur when your rescue
infrastructure is shared across devices (in the context of df2cb6daa4,
current->bio_list contains bios from multiple devices). 

If a single splitting bio_set were shared across devices there would be
no guarantee of forward progress with complex stacked devices (one or
more devices could exhaust the reserve and starve out other devices in
the stack).  So keeping the bio_set per request_queue isn't prone to
failure like a shared bio_set might be.

Mike

[toc] | [prev] | [next] | [standalone]


#1341261

FromMing Lei <ming.lei@canonical.com>
Date2016-02-24 03:50 +0100
Message-ID<r5xGx-6lj-3@gated-at.bofh.it>
In reply to#1340730
On Tue, Feb 23, 2016 at 10:54 PM, Mike Snitzer <snitzer@redhat.com> wrote:
> On Mon, Feb 22 2016 at  9:55pm -0500,
> Ming Lei <ming.lei@canonical.com> wrote:
>
>> On Tue, Feb 23, 2016 at 6:58 AM, Kent Overstreet
>> <kent.overstreet@gmail.com> wrote:
>> > On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote:
>> >> On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote:
>> >> >>-----Original Message-----
>> >> >
>> >> > So it's almost already "per request_queue"
>> >>
>> >> Yes, that is because of the following line:
>> >>
>> >> q->bio_split = bioset_create(BIO_POOL_SIZE, 0);
>> >>
>> >> in blk_alloc_queue_node().
>> >>
>> >> Looks like this bio_set doesn't need to be per-request_queue, and
>> >> now it is only used for fast-cloning bio for splitting, and one global
>> >> split bio_set should be enough.
>> >
>> > It does have to be per request queue for stacking block devices (which includes
>> > loopback).
>>
>> In commit df2cb6daa4(block: Avoid deadlocks with bio allocation by
>> stacking drivers), deadlock in this situation has been avoided already.
>> Or are there other issues with global bio_set? I appreciate if you may
>> explain it a bit if there are.
>
> Even with commit df2cb6daa4 there is still risk of deadlocks (even
> without low memory condition), see:
> https://patchwork.kernel.org/patch/7398411/

That is definitely another problem which isn't related with low memory,
and I guess Kent means there might be deadlock risk in case of shared
bio_set.

>
> (you may recall you blocked this patch with concerns about performance,
> context switches, plug merging being compromised, etc.. to which I never
> circled back to verify your concerns)

I still remember that problem:

1) Process A
     - two bio(a, b) are splitted in dm's make_request funtion
     - bio(a) is submitted via generic_make_request(), so it is staged
       in current->bio_list
     - time t1
     - before bio(b) is submitted, down_write(&s->lock) is run and
      never return

2) Process B:
     - just during time t1, wait completion of bio(a) by down_write(&s->lock)

Then Process A waits the lock which is acquired by B first, and the
two bio(a, b)
can't reach to driver/device at all.

Looks that current->bio_list is fragile to locks from make_request function,
and moving the lock into workqueue context should be helpful.

And I am happy to continue to discuss this issue further.

>
> But it illustrates the type of problems that can occur when your rescue
> infrastructure is shared across devices (in the context of df2cb6daa4,
> current->bio_list contains bios from multiple devices).
>
> If a single splitting bio_set were shared across devices there would be
> no guarantee of forward progress with complex stacked devices (one or
> more devices could exhaust the reserve and starve out other devices in
> the stack).  So keeping the bio_set per request_queue isn't prone to
> failure like a shared bio_set might be.

Not consider the dm lock problem, from Kent's commit(df2cb6daa4) log and
the patch, looks forward progress can be guaranteed for stacked devices
with same bio_set, but better to get Kent's clarification.

If forward progress can be guaranteed, percpu mempool might avoid
easy exhausting, because it is reasonable to assume that one CPU can only
provide a certain amount of bandwidth wrt. block transfer.

Thanks
Ming

[toc] | [prev] | [next] | [standalone]


#1341267

FromKent Overstreet <kent.overstreet@gmail.com>
Date2016-02-24 04:30 +0100
Message-ID<r5yjg-6Uj-1@gated-at.bofh.it>
In reply to#1341261
On Wed, Feb 24, 2016 at 10:48:10AM +0800, Ming Lei wrote:
> On Tue, Feb 23, 2016 at 10:54 PM, Mike Snitzer <snitzer@redhat.com> wrote:
> > On Mon, Feb 22 2016 at  9:55pm -0500,
> > Ming Lei <ming.lei@canonical.com> wrote:
> >
> >> On Tue, Feb 23, 2016 at 6:58 AM, Kent Overstreet
> >> <kent.overstreet@gmail.com> wrote:
> >> > On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote:
> >> >> On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote:
> >> >> >>-----Original Message-----
> >> >> >
> >> >> > So it's almost already "per request_queue"
> >> >>
> >> >> Yes, that is because of the following line:
> >> >>
> >> >> q->bio_split = bioset_create(BIO_POOL_SIZE, 0);
> >> >>
> >> >> in blk_alloc_queue_node().
> >> >>
> >> >> Looks like this bio_set doesn't need to be per-request_queue, and
> >> >> now it is only used for fast-cloning bio for splitting, and one global
> >> >> split bio_set should be enough.
> >> >
> >> > It does have to be per request queue for stacking block devices (which includes
> >> > loopback).
> >>
> >> In commit df2cb6daa4(block: Avoid deadlocks with bio allocation by
> >> stacking drivers), deadlock in this situation has been avoided already.
> >> Or are there other issues with global bio_set? I appreciate if you may
> >> explain it a bit if there are.
> >
> > Even with commit df2cb6daa4 there is still risk of deadlocks (even
> > without low memory condition), see:
> > https://patchwork.kernel.org/patch/7398411/
> 
> That is definitely another problem which isn't related with low memory,
> and I guess Kent means there might be deadlock risk in case of shared
> bio_set.
> 
> >
> > (you may recall you blocked this patch with concerns about performance,
> > context switches, plug merging being compromised, etc.. to which I never
> > circled back to verify your concerns)
> 
> I still remember that problem:
> 
> 1) Process A
>      - two bio(a, b) are splitted in dm's make_request funtion
>      - bio(a) is submitted via generic_make_request(), so it is staged
>        in current->bio_list
>      - time t1
>      - before bio(b) is submitted, down_write(&s->lock) is run and
>       never return
> 
> 2) Process B:
>      - just during time t1, wait completion of bio(a) by down_write(&s->lock)
> 
> Then Process A waits the lock which is acquired by B first, and the
> two bio(a, b)
> can't reach to driver/device at all.
> 
> Looks that current->bio_list is fragile to locks from make_request function,
> and moving the lock into workqueue context should be helpful.
> 
> And I am happy to continue to discuss this issue further.
> 
> >
> > But it illustrates the type of problems that can occur when your rescue
> > infrastructure is shared across devices (in the context of df2cb6daa4,
> > current->bio_list contains bios from multiple devices).
> >
> > If a single splitting bio_set were shared across devices there would be
> > no guarantee of forward progress with complex stacked devices (one or
> > more devices could exhaust the reserve and starve out other devices in
> > the stack).  So keeping the bio_set per request_queue isn't prone to
> > failure like a shared bio_set might be.
> 
> Not consider the dm lock problem, from Kent's commit(df2cb6daa4) log and
> the patch, looks forward progress can be guaranteed for stacked devices
> with same bio_set, but better to get Kent's clarification.
> 
> If forward progress can be guaranteed, percpu mempool might avoid
> easy exhausting, because it is reasonable to assume that one CPU can only
> provide a certain amount of bandwidth wrt. block transfer.

Generally speaking, with potential deadlocks like this I don't bother to work
out the specific scenario, it's enough to know that there's a shared resource
and multiple users that depend on each other... if you've got that, you'll have
a deadlock.

But, if you're curious: say we've got block devices a and b, when you submit to
a the bio will get passed down to b:

for the bioset itself: if a bio gets split when submitted to a, then needs to be
split again when it's submitted to b - you're allocating twice from the same
mempool, and the first allocation can't be freed until the original bio
completes. deadlock.

with the rescuer threads it's more subtle, but you just need a scenario where
the rescuer is required twice in a row. I'm not going to bother trying to work
out the details, but it's the same principle - you can end up in a situation
where you're blocked, and you need the rescuer thread to make forward progress
(or you'd deadlock - that's why it exists, right?) - well, what happens if that
happens twice in a row, and the second time you're running out of the rescuer
thread? oops.

[toc] | [prev] | [next] | [standalone]


#1341021

FromPavel Machek <pavel@ucw.cz>
Date2016-02-23 21:50 +0100
Message-ID<r5s4b-2dy-23@gated-at.bofh.it>
In reply to#1340046
On Mon 2016-02-22 13:58:18, Kent Overstreet wrote:
> On Sun, Feb 21, 2016 at 05:40:59PM +0800, Ming Lei wrote:
> > On Sun, Feb 21, 2016 at 2:43 PM, Ming Lin-SSI <ming.l@ssi.samsung.com> wrote:
> > >>-----Original Message-----
> > >
> > > So it's almost already "per request_queue"
> > 
> > Yes, that is because of the following line:
> > 
> > q->bio_split = bioset_create(BIO_POOL_SIZE, 0);
> > 
> > in blk_alloc_queue_node().
> > 
> > Looks like this bio_set doesn't need to be per-request_queue, and
> > now it is only used for fast-cloning bio for splitting, and one global
> > split bio_set should be enough.
> 
> It does have to be per request queue for stacking block devices (which includes
> loopback).

Could we only allocate request queues for devices that are not even
opened? I have these in my system:

loop0  loop2  loop4  loop6  md0   nbd1	 nbd11	nbd13  nbd15  nbd3 nbd5  nbd7    nbd9  sda1  sda3
loop1  loop3  loop5  loop7  nbd0  nbd10  nbd12	nbd14  nbd2   nbd4 nbd6  nbd8    sda   sda2  sda4

...but nbd is never used, loop1+ is never used, and loop0 is only used
once in a blue moon. Each process takes 8K+...

								Pavel
-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web