Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1232360 > unrolled thread

Re: fuse scalability part 1

Started byAshish Samant <ashish.samant@oracle.com>
First post2015-09-24 21:20 +0200
Last post2015-09-29 08:20 +0200
Articles 4 — 3 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: fuse scalability part 1 Ashish Samant <ashish.samant@oracle.com> - 2015-09-24 21:20 +0200
    Re: fuse scalability part 1 Miklos Szeredi <miklos@szeredi.hu> - 2015-09-25 14:20 +0200
      Re: fuse scalability part 1 Ashish Samant <ashish.samant@oracle.com> - 2015-09-25 20:00 +0200
      Re: fuse scalability part 1 Srinivas Eeda <srinivas.eeda@oracle.com> - 2015-09-29 08:20 +0200

#1232360 — Re: fuse scalability part 1

FromAshish Samant <ashish.samant@oracle.com>
Date2015-09-24 21:20 +0200
SubjectRe: fuse scalability part 1
Message-ID<qckdI-3Kz-15@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

On 05/18/2015 08:13 AM, Miklos Szeredi wrote:
> This part splits out an "input queue" and a "processing queue" from the
> monolithic "fuse connection", each of those having their own spinlock.
>
> The end of the patchset adds the ability to "clone" a fuse connection.  This
> means, that instead of having to read/write requests/answers on a single fuse
> device fd, the fuse daemon can have multiple distinct file descriptors open.
> Each of those can be used to receive requests and send answers, currently the
> only constraint is that a request must be answered on the same fd as it was read
> from.
>
> This can be extended further to allow binding a device clone to a specific CPU
> or NUMA node.
>
> Patchset is available here:
>
>    git://git.kernel.org/pub/scm/linux/kernel/git/mszeredi/fuse.git for-next
>
> Libfuse patches adding support for "clone_fd" option:
>
>    git://git.code.sf.net/p/fuse/fuse clone_fd
>
> Thanks,
> Miklos
>
>
Resending the numbers as attachments because my email client messes the 
formatting of the message. Sorry for the noise.

We did some performance testing without these patches and with these 
patches (with -o clone_fd  option specified). We did 2 types of tests:

1. Throughput test : We did some parallel dd tests to read/write to FUSE 
based database fs on a system with 8 numa nodes and 288 cpus. The 
performance here is almost equal to the the per-numa patches we 
submitted a while back.Please find results attached.

2. Spinlock access times test: We also ran some tests within the kernel 
to check the time spent in accessing the spinlocks per request in both 
cases. As can be seen, the time taken per request to access the spinlock 
in the kernel code throughout the lifetime of the request is 30X to 100X 
better in the 2nd case (with patchset). Please find results attached.

Thanks,
Ashish


[toc] | [next] | [standalone]


#1232750

FromMiklos Szeredi <miklos@szeredi.hu>
Date2015-09-25 14:20 +0200
Message-ID<qcA8P-14l-25@gated-at.bofh.it>
In reply to#1232360
On Thu, Sep 24, 2015 at 9:17 PM, Ashish Samant <ashish.samant@oracle.com> wrote:

> We did some performance testing without these patches and with these patches
> (with -o clone_fd  option specified). We did 2 types of tests:
>
> 1. Throughput test : We did some parallel dd tests to read/write to FUSE
> based database fs on a system with 8 numa nodes and 288 cpus. The
> performance here is almost equal to the the per-numa patches we submitted a
> while back.Please find results attached.

Interesting.  This means, that serving the request on a different NUMA
node as the one where the request originated doesn't appear to make
the performance much worse.

Thanks,
Miklos
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1232976

FromAshish Samant <ashish.samant@oracle.com>
Date2015-09-25 20:00 +0200
Message-ID<qcFrP-8B-5@gated-at.bofh.it>
In reply to#1232750
On 09/25/2015 05:11 AM, Miklos Szeredi wrote:
> On Thu, Sep 24, 2015 at 9:17 PM, Ashish Samant <ashish.samant@oracle.com> wrote:
>
>> We did some performance testing without these patches and with these patches
>> (with -o clone_fd  option specified). We did 2 types of tests:
>>
>> 1. Throughput test : We did some parallel dd tests to read/write to FUSE
>> based database fs on a system with 8 numa nodes and 288 cpus. The
>> performance here is almost equal to the the per-numa patches we submitted a
>> while back.Please find results attached.
> Interesting.  This means, that serving the request on a different NUMA
> node as the one where the request originated doesn't appear to make
> the performance much worse.
>
> Thanks,
> Miklos
Yes. The main performance gain is due to the reduced contention on one 
spinlock(fc->lock) , especially with a large number of requests.
Splitting fc->fiq per cloned device will definitely improve performance 
further and we can  experiment further with per numa / cpu cloned device.

Thanks,
Ashish

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1234733

FromSrinivas Eeda <srinivas.eeda@oracle.com>
Date2015-09-29 08:20 +0200
Message-ID<qdWqC-6Bt-1@gated-at.bofh.it>
In reply to#1232750
Hi Miklos,

On 09/25/2015 05:11 AM, Miklos Szeredi wrote:
> On Thu, Sep 24, 2015 at 9:17 PM, Ashish Samant <ashish.samant@oracle.com> wrote:
>
>> We did some performance testing without these patches and with these patches
>> (with -o clone_fd  option specified). We did 2 types of tests:
>>
>> 1. Throughput test : We did some parallel dd tests to read/write to FUSE
>> based database fs on a system with 8 numa nodes and 288 cpus. The
>> performance here is almost equal to the the per-numa patches we submitted a
>> while back.Please find results attached.
> Interesting.  This means, that serving the request on a different NUMA
> node as the one where the request originated doesn't appear to make
> the performance much worse.
with the new change, contention of spinlock is significantly reduced, 
hence the latency caused by NUMA is not visible. Even in earlier case, 
the scalability was not a big problem if we bind all processes(fuse 
worker and user (dd threads)) to a single NUMA node. The problem was 
only seen when threads spread out across numa nodes and contend for the 
spin lock.


>
> Thanks,
> Miklos

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web