Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1463519 > unrolled thread

Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

Started byXin Long <lucien.xin@gmail.com>
First post2016-08-16 10:10 +0200
Last post2016-08-16 20:40 +0200
Articles 20 on this page of 27 — 3 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression Xin Long <lucien.xin@gmail.com> - 2016-08-16 10:10 +0200
    Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-16 10:40 +0200
    Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-16 11:00 +0200
      Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression Xin Long <lucien.xin@gmail.com> - 2016-08-16 12:00 +0200
        Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-17 07:10 +0200
          Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-17 07:40 +0200
            Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression Xin Long <lucien.xin@gmail.com> - 2016-08-17 07:50 +0200
              Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-17 08:20 +0200
                Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-17 08:40 +0200
                  Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-17 08:50 +0200
                  Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression Xin Long <lucien.xin@gmail.com> - 2016-08-17 09:40 +0200
                    Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-17 09:50 +0200
                      Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-17 10:00 +0200
                      Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression Xin Long <lucien.xin@gmail.com> - 2016-08-17 10:10 +0200
                        Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-17 10:50 +0200
                          Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression Xin Long <lucien.xin@gmail.com> - 2016-08-17 11:10 +0200
                            Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-17 11:30 +0200
                              Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression Xin Long <lucien.xin@gmail.com> - 2016-08-17 20:10 +0200
                                Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-18 05:30 +0200
                                  Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression Xin Long <lucien.xin@gmail.com> - 2016-08-18 14:50 +0200
                                    Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-19 07:30 +0200
                                      Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Marcelo Ricardo Leitner <marcelo.leitner@gmail.com> - 2016-08-19 09:30 +0200
                                        Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-19 09:30 +0200
                                          Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Marcelo Ricardo Leitner <marcelo.leitner@gmail.com> - 2016-08-22 23:50 +0200
                                            Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2%  regression Aaron Lu <aaron.lu@intel.com> - 2016-08-23 11:30 +0200
          Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression Xin Long <lucien.xin@gmail.com> - 2016-08-17 07:40 +0200
      Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression Xin Long <lucien.xin@gmail.com> - 2016-08-16 20:40 +0200

Page 1 of 2  [1] 2  Next page →


#1463519 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromXin Long <lucien.xin@gmail.com>
Date2016-08-16 10:10 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s6HBE-ym-15@gated-at.bofh.it>
>
> I'm testing on Linus' master, can we all use that please?
>

[git] git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git

[mechine]
Intel(R) Xeon(R) CPU E5-2690 v2 @ 3.00GHz
mem 62G (66000220K)

[system]
# cat /etc/redhat-release
Red Hat Enterprise Linux Server release 7.3 Beta (Maipo)

[commit 3684b03]
[root@hp-dl380pg8-11 lxin]# uname -r
4.8.0-rc2.3684b03
[root@hp-dl380pg8-11 lxin]# cat test.sh
killall -0 netserver || netserver -4 &
netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1
[root@hp-dl380pg8-11 lxin]# sh test.sh
SCTP 1-TO-MANY STREAM TEST from 0.0.0.0 (0.0.0.0) port 0 AF_INET to
127.0.0.1 () port 0 AF_INET
Recv   Send    Send                          Utilization       Service Demand
Socket Socket  Message  Elapsed              Send     Recv     Send    Recv
Size   Size    Size     Time     Throughput  local    remote   local   remote
bytes  bytes   bytes    secs.    10^6bits/s  % S      % S      us/KB   us/KB

212992 212992  10240    300.00     16914.99   3.28     3.28     0.636   0.636

[commit f959fb4]
[root@localhost lxin]# uname -r
4.7.0-rc6.f959fb4
[root@localhost lxin]# cat test.sh
killall -0 netserver || netserver -4 &
netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1
[root@localhost lxin]# sh test.sh
SCTP 1-TO-MANY STREAM TEST from 0.0.0.0 (0.0.0.0) port 0 AF_INET to
127.0.0.1 () port 0 AF_INET
Recv   Send    Send                          Utilization       Service Demand
Socket Socket  Message  Elapsed              Send     Recv     Send    Recv
Size   Size    Size     Time     Throughput  local    remote   local   remote
bytes  bytes   bytes    secs.    10^6bits/s  % S      % S      us/KB   us/KB

212992 212992  10240    300.00     12975.32   3.35     3.35     0.847   0.846


Still, in my env, the latest kernel is better than old one.
Sorry, I'm not sure why it's so different in your env.

Could you do 'netperf' test manually, instead of lkp-tests, then check again.
Pls show you system's distros as well, like rhel, ubuntu or arch ?

[toc] | [next] | [standalone]


#1463539 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-16 10:40 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s6I4F-IW-3@gated-at.bofh.it>
In reply to#1463519
On 08/16/2016 04:02 PM, Xin Long wrote:
>>
>> I'm testing on Linus' master, can we all use that please?
>>
> 
> [git] git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
> 
> [mechine]
> Intel(R) Xeon(R) CPU E5-2690 v2 @ 3.00GHz
> mem 62G (66000220K)
> 
> [system]
> # cat /etc/redhat-release
> Red Hat Enterprise Linux Server release 7.3 Beta (Maipo)
> 
> [commit 3684b03]
> [root@hp-dl380pg8-11 lxin]# uname -r
> 4.8.0-rc2.3684b03
> [root@hp-dl380pg8-11 lxin]# cat test.sh
> killall -0 netserver || netserver -4 &
> netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1
> [root@hp-dl380pg8-11 lxin]# sh test.sh
> SCTP 1-TO-MANY STREAM TEST from 0.0.0.0 (0.0.0.0) port 0 AF_INET to
> 127.0.0.1 () port 0 AF_INET
> Recv   Send    Send                          Utilization       Service Demand
> Socket Socket  Message  Elapsed              Send     Recv     Send    Recv
> Size   Size    Size     Time     Throughput  local    remote   local   remote
> bytes  bytes   bytes    secs.    10^6bits/s  % S      % S      us/KB   us/KB
> 
> 212992 212992  10240    300.00     16914.99   3.28     3.28     0.636   0.636
> 
> [commit f959fb4]
> [root@localhost lxin]# uname -r
> 4.7.0-rc6.f959fb4
> [root@localhost lxin]# cat test.sh
> killall -0 netserver || netserver -4 &
> netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1
> [root@localhost lxin]# sh test.sh
> SCTP 1-TO-MANY STREAM TEST from 0.0.0.0 (0.0.0.0) port 0 AF_INET to
> 127.0.0.1 () port 0 AF_INET
> Recv   Send    Send                          Utilization       Service Demand
> Socket Socket  Message  Elapsed              Send     Recv     Send    Recv
> Size   Size    Size     Time     Throughput  local    remote   local   remote
> bytes  bytes   bytes    secs.    10^6bits/s  % S      % S      us/KB   us/KB
> 
> 212992 212992  10240    300.00     12975.32   3.35     3.35     0.847   0.846
> 
> 
> Still, in my env, the latest kernel is better than old one.
> Sorry, I'm not sure why it's so different in your env.

Could the test have anything to do with the hardware? i.e. yours is Xeon
E5-2690 while mine is IVB i3?

> 
> Could you do 'netperf' test manually, instead of lkp-tests, then check again.

Manually test under LKP is not easy as those test machines are all doing
things automatically. But if you think that is necessary, I can do that.

> Pls show you system's distros as well, like rhel, ubuntu or arch ?

We do not use any of these distros.
The rootfs is derived from debian:
https://github.com/fengguang/reproduce-kernel-bug/blob/master/debian/debian-x86_64-2015-02-07.cgz

Thanks,
Aaron

[toc] | [prev] | [next] | [standalone]


#1463553 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-16 11:00 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s6Io1-R5-5@gated-at.bofh.it>
In reply to#1463519
On 08/16/2016 04:02 PM, Xin Long wrote:
>>
>> I'm testing on Linus' master, can we all use that please?
>>
> 
> [git] git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
> 
> [mechine]
> Intel(R) Xeon(R) CPU E5-2690 v2 @ 3.00GHz
> mem 62G (66000220K)
> 
> [system]
> # cat /etc/redhat-release
> Red Hat Enterprise Linux Server release 7.3 Beta (Maipo)
> 
> [commit 3684b03]
> [root@hp-dl380pg8-11 lxin]# uname -r
> 4.8.0-rc2.3684b03
> [root@hp-dl380pg8-11 lxin]# cat test.sh
> killall -0 netserver || netserver -4 &
> netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1

I just realized the test we are doing is not exactly the same.
As the original report says:
        ip: ipv4
        runtime: 300s
        nr_threads: 200%
        cluster: cs-localhost
        send_size: 10K
        test: SCTP_STREAM_MANY
        cpufreq_governor: performance

Note the nr_threads: 200%, which means to start 2 times of CPU number
processes of netperf.

In our IVB i3(2 cores, 2 threads per core) case, 8 netperf processes
are started concurrently:

2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &

The throughput is the average of those runs.

And I think we should be doing test on:
commit a6c2f79287 ("sctp: implement prsctp TTL policy") (the bisected one)
and
commit 826d253d57 ("sctp: add SCTP_PR_ASSOC_STATUS on sctp sockopt") (its immediate parent)
instead of Linus' master HEAD to avoid other factors.

Thanks,
Aaron

[toc] | [prev] | [next] | [standalone]


#1463634

FromXin Long <lucien.xin@gmail.com>
Date2016-08-16 12:00 +0200
Message-ID<s6Jk7-1sx-59@gated-at.bofh.it>
In reply to#1463553
>>>
>>> I'm testing on Linus' master, can we all use that please?
>>>
>>
>> [git] git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
>>
>> [mechine]
>> Intel(R) Xeon(R) CPU E5-2690 v2 @ 3.00GHz
>> mem 62G (66000220K)
>>
>> [system]
>> # cat /etc/redhat-release
>> Red Hat Enterprise Linux Server release 7.3 Beta (Maipo)
>>
>> [commit 3684b03]
>> [root@hp-dl380pg8-11 lxin]# uname -r
>> 4.8.0-rc2.3684b03
>> [root@hp-dl380pg8-11 lxin]# cat test.sh
>> killall -0 netserver || netserver -4 &
>> netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1
>
> I just realized the test we are doing is not exactly the same.
> As the original report says:
>         ip: ipv4
>         runtime: 300s
>         nr_threads: 200%
>         cluster: cs-localhost
>         send_size: 10K
>         test: SCTP_STREAM_MANY
>         cpufreq_governor: performance
>
> Note the nr_threads: 200%, which means to start 2 times of CPU number
> processes of netperf.
>
> In our IVB i3(2 cores, 2 threads per core) case, 8 netperf processes
> are started concurrently:
OK, understand.

>
> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>
> The throughput is the average of those runs.
>
> And I think we should be doing test on:
> commit a6c2f79287 ("sctp: implement prsctp TTL policy") (the bisected one)
> and
> commit 826d253d57 ("sctp: add SCTP_PR_ASSOC_STATUS on sctp sockopt") (its immediate parent)
> instead of Linus' master HEAD to avoid other factors.
>
OK, I will do tests as your suggestion now,  but need to rebuild again :D

can you disable pr_enable with "sysctl -w net.sctp.prsctp_enable=0",
then try again?

[toc] | [prev] | [next] | [standalone]


#1464315 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-17 07:10 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s71gZ-4UF-11@gated-at.bofh.it>
In reply to#1463634
On 08/16/2016 05:56 PM, Xin Long wrote:
>>>>
>>>> I'm testing on Linus' master, can we all use that please?
>>>>
>>>
>>> [git] git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
>>>
>>> [mechine]
>>> Intel(R) Xeon(R) CPU E5-2690 v2 @ 3.00GHz
>>> mem 62G (66000220K)
>>>
>>> [system]
>>> # cat /etc/redhat-release
>>> Red Hat Enterprise Linux Server release 7.3 Beta (Maipo)
>>>
>>> [commit 3684b03]
>>> [root@hp-dl380pg8-11 lxin]# uname -r
>>> 4.8.0-rc2.3684b03
>>> [root@hp-dl380pg8-11 lxin]# cat test.sh
>>> killall -0 netserver || netserver -4 &
>>> netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1
>>
>> I just realized the test we are doing is not exactly the same.
>> As the original report says:
>>         ip: ipv4
>>         runtime: 300s
>>         nr_threads: 200%
>>         cluster: cs-localhost
>>         send_size: 10K
>>         test: SCTP_STREAM_MANY
>>         cpufreq_governor: performance
>>
>> Note the nr_threads: 200%, which means to start 2 times of CPU number
>> processes of netperf.
>>
>> In our IVB i3(2 cores, 2 threads per core) case, 8 netperf processes
>> are started concurrently:
> OK, understand.
> 
>>
>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>>
>> The throughput is the average of those runs.
>>
>> And I think we should be doing test on:
>> commit a6c2f79287 ("sctp: implement prsctp TTL policy") (the bisected one)
>> and
>> commit 826d253d57 ("sctp: add SCTP_PR_ASSOC_STATUS on sctp sockopt") (its immediate parent)
>> instead of Linus' master HEAD to avoid other factors.
>>
> OK, I will do tests as your suggestion now,  but need to rebuild again :D
> 
> can you disable pr_enable with "sysctl -w net.sctp.prsctp_enable=0",
> then try again?

For commit a6c2f79287 ("sctp: implement prsctp TTL policy"), no matter
the value of net.sctp.prsctp_enable, the throughput is almost the same:

net.sctp.prsctp_enable = 0
{
  "netperf.Throughput_Mbps": [
    2353.3112499999997
  ]
}

net.sctp.prsctp_enable = 1
{
  "netperf.Throughput_Mbps": [
    2371.5862500000003
  ]
}

For its immediate parent:
commit 826d253d57 ("sctp: add SCTP_PR_ASSOC_STATUS on sctp sockopt")
No matter the value of net.sctp.prsctp_enable, the throughput is again
almost the same:

net.sctp.prsctp_enable = 0
{
  "netperf.Throughput_Mbps": [
    3838.8300000000004
  ]
}

net.sctp.prsctp_enable = 1
{
  "netperf.Throughput_Mbps": [
    3751.4600000000005
  ]
}

Does this result give any hint?

Thanks,
Aaron

[toc] | [prev] | [next] | [standalone]


#1464326 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-17 07:40 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s71K1-57o-3@gated-at.bofh.it>
In reply to#1464315

[Multipart message — attachments visible in raw view] — view raw

On 08/17/2016 01:04 PM, Aaron Lu wrote:
> On 08/16/2016 05:56 PM, Xin Long wrote:
>>>>>
>>>>> I'm testing on Linus' master, can we all use that please?
>>>>>
>>>>
>>>> [git] git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
>>>>
>>>> [mechine]
>>>> Intel(R) Xeon(R) CPU E5-2690 v2 @ 3.00GHz
>>>> mem 62G (66000220K)
>>>>
>>>> [system]
>>>> # cat /etc/redhat-release
>>>> Red Hat Enterprise Linux Server release 7.3 Beta (Maipo)
>>>>
>>>> [commit 3684b03]
>>>> [root@hp-dl380pg8-11 lxin]# uname -r
>>>> 4.8.0-rc2.3684b03
>>>> [root@hp-dl380pg8-11 lxin]# cat test.sh
>>>> killall -0 netserver || netserver -4 &
>>>> netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1
>>>
>>> I just realized the test we are doing is not exactly the same.
>>> As the original report says:
>>>         ip: ipv4
>>>         runtime: 300s
>>>         nr_threads: 200%
>>>         cluster: cs-localhost
>>>         send_size: 10K
>>>         test: SCTP_STREAM_MANY
>>>         cpufreq_governor: performance
>>>
>>> Note the nr_threads: 200%, which means to start 2 times of CPU number
>>> processes of netperf.
>>>
>>> In our IVB i3(2 cores, 2 threads per core) case, 8 netperf processes
>>> are started concurrently:
>> OK, understand.
>>
>>>
>>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>>> 2016-07-27 03:48:09 netperf -4 -t SCTP_STREAM_MANY -c -C -l 300 -- -m 10K -H 127.0.0.1 &
>>>
>>> The throughput is the average of those runs.
>>>
>>> And I think we should be doing test on:
>>> commit a6c2f79287 ("sctp: implement prsctp TTL policy") (the bisected one)
>>> and
>>> commit 826d253d57 ("sctp: add SCTP_PR_ASSOC_STATUS on sctp sockopt") (its immediate parent)
>>> instead of Linus' master HEAD to avoid other factors.
>>>
>> OK, I will do tests as your suggestion now,  but need to rebuild again :D
>>
>> can you disable pr_enable with "sysctl -w net.sctp.prsctp_enable=0",
>> then try again?
> 
> For commit a6c2f79287 ("sctp: implement prsctp TTL policy"), no matter
> the value of net.sctp.prsctp_enable, the throughput is almost the same:

The perf-profile data for the two commits are attached(for the case of
prsctp_enable=1, the perf-profile data doesn't get collected for the 0
case for some reason, I'm checking the problem now).

The CPU gets much more idle time in the bisected commit a6c2f79287:

    68.89%     0.70%  [kernel.kallsyms]   [k] entry_SYSCALL_64_fastpath
    49.32%     0.12%  [kernel.kallsyms]   [k] sys_sendmsg
    49.17%     0.12%  [kernel.kallsyms]   [k] __sys_sendmsg
    48.58%     0.22%  [kernel.kallsyms]   [k] ___sys_sendmsg
    46.69%     0.06%  [kernel.kallsyms]   [k] sock_sendmsg
    46.31%     0.16%  [kernel.kallsyms]   [k] inet_sendmsg
    45.90%     0.98%  [kernel.kallsyms]   [k] sctp_sendmsg
    29.66%     0.45%  [kernel.kallsyms]   [k] sctp_do_sm
    29.54%     0.23%  [kernel.kallsyms]   [k] cpu_startup_entry
    28.81%     0.68%  [kernel.kallsyms]   [k] sctp_cmd_interpreter.isra.24
    26.20%     0.00%  [kernel.kallsyms]   [k] start_secondary
    23.04%     0.09%  [kernel.kallsyms]   [k] sctp_inq_push
    23.03%     0.08%  [kernel.kallsyms]   [k] call_cpuidle
    22.94%     0.00%  [kernel.kallsyms]   [k] cpuidle_enter
    22.60%     0.18%  [kernel.kallsyms]   [k] cpuidle_enter_state
    21.99%    21.99%  [kernel.kallsyms]   [k] intel_idle
... ...

While its immediate parent commit 826d253d57 is mostly busy working:

    98.53%     0.83%  [kernel.kallsyms]   [k] entry_SYSCALL_64_fastpath
    78.13%     0.12%  [kernel.kallsyms]   [k] sys_sendmsg
    78.03%     0.16%  [kernel.kallsyms]   [k] __sys_sendmsg
    77.08%     0.28%  [kernel.kallsyms]   [k] ___sys_sendmsg
    74.44%     0.08%  [kernel.kallsyms]   [k] sock_sendmsg
    73.82%     0.13%  [kernel.kallsyms]   [k] inet_sendmsg
    73.34%     1.44%  [kernel.kallsyms]   [k] sctp_sendmsg
    47.52%     0.75%  [kernel.kallsyms]   [k] sctp_do_sm
    46.19%     0.90%  [kernel.kallsyms]   [k] sctp_cmd_interpreter.isra.24
    37.17%     1.43%  [kernel.kallsyms]   [k] sctp_outq_flush
    36.93%     0.08%  [kernel.kallsyms]   [k] sctp_outq_uncork
    34.24%     0.15%  [kernel.kallsyms]   [k] sctp_inq_push
... ...
No idle related function above 1%.

Will the bisected commit make the idle possible?

Thanks,
Aaron

[toc] | [prev] | [next] | [standalone]


#1464332

FromXin Long <lucien.xin@gmail.com>
Date2016-08-17 07:50 +0200
Message-ID<s71TI-5aR-13@gated-at.bofh.it>
In reply to#1464326
> The perf-profile data for the two commits are attached(for the case of
> prsctp_enable=1, the perf-profile data doesn't get collected for the 0
> case for some reason, I'm checking the problem now).
>
> The CPU gets much more idle time in the bisected commit a6c2f79287:
>
>     68.89%     0.70%  [kernel.kallsyms]   [k] entry_SYSCALL_64_fastpath
>     49.32%     0.12%  [kernel.kallsyms]   [k] sys_sendmsg
>     49.17%     0.12%  [kernel.kallsyms]   [k] __sys_sendmsg
>     48.58%     0.22%  [kernel.kallsyms]   [k] ___sys_sendmsg
>     46.69%     0.06%  [kernel.kallsyms]   [k] sock_sendmsg
>     46.31%     0.16%  [kernel.kallsyms]   [k] inet_sendmsg
>     45.90%     0.98%  [kernel.kallsyms]   [k] sctp_sendmsg
>     29.66%     0.45%  [kernel.kallsyms]   [k] sctp_do_sm
>     29.54%     0.23%  [kernel.kallsyms]   [k] cpu_startup_entry
>     28.81%     0.68%  [kernel.kallsyms]   [k] sctp_cmd_interpreter.isra.24
>     26.20%     0.00%  [kernel.kallsyms]   [k] start_secondary
>     23.04%     0.09%  [kernel.kallsyms]   [k] sctp_inq_push
>     23.03%     0.08%  [kernel.kallsyms]   [k] call_cpuidle
>     22.94%     0.00%  [kernel.kallsyms]   [k] cpuidle_enter
>     22.60%     0.18%  [kernel.kallsyms]   [k] cpuidle_enter_state
>     21.99%    21.99%  [kernel.kallsyms]   [k] intel_idle
> ... ...
>
> While its immediate parent commit 826d253d57 is mostly busy working:
>
>     98.53%     0.83%  [kernel.kallsyms]   [k] entry_SYSCALL_64_fastpath
>     78.13%     0.12%  [kernel.kallsyms]   [k] sys_sendmsg
>     78.03%     0.16%  [kernel.kallsyms]   [k] __sys_sendmsg
>     77.08%     0.28%  [kernel.kallsyms]   [k] ___sys_sendmsg
>     74.44%     0.08%  [kernel.kallsyms]   [k] sock_sendmsg
>     73.82%     0.13%  [kernel.kallsyms]   [k] inet_sendmsg
>     73.34%     1.44%  [kernel.kallsyms]   [k] sctp_sendmsg
>     47.52%     0.75%  [kernel.kallsyms]   [k] sctp_do_sm
>     46.19%     0.90%  [kernel.kallsyms]   [k] sctp_cmd_interpreter.isra.24
>     37.17%     1.43%  [kernel.kallsyms]   [k] sctp_outq_flush
>     36.93%     0.08%  [kernel.kallsyms]   [k] sctp_outq_uncork
>     34.24%     0.15%  [kernel.kallsyms]   [k] sctp_inq_push
> ... ...
> No idle related function above 1%.
>
> Will the bisected commit make the idle possible?
No, not at all. :)

pls help to debug as I said in the last reply.

[toc] | [prev] | [next] | [standalone]


#1464335 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-17 08:20 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s72mJ-5By-5@gated-at.bofh.it>
In reply to#1464332
On Wed, Aug 17, 2016 at 01:41:04PM +0800, Xin Long wrote:
> > The perf-profile data for the two commits are attached(for the case of
> > prsctp_enable=1, the perf-profile data doesn't get collected for the 0
> > case for some reason, I'm checking the problem now).
> >
> > The CPU gets much more idle time in the bisected commit a6c2f79287:
> >
> >     68.89%     0.70%  [kernel.kallsyms]   [k] entry_SYSCALL_64_fastpath
> >     49.32%     0.12%  [kernel.kallsyms]   [k] sys_sendmsg
> >     49.17%     0.12%  [kernel.kallsyms]   [k] __sys_sendmsg
> >     48.58%     0.22%  [kernel.kallsyms]   [k] ___sys_sendmsg
> >     46.69%     0.06%  [kernel.kallsyms]   [k] sock_sendmsg
> >     46.31%     0.16%  [kernel.kallsyms]   [k] inet_sendmsg
> >     45.90%     0.98%  [kernel.kallsyms]   [k] sctp_sendmsg
> >     29.66%     0.45%  [kernel.kallsyms]   [k] sctp_do_sm
> >     29.54%     0.23%  [kernel.kallsyms]   [k] cpu_startup_entry
> >     28.81%     0.68%  [kernel.kallsyms]   [k] sctp_cmd_interpreter.isra.24
> >     26.20%     0.00%  [kernel.kallsyms]   [k] start_secondary
> >     23.04%     0.09%  [kernel.kallsyms]   [k] sctp_inq_push
> >     23.03%     0.08%  [kernel.kallsyms]   [k] call_cpuidle
> >     22.94%     0.00%  [kernel.kallsyms]   [k] cpuidle_enter
> >     22.60%     0.18%  [kernel.kallsyms]   [k] cpuidle_enter_state
> >     21.99%    21.99%  [kernel.kallsyms]   [k] intel_idle
> > ... ...
> >
> > While its immediate parent commit 826d253d57 is mostly busy working:
> >
> >     98.53%     0.83%  [kernel.kallsyms]   [k] entry_SYSCALL_64_fastpath
> >     78.13%     0.12%  [kernel.kallsyms]   [k] sys_sendmsg
> >     78.03%     0.16%  [kernel.kallsyms]   [k] __sys_sendmsg
> >     77.08%     0.28%  [kernel.kallsyms]   [k] ___sys_sendmsg
> >     74.44%     0.08%  [kernel.kallsyms]   [k] sock_sendmsg
> >     73.82%     0.13%  [kernel.kallsyms]   [k] inet_sendmsg
> >     73.34%     1.44%  [kernel.kallsyms]   [k] sctp_sendmsg
> >     47.52%     0.75%  [kernel.kallsyms]   [k] sctp_do_sm
> >     46.19%     0.90%  [kernel.kallsyms]   [k] sctp_cmd_interpreter.isra.24
> >     37.17%     1.43%  [kernel.kallsyms]   [k] sctp_outq_flush
> >     36.93%     0.08%  [kernel.kallsyms]   [k] sctp_outq_uncork
> >     34.24%     0.15%  [kernel.kallsyms]   [k] sctp_inq_push
> > ... ...
> > No idle related function above 1%.
> >
> > Will the bisected commit make the idle possible?
> No, not at all. :)
> 
> pls help to debug as I said in the last reply.

OK, will see how to do that.

In the meantime, I just tried to reproduce on my own desktop:
Sandybridge i7-2600 CPU @ 3.40GHz and it reproduced:
$ cat 4.7.0-rc6-01198-ga6c2f792873a/0/netperf.json
{
  "netperf.Throughput_Mbps": [
   752.9450000000002
  ]
}
$ cat 4.7.0-rc6-01197-g826d253d57b1/0/netperf.json
{
  "netperf.Throughput_Mbps": [
   1068.5556249999997
  ]
}

[toc] | [prev] | [next] | [standalone]


#1464344 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-17 08:40 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s72G6-5IV-13@gated-at.bofh.it>
In reply to#1464335
On Wed, Aug 17, 2016 at 02:14:05PM +0800, Aaron Lu wrote:
> On Wed, Aug 17, 2016 at 01:41:04PM +0800, Xin Long wrote:
> > > The perf-profile data for the two commits are attached(for the case of
> > > prsctp_enable=1, the perf-profile data doesn't get collected for the 0
> > > case for some reason, I'm checking the problem now).
> > >
> > > The CPU gets much more idle time in the bisected commit a6c2f79287:
> > >
> > >     68.89%     0.70%  [kernel.kallsyms]   [k] entry_SYSCALL_64_fastpath
> > >     49.32%     0.12%  [kernel.kallsyms]   [k] sys_sendmsg
> > >     49.17%     0.12%  [kernel.kallsyms]   [k] __sys_sendmsg
> > >     48.58%     0.22%  [kernel.kallsyms]   [k] ___sys_sendmsg
> > >     46.69%     0.06%  [kernel.kallsyms]   [k] sock_sendmsg
> > >     46.31%     0.16%  [kernel.kallsyms]   [k] inet_sendmsg
> > >     45.90%     0.98%  [kernel.kallsyms]   [k] sctp_sendmsg
> > >     29.66%     0.45%  [kernel.kallsyms]   [k] sctp_do_sm
> > >     29.54%     0.23%  [kernel.kallsyms]   [k] cpu_startup_entry
> > >     28.81%     0.68%  [kernel.kallsyms]   [k] sctp_cmd_interpreter.isra.24
> > >     26.20%     0.00%  [kernel.kallsyms]   [k] start_secondary
> > >     23.04%     0.09%  [kernel.kallsyms]   [k] sctp_inq_push
> > >     23.03%     0.08%  [kernel.kallsyms]   [k] call_cpuidle
> > >     22.94%     0.00%  [kernel.kallsyms]   [k] cpuidle_enter
> > >     22.60%     0.18%  [kernel.kallsyms]   [k] cpuidle_enter_state
> > >     21.99%    21.99%  [kernel.kallsyms]   [k] intel_idle
> > > ... ...
> > >
> > > While its immediate parent commit 826d253d57 is mostly busy working:
> > >
> > >     98.53%     0.83%  [kernel.kallsyms]   [k] entry_SYSCALL_64_fastpath
> > >     78.13%     0.12%  [kernel.kallsyms]   [k] sys_sendmsg
> > >     78.03%     0.16%  [kernel.kallsyms]   [k] __sys_sendmsg
> > >     77.08%     0.28%  [kernel.kallsyms]   [k] ___sys_sendmsg
> > >     74.44%     0.08%  [kernel.kallsyms]   [k] sock_sendmsg
> > >     73.82%     0.13%  [kernel.kallsyms]   [k] inet_sendmsg
> > >     73.34%     1.44%  [kernel.kallsyms]   [k] sctp_sendmsg
> > >     47.52%     0.75%  [kernel.kallsyms]   [k] sctp_do_sm
> > >     46.19%     0.90%  [kernel.kallsyms]   [k] sctp_cmd_interpreter.isra.24
> > >     37.17%     1.43%  [kernel.kallsyms]   [k] sctp_outq_flush
> > >     36.93%     0.08%  [kernel.kallsyms]   [k] sctp_outq_uncork
> > >     34.24%     0.15%  [kernel.kallsyms]   [k] sctp_inq_push
> > > ... ...
> > > No idle related function above 1%.
> > >
> > > Will the bisected commit make the idle possible?
> > No, not at all. :)
> > 
> > pls help to debug as I said in the last reply.
> 
> OK, will see how to do that.
> 
> In the meantime, I just tried to reproduce on my own desktop:
> Sandybridge i7-2600 CPU @ 3.40GHz and it reproduced:
> $ cat 4.7.0-rc6-01198-ga6c2f792873a/0/netperf.json
> {
>   "netperf.Throughput_Mbps": [
>    752.9450000000002
>   ]
> }
> $ cat 4.7.0-rc6-01197-g826d253d57b1/0/netperf.json
> {
>   "netperf.Throughput_Mbps": [
>    1068.5556249999997
>   ]
> }

On top of
commit 826d253d57b1 ("sctp: add SCTP_PR_ASSOC_STATUS on sctp sockopt")
I applied the below commit:

From 98dd2532b14e29dcc2ab40a7348755531afa79e4 Mon Sep 17 00:00:00 2001
From: Aaron Lu <aaron.lu@intel.com>
Date: Wed, 17 Aug 2016 14:20:00 +0800
Subject: [PATCH] sctp: test

---
 include/net/sctp/structs.h | 3 +++
 1 file changed, 3 insertions(+)

diff --git a/include/net/sctp/structs.h b/include/net/sctp/structs.h
index d8e464aacb20..932f2780d3a4 100644
--- a/include/net/sctp/structs.h
+++ b/include/net/sctp/structs.h
@@ -602,6 +602,9 @@ struct sctp_chunk {
 	/* This needs to be recoverable for SCTP_SEND_FAILED events. */
 	struct sctp_sndrcvinfo sinfo;
 
+	unsigned long prsctp_param;
+	int sent_count;
+
 	/* Which association does this belong to?  */
 	struct sctp_association *asoc;
 
-- 
2.5.5

Then the performance dropped to the same as the bisected commit
a6c2f792873a:
$ cat 4.7.0-rc6-01198-g98dd2532b14e/0/netperf.json
{
  "netperf.Throughput_Mbps": [
   754.494375
  ]
}

I think this agrees with the perf data in that the newly added function
doesn't show up in the perf-profile but still, the performance drops.
So the only possible reason is the newly added fields to the sctp_chunk
structure.

Is this expected?

Thanks,
Aaron

[toc] | [prev] | [next] | [standalone]


#1464348 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-17 08:50 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s72PM-5MN-5@gated-at.bofh.it>
In reply to#1464344
On Wed, Aug 17, 2016 at 02:37:19PM +0800, Aaron Lu wrote:
> On Wed, Aug 17, 2016 at 02:14:05PM +0800, Aaron Lu wrote:
> > On Wed, Aug 17, 2016 at 01:41:04PM +0800, Xin Long wrote:
> > > > The perf-profile data for the two commits are attached(for the case of
> > > > prsctp_enable=1, the perf-profile data doesn't get collected for the 0
> > > > case for some reason, I'm checking the problem now).
> > > >
> > > > The CPU gets much more idle time in the bisected commit a6c2f79287:
> > > >
> > > >     68.89%     0.70%  [kernel.kallsyms]   [k] entry_SYSCALL_64_fastpath
> > > >     49.32%     0.12%  [kernel.kallsyms]   [k] sys_sendmsg
> > > >     49.17%     0.12%  [kernel.kallsyms]   [k] __sys_sendmsg
> > > >     48.58%     0.22%  [kernel.kallsyms]   [k] ___sys_sendmsg
> > > >     46.69%     0.06%  [kernel.kallsyms]   [k] sock_sendmsg
> > > >     46.31%     0.16%  [kernel.kallsyms]   [k] inet_sendmsg
> > > >     45.90%     0.98%  [kernel.kallsyms]   [k] sctp_sendmsg
> > > >     29.66%     0.45%  [kernel.kallsyms]   [k] sctp_do_sm
> > > >     29.54%     0.23%  [kernel.kallsyms]   [k] cpu_startup_entry
> > > >     28.81%     0.68%  [kernel.kallsyms]   [k] sctp_cmd_interpreter.isra.24
> > > >     26.20%     0.00%  [kernel.kallsyms]   [k] start_secondary
> > > >     23.04%     0.09%  [kernel.kallsyms]   [k] sctp_inq_push
> > > >     23.03%     0.08%  [kernel.kallsyms]   [k] call_cpuidle
> > > >     22.94%     0.00%  [kernel.kallsyms]   [k] cpuidle_enter
> > > >     22.60%     0.18%  [kernel.kallsyms]   [k] cpuidle_enter_state
> > > >     21.99%    21.99%  [kernel.kallsyms]   [k] intel_idle
> > > > ... ...
> > > >
> > > > While its immediate parent commit 826d253d57 is mostly busy working:
> > > >
> > > >     98.53%     0.83%  [kernel.kallsyms]   [k] entry_SYSCALL_64_fastpath
> > > >     78.13%     0.12%  [kernel.kallsyms]   [k] sys_sendmsg
> > > >     78.03%     0.16%  [kernel.kallsyms]   [k] __sys_sendmsg
> > > >     77.08%     0.28%  [kernel.kallsyms]   [k] ___sys_sendmsg
> > > >     74.44%     0.08%  [kernel.kallsyms]   [k] sock_sendmsg
> > > >     73.82%     0.13%  [kernel.kallsyms]   [k] inet_sendmsg
> > > >     73.34%     1.44%  [kernel.kallsyms]   [k] sctp_sendmsg
> > > >     47.52%     0.75%  [kernel.kallsyms]   [k] sctp_do_sm
> > > >     46.19%     0.90%  [kernel.kallsyms]   [k] sctp_cmd_interpreter.isra.24
> > > >     37.17%     1.43%  [kernel.kallsyms]   [k] sctp_outq_flush
> > > >     36.93%     0.08%  [kernel.kallsyms]   [k] sctp_outq_uncork
> > > >     34.24%     0.15%  [kernel.kallsyms]   [k] sctp_inq_push
> > > > ... ...
> > > > No idle related function above 1%.
> > > >
> > > > Will the bisected commit make the idle possible?
> > > No, not at all. :)
> > > 
> > > pls help to debug as I said in the last reply.
> > 
> > OK, will see how to do that.
> > 
> > In the meantime, I just tried to reproduce on my own desktop:
> > Sandybridge i7-2600 CPU @ 3.40GHz and it reproduced:
> > $ cat 4.7.0-rc6-01198-ga6c2f792873a/0/netperf.json
> > {
> >   "netperf.Throughput_Mbps": [
> >    752.9450000000002
> >   ]
> > }
> > $ cat 4.7.0-rc6-01197-g826d253d57b1/0/netperf.json
> > {
> >   "netperf.Throughput_Mbps": [
> >    1068.5556249999997
> >   ]
> > }
> 
> On top of
> commit 826d253d57b1 ("sctp: add SCTP_PR_ASSOC_STATUS on sctp sockopt")
> I applied the below commit:
> 
> From 98dd2532b14e29dcc2ab40a7348755531afa79e4 Mon Sep 17 00:00:00 2001
> From: Aaron Lu <aaron.lu@intel.com>
> Date: Wed, 17 Aug 2016 14:20:00 +0800
> Subject: [PATCH] sctp: test
> 
> ---
>  include/net/sctp/structs.h | 3 +++
>  1 file changed, 3 insertions(+)
> 
> diff --git a/include/net/sctp/structs.h b/include/net/sctp/structs.h
> index d8e464aacb20..932f2780d3a4 100644
> --- a/include/net/sctp/structs.h
> +++ b/include/net/sctp/structs.h
> @@ -602,6 +602,9 @@ struct sctp_chunk {
>  	/* This needs to be recoverable for SCTP_SEND_FAILED events. */
>  	struct sctp_sndrcvinfo sinfo;
>  
> +	unsigned long prsctp_param;
> +	int sent_count;
> +
>  	/* Which association does this belong to?  */
>  	struct sctp_association *asoc;
>  
> -- 
> 2.5.5
> 
> Then the performance dropped to the same as the bisected commit
> a6c2f792873a:
> $ cat 4.7.0-rc6-01198-g98dd2532b14e/0/netperf.json
> {
>   "netperf.Throughput_Mbps": [
>    754.494375
>   ]
> }
> 
> I think this agrees with the perf data in that the newly added function

Actually, I mean the modified functions like sctp_chunk_abandoned and
__sctp_packet_append_chunk, etc.

> doesn't show up in the perf-profile but still, the performance drops.
> So the only possible reason is the newly added fields to the sctp_chunk
> structure.
> 
> Is this expected?
> 
> Thanks,
> Aaron

[toc] | [prev] | [next] | [standalone]


#1464386

FromXin Long <lucien.xin@gmail.com>
Date2016-08-17 09:40 +0200
Message-ID<s73C9-6mv-9@gated-at.bofh.it>
In reply to#1464344
>  include/net/sctp/structs.h | 3 +++
>  1 file changed, 3 insertions(+)
>
> diff --git a/include/net/sctp/structs.h b/include/net/sctp/structs.h
> index d8e464aacb20..932f2780d3a4 100644
> --- a/include/net/sctp/structs.h
> +++ b/include/net/sctp/structs.h
> @@ -602,6 +602,9 @@ struct sctp_chunk {
>         /* This needs to be recoverable for SCTP_SEND_FAILED events. */
>         struct sctp_sndrcvinfo sinfo;
>
> +       unsigned long prsctp_param;
> +       int sent_count;
> +
>         /* Which association does this belong to?  */
>         struct sctp_association *asoc;
>
> --
> 2.5.5
>
> Then the performance dropped to the same as the bisected commit
> a6c2f792873a:
> $ cat 4.7.0-rc6-01198-g98dd2532b14e/0/netperf.json
> {
>   "netperf.Throughput_Mbps": [
>    754.494375
>   ]
> }
>
> I think this agrees with the perf data in that the newly added function
> doesn't show up in the perf-profile but still, the performance drops.
> So the only possible reason is the newly added fields to the sctp_chunk
> structure.
>
> Is this expected?
interesting , you didn't include the modification of the functions
parts, right ?
you mean only this two line:
> +       unsigned long prsctp_param;
> +       int sent_count;ca;

caused the performance issue ?

[toc] | [prev] | [next] | [standalone]


#1464393 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-17 09:50 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s73LP-6q3-11@gated-at.bofh.it>
In reply to#1464386
On 08/17/2016 03:35 PM, Xin Long wrote:
>>  include/net/sctp/structs.h | 3 +++
>>  1 file changed, 3 insertions(+)
>>
>> diff --git a/include/net/sctp/structs.h b/include/net/sctp/structs.h
>> index d8e464aacb20..932f2780d3a4 100644
>> --- a/include/net/sctp/structs.h
>> +++ b/include/net/sctp/structs.h
>> @@ -602,6 +602,9 @@ struct sctp_chunk {
>>         /* This needs to be recoverable for SCTP_SEND_FAILED events. */
>>         struct sctp_sndrcvinfo sinfo;
>>
>> +       unsigned long prsctp_param;
>> +       int sent_count;
>> +
>>         /* Which association does this belong to?  */
>>         struct sctp_association *asoc;
>>
>> --
>> 2.5.5
>>
>> Then the performance dropped to the same as the bisected commit
>> a6c2f792873a:
>> $ cat 4.7.0-rc6-01198-g98dd2532b14e/0/netperf.json
>> {
>>   "netperf.Throughput_Mbps": [
>>    754.494375
>>   ]
>> }
>>
>> I think this agrees with the perf data in that the newly added function
>> doesn't show up in the perf-profile but still, the performance drops.
>> So the only possible reason is the newly added fields to the sctp_chunk
>> structure.
>>
>> Is this expected?
> interesting , you didn't include the modification of the functions
> parts, right ?

Yes.

> you mean only this two line:
>> +       unsigned long prsctp_param;
>> +       int sent_count;ca;
> 
> caused the performance issue ?
 
Right.

[toc] | [prev] | [next] | [standalone]


#1464395 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-17 10:00 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s73Vv-6tw-5@gated-at.bofh.it>
In reply to#1464393
On Wed, Aug 17, 2016 at 03:42:34PM +0800, Aaron Lu wrote:
> On 08/17/2016 03:35 PM, Xin Long wrote:
> >>  include/net/sctp/structs.h | 3 +++
> >>  1 file changed, 3 insertions(+)
> >>
> >> diff --git a/include/net/sctp/structs.h b/include/net/sctp/structs.h
> >> index d8e464aacb20..932f2780d3a4 100644
> >> --- a/include/net/sctp/structs.h
> >> +++ b/include/net/sctp/structs.h
> >> @@ -602,6 +602,9 @@ struct sctp_chunk {
> >>         /* This needs to be recoverable for SCTP_SEND_FAILED events. */
> >>         struct sctp_sndrcvinfo sinfo;
> >>
> >> +       unsigned long prsctp_param;
> >> +       int sent_count;
> >> +
> >>         /* Which association does this belong to?  */
> >>         struct sctp_association *asoc;
> >>
> >> --
> >> 2.5.5
> >>
> >> Then the performance dropped to the same as the bisected commit
> >> a6c2f792873a:
> >> $ cat 4.7.0-rc6-01198-g98dd2532b14e/0/netperf.json
> >> {
> >>   "netperf.Throughput_Mbps": [
> >>    754.494375
> >>   ]
> >> }
> >>
> >> I think this agrees with the perf data in that the newly added function
> >> doesn't show up in the perf-profile but still, the performance drops.
> >> So the only possible reason is the newly added fields to the sctp_chunk
> >> structure.
> >>
> >> Is this expected?
> > interesting , you didn't include the modification of the functions
> > parts, right ?
> 
> Yes.
> 
> > you mean only this two line:
> >> +       unsigned long prsctp_param;
> >> +       int sent_count;ca;
> > 
> > caused the performance issue ?
>  
> Right.

Note the test is done on my own Sandybridge desktop, I'll queue a job to
run on the Ivybridge test box now.

[toc] | [prev] | [next] | [standalone]


#1464397

FromXin Long <lucien.xin@gmail.com>
Date2016-08-17 10:10 +0200
Message-ID<s745c-6M0-7@gated-at.bofh.it>
In reply to#1464393
>> you mean only this two line:
>>> +       unsigned long prsctp_param;
>>> +       int sent_count;ca;
>>
>> caused the performance issue ?
>
> Right.
OK, can you remove this line from your patch
+       int sent_count;

then test again, thanks.

[toc] | [prev] | [next] | [standalone]


#1464414 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-17 10:50 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s74HT-71z-5@gated-at.bofh.it>
In reply to#1464397
On Wed, Aug 17, 2016 at 04:02:45PM +0800, Xin Long wrote:
> >> you mean only this two line:
> >>> +       unsigned long prsctp_param;
> >>> +       int sent_count;ca;
> >>
> >> caused the performance issue ?
> >
> > Right.
> OK, can you remove this line from your patch
> +       int sent_count;
> 
> then test again, thanks.

It doesn't change on my desktop Sandybridge.

$ cat 4.7.0-rc6-01199-g116558d316e8/0/netperf.json
{
  "netperf.Throughput_Mbps": [
   748.2056249999998
  ]
}

Where commit 116558d316e8 is based on top of the last test commit
98dd2532b14e with the sent_count removed.

Thanks,
Aaron

[toc] | [prev] | [next] | [standalone]


#1464430

FromXin Long <lucien.xin@gmail.com>
Date2016-08-17 11:10 +0200
Message-ID<s751g-7ok-27@gated-at.bofh.it>
In reply to#1464414
>
> It doesn't change on my desktop Sandybridge.
>
> $ cat 4.7.0-rc6-01199-g116558d316e8/0/netperf.json
> {
>   "netperf.Throughput_Mbps": [
>    748.2056249999998
>   ]
> }
>
> Where commit 116558d316e8 is based on top of the last test commit
> 98dd2532b14e with the sent_count removed.
Nice job
I guess it may be because of your system memory limitation

sctp_chunk size is bigger than before, netpref produced a lot
of sctp_chunk in send queue.

can you check the memory of your systems when the test is
running,  to see if memory is the bottle neck of this test ?

>
> Thanks,
> Aaron

[toc] | [prev] | [next] | [standalone]


#1464441 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-17 11:30 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s75kB-7vR-29@gated-at.bofh.it>
In reply to#1464430
On 08/17/2016 04:58 PM, Xin Long wrote:
>>
>> It doesn't change on my desktop Sandybridge.
>>
>> $ cat 4.7.0-rc6-01199-g116558d316e8/0/netperf.json
>> {
>>   "netperf.Throughput_Mbps": [
>>    748.2056249999998
>>   ]
>> }
>>
>> Where commit 116558d316e8 is based on top of the last test commit
>> 98dd2532b14e with the sent_count removed.
> Nice job
> I guess it may be because of your system memory limitation
> 
> sctp_chunk size is bigger than before, netpref produced a lot
> of sctp_chunk in send queue.
> 
> can you check the memory of your systems when the test is
> running,  to see if memory is the bottle neck of this test ?

We have a monitor to dump /proc/meminfo every second during the run.

On my desktop, the result is -

At start:
time: 1471413103.386122645
MemTotal:       14193468 kB
MemFree:        13849136 kB
MemAvailable:   13789204 kB

In the middle of the run:
time: 1471413254.363430637
MemTotal:       14193468 kB
MemFree:        13811732 kB
MemAvailable:   13756376 kB

When the test is about to finish:
time: 1471413391.294215121
MemTotal:       14193468 kB
MemFree:        13286080 kB
MemAvailable:   13749416 kB

It doesn't seem memory is an issue.

The whole dump is about the same.
The MemFree and MemAvailable doesn't change much.

Thanks,
Aaron

[toc] | [prev] | [next] | [standalone]


#1464698

FromXin Long <lucien.xin@gmail.com>
Date2016-08-17 20:10 +0200
Message-ID<s7drP-4Al-13@gated-at.bofh.it>
In reply to#1464441

[Multipart message — attachments visible in raw view] — view raw

>
> It doesn't seem memory is an issue.
>
> The whole dump is about the same.
> The MemFree and MemAvailable doesn't change much.
>
Hi, Aaron

1)
I talked with Marcelo about this one.
He said it might be related with cacheline.  the  new field distroyed
the prior cacheline. So on top of commit 826d253d57b1, pls only add
+       unsigned long prsctp_param;

to the end of struct sctp_chunk, then try.


2)
if 1) still doesn't work, I may think about to drop prsctp_param in
sctp_chunk, and reuse msg->expire_at. as for sent_count, I will
put it to the mem hole of sctp_chunk.

So pls also try the attachment patch,  on top of commit a6c2f792873a

Thanks.

[toc] | [prev] | [next] | [standalone]


#1464910 — Re: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression

FromAaron Lu <aaron.lu@intel.com>
Date2016-08-18 05:30 +0200
SubjectRe: [LKP] [lkp] [sctp] a6c2f79287: netperf.Throughput_Mbps -37.2% regression
Message-ID<s7mbL-2j1-3@gated-at.bofh.it>
In reply to#1464698
On Thu, Aug 18, 2016 at 02:06:50AM +0800, Xin Long wrote:
> >
> > It doesn't seem memory is an issue.
> >
> > The whole dump is about the same.
> > The MemFree and MemAvailable doesn't change much.
> >
> Hi, Aaron
> 
> 1)
> I talked with Marcelo about this one.
> He said it might be related with cacheline.  the  new field distroyed
> the prior cacheline. So on top of commit 826d253d57b1, pls only add
> +       unsigned long prsctp_param;
> 
> to the end of struct sctp_chunk, then try.

This doesn't work.
 
> 2)
> if 1) still doesn't work, I may think about to drop prsctp_param in
> sctp_chunk, and reuse msg->expire_at. as for sent_count, I will
> put it to the mem hole of sctp_chunk.
> 
> So pls also try the attachment patch,  on top of commit a6c2f792873a

Good news, this brings the performance back on my Sandybridge desktop :)
I have queued jobs to the Ivybridge test box but I guess the result is
the same, but will let you know if it isn't.

It looks like the size of the structure plays a role here, but not clear
to me what happened underneath. Do you know why?

Thanks,
Aaron

> 
> Thanks.

> diff --git a/include/net/sctp/structs.h b/include/net/sctp/structs.h
> index 6bcda71..008cb76 100644
> --- a/include/net/sctp/structs.h
> +++ b/include/net/sctp/structs.h
> @@ -524,7 +524,11 @@ struct sctp_datamsg {
>  	struct list_head chunks;
>  	/* Reference counting. */
>  	atomic_t refcnt;
> -	/* When is this message no longer interesting to the peer? */
> +	/* Re-use this field to record param for prsctp policies,
> +	 * for TTL policy, it is the time_to_drop of this chunk,
> +	 * for RTX policy, it is the max_sent_count of this chunk,
> +	 * for PRIO policy, it is the priority of this chunk.
> +	 */
>  	unsigned long expires_at;
>  	/* Did the messenge fail to send? */
>  	int send_error;
> @@ -553,6 +557,9 @@ struct sctp_chunk {
>  
>  	atomic_t refcnt;
>  
> +	/* How many times this chunk have been sent, for prsctp RTX policy */
> +	int sent_count;
> +
>  	/* This is our link to the per-transport transmitted list.  */
>  	struct list_head transmitted_list;
>  
> @@ -602,16 +609,6 @@ struct sctp_chunk {
>  	/* This needs to be recoverable for SCTP_SEND_FAILED events. */
>  	struct sctp_sndrcvinfo sinfo;
>  
> -	/* We use this field to record param for prsctp policies,
> -	 * for TTL policy, it is the time_to_drop of this chunk,
> -	 * for RTX policy, it is the max_sent_count of this chunk,
> -	 * for PRIO policy, it is the priority of this chunk.
> -	 */
> -	unsigned long prsctp_param;
> -
> -	/* How many times this chunk have been sent, for prsctp RTX policy */
> -	int sent_count;
> -
>  	/* Which association does this belong to?  */
>  	struct sctp_association *asoc;
>  
> diff --git a/net/sctp/chunk.c b/net/sctp/chunk.c
> index 2698d12..0c53d64 100644
> --- a/net/sctp/chunk.c
> +++ b/net/sctp/chunk.c
> @@ -349,7 +349,7 @@ int sctp_chunk_abandoned(struct sctp_chunk *chunk)
>  	}
>  
>  	if (SCTP_PR_TTL_ENABLED(chunk->sinfo.sinfo_flags) &&
> -	    time_after(jiffies, chunk->prsctp_param)) {
> +	    time_after(jiffies, chunk->msg->expires_at)) {
>  		if (chunk->sent_count)
>  			chunk->asoc->abandoned_sent[SCTP_PR_INDEX(TTL)]++;
>  		else
> diff --git a/net/sctp/sm_make_chunk.c b/net/sctp/sm_make_chunk.c
> index 2c431ee..c7110a9 100644
> --- a/net/sctp/sm_make_chunk.c
> +++ b/net/sctp/sm_make_chunk.c
> @@ -718,7 +718,7 @@ static void sctp_set_prsctp_policy(struct sctp_chunk *chunk,
>  		return;
>  
>  	if (SCTP_PR_TTL_ENABLED(sinfo->sinfo_flags))
> -		chunk->prsctp_param =
> +		chunk->msg->expires_at =
>  			jiffies + msecs_to_jiffies(sinfo->sinfo_timetolive);
>  }
>  

[toc] | [prev] | [next] | [standalone]


#1465226

FromXin Long <lucien.xin@gmail.com>
Date2016-08-18 14:50 +0200
Message-ID<s7uVH-8hT-5@gated-at.bofh.it>
In reply to#1464910
>> Hi, Aaron
>>
>> 1)
>> I talked with Marcelo about this one.
>> He said it might be related with cacheline.  the  new field distroyed
>> the prior cacheline. So on top of commit 826d253d57b1, pls only add
>> +       unsigned long prsctp_param;
>>
>> to the end of struct sctp_chunk, then try.
>
> This doesn't work.
>

If it's because of cache lines changed, I'm not sure this, either.
Maybe 2) is a good way to fix it.

Thanks Aaron.

>> 2)
>> if 1) still doesn't work, I may think about to drop prsctp_param in
>> sctp_chunk, and reuse msg->expire_at. as for sent_count, I will
>> put it to the mem hole of sctp_chunk.
>>
>> So pls also try the attachment patch,  on top of commit a6c2f792873a
>
> Good news, this brings the performance back on my Sandybridge desktop :)
> I have queued jobs to the Ivybridge test box but I guess the result is
> the same, but will let you know if it isn't.
>
> It looks like the size of the structure plays a role here, but not clear
> to me what happened underneath. Do you know why?
>
> Thanks,
> Aaron
>
>>
>> Thanks.
>
>> diff --git a/include/net/sctp/structs.h b/include/net/sctp/structs.h
>> index 6bcda71..008cb76 100644
>> --- a/include/net/sctp/structs.h
>> +++ b/include/net/sctp/structs.h
>> @@ -524,7 +524,11 @@ struct sctp_datamsg {
>>       struct list_head chunks;
>>       /* Reference counting. */
>>       atomic_t refcnt;
>> -     /* When is this message no longer interesting to the peer? */
>> +     /* Re-use this field to record param for prsctp policies,
>> +      * for TTL policy, it is the time_to_drop of this chunk,
>> +      * for RTX policy, it is the max_sent_count of this chunk,
>> +      * for PRIO policy, it is the priority of this chunk.
>> +      */
>>       unsigned long expires_at;
>>       /* Did the messenge fail to send? */
>>       int send_error;
>> @@ -553,6 +557,9 @@ struct sctp_chunk {
>>
>>       atomic_t refcnt;
>>
>> +     /* How many times this chunk have been sent, for prsctp RTX policy */
>> +     int sent_count;
>> +
>>       /* This is our link to the per-transport transmitted list.  */
>>       struct list_head transmitted_list;
>>
>> @@ -602,16 +609,6 @@ struct sctp_chunk {
>>       /* This needs to be recoverable for SCTP_SEND_FAILED events. */
>>       struct sctp_sndrcvinfo sinfo;
>>
>> -     /* We use this field to record param for prsctp policies,
>> -      * for TTL policy, it is the time_to_drop of this chunk,
>> -      * for RTX policy, it is the max_sent_count of this chunk,
>> -      * for PRIO policy, it is the priority of this chunk.
>> -      */
>> -     unsigned long prsctp_param;
>> -
>> -     /* How many times this chunk have been sent, for prsctp RTX policy */
>> -     int sent_count;
>> -
>>       /* Which association does this belong to?  */
>>       struct sctp_association *asoc;
>>
>> diff --git a/net/sctp/chunk.c b/net/sctp/chunk.c
>> index 2698d12..0c53d64 100644
>> --- a/net/sctp/chunk.c
>> +++ b/net/sctp/chunk.c
>> @@ -349,7 +349,7 @@ int sctp_chunk_abandoned(struct sctp_chunk *chunk)
>>       }
>>
>>       if (SCTP_PR_TTL_ENABLED(chunk->sinfo.sinfo_flags) &&
>> -         time_after(jiffies, chunk->prsctp_param)) {
>> +         time_after(jiffies, chunk->msg->expires_at)) {
>>               if (chunk->sent_count)
>>                       chunk->asoc->abandoned_sent[SCTP_PR_INDEX(TTL)]++;
>>               else
>> diff --git a/net/sctp/sm_make_chunk.c b/net/sctp/sm_make_chunk.c
>> index 2c431ee..c7110a9 100644
>> --- a/net/sctp/sm_make_chunk.c
>> +++ b/net/sctp/sm_make_chunk.c
>> @@ -718,7 +718,7 @@ static void sctp_set_prsctp_policy(struct sctp_chunk *chunk,
>>               return;
>>
>>       if (SCTP_PR_TTL_ENABLED(sinfo->sinfo_flags))
>> -             chunk->prsctp_param =
>> +             chunk->msg->expires_at =
>>                       jiffies + msecs_to_jiffies(sinfo->sinfo_timetolive);
>>  }
>>
>

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.kernel


csiph-web