Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1215733 > unrolled thread

[PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking all the percpu data at once

Started byRaghavendra K T <raghavendra.kt@linux.vnet.ibm.com>
First post2015-08-29 11:10 +0200
Last post2015-08-29 19:30 +0200
Articles 4 — 3 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking all the percpu data at once Raghavendra K T <raghavendra.kt@linux.vnet.ibm.com> - 2015-08-29 11:10 +0200
    Re: [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by  walking all the percpu data at once Eric Dumazet <eric.dumazet@gmail.com> - 2015-08-29 16:40 +0200
      Re: [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by  walking all the percpu data at once Joe Perches <joe@perches.com> - 2015-08-29 17:30 +0200
        Re: [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking  all the percpu data at once Raghavendra K T <raghavendra.kt@linux.vnet.ibm.com> - 2015-08-29 19:30 +0200

#1215733 — [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking all the percpu data at once

FromRaghavendra K T <raghavendra.kt@linux.vnet.ibm.com>
Date2015-08-29 11:10 +0200
Subject[PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking all the percpu data at once
Message-ID<q2Kj7-5Iz-11@gated-at.bofh.it>
Docker container creation linearly increased from around 1.6 sec to 7.5 sec
(at 1000 containers) and perf data showed 50% ovehead in snmp_fold_field.

reason: currently __snmp6_fill_stats64 calls snmp_fold_field that walks
through per cpu data of an item (iteratively for around 36 items).

idea: This patch tries to aggregate the statistics by going through
all the items of each cpu sequentially which is reducing cache
misses.

Docker creation got faster by more than 2x after the patch.

Result:
                       Before           After
Docker creation time   6.836s           3.25s
cache miss             2.7%             1.41%

perf before:
    50.73%  docker           [kernel.kallsyms]       [k] snmp_fold_field
     9.07%  swapper          [kernel.kallsyms]       [k] snooze_loop
     3.49%  docker           [kernel.kallsyms]       [k] veth_stats_one
     2.85%  swapper          [kernel.kallsyms]       [k] _raw_spin_lock

perf after:
    10.57%  docker           docker                [.] scanblock
     8.37%  swapper          [kernel.kallsyms]     [k] snooze_loop
     6.91%  docker           [kernel.kallsyms]     [k] snmp_get_cpu_field
     6.67%  docker           [kernel.kallsyms]     [k] veth_stats_one

changes/ideas suggested:
Using buffer in stack (Eric), Usage of memset (David), Using memcpy in
place of unaligned_put (Joe).

Signed-off-by: Raghavendra K T <raghavendra.kt@linux.vnet.ibm.com>
---
 net/ipv6/addrconf.c | 22 +++++++++++++---------
 1 file changed, 13 insertions(+), 9 deletions(-)

Changes in V3:
 - use memset to initialize temp buffer in leaf function. (David)
 - use memcpy to copy the buffer data to stat instead of unalign_pu (Joe)
 - Move buffer definition to leaf function __snmp6_fill_stats64() (Eric)
 -
Changes in V2:
 - Allocate the stat calculation buffer in stack. (Eric)

diff --git a/net/ipv6/addrconf.c b/net/ipv6/addrconf.c
index 21c2c81..379619a 100644
--- a/net/ipv6/addrconf.c
+++ b/net/ipv6/addrconf.c
@@ -4624,18 +4624,22 @@ static inline void __snmp6_fill_statsdev(u64 *stats, atomic_long_t *mib,
 }
 
 static inline void __snmp6_fill_stats64(u64 *stats, void __percpu *mib,
-				      int items, int bytes, size_t syncpoff)
+					int items, int bytes, size_t syncpoff)
 {
-	int i;
+	int i, c;
 	int pad = bytes - sizeof(u64) * items;
+	u64 buff[items];
+
 	BUG_ON(pad < 0);
 
-	/* Use put_unaligned() because stats may not be aligned for u64. */
-	put_unaligned(items, &stats[0]);
-	for (i = 1; i < items; i++)
-		put_unaligned(snmp_fold_field64(mib, i, syncpoff), &stats[i]);
+	memset(buff, 0, sizeof(buff));
+	buff[0] = items;
 
-	memset(&stats[items], 0, pad);
+	for_each_possible_cpu(c)
+		for (i = 1; i < items; i++)
+			buff[i] += snmp_get_cpu_field64(mib, c, i, syncpoff);
+
+	memcpy(stats, buff, items * sizeof(u64));
 }
 
 static void snmp6_fill_stats(u64 *stats, struct inet6_dev *idev, int attrtype,
@@ -4643,8 +4647,8 @@ static void snmp6_fill_stats(u64 *stats, struct inet6_dev *idev, int attrtype,
 {
 	switch (attrtype) {
 	case IFLA_INET6_STATS:
-		__snmp6_fill_stats64(stats, idev->stats.ipv6,
-				     IPSTATS_MIB_MAX, bytes, offsetof(struct ipstats_mib, syncp));
+		__snmp6_fill_stats64(stats, idev->stats.ipv6, IPSTATS_MIB_MAX,
+				     bytes, offsetof(struct ipstats_mib, syncp));
 		break;
 	case IFLA_INET6_ICMP6STATS:
 		__snmp6_fill_statsdev(stats, idev->stats.icmpv6dev->mibs, ICMP6_MIB_MAX, bytes);
-- 
1.7.11.7

--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [next] | [standalone]


#1215781 — Re: [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking all the percpu data at once

FromEric Dumazet <eric.dumazet@gmail.com>
Date2015-08-29 16:40 +0200
SubjectRe: [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking all the percpu data at once
Message-ID<q2Pst-4AG-1@gated-at.bofh.it>
In reply to#1215733
On Sat, 2015-08-29 at 14:37 +0530, Raghavendra K T wrote:

>  
>  static inline void __snmp6_fill_stats64(u64 *stats, void __percpu *mib,
> -				      int items, int bytes, size_t syncpoff)
> +					int items, int bytes, size_t syncpoff)
>  {
> -	int i;
> +	int i, c;
>  	int pad = bytes - sizeof(u64) * items;
> +	u64 buff[items];
> +

One last comment : using variable length arrays is confusing for the
reader, and sparse as well.

$ make C=2 net/ipv6/addrconf.o
...
  CHECK   net/ipv6/addrconf.c
net/ipv6/addrconf.c:4733:18: warning: Variable length array is used.
net/ipv6/addrconf.c:4737:25: error: cannot size expression


I suggest you remove 'items' parameter and replace it by
IPSTATS_MIB_MAX, as __snmp6_fill_stats64() is called exactly once.



--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1215803 — Re: [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking all the percpu data at once

FromJoe Perches <joe@perches.com>
Date2015-08-29 17:30 +0200
SubjectRe: [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking all the percpu data at once
Message-ID<q2QeS-5LL-1@gated-at.bofh.it>
In reply to#1215781
On Sat, 2015-08-29 at 07:32 -0700, Eric Dumazet wrote:
> On Sat, 2015-08-29 at 14:37 +0530, Raghavendra K T wrote:
> 
> >  
> >  static inline void __snmp6_fill_stats64(u64 *stats, void __percpu *mib,
> > -				      int items, int bytes, size_t syncpoff)
> > +					int items, int bytes, size_t syncpoff)
> >  {
> > -	int i;
> > +	int i, c;
> >  	int pad = bytes - sizeof(u64) * items;
> > +	u64 buff[items];
> > +
> 
> One last comment : using variable length arrays is confusing for the
> reader, and sparse as well.
> 
> $ make C=2 net/ipv6/addrconf.o
> ...
>   CHECK   net/ipv6/addrconf.c
> net/ipv6/addrconf.c:4733:18: warning: Variable length array is used.
> net/ipv6/addrconf.c:4737:25: error: cannot size expression
> 
> 
> I suggest you remove 'items' parameter and replace it by
> IPSTATS_MIB_MAX, as __snmp6_fill_stats64() is called exactly once.

If you respin, I suggest:

o remove "items" from the __snmp6_fill_stats64 arguments
  and use IPSTATS_MIB_MAX in the function instead

o add braces around the for_each_possible_cpu loop

	for_each_possible_cpu(c) {
		for (i = 1; i < items; i++)
			buff[i] += snmp_get_cpu_field64(mib, c, i, syncpoff);
	}


--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1215820 — Re: [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking all the percpu data at once

FromRaghavendra K T <raghavendra.kt@linux.vnet.ibm.com>
Date2015-08-29 19:30 +0200
SubjectRe: [PATCH RFC V3 2/2] net: Optimize snmp stat aggregation by walking all the percpu data at once
Message-ID<q2S70-8t9-7@gated-at.bofh.it>
In reply to#1215803
On 08/29/2015 08:51 PM, Joe Perches wrote:
> On Sat, 2015-08-29 at 07:32 -0700, Eric Dumazet wrote:
>> On Sat, 2015-08-29 at 14:37 +0530, Raghavendra K T wrote:
>>
>>>
>>>   static inline void __snmp6_fill_stats64(u64 *stats, void __percpu *mib,
>>> -				      int items, int bytes, size_t syncpoff)
>>> +					int items, int bytes, size_t syncpoff)
>>>   {
>>> -	int i;
>>> +	int i, c;
>>>   	int pad = bytes - sizeof(u64) * items;
>>> +	u64 buff[items];
>>> +
>>
>> One last comment : using variable length arrays is confusing for the
>> reader, and sparse as well.
>>
>> $ make C=2 net/ipv6/addrconf.o
>> ...
>>    CHECK   net/ipv6/addrconf.c
>> net/ipv6/addrconf.c:4733:18: warning: Variable length array is used.
>> net/ipv6/addrconf.c:4737:25: error: cannot size expression
>>
>>
>> I suggest you remove 'items' parameter and replace it by
>> IPSTATS_MIB_MAX, as __snmp6_fill_stats64() is called exactly once.
>
> If you respin, I suggest:
>
> o remove "items" from the __snmp6_fill_stats64 arguments
>    and use IPSTATS_MIB_MAX in the function instead

Yes, as also suggested by Eric.

> o add braces around the for_each_possible_cpu loop
>
> 	for_each_possible_cpu(c) {
> 		for (i = 1; i < items; i++)
> 			buff[i] += snmp_get_cpu_field64(mib, c, i, syncpoff);
> 	}
>

Sure. It makes it more readable.
will respin V4 with these changes (+ memset 0 for pad which I realized).



--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web