Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1600743 > unrolled thread
| Started by | Shannon Nelson <shannon.nelson@oracle.com> |
|---|---|
| First post | 2017-03-14 18:40 +0100 |
| Last post | 2017-03-16 19:40 +0100 |
| Articles | 5 — 3 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
[PATCH v2 net-next 4/5] sunvnet: count multicast packets Shannon Nelson <shannon.nelson@oracle.com> - 2017-03-14 18:40 +0100
RE: [PATCH v2 net-next 4/5] sunvnet: count multicast packets David Laight <David.Laight@ACULAB.COM> - 2017-03-15 10:00 +0100
Re: [PATCH v2 net-next 4/5] sunvnet: count multicast packets Shannon Nelson <shannon.nelson@oracle.com> - 2017-03-16 01:20 +0100
RE: [PATCH v2 net-next 4/5] sunvnet: count multicast packets David Laight <David.Laight@ACULAB.COM> - 2017-03-16 13:20 +0100
Re: [PATCH v2 net-next 4/5] sunvnet: count multicast packets David Miller <davem@davemloft.net> - 2017-03-16 19:40 +0100
| From | Shannon Nelson <shannon.nelson@oracle.com> |
|---|---|
| Date | 2017-03-14 18:40 +0100 |
| Subject | [PATCH v2 net-next 4/5] sunvnet: count multicast packets |
| Message-ID | <tkYAq-3tR-25@gated-at.bofh.it> |
Make sure multicast packets get counted in the device. Orabug: 25190537 Signed-off-by: Shannon Nelson <shannon.nelson@oracle.com> --- drivers/net/ethernet/sun/sunvnet_common.c | 2 ++ 1 files changed, 2 insertions(+), 0 deletions(-) diff --git a/drivers/net/ethernet/sun/sunvnet_common.c b/drivers/net/ethernet/sun/sunvnet_common.c index 5e1d016..0c35a9a 100644 --- a/drivers/net/ethernet/sun/sunvnet_common.c +++ b/drivers/net/ethernet/sun/sunvnet_common.c @@ -409,6 +409,8 @@ static int vnet_rx_one(struct vnet_port *port, struct vio_net_desc *desc) skb->ip_summed = port->switch_port ? CHECKSUM_NONE : CHECKSUM_PARTIAL; + if (unlikely(is_multicast_ether_addr(eth_hdr(skb)->h_dest))) + dev->stats.multicast++; dev->stats.rx_packets++; dev->stats.rx_bytes += len; port->stats.rx_packets++; -- 1.7.1
[toc] | [next] | [standalone]
| From | David Laight <David.Laight@ACULAB.COM> |
|---|---|
| Date | 2017-03-15 10:00 +0100 |
| Message-ID | <tlcWJ-5bE-3@gated-at.bofh.it> |
| In reply to | #1600743 |
From: Shannon Nelson > Sent: 14 March 2017 17:25 ... > + if (unlikely(is_multicast_ether_addr(eth_hdr(skb)->h_dest))) > + dev->stats.multicast++; I'd guess that: dev->stats.multicast += is_multicast_ether_addr(eth_hdr(skb)->h_dest); generates faster code. Especially if is_multicast_ether_addr(x) is (*x >> 7). David
[toc] | [prev] | [next] | [standalone]
| From | Shannon Nelson <shannon.nelson@oracle.com> |
|---|---|
| Date | 2017-03-16 01:20 +0100 |
| Message-ID | <tlrj4-6XI-3@gated-at.bofh.it> |
| In reply to | #1601146 |
On 3/15/2017 1:50 AM, David Laight wrote:
> From: Shannon Nelson
>> Sent: 14 March 2017 17:25
> ...
>> + if (unlikely(is_multicast_ether_addr(eth_hdr(skb)->h_dest)))
>> + dev->stats.multicast++;
>
> I'd guess that:
> dev->stats.multicast += is_multicast_ether_addr(eth_hdr(skb)->h_dest);
> generates faster code.
> Especially if is_multicast_ether_addr(x) is (*x >> 7).
>
> David
Hi David, thanks for the comment. My local instruction level
performance guru is on vacation this week so I can't do a quick check
with him today on this. However, I"m not too worried here since the
inline code for is_multicast_ether_addr() is simply
return 0x01 & addr[0];
and objdump tells me that on sparc it compiles down to a simple single
byte load and compare:
325c: c2 08 80 03 ldub [ %g2 + %g3 ], %g1
3260: 80 88 60 01 btst 1, %g1
3264: 32 60 00 b3 bne,a,pn %xcc, 3530 <vnet_rx_one+0x430>
3268: c2 5c 61 68 ldx [ %l1 + 0x168 ], %g1
dev->stats.multicast++;
I don't think this driver will ever be used on anything bug sparc, so
I'm not worried about how x86 might compile this.
sln
[toc] | [prev] | [next] | [standalone]
| From | David Laight <David.Laight@ACULAB.COM> |
|---|---|
| Date | 2017-03-16 13:20 +0100 |
| Message-ID | <tlCxP-6BJ-3@gated-at.bofh.it> |
| In reply to | #1601821 |
From: Shannon Nelson > Sent: 16 March 2017 00:18 > To: David Laight; netdev@vger.kernel.org; davem@davemloft.net > On 3/15/2017 1:50 AM, David Laight wrote: > > From: Shannon Nelson > >> Sent: 14 March 2017 17:25 > > ... > >> + if (unlikely(is_multicast_ether_addr(eth_hdr(skb)->h_dest))) > >> + dev->stats.multicast++; > > > > I'd guess that: > > dev->stats.multicast += is_multicast_ether_addr(eth_hdr(skb)->h_dest); > > generates faster code. > > Especially if is_multicast_ether_addr(x) is (*x >> 7). I'd clearly got brain-fade there, mcast bit is the first transmitted bit (on ethernet) but the bytes are sent LSB first (like async). > > David > > Hi David, thanks for the comment. My local instruction level > performance guru is on vacation this week so I can't do a quick check > with him today on this. However, I"m not too worried here since the > inline code for is_multicast_ether_addr() is simply > > return 0x01 & addr[0]; > > and objdump tells me that on sparc it compiles down to a simple single > byte load and compare: > > 325c: c2 08 80 03 ldub [ %g2 + %g3 ], %g1 > 3260: 80 88 60 01 btst 1, %g1 > 3264: 32 60 00 b3 bne,a,pn %xcc, 3530 <vnet_rx_one+0x430> > 3268: c2 5c 61 68 ldx [ %l1 + 0x168 ], %g1 > dev->stats.multicast++; Followed by a branch that might be marked 'assume taken' so the normal path takes the branch. I guess that is followed by 'add 1 to %g1', 'stx %g1, [ %l1 + 0x168 ]' and a branch to 3530. GCC must be using that condition to generate get the bottom of a loop to 'fallthrough' to its top! My version should generate something like: ldub [ %g2 + %g3 ], %g1 ldx [ %l1 + 0x168 ], %g2 and 1, %g1 add %g1, %g2, %g2 stx %g2, [ %l1 + 0x168 ] While this looks like 5 instructions (rather than 2) it has fewer pipeline stalls and can be 'spread out' into the surrounding lines of code to reduce the stalls further. > I don't think this driver will ever be used on anything bug sparc, so > I'm not worried about how x86 might compile this. On x86 gcc is likely to ignore the 'unlikely' and generate a forwards (predicted not taken) branch around the increment. I've had to but asm comments in the else part of conditionals like that to force gcc to generate a forwards jump to the 'unlikely' statements. David
[toc] | [prev] | [next] | [standalone]
| From | David Miller <davem@davemloft.net> |
|---|---|
| Date | 2017-03-16 19:40 +0100 |
| Message-ID | <tlItB-2iT-41@gated-at.bofh.it> |
| In reply to | #1602216 |
From: David Laight <David.Laight@ACULAB.COM> Date: Thu, 16 Mar 2017 12:12:06 +0000 > From: Shannon Nelson >> Sent: 16 March 2017 00:18 >> To: David Laight; netdev@vger.kernel.org; davem@davemloft.net >> On 3/15/2017 1:50 AM, David Laight wrote: >> > From: Shannon Nelson >> >> Sent: 14 March 2017 17:25 >> > ... >> >> + if (unlikely(is_multicast_ether_addr(eth_hdr(skb)->h_dest))) >> >> + dev->stats.multicast++; >> > >> > I'd guess that: >> > dev->stats.multicast += is_multicast_ether_addr(eth_hdr(skb)->h_dest); >> > generates faster code. >> > Especially if is_multicast_ether_addr(x) is (*x >> 7). > > I'd clearly got brain-fade there, mcast bit is the first transmitted bit > (on ethernet) but the bytes are sent LSB first (like async). >> > David >> >> Hi David, thanks for the comment. My local instruction level >> performance guru is on vacation this week so I can't do a quick check >> with him today on this. However, I"m not too worried here since the >> inline code for is_multicast_ether_addr() is simply >> >> return 0x01 & addr[0]; >> >> and objdump tells me that on sparc it compiles down to a simple single >> byte load and compare: >> >> 325c: c2 08 80 03 ldub [ %g2 + %g3 ], %g1 >> 3260: 80 88 60 01 btst 1, %g1 >> 3264: 32 60 00 b3 bne,a,pn %xcc, 3530 <vnet_rx_one+0x430> >> 3268: c2 5c 61 68 ldx [ %l1 + 0x168 ], %g1 >> dev->stats.multicast++; > > Followed by a branch that might be marked 'assume taken' so the > normal path takes the branch. The branch is predicted not taken, so the fallthrough happens most often. And this is optimal for most Niagara parts as taken branches make the cpu thread yield whereas non-taken branches do not. But this is such a petty thing to be discussing compared to the substance of this person's changes. David, I really wish you wouldn't waste people's time with this stuff. Maybe if you had to review hundreds of networking patches every day like I do, you would start to understand the costs of the interference you place into the review process when you bring up such small matters like this all the time. I'd much rather you review the substance of a person's changes, because that actually helps things more forward. If you want to micro optimize then _do it on your own time_, submit patches that do the micro optimization, and have it go through the review process like everyone else's changes. I very much appreciate your cooperation on this matter. Thanks.
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web