Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1303420 > unrolled thread
| Started by | Vitaly Kuznetsov <vkuznets@redhat.com> |
|---|---|
| First post | 2016-01-07 10:40 +0100 |
| Last post | 2016-01-14 19:00 +0100 |
| Articles | 5 on this page of 25 — 8 participants |
Back to article view | Back to linux.kernel
[PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Vitaly Kuznetsov <vkuznets@redhat.com> - 2016-01-07 10:40 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Eric Dumazet <eric.dumazet@gmail.com> - 2016-01-07 14:00 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Vitaly Kuznetsov <vkuznets@redhat.com> - 2016-01-07 14:30 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout John Fastabend <john.fastabend@gmail.com> - 2016-01-08 02:10 +0100
RE: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout KY Srinivasan <kys@microsoft.com> - 2016-01-08 04:50 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout John Fastabend <john.fastabend@gmail.com> - 2016-01-08 07:20 +0100
RE: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout KY Srinivasan <kys@microsoft.com> - 2016-01-08 19:10 +0100
RE: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Haiyang Zhang <haiyangz@microsoft.com> - 2016-01-08 22:10 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Tom Herbert <tom@herbertland.com> - 2016-01-09 01:20 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout David Miller <davem@davemloft.net> - 2016-01-10 23:30 +0100
RE: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Haiyang Zhang <haiyangz@microsoft.com> - 2016-01-14 00:30 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout David Miller <davem@davemloft.net> - 2016-01-14 06:00 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Tom Herbert <tom@herbertland.com> - 2016-01-14 18:20 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout One Thousand Gnomes <gnomes@lxorguk.ukuu.org.uk> - 2016-01-14 19:00 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Eric Dumazet <eric.dumazet@gmail.com> - 2016-01-14 19:30 +0100
RE: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Haiyang Zhang <haiyangz@microsoft.com> - 2016-01-14 19:40 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Tom Herbert <tom@herbertland.com> - 2016-01-14 19:50 +0100
RE: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Haiyang Zhang <haiyangz@microsoft.com> - 2016-01-14 20:20 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Tom Herbert <tom@herbertland.com> - 2016-01-14 20:50 +0100
RE: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Haiyang Zhang <haiyangz@microsoft.com> - 2016-01-14 21:30 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Tom Herbert <tom@herbertland.com> - 2016-01-14 22:50 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout David Miller <davem@davemloft.net> - 2016-01-14 23:10 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Eric Dumazet <eric.dumazet@gmail.com> - 2016-01-14 23:10 +0100
RE: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Haiyang Zhang <haiyangz@microsoft.com> - 2016-01-14 23:50 +0100
Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout Eric Dumazet <eric.dumazet@gmail.com> - 2016-01-14 19:00 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Tom Herbert <tom@herbertland.com> |
|---|---|
| Date | 2016-01-14 22:50 +0100 |
| Subject | Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout |
| Message-ID | <qQXWh-4xw-13@gated-at.bofh.it> |
| In reply to | #1309636 |
On Thu, Jan 14, 2016 at 12:23 PM, Haiyang Zhang <haiyangz@microsoft.com> wrote: > > >> -----Original Message----- >> From: Tom Herbert [mailto:tom@herbertland.com] >> Sent: Thursday, January 14, 2016 2:41 PM >> To: Haiyang Zhang <haiyangz@microsoft.com> >> Cc: Eric Dumazet <eric.dumazet@gmail.com>; One Thousand Gnomes >> <gnomes@lxorguk.ukuu.org.uk>; David Miller <davem@davemloft.net>; >> vkuznets@redhat.com; netdev@vger.kernel.org; KY Srinivasan >> <kys@microsoft.com>; devel@linuxdriverproject.org; linux- >> kernel@vger.kernel.org >> Subject: Re: [PATCH net-next] hv_netvsc: don't make assumptions on >> struct flow_keys layout >> >> On Thu, Jan 14, 2016 at 11:15 AM, Haiyang Zhang <haiyangz@microsoft.com> >> wrote: >> > >> > >> >> -----Original Message----- >> >> From: Tom Herbert [mailto:tom@herbertland.com] >> >> Sent: Thursday, January 14, 2016 1:49 PM >> >> To: Haiyang Zhang <haiyangz@microsoft.com> >> >> Cc: Eric Dumazet <eric.dumazet@gmail.com>; One Thousand Gnomes >> >> <gnomes@lxorguk.ukuu.org.uk>; David Miller <davem@davemloft.net>; >> >> vkuznets@redhat.com; netdev@vger.kernel.org; KY Srinivasan >> >> <kys@microsoft.com>; devel@linuxdriverproject.org; linux- >> >> kernel@vger.kernel.org >> >> Subject: Re: [PATCH net-next] hv_netvsc: don't make assumptions on >> >> struct flow_keys layout >> >> >> >> On Thu, Jan 14, 2016 at 10:35 AM, Haiyang Zhang >> <haiyangz@microsoft.com> >> >> wrote: >> >> > >> >> > >> >> >> -----Original Message----- >> >> >> From: Eric Dumazet [mailto:eric.dumazet@gmail.com] >> >> >> Sent: Thursday, January 14, 2016 1:24 PM >> >> >> To: One Thousand Gnomes <gnomes@lxorguk.ukuu.org.uk> >> >> >> Cc: Tom Herbert <tom@herbertland.com>; Haiyang Zhang >> >> >> <haiyangz@microsoft.com>; David Miller <davem@davemloft.net>; >> >> >> vkuznets@redhat.com; netdev@vger.kernel.org; KY Srinivasan >> >> >> <kys@microsoft.com>; devel@linuxdriverproject.org; linux- >> >> >> kernel@vger.kernel.org >> >> >> Subject: Re: [PATCH net-next] hv_netvsc: don't make assumptions on >> >> >> struct flow_keys layout >> >> >> >> >> >> On Thu, 2016-01-14 at 17:53 +0000, One Thousand Gnomes wrote: >> >> >> > > These results for Toeplitz are not plausible. Given random >> input >> >> you >> >> >> > > cannot expect any hash function to produce such uniform >> results. >> >> I >> >> >> > > suspect either your input data is biased or how your applying >> the >> >> >> hash >> >> >> > > is. >> >> >> > > >> >> >> > > When I run 64 random IPv4 3-tuples through Toeplitz and >> Jenkins I >> >> >> get >> >> >> > > something more reasonable: >> >> >> > >> >> >> > IPv4 address patterns are not random. Nothing like it. A long >> long >> >> >> time >> >> >> > ago we did do a bunch of tuning for network hashes using big >> porn >> >> site >> >> >> > data sets. Random it was not. >> >> >> > >> >> >> >> >> >> I ran my tests with non random IPV4 addresses, as I had 2 hosts, >> >> >> one server, one client. (typical benchmark stuff) >> >> >> >> >> >> The only 'random' part was the ports, so maybe ~20 bits of entropy, >> >> >> considering how we allocate ports during connect() to a given >> >> >> destination to avoid port reuse. >> >> >> >> >> >> > It's probably hard to repeat that exercise now with geo specific >> >> >> routing, >> >> >> > and all the front end caches and redirectors on big sites but >> I'd >> >> >> > strongly suggest random input is not a good test, and also that >> you >> >> >> need >> >> >> > to worry more about hash attacks than perfect distributions. >> >> >> >> >> >> Anyway, the exercise is not to find a hash that exactly splits 128 >> >> flows >> >> >> into 16 buckets, according to the number of flows per bucket. >> >> >> >> >> >> Maybe only 4 flows are sending at 3Gbits, and others are sending >> at >> >> 100 >> >> >> kbits. There is no way the driver can predict the future. >> >> >> >> >> >> This is why we prefer to select a queue given the cpu sending the >> >> >> packet. This permits a natural shift based on actual load, and is >> the >> >> >> default on linux (see XPS in Documentation/networking/scaling.txt) >> >> >> >> >> >> Only this driver has a selection based on a flow 'hash'. >> >> > >> >> > Also, the port number selection may not be random either. For >> example, >> >> > the well-known network throughput test tool, iperf, use port >> numbers >> >> with >> >> > equal increment among them. We tested these non-random cases, and >> >> found >> >> > the Toeplitz hash has distributed evenly, but Jenkins hash has non- >> >> even >> >> > distribution. >> >> > >> >> > I'm aware of the test from Tom Herbert <tom@herbertland.com>, which >> >> > showing similar results of Toeplitz v.s. Jenkins with random inputs. >> >> > >> >> > In summary, the Toeplitz performs better in case of non-random >> inputs, >> >> > and performs similar to Jenkins in random inputs (which may not be >> the >> >> > case in real world). So we still prefer to use Toeplitz hash. >> >> > >> >> You are basing your conclusions on one toy benchmark. I don't believe >> >> that an realistically loaded web server is going to consistently give >> >> you tuples that happen to somehow fit into a nice model so that the >> >> bias benefits your load distribution. >> >> >> >> > To minimize the computational overhead, we may consider put the >> hash >> >> > in a per-connection cache in TCP layer, so it only needs one time >> >> > computation. But, even with the computation overhead at this moment, >> >> > the throughput based on Toeplitz hash is better than Jenkins: >> >> > Throughput (Gbps) comparison: >> >> > #conn Toeplitz Jenkins >> >> > 32 26.6 23.2 >> >> > 64 32.1 23.4 >> >> > 128 29.1 24.1 >> >> > >> >> You don't need to do that. We already store a random hash value in >> the >> >> connection context. If you want to make it non-random then just >> >> replace that with a simple global counter. This will have the exact >> >> same effect that you see in your tests without needing any expensive >> >> computation. >> > >> > Could you point me to the data field of connection context where this >> > hash value is stored? Is it computed only one time? >> > >> sk_txhash in struct sock. It is set to a random number on TCP or UDP >> connect call, It can be reset to a different random value when >> connection is seen to be have trouble (sk_rethink_txhash). >> >> Also when you say "Toeplitz performs better in case of non-random >> inputs" please quantify exactly how your input data is not random. >> What header changes with each connection in your test... > > Thank you for the info! > > For non-random inputs, I used the port selection of iperf that increases > the port number by 2 for each connection. Only send-port numbers are > different, other values are the same. I also tested some other fixed > increment, Toeplitz spreads the connections evenly. For real applications, > if the load came from local area, then the IP/port combinations are > likely to have some non-random patterns. > Okay, by only changing source port I can produce the same uniformity: 64 connections with a step of 2 for changing source port gives: Buckets: 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 _but_, I can also find also make steps that severely mess up load distribution. Step 1024 gives: Buckets: 0 8 8 0 0 8 8 0 8 0 0 8 8 0 0 8 The fact that we can negatively affect the output of Toeplitz so predictably is actually a liability and not a benefit. This sort of thing can be the basis of a DOS attack and is why we kicked out XOR hash in favor of Jenkins. > For our driver, we are thinking to put the Toeplitz hash to the sk_txhash, > so it only needs to be computed only once, or during sk_rethink_txhash. > So, the computational overhead happens almost only once. > > Thanks, > - Haiyang > >
[toc] | [prev] | [next] | [standalone]
| From | David Miller <davem@davemloft.net> |
|---|---|
| Date | 2016-01-14 23:10 +0100 |
| Subject | Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout |
| Message-ID | <qQYfF-533-3@gated-at.bofh.it> |
| In reply to | #1309667 |
From: Tom Herbert <tom@herbertland.com> Date: Thu, 14 Jan 2016 13:44:24 -0800 > The fact that we can negatively affect the output of Toeplitz so > predictably is actually a liability and not a benefit. This sort of > thing can be the basis of a DOS attack and is why we kicked out XOR > hash in favor of Jenkins. +1 Toeplitz should not be used for any software calculated flow hash whatsoever.
[toc] | [prev] | [next] | [standalone]
| From | Eric Dumazet <eric.dumazet@gmail.com> |
|---|---|
| Date | 2016-01-14 23:10 +0100 |
| Subject | Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout |
| Message-ID | <qQYfF-533-15@gated-at.bofh.it> |
| In reply to | #1309636 |
On Thu, 2016-01-14 at 20:23 +0000, Haiyang Zhang wrote:
>
> For non-random inputs, I used the port selection of iperf that increases
> the port number by 2 for each connection. Only send-port numbers are
> different, other values are the same. I also tested some other fixed
> increment, Toeplitz spreads the connections evenly. For real applications,
> if the load came from local area, then the IP/port combinations are
> likely to have some non-random patterns.
We are not putting code in core networking stack favoring non secure
behavior.
The +2 behavior for connections from A to B:<fixed port> is something
that we will eventually remove in the future. It used to be +1 not a
long time ago...
Say if we implement the following,
https://tools.ietf.org/html/rfc6056#section-3.3.4
The fact that Toeplitz hash has this linear property should not be a
valid reason to help hackers to exploit vulnerabilities.
In my tests I was using netperf, which randomizes both source &
destination ports.
This is why I could not reproduce your results based on iperf, which
generates 5-tuple in a totally predictable way.
This reminds me some drivers had a well known Toeplitz RSS key, allowing
attackers to direct their attack on a single queue.
I guess we could replace sk_txhash generator by a simple linear
allocator and boom, your driver will be pleased.
But this is only for a very specific workload.
diff --git a/include/net/sock.h b/include/net/sock.h
index e830c1006935..949527413cfb 100644
--- a/include/net/sock.h
+++ b/include/net/sock.h
@@ -1689,7 +1689,8 @@ unsigned long sock_i_ino(struct sock *sk);
static inline u32 net_tx_rndhash(void)
{
- u32 v = prandom_u32();
+ static u32 last_hash;
+ u32 v = ++last_hash; // do not care about SMP races.
return v ?: 1;
}
[toc] | [prev] | [next] | [standalone]
| From | Haiyang Zhang <haiyangz@microsoft.com> |
|---|---|
| Date | 2016-01-14 23:50 +0100 |
| Subject | RE: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout |
| Message-ID | <qQYSm-5iF-19@gated-at.bofh.it> |
| In reply to | #1309701 |
> -----Original Message-----
> From: Eric Dumazet [mailto:eric.dumazet@gmail.com]
> Sent: Thursday, January 14, 2016 5:08 PM
> To: Haiyang Zhang <haiyangz@microsoft.com>
> Cc: Tom Herbert <tom@herbertland.com>; One Thousand Gnomes
> <gnomes@lxorguk.ukuu.org.uk>; David Miller <davem@davemloft.net>;
> vkuznets@redhat.com; netdev@vger.kernel.org; KY Srinivasan
> <kys@microsoft.com>; devel@linuxdriverproject.org; linux-
> kernel@vger.kernel.org
> Subject: Re: [PATCH net-next] hv_netvsc: don't make assumptions on
> struct flow_keys layout
>
> On Thu, 2016-01-14 at 20:23 +0000, Haiyang Zhang wrote:
> >
>
>
> > For non-random inputs, I used the port selection of iperf that
> increases
> > the port number by 2 for each connection. Only send-port numbers are
> > different, other values are the same. I also tested some other fixed
> > increment, Toeplitz spreads the connections evenly. For real
> applications,
> > if the load came from local area, then the IP/port combinations are
> > likely to have some non-random patterns.
>
> We are not putting code in core networking stack favoring non secure
> behavior.
>
> The +2 behavior for connections from A to B:<fixed port> is something
> that we will eventually remove in the future. It used to be +1 not a
> long time ago...
>
> Say if we implement the following,
>
> https://na01.safelinks.protection.outlook.com/?url=https%3a%2f%2ftools.i
> etf.org%2fhtml%2frfc6056%23section-
> 3.3.4&data=01%7c01%7chaiyangz%40microsoft.com%7ced5f98ae23a843df05c408d3
> 1d2f3028%7c72f988bf86f141af91ab2d7cd011db47%7c1&sdata=uPo0Rdme20vZX%2b%2
> frcwe1iE0mKGZYl%2fMdeaF1wld%2fgbQ%3d
>
>
> The fact that Toeplitz hash has this linear property should not be a
> valid reason to help hackers to exploit vulnerabilities.
>
> In my tests I was using netperf, which randomizes both source &
> destination ports.
>
> This is why I could not reproduce your results based on iperf, which
> generates 5-tuple in a totally predictable way.
>
> This reminds me some drivers had a well known Toeplitz RSS key, allowing
> attackers to direct their attack on a single queue.
>
> I guess we could replace sk_txhash generator by a simple linear
> allocator and boom, your driver will be pleased.
>
> But this is only for a very specific workload.
>
> diff --git a/include/net/sock.h b/include/net/sock.h
> index e830c1006935..949527413cfb 100644
> --- a/include/net/sock.h
> +++ b/include/net/sock.h
> @@ -1689,7 +1689,8 @@ unsigned long sock_i_ino(struct sock *sk);
>
> static inline u32 net_tx_rndhash(void)
> {
> - u32 v = prandom_u32();
> + static u32 last_hash;
> + u32 v = ++last_hash; // do not care about SMP races.
>
> return v ?: 1;
> }
Tom, Thanks for your test -- I was not able to reproduce the
"0 8 8 0 0 8 8 0 8 0 0 8 8 0 0 8" distribution, but I did see some
predictable patterns by using some increments like 512...
Tom, Dave, and Eric -- I share your concerns on potential DoS attack
on predictable patterns. We will re-think about this.
Thanks,
- Haiyang
[toc] | [prev] | [next] | [standalone]
| From | Eric Dumazet <eric.dumazet@gmail.com> |
|---|---|
| Date | 2016-01-14 19:00 +0100 |
| Subject | Re: [PATCH net-next] hv_netvsc: don't make assumptions on struct flow_keys layout |
| Message-ID | <qQUlI-1WT-17@gated-at.bofh.it> |
| In reply to | #1308896 |
On Wed, 2016-01-13 at 23:10 +0000, Haiyang Zhang wrote: > I have done a comparison of the Toeplitz v.s. Jenkins Hash algorithms, > and found that the Toeplitz provides much better distribution of the > connections into send-indirection-table entries. See the data below -- > showing how many TCP connections are distributed into each of the > sixteen table entries. The Toeplitz hash distributes the connections > almost perfectly evenly, but the Jenkins hash distributes them unevenly. > For example, in case of 64 connections, some entries are 0 or 1, some > other entries are 8. This could cause too many connections in one VMBus > channel and slow down the throughput. So a VMBus channel has a limit of number of flows ? Why is it so ? What happens with 1000 flows ? > This is consistent to our test > which showing slower performance while using the generic skb_get_hash > (Jenkins) than using Toeplitz hash (see perf numbers below). > > > #connections:32: > Toeplitz:2,2,2,2,2,1,2,2,2,2,2,3,2,2,2,2, > Jenkins:3,2,2,4,1,1,0,2,1,1,4,3,2,5,1,0, > #connections:64: > Toeplitz:4,4,5,4,4,3,4,4,4,4,4,4,4,4,4,4, > Jenkins:4,5,4,6,3,5,0,6,1,2,8,3,6,8,2,1, > #connections:128: > Toeplitz:8,8,8,8,8,7,9,8,8,8,8,8,8,8,8,8, > Jenkins:8,12,10,9,7,8,3,10,6,8,9,8,10,11,6,3, > > Throughput (Gbps) comparison: > #conn Toeplitz Jenkins > 32 26.6 23.2 > 64 32.1 23.4 > 128 29.1 24.1 > > For long term solution, I think we should put the Toeplitz hash as > another option to the generic hash function in kernel... But, for the > time being, can you accept this patch to fix the assumptions on > struct flow_keys layout? I find your Toeplitz distribution has an anomaly. Having 128 connections distributed almost _perfectly_ into 16 buckets is telling something how the source/destination ports where allocated maybe, knowing the RSS key or something ? It looks too _perfect_ to be true. Here what I get here from 20 runs of 128 sessions using prandom_u32() hash, distributed to 16 buckets (hash % 16) : 6,9,9,6,11,8,9,7,7,7,9,8,8,7,9,8 : 6,9,6,6,6,9,8,5,12,10,7,7,9,7,13,8 : 7,4,9,9,10,9,8,7,15,4,8,8,11,10,2,7 : 12,5,10,6,7,4,10,10,6,5,10,14,8,8,5,8 : 4,8,5,13,7,4,7,9,7,6,6,9,6,11,17,9 : 10,10,8,5,7,4,5,14,6,9,9,7,8,9,7,10 : 6,4,9,10,13,8,8,7,6,5,8,9,7,5,15,8 : 11,13,7,4,8,6,6,9,10,8,8,5,6,6,11,10 : 8,8,11,7,12,13,5,8,9,6,8,10,5,4,9,5 : 13,5,5,4,5,11,8,8,11,8,9,10,10,6,9,6 : 13,6,12,6,6,7,4,9,5,14,9,12,9,4,4,8 : 4,9,10,12,10,4,8,6,8,5,14,10,5,8,8,7 : 7,7,6,6,12,13,8,12,7,6,8,9,6,5,12,4 : 4,12,9,10,2,12,10,13,5,8,4,6,8,10,4,11 : 5,6,10,10,10,9,16,8,8,7,4,10,7,6,6,6 : 9,13,10,11,6,9,4,7,7,9,7,6,9,9,7,5 : 8,7,4,8,6,9,9,8,7,10,8,10,17,7,5,5 : 10,5,10,8,9,5,9,6,12,8,5,8,7,9,7,10 : 8,10,10,7,10,7,13,3,9,5,7,2,10,9,12,6 : 4,6,13,6,6,6,12,9,11,5,7,10,9,8,11,5 This looks more 'random' to me, and _if_ I use Jenkins hash I have the same distribution. Sure, it is not 'perfectly spread', but who said that all flows are sending the same amount of traffic in the real world ? Using Toeplitz hash is adding a cost of 300 ns per IPV6 packet. TCP_RR (small RPC) workload would certainly not like to compute Toeplitz for every packet. I would like we do not add complexity just to make some benchmark better.
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web