Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1398143 > unrolled thread
| Started by | Paolo Abeni <pabeni@redhat.com> |
|---|---|
| First post | 2016-05-10 16:20 +0200 |
| Last post | 2016-05-10 22:50 +0200 |
| Articles | 20 on this page of 37 — 8 participants |
Back to article view | Back to linux.kernel
[RFC PATCH 0/2] net: threadable napi poll loop Paolo Abeni <pabeni@redhat.com> - 2016-05-10 16:20 +0200
[RFC PATCH 2/2] net: add sysfs attribute to control napi threaded mode Paolo Abeni <pabeni@redhat.com> - 2016-05-10 16:20 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-10 16:40 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop David Miller <davem@davemloft.net> - 2016-05-10 18:00 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-10 18:10 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Paolo Abeni <pabeni@redhat.com> - 2016-05-10 22:30 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop David Miller <davem@davemloft.net> - 2016-05-10 22:50 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop David Miller <davem@davemloft.net> - 2016-05-10 23:00 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Rik van Riel <riel@redhat.com> - 2016-05-10 23:10 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Rik van Riel <riel@redhat.com> - 2016-05-10 23:00 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Paolo Abeni <pabeni@redhat.com> - 2016-05-10 18:10 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Hannes Frederic Sowa <hannes@stressinduktion.org> - 2016-05-10 22:50 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <edumazet@google.com> - 2016-05-10 23:10 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-10 23:40 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Rik van Riel <riel@redhat.com> - 2016-05-10 23:40 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-11 00:00 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Rik van Riel <riel@redhat.com> - 2016-05-11 00:10 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-11 00:10 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-11 00:50 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-11 20:00 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Hannes Frederic Sowa <hannes@stressinduktion.org> - 2016-05-11 00:40 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-11 01:00 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Peter Zijlstra <peterz@infradead.org> - 2016-05-11 09:00 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Hannes Frederic Sowa <hannes@stressinduktion.org> - 2016-05-11 15:20 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <edumazet@google.com> - 2016-05-11 16:50 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Rik van Riel <riel@redhat.com> - 2016-05-11 17:10 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-11 18:00 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-12 00:00 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Paolo Abeni <pabeni@redhat.com> - 2016-05-11 11:50 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <eric.dumazet@gmail.com> - 2016-05-11 15:10 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Hannes Frederic Sowa <hannes@stressinduktion.org> - 2016-05-11 15:40 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Hannes Frederic Sowa <hannes@stressinduktion.org> - 2016-05-11 15:50 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Paolo Abeni <pabeni@redhat.com> - 2016-05-11 16:40 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Eric Dumazet <edumazet@google.com> - 2016-05-11 16:50 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Hannes Frederic Sowa <hannes@stressinduktion.org> - 2016-05-12 00:50 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Thomas Gleixner <tglx@linutronix.de> - 2016-05-10 18:00 +0200
Re: [RFC PATCH 0/2] net: threadable napi poll loop Paolo Abeni <pabeni@redhat.com> - 2016-05-10 22:50 +0200
Page 1 of 2 [1] 2 Next page →
| From | Paolo Abeni <pabeni@redhat.com> |
|---|---|
| Date | 2016-05-10 16:20 +0200 |
| Subject | [RFC PATCH 0/2] net: threadable napi poll loop |
| Message-ID | <rxgFY-3QP-7@gated-at.bofh.it> |
Currently, the softirq loop can be scheduled both inside the ksofirqd kernel thread and inside any running process. This makes nearly impossible for the process scheduler to balance in a fair way the amount of time that a given core spends performing the softirq loop. Under high network load, the softirq loop can take nearly 100% of a given CPU, leaving very little time for use space processing. On single core hosts, this means that the user space can nearly starve; for example super_netperf UDP_STREAM tests towards a remote single core vCPU guest[1] can measure an aggregated throughput of a few thousands pps, and the same behavior can be reproduced even on bare-metal, eventually simulating a single core with taskset and/or sysfs configuration. This patch series allows the administrator to let the napi poll loop run inside its own kernel thread, a thread for each napi instance, while retaining the default, softirq-based behavior. The RPS mechanism is currently not affected. When the napi poll loop is run inside a proper kernel thread, the process scheduler can fairly balance the rx job between the user space application and the kernel and give the administrator the ability to manage the network workload with scheduler tools and configuration. With the default scheduling policy, the starvation issue observed on single vCPU guest under UDP flood is solved and the throughput measured under heavy overload is quite stable around the peak performance. In the remote host to VM scenario, running even the hypervisor napi poll loop in threaded mode gives additional benefit, since the process scheduler can more easily avoid cpu conflict between the VM process and the kernel thread processing the rx packets. The raw numbers, obtained with the super_neterf UDP_STREAM test, in a remote host to VM scenario, using a tun device with a noqueue qdisc in the hypervisor and using 'sdfn' for the rx flow hash on the ingress device, are as follow: vanilla guest threaded both hypevisor and guest threaded size/flow kpps kpps/delta kpps/delta 1/1 746 901/+20% 1024/+37% 1/25 185 585/+215% 789/+325% 1/50 330 642/+94% 843/+155% 1/100 180 662/+267% 872/+383% 1/200 177 672/+279% 812/+358% 64/1 707 1042/+47% 1062/+50% 64/25 320 586/+83% 746/+132% 64/50 195 648/+232% 761/+290% 64/100 221 666/+200% 787/+255% 64/200 186 688/+268% 793/+325% 256/1 475 777/+63% 809/+70% 256/25 303 589/+83% 860/+183% 256/50 308 584/+89% 825/+168% 256/100 268 698/+159% 785/+191% 256/200 186 656/+398% 795/+503% 1438/1 619 664/+7% 640/+3% 1438/25 519 766/+47% 829/+59% 1438/50 451 712/+57% 820/+81% 1438/100 294 759/+158% 797/+170% 1438/200 262 728/+177% 769/+193% 4096/1 176 207/+17% 200/+13% 4096/25 225 275/+22% 286/+27% 4096/50 212 272/+28% 283/+33% 4096/100 168 264/+57% 283/+68% 4096/200 134 240/+78% 273/+102% 64000/1 16 18/+13% 18/+13% 64000/25 18 18/0 18/0 64000/50 18 18/0 18/0 64000/100 18 18/0 18/0 64000/200 15 15/0 15/0 This patchset is a first RFC but in the long run we would like to move more and more NAPI instances into kthreads. The kthread approach should give a lot of new advantages over the softirq based approach: * moving into a more dpdk-alike busy poll packet processing direction: we can even use busy polling without the need of a connected UDP or TCP socket and can leverage busy polling for forwarding setups. This could very well increase latency and packet throughput without hurting other processes if the networking stack gets more and more preemptive in the future. * possibility to acquire mutexes in the networking processing path: e.g. we would need that to configure hw_breakpoints if we want to add watchpoints in the memory based on some rules in the kernel * more and better tooling to adjust the weight of the networking kthreads, preferring certain networking cards or setting cpus affinity on packet processing threads. Maybe also using deadline scheduling or other scheduler features might be worthwhile. * scheduler statistics can be used to observe network packet processing At this point we are not really sure if we should go with this simpler approach by putting NAPI itself into kthreads or leverage the threadirqs function by putting the whole interrupt into a thread and signaling NAPI that it does not reschedule itself in a softirq but to simply run at this particular context of the interrupt handler. While the threaded irq way seems to better integrate into the kernel and also other devices could move their interrupts into the threads easily on a common policy, we don't know how to really express the necessary knobs with the current device driver model (module parameters, sysfs attributes, etc.). This is where we would like to hear some opinions. NAPI would e.g. have to query the kernel if the particular IRQ/MSI if it should be scheduled in a softirq or in a thread, so we don't have to rewrite all device drivers. This might even be needed on a per rx-queue granularity. [1] when the flows are processed by the hypervisor on different rx queues, i.e. the flows use different source/destination IPs or the hypervisor uses the L4 header to compute the rx hash. Paolo Abeni (2): net: implement threaded-able napi poll loop support net: add sysfs attribute to control napi threaded mode include/linux/netdevice.h | 4 ++ net/core/dev.c | 113 ++++++++++++++++++++++++++++++++++++++++++++++ net/core/net-sysfs.c | 102 +++++++++++++++++++++++++++++++++++++++++ 3 files changed, 219 insertions(+) -- 1.8.3.1
[toc] | [next] | [standalone]
| From | Paolo Abeni <pabeni@redhat.com> |
|---|---|
| Date | 2016-05-10 16:20 +0200 |
| Subject | [RFC PATCH 2/2] net: add sysfs attribute to control napi threaded mode |
| Message-ID | <rxgFZ-3QP-23@gated-at.bofh.it> |
| In reply to | #1398143 |
this patch addis a new sysfs attribute to the network
device class. Said attribute is a bitmask that allows controlling
the threaded mode for all the napi instances of the given
network device.
The threaded mode can be switched only if related network device
is down.
Signed-off-by: Paolo Abeni <pabeni@redhat.com>
Signed-off-by: Hannes Frederic Sowa <hannes@stressinduktion.org>
---
net/core/net-sysfs.c | 102 +++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 102 insertions(+)
diff --git a/net/core/net-sysfs.c b/net/core/net-sysfs.c
index 2b3f76f..60bc768 100644
--- a/net/core/net-sysfs.c
+++ b/net/core/net-sysfs.c
@@ -489,6 +489,107 @@ static ssize_t phys_switch_id_show(struct device *dev,
}
static DEVICE_ATTR_RO(phys_switch_id);
+unsigned long *__alloc_thread_bitmap(struct net_device *netdev, int *bits)
+{
+ struct napi_struct *n;
+
+ *bits = 0;
+ list_for_each_entry(n, &netdev->napi_list, dev_list)
+ (*bits)++;
+
+ return kmalloc_array(BITS_TO_LONGS(*bits), sizeof(unsigned long),
+ GFP_ATOMIC | __GFP_ZERO);
+}
+
+static ssize_t threaded_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
+{
+ struct net_device *netdev = to_net_dev(dev);
+ struct napi_struct *n;
+ unsigned long *bmap;
+ size_t count = 0;
+ int i, bits;
+
+ if (!rtnl_trylock())
+ return restart_syscall();
+
+ if (!dev_isalive(netdev))
+ goto unlock;
+
+ bmap = __alloc_thread_bitmap(netdev, &bits);
+ if (!bmap) {
+ count = -ENOMEM;
+ goto unlock;
+ }
+
+ i = 0;
+ list_for_each_entry(n, &netdev->napi_list, dev_list) {
+ if (test_bit(NAPI_STATE_THREADED, &n->state))
+ set_bit(i, bmap);
+ i++;
+ }
+
+ count = bitmap_print_to_pagebuf(true, buf, bmap, bits);
+ kfree(bmap);
+
+unlock:
+ rtnl_unlock();
+
+ return count;
+}
+
+static ssize_t threaded_store(struct device *dev,
+ struct device_attribute *attr,
+ const char *buf, size_t len)
+{
+ struct net_device *netdev = to_net_dev(dev);
+ struct napi_struct *n;
+ unsigned long *bmap;
+ int i, bits;
+ size_t ret;
+
+ if (!capable(CAP_NET_ADMIN))
+ return -EPERM;
+
+ if (!rtnl_trylock())
+ return restart_syscall();
+
+ if (!dev_isalive(netdev)) {
+ ret = len;
+ goto unlock;
+ }
+
+ if (netdev->flags & IFF_UP) {
+ ret = -EBUSY;
+ goto unlock;
+ }
+
+ bmap = __alloc_thread_bitmap(netdev, &bits);
+ if (!bmap) {
+ ret = -ENOMEM;
+ goto unlock;
+ }
+
+ ret = bitmap_parselist(buf, bmap, bits);
+ if (ret)
+ goto free_unlock;
+
+ i = 0;
+ list_for_each_entry(n, &netdev->napi_list, dev_list) {
+ napi_set_threaded(n, test_bit(i, bmap));
+ i++;
+ }
+ ret = len;
+
+free_unlock:
+ kfree(bmap);
+
+unlock:
+ rtnl_unlock();
+ return ret;
+}
+static DEVICE_ATTR_RW(threaded);
+
static struct attribute *net_class_attrs[] = {
&dev_attr_netdev_group.attr,
&dev_attr_type.attr,
@@ -517,6 +618,7 @@ static struct attribute *net_class_attrs[] = {
&dev_attr_phys_port_name.attr,
&dev_attr_phys_switch_id.attr,
&dev_attr_proto_down.attr,
+ &dev_attr_threaded.attr,
NULL,
};
ATTRIBUTE_GROUPS(net_class);
--
1.8.3.1
[toc] | [prev] | [next] | [standalone]
| From | Eric Dumazet <eric.dumazet@gmail.com> |
|---|---|
| Date | 2016-05-10 16:40 +0200 |
| Message-ID | <rxgZk-42L-29@gated-at.bofh.it> |
| In reply to | #1398143 |
On Tue, 2016-05-10 at 16:11 +0200, Paolo Abeni wrote: > Currently, the softirq loop can be scheduled both inside the ksofirqd kernel > thread and inside any running process. This makes nearly impossible for the > process scheduler to balance in a fair way the amount of time that > a given core spends performing the softirq loop. > > Under high network load, the softirq loop can take nearly 100% of a given CPU, > leaving very little time for use space processing. On single core hosts, this > means that the user space can nearly starve; for example super_netperf > UDP_STREAM tests towards a remote single core vCPU guest[1] can measure an > aggregated throughput of a few thousands pps, and the same behavior can be > reproduced even on bare-metal, eventually simulating a single core with taskset > and/or sysfs configuration. I hate these patches and ideas guys, sorry. That is before my breakfast, but still... I have enough hard time dealing with loads where ksoftirqd has to compete with user threads that thought that playing with priorities was a nice idea. Guess what, when they lose networking they complain. We already have ksoftirqd to normally cope with the case you are describing. If it is not working as intended, please identify the bugs and fix them, instead of adding yet another tests in fast path and extra complexity in the stack. In the one vcpu case, allowing the user thread to consume more UDP packets from the target UDP socket will also make your NIC drop more packets, that are not necessarily packets for the same socket. So you are shifting the attack to a different target, at the expense of more kernel bloat.
[toc] | [prev] | [next] | [standalone]
| From | David Miller <davem@davemloft.net> |
|---|---|
| Date | 2016-05-10 18:00 +0200 |
| Message-ID | <rxieJ-5nU-1@gated-at.bofh.it> |
| In reply to | #1398158 |
From: Eric Dumazet <eric.dumazet@gmail.com> Date: Tue, 10 May 2016 07:29:50 -0700 > We already have ksoftirqd to normally cope with the case you are > describing. > > If it is not working as intended, please identify the bugs and fix them, > instead of adding yet another tests in fast path and extra complexity in > the stack. +1 Indeed, if ksoftirqd is not doing it's job, please fix it. It is designed exactly to deal with the problems described here.
[toc] | [prev] | [next] | [standalone]
| From | Eric Dumazet <eric.dumazet@gmail.com> |
|---|---|
| Date | 2016-05-10 18:10 +0200 |
| Message-ID | <rxior-5TU-35@gated-at.bofh.it> |
| In reply to | #1398158 |
On Tue, 2016-05-10 at 18:03 +0200, Paolo Abeni wrote: > If a single core host is under network flood, i.e. ksoftirqd is > scheduled and it eventually (after processing ~640 packets) will let the > user space process run. The latter will execute a syscall to receive a > packet, which will have to disable/enable bh at least once and that will > cause the processing of another ~640 packets. To receive a single packet > in user space, the kernel has to process more than one thousand packets. Looks you found the bug then. Have you tried to fix it ?
[toc] | [prev] | [next] | [standalone]
| From | Paolo Abeni <pabeni@redhat.com> |
|---|---|
| Date | 2016-05-10 22:30 +0200 |
| Message-ID | <rxms2-1tG-3@gated-at.bofh.it> |
| In reply to | #1398282 |
On Tue, 2016-05-10 at 09:08 -0700, Eric Dumazet wrote: > On Tue, 2016-05-10 at 18:03 +0200, Paolo Abeni wrote: > > > If a single core host is under network flood, i.e. ksoftirqd is > > scheduled and it eventually (after processing ~640 packets) will let the > > user space process run. The latter will execute a syscall to receive a > > packet, which will have to disable/enable bh at least once and that will > > cause the processing of another ~640 packets. To receive a single packet > > in user space, the kernel has to process more than one thousand packets. > > Looks you found the bug then. Have you tried to fix it ? The core functionality is implemented in ~100 lines of code, is that the kind of bloat that do concerns you ? That could probably be improved removing some code duplication, i.e. factorizing napi_thread_wait() with irq_wait_for_interrupt() and possibly napi_threaded_poll() with net_rx_action(). If the additional test inside napi_schedule() is really scaring, it can be guarded with a static_key. The ksoftirq and the local_bh_enable() design are the root of the problem, they need to be touched/affected to solve it. We actually experimented several different options. Limiting the amount of work performed by local_bh_enable() somewhat mitigate the issue, but it adds just another kernel parameter difficult to be tuned. Running the softirq loop exclusively inside the ksoftirqd will solve the issue, but this is a very invasive approach, affecting all others subsystem. The above can be restricted to the net_rx_action only (i.e. running net_rx_action always in ksoftirqd context). The related patch isn't really much simpler than this and will add at least the same number of additional tests in fast path. Running the napi loop in a thread that can be migrated gives additional benefit in the hyper-visor/VM scenario, which can't be achieved elsewhere. Would you consider the threaded irq alternative more viable ? Cheers, Paolo
[toc] | [prev] | [next] | [standalone]
| From | David Miller <davem@davemloft.net> |
|---|---|
| Date | 2016-05-10 22:50 +0200 |
| Message-ID | <rxmLo-1CV-11@gated-at.bofh.it> |
| In reply to | #1398490 |
From: Paolo Abeni <pabeni@redhat.com> Date: Tue, 10 May 2016 22:22:50 +0200 > On Tue, 2016-05-10 at 09:08 -0700, Eric Dumazet wrote: >> On Tue, 2016-05-10 at 18:03 +0200, Paolo Abeni wrote: >> >> > If a single core host is under network flood, i.e. ksoftirqd is >> > scheduled and it eventually (after processing ~640 packets) will let the >> > user space process run. The latter will execute a syscall to receive a >> > packet, which will have to disable/enable bh at least once and that will >> > cause the processing of another ~640 packets. To receive a single packet >> > in user space, the kernel has to process more than one thousand packets. >> >> Looks you found the bug then. Have you tried to fix it ? ... > The ksoftirq and the local_bh_enable() design are the root of the > problem, they need to be touched/affected to solve it. That's not what I read from your description, processing 640 packets before going to ksoftirqd seems to the be the absolute root problem.
[toc] | [prev] | [next] | [standalone]
| From | David Miller <davem@davemloft.net> |
|---|---|
| Date | 2016-05-10 23:00 +0200 |
| Message-ID | <rxmV5-1JH-13@gated-at.bofh.it> |
| In reply to | #1398516 |
From: Rik van Riel <riel@redhat.com> Date: Tue, 10 May 2016 16:50:56 -0400 > On Tue, 2016-05-10 at 16:45 -0400, David Miller wrote: >> From: Paolo Abeni <pabeni@redhat.com> >> Date: Tue, 10 May 2016 22:22:50 +0200 >> >> > On Tue, 2016-05-10 at 09:08 -0700, Eric Dumazet wrote: >> >> On Tue, 2016-05-10 at 18:03 +0200, Paolo Abeni wrote: >> >> >> >> > If a single core host is under network flood, i.e. ksoftirqd is >> >> > scheduled and it eventually (after processing ~640 packets) will >> let the >> >> > user space process run. The latter will execute a syscall to >> receive a >> >> > packet, which will have to disable/enable bh at least once and >> that will >> >> > cause the processing of another ~640 packets. To receive a >> single packet >> >> > in user space, the kernel has to process more than one thousand >> packets. >> >> >> >> Looks you found the bug then. Have you tried to fix it ? >> ... >> > The ksoftirq and the local_bh_enable() design are the root of the >> > problem, they need to be touched/affected to solve it. >> >> That's not what I read from your description, processing 640 packets >> before going to ksoftirqd seems to the be the absolute root problem. > > What would a fix for that look like? > > Keep track of the number of processed incoming packets, > and the number of packets handed off, and defer to > ksoftirqd earlier if the statistics suggest packets are > getting dropped on the floor? Not by packet count but by something more easily to measure and scalable to fairness like processing time.
[toc] | [prev] | [next] | [standalone]
| From | Rik van Riel <riel@redhat.com> |
|---|---|
| Date | 2016-05-10 23:10 +0200 |
| Message-ID | <rxn4K-2dm-19@gated-at.bofh.it> |
| In reply to | #1398521 |
[Multipart message — attachments visible in raw view] — view raw
On Tue, 2016-05-10 at 16:52 -0400, David Miller wrote: > From: Rik van Riel <riel@redhat.com> > Date: Tue, 10 May 2016 16:50:56 -0400 > > > On Tue, 2016-05-10 at 16:45 -0400, David Miller wrote: > >> From: Paolo Abeni <pabeni@redhat.com> > >> Date: Tue, 10 May 2016 22:22:50 +0200 > >> > >> > On Tue, 2016-05-10 at 09:08 -0700, Eric Dumazet wrote: > >> >> On Tue, 2016-05-10 at 18:03 +0200, Paolo Abeni wrote: > >> >> > >> >> > If a single core host is under network flood, i.e. ksoftirqd > is > >> >> > scheduled and it eventually (after processing ~640 packets) > will > >> let the > >> >> > user space process run. The latter will execute a syscall to > >> receive a > >> >> > packet, which will have to disable/enable bh at least once > and > >> that will > >> >> > cause the processing of another ~640 packets. To receive a > >> single packet > >> >> > in user space, the kernel has to process more than one > thousand > >> packets. > >> >> > >> >> Looks you found the bug then. Have you tried to fix it ? > >> ... > >> > The ksoftirq and the local_bh_enable() design are the root of > the > >> > problem, they need to be touched/affected to solve it. > >> > >> That's not what I read from your description, processing 640 > packets > >> before going to ksoftirqd seems to the be the absolute root > problem. > > > > What would a fix for that look like? > > > > Keep track of the number of processed incoming packets, > > and the number of packets handed off, and defer to > > ksoftirqd earlier if the statistics suggest packets are > > getting dropped on the floor? > > Not by packet count but by something more easily to measure and > scalable to fairness like processing time. I need to get back to fixing irq & softirq time accounting, which does not currently work correctly in all time keeping modes... -- All Rights Reversed.
[toc] | [prev] | [next] | [standalone]
| From | Rik van Riel <riel@redhat.com> |
|---|---|
| Date | 2016-05-10 23:00 +0200 |
| Message-ID | <rxmV5-1JH-15@gated-at.bofh.it> |
| In reply to | #1398516 |
[Multipart message — attachments visible in raw view] — view raw
On Tue, 2016-05-10 at 16:45 -0400, David Miller wrote: > From: Paolo Abeni <pabeni@redhat.com> > Date: Tue, 10 May 2016 22:22:50 +0200 > > > On Tue, 2016-05-10 at 09:08 -0700, Eric Dumazet wrote: > >> On Tue, 2016-05-10 at 18:03 +0200, Paolo Abeni wrote: > >> > >> > If a single core host is under network flood, i.e. ksoftirqd is > >> > scheduled and it eventually (after processing ~640 packets) will > let the > >> > user space process run. The latter will execute a syscall to > receive a > >> > packet, which will have to disable/enable bh at least once and > that will > >> > cause the processing of another ~640 packets. To receive a > single packet > >> > in user space, the kernel has to process more than one thousand > packets. > >> > >> Looks you found the bug then. Have you tried to fix it ? > ... > > The ksoftirq and the local_bh_enable() design are the root of the > > problem, they need to be touched/affected to solve it. > > That's not what I read from your description, processing 640 packets > before going to ksoftirqd seems to the be the absolute root problem. What would a fix for that look like? Keep track of the number of processed incoming packets, and the number of packets handed off, and defer to ksoftirqd earlier if the statistics suggest packets are getting dropped on the floor? Is there a cheap way to do that kind of thing? -- All Rights Reversed.
[toc] | [prev] | [next] | [standalone]
| From | Paolo Abeni <pabeni@redhat.com> |
|---|---|
| Date | 2016-05-10 18:10 +0200 |
| Message-ID | <rxior-5TU-37@gated-at.bofh.it> |
| In reply to | #1398158 |
Hi, On Tue, 2016-05-10 at 07:29 -0700, Eric Dumazet wrote: > On Tue, 2016-05-10 at 16:11 +0200, Paolo Abeni wrote: > > Currently, the softirq loop can be scheduled both inside the ksofirqd kernel > > thread and inside any running process. This makes nearly impossible for the > > process scheduler to balance in a fair way the amount of time that > > a given core spends performing the softirq loop. > > > > Under high network load, the softirq loop can take nearly 100% of a given CPU, > > leaving very little time for use space processing. On single core hosts, this > > means that the user space can nearly starve; for example super_netperf > > UDP_STREAM tests towards a remote single core vCPU guest[1] can measure an > > aggregated throughput of a few thousands pps, and the same behavior can be > > reproduced even on bare-metal, eventually simulating a single core with taskset > > and/or sysfs configuration. > > I hate these patches and ideas guys, sorry. That is before my breakfast, > but still... I'm sorry, I did not meant to spoil your breakfast ;-) > I have enough hard time dealing with loads where ksoftirqd has to > compete with user threads that thought that playing with priorities was > a nice idea. I fear there is a misunderstanding. I'm not suggesting to fiddle with priorities; the above 'taskset' reference was just an hint to replicate the starvation issue on bare-metal in the lack of a single core host. > > Guess what, when they lose networking they complain. > > We already have ksoftirqd to normally cope with the case you are > describing. > > If it is not working as intended, please identify the bugs and fix them, > instead of adding yet another tests in fast path and extra complexity in > the stack. The idea it exactly that: the problem is how the softirq loop is scheduled and executed, i.e. the current ksoftirqd/"inline loop" model. If a single core host is under network flood, i.e. ksoftirqd is scheduled and it eventually (after processing ~640 packets) will let the user space process run. The latter will execute a syscall to receive a packet, which will have to disable/enable bh at least once and that will cause the processing of another ~640 packets. To receive a single packet in user space, the kernel has to process more than one thousand packets. AFAICS it can't be solved without changing how the net_rx_action is served. The user space starvation issue don't affect large server, but AFAIK many small devices have a lot of out-of-tree hacks to cope with this sort of issues. In the VM scenario, the starvation issue was not a real concern up to a little time ago because the vhost/tun device was not able to push packets fast enough into the guest to trigger the issue. Recent improvements have changed the situation. Also, the scheduler's ability to migrate the napi threads is quite beneficial for hypervisor when the VMs are receiving a lot of network traffic. Please have a look at the performance numbers. The current patch adds a single, simple, test per napi_schedule invocation, and with minimal changes, the kernel won't access any additional cache-line when the napi thread is disabled. Even in the current form, in my tests no regression is seen with the patched kernel when the napi thread mode is disabled. > In the one vcpu case, allowing the user thread to consume more UDP > packets from the target UDP socket will also make your NIC drop more > packets, that are not necessarily packets for the same socket. That is true. But the threaded napi will not starve, i.e. the forwarding process, to a nearly zero packet rate, while with the current code the reverse scenario can happen. Cheers, Paolo > > So you are shifting the attack to a different target, > at the expense of more kernel bloat. > > >
[toc] | [prev] | [next] | [standalone]
| From | Hannes Frederic Sowa <hannes@stressinduktion.org> |
|---|---|
| Date | 2016-05-10 22:50 +0200 |
| Message-ID | <rxmLo-1CV-7@gated-at.bofh.it> |
| In reply to | #1398158 |
Hello, On 10.05.2016 16:29, Eric Dumazet wrote: > On Tue, 2016-05-10 at 16:11 +0200, Paolo Abeni wrote: >> Currently, the softirq loop can be scheduled both inside the ksofirqd kernel >> thread and inside any running process. This makes nearly impossible for the >> process scheduler to balance in a fair way the amount of time that >> a given core spends performing the softirq loop. >> >> Under high network load, the softirq loop can take nearly 100% of a given CPU, >> leaving very little time for use space processing. On single core hosts, this >> means that the user space can nearly starve; for example super_netperf >> UDP_STREAM tests towards a remote single core vCPU guest[1] can measure an >> aggregated throughput of a few thousands pps, and the same behavior can be >> reproduced even on bare-metal, eventually simulating a single core with taskset >> and/or sysfs configuration. > > I hate these patches and ideas guys, sorry. That is before my breakfast, > but still... :) > I have enough hard time dealing with loads where ksoftirqd has to > compete with user threads that thought that playing with priorities was > a nice idea. We tried a lot of approaches so far and this seemed to be the best architectural RFC we could post. I was quite surprised to see such good performance numbers with threaded NAPI, thus I think it could be a way forward. Your mentioned problem above seems to be a configuration mistake, no? Otherwise isn't that something user space/cgroups might solve? > Guess what, when they lose networking they complain. > > We already have ksoftirqd to normally cope with the case you are > describing. Indeed, but the time until we wake up ksoftirqd can be already quite long and for every packet we get in udp_recvmsg the local_bh_enable call let's us pick up quite a lot of new packets, which we drop before user space can make any progress. By being more fair between user space and "napid" we hoped to solve this. We also want more feedback from the scheduler people, so we Cc'ed them also. > If it is not working as intended, please identify the bugs and fix them, > instead of adding yet another tests in fast path and extra complexity in > the stack. We could use _local_bh_enable instead of local_bh_enable in udp_recvmsg, which certainly wouldn't branch down to softirqs as often, but this feels wrong to me and certainly is. After the discussion on netdev@ with Peter Hurley here [1] about "Softirq priority inversion from "softirq: reduce latencies"", we didn't want to propose some patch looking like this again, but this could help. The idea would be to limit the number we recheck for softirqs but give back control to user space. [1] https://lkml.org/lkml/2016/2/27/152 If I remember local_bh_enable in kernel-rt processes one softirq directly and defers its work to ksoftirqd much more quickly. > In the one vcpu case, allowing the user thread to consume more UDP > packets from the target UDP socket will also make your NIC drop more > packets, that are not necessarily packets for the same socket. > > So you are shifting the attack to a different target, > at the expense of more kernel bloat. I agree here, but I don't think this patch particularly is a lot of bloat and something very interesting people can play with and extend upon. Thanks, Hannes
[toc] | [prev] | [next] | [standalone]
| From | Eric Dumazet <edumazet@google.com> |
|---|---|
| Date | 2016-05-10 23:10 +0200 |
| Message-ID | <rxn4J-2dm-1@gated-at.bofh.it> |
| In reply to | #1398515 |
On Tue, May 10, 2016 at 1:46 PM, Hannes Frederic Sowa <hannes@stressinduktion.org> wrote: > I agree here, but I don't think this patch particularly is a lot of > bloat and something very interesting people can play with and extend upon. > Sure, very rarely patch authors think their stuff is bloat. I prefer to fix kernel softirq.c, or at least show me that you tried hard enough. I am pretty sure that the following would work : When ksoftirqd is scheduled, remember this in a per cpu variable (ksoftiqd_scheduled) When enabling BH , do not call do_softirq() if this variable is set. ksoftirqd would clear the variable at the right place (probably in run_ksoftirqd()) Sure, this might add a lot of latency regressions, but lets fix them.
[toc] | [prev] | [next] | [standalone]
| From | Eric Dumazet <eric.dumazet@gmail.com> |
|---|---|
| Date | 2016-05-10 23:40 +0200 |
| Message-ID | <rxnxM-2xf-3@gated-at.bofh.it> |
| In reply to | #1398523 |
On Tue, 2016-05-10 at 14:09 -0700, Eric Dumazet wrote:
> On Tue, May 10, 2016 at 1:46 PM, Hannes Frederic Sowa
> <hannes@stressinduktion.org> wrote:
>
> > I agree here, but I don't think this patch particularly is a lot of
> > bloat and something very interesting people can play with and extend upon.
> >
>
> Sure, very rarely patch authors think their stuff is bloat.
>
> I prefer to fix kernel softirq.c, or at least show me that you tried
> hard enough.
>
> I am pretty sure that the following would work :
>
> When ksoftirqd is scheduled, remember this in a per cpu variable
> (ksoftiqd_scheduled)
>
> When enabling BH , do not call do_softirq() if this variable is set.
>
> ksoftirqd would clear the variable at the right place (probably in
> run_ksoftirqd())
>
> Sure, this might add a lot of latency regressions, but lets fix them.
Only to give the idea (it is completely untested and probably buggy)
diff --git a/kernel/softirq.c b/kernel/softirq.c
index 17caf4b63342..cb30cfd76687 100644
--- a/kernel/softirq.c
+++ b/kernel/softirq.c
@@ -56,6 +56,7 @@ EXPORT_SYMBOL(irq_stat);
static struct softirq_action softirq_vec[NR_SOFTIRQS] __cacheline_aligned_in_smp;
DEFINE_PER_CPU(struct task_struct *, ksoftirqd);
+DEFINE_PER_CPU(bool, ksoftirqd_scheduled);
const char * const softirq_to_name[NR_SOFTIRQS] = {
"HI", "TIMER", "NET_TX", "NET_RX", "BLOCK", "BLOCK_IOPOLL",
@@ -73,8 +74,10 @@ static void wakeup_softirqd(void)
/* Interrupts are disabled: no need to stop preemption */
struct task_struct *tsk = __this_cpu_read(ksoftirqd);
- if (tsk && tsk->state != TASK_RUNNING)
+ if (tsk && tsk->state != TASK_RUNNING) {
+ __this_cpu_write(ksoftirqd_scheduled, true);
wake_up_process(tsk);
+ }
}
/*
@@ -162,7 +165,9 @@ void __local_bh_enable_ip(unsigned long ip, unsigned int cnt)
*/
preempt_count_sub(cnt - 1);
- if (unlikely(!in_interrupt() && local_softirq_pending())) {
+ if (unlikely(!in_interrupt() &&
+ local_softirq_pending() &&
+ !__this_cpu_read(ksoftirqd_scheduled))) {
/*
* Run softirq if any pending. And do it in its own stack
* as we may be calling this deep in a task call stack already.
@@ -660,6 +665,8 @@ static void run_ksoftirqd(unsigned int cpu)
* in the task stack here.
*/
__do_softirq();
+ if (!local_softirq_pending())
+ __this_cpu_write(ksoftirqd_scheduled, false);
local_irq_enable();
cond_resched_rcu_qs();
return;
[toc] | [prev] | [next] | [standalone]
| From | Rik van Riel <riel@redhat.com> |
|---|---|
| Date | 2016-05-10 23:40 +0200 |
| Message-ID | <rxnxM-2xf-19@gated-at.bofh.it> |
| In reply to | #1398537 |
[Multipart message — attachments visible in raw view] — view raw
On Tue, 2016-05-10 at 14:31 -0700, Eric Dumazet wrote:
> On Tue, 2016-05-10 at 14:09 -0700, Eric Dumazet wrote:
> >
> > On Tue, May 10, 2016 at 1:46 PM, Hannes Frederic Sowa
> > <hannes@stressinduktion.org> wrote:
> >
> > >
> > > I agree here, but I don't think this patch particularly is a lot
> > > of
> > > bloat and something very interesting people can play with and
> > > extend upon.
> > >
> > Sure, very rarely patch authors think their stuff is bloat.
> >
> > I prefer to fix kernel softirq.c, or at least show me that you
> > tried
> > hard enough.
> >
> > I am pretty sure that the following would work :
> >
> > When ksoftirqd is scheduled, remember this in a per cpu variable
> > (ksoftiqd_scheduled)
> >
> > When enabling BH , do not call do_softirq() if this variable is
> > set.
> >
> > ksoftirqd would clear the variable at the right place (probably in
> > run_ksoftirqd())
> >
> > Sure, this might add a lot of latency regressions, but lets fix
> > them.
> Only to give the idea (it is completely untested and probably buggy)
>
> diff --git a/kernel/softirq.c b/kernel/softirq.c
> index 17caf4b63342..cb30cfd76687 100644
> --- a/kernel/softirq.c
> +++ b/kernel/softirq.c
> @@ -56,6 +56,7 @@ EXPORT_SYMBOL(irq_stat);
> static struct softirq_action softirq_vec[NR_SOFTIRQS]
> __cacheline_aligned_in_smp;
>
> DEFINE_PER_CPU(struct task_struct *, ksoftirqd);
> +DEFINE_PER_CPU(bool, ksoftirqd_scheduled);
>
> const char * const softirq_to_name[NR_SOFTIRQS] = {
> "HI", "TIMER", "NET_TX", "NET_RX", "BLOCK", "BLOCK_IOPOLL",
> @@ -73,8 +74,10 @@ static void wakeup_softirqd(void)
> /* Interrupts are disabled: no need to stop preemption */
> struct task_struct *tsk = __this_cpu_read(ksoftirqd);
>
> - if (tsk && tsk->state != TASK_RUNNING)
> + if (tsk && tsk->state != TASK_RUNNING) {
> + __this_cpu_write(ksoftirqd_scheduled, true);
> wake_up_process(tsk);
> + }
> }
>
> /*
> @@ -162,7 +165,9 @@ void __local_bh_enable_ip(unsigned long ip,
> unsigned int cnt)
> */
> preempt_count_sub(cnt - 1);
>
> - if (unlikely(!in_interrupt() && local_softirq_pending())) {
> + if (unlikely(!in_interrupt() &&
> + local_softirq_pending() &&
> + !__this_cpu_read(ksoftirqd_scheduled))) {
> /*
>
You might need another one of these in invoke_softirq()
--
All Rights Reversed.
[toc] | [prev] | [next] | [standalone]
| From | Eric Dumazet <eric.dumazet@gmail.com> |
|---|---|
| Date | 2016-05-11 00:00 +0200 |
| Message-ID | <rxnR8-2Ki-3@gated-at.bofh.it> |
| In reply to | #1398539 |
On Tue, 2016-05-10 at 17:35 -0400, Rik van Riel wrote: > You might need another one of these in invoke_softirq() > Excellent. I gave it a quick try (without your suggestion), and host seems to survive a stress test. Of course we do have to fix these problems : [ 147.781629] NOHZ: local_softirq_pending 48 [ 147.785546] NOHZ: local_softirq_pending 48 [ 147.788344] NOHZ: local_softirq_pending 48 [ 147.788992] NOHZ: local_softirq_pending 48 [ 147.790943] NOHZ: local_softirq_pending 48 [ 147.791232] NOHZ: local_softirq_pending 24a [ 147.791258] NOHZ: local_softirq_pending 48 [ 147.791366] NOHZ: local_softirq_pending 48 [ 147.792118] NOHZ: local_softirq_pending 48 [ 147.793428] NOHZ: local_softirq_pending 48 Thanks.
[toc] | [prev] | [next] | [standalone]
| From | Rik van Riel <riel@redhat.com> |
|---|---|
| Date | 2016-05-11 00:10 +0200 |
| Message-ID | <rxo0N-3aD-1@gated-at.bofh.it> |
| In reply to | #1398548 |
[Multipart message — attachments visible in raw view] — view raw
On Tue, 2016-05-10 at 14:53 -0700, Eric Dumazet wrote: > On Tue, 2016-05-10 at 17:35 -0400, Rik van Riel wrote: > > > > > You might need another one of these in invoke_softirq() > > > Excellent. > > I gave it a quick try (without your suggestion), and host seems to > survive a stress test. > > Of course we do have to fix these problems : > > [ 147.781629] NOHZ: local_softirq_pending 48 > [ 147.785546] NOHZ: local_softirq_pending 48 > [ 147.788344] NOHZ: local_softirq_pending 48 > [ 147.788992] NOHZ: local_softirq_pending 48 > [ 147.790943] NOHZ: local_softirq_pending 48 > [ 147.791232] NOHZ: local_softirq_pending 24a > [ 147.791258] NOHZ: local_softirq_pending 48 > [ 147.791366] NOHZ: local_softirq_pending 48 > [ 147.792118] NOHZ: local_softirq_pending 48 > [ 147.793428] NOHZ: local_softirq_pending 48 As long as ksoftirqd is running, that should not be an actual problem, just a false positive. -- All Rights Reversed.
[toc] | [prev] | [next] | [standalone]
| From | Eric Dumazet <eric.dumazet@gmail.com> |
|---|---|
| Date | 2016-05-11 00:10 +0200 |
| Message-ID | <rxo0O-3aD-13@gated-at.bofh.it> |
| In reply to | #1398548 |
On Tue, 2016-05-10 at 14:53 -0700, Eric Dumazet wrote: > On Tue, 2016-05-10 at 17:35 -0400, Rik van Riel wrote: > > > You might need another one of these in invoke_softirq() > > > > Excellent. > > I gave it a quick try (without your suggestion), and host seems to > survive a stress test. > > Of course we do have to fix these problems : > > [ 147.781629] NOHZ: local_softirq_pending 48 > [ 147.785546] NOHZ: local_softirq_pending 48 > [ 147.788344] NOHZ: local_softirq_pending 48 > [ 147.788992] NOHZ: local_softirq_pending 48 > [ 147.790943] NOHZ: local_softirq_pending 48 > [ 147.791232] NOHZ: local_softirq_pending 24a > [ 147.791258] NOHZ: local_softirq_pending 48 > [ 147.791366] NOHZ: local_softirq_pending 48 > [ 147.792118] NOHZ: local_softirq_pending 48 > [ 147.793428] NOHZ: local_softirq_pending 48 Well, with your suggestion, these warnings disappear ;)
[toc] | [prev] | [next] | [standalone]
| From | Eric Dumazet <eric.dumazet@gmail.com> |
|---|---|
| Date | 2016-05-11 00:50 +0200 |
| Message-ID | <rxoDv-3zC-15@gated-at.bofh.it> |
| In reply to | #1398553 |
On Tue, 2016-05-10 at 15:02 -0700, Eric Dumazet wrote: > On Tue, 2016-05-10 at 14:53 -0700, Eric Dumazet wrote: > > On Tue, 2016-05-10 at 17:35 -0400, Rik van Riel wrote: > > > > > You might need another one of these in invoke_softirq() > > > > > > > Excellent. > > > > I gave it a quick try (without your suggestion), and host seems to > > survive a stress test. > > > > Of course we do have to fix these problems : > > > > [ 147.781629] NOHZ: local_softirq_pending 48 > > [ 147.785546] NOHZ: local_softirq_pending 48 > > [ 147.788344] NOHZ: local_softirq_pending 48 > > [ 147.788992] NOHZ: local_softirq_pending 48 > > [ 147.790943] NOHZ: local_softirq_pending 48 > > [ 147.791232] NOHZ: local_softirq_pending 24a > > [ 147.791258] NOHZ: local_softirq_pending 48 > > [ 147.791366] NOHZ: local_softirq_pending 48 > > [ 147.792118] NOHZ: local_softirq_pending 48 > > [ 147.793428] NOHZ: local_softirq_pending 48 > > > Well, with your suggestion, these warnings disappear ;) This is really nice. Under stress number of context switches is really small. ksoftirqd and my netserver compete equally to get the cpu cycles (on CPU0) lpaa23:~# vmstat 1 10 procs -----------memory---------- ---swap-- -----io---- -system-- ----cpu---- r b swpd free buff cache si so bi bo in cs us sy id wa 2 0 0 260668416 37240 2414428 0 0 21 0 329 349 0 3 96 0 1 0 0 260667904 37240 2414428 0 0 0 12 193126 1050 0 2 98 0 1 0 0 260667904 37240 2414428 0 0 0 0 194354 1056 0 2 98 0 1 0 0 260669104 37240 2414492 0 0 0 0 200897 1095 0 2 98 0 1 0 0 260668592 37240 2414492 0 0 0 0 205731 964 0 2 98 0 1 0 0 260678832 37240 2414492 0 0 0 0 201689 981 0 2 98 0 1 0 0 260678832 37240 2414492 0 0 0 0 204899 742 0 2 98 0 1 0 0 260678320 37240 2414492 0 0 0 0 199148 792 0 3 97 0 1 0 0 260678832 37240 2414492 0 0 0 0 196398 766 0 2 98 0 1 0 0 260678832 37240 2414492 0 0 0 0 201930 858 0 2 98 0 And we can see that ksoftirqd/0 runs for longer periods (~500 usec), instead of stupid 4 usec before the patch. Less overhead. lpaa23:~# cat /proc/3/sched ksoftirqd/0 (3, #threads: 1) ------------------------------------------------------------------- se.exec_start : 1552401.399526 se.vruntime : 237599.421560 se.sum_exec_runtime : 75432.494199 se.nr_migrations : 0 nr_switches : 144333 nr_voluntary_switches : 143828 nr_involuntary_switches : 505 se.load.weight : 1024 se.avg.load_sum : 10445 se.avg.util_sum : 10445 se.avg.load_avg : 0 se.avg.util_avg : 0 se.avg.last_update_time : 1552401399526 policy : 0 prio : 120 clock-delta : 47 lpaa23:~# echo 75432.494199/144333|bc -l .52262818758703830724 And yes indeed, user space can progress way faster under flood. lpaa23:~# nstat >/dev/null;sleep 1;nstat | grep Udp UdpInDatagrams 186132 0.0 UdpInErrors 735462 0.0 UdpOutDatagrams 10 0.0 UdpRcvbufErrors 735461 0.0
[toc] | [prev] | [next] | [standalone]
| From | Eric Dumazet <eric.dumazet@gmail.com> |
|---|---|
| Date | 2016-05-11 20:00 +0200 |
| Message-ID | <rxGAp-4yw-9@gated-at.bofh.it> |
| In reply to | #1398548 |
On Tue, 2016-05-10 at 14:53 -0700, Eric Dumazet wrote:
> On Tue, 2016-05-10 at 17:35 -0400, Rik van Riel wrote:
>
> > You might need another one of these in invoke_softirq()
> >
>
> Excellent.
>
> I gave it a quick try (without your suggestion), and host seems to
> survive a stress test.
Well, we instantly trigger rcu issues.
How to reproduce :
netserver &
for i in `seq 1 100`
do
netperf -H 127.0.0.1 -t TCP_RR -l 1000 &
done
# local hack to enable the new behavior
# without having to add a new sysctl, but hacking an existing one
echo 1001 >/proc/sys/net/core/netdev_max_backlog
<bang :>
[ 236.977511] INFO: rcu_sched self-detected stall on CPU
[ 236.977512] INFO: rcu_sched self-detected stall on CPU
[ 236.977515] INFO: rcu_sched self-detected stall on CPU
[ 236.977518] INFO: rcu_sched self-detected stall on CPU
[ 236.977519] INFO: rcu_sched self-detected stall on CPU
[ 236.977521] INFO: rcu_sched self-detected stall on CPU
[ 236.977522] INFO: rcu_sched self-detected stall on CPU
[ 236.977523] INFO: rcu_sched self-detected stall on CPU
[ 236.977525] INFO: rcu_sched self-detected stall on CPU
[ 236.977526] INFO: rcu_sched self-detected stall on CPU
[ 236.977527] INFO: rcu_sched self-detected stall on CPU
[ 236.977529] INFO: rcu_sched self-detected stall on CPU
[ 236.977530] INFO: rcu_sched self-detected stall on CPU
[ 236.977532] INFO: rcu_sched self-detected stall on CPU
[ 236.977532] 47-...: (1 GPs behind) idle=8d1/1/0 softirq=2500/2506 fqs=1
[ 236.977535] INFO: rcu_sched self-detected stall on CPU
[ 236.977536] INFO: rcu_sched self-detected stall on CPU
[ 236.977540] 36-...: (1 GPs behind) idle=d05/1/0 softirq=2637/2644 fqs=1
[ 236.977546]
[ 236.977546] 38-...: (1 GPs behind) idle=a5b/1/0 softirq=2612/2618 fqs=1
[ 236.977549] 0-...: (1 GPs behind) idle=c39/1/0 softirq=15315/15321 fqs=1
[ 236.977551] 24-...: (1 GPs behind) idle=ea3/1/0 softirq=2455/2461 fqs=1
[ 236.977554] 18-...: (20995 ticks this GP) idle=ef5/1/0 softirq=8530/8530 fqs=1
[ 236.977556] 39-...: (1 GPs behind) idle=f9d/1/0 softirq=2144/2150 fqs=1
[ 236.977558]
[ 236.977558] 22-...: (1 GPs behind) idle=5a7/1/0 softirq=10238/10244 fqs=1
[ 236.977561] 7-...: (1 GPs behind) idle=323/1/0 softirq=5279/5285 fqs=1
[ 236.977563] 31-...: (1 GPs behind) idle=47d/1/0 softirq=2526/2532 fqs=1
[ 236.977565] 33-...: (1 GPs behind) idle=175/1/0 softirq=2060/2066 fqs=1
[ 236.977568] 10-...: (1 GPs behind) idle=c3d/1/0 softirq=4864/4870 fqs=1
[ 236.977570] 34-...: (20995 ticks this GP) idle=dd5/1/0 softirq=2243/2243 fqs=1
[ 236.977574]
[ 236.977574] 37-...: (1 GPs behind) idle=aef/1/0 softirq=2660/2666 fqs=1
[ 236.977576] 13-...: (1 GPs behind) idle=a2b/1/0 softirq=9928/9934 fqs=1
[ 236.977578]
[ 236.977578]
[ 236.977579]
[ 236.977580]
[ 236.977582]
[ 236.977583]
[ 236.977583]
[ 236.977584]
[ 236.977584]
[ 236.977586]
[ 236.977587] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977588]
[ 236.977589]
[ 236.977595] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977603] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977607] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977609] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977610] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977612] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977614] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977616] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977618] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977619] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977620] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977622] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977626] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.977627] rcu_sched kthread starved for 20997 jiffies! g33049 c33048 f0x0 RCU_GP_WAIT_FQS(3) ->state=0x1
[ 236.978512] INFO: rcu_sched self-detected stall on CPU
[ 236.978512] INFO: rcu_sched self-detected stall on CPU
[ 236.978514] INFO: rcu_sched self-detected stall on CPU
[ 236.978516] INFO: rcu_sched self-detected stall on CPU
[ 236.978517] INFO: rcu_sched self-detected stall on CPU
[ 236.978518] INFO: rcu_sched self-detected stall on CPU
[ 236.978519] INFO: rcu_sched self-detected stall on CPU
[ 236.978520] INFO: rcu_sched self-detected stall on CPU
[ 236.978521] INFO: rcu_sched self-detected stall on CPU
[ 236.978522] INFO: rcu_sched self-detected stall on CPU
[ 236.978523] INFO: rcu_sched self-detected stall on CPU
[ 236.978524] INFO: rcu_sched self-detected stall on CPU
[ 236.978532] 45-...: (1 GPs behind) idle=8ed/1/0 softirq=3047/3053 fqs=1
[ 236.978534] 19-...: (20996 ticks this GP) idle=b5d/1/0 softirq=8157/8157 fqs=1
[ 236.978538] 17-...: (1 GPs behind) idle=5ad/1/0 softirq=7839/7845 fqs=1
[ 236.978539] 41-...: (1 GPs behind) idle=f4f/1/0 softirq=2345/2351 fqs=1
[ 236.978542] 6-...: (1 GPs behind) idle=a39/1/0 softirq=5492/5498 fqs=1
[ 236.978544] 30-...: (1 GPs behind) idle=c51/1/0 softirq=2499/2505 fqs=1
[ 236.978546] 5-...: (1 GPs behind) idle=917/1/0 softirq=5196/5202 fqs=1
[ 236.978548] 26-...: (20996 ticks this GP) idle=c61/1/0 softirq=2863/2863 fqs=1
[ 236.978550] 32-...: (1 GPs behind) idle=8db/1/0 softirq=2588/2594 fqs=1
[ 236.978552] 35-...: (1 GPs behind) idle=351/1/0 softirq=1869/1875 fqs=1
[ 236.978554] 8-...: (1 GPs behind) idle=221/1/0 softirq=5192/5198 fqs=1
[ 236.978556] 11-...: (1 GPs behind) idle=485/1/0 softirq=4480/4486 fqs=1
[ 236.978557]
[ 236.978558]
[ 236.978559]
[ 236.978560]
[ 236.978561]
Tentative proto / patch (not including Peter suggestions yet)
diff --git a/kernel/softirq.c b/kernel/softirq.c
index 17caf4b63342..be94e0241a70 100644
--- a/kernel/softirq.c
+++ b/kernel/softirq.c
@@ -56,6 +56,14 @@ EXPORT_SYMBOL(irq_stat);
static struct softirq_action softirq_vec[NR_SOFTIRQS] __cacheline_aligned_in_smp;
DEFINE_PER_CPU(struct task_struct *, ksoftirqd);
+DEFINE_PER_CPU(bool, ksoftirqd_scheduled);
+
+static inline bool ksoftirqd_running(void)
+{
+ extern int netdev_max_backlog; /* temp hack */
+
+ return (netdev_max_backlog & 1) && __this_cpu_read(ksoftirqd_scheduled);
+}
const char * const softirq_to_name[NR_SOFTIRQS] = {
"HI", "TIMER", "NET_TX", "NET_RX", "BLOCK", "BLOCK_IOPOLL",
@@ -73,8 +81,10 @@ static void wakeup_softirqd(void)
/* Interrupts are disabled: no need to stop preemption */
struct task_struct *tsk = __this_cpu_read(ksoftirqd);
- if (tsk && tsk->state != TASK_RUNNING)
+ if (tsk && tsk->state != TASK_RUNNING) {
+ __this_cpu_write(ksoftirqd_scheduled, true);
wake_up_process(tsk);
+ }
}
/*
@@ -313,7 +323,7 @@ asmlinkage __visible void do_softirq(void)
pending = local_softirq_pending();
- if (pending)
+ if (pending && !ksoftirqd_running())
do_softirq_own_stack();
local_irq_restore(flags);
@@ -340,6 +350,9 @@ void irq_enter(void)
static inline void invoke_softirq(void)
{
+ if (ksoftirqd_running())
+ return;
+
if (!force_irqthreads) {
#ifdef CONFIG_HAVE_IRQ_EXIT_ON_IRQ_STACK
/*
@@ -660,6 +673,8 @@ static void run_ksoftirqd(unsigned int cpu)
* in the task stack here.
*/
__do_softirq();
+ if (!local_softirq_pending())
+ __this_cpu_write(ksoftirqd_scheduled, false);
local_irq_enable();
cond_resched_rcu_qs();
return;
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web