Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1308505 > unrolled thread

[PATCH RT] net: move xmit_recursion to per-task variable on -RT

Started bySebastian Andrzej Siewior <bigeasy@linutronix.de>
First post2016-01-13 16:30 +0100
Last post2016-01-15 10:40 +0100
Articles 8 — 4 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH RT] net: move xmit_recursion to per-task variable on -RT Sebastian Andrzej Siewior <bigeasy@linutronix.de> - 2016-01-13 16:30 +0100
    Re: [PATCH RT] net: move xmit_recursion to per-task variable on  -RT Thomas Gleixner <tglx@linutronix.de> - 2016-01-13 18:40 +0100
      Re: [PATCH RT] net: move xmit_recursion to per-task variable on -RT Sebastian Andrzej Siewior <bigeasy@linutronix.de> - 2016-01-14 16:00 +0100
        Re: [PATCH RT] net: move xmit_recursion to per-task variable on -RT Hannes Frederic Sowa <hannes@stressinduktion.org> - 2016-01-14 23:10 +0100
          Re: [PATCH RT] net: move xmit_recursion to per-task variable on -RT Eric Dumazet <eric.dumazet@gmail.com> - 2016-01-14 23:30 +0100
            Re: [PATCH RT] net: move xmit_recursion to per-task variable on -RT Hannes Frederic Sowa <hannes@stressinduktion.org> - 2016-01-15 00:10 +0100
              Re: [PATCH RT] net: move xmit_recursion to per-task variable on  -RT Thomas Gleixner <tglx@linutronix.de> - 2016-01-15 09:30 +0100
                Re: [PATCH RT] net: move xmit_recursion to per-task variable on -RT Hannes Frederic Sowa <hannes@stressinduktion.org> - 2016-01-15 10:40 +0100

#1308505 — [PATCH RT] net: move xmit_recursion to per-task variable on -RT

FromSebastian Andrzej Siewior <bigeasy@linutronix.de>
Date2016-01-13 16:30 +0100
Subject[PATCH RT] net: move xmit_recursion to per-task variable on -RT
Message-ID<qQvx1-1Cf-19@gated-at.bofh.it>
A softirq on -RT can be preempted. That means one task is in
__dev_queue_xmit(), gets preempted and another task may enter
__dev_queue_xmit() aw well. netperf together with a bridge device
will then trigger the `recursion alert` because each task increments
the xmit_recursion variable which is per-CPU.
A virtual device like br0 is required to trigger this warning.

This patch moves the counter to per task instead per-CPU so it counts
the recursion properly on -RT.

Cc: stable-rt@vger.kernel.org
Signed-off-by: Sebastian Andrzej Siewior <bigeasy@linutronix.de>
---
 include/linux/netdevice.h |  9 +++++++++
 include/linux/sched.h     |  3 +++
 net/core/dev.c            | 41 ++++++++++++++++++++++++++++++++++++++---
 3 files changed, 50 insertions(+), 3 deletions(-)

diff --git a/include/linux/netdevice.h b/include/linux/netdevice.h
index f14e39cb897c..4a8d3429dc12 100644
--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -2249,11 +2249,20 @@ void netdev_freemem(struct net_device *dev);
 void synchronize_net(void);
 int init_dummy_netdev(struct net_device *dev);
 
+#ifdef CONFIG_PREEMPT_RT_FULL
+static inline int dev_recursion_level(void)
+{
+	return atomic_read(&current->xmit_recursion);
+}
+
+#else
+
 DECLARE_PER_CPU(int, xmit_recursion);
 static inline int dev_recursion_level(void)
 {
 	return this_cpu_read(xmit_recursion);
 }
+#endif
 
 struct net_device *dev_get_by_index(struct net *net, int ifindex);
 struct net_device *__dev_get_by_index(struct net *net, int ifindex);
diff --git a/include/linux/sched.h b/include/linux/sched.h
index 04eb2f8bc274..5d36818107b0 100644
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -1855,6 +1855,9 @@ struct task_struct {
 	pte_t kmap_pte[KM_TYPE_NR];
 # endif
 #endif
+#ifdef CONFIG_PREEMPT_RT_FULL
+	atomic_t xmit_recursion;
+#endif
 #ifdef CONFIG_DEBUG_ATOMIC_SLEEP
 	unsigned long	task_state_change;
 #endif
diff --git a/net/core/dev.c b/net/core/dev.c
index ae4a67e7e654..1f6a7e9a22c4 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -2946,9 +2946,44 @@ static void skb_update_prio(struct sk_buff *skb)
 #define skb_update_prio(skb)
 #endif
 
+#ifdef CONFIG_PREEMPT_RT_FULL
+
+static inline int xmit_rec_read(void)
+{
+       return atomic_read(&current->xmit_recursion);
+}
+
+static inline void xmit_rec_inc(void)
+{
+       atomic_inc(&current->xmit_recursion);
+}
+
+static inline void xmit_rec_dec(void)
+{
+       atomic_dec(&current->xmit_recursion);
+}
+
+#else
+
 DEFINE_PER_CPU(int, xmit_recursion);
 EXPORT_SYMBOL(xmit_recursion);
 
+static inline int xmit_rec_read(void)
+{
+	return __this_cpu_read(xmit_recursion);
+}
+
+static inline void xmit_rec_inc(void)
+{
+	__this_cpu_inc(xmit_recursion);
+}
+
+static inline int xmit_rec_dec(void)
+{
+	__this_cpu_dec(xmit_recursion);
+}
+#endif
+
 #define RECURSION_LIMIT 10
 
 /**
@@ -3141,7 +3176,7 @@ static int __dev_queue_xmit(struct sk_buff *skb, void *accel_priv)
 
 		if (txq->xmit_lock_owner != cpu) {
 
-			if (__this_cpu_read(xmit_recursion) > RECURSION_LIMIT)
+			if (xmit_rec_read() > RECURSION_LIMIT)
 				goto recursion_alert;
 
 			skb = validate_xmit_skb(skb, dev);
@@ -3151,9 +3186,9 @@ static int __dev_queue_xmit(struct sk_buff *skb, void *accel_priv)
 			HARD_TX_LOCK(dev, txq, cpu);
 
 			if (!netif_xmit_stopped(txq)) {
-				__this_cpu_inc(xmit_recursion);
+				xmit_rec_inc();
 				skb = dev_hard_start_xmit(skb, dev, txq, &rc);
-				__this_cpu_dec(xmit_recursion);
+				xmit_rec_dec();
 				if (dev_xmit_complete(rc)) {
 					HARD_TX_UNLOCK(dev, txq);
 					goto out;
-- 
2.7.0.rc3

[toc] | [next] | [standalone]


#1308675 — Re: [PATCH RT] net: move xmit_recursion to per-task variable on -RT

FromThomas Gleixner <tglx@linutronix.de>
Date2016-01-13 18:40 +0100
SubjectRe: [PATCH RT] net: move xmit_recursion to per-task variable on -RT
Message-ID<qQxyP-300-55@gated-at.bofh.it>
In reply to#1308505
On Wed, 13 Jan 2016, Sebastian Andrzej Siewior wrote:
> +#ifdef CONFIG_PREEMPT_RT_FULL
> +static inline int dev_recursion_level(void)
> +{
> +	return atomic_read(&current->xmit_recursion);

Why would you need an atomic here. current does hardly race against itself.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1309355

FromSebastian Andrzej Siewior <bigeasy@linutronix.de>
Date2016-01-14 16:00 +0100
Message-ID<qQRxx-8vf-29@gated-at.bofh.it>
In reply to#1308675
* Thomas Gleixner | 2016-01-13 18:31:46 [+0100]:

>On Wed, 13 Jan 2016, Sebastian Andrzej Siewior wrote:
>> +#ifdef CONFIG_PREEMPT_RT_FULL
>> +static inline int dev_recursion_level(void)
>> +{
>> +	return atomic_read(&current->xmit_recursion);
>
>Why would you need an atomic here. current does hardly race against itself.

right.

--- a/include/linux/netdevice.h
+++ b/include/linux/netdevice.h
@@ -2249,11 +2249,20 @@ void netdev_freemem(struct net_device *d
 void synchronize_net(void);
 int init_dummy_netdev(struct net_device *dev);
 
+#ifdef CONFIG_PREEMPT_RT_FULL
+static inline int dev_recursion_level(void)
+{
+	return current->xmit_recursion;
+}
+
+#else
+
 DECLARE_PER_CPU(int, xmit_recursion);
 static inline int dev_recursion_level(void)
 {
 	return this_cpu_read(xmit_recursion);
 }
+#endif
 
 struct net_device *dev_get_by_index(struct net *net, int ifindex);
 struct net_device *__dev_get_by_index(struct net *net, int ifindex);
--- a/include/linux/sched.h
+++ b/include/linux/sched.h
@@ -1851,6 +1851,9 @@ struct task_struct {
 #ifdef CONFIG_DEBUG_ATOMIC_SLEEP
 	unsigned long	task_state_change;
 #endif
+#ifdef CONFIG_PREEMPT_RT_FULL
+	int xmit_recursion;
+#endif
 	int pagefault_disabled;
 /* CPU-specific state of this task */
 	struct thread_struct thread;
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -2940,9 +2940,44 @@ static void skb_update_prio(struct sk_bu
 #define skb_update_prio(skb)
 #endif
 
+#ifdef CONFIG_PREEMPT_RT_FULL
+
+static inline int xmit_rec_read(void)
+{
+       return current->xmit_recursion;
+}
+
+static inline void xmit_rec_inc(void)
+{
+       current->xmit_recursion++;
+}
+
+static inline void xmit_rec_dec(void)
+{
+       current->xmit_recursion--;
+}
+
+#else
+
 DEFINE_PER_CPU(int, xmit_recursion);
 EXPORT_SYMBOL(xmit_recursion);
 
+static inline int xmit_rec_read(void)
+{
+	return __this_cpu_read(xmit_recursion);
+}
+
+static inline void xmit_rec_inc(void)
+{
+	__this_cpu_inc(xmit_recursion);
+}
+
+static inline int xmit_rec_dec(void)
+{
+	__this_cpu_dec(xmit_recursion);
+}
+#endif
+
 #define RECURSION_LIMIT 10
 
 /**
@@ -3135,7 +3170,7 @@ static int __dev_queue_xmit(struct sk_bu
 
 		if (txq->xmit_lock_owner != cpu) {
 
-			if (__this_cpu_read(xmit_recursion) > RECURSION_LIMIT)
+			if (xmit_rec_read() > RECURSION_LIMIT)
 				goto recursion_alert;
 
 			skb = validate_xmit_skb(skb, dev);
@@ -3145,9 +3180,9 @@ static int __dev_queue_xmit(struct sk_bu
 			HARD_TX_LOCK(dev, txq, cpu);
 
 			if (!netif_xmit_stopped(txq)) {
-				__this_cpu_inc(xmit_recursion);
+				xmit_rec_inc();
 				skb = dev_hard_start_xmit(skb, dev, txq, &rc);
-				__this_cpu_dec(xmit_recursion);
+				xmit_rec_dec();
 				if (dev_xmit_complete(rc)) {
 					HARD_TX_UNLOCK(dev, txq);
 					goto out;

>Thanks,
>
>	tglx

Sebastian

[toc] | [prev] | [next] | [standalone]


#1309702

FromHannes Frederic Sowa <hannes@stressinduktion.org>
Date2016-01-14 23:10 +0100
Message-ID<qQYfF-533-21@gated-at.bofh.it>
In reply to#1309355
On 14.01.2016 15:50, Sebastian Andrzej Siewior wrote:
> * Thomas Gleixner | 2016-01-13 18:31:46 [+0100]:
>
>> On Wed, 13 Jan 2016, Sebastian Andrzej Siewior wrote:
>>> +#ifdef CONFIG_PREEMPT_RT_FULL
>>> +static inline int dev_recursion_level(void)
>>> +{
>>> +	return atomic_read(&current->xmit_recursion);
>>
>> Why would you need an atomic here. current does hardly race against itself.
>
> right.

We are just adding a second recursion limit solely to openvswitch which 
has the same problem:

https://patchwork.ozlabs.org/patch/566769/

This time also we depend on rcu_read_lock marking the section being 
nonpreemptible. Nice would be a more generic solution here which doesn't 
need to always add something to *current.

Thanks,
Hannes

[toc] | [prev] | [next] | [standalone]


#1309722

FromEric Dumazet <eric.dumazet@gmail.com>
Date2016-01-14 23:30 +0100
Message-ID<qQYz2-5b5-59@gated-at.bofh.it>
In reply to#1309702
On Thu, 2016-01-14 at 23:02 +0100, Hannes Frederic Sowa wrote:

> We are just adding a second recursion limit solely to openvswitch which 
> has the same problem:
> 
> https://patchwork.ozlabs.org/patch/566769/
> 
> This time also we depend on rcu_read_lock marking the section being 
> nonpreemptible. Nice would be a more generic solution here which doesn't 
> need to always add something to *current.


Note that rcu_read_lock() does not imply that preemption is disabled.

[toc] | [prev] | [next] | [standalone]


#1309748

FromHannes Frederic Sowa <hannes@stressinduktion.org>
Date2016-01-15 00:10 +0100
Message-ID<qQZbI-5Fp-13@gated-at.bofh.it>
In reply to#1309722
On 14.01.2016 23:20, Eric Dumazet wrote:
> On Thu, 2016-01-14 at 23:02 +0100, Hannes Frederic Sowa wrote:
>
>> We are just adding a second recursion limit solely to openvswitch which
>> has the same problem:
>>
>> https://patchwork.ozlabs.org/patch/566769/
>>
>> This time also we depend on rcu_read_lock marking the section being
>> nonpreemptible. Nice would be a more generic solution here which doesn't
>> need to always add something to *current.
>
>
> Note that rcu_read_lock() does not imply that preemption is disabled.

Exactly, it is conditional on CONFIG_PREEMPT_CPU/CONFIG_PREMPT_COUNT but 
haven't thought about exactly that in this moment.

I will resend this patch with better protection.

Thanks Eric!

[toc] | [prev] | [next] | [standalone]


#1309948 — Re: [PATCH RT] net: move xmit_recursion to per-task variable on -RT

FromThomas Gleixner <tglx@linutronix.de>
Date2016-01-15 09:30 +0100
SubjectRe: [PATCH RT] net: move xmit_recursion to per-task variable on -RT
Message-ID<qR7VE-3qX-21@gated-at.bofh.it>
In reply to#1309748
On Fri, 15 Jan 2016, Hannes Frederic Sowa wrote:
> On 14.01.2016 23:20, Eric Dumazet wrote:
> > On Thu, 2016-01-14 at 23:02 +0100, Hannes Frederic Sowa wrote:
> > 
> > > We are just adding a second recursion limit solely to openvswitch which
> > > has the same problem:
> > > 
> > > https://patchwork.ozlabs.org/patch/566769/
> > > 
> > > This time also we depend on rcu_read_lock marking the section being
> > > nonpreemptible. Nice would be a more generic solution here which doesn't
> > > need to always add something to *current.
> > 
> > 
> > Note that rcu_read_lock() does not imply that preemption is disabled.
> 
> Exactly, it is conditional on CONFIG_PREEMPT_CPU/CONFIG_PREMPT_COUNT but
> haven't thought about exactly that in this moment.

Wrong. CONFIG_PREEMPT_RCU makes RCU preemptible.

If that is not set then it fiddles with preempt_count when
CONFIG_PREEMPT_COUNT=y. If CONFIG_PREEMPT_COUNT=n then you have a non
preemptible system anyway.

So you cannot assume that rcu_read_lock() disables preemption.

Thanks,

	tglx

[toc] | [prev] | [next] | [standalone]


#1309992

FromHannes Frederic Sowa <hannes@stressinduktion.org>
Date2016-01-15 10:40 +0100
Message-ID<qR91o-49P-17@gated-at.bofh.it>
In reply to#1309948
On 15.01.2016 09:21, Thomas Gleixner wrote:
> On Fri, 15 Jan 2016, Hannes Frederic Sowa wrote:
>> On 14.01.2016 23:20, Eric Dumazet wrote:
>>> On Thu, 2016-01-14 at 23:02 +0100, Hannes Frederic Sowa wrote:
>>>
>>>> We are just adding a second recursion limit solely to openvswitch which
>>>> has the same problem:
>>>>
>>>> https://patchwork.ozlabs.org/patch/566769/
>>>>
>>>> This time also we depend on rcu_read_lock marking the section being
>>>> nonpreemptible. Nice would be a more generic solution here which doesn't
>>>> need to always add something to *current.
>>>
>>>
>>> Note that rcu_read_lock() does not imply that preemption is disabled.
>>
>> Exactly, it is conditional on CONFIG_PREEMPT_CPU/CONFIG_PREMPT_COUNT but
>> haven't thought about exactly that in this moment.
>
> Wrong. CONFIG_PREEMPT_RCU makes RCU preemptible.
>
> If that is not set then it fiddles with preempt_count when
> CONFIG_PREEMPT_COUNT=y. If CONFIG_PREEMPT_COUNT=n then you have a non
> preemptible system anyway.
>
> So you cannot assume that rcu_read_lock() disables preemption.

Sorry for maybe writing it misleading but that is exactly what I wanted 
to say here. Yes, I agree, I didn't really check because of _bh and 
rcu_read_lock. This was a mistake. ;)

I already send out an updated patch with added preemption guards.

Thanks,
Hannes

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web