Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1440446 > unrolled thread
| Started by | Jan Kara <jack@suse.cz> |
|---|---|
| First post | 2016-07-11 12:30 +0200 |
| Last post | 2016-07-11 21:10 +0200 |
| Articles | 7 on this page of 47 — 8 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [Query] Preemption (hogging) of the work handler Jan Kara <jack@suse.cz> - 2016-07-11 12:30 +0200
Re: [Query] Preemption (hogging) of the work handler Sergey Senozhatsky <sergey.senozhatsky@gmail.com> - 2016-07-11 17:50 +0200
Re: [Query] Preemption (hogging) of the work handler "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-07-12 00:40 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-12 00:50 +0200
Re: [Query] Preemption (hogging) of the work handler "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-07-12 14:30 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-12 15:10 +0200
Re: [Query] Preemption (hogging) of the work handler Petr Mladek <pmladek@suse.com> - 2016-07-12 16:00 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-12 16:10 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-12 00:40 +0200
Re: [Query] Preemption (hogging) of the work handler Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-07-12 11:40 +0200
Re: [Query] Preemption (hogging) of the work handler Petr Mladek <pmladek@suse.com> - 2016-07-12 15:00 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-12 15:20 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-12 19:20 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-12 22:10 +0200
Re: [Query] Preemption (hogging) of the work handler "Rafael J. Wysocki" <rafael@kernel.org> - 2016-07-12 22:10 +0200
Re: [Query] Preemption (hogging) of the work handler Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-07-13 09:10 +0200
Re: [Query] Preemption (hogging) of the work handler "Rafael J. Wysocki" <rafael@kernel.org> - 2016-07-13 14:10 +0200
Re: [Query] Preemption (hogging) of the work handler Sergey Senozhatsky <sergey.senozhatsky@gmail.com> - 2016-07-13 15:00 +0200
Re: [Query] Preemption (hogging) of the work handler "Rafael J. Wysocki" <rafael@kernel.org> - 2016-07-13 15:30 +0200
Re: [Query] Preemption (hogging) of the work handler Sergey Senozhatsky <sergey.senozhatsky@gmail.com> - 2016-07-12 16:10 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-15 02:00 +0200
Re: [Query] Preemption (hogging) of the work handler Sergey Senozhatsky <sergey.senozhatsky@gmail.com> - 2016-07-15 15:20 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-15 18:00 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-13 01:30 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-13 02:20 +0200
Re: [Query] Preemption (hogging) of the work handler Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-07-13 07:50 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-13 17:50 +0200
Re: [Query] Preemption (hogging) of the work handler "Rafael J. Wysocki" <rafael@kernel.org> - 2016-07-14 01:10 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-14 01:20 +0200
Re: [Query] Preemption (hogging) of the work handler Greg Kroah-Hartman <gregkh@linuxfoundation.org> - 2016-07-14 01:40 +0200
Re: [Query] Preemption (hogging) of the work handler Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-07-14 03:00 +0200
Re: [Query] Preemption (hogging) of the work handler "Rafael J. Wysocki" <rafael@kernel.org> - 2016-07-14 03:10 +0200
Re: [Query] Preemption (hogging) of the work handler Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2016-07-14 03:40 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-15 00:00 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-15 00:00 +0200
Re: [Query] Preemption (hogging) of the work handler Jan Kara <jack@suse.cz> - 2016-07-14 16:20 +0200
Re: [Query] Preemption (hogging) of the work handler "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-07-14 16:30 +0200
Re: [Query] Preemption (hogging) of the work handler Jan Kara <jack@suse.cz> - 2016-07-14 16:40 +0200
Re: [Query] Preemption (hogging) of the work handler "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-07-14 16:50 +0200
Re: [Query] Preemption (hogging) of the work handler Jan Kara <jack@suse.cz> - 2016-07-14 17:00 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-15 00:20 +0200
Re: [Query] Preemption (hogging) of the work handler Sergey Senozhatsky <sergey.senozhatsky@gmail.com> - 2016-07-14 16:40 +0200
Re: [Query] Preemption (hogging) of the work handler Jan Kara <jack@suse.cz> - 2016-07-14 17:10 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-15 00:20 +0200
Re: [Query] Preemption (hogging) of the work handler Jan Kara <jack@suse.cz> - 2016-07-18 13:10 +0200
Re: [Query] Preemption (hogging) of the work handler "Rafael J. Wysocki" <rjw@rjwysocki.net> - 2016-07-18 13:50 +0200
Re: [Query] Preemption (hogging) of the work handler Viresh Kumar <viresh.kumar@linaro.org> - 2016-07-11 21:10 +0200
Page 3 of 3 — ← Prev page 1 2 [3]
| From | Viresh Kumar <viresh.kumar@linaro.org> |
|---|---|
| Date | 2016-07-15 00:20 +0200 |
| Message-ID | <rUX9c-72G-153@gated-at.bofh.it> |
| In reply to | #1443503 |
On 14-07-16, 16:55, Jan Kara wrote: > Agree on that - but that seems to be a problem of a particular wakeup > implementation of the 3.10 kernel Viresh is using, not a problem of the > upstream kernel. I think we can get it to trigger on mainline as well. Also to mention that I also don't see it on every suspend. Probably hrtimer_active() call returns true in the sequence somewhere earlier, and we don't have to go activate a hrtimer. -- viresh
[toc] | [prev] | [next] | [standalone]
| From | Sergey Senozhatsky <sergey.senozhatsky@gmail.com> |
|---|---|
| Date | 2016-07-14 16:40 +0200 |
| Message-ID | <rUPXY-2rx-11@gated-at.bofh.it> |
| In reply to | #1443476 |
Hello Jan,
On (07/14/16 16:12), Jan Kara wrote:
[..]
> > *** a printk() call from here will kill the system. either it will
> > recurse printk(), or spin forever in 'nested' printk() on one of
> > the already taken spin locks.
[..]
> And with sync printk the above deadlock doesn't trigger only by chance - if
> there happened to be a waiter on console_sem while we suspend, the same
> deadlock would trigger because up(&console_sem) will try to wake him up and
> the warning in timekeeping code will cause recursive printk.
>
> So I think your patch doesn't really address the real issue - it only
> works around the particular WARN_ON(timekeeping_enabled) warning but if
> there was a different warning in timekeeping code which would trigger, it
> has a potential for causing recursive printk deadlock (and indeed we had
> such issues previously - see e.g. 504d58745c9c "timer: Fix lock inversion
> between hrtimer_bases.lock and scheduler locks").
we switch to sync printk in suspend_console(), that is happening
long before we start bringing cpu downs
suspend_devices_and_enter()
suspend_console()
...
suspend_enter()
...
dpm_suspend_late
...
disable_nonboot_cpus
and cpu_down() in printk does
static int console_cpu_notify(struct notifier_block *self,
unsigned long action, void *hcpu)
{
switch (action) {
case CPU_ONLINE:
case CPU_DEAD:
case CPU_DOWN_FAILED:
case CPU_UP_CANCELED:
console_lock();
console_unlock();
}
return NOTIFY_OK;
}
so I think this console_lock() sort of guarantees that there should be
no sleeping tasks in console semaphore wait list. or am I missing something?
> So there are IMHO two issues here worth looking at:
>
> 1) I didn't find how a wakeup would would lead to calling to ktime_get() in
> the current upstream kernel or even current RT kernel. Maybe this is a
> problem specific to the 3.10 kernel you are using? If yes, we don't have to
> do anything for current upstream AFAIU.
I personally suspect it's an in-hose (custom) code.
-ss
> If I just missed how wakeup can call into ktime_get() in current upstream,
> there is another question:
>
> 2) Is it OK that printk calls wakeup so late during suspend? I believe it
> is but I'm neither scheduler nor suspend expert. If it is OK, and wakeup
> can lead to ktime_get() in current upstream, then this contradicts the
> check WARN_ON(timekeeping_suspended) in ktime_get() and something is wrong.
>
> Adding Thomas to CC as timer / RT expert...
>
> Honza
>
> > so... I think we can switch to sync printk mode in suspend_console() and
> > enable async printk from resume_console(). IOW, suspend/kexec are now
> > executed under sync printk mode.
> >
> > we already call console_unlock() during suspend, which is synchronous,
> > many times (e.g. console_cpu_notify()).
> >
> >
> > something like below, perhaps. will this work for you?
> >
> > ---
> > kernel/printk/printk.c | 12 +++++++++++-
> > 1 file changed, 11 insertions(+), 1 deletion(-)
> >
> > diff --git a/kernel/printk/printk.c b/kernel/printk/printk.c
> > index bbb4180..786690e 100644
> > --- a/kernel/printk/printk.c
> > +++ b/kernel/printk/printk.c
> > @@ -288,6 +288,11 @@ static u32 log_buf_len = __LOG_BUF_LEN;
> >
> > /* Control whether printing to console must be synchronous. */
> > static bool __read_mostly printk_sync = true;
> > +/*
> > + * Force sync printk mode during suspend/kexec, regardless whether
> > + * console_suspend_enabled permits console suspend.
> > + */
> > +static bool __read_mostly force_printk_sync;
> > /* Printing kthread for async printk */
> > static struct task_struct *printk_kthread;
> > /* When `true' printing thread has messages to print */
> > @@ -295,7 +300,7 @@ static bool printk_kthread_need_flush_console;
> >
> > static inline bool can_printk_async(void)
> > {
> > - return !printk_sync && printk_kthread;
> > + return !printk_sync && printk_kthread && !force_printk_sync;
> > }
> >
> > /* Return log buffer address */
> > @@ -2027,6 +2032,7 @@ static bool suppress_message_printing(int level) { return false; }
> >
> > /* Still needs to be defined for users */
> > DEFINE_PER_CPU(printk_func_t, printk_func);
> > +static bool __read_mostly force_printk_sync;
> >
> > #endif /* CONFIG_PRINTK */
> >
> > @@ -2163,6 +2169,8 @@ MODULE_PARM_DESC(console_suspend, "suspend console during suspend"
> > */
> > void suspend_console(void)
> > {
> > + force_printk_sync = true;
> > +
> > if (!console_suspend_enabled)
> > return;
> > printk("Suspending console(s) (use no_console_suspend to debug)\n");
> > @@ -2173,6 +2181,8 @@ void suspend_console(void)
> >
> > void resume_console(void)
> > {
> > + force_printk_sync = false;
> > +
> > if (!console_suspend_enabled)
> > return;
> > down_console_sem();
> > --
> > 2.9.0.rc1
> >
> --
> Jan Kara <jack@suse.com>
> SUSE Labs, CR
>
[toc] | [prev] | [next] | [standalone]
| From | Jan Kara <jack@suse.cz> |
|---|---|
| Date | 2016-07-14 17:10 +0200 |
| Message-ID | <rUQr0-2Rf-9@gated-at.bofh.it> |
| In reply to | #1443486 |
On Thu 14-07-16 23:34:50, Sergey Senozhatsky wrote:
> Hello Jan,
>
> On (07/14/16 16:12), Jan Kara wrote:
> [..]
> > > *** a printk() call from here will kill the system. either it will
> > > recurse printk(), or spin forever in 'nested' printk() on one of
> > > the already taken spin locks.
> [..]
> > And with sync printk the above deadlock doesn't trigger only by chance - if
> > there happened to be a waiter on console_sem while we suspend, the same
> > deadlock would trigger because up(&console_sem) will try to wake him up and
> > the warning in timekeeping code will cause recursive printk.
> >
> > So I think your patch doesn't really address the real issue - it only
> > works around the particular WARN_ON(timekeeping_enabled) warning but if
> > there was a different warning in timekeeping code which would trigger, it
> > has a potential for causing recursive printk deadlock (and indeed we had
> > such issues previously - see e.g. 504d58745c9c "timer: Fix lock inversion
> > between hrtimer_bases.lock and scheduler locks").
>
> we switch to sync printk in suspend_console(), that is happening
> long before we start bringing cpu downs
>
> suspend_devices_and_enter()
> suspend_console()
> ...
> suspend_enter()
> ...
> dpm_suspend_late
> ...
> disable_nonboot_cpus
>
>
>
> and cpu_down() in printk does
>
> static int console_cpu_notify(struct notifier_block *self,
> unsigned long action, void *hcpu)
> {
> switch (action) {
> case CPU_ONLINE:
> case CPU_DEAD:
> case CPU_DOWN_FAILED:
> case CPU_UP_CANCELED:
> console_lock();
> console_unlock();
> }
> return NOTIFY_OK;
> }
>
> so I think this console_lock() sort of guarantees that there should be
> no sleeping tasks in console semaphore wait list. or am I missing something?
No, probably you're right - unless there would be a CPU notifier executed
after console_cpu_notify() which would try to acquire console_sem for some
reason. But that is a wild speculation and I tend to agree that in
synchronous printk case and current code the wakeup cannot happen.
But my point really is that I don't see why changing process state (which
is what wakeup actually is) should be problematic even this late during
suspend...
Honza
--
Jan Kara <jack@suse.com>
SUSE Labs, CR
[toc] | [prev] | [next] | [standalone]
| From | Viresh Kumar <viresh.kumar@linaro.org> |
|---|---|
| Date | 2016-07-15 00:20 +0200 |
| Message-ID | <rUX9a-72G-87@gated-at.bofh.it> |
| In reply to | #1443476 |
On 14-07-16, 16:12, Jan Kara wrote:
> Exactly. Calling printk() from certain parts of the kernel (like scheduler
> code or timer code) has been always unsafe because printk itself uses these
> parts and so it can lead to deadlocks. That's why printk_deffered() has
> been introduced as you mention below.
>
> And with sync printk the above deadlock doesn't trigger only by chance - if
> there happened to be a waiter on console_sem while we suspend, the same
> deadlock would trigger because up(&console_sem) will try to wake him up and
> the warning in timekeeping code will cause recursive printk.
>
> So I think your patch doesn't really address the real issue - it only
> works around the particular WARN_ON(timekeeping_enabled) warning but if
> there was a different warning in timekeeping code which would trigger, it
> has a potential for causing recursive printk deadlock (and indeed we had
> such issues previously - see e.g. 504d58745c9c "timer: Fix lock inversion
> between hrtimer_bases.lock and scheduler locks").
>
> So there are IMHO two issues here worth looking at:
>
> 1) I didn't find how a wakeup would would lead to calling to ktime_get() in
> the current upstream kernel or even current RT kernel. Maybe this is a
> problem specific to the 3.10 kernel you are using? If yes, we don't have to
> do anything for current upstream AFAIU.
I haven't checked that earlier, but I see the path in both 3.10 and mainline.
vprintk_emit
-> wake_up_process
-> try_to_wake_up
-> ttwu_queue
-> ttwu_do_activate
-> ttwu_activate
-> activate_task
-> enqueue_task (sched/core.c)
-> enqueue_task_rt (rt.c)
-> enqueue_rt_entity
-> __enqueue_rt_entity
-> inc_rt_tasks
-> inc_rt_group
-> start_rt_bandwidth
-> start_bandwidth_timer
-> __hrtimer_start_range_ns
-> ktime_get()
> If I just missed how wakeup can call into ktime_get() in current upstream,
> there is another question:
>
> 2) Is it OK that printk calls wakeup so late during suspend?
To clarify again to everybody, we are talking about the place where all non-boot
CPUs are already hot-unplugged and the last running one has disabled interrupts.
I believe that we can't do migration at all now, right? What will we get by
calling wake_up_process() now anyway ?
> I believe it
> is but I'm neither scheduler nor suspend expert. If it is OK, and wakeup
> can lead to ktime_get() in current upstream, then this contradicts the
> check WARN_ON(timekeeping_suspended) in ktime_get() and something is wrong.
>
> Adding Thomas to CC as timer / RT expert...
Thanks.
--
viresh
[toc] | [prev] | [next] | [standalone]
| From | Jan Kara <jack@suse.cz> |
|---|---|
| Date | 2016-07-18 13:10 +0200 |
| Message-ID | <rWeAW-5p6-13@gated-at.bofh.it> |
| In reply to | #1443810 |
On Thu 14-07-16 15:12:51, Viresh Kumar wrote: > On 14-07-16, 16:12, Jan Kara wrote: > > Exactly. Calling printk() from certain parts of the kernel (like scheduler > > code or timer code) has been always unsafe because printk itself uses these > > parts and so it can lead to deadlocks. That's why printk_deffered() has > > been introduced as you mention below. > > > > And with sync printk the above deadlock doesn't trigger only by chance - if > > there happened to be a waiter on console_sem while we suspend, the same > > deadlock would trigger because up(&console_sem) will try to wake him up and > > the warning in timekeeping code will cause recursive printk. > > > > So I think your patch doesn't really address the real issue - it only > > works around the particular WARN_ON(timekeeping_enabled) warning but if > > there was a different warning in timekeeping code which would trigger, it > > has a potential for causing recursive printk deadlock (and indeed we had > > such issues previously - see e.g. 504d58745c9c "timer: Fix lock inversion > > between hrtimer_bases.lock and scheduler locks"). > > > > So there are IMHO two issues here worth looking at: > > > > 1) I didn't find how a wakeup would would lead to calling to ktime_get() in > > the current upstream kernel or even current RT kernel. Maybe this is a > > problem specific to the 3.10 kernel you are using? If yes, we don't have to > > do anything for current upstream AFAIU. > > I haven't checked that earlier, but I see the path in both 3.10 and mainline. > > vprintk_emit > -> wake_up_process > -> try_to_wake_up > -> ttwu_queue > -> ttwu_do_activate > -> ttwu_activate > -> activate_task > -> enqueue_task (sched/core.c) > -> enqueue_task_rt (rt.c) > -> enqueue_rt_entity > -> __enqueue_rt_entity > -> inc_rt_tasks > -> inc_rt_group > -> start_rt_bandwidth > -> start_bandwidth_timer > -> __hrtimer_start_range_ns > -> ktime_get() Yeah, you are right. > > If I just missed how wakeup can call into ktime_get() in current upstream, > > there is another question: > > > > 2) Is it OK that printk calls wakeup so late during suspend? > > To clarify again to everybody, we are talking about the place where all > non-boot CPUs are already hot-unplugged and the last running one has > disabled interrupts. > > I believe that we can't do migration at all now, right? What will we get by > calling wake_up_process() now anyway ? As I already wrote to Rafael, wake_up_process() will change the process state to TASK_RUNNING so that it can run after we resume from suspend. But seeing that the same problem is in upstream I guess what Sergey did makes more sense if it works for you. If Sergey's fix does not work for you due to too many messages being printed during device suspend, then we will have to try something else... Honza -- Jan Kara <jack@suse.com> SUSE Labs, CR
[toc] | [prev] | [next] | [standalone]
| From | "Rafael J. Wysocki" <rjw@rjwysocki.net> |
|---|---|
| Date | 2016-07-18 13:50 +0200 |
| Message-ID | <rWfdE-5Cg-5@gated-at.bofh.it> |
| In reply to | #1445381 |
On Monday, July 18, 2016 01:01:34 PM Jan Kara wrote: > On Thu 14-07-16 15:12:51, Viresh Kumar wrote: > > On 14-07-16, 16:12, Jan Kara wrote: > > > Exactly. Calling printk() from certain parts of the kernel (like scheduler > > > code or timer code) has been always unsafe because printk itself uses these > > > parts and so it can lead to deadlocks. That's why printk_deffered() has > > > been introduced as you mention below. > > > > > > And with sync printk the above deadlock doesn't trigger only by chance - if > > > there happened to be a waiter on console_sem while we suspend, the same > > > deadlock would trigger because up(&console_sem) will try to wake him up and > > > the warning in timekeeping code will cause recursive printk. > > > > > > So I think your patch doesn't really address the real issue - it only > > > works around the particular WARN_ON(timekeeping_enabled) warning but if > > > there was a different warning in timekeeping code which would trigger, it > > > has a potential for causing recursive printk deadlock (and indeed we had > > > such issues previously - see e.g. 504d58745c9c "timer: Fix lock inversion > > > between hrtimer_bases.lock and scheduler locks"). > > > > > > So there are IMHO two issues here worth looking at: > > > > > > 1) I didn't find how a wakeup would would lead to calling to ktime_get() in > > > the current upstream kernel or even current RT kernel. Maybe this is a > > > problem specific to the 3.10 kernel you are using? If yes, we don't have to > > > do anything for current upstream AFAIU. > > > > I haven't checked that earlier, but I see the path in both 3.10 and mainline. > > > > vprintk_emit > > -> wake_up_process > > -> try_to_wake_up > > -> ttwu_queue > > -> ttwu_do_activate > > -> ttwu_activate > > -> activate_task > > -> enqueue_task (sched/core.c) > > -> enqueue_task_rt (rt.c) > > -> enqueue_rt_entity > > -> __enqueue_rt_entity > > -> inc_rt_tasks > > -> inc_rt_group > > -> start_rt_bandwidth > > -> start_bandwidth_timer > > -> __hrtimer_start_range_ns > > -> ktime_get() > > Yeah, you are right. > > > > If I just missed how wakeup can call into ktime_get() in current upstream, > > > there is another question: > > > > > > 2) Is it OK that printk calls wakeup so late during suspend? > > > > To clarify again to everybody, we are talking about the place where all > > non-boot CPUs are already hot-unplugged and the last running one has > > disabled interrupts. > > > > I believe that we can't do migration at all now, right? What will we get by > > calling wake_up_process() now anyway ? > > As I already wrote to Rafael, wake_up_process() will change the process > state to TASK_RUNNING so that it can run after we resume from suspend. > > But seeing that the same problem is in upstream I guess what Sergey did > makes more sense if it works for you. If Sergey's fix does not work for you > due to too many messages being printed during device suspend, then we will > have to try something else... Which is exactly my point. :-) Thanks, Rafael
[toc] | [prev] | [next] | [standalone]
| From | Viresh Kumar <viresh.kumar@linaro.org> |
|---|---|
| Date | 2016-07-11 21:10 +0200 |
| Message-ID | <rTOKC-2Ii-19@gated-at.bofh.it> |
| In reply to | #1440446 |
Hi Jan, On 11-07-16, 12:26, Jan Kara wrote: > Yes. We have similar problems as you observe on machines when they do a lot > of printing (usually due to device discovery or similar reasons). The > problem is not fully solved even upstream as Andrew is reluctant to merge > the patches. Sergey (added to CC) has the latest version of the series [1]. Yeah, I saw these patches on last Thursday. I backported all printk patches from 3.10 to mainline to my 3.10 branch and applied your patches on the top. It did work for my case (thanks) and I wanted to give a Tested-by on the thread [1], but by that time it was late Friday for me :) Though I saw a issue with that. [ 12.874909] sched: RT throttling activated for rt_rq ffffffc0ac13fcd0 (cpu 0) [ 12.874909] potential CPU hogs: [ 12.874909] printk (292) On my system, the excessive printing happens during suspend/resume and this happened after all the non-boot CPUs were offlined. So, only CPU 0 was left and that was doing printing for a long time and so these errors :) It resulted in missing some print messages eventually as the scheduler probably didn't schedule this thread for sometime after that. Will it be fine to get the priority of this kthread to a somewhat lower value, etc ? > If you are interested, I can send you the patches for 3.12 kernel which we > carry in SLES kernels and which fixes the issue for us. It is significanly > different from current upstream version but it works good enough for us. Thanks, that will be a good thing to have. I am currently backport 100+ patches from 3.10 to mainline for printk :) Please send them to me (please make sure that you send all the patches touching drivers/printk/ after 3.12, so that I am not left solving merge conflicts for ever :). -- viresh
[toc] | [prev] | [standalone]
Page 3 of 3 — ← Prev page 1 2 [3]
Back to top | Article view | linux.kernel
csiph-web