Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1303855 > unrolled thread
| Started by | Jacob Pan <jacob.jun.pan@linux.intel.com> |
|---|---|
| First post | 2016-01-07 21:00 +0100 |
| Last post | 2016-01-14 16:40 +0100 |
| Articles | 8 — 2 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH v3 22/22] thermal/intel_powerclamp: Convert the kthread to kthread worker API Jacob Pan <jacob.jun.pan@linux.intel.com> - 2016-01-07 21:00 +0100
Re: [PATCH v3 22/22] thermal/intel_powerclamp: Convert the kthread to kthread worker API Petr Mladek <pmladek@suse.com> - 2016-01-08 17:50 +0100
Re: [PATCH v3 22/22] thermal/intel_powerclamp: Convert the kthread to kthread worker API Jacob Pan <jacob.jun.pan@linux.intel.com> - 2016-01-12 03:30 +0100
Re: [PATCH v3 22/22] thermal/intel_powerclamp: Convert the kthread to kthread worker API Petr Mladek <pmladek@suse.com> - 2016-01-12 11:20 +0100
Re: [PATCH v3 22/22] thermal/intel_powerclamp: Convert the kthread to kthread worker API Jacob Pan <jacob.jun.pan@linux.intel.com> - 2016-01-12 17:30 +0100
Re: [PATCH v3 22/22] thermal/intel_powerclamp: Convert the kthread to kthread worker API Petr Mladek <pmladek@suse.com> - 2016-01-13 11:20 +0100
Re: [PATCH v3 22/22] thermal/intel_powerclamp: Convert the kthread to kthread worker API Jacob Pan <jacob.jun.pan@linux.intel.com> - 2016-01-13 19:00 +0100
Re: [PATCH v3 22/22] thermal/intel_powerclamp: Convert the kthread to kthread worker API Petr Mladek <pmladek@suse.com> - 2016-01-14 16:40 +0100
| From | Jacob Pan <jacob.jun.pan@linux.intel.com> |
|---|---|
| Date | 2016-01-07 21:00 +0100 |
| Subject | Re: [PATCH v3 22/22] thermal/intel_powerclamp: Convert the kthread to kthread worker API |
| Message-ID | <qOoSZ-4ij-3@gated-at.bofh.it> |
On Wed, 18 Nov 2015 14:25:27 +0100 Petr Mladek <pmladek@suse.com> wrote: > From: Petr Mladek <pmladek@suse.com> > To: Andrew Morton <akpm@linux-foundation.org>, Oleg Nesterov > <oleg@redhat.com>, Tejun Heo <tj@kernel.org>, Ingo Molnar > <mingo@redhat.com>, Peter Zijlstra <peterz@infradead.org> Cc: Steven > Rostedt <rostedt@goodmis.org>, "Paul E. McKenney" > <paulmck@linux.vnet.ibm.com>, Josh Triplett <josh@joshtriplett.org>, > Thomas Gleixner <tglx@linutronix.de>, Linus Torvalds > <torvalds@linux-foundation.org>, Jiri Kosina <jkosina@suse.cz>, > Borislav Petkov <bp@suse.de>, Michal Hocko <mhocko@suse.cz>, > linux-mm@kvack.org, Vlastimil Babka <vbabka@suse.cz>, > linux-api@vger.kernel.org, linux-kernel@vger.kernel.org, Petr Mladek > <pmladek@suse.com>, Zhang Rui <rui.zhang@intel.com>, Eduardo Valentin > <edubezval@gmail.com>, Jacob Pan <jacob.jun.pan@linux.intel.com>, > linux-pm@vger.kernel.org Subject: [PATCH v3 22/22] > thermal/intel_powerclamp: Convert the kthread to kthread worker API > Date: Wed, 18 Nov 2015 14:25:27 +0100 X-Mailer: git-send-email 1.8.5.6 > > Kthreads are currently implemented as an infinite loop. Each > has its own variant of checks for terminating, freezing, > awakening. In many cases it is unclear to say in which state > it is and sometimes it is done a wrong way. > > The plan is to convert kthreads into kthread_worker or workqueues > API. It allows to split the functionality into separate operations. > It helps to make a better structure. Also it defines a clean state > where no locks are taken, IRQs blocked, the kthread might sleep > or even be safely migrated. > > The kthread worker API is useful when we want to have a dedicated > single thread for the work. It helps to make sure that it is > available when needed. Also it allows a better control, e.g. > define a scheduling priority. > > This patch converts the intel powerclamp kthreads into the kthread > worker because they need to have a good control over the assigned > CPUs. > I have tested this patchset and found no obvious issues in terms of functionality, power and performance. Tested CPU online/offline, suspend resume, freeze etc. Power numbers are comparable too. e.g. on IVB 8C system. Inject idle from 5 to 50% and read package power while running CPU bound workload. Before: IdlePct Perf RAPL WallPower 5 256.28 16.50 0.0 10 248.86 15.64 0.0 15 209.01 14.57 0.0 20 176.17 13.88 0.0 25 161.25 13.37 0.0 30 165.62 13.38 0.0 35 150.94 12.89 0.0 40 137.45 12.47 0.0 45 123.80 11.83 0.0 50 137.59 11.80 0.0 After: (deb_chroot)root@ubuntu-jp-nfs:~/powercap-power# ./test.py -c 5 IdlePct Perf RAPL WallPower 5 266.30 16.34 0.0 10 226.32 15.27 0.0 15 195.52 14.29 0.0 20 200.96 13.98 0.0 25 174.77 13.08 0.0 30 162.05 13.04 0.0 35 166.70 12.90 0.0 40 134.78 12.12 0.0 45 128.08 11.70 0.0 50 117.74 11.74 0.0 > IMHO, the most natural way is to split one cycle into two works. > First one does some balancing and let the CPU work normal > way for some time. The second work checks what the CPU has done > in the meantime and put it into C-state to reach the required > idle time ratio. The delay between the two works is achieved > by the delayed kthread work. > > The two works have to share some data that used to be local > variables of the single kthread function. This is achieved > by the new per-CPU struct kthread_worker_data. It might look > as a complication. On the other hand, the long original kthread > function was not nice either. > > The patch tries to avoid extra init and cleanup works. All the > actions might be done outside the thread. They are moved > to the functions that create or destroy the worker. Especially, > I checked that the timers are assigned to the right CPU. > > The two works are queuing each other. It makes it a bit tricky to > break it when we want to stop the worker. We use the global and > per-worker "clamping" variables to make sure that the re-queuing > eventually stops. We also cancel the works to make it faster. > Note that the canceling is not reliable because the handling > of the two variables and queuing is not synchronized via a lock. > But it is not a big deal because it is just an optimization. > The job is stopped faster than before in most cases. I am not convinced this added complexity is necessary, here are my concerns by breaking down into two work items. - overhead of queuing, per cpu data as you already mentioned. - since we need to have very tight timing control, two items may limit our turnaround time. Wouldn't it take one extra tick for the scheduler to run the balance work then add delay? as opposed to just schedule_timeout()? - vulnerable to future changes of queuing work Jacob
[toc] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-01-08 17:50 +0100 |
| Message-ID | <qOIoH-UX-15@gated-at.bofh.it> |
| In reply to | #1303855 |
On Thu 2016-01-07 11:55:31, Jacob Pan wrote:
> On Wed, 18 Nov 2015 14:25:27 +0100
> Petr Mladek <pmladek@suse.com> wrote:
> I have tested this patchset and found no obvious issues in terms of
> functionality, power and performance. Tested CPU online/offline,
> suspend resume, freeze etc.
> Power numbers are comparable too. e.g. on IVB 8C system. Inject idle
> from 5 to 50% and read package power while running CPU bound workload.
Great news. Thanks a lot for testing.
> > IMHO, the most natural way is to split one cycle into two works.
> > First one does some balancing and let the CPU work normal
> > way for some time. The second work checks what the CPU has done
> > in the meantime and put it into C-state to reach the required
> > idle time ratio. The delay between the two works is achieved
> > by the delayed kthread work.
> >
> > The two works have to share some data that used to be local
> > variables of the single kthread function. This is achieved
> > by the new per-CPU struct kthread_worker_data. It might look
> > as a complication. On the other hand, the long original kthread
> > function was not nice either.
> >
> > The two works are queuing each other. It makes it a bit tricky to
> > break it when we want to stop the worker. We use the global and
> > per-worker "clamping" variables to make sure that the re-queuing
> > eventually stops. We also cancel the works to make it faster.
> > Note that the canceling is not reliable because the handling
> > of the two variables and queuing is not synchronized via a lock.
> > But it is not a big deal because it is just an optimization.
> > The job is stopped faster than before in most cases.
> I am not convinced this added complexity is necessary, here are my
> concerns by breaking down into two work items.
I am not super happy with the split either. But the current state has
its drawback as well.
> - overhead of queuing,
Good question. Here is a rather typical snippet from function_graph
tracer of the clamp_balancing func:
31) | clamp_balancing_func() {
31) | queue_delayed_kthread_work() {
31) | __queue_delayed_kthread_work() {
31) | add_timer() {
31) 4.906 us | }
31) 5.959 us | }
31) 9.702 us | }
31) + 10.878 us | }
On one hand it spends most of the time (10 of 11 secs) in queueing
the work. On the other hand, half of this time is spent on adding
the timer. schedule_timeout() would need to setup the timer as well.
Here is a snippet from clamp_idle_injection_func()
31) | clamp_idle_injection_func() {
31) | smp_apic_timer_interrupt() {
31) + 67.523 us | }
31) | smp_apic_timer_interrupt() {
31) + 59.946 us | }
...
31) | queue_kthread_work() {
31) 4.314 us | }
31) * 24075.11 us | }
Of course, it spends most of the time in the idle state. Anyway, the
time spent on queuing is negligible in compare with the time spent
in the several timer interrupt handlers.
> per cpu data as you already mentioned.
On the other hand, the variables need to be stored somewhere.
Also it helps to split the rather long function into more pieces.
> - since we need to have very tight timing control, two items may limit
> our turnaround time. Wouldn't it take one extra tick for the scheduler
> to run the balance work then add delay? as opposed to just
> schedule_timeout()?
Kthread worker processes works until the queue is empty. It calls
try_to_freeze() and __preempt_schedule() between the works.
Where __preempt_schedule() is hidden in the spin_unlock_irq().
try_to_freeze() is in the original code as well.
Is the __preempt_schedule() a problem? It allows to switch the process
when needed. I thought that it was safe because try_to_freeze() might
have slept as well.
> - vulnerable to future changes of queuing work
The question is if it is safe to sleep, freeze, or even migrate
the system between the works. It looks like because of the
try_to_freeze() and schedule_interrupt() calls in the original code.
BTW: I wonder if the original code correctly handle freezing after
the schedule_timeout(). It does not call try_to_freeze()
there and the forced idle states might block freezing.
I think that the small overhead of kthread works is worth
solving such bugs. It makes it easier to maintain these
sleeping states.
Thanks a lot for feedback,
Petr
[toc] | [prev] | [next] | [standalone]
| From | Jacob Pan <jacob.jun.pan@linux.intel.com> |
|---|---|
| Date | 2016-01-12 03:30 +0100 |
| Message-ID | <qPWSC-2Tc-11@gated-at.bofh.it> |
| In reply to | #1304723 |
On Fri, 8 Jan 2016 17:49:31 +0100 Petr Mladek <pmladek@suse.com> wrote: > Is the __preempt_schedule() a problem? It allows to switch the process > when needed. I thought that it was safe because try_to_freeze() might > have slept as well. > not a problem. i originally thought queue_kthread_work() may add delay but it doesn't since there is no other work on this kthread. > > > - vulnerable to future changes of queuing work > > The question is if it is safe to sleep, freeze, or even migrate > the system between the works. It looks like because of the > try_to_freeze() and schedule_interrupt() calls in the original code. > > BTW: I wonder if the original code correctly handle freezing after > the schedule_timeout(). It does not call try_to_freeze() > there and the forced idle states might block freezing. > I think that the small overhead of kthread works is worth > solving such bugs. It makes it easier to maintain these > sleeping states. it is in a while loop, so try_to_freeze() gets called. Am I missing something? Thanks, Jacob
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-01-12 11:20 +0100 |
| Message-ID | <qQ4dv-7Y1-45@gated-at.bofh.it> |
| In reply to | #1306933 |
On Mon 2016-01-11 18:17:18, Jacob Pan wrote:
> On Fri, 8 Jan 2016 17:49:31 +0100
> Petr Mladek <pmladek@suse.com> wrote:
>
> > Is the __preempt_schedule() a problem? It allows to switch the process
> > when needed. I thought that it was safe because try_to_freeze() might
> > have slept as well.
> >
> not a problem. i originally thought queue_kthread_work() may add
> delay but it doesn't since there is no other work on this kthread.
Great.
> > > - vulnerable to future changes of queuing work
> >
> > The question is if it is safe to sleep, freeze, or even migrate
> > the system between the works. It looks like because of the
> > try_to_freeze() and schedule_interrupt() calls in the original code.
> >
> > BTW: I wonder if the original code correctly handle freezing after
> > the schedule_timeout(). It does not call try_to_freeze()
> > there and the forced idle states might block freezing.
> > I think that the small overhead of kthread works is worth
> > solving such bugs. It makes it easier to maintain these
> > sleeping states.
> it is in a while loop, so try_to_freeze() gets called. Am I missing
> something?
But it might take some time until try_to_freeze() is called.
If I get it correctly. try_to_freeze_tasks() wakes freezable
tasks to get them into the fridge. If clamp_thread() is waken
from that schedule_timeout_interruptible(), it still might inject
the idle state before calling try_to_freeze(). It means that freezer
needs to wait "quite" some time until the kthread ends up in the
fridge.
Hmm, even my conversion does not solve this entirely. We might
need to call freezing(current) in the
while (time_before(jiffies, target_jiffies)) {
cycle. And break injecting the idle state when freezing is requested.
Or do I miss something, please?
Best Regards,
Petr
Best Regards,
Petr
[toc] | [prev] | [next] | [standalone]
| From | Jacob Pan <jacob.jun.pan@linux.intel.com> |
|---|---|
| Date | 2016-01-12 17:30 +0100 |
| Message-ID | <qQ9Zx-3q9-37@gated-at.bofh.it> |
| In reply to | #1307218 |
On Tue, 12 Jan 2016 11:11:29 +0100
Petr Mladek <pmladek@suse.com> wrote:
> > > BTW: I wonder if the original code correctly handle freezing after
> > > the schedule_timeout(). It does not call try_to_freeze()
> > > there and the forced idle states might block freezing.
> > > I think that the small overhead of kthread works is worth
> > > solving such bugs. It makes it easier to maintain these
> > > sleeping states.
> > it is in a while loop, so try_to_freeze() gets called. Am I missing
> > something?
>
> But it might take some time until try_to_freeze() is called.
> If I get it correctly. try_to_freeze_tasks() wakes freezable
> tasks to get them into the fridge. If clamp_thread() is waken
> from that schedule_timeout_interruptible(), it still might inject
> the idle state before calling try_to_freeze(). It means that freezer
> needs to wait "quite" some time until the kthread ends up in the
> fridge.
>
> Hmm, even my conversion does not solve this entirely. We might
> need to call freezing(current) in the
>
> while (time_before(jiffies, target_jiffies)) {
>
> cycle. And break injecting the idle state when freezing is requested.
The injection time for each period is very short, default 6ms. While on
the other side the default freeze timeout is 20 sec. So I think task
freeze can wait :)
i.e.
unsigned int __read_mostly freeze_timeout_msecs = 20 * MSEC_PER_SEC;
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-01-13 11:20 +0100 |
| Message-ID | <qQqGZ-6FS-1@gated-at.bofh.it> |
| In reply to | #1307610 |
On Tue 2016-01-12 08:20:21, Jacob Pan wrote:
> On Tue, 12 Jan 2016 11:11:29 +0100
> Petr Mladek <pmladek@suse.com> wrote:
>
> > > > BTW: I wonder if the original code correctly handle freezing after
> > > > the schedule_timeout(). It does not call try_to_freeze()
> > > > there and the forced idle states might block freezing.
> > > > I think that the small overhead of kthread works is worth
> > > > solving such bugs. It makes it easier to maintain these
> > > > sleeping states.
> > > it is in a while loop, so try_to_freeze() gets called. Am I missing
> > > something?
> >
> > But it might take some time until try_to_freeze() is called.
> > If I get it correctly. try_to_freeze_tasks() wakes freezable
> > tasks to get them into the fridge. If clamp_thread() is waken
> > from that schedule_timeout_interruptible(), it still might inject
> > the idle state before calling try_to_freeze(). It means that freezer
> > needs to wait "quite" some time until the kthread ends up in the
> > fridge.
> >
> > Hmm, even my conversion does not solve this entirely. We might
> > need to call freezing(current) in the
> >
> > while (time_before(jiffies, target_jiffies)) {
> >
> > cycle. And break injecting the idle state when freezing is requested.
>
> The injection time for each period is very short, default 6ms. While on
> the other side the default freeze timeout is 20 sec. So I think task
> freeze can wait :)
> i.e.
> unsigned int __read_mostly freeze_timeout_msecs = 20 * MSEC_PER_SEC;
You are right. And it does not make sense to add an extra
freezer-specific code if not really necessary.
Otherwise, I will keep the conversion into the kthread worker as is
for now. Please, let me know if you are strongly against the split
into the two works.
Best Regards,
Petr
[toc] | [prev] | [next] | [standalone]
| From | Jacob Pan <jacob.jun.pan@linux.intel.com> |
|---|---|
| Date | 2016-01-13 19:00 +0100 |
| Message-ID | <qQxSa-37P-9@gated-at.bofh.it> |
| In reply to | #1308245 |
On Wed, 13 Jan 2016 11:18:31 +0100 Petr Mladek <pmladek@suse.com> wrote: > > unsigned int __read_mostly freeze_timeout_msecs = 20 * > > MSEC_PER_SEC; > > You are right. And it does not make sense to add an extra > freezer-specific code if not really necessary. > > Otherwise, I will keep the conversion into the kthread worker as is > for now. Please, let me know if you are strongly against the split > into the two works. I am fine with the split now. Another question, are you planning to convert acpi_pad.c as well? It uses kthread similar way. Jacob
[toc] | [prev] | [next] | [standalone]
| From | Petr Mladek <pmladek@suse.com> |
|---|---|
| Date | 2016-01-14 16:40 +0100 |
| Message-ID | <qQSae-yO-39@gated-at.bofh.it> |
| In reply to | #1308687 |
On Wed 2016-01-13 09:53:53, Jacob Pan wrote: > On Wed, 13 Jan 2016 11:18:31 +0100 > Petr Mladek <pmladek@suse.com> wrote: > > Otherwise, I will keep the conversion into the kthread worker as is > > for now. Please, let me know if you are strongly against the split > > into the two works. > I am fine with the split now. Great. > Another question, are you planning to convert acpi_pad.c as well? It > uses kthread similar way. Yup. I would like to convert as many kthreads as possible either to the kthread worker or workqueue APIs in the long term. Best Regards, Petr
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web