Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1310019 > unrolled thread
| Started by | Ingo Molnar <mingo@kernel.org> |
|---|---|
| First post | 2016-01-15 11:20 +0100 |
| Last post | 2016-01-19 11:40 +0100 |
| Articles | 5 — 4 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH v2] reboot: Backup orderly_poweroff Ingo Molnar <mingo@kernel.org> - 2016-01-15 11:20 +0100
Re: [PATCH v2] reboot: Backup orderly_poweroff Grygorii Strashko <grygorii.strashko@ti.com> - 2016-01-15 14:40 +0100
Re: [PATCH v2] reboot: Backup orderly_poweroff Russell King - ARM Linux <linux@arm.linux.org.uk> - 2016-01-15 15:20 +0100
Re: [PATCH v2] reboot: Backup orderly_poweroff Ingo Molnar <mingo@kernel.org> - 2016-01-19 10:10 +0100
Re: [PATCH v2] reboot: Backup orderly_poweroff Keerthy <a0393675@ti.com> - 2016-01-19 11:40 +0100
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2016-01-15 11:20 +0100 |
| Subject | Re: [PATCH v2] reboot: Backup orderly_poweroff |
| Message-ID | <qR9E7-4Gx-31@gated-at.bofh.it> |
* One Thousand Gnomes <gnomes@lxorguk.ukuu.org.uk> wrote: > > If kernel_power_off() is called then the system should power off. No ifs and > > whens. > > Even if it doesn't the watchdog should kill it. > > That is broken on some platforms on the watchdog side as the > watchdog shuts down during our power off callbacks - because the system > firmware is too stupid to reset the watchdog as it powers back up (so > keeps rebooting). > > If you watchdog and firmware function properly you shouldn't even have to > care if you crash during the kernel power off. That's a good point as well - if the system is 'stuck' for some notion of stuck, then watchdog drivers can help. Here it's unclear whether user-space even called the sys_reboot() system call. Thanks, Ingo
[toc] | [next] | [standalone]
| From | Grygorii Strashko <grygorii.strashko@ti.com> |
|---|---|
| Date | 2016-01-15 14:40 +0100 |
| Message-ID | <qRcLE-6Jg-19@gated-at.bofh.it> |
| In reply to | #1310019 |
On 01/15/2016 12:14 PM, Ingo Molnar wrote:
>
> * One Thousand Gnomes <gnomes@lxorguk.ukuu.org.uk> wrote:
>
>>> If kernel_power_off() is called then the system should power off. No ifs and
>>> whens.
>>
>> Even if it doesn't the watchdog should kill it.
>>
>> That is broken on some platforms on the watchdog side as the
>> watchdog shuts down during our power off callbacks - because the system
>> firmware is too stupid to reset the watchdog as it powers back up (so
>> keeps rebooting).
>>
>> If you watchdog and firmware function properly you shouldn't even have to
>> care if you crash during the kernel power off.
>
> That's a good point as well - if the system is 'stuck' for some notion of stuck,
> then watchdog drivers can help.
>
Seems ARM doesn't have endless loop implemented in machine_power_off() - so,
not too much chances for Watchdog to fire.
void machine_power_off(void)
{
local_irq_disable();
smp_send_stop();
if (pm_power_off)
pm_power_off();
--- endless loop ?
--- or restart ?
}
[and even if it will be there - 20-30sec is usual timeout for Watchdog and this
enough time to burn the system in case of thermal emergency poweroff :(]
> Here it's unclear whether user-space even called the sys_reboot() system call.
>
That's true - original log [1] has
Nov 30 11:19:22 [ 5.942769] thermal thermal_zone3: critical temperature reached(108 C),shutting down
[...]
Nov 30 11:19:24 [ 7.387900] ahci 4a140000.sata: flags: 64bit ncq sntf stag pm led clo only pmp pio slum part ccc apst
Nov 30 11:19:24 INIT: Switching to runlevel: 0
Nov 30 11:19:24 INIT: Sending processes the TERM signal
and there are no
[ 220.004522] reboot: Power down
Also, It's not the first time this part of code is discussed (thermal emergency poweroff) [2],
so the good question, as for me, is it really required and safe to use orderly_poweroff() in
case of thermal emergency poweroff ([3] as example)?
In general, this kind of use case can be simulated using SysRq on any arch
- [3.290034] Freeing unused kernel memory: 492K (c0a67000 - c0ae2000)
INIT: version 2.88 booting
Starting udev
^^ The issue most probably might happens when system in the process of loading modules
So, once modules loading process is started - fire Sysrq "poweroff(o)"
[1] http://pastebin.ubuntu.com/14326688/
[2] https://lkml.org/lkml/2012/9/18/577
[3] http://review.omapzoom.org/#/c/34898/
--
regards,
-grygorii
[toc] | [prev] | [next] | [standalone]
| From | Russell King - ARM Linux <linux@arm.linux.org.uk> |
|---|---|
| Date | 2016-01-15 15:20 +0100 |
| Message-ID | <qRdom-7gx-11@gated-at.bofh.it> |
| In reply to | #1310114 |
On Fri, Jan 15, 2016 at 03:29:04PM +0200, Grygorii Strashko wrote:
> Seems ARM doesn't have endless loop implemented in machine_power_off() - so,
> not too much chances for Watchdog to fire.
> void machine_power_off(void)
> {
> local_irq_disable();
> smp_send_stop();
>
> if (pm_power_off)
> pm_power_off();
>
> --- endless loop ?
> --- or restart ?
> }
> [and even if it will be there - 20-30sec is usual timeout for Watchdog
> and this enough time to burn the system in case of thermal emergency
> poweroff :(]
I covered this in my reply to Ingo yesterday. The result is that a
failed or unimplemented call drops through to do_exit(0) on behalf of
the calling process, terminating that process. However, as I said
in that same email, I don't think you're getting anywhere near this
code.
> That's true - original log [1] has
> Nov 30 11:19:22 [ 5.942769] thermal thermal_zone3: critical temperature reached(108 C),shutting down
> [...]
> Nov 30 11:19:24 [ 7.387900] ahci 4a140000.sata: flags: 64bit ncq sntf stag pm led clo only pmp pio slum part ccc apst
> Nov 30 11:19:24 INIT: Switching to runlevel: 0
> Nov 30 11:19:24 INIT: Sending processes the TERM signal
>
> and there are no
> [ 220.004522] reboot: Power down
Right, so things are stuck in userspace, which means the system is still
in an active runnable state.
As I mentioned (again) in my email, the issue appears to be that the 'rc'
script is stuck waiting on a FIFO.
The init daemon is trying to do an orderly shutdown. As part of that,
it's executing the 'rc' script, which in systems I've seen, runs through
a set of scripts in the /etc/rc?.d directory in order, which normally
bring up or take down services and perform other sequenced actions.
If this script hangs (as it seems to be doing) it won't get to running
/sbin/poweroff or similar, and that means machine_power_off() won't be
called.
> In general, this kind of use case can be simulated using SysRq on any arch
> - [3.290034] Freeing unused kernel memory: 492K (c0a67000 - c0ae2000)
> INIT: version 2.88 booting
> Starting udev
> ^^ The issue most probably might happens when system in the process of
> loading modules
> So, once modules loading process is started - fire Sysrq "poweroff(o)"
This suggests it could be a udev issue - but without knowing what's
happening inside sysvinit's scripts, it's hard to know for certain.
Adding some debug to the 'rc' script (make sure it works without
rebooting or changing the run level, or have a way of restoring the
file if it fails to boot) so that it's possible to see what it's doing
may be a good idea - the simplest approach may be to just add
set -x
towards the top of the file - which will make it very noisy.
--
RMK's Patch system: http://www.arm.linux.org.uk/developer/patches/
FTTC broadband for 0.8mile line: currently at 9.6Mbps down 400kbps up
according to speedtest.net.
[toc] | [prev] | [next] | [standalone]
| From | Ingo Molnar <mingo@kernel.org> |
|---|---|
| Date | 2016-01-19 10:10 +0100 |
| Message-ID | <qSAsz-5F1-29@gated-at.bofh.it> |
| In reply to | #1310114 |
* Grygorii Strashko <grygorii.strashko@ti.com> wrote:
> On 01/15/2016 12:14 PM, Ingo Molnar wrote:
> >
> > * One Thousand Gnomes <gnomes@lxorguk.ukuu.org.uk> wrote:
> >
> >>> If kernel_power_off() is called then the system should power off. No ifs and
> >>> whens.
> >>
> >> Even if it doesn't the watchdog should kill it.
> >>
> >> That is broken on some platforms on the watchdog side as the
> >> watchdog shuts down during our power off callbacks - because the system
> >> firmware is too stupid to reset the watchdog as it powers back up (so
> >> keeps rebooting).
> >>
> >> If you watchdog and firmware function properly you shouldn't even have to
> >> care if you crash during the kernel power off.
> >
> > That's a good point as well - if the system is 'stuck' for some notion of stuck,
> > then watchdog drivers can help.
> >
>
> Seems ARM doesn't have endless loop implemented in machine_power_off() - so,
> not too much chances for Watchdog to fire.
> void machine_power_off(void)
> {
> local_irq_disable();
> smp_send_stop();
>
> if (pm_power_off)
> pm_power_off();
>
> --- endless loop ?
> --- or restart ?
> }
> [and even if it will be there - 20-30sec is usual timeout for Watchdog and this
> enough time to burn the system in case of thermal emergency poweroff :(]
>
> > Here it's unclear whether user-space even called the sys_reboot() system call.
> >
>
> That's true - original log [1] has
> Nov 30 11:19:22 [ 5.942769] thermal thermal_zone3: critical temperature reached(108 C),shutting down
> [...]
> Nov 30 11:19:24 [ 7.387900] ahci 4a140000.sata: flags: 64bit ncq sntf stag pm led clo only pmp pio slum part ccc apst
> Nov 30 11:19:24 INIT: Switching to runlevel: 0
> Nov 30 11:19:24 INIT: Sending processes the TERM signal
>
> and there are no
> [ 220.004522] reboot: Power down
>
>
> Also, It's not the first time this part of code is discussed (thermal emergency poweroff) [2],
> so the good question, as for me, is it really required and safe to use orderly_poweroff() in
> case of thermal emergency poweroff ([3] as example)?
>
> In general, this kind of use case can be simulated using SysRq on any arch
> - [3.290034] Freeing unused kernel memory: 492K (c0a67000 - c0ae2000)
> INIT: version 2.88 booting
> Starting udev
> ^^ The issue most probably might happens when system in the process of loading modules
> So, once modules loading process is started - fire Sysrq "poweroff(o)"
So I'd say emergency poweroff should be named accordingly - and the
orderly_poweroff() name suggest anything but an emergency, right?
So I'd be fine with the following:
- introduce a poweroff_emergency() core kernel function call
- use it in drivers where it's justified
- poweroff_emergency() has a configurable timeout value. If the timeout value is
set to 0 then it powers the system off immediately.
Functionally it would be mostly equivalent to your current patch (except the '0'
immediate poweroff functionality).
Thanks,
Ingo
[toc] | [prev] | [next] | [standalone]
| From | Keerthy <a0393675@ti.com> |
|---|---|
| Date | 2016-01-19 11:40 +0100 |
| Message-ID | <qSBRF-6uw-25@gated-at.bofh.it> |
| In reply to | #1312027 |
Hi Ingo,
On Tuesday 19 January 2016 02:36 PM, Ingo Molnar wrote:
>
> * Grygorii Strashko <grygorii.strashko@ti.com> wrote:
>
>> On 01/15/2016 12:14 PM, Ingo Molnar wrote:
>>>
>>> * One Thousand Gnomes <gnomes@lxorguk.ukuu.org.uk> wrote:
>>>
>>>>> If kernel_power_off() is called then the system should power off. No ifs and
>>>>> whens.
>>>>
>>>> Even if it doesn't the watchdog should kill it.
>>>>
>>>> That is broken on some platforms on the watchdog side as the
>>>> watchdog shuts down during our power off callbacks - because the system
>>>> firmware is too stupid to reset the watchdog as it powers back up (so
>>>> keeps rebooting).
>>>>
>>>> If you watchdog and firmware function properly you shouldn't even have to
>>>> care if you crash during the kernel power off.
>>>
>>> That's a good point as well - if the system is 'stuck' for some notion of stuck,
>>> then watchdog drivers can help.
>>>
>>
>> Seems ARM doesn't have endless loop implemented in machine_power_off() - so,
>> not too much chances for Watchdog to fire.
>> void machine_power_off(void)
>> {
>> local_irq_disable();
>> smp_send_stop();
>>
>> if (pm_power_off)
>> pm_power_off();
>>
>> --- endless loop ?
>> --- or restart ?
>> }
>> [and even if it will be there - 20-30sec is usual timeout for Watchdog and this
>> enough time to burn the system in case of thermal emergency poweroff :(]
>>
>>> Here it's unclear whether user-space even called the sys_reboot() system call.
>>>
>>
>> That's true - original log [1] has
>> Nov 30 11:19:22 [ 5.942769] thermal thermal_zone3: critical temperature reached(108 C),shutting down
>> [...]
>> Nov 30 11:19:24 [ 7.387900] ahci 4a140000.sata: flags: 64bit ncq sntf stag pm led clo only pmp pio slum part ccc apst
>> Nov 30 11:19:24 INIT: Switching to runlevel: 0
>> Nov 30 11:19:24 INIT: Sending processes the TERM signal
>>
>> and there are no
>> [ 220.004522] reboot: Power down
>>
>>
>> Also, It's not the first time this part of code is discussed (thermal emergency poweroff) [2],
>> so the good question, as for me, is it really required and safe to use orderly_poweroff() in
>> case of thermal emergency poweroff ([3] as example)?
>>
>> In general, this kind of use case can be simulated using SysRq on any arch
>> - [3.290034] Freeing unused kernel memory: 492K (c0a67000 - c0ae2000)
>> INIT: version 2.88 booting
>> Starting udev
>> ^^ The issue most probably might happens when system in the process of loading modules
>> So, once modules loading process is started - fire Sysrq "poweroff(o)"
>
> So I'd say emergency poweroff should be named accordingly - and the
> orderly_poweroff() name suggest anything but an emergency, right?
>
> So I'd be fine with the following:
>
> - introduce a poweroff_emergency() core kernel function call
>
> - use it in drivers where it's justified
>
> - poweroff_emergency() has a configurable timeout value. If the timeout value is
> set to 0 then it powers the system off immediately.
>
> Functionally it would be mostly equivalent to your current patch (except the '0'
> immediate poweroff functionality).
Thanks for the suggestion. I will work on this and get back.
Best Regards,
Keerthy
>
> Thanks,
>
> Ingo
>
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web