Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1585779 > unrolled thread

v4.10: kernel stack frame pointer .. has bad value (null)

Started byPavel Machek <pavel@ucw.cz>
First post2017-02-21 23:20 +0100
Last post2017-02-25 06:30 +0100
Articles 12 — 3 participants

Back to article view | Back to linux.kernel


Contents

  v4.10: kernel stack frame pointer .. has bad value (null) Pavel Machek <pavel@ucw.cz> - 2017-02-21 23:20 +0100
    Re: v4.10: kernel stack frame pointer .. has bad value (null) "H. Peter Anvin" <hpa@zytor.com> - 2017-02-22 00:20 +0100
      Re: v4.10: kernel stack frame pointer .. has bad value (null) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-02-22 18:00 +0100
        Re: v4.10: kernel stack frame pointer .. has bad value (null) "H. Peter Anvin" <hpa@zytor.com> - 2017-02-22 22:00 +0100
          Re: v4.10: kernel stack frame pointer .. has bad value (null) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-02-22 22:20 +0100
    Re: v4.10: kernel stack frame pointer .. has bad value (null) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-02-22 00:20 +0100
      Re: v4.10: kernel stack frame pointer .. has bad value (null) Pavel Machek <pavel@ucw.cz> - 2017-02-22 22:10 +0100
        Re: v4.10: kernel stack frame pointer .. has bad value (null) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-02-22 22:30 +0100
          Re: v4.10: kernel stack frame pointer .. has bad value (null) Pavel Machek <pavel@ucw.cz> - 2017-02-22 23:50 +0100
            Re: v4.10: kernel stack frame pointer .. has bad value (null) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-02-23 00:10 +0100
              Re: v4.10: kernel stack frame pointer .. has bad value (null) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-02-23 00:20 +0100
                Re: v4.10: kernel stack frame pointer .. has bad value (null) Josh Poimboeuf <jpoimboe@redhat.com> - 2017-02-25 06:30 +0100

#1585779 — v4.10: kernel stack frame pointer .. has bad value (null)

FromPavel Machek <pavel@ucw.cz>
Date2017-02-21 23:20 +0100
Subjectv4.10: kernel stack frame pointer .. has bad value (null)
Message-ID<tdqWT-l3-33@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Hi!

Thinkpad X220, in 32 bit mode... and I'm getting rather scary messages
from kernel during boot:

Git blame says that message comes from commit

commit 24d86f59093b0bcb3756cdf47f2db10ff4e90dbb
Author: Josh Poimboeuf <jpoimboe@redhat.com>
Date:   Thu Oct 27 08:10:58 2016 -0500

    x86/unwind: Ensure stack grows down

    Add a sanity check to ensure the stack only grows down, and print
    a
        warning if the check fails.

Any ideas?
							Pavel

[    1.047295] [drm] Memory usable by graphics device = 2048M
[    1.047356] [drm] Replacing VGA console driver
[    1.048029] Console: switching to colour dummy device 80x25
[    1.048348] WARNING: kernel stack frame pointer at f50cdf98 in
swapper/2:0 has bad value   (null)
[    1.048349] unwind stack type:0 next_sp:  (null) mask:a graph_idx:0
[    1.048352] f50cdebc: 00000000f50cdec4 (0xf50cdec4)
[    1.048356] f50cdec0: 00000000c40489b7 (irq_exit+0x87/0xa0)
[    1.048357] f50cdec4: 00000000f50cded0 (0xf50cded0)
[    1.048361] f50cdec8: 00000000c402f6d3
(smp_apic_timer_interrupt+0x33/0x40)
[    1.048362] f50cdecc: 000000003e7c28eb (0x3e7c28eb)
[    1.048363] f50cded0: 00000000f50cded9 (0xf50cded9)
[    1.048366] f50cded4: 00000000c4b7ac8e
(apic_timer_interrupt+0x36/0x3c)
[    1.048367] f50cded8: 000000003e7c28eb (0x3e7c28eb)
[    1.048368] f50cdedc: 0000000000000001 (0x1)
[    1.048370] f50cdee0: 00000000f57c41c0 (0xf57c41c0)
[    1.048370] f50cdee4: 0000000000000000 ...
[    1.048371] f50cdee8: 0000000000000002 (0x2)
[    1.048372] f50cdeec: 00000000f50cdf30 (0xf50cdf30)
[    1.048373] f50cdef0: 000000003e7c28eb (0x3e7c28eb)
[    1.048376] f50cdef4: 00000000c506007b (cltrack_prog+0x7fb/0x1000)
[    1.048377] f50cdef8: 000000000000007b (0x7b)
[    1.048378] f50cdefc: 00000000000000d8 (0xd8)
[    1.048379] f50cdf00: 00000000c50600e0 (cltrack_prog+0x860/0x1000)
[    1.048380] f50cdf04: 00000000ffffff10 (0xffffff10)
[    1.048384] f50cdf08: 00000000c48d05b5
(cpuidle_enter_state+0xf5/0x220)
[    1.048384] f50cdf0c: 0000000000000060 (0x60)
[    1.048385] f50cdf10: 0000000000200286 (0x200286)
[    1.048386] f50cdf14: 0000000000000000 ...
[    1.048387] f50cdf18: 00000000ffb8fad0 (0xffb8fad0)
[    1.048388] f50cdf1c: 000000003e6bbd2b (0x3e6bbd2b)
[    1.048389] f50cdf20: 0000000000000000 ...
[    1.048391] f50cdf24: 00000000c506c680 (max_cstate+0x2c/0x2c)
[    1.048394] f50cdf28: 00000000c506c680 (max_cstate+0x2c/0x2c)
[    1.048394] f50cdf2c: 00000000f50a2040 (0xf50a2040)
[    1.048395] f50cdf30: 00000000f50cdf3c (0xf50cdf3c)
[    1.048397] f50cdf34: 00000000c48d06ff (cpuidle_enter+0xf/0x20)
[    1.048398] f50cdf38: 0000000000000000 ...
[    1.048399] f50cdf3c: 00000000f50cdf48 (0xf50cdf48)
[    1.048402] f50cdf40: 00000000c4083113 (call_cpuidle+0x23/0x40)
[    1.048403] f50cdf44: 00000000ffb8fad0 (0xffb8fad0)
[    1.048404] f50cdf48: 00000000f50cdf60 (0xf50cdf60)
[    1.048406] f50cdf4c: 00000000c4083364 (do_idle+0x174/0x1d0)
[    1.048407] f50cdf50: 00000000f50a2040 (0xf50a2040)
[    1.048408] f50cdf54: 000000008c0ed583 (0x8c0ed583)
[    1.048409] f50cdf58: 0000000000000087 (0x87)
[    1.048410] f50cdf5c: 000000009779245f (0x9779245f)
[    1.048411] f50cdf60: 00000000f50cdf78 (0xf50cdf78)
[    1.048413] f50cdf64: 00000000c408361d
(cpu_startup_entry+0x5d/0x60)
[    1.048414] f50cdf68: 0000000033ee4c15 (0x33ee4c15)
[    1.048415] f50cdf6c: 00000000c9d84398 (0xc9d84398)
[    1.048417] f50cdf70: 0000000002100800 (0x2100800)
[    1.048418] f50cdf74: 000000000fdc6fb2 (0xfdc6fb2)
[    1.048418] f50cdf78: 00000000f50cdf98 (0xf50cdf98)
[    1.048421] f50cdf7c: 00000000c402d216
(start_secondary+0x176/0x1c0)
[    1.048422] f50cdf80: 000000000fdc6fb2 (0xfdc6fb2)
[    1.048423] f50cdf84: 00000000d5425adc (0xd5425adc)
[    1.048424] f50cdf88: 0000000000000002 (0x2)
[    1.048425] f50cdf8c: 00000000ffffffff (0xffffffff)
[    1.048425] f50cdf90: 0000000000000000 ...
[    1.048426] f50cdf94: 00000000f50cdfac (0xf50cdfac)
[    1.048427] f50cdf98: 0000000000000000 ...
[    1.048429] f50cdf9c: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
[    1.048429] f50cdfa0: 0000000000200002 (0x200002)
[    1.048430] f50cdfa4: 0000000000000000 ...
[    1.048432] f50cdfa8: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
[    1.048432] f50cdfac: 0000000000000000 ...
[    1.048433] f50cdff4: 0000000000000100 (0x100)
[    1.048434] f50cdff8: 0000000000000200 (0x200)
[    1.048435] f50cdffc: 0000000000000000 ...
[    1.060368] [drm] Supports vblank timestamp caching Rev 2
(21.10.2013).
[    1.060373] [drm] Driver supports precise vblank timestamp query.
[    1.061668] i915 0000:00:02.0: vgaarb: changed VGA decodes:
olddecodes=io+mem,decodes=io+mem:owns=io+mem
[    1.146917] ACPI: Video Device [VID] (multi-head: yes  rom: no
post: no)

-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [next] | [standalone]


#1585796

From"H. Peter Anvin" <hpa@zytor.com>
Date2017-02-22 00:20 +0100
Message-ID<tdrSV-11G-1@gated-at.bofh.it>
In reply to#1585779
On 02/21/17 15:12, Josh Poimboeuf wrote:
>>
>> commit 24d86f59093b0bcb3756cdf47f2db10ff4e90dbb
>> Author: Josh Poimboeuf <jpoimboe@redhat.com>
>> Date:   Thu Oct 27 08:10:58 2016 -0500
>>
>>     x86/unwind: Ensure stack grows down
>>
>>     Add a sanity check to ensure the stack only grows down, and print
>>     a
>>         warning if the check fails.
>>
>> Any ideas?
> 
> Hi Pavel,
> 
> I don't think I've seen this one.  Any chance this came after resuming
> from a hibernation or suspend?
> 
> 
>> [    1.047295] [drm] Memory usable by graphics device = 2048M
>> [    1.047356] [drm] Replacing VGA console driver
>> [    1.048029] Console: switching to colour dummy device 80x25
>> [    1.048348] WARNING: kernel stack frame pointer at f50cdf98 in
>> swapper/2:0 has bad value   (null)
>> [    1.048349] unwind stack type:0 next_sp:  (null) mask:a graph_idx:0
>> [    1.048352] f50cdebc: 00000000f50cdec4 (0xf50cdec4)
                            ^^^^^^^^^^^^^^^^

FWIW, it would be really darned nice to not have all those zeroes in a
32-bit stack frame dump.

Is not a zero stack frame pointer value an end of stack token?

	-hpa

[toc] | [prev] | [next] | [standalone]


#1586317

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2017-02-22 18:00 +0100
Message-ID<tdIqK-4Fw-13@gated-at.bofh.it>
In reply to#1585796
On Tue, Feb 21, 2017 at 03:15:36PM -0800, H. Peter Anvin wrote:
> On 02/21/17 15:12, Josh Poimboeuf wrote:
> >>
> >> commit 24d86f59093b0bcb3756cdf47f2db10ff4e90dbb
> >> Author: Josh Poimboeuf <jpoimboe@redhat.com>
> >> Date:   Thu Oct 27 08:10:58 2016 -0500
> >>
> >>     x86/unwind: Ensure stack grows down
> >>
> >>     Add a sanity check to ensure the stack only grows down, and print
> >>     a
> >>         warning if the check fails.
> >>
> >> Any ideas?
> > 
> > Hi Pavel,
> > 
> > I don't think I've seen this one.  Any chance this came after resuming
> > from a hibernation or suspend?
> > 
> > 
> >> [    1.047295] [drm] Memory usable by graphics device = 2048M
> >> [    1.047356] [drm] Replacing VGA console driver
> >> [    1.048029] Console: switching to colour dummy device 80x25
> >> [    1.048348] WARNING: kernel stack frame pointer at f50cdf98 in
> >> swapper/2:0 has bad value   (null)
> >> [    1.048349] unwind stack type:0 next_sp:  (null) mask:a graph_idx:0
> >> [    1.048352] f50cdebc: 00000000f50cdec4 (0xf50cdec4)
>                             ^^^^^^^^^^^^^^^^
> 
> FWIW, it would be really darned nice to not have all those zeroes in a
> 32-bit stack frame dump.

Yeah, I'll fix that.

> Is not a zero stack frame pointer value an end of stack token?

There's no end of stack "token" per se, though any frame pointer value
outside the bounds of the stack will terminate the stack trace (and that
still happened here).

The warning is because the stack trace didn't make it all the way to the
"end" location of the stack (right before the syscall pt_regs location).
The warning is part of the effort to ensure reliable stacks.

-- 
Josh

[toc] | [prev] | [next] | [standalone]


#1586470

From"H. Peter Anvin" <hpa@zytor.com>
Date2017-02-22 22:00 +0100
Message-ID<tdMb0-7mE-17@gated-at.bofh.it>
In reply to#1586317
On 02/22/17 08:45, Josh Poimboeuf wrote:
>>
>> FWIW, it would be really darned nice to not have all those zeroes in a
>> 32-bit stack frame dump.
> 
> Yeah, I'll fix that.
> 
>> Is not a zero stack frame pointer value an end of stack token?
> 
> There's no end of stack "token" per se, though any frame pointer value
> outside the bounds of the stack will terminate the stack trace (and that
> still happened here).
> 

Well, my understanding is that at least gdb and perhaps other unwinders
consider a zero stack frame pointer to be an indicator that the stack
has reached its end.  That's why I'm wondering if this is possible in
this case or if it is unlikely because of the value.

> The warning is because the stack trace didn't make it all the way to the
> "end" location of the stack (right before the syscall pt_regs location).
> The warning is part of the effort to ensure reliable stacks.

It would be useful to get an understanding why...

	-hpa

[toc] | [prev] | [next] | [standalone]


#1586494

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2017-02-22 22:20 +0100
Message-ID<tdMun-7Lo-29@gated-at.bofh.it>
In reply to#1586470
On Wed, Feb 22, 2017 at 12:51:11PM -0800, H. Peter Anvin wrote:
> On 02/22/17 08:45, Josh Poimboeuf wrote:
> >>
> >> FWIW, it would be really darned nice to not have all those zeroes in a
> >> 32-bit stack frame dump.
> > 
> > Yeah, I'll fix that.
> > 
> >> Is not a zero stack frame pointer value an end of stack token?
> > 
> > There's no end of stack "token" per se, though any frame pointer value
> > outside the bounds of the stack will terminate the stack trace (and that
> > still happened here).
> > 
> 
> Well, my understanding is that at least gdb and perhaps other unwinders
> consider a zero stack frame pointer to be an indicator that the stack
> has reached its end.  That's why I'm wondering if this is possible in
> this case or if it is unlikely because of the value.

I'm not sure I follow your question.  The frame pointer was zero, and
that did cause the unwinder to stop the stack trace.  The warning was
because it ended in an unexpected place.

> > The warning is because the stack trace didn't make it all the way to the
> > "end" location of the stack (right before the syscall pt_regs location).
> > The warning is part of the effort to ensure reliable stacks.
> 
> It would be useful to get an understanding why...

Agreed...

-- 
Josh

[toc] | [prev] | [next] | [standalone]


#1585800

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2017-02-22 00:20 +0100
Message-ID<tdrSV-11G-3@gated-at.bofh.it>
In reply to#1585779
On Tue, Feb 21, 2017 at 11:14:18PM +0100, Pavel Machek wrote:
> Hi!
> 
> Thinkpad X220, in 32 bit mode... and I'm getting rather scary messages
> from kernel during boot:
> 
> Git blame says that message comes from commit
> 
> commit 24d86f59093b0bcb3756cdf47f2db10ff4e90dbb
> Author: Josh Poimboeuf <jpoimboe@redhat.com>
> Date:   Thu Oct 27 08:10:58 2016 -0500
> 
>     x86/unwind: Ensure stack grows down
> 
>     Add a sanity check to ensure the stack only grows down, and print
>     a
>         warning if the check fails.
> 
> Any ideas?

Hi Pavel,

I don't think I've seen this one.  Any chance this came after resuming
from a hibernation or suspend?


> [    1.047295] [drm] Memory usable by graphics device = 2048M
> [    1.047356] [drm] Replacing VGA console driver
> [    1.048029] Console: switching to colour dummy device 80x25
> [    1.048348] WARNING: kernel stack frame pointer at f50cdf98 in
> swapper/2:0 has bad value   (null)
> [    1.048349] unwind stack type:0 next_sp:  (null) mask:a graph_idx:0
> [    1.048352] f50cdebc: 00000000f50cdec4 (0xf50cdec4)
> [    1.048356] f50cdec0: 00000000c40489b7 (irq_exit+0x87/0xa0)
> [    1.048357] f50cdec4: 00000000f50cded0 (0xf50cded0)
> [    1.048361] f50cdec8: 00000000c402f6d3
> (smp_apic_timer_interrupt+0x33/0x40)
> [    1.048362] f50cdecc: 000000003e7c28eb (0x3e7c28eb)
> [    1.048363] f50cded0: 00000000f50cded9 (0xf50cded9)
> [    1.048366] f50cded4: 00000000c4b7ac8e
> (apic_timer_interrupt+0x36/0x3c)
> [    1.048367] f50cded8: 000000003e7c28eb (0x3e7c28eb)
> [    1.048368] f50cdedc: 0000000000000001 (0x1)
> [    1.048370] f50cdee0: 00000000f57c41c0 (0xf57c41c0)
> [    1.048370] f50cdee4: 0000000000000000 ...
> [    1.048371] f50cdee8: 0000000000000002 (0x2)
> [    1.048372] f50cdeec: 00000000f50cdf30 (0xf50cdf30)
> [    1.048373] f50cdef0: 000000003e7c28eb (0x3e7c28eb)
> [    1.048376] f50cdef4: 00000000c506007b (cltrack_prog+0x7fb/0x1000)
> [    1.048377] f50cdef8: 000000000000007b (0x7b)
> [    1.048378] f50cdefc: 00000000000000d8 (0xd8)
> [    1.048379] f50cdf00: 00000000c50600e0 (cltrack_prog+0x860/0x1000)
> [    1.048380] f50cdf04: 00000000ffffff10 (0xffffff10)
> [    1.048384] f50cdf08: 00000000c48d05b5
> (cpuidle_enter_state+0xf5/0x220)
> [    1.048384] f50cdf0c: 0000000000000060 (0x60)
> [    1.048385] f50cdf10: 0000000000200286 (0x200286)
> [    1.048386] f50cdf14: 0000000000000000 ...
> [    1.048387] f50cdf18: 00000000ffb8fad0 (0xffb8fad0)
> [    1.048388] f50cdf1c: 000000003e6bbd2b (0x3e6bbd2b)
> [    1.048389] f50cdf20: 0000000000000000 ...
> [    1.048391] f50cdf24: 00000000c506c680 (max_cstate+0x2c/0x2c)
> [    1.048394] f50cdf28: 00000000c506c680 (max_cstate+0x2c/0x2c)
> [    1.048394] f50cdf2c: 00000000f50a2040 (0xf50a2040)
> [    1.048395] f50cdf30: 00000000f50cdf3c (0xf50cdf3c)
> [    1.048397] f50cdf34: 00000000c48d06ff (cpuidle_enter+0xf/0x20)
> [    1.048398] f50cdf38: 0000000000000000 ...
> [    1.048399] f50cdf3c: 00000000f50cdf48 (0xf50cdf48)
> [    1.048402] f50cdf40: 00000000c4083113 (call_cpuidle+0x23/0x40)
> [    1.048403] f50cdf44: 00000000ffb8fad0 (0xffb8fad0)
> [    1.048404] f50cdf48: 00000000f50cdf60 (0xf50cdf60)
> [    1.048406] f50cdf4c: 00000000c4083364 (do_idle+0x174/0x1d0)
> [    1.048407] f50cdf50: 00000000f50a2040 (0xf50a2040)
> [    1.048408] f50cdf54: 000000008c0ed583 (0x8c0ed583)
> [    1.048409] f50cdf58: 0000000000000087 (0x87)
> [    1.048410] f50cdf5c: 000000009779245f (0x9779245f)
> [    1.048411] f50cdf60: 00000000f50cdf78 (0xf50cdf78)
> [    1.048413] f50cdf64: 00000000c408361d
> (cpu_startup_entry+0x5d/0x60)
> [    1.048414] f50cdf68: 0000000033ee4c15 (0x33ee4c15)
> [    1.048415] f50cdf6c: 00000000c9d84398 (0xc9d84398)
> [    1.048417] f50cdf70: 0000000002100800 (0x2100800)
> [    1.048418] f50cdf74: 000000000fdc6fb2 (0xfdc6fb2)
> [    1.048418] f50cdf78: 00000000f50cdf98 (0xf50cdf98)
> [    1.048421] f50cdf7c: 00000000c402d216
> (start_secondary+0x176/0x1c0)
> [    1.048422] f50cdf80: 000000000fdc6fb2 (0xfdc6fb2)
> [    1.048423] f50cdf84: 00000000d5425adc (0xd5425adc)
> [    1.048424] f50cdf88: 0000000000000002 (0x2)
> [    1.048425] f50cdf8c: 00000000ffffffff (0xffffffff)
> [    1.048425] f50cdf90: 0000000000000000 ...
> [    1.048426] f50cdf94: 00000000f50cdfac (0xf50cdfac)
> [    1.048427] f50cdf98: 0000000000000000 ...
> [    1.048429] f50cdf9c: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> [    1.048429] f50cdfa0: 0000000000200002 (0x200002)
> [    1.048430] f50cdfa4: 0000000000000000 ...
> [    1.048432] f50cdfa8: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> [    1.048432] f50cdfac: 0000000000000000 ...
> [    1.048433] f50cdff4: 0000000000000100 (0x100)
> [    1.048434] f50cdff8: 0000000000000200 (0x200)
> [    1.048435] f50cdffc: 0000000000000000 ...
> [    1.060368] [drm] Supports vblank timestamp caching Rev 2
> (21.10.2013).
> [    1.060373] [drm] Driver supports precise vblank timestamp query.
> [    1.061668] i915 0000:00:02.0: vgaarb: changed VGA decodes:
> olddecodes=io+mem,decodes=io+mem:owns=io+mem
> [    1.146917] ACPI: Video Device [VID] (multi-head: yes  rom: no
> post: no)
> 
> -- 
> (english) http://www.livejournal.com/~pavelmachek
> (cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html



-- 
Josh

[toc] | [prev] | [next] | [standalone]


#1586480

FromPavel Machek <pavel@ucw.cz>
Date2017-02-22 22:10 +0100
Message-ID<tdMkG-7Gh-15@gated-at.bofh.it>
In reply to#1585800

[Multipart message — attachments visible in raw view] — view raw

Hi!

> > Thinkpad X220, in 32 bit mode... and I'm getting rather scary messages
> > from kernel during boot:
> > 
> > Git blame says that message comes from commit
> > 
> > commit 24d86f59093b0bcb3756cdf47f2db10ff4e90dbb
> > Author: Josh Poimboeuf <jpoimboe@redhat.com>
> > Date:   Thu Oct 27 08:10:58 2016 -0500
> > 
> >     x86/unwind: Ensure stack grows down
> > 
> >     Add a sanity check to ensure the stack only grows down, and print
> >     a
> >         warning if the check fails.
> > 
> > Any ideas?
> 
> Hi Pavel,
> 
> I don't think I've seen this one.  Any chance this came after resuming
> from a hibernation or suspend?

No, it was during the boot. Notice the timestamps...
								Pavel

> > [    1.047295] [drm] Memory usable by graphics device = 2048M
> > [    1.047356] [drm] Replacing VGA console driver
> > [    1.048029] Console: switching to colour dummy device 80x25
> > [    1.048348] WARNING: kernel stack frame pointer at f50cdf98 in
> > swapper/2:0 has bad value   (null)
> > [    1.048349] unwind stack type:0 next_sp:  (null) mask:a graph_idx:0
> > [    1.048352] f50cdebc: 00000000f50cdec4 (0xf50cdec4)
> > [    1.048356] f50cdec0: 00000000c40489b7 (irq_exit+0x87/0xa0)
> > [    1.048357] f50cdec4: 00000000f50cded0 (0xf50cded0)
> > [    1.048361] f50cdec8: 00000000c402f6d3
> > (smp_apic_timer_interrupt+0x33/0x40)
> > [    1.048362] f50cdecc: 000000003e7c28eb (0x3e7c28eb)
> > [    1.048363] f50cded0: 00000000f50cded9 (0xf50cded9)
> > [    1.048366] f50cded4: 00000000c4b7ac8e
> > (apic_timer_interrupt+0x36/0x3c)
> > [    1.048367] f50cded8: 000000003e7c28eb (0x3e7c28eb)
> > [    1.048368] f50cdedc: 0000000000000001 (0x1)
> > [    1.048370] f50cdee0: 00000000f57c41c0 (0xf57c41c0)
> > [    1.048370] f50cdee4: 0000000000000000 ...
> > [    1.048371] f50cdee8: 0000000000000002 (0x2)
> > [    1.048372] f50cdeec: 00000000f50cdf30 (0xf50cdf30)
> > [    1.048373] f50cdef0: 000000003e7c28eb (0x3e7c28eb)
> > [    1.048376] f50cdef4: 00000000c506007b (cltrack_prog+0x7fb/0x1000)
> > [    1.048377] f50cdef8: 000000000000007b (0x7b)
> > [    1.048378] f50cdefc: 00000000000000d8 (0xd8)
> > [    1.048379] f50cdf00: 00000000c50600e0 (cltrack_prog+0x860/0x1000)
> > [    1.048380] f50cdf04: 00000000ffffff10 (0xffffff10)
> > [    1.048384] f50cdf08: 00000000c48d05b5
> > (cpuidle_enter_state+0xf5/0x220)
> > [    1.048384] f50cdf0c: 0000000000000060 (0x60)
> > [    1.048385] f50cdf10: 0000000000200286 (0x200286)
> > [    1.048386] f50cdf14: 0000000000000000 ...
> > [    1.048387] f50cdf18: 00000000ffb8fad0 (0xffb8fad0)
> > [    1.048388] f50cdf1c: 000000003e6bbd2b (0x3e6bbd2b)
> > [    1.048389] f50cdf20: 0000000000000000 ...
> > [    1.048391] f50cdf24: 00000000c506c680 (max_cstate+0x2c/0x2c)
> > [    1.048394] f50cdf28: 00000000c506c680 (max_cstate+0x2c/0x2c)
> > [    1.048394] f50cdf2c: 00000000f50a2040 (0xf50a2040)
> > [    1.048395] f50cdf30: 00000000f50cdf3c (0xf50cdf3c)
> > [    1.048397] f50cdf34: 00000000c48d06ff (cpuidle_enter+0xf/0x20)
> > [    1.048398] f50cdf38: 0000000000000000 ...
> > [    1.048399] f50cdf3c: 00000000f50cdf48 (0xf50cdf48)
> > [    1.048402] f50cdf40: 00000000c4083113 (call_cpuidle+0x23/0x40)
> > [    1.048403] f50cdf44: 00000000ffb8fad0 (0xffb8fad0)
> > [    1.048404] f50cdf48: 00000000f50cdf60 (0xf50cdf60)
> > [    1.048406] f50cdf4c: 00000000c4083364 (do_idle+0x174/0x1d0)
> > [    1.048407] f50cdf50: 00000000f50a2040 (0xf50a2040)
> > [    1.048408] f50cdf54: 000000008c0ed583 (0x8c0ed583)
> > [    1.048409] f50cdf58: 0000000000000087 (0x87)
> > [    1.048410] f50cdf5c: 000000009779245f (0x9779245f)
> > [    1.048411] f50cdf60: 00000000f50cdf78 (0xf50cdf78)
> > [    1.048413] f50cdf64: 00000000c408361d
> > (cpu_startup_entry+0x5d/0x60)
> > [    1.048414] f50cdf68: 0000000033ee4c15 (0x33ee4c15)
> > [    1.048415] f50cdf6c: 00000000c9d84398 (0xc9d84398)
> > [    1.048417] f50cdf70: 0000000002100800 (0x2100800)
> > [    1.048418] f50cdf74: 000000000fdc6fb2 (0xfdc6fb2)
> > [    1.048418] f50cdf78: 00000000f50cdf98 (0xf50cdf98)
> > [    1.048421] f50cdf7c: 00000000c402d216
> > (start_secondary+0x176/0x1c0)
> > [    1.048422] f50cdf80: 000000000fdc6fb2 (0xfdc6fb2)
> > [    1.048423] f50cdf84: 00000000d5425adc (0xd5425adc)
> > [    1.048424] f50cdf88: 0000000000000002 (0x2)
> > [    1.048425] f50cdf8c: 00000000ffffffff (0xffffffff)
> > [    1.048425] f50cdf90: 0000000000000000 ...
> > [    1.048426] f50cdf94: 00000000f50cdfac (0xf50cdfac)
> > [    1.048427] f50cdf98: 0000000000000000 ...
> > [    1.048429] f50cdf9c: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> > [    1.048429] f50cdfa0: 0000000000200002 (0x200002)
> > [    1.048430] f50cdfa4: 0000000000000000 ...
> > [    1.048432] f50cdfa8: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> > [    1.048432] f50cdfac: 0000000000000000 ...
> > [    1.048433] f50cdff4: 0000000000000100 (0x100)
> > [    1.048434] f50cdff8: 0000000000000200 (0x200)
> > [    1.048435] f50cdffc: 0000000000000000 ...
> > [    1.060368] [drm] Supports vblank timestamp caching Rev 2
> > (21.10.2013).
> > [    1.060373] [drm] Driver supports precise vblank timestamp query.
> > [    1.061668] i915 0000:00:02.0: vgaarb: changed VGA decodes:
> > olddecodes=io+mem,decodes=io+mem:owns=io+mem
> > [    1.146917] ACPI: Video Device [VID] (multi-head: yes  rom: no
> > post: no)
> > 
> > -- 
> > (english) http://www.livejournal.com/~pavelmachek
> > (cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html
> 
> 
> 

-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1586500

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2017-02-22 22:30 +0100
Message-ID<tdME2-7PD-19@gated-at.bofh.it>
In reply to#1586480
On Wed, Feb 22, 2017 at 10:05:48PM +0100, Pavel Machek wrote:
> Hi!
> 
> > > Thinkpad X220, in 32 bit mode... and I'm getting rather scary messages
> > > from kernel during boot:
> > > 
> > > Git blame says that message comes from commit
> > > 
> > > commit 24d86f59093b0bcb3756cdf47f2db10ff4e90dbb
> > > Author: Josh Poimboeuf <jpoimboe@redhat.com>
> > > Date:   Thu Oct 27 08:10:58 2016 -0500
> > > 
> > >     x86/unwind: Ensure stack grows down
> > > 
> > >     Add a sanity check to ensure the stack only grows down, and print
> > >     a
> > >         warning if the check fails.
> > > 
> > > Any ideas?
> > 
> > Hi Pavel,
> > 
> > I don't think I've seen this one.  Any chance this came after resuming
> > from a hibernation or suspend?
> 
> No, it was during the boot. Notice the timestamps...

Right, but doesn't waking from hibernation initially start with a
timestamp of zero?

The reason I asked is because of the following part of the stack dump:

> > > [    1.048429] f50cdf9c: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> > > [    1.048429] f50cdfa0: 0000000000200002 (0x200002)
> > > [    1.048430] f50cdfa4: 0000000000000000 ...
> > > [    1.048432] f50cdfa8: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> > > [    1.048432] f50cdfac: 0000000000000000 ...
> > > [    1.048433] f50cdff4: 0000000000000100 (0x100)
> > > [    1.048434] f50cdff8: 0000000000000200 (0x200)
> > > [    1.048435] f50cdffc: 0000000000000000 ...
> > > [    1.060368] [drm] Supports vblank timestamp caching Rev 2

Somehow, startup_32_smp() is on the stack twice.  The stack unwind led
to the startup_32_smp() frame at 0xf50cdf9c rather than the one at
0xf50cdfa8 (which is where it should normally be).  So the question is
how startup_32_smp() got executed the second time, with the wrong stack
offset.

-- 
Josh

[toc] | [prev] | [next] | [standalone]


#1586533

FromPavel Machek <pavel@ucw.cz>
Date2017-02-22 23:50 +0100
Message-ID<tdNTs-eQ-19@gated-at.bofh.it>
In reply to#1586500

[Multipart message — attachments visible in raw view] — view raw

Hi!

> > > > Thinkpad X220, in 32 bit mode... and I'm getting rather scary messages
> > > > from kernel during boot:
> > > > 
> > > > Git blame says that message comes from commit
> > > > 
> > > > commit 24d86f59093b0bcb3756cdf47f2db10ff4e90dbb
> > > > Author: Josh Poimboeuf <jpoimboe@redhat.com>
> > > > Date:   Thu Oct 27 08:10:58 2016 -0500
> > > > 
> > > >     x86/unwind: Ensure stack grows down
> > > > 
> > > >     Add a sanity check to ensure the stack only grows down, and print
> > > >     a
> > > >         warning if the check fails.
> > > > 
> > > > Any ideas?
> > > 
> > > I don't think I've seen this one.  Any chance this came after resuming
> > > from a hibernation or suspend?
> > 
> > No, it was during the boot. Notice the timestamps...
> 
> Right, but doesn't waking from hibernation initially start with a
> timestamp of zero?

Aha, ok, I guess so. Anyway... no hibernation was involved.

> The reason I asked is because of the following part of the stack
> dump:

> 
> > > > [    1.048429] f50cdf9c: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> > > > [    1.048429] f50cdfa0: 0000000000200002 (0x200002)
> > > > [    1.048430] f50cdfa4: 0000000000000000 ...
> > > > [    1.048432] f50cdfa8: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> > > > [    1.048432] f50cdfac: 0000000000000000 ...
> > > > [    1.048433] f50cdff4: 0000000000000100 (0x100)
> > > > [    1.048434] f50cdff8: 0000000000000200 (0x200)
> > > > [    1.048435] f50cdffc: 0000000000000000 ...
> > > > [    1.060368] [drm] Supports vblank timestamp caching Rev 2
> 
> Somehow, startup_32_smp() is on the stack twice.  The stack unwind led
> to the startup_32_smp() frame at 0xf50cdf9c rather than the one at
> 0xf50cdfa8 (which is where it should normally be).  So the question is
> how startup_32_smp() got executed the second time, with the wrong stack
> offset.

Not much idea... but this is stack dump, right? Just because some
value is on the stack does not mean it is a return address, no?

And .... startup_32_smp is kind of "interesting" function. Take a
look...

									Pavel

-- 
(english) http://www.livejournal.com/~pavelmachek
(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html

[toc] | [prev] | [next] | [standalone]


#1586540

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2017-02-23 00:10 +0100
Message-ID<tdOcN-DJ-5@gated-at.bofh.it>
In reply to#1586533
On Wed, Feb 22, 2017 at 11:47:55PM +0100, Pavel Machek wrote:
> Hi!
> 
> > > > > Thinkpad X220, in 32 bit mode... and I'm getting rather scary messages
> > > > > from kernel during boot:
> > > > > 
> > > > > Git blame says that message comes from commit
> > > > > 
> > > > > commit 24d86f59093b0bcb3756cdf47f2db10ff4e90dbb
> > > > > Author: Josh Poimboeuf <jpoimboe@redhat.com>
> > > > > Date:   Thu Oct 27 08:10:58 2016 -0500
> > > > > 
> > > > >     x86/unwind: Ensure stack grows down
> > > > > 
> > > > >     Add a sanity check to ensure the stack only grows down, and print
> > > > >     a
> > > > >         warning if the check fails.
> > > > > 
> > > > > Any ideas?
> > > > 
> > > > I don't think I've seen this one.  Any chance this came after resuming
> > > > from a hibernation or suspend?
> > > 
> > > No, it was during the boot. Notice the timestamps...
> > 
> > Right, but doesn't waking from hibernation initially start with a
> > timestamp of zero?
> 
> Aha, ok, I guess so. Anyway... no hibernation was involved.
> 
> > The reason I asked is because of the following part of the stack
> > dump:
> 
> > 
> > > > > [    1.048429] f50cdf9c: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> > > > > [    1.048429] f50cdfa0: 0000000000200002 (0x200002)
> > > > > [    1.048430] f50cdfa4: 0000000000000000 ...
> > > > > [    1.048432] f50cdfa8: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> > > > > [    1.048432] f50cdfac: 0000000000000000 ...
> > > > > [    1.048433] f50cdff4: 0000000000000100 (0x100)
> > > > > [    1.048434] f50cdff8: 0000000000000200 (0x200)
> > > > > [    1.048435] f50cdffc: 0000000000000000 ...
> > > > > [    1.060368] [drm] Supports vblank timestamp caching Rev 2
> > 
> > Somehow, startup_32_smp() is on the stack twice.  The stack unwind led
> > to the startup_32_smp() frame at 0xf50cdf9c rather than the one at
> > 0xf50cdfa8 (which is where it should normally be).  So the question is
> > how startup_32_smp() got executed the second time, with the wrong stack
> > offset.
> 
> Not much idea... but this is stack dump, right? Just because some
> value is on the stack does not mean it is a return address, no?

Right, but the one at 0xf50cdfa8 is where the startup_32_smp() is
*supposed* to be.  If the unwinder had unwinded to that one, it wouldn't
have complained.  So it looks to me like the CPU somehow booted twice:
the first time at the right stack address, and the second time it
somehow ended up with a different stack address.

> And .... startup_32_smp is kind of "interesting" function. Take a
> look...

Yes, it's used in bringing up the CPU.

-- 
Josh

[toc] | [prev] | [next] | [standalone]


#1586543

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2017-02-23 00:20 +0100
Message-ID<tdOmt-HE-9@gated-at.bofh.it>
In reply to#1586540
On Wed, Feb 22, 2017 at 04:56:14PM -0600, Josh Poimboeuf wrote:
> On Wed, Feb 22, 2017 at 11:47:55PM +0100, Pavel Machek wrote:
> > Hi!
> > 
> > > > > > Thinkpad X220, in 32 bit mode... and I'm getting rather scary messages
> > > > > > from kernel during boot:
> > > > > > 
> > > > > > Git blame says that message comes from commit
> > > > > > 
> > > > > > commit 24d86f59093b0bcb3756cdf47f2db10ff4e90dbb
> > > > > > Author: Josh Poimboeuf <jpoimboe@redhat.com>
> > > > > > Date:   Thu Oct 27 08:10:58 2016 -0500
> > > > > > 
> > > > > >     x86/unwind: Ensure stack grows down
> > > > > > 
> > > > > >     Add a sanity check to ensure the stack only grows down, and print
> > > > > >     a
> > > > > >         warning if the check fails.
> > > > > > 
> > > > > > Any ideas?
> > > > > 
> > > > > I don't think I've seen this one.  Any chance this came after resuming
> > > > > from a hibernation or suspend?
> > > > 
> > > > No, it was during the boot. Notice the timestamps...
> > > 
> > > Right, but doesn't waking from hibernation initially start with a
> > > timestamp of zero?
> > 
> > Aha, ok, I guess so. Anyway... no hibernation was involved.
> > 
> > > The reason I asked is because of the following part of the stack
> > > dump:
> > 
> > > 
> > > > > > [    1.048429] f50cdf9c: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> > > > > > [    1.048429] f50cdfa0: 0000000000200002 (0x200002)
> > > > > > [    1.048430] f50cdfa4: 0000000000000000 ...
> > > > > > [    1.048432] f50cdfa8: 00000000c4000237 (startup_32_smp+0x16b/0x16d)
> > > > > > [    1.048432] f50cdfac: 0000000000000000 ...
> > > > > > [    1.048433] f50cdff4: 0000000000000100 (0x100)
> > > > > > [    1.048434] f50cdff8: 0000000000000200 (0x200)
> > > > > > [    1.048435] f50cdffc: 0000000000000000 ...
> > > > > > [    1.060368] [drm] Supports vblank timestamp caching Rev 2
> > > 
> > > Somehow, startup_32_smp() is on the stack twice.  The stack unwind led
> > > to the startup_32_smp() frame at 0xf50cdf9c rather than the one at
> > > 0xf50cdfa8 (which is where it should normally be).  So the question is
> > > how startup_32_smp() got executed the second time, with the wrong stack
> > > offset.
> > 
> > Not much idea... but this is stack dump, right? Just because some
> > value is on the stack does not mean it is a return address, no?
> 
> Right, but the one at 0xf50cdfa8 is where the startup_32_smp() is
> *supposed* to be.  If the unwinder had unwinded to that one, it wouldn't
> have complained.  So it looks to me like the CPU somehow booted twice:
> the first time at the right stack address, and the second time it
> somehow ended up with a different stack address.
> 
> > And .... startup_32_smp is kind of "interesting" function. Take a
> > look...
> 
> Yes, it's used in bringing up the CPU.

Can you share your .config?  

-- 
Josh

[toc] | [prev] | [next] | [standalone]


#1588064

FromJosh Poimboeuf <jpoimboe@redhat.com>
Date2017-02-25 06:30 +0100
Message-ID<teD5D-33n-3@gated-at.bofh.it>
In reply to#1586543
On Thu, Feb 23, 2017 at 09:10:39PM +0100, Pavel Machek wrote:
> Hi!
> 
> 
> > > > > Somehow, startup_32_smp() is on the stack twice.  The stack unwind led
> > > > > to the startup_32_smp() frame at 0xf50cdf9c rather than the one at
> > > > > 0xf50cdfa8 (which is where it should normally be).  So the question is
> > > > > how startup_32_smp() got executed the second time, with the wrong stack
> > > > > offset.
> > > > 
> > > > Not much idea... but this is stack dump, right? Just because some
> > > > value is on the stack does not mean it is a return address, no?
> > > 
> > > Right, but the one at 0xf50cdfa8 is where the startup_32_smp() is
> > > *supposed* to be.  If the unwinder had unwinded to that one, it wouldn't
> > > have complained.  So it looks to me like the CPU somehow booted twice:
> > > the first time at the right stack address, and the second time it
> > > somehow ended up with a different stack address.
> > > 
> > > > And .... startup_32_smp is kind of "interesting" function. Take a
> > > > look...
> > > 
> > > Yes, it's used in bringing up the CPU.
> > 
> > Can you share your .config?  
> 
> Here you go...

What version of gcc are you using?

Can you post a disassembly of the first 10 instructions of
start_secondary()?

-- 
Josh

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web