Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > comp.os.linux.hardware > #2913 > unrolled thread

Kernel panics

Started bybuck <buck@private.mil>
First post2015-09-17 19:36 +0000
Last post2015-09-29 20:19 +0000
Articles 8 — 4 participants

Back to article view | Back to comp.os.linux.hardware


Contents

  Kernel panics buck <buck@private.mil> - 2015-09-17 19:36 +0000
    Re: Kernel panics Poutnik <poutnik4nntp@gmail.com> - 2015-09-17 23:19 +0200
      Re: Kernel panics buck <buck@private.mil> - 2015-09-18 00:47 +0000
        Re: Kernel panics Poutnik <poutnik4nntp@gmail.com> - 2015-09-18 07:50 +0200
          Re: Kernel panics Poutnik <poutnik4nntp@gmail.com> - 2015-09-18 08:37 +0200
    Re: Kernel panics Frank Miles <fpm@u.washington.edu> - 2015-09-29 04:26 +0000
      Re: Kernel panics Bobbie Sellers <bliss-sf4ever@dslextreme.com> - 2015-09-28 21:41 -0700
      Re: Kernel panics buck <buck@private.mil> - 2015-09-29 20:19 +0000

#2913 — Kernel panics

Frombuck <buck@private.mil>
Date2015-09-17 19:36 +0000
SubjectKernel panics
Message-ID<mtf4nj617m4@news7.newsguy.com>
NOTE:  I am unable to convince Xnews to set the FollowUp-To to 
alt.os.linux.slackware. Please reply there.

Problem:
My internet-facing computer is experiencing kernel panics on an 
increasingly frequent basis.

Scenario:
The box was set up in July 2001.  AMD Duron CPU, 4 IDE drives in a 
software RAID-5 array, 2 NICs (LAN and WAN) and a Diamond Stealth 
video card.  When the kernel panics began to become more frequent, I 
added append = "panic=10" to lilo.conf.

General Info:
My understanding of "In interrupt handler - not syncing" is that some 
hardware is at fault, but that seems not to be the case here.  I have 
2 other boxes with Intel P4 CPU, IDE interfaces, PS2 kbd & mouse, a 
serial and a parallel port.  I changed the CPU in .config from 
CONFIG_MK7 to CONFIG_MPENTIUM4 (for the switch from AMD to Intel) and 
built bzImage and modules.  Then I removed the drives from the AMD 
computer and put them into one of the 2 Intel boxes.  Kernel panics 
were just as frequent.  On the off chance that that Intel box was also 
bad, I moved the RAID array to the other Intel box and kernel panics 
still continue.

Each of the 2 Intel boxes contains one of the NICs that were in the 
AMD box: the SMC uses the tulip driver and the 3com uses the 3c59x 
(3com "vortex") driver.  That NIC is connected to the cable modem.  
There is no other commonality other than the connections: power cord, 
keyboard, mouse, parallel printer, and serial port (battery backup but 
not monitored).

There is nothing in the logs to indicate any problem.

Request for a suggested course of action:
Should I replace one drive at a time in the array?  If so, why?
Should I reinstall the packages?  If so, why?
if this were your computer, how would you approach a solution?

Comments are also welome.  Please help if you can.
-- 
buck

[toc] | [next] | [standalone]


#2914

FromPoutnik <poutnik4nntp@gmail.com>
Date2015-09-17 23:19 +0200
Message-ID<mtfalb$kbb$1@dont-email.me>
In reply to#2913
Dne 17/09/2015 v 21:36 buck napsal(a):
> NOTE:  I am unable to convince Xnews to set the FollowUp-To to 
> alt.os.linux.slackware. Please reply there.
> 
> Problem:
> My internet-facing computer is experiencing kernel panics on an 
> increasingly frequent basis.
> 

Is possible an occasional RAM hardware faults is the reason ?
What about long term RAM checking ?


-- 
Poutnik ( the Czech word for a wanderer )

Knowledge makes great men humble, but small men arrogant.

[toc] | [prev] | [next] | [standalone]


#2916

Frombuck <buck@private.mil>
Date2015-09-18 00:47 +0000
Message-ID<mtfmvj01ot8@news7.newsguy.com>
In reply to#2914
Poutnik <poutnik4nntp@gmail.com> wrote in news:mtfalb$kbb$1@dont-
email.me:

> Dne 17/09/2015 v 21:36 buck napsal(a):
>> NOTE:  I am unable to convince Xnews to set the FollowUp-To to 
>> alt.os.linux.slackware. Please reply there.
>> 
>> Problem:
>> My internet-facing computer is experiencing kernel panics on an 
>> increasingly frequent basis.
>> 
> 
> Is possible an occasional RAM hardware faults is the reason ?
> What about long term RAM checking ?

On 3 different machines?  I don't think all 3 could have bad RAM?
-- 
buck

[toc] | [prev] | [next] | [standalone]


#2919

FromPoutnik <poutnik4nntp@gmail.com>
Date2015-09-18 07:50 +0200
Message-ID<mtg8jh$8q2$2@dont-email.me>
In reply to#2916
Dne 18/09/2015 v 02:47 buck napsal(a):
> Poutnik <poutnik4nntp@gmail.com> wrote in news:mtfalb$kbb$1@dont-
> email.me:
> 
>> Dne 17/09/2015 v 21:36 buck napsal(a):
>>> NOTE:  I am unable to convince Xnews to set the FollowUp-To to 
>>> alt.os.linux.slackware. Please reply there.
>>>
>>> Problem:
>>> My internet-facing computer is experiencing kernel panics on an 
>>> increasingly frequent basis.
>>>
>>
>> Is possible an occasional RAM hardware faults is the reason ?
>> What about long term RAM checking ?
> 
> On 3 different machines?  I don't think all 3 could have bad RAM?
> 
Opps, of course not, I have misread the info as on 1 on 3.

-- 
Poutnik ( the Czech word for a wanderer )

Knowledge makes great men humble, but small men arrogant.

[toc] | [prev] | [next] | [standalone]


#2920

FromPoutnik <poutnik4nntp@gmail.com>
Date2015-09-18 08:37 +0200
Message-ID<mtgbc9$h1l$1@dont-email.me>
In reply to#2919
Dne 18/09/2015 v 07:50 Poutnik napsal(a):
> Dne 18/09/2015 v 02:47 buck napsal(a):
>> Poutnik <poutnik4nntp@gmail.com> wrote in news:mtfalb$kbb$1@dont-
>> email.me:
>>
>>> Dne 17/09/2015 v 21:36 buck napsal(a):
>>>> NOTE:  I am unable to convince Xnews to set the FollowUp-To to 
>>>> alt.os.linux.slackware. Please reply there.
>>>>
>>>> Problem:
>>>> My internet-facing computer is experiencing kernel panics on an 
>>>> increasingly frequent basis.
>>>>
>>>
>>> Is possible an occasional RAM hardware faults is the reason ?
>>> What about long term RAM checking ?
>>
>> On 3 different machines?  I don't think all 3 could have bad RAM?
>>
> Opps, of course not, I have misread the info as on 1 on 3.
> 
OTOH, the progressiveness of error frequency gives the hint
there may be rather HW than SW reasons.

Do the machines have similar HW config ?
Some problematic component manufacturer serie ?

If possible, replacing of systems, perhaps just for testing,
may distinguish SW and HW reasons.

-- 
Poutnik ( the Czech word for a wanderer )

Knowledge makes great men humble, but small men arrogant.

[toc] | [prev] | [next] | [standalone]


#2922

FromFrank Miles <fpm@u.washington.edu>
Date2015-09-29 04:26 +0000
Message-ID<mud3uc$277$1@dont-email.me>
In reply to#2913
On Thu, 17 Sep 2015 19:36:19 +0000, buck wrote:

> NOTE:  I am unable to convince Xnews to set the FollowUp-To to
> alt.os.linux.slackware. Please reply there.
> 
> Problem:
> My internet-facing computer is experiencing kernel panics on an
> increasingly frequent basis.
> 
> Scenario:
> The box was set up in July 2001.  AMD Duron CPU, 4 IDE drives in a
> software RAID-5 array, 2 NICs (LAN and WAN) and a Diamond Stealth video
> card.  When the kernel panics began to become more frequent, I added
> append = "panic=10" to lilo.conf.
> 
> General Info:
> My understanding of "In interrupt handler - not syncing" is that some
> hardware is at fault, but that seems not to be the case here.  I have 2
> other boxes with Intel P4 CPU, IDE interfaces, PS2 kbd & mouse, a serial
> and a parallel port.  I changed the CPU in .config from CONFIG_MK7 to
> CONFIG_MPENTIUM4 (for the switch from AMD to Intel) and built bzImage
> and modules.  Then I removed the drives from the AMD computer and put
> them into one of the 2 Intel boxes.  Kernel panics were just as
> frequent.  On the off chance that that Intel box was also bad, I moved
> the RAID array to the other Intel box and kernel panics still continue.
> 
> Each of the 2 Intel boxes contains one of the NICs that were in the AMD
> box: the SMC uses the tulip driver and the 3com uses the 3c59x (3com
> "vortex") driver.  That NIC is connected to the cable modem. There is no
> other commonality other than the connections: power cord, keyboard,
> mouse, parallel printer, and serial port (battery backup but not
> monitored).
> 
> There is nothing in the logs to indicate any problem.
> 
> Request for a suggested course of action:
> Should I replace one drive at a time in the array?  If so, why? Should I
> reinstall the packages?  If so, why?
> if this were your computer, how would you approach a solution?
> 
> Comments are also welome.  Please help if you can.

What constitutes "increasingly frequent"?

I've had keyboard that killed a system (wasn't intermittent in my case).
So I'd check the remaining pieces, especially if you've got spares handy.
Power cord included.

Do all of the machines use the same distribution/version OS?  Preferably
something newer than 14 years old?

Intermittents are such a pain!

[toc] | [prev] | [next] | [standalone]


#2923

FromBobbie Sellers <bliss-sf4ever@dslextreme.com>
Date2015-09-28 21:41 -0700
Message-ID<mud4m9$648$1@dont-email.me>
In reply to#2922
On 09/28/2015 09:26 PM, Frank Miles wrote:
> On Thu, 17 Sep 2015 19:36:19 +0000, buck wrote:
>
>> NOTE:  I am unable to convince Xnews to set the FollowUp-To to
>> alt.os.linux.slackware. Please reply there.
>>
>> Problem:
>> My internet-facing computer is experiencing kernel panics on an
>> increasingly frequent basis.
>>
>> Scenario:
>> The box was set up in July 2001.  AMD Duron CPU, 4 IDE drives in a
>> software RAID-5 array, 2 NICs (LAN and WAN) and a Diamond Stealth video
>> card.  When the kernel panics began to become more frequent, I added
>> append = "panic=10" to lilo.conf.
>>
>> General Info:
>> My understanding of "In interrupt handler - not syncing" is that some
>> hardware is at fault, but that seems not to be the case here.  I have 2
>> other boxes with Intel P4 CPU, IDE interfaces, PS2 kbd & mouse, a serial
>> and a parallel port.  I changed the CPU in .config from CONFIG_MK7 to
>> CONFIG_MPENTIUM4 (for the switch from AMD to Intel) and built bzImage
>> and modules.  Then I removed the drives from the AMD computer and put
>> them into one of the 2 Intel boxes.  Kernel panics were just as
>> frequent.  On the off chance that that Intel box was also bad, I moved
>> the RAID array to the other Intel box and kernel panics still continue.
>>
>> Each of the 2 Intel boxes contains one of the NICs that were in the AMD
>> box: the SMC uses the tulip driver and the 3com uses the 3c59x (3com
>> "vortex") driver.  That NIC is connected to the cable modem. There is no
>> other commonality other than the connections: power cord, keyboard,
>> mouse, parallel printer, and serial port (battery backup but not
>> monitored).
>>
>> There is nothing in the logs to indicate any problem.
>>
>> Request for a suggested course of action:
>> Should I replace one drive at a time in the array?  If so, why? Should I
>> reinstall the packages?  If so, why?
>> if this were your computer, how would you approach a solution?
>>
>> Comments are also welome.  Please help if you can.
>
> What constitutes "increasingly frequent"?
>
> I've had keyboard that killed a system (wasn't intermittent in my case).
> So I'd check the remaining pieces, especially if you've got spares handy.
> Power cord included.
>
> Do all of the machines use the same distribution/version OS?  Preferably
> something newer than 14 years old?
>
> Intermittents are such a pain!
>


	So you moved the drives to another box and had the same symptoms?
	I would do a test of the hard drives.
	There are plenty of rescue OSes out there which should
have this sort of test.

	14 year old?  These toys do not last forever, not any part
but the cases (assuming desktops).

	bliss

[toc] | [prev] | [next] | [standalone]


#2924

Frombuck <buck@private.mil>
Date2015-09-29 20:19 +0000
Message-ID<muerpc0eg@news4.newsguy.com>
In reply to#2922
Frank Miles <fpm@u.washington.edu> wrote in
news:mud3uc$277$1@dont-email.me: 

> 
> What constitutes "increasingly frequent"?
> 
> I've had keyboard that killed a system (wasn't intermittent in my
> case). So I'd check the remaining pieces, especially if you've got
> spares handy. Power cord included.
> 
> Do all of the machines use the same distribution/version OS? 
> Preferably something newer than 14 years old?
> 
> Intermittents are such a pain!

Thanks for the input.

Incredibly, what fixed this was the removal of a few linefeeds from 
panic.c.

The AMD kernel would not boot, so I changed the CPU to Pentium 4.  
That build failed to include iptables' connlimit and rebooted 3 times 
in 24 hours when the RAID array was mounted in the first of 2 Intel 
boxes.

Ran runme extra in patch-o-matic-ng-20040302 to replace connlimit and 
that also rebooted several times (on both Intel boxes).

Fearing that the panic text would scroll important information, I 
edited panic.c and removed \n from some printk lines.  VOILA stable - 
for unfathomable reasons.

Because the clock on the second Intel box drifted, the array was put 
into the first Intel box and has been stable since Sep 16 15:12.
-- 
buck

[toc] | [prev] | [standalone]


Back to top | Article view | comp.os.linux.hardware


csiph-web