Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1380837
| From | Bob Tracy <rct@gherkin.frus.com> |
|---|---|
| Newsgroups | linux.kernel, linux.debian.ports.alpha |
| Subject | [BUG] machine check Oops on Alpha |
| Date | 2016-04-17 23:40 +0200 |
| Message-ID | <rp2Aa-2oz-3@gated-at.bofh.it> (permalink) |
| Organization | linux.* mail to news gateway |
Cross-posted to 2 groups.
[Multipart message — attachments visible in raw view] - view raw
Apologies in advance for the "poor" quality of this bug report. No idea how to proceed, because the issue historically has been intermittent to non-existant for reasons unknown. Within 24 hours of booting my Alpha (PWS 433au), I'm pretty much guaranteed to see a "machine check" Oops which typically will occur during a period of high disk activity (for example, during an "apt-get update / upgrade". If I want a huge mess to clean up afterward, "git pull" on the kernel source tree will generally suffice as well :-(. As long as the "Oops" trace doesn't include evidence of filesystem write activity (calls to ext3/4 functions), the machine is perfectly stable afterward for as long as I care to let it run -- days, weeks, whatever -- no further Oopses will occur, regardless of how hard I flog the machine. A "bad" Oops will cause an immediate system lockup if any process attempts to access the region of disk that was active at the time the Oops occurred. While a "machine check" is normally indicative of an underlying hardware issue, the fact this is a one-time-per-boot issue has me thinking otherwise. I suspect a code path being traversed prior to the Oops that gets bypassed afterward. As previously mentioned, there have been months- long intervals in the past where the issue has either been masked or non- existent. Currently, the issue has persisted through several 4.X kernel release candidates and releases. Attached is an example of precisely what I'm talking about as far as a "good" Oops. It occurred within a day of the last reboot, and the machine has been running fine since. Been flogging the devil out of it, too: lots of updates (hundreds of megabytes), kernel builds, etc. While any and all help tracking this down will be appreciated, please know that kernel rebuilds (to turn on debugging or for whatever reason) are an overnight affair on this system. In other words, turnaround time on diagnostic iterations involving kernel modifications will be slow. --Bob
Back to linux.kernel | Previous | Next — Next in thread | Find similar | Unroll thread
[BUG] machine check Oops on Alpha Bob Tracy <rct@gherkin.frus.com> - 2016-04-17 23:40 +0200
Re: [BUG] machine check Oops on Alpha "Maciej W. Rozycki" <macro@linux-mips.org> - 2016-04-18 03:40 +0200
Re: [BUG] machine check Oops on Alpha Bob Tracy <rct@gherkin.frus.com> - 2016-04-18 06:00 +0200
Re: [BUG] machine check Oops on Alpha Bob Tracy <rct@gherkin.frus.com> - 2016-04-18 14:40 +0200
Re: [BUG] machine check Oops on Alpha "Maciej W. Rozycki" <macro@linux-mips.org> - 2016-04-18 15:50 +0200
Re: [BUG] machine check Oops on Alpha Bob Tracy <rct@gherkin.frus.com> - 2016-04-19 05:00 +0200
Re: [BUG] machine check Oops on Alpha Bob Tracy <rct@gherkin.frus.com> - 2016-04-20 02:00 +0200
Re: [BUG] machine check Oops on Alpha "Maciej W. Rozycki" <macro@linux-mips.org> - 2016-04-20 02:50 +0200
Re: [BUG] machine check Oops on Alpha Bob Tracy <rct@gherkin.frus.com> - 2016-04-20 06:00 +0200
csiph-web