Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1668368 > unrolled thread
| Started by | Laszlo Ersek <lersek@redhat.com> |
|---|---|
| First post | 2017-06-17 19:00 +0200 |
| Last post | 2017-06-18 22:00 +0200 |
| Articles | 4 — 3 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [Qemu-devel] [RFH] qemu-2.6 memory corruption with OVMF and linux-4.9 Laszlo Ersek <lersek@redhat.com> - 2017-06-17 19:00 +0200
Re: [Qemu-devel] [RFH] qemu-2.6 memory corruption with OVMF and linux-4.9 Philipp Hahn <hahn@univention.de> - 2017-06-18 20:30 +0200
Re: [Qemu-devel] [RFH] qemu-2.6 memory corruption with OVMF and linux-4.9 "Dr. David Alan Gilbert" <linux@treblig.org> - 2017-06-18 21:10 +0200
Re: [Qemu-devel] [RFH] qemu-2.6 memory corruption with OVMF and linux-4.9 Philipp Hahn <hahn@univention.de> - 2017-06-18 22:00 +0200
| From | Laszlo Ersek <lersek@redhat.com> |
|---|---|
| Date | 2017-06-17 19:00 +0200 |
| Subject | Re: [Qemu-devel] [RFH] qemu-2.6 memory corruption with OVMF and linux-4.9 |
| Message-ID | <tTpeN-1pE-1@gated-at.bofh.it> |
On 06/16/17 19:03, Philipp Hahn wrote:
> Comparing the corrupted (left) with the supposed (right) driver shows
> the following pattern:
>> /tmp/uefi.bin [+] 15038,1 Alles /tmp/uefi.ko [+] 15038,1 Alles
>> 003ac00: e801 0000 0000 0000 3c00 0000 1700 0000 ........<....... | 003ac00: e801 0000 0000 0000 5e8c 0000 1000 f1ff ........^.......
>> 003ac10: 785b 3e8a 0000 0000 3c00 0000 0700 0000 x[>.....<....... | 003ac10: 785b 3e8a 0000 0000 0000 0000 0000 0000 x[>.............
>> 003ac20: 778c 0000 1200 0200 3c00 0000 0700 0000 w.......<....... | 003ac20: 778c 0000 1200 0200 f018 0000 0000 0000 w...............
>> 003ac30: 1e00 0000 0000 0000 3c00 0000 1700 0000 ........<....... | 003ac30: 1e00 0000 0000 0000 8c8c 0000 1200 0200 ................
>> 003ac40: 7007 0000 0000 0000 3c00 0000 0700 0000 p.......<....... | 003ac40: 7007 0000 0000 0000 1400 0000 0000 0000 p...............
>> 003ac50: 9c8c 0000 1200 0200 3c00 0000 0700 0000 ........<....... | 003ac50: 9c8c 0000 1200 0200 0022 0000 0000 0000 ........."......
>> 003ac60: 4000 0000 0000 0000 3c00 0000 1700 0000 @.......<....... | 003ac60: 4000 0000 0000 0000 ac8c 0000 1000 f1ff @...............
Let me give you a different visual representation. First good, then bad.
(I also recommend using the "vbindiff" tool for such problems, it is
great for picking out patterns.)
** ** ** ** ** ** ** ** 8 9 ** ** ** 13 14 15
-- -- -- -- -- -- -- -- -- -- -- -- -- -- -- --
00000000 01 e8 00 00 00 00 00 00 8c 5e 00 00 00 10 ff f1
00000010 5b 78 8a 3e 00 00 00 00 00 00 00 00 00 00 00 00
00000020 8c 77 00 00 00 12 00 02 18 f0 00 00 00 00 00 00
00000030 00 1e 00 00 00 00 00 00 8c 8c 00 00 00 12 00 02
00000040 07 70 00 00 00 00 00 00 00 14 00 00 00 00 00 00
00000050 8c 9c 00 00 00 12 00 02 22 00 00 00 00 00 00 00
00000060 00 40 00 00 00 00 00 00 8c ac 00 00 00 10 ff f1
00000000 01 e8 00 00 00 00 00 00 00 3c 00 00 00 17 00 00
00000010 5b 78 8a 3e 00 00 00 00 00 3c 00 00 00 07 00 00
00000020 8c 77 00 00 00 12 00 02 00 3c 00 00 00 07 00 00
00000030 00 1e 00 00 00 00 00 00 00 3c 00 00 00 17 00 00
00000040 07 70 00 00 00 00 00 00 00 3c 00 00 00 07 00 00
00000050 8c 9c 00 00 00 12 00 02 00 3c 00 00 00 07 00 00
00000060 00 40 00 00 00 00 00 00 00 3c 00 00 00 17 00 00
-- -- -- -- -- -- -- -- -- -- -- -- -- -- -- --
** ** ** ** ** ** ** ** 8 9 ** ** ** 13 14 15
The columns that I marked with "**" are identical between "good" and
"bad". (These are columns 0-7, 10-12.)
Column 8 is overwritten by zeros (every 16th byte).
Column 9 is overwritten by 0x3c (every 16th byte).
Column 13 is super interesting. The most significant nibble in that
column is not disturbed. And, in the least significant nibble, the least
significant three bits are turned on. Basically, the corruption could be
described, for this column (i.e., every 16th byte), as
bad = good | 0x7
Column 14 is overwritten by zeros (every 16th byte).
Column 15 is overwritten by zeros (every 16th byte).
My take is that your host machine has faulty RAM. Please run memtest86+
or something similar.
Thanks
Laszlo
[toc] | [next] | [standalone]
| From | Philipp Hahn <hahn@univention.de> |
|---|---|
| Date | 2017-06-18 20:30 +0200 |
| Message-ID | <tTN7s-cP-9@gated-at.bofh.it> |
| In reply to | #1668368 |
Hello, Am 17.06.2017 um 18:51 schrieb Laszlo Ersek: > (I also recommend using the "vbindiff" tool for such problems, it is > great for picking out patterns.) > > ** ** ** ** ** ** ** ** 8 9 ** ** ** 13 14 15 > -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- > 00000000 01 e8 00 00 00 00 00 00 8c 5e 00 00 00 10 ff f1 > 00000010 5b 78 8a 3e 00 00 00 00 00 00 00 00 00 00 00 00 > 00000020 8c 77 00 00 00 12 00 02 18 f0 00 00 00 00 00 00 > 00000030 00 1e 00 00 00 00 00 00 8c 8c 00 00 00 12 00 02 > 00000040 07 70 00 00 00 00 00 00 00 14 00 00 00 00 00 00 > 00000050 8c 9c 00 00 00 12 00 02 22 00 00 00 00 00 00 00 > 00000060 00 40 00 00 00 00 00 00 8c ac 00 00 00 10 ff f1 > > 00000000 01 e8 00 00 00 00 00 00 00 3c 00 00 00 17 00 00 > 00000010 5b 78 8a 3e 00 00 00 00 00 3c 00 00 00 07 00 00 > 00000020 8c 77 00 00 00 12 00 02 00 3c 00 00 00 07 00 00 > 00000030 00 1e 00 00 00 00 00 00 00 3c 00 00 00 17 00 00 > 00000040 07 70 00 00 00 00 00 00 00 3c 00 00 00 07 00 00 > 00000050 8c 9c 00 00 00 12 00 02 00 3c 00 00 00 07 00 00 > 00000060 00 40 00 00 00 00 00 00 00 3c 00 00 00 17 00 00 > -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- > ** ** ** ** ** ** ** ** 8 9 ** ** ** 13 14 15 > > The columns that I marked with "**" are identical between "good" and > "bad". (These are columns 0-7, 10-12.) > > Column 8 is overwritten by zeros (every 16th byte). > > Column 9 is overwritten by 0x3c (every 16th byte). > > Column 13 is super interesting. The most significant nibble in that > column is not disturbed. And, in the least significant nibble, the least > significant three bits are turned on. Basically, the corruption could be > described, for this column (i.e., every 16th byte), as > > bad = good | 0x7 > > Column 14 is overwritten by zeros (every 16th byte). > > Column 15 is overwritten by zeros (every 16th byte). > > My take is that your host machine has faulty RAM. Please run memtest86+ > or something similar. I will do so, but for me very unlikely: - it never happens with BIOS, only with OVMF - for each test I start q new QEMU process, which should use a different memory region - it repeatedly hits e1000 or libata.ko After updating from OVMF to 0~20161202.7bbe0b3e-1 from (0~20160813.de74668f-2 it has not yet happened again. Anyway, thank you for your help. Philipp
[toc] | [prev] | [next] | [standalone]
| From | "Dr. David Alan Gilbert" <linux@treblig.org> |
|---|---|
| Date | 2017-06-18 21:10 +0200 |
| Message-ID | <tTNKa-G9-3@gated-at.bofh.it> |
| In reply to | #1668640 |
* Philipp Hahn (hahn@univention.de) wrote: > Hello, > > Am 17.06.2017 um 18:51 schrieb Laszlo Ersek: > > (I also recommend using the "vbindiff" tool for such problems, it is > > great for picking out patterns.) > > > > ** ** ** ** ** ** ** ** 8 9 ** ** ** 13 14 15 > > -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- > > 00000000 01 e8 00 00 00 00 00 00 8c 5e 00 00 00 10 ff f1 > > 00000010 5b 78 8a 3e 00 00 00 00 00 00 00 00 00 00 00 00 > > 00000020 8c 77 00 00 00 12 00 02 18 f0 00 00 00 00 00 00 > > 00000030 00 1e 00 00 00 00 00 00 8c 8c 00 00 00 12 00 02 > > 00000040 07 70 00 00 00 00 00 00 00 14 00 00 00 00 00 00 > > 00000050 8c 9c 00 00 00 12 00 02 22 00 00 00 00 00 00 00 > > 00000060 00 40 00 00 00 00 00 00 8c ac 00 00 00 10 ff f1 > > > > 00000000 01 e8 00 00 00 00 00 00 00 3c 00 00 00 17 00 00 > > 00000010 5b 78 8a 3e 00 00 00 00 00 3c 00 00 00 07 00 00 > > 00000020 8c 77 00 00 00 12 00 02 00 3c 00 00 00 07 00 00 > > 00000030 00 1e 00 00 00 00 00 00 00 3c 00 00 00 17 00 00 > > 00000040 07 70 00 00 00 00 00 00 00 3c 00 00 00 07 00 00 > > 00000050 8c 9c 00 00 00 12 00 02 00 3c 00 00 00 07 00 00 > > 00000060 00 40 00 00 00 00 00 00 00 3c 00 00 00 17 00 00 > > -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- > > ** ** ** ** ** ** ** ** 8 9 ** ** ** 13 14 15 > > > > The columns that I marked with "**" are identical between "good" and > > "bad". (These are columns 0-7, 10-12.) > > > > Column 8 is overwritten by zeros (every 16th byte). > > > > Column 9 is overwritten by 0x3c (every 16th byte). > > > > Column 13 is super interesting. The most significant nibble in that > > column is not disturbed. And, in the least significant nibble, the least > > significant three bits are turned on. Basically, the corruption could be > > described, for this column (i.e., every 16th byte), as > > > > bad = good | 0x7 > > > > Column 14 is overwritten by zeros (every 16th byte). > > > > Column 15 is overwritten by zeros (every 16th byte). > > > > My take is that your host machine has faulty RAM. Please run memtest86+ > > or something similar. > > I will do so, but for me very unlikely: > - it never happens with BIOS, only with OVMF > - for each test I start q new QEMU process, which should use a different > memory region > - it repeatedly hits e1000 or libata.ko > > After updating from OVMF to 0~20161202.7bbe0b3e-1 from > (0~20160813.de74668f-2 it has not yet happened again. > > Anyway, thank you for your help. What host CPU are you using? Dave > > Philipp -- -----Open up your eyes, open up your mind, open up your code ------- / Dr. David Alan Gilbert | Running GNU/Linux | Happy \ \ dave @ treblig.org | | In Hex / \ _________________________|_____ http://www.treblig.org |_______/
[toc] | [prev] | [next] | [standalone]
| From | Philipp Hahn <hahn@univention.de> |
|---|---|
| Date | 2017-06-18 22:00 +0200 |
| Message-ID | <tTOwy-ZC-15@gated-at.bofh.it> |
| In reply to | #1668647 |
Am 18.06.2017 um 20:27 schrieb Dr. David Alan Gilbert: > * Philipp Hahn (hahn@univention.de) wrote: >> Hello, >> >> Am 17.06.2017 um 18:51 schrieb Laszlo Ersek: >>> (I also recommend using the "vbindiff" tool for such problems, it is >>> great for picking out patterns.) >>> >>> ** ** ** ** ** ** ** ** 8 9 ** ** ** 13 14 15 >>> -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- >>> 00000000 01 e8 00 00 00 00 00 00 8c 5e 00 00 00 10 ff f1 >>> 00000010 5b 78 8a 3e 00 00 00 00 00 00 00 00 00 00 00 00 >>> 00000020 8c 77 00 00 00 12 00 02 18 f0 00 00 00 00 00 00 >>> 00000030 00 1e 00 00 00 00 00 00 8c 8c 00 00 00 12 00 02 >>> 00000040 07 70 00 00 00 00 00 00 00 14 00 00 00 00 00 00 >>> 00000050 8c 9c 00 00 00 12 00 02 22 00 00 00 00 00 00 00 >>> 00000060 00 40 00 00 00 00 00 00 8c ac 00 00 00 10 ff f1 >>> >>> 00000000 01 e8 00 00 00 00 00 00 00 3c 00 00 00 17 00 00 >>> 00000010 5b 78 8a 3e 00 00 00 00 00 3c 00 00 00 07 00 00 >>> 00000020 8c 77 00 00 00 12 00 02 00 3c 00 00 00 07 00 00 >>> 00000030 00 1e 00 00 00 00 00 00 00 3c 00 00 00 17 00 00 >>> 00000040 07 70 00 00 00 00 00 00 00 3c 00 00 00 07 00 00 >>> 00000050 8c 9c 00 00 00 12 00 02 00 3c 00 00 00 07 00 00 >>> 00000060 00 40 00 00 00 00 00 00 00 3c 00 00 00 17 00 00 >>> -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- -- >>> ** ** ** ** ** ** ** ** 8 9 ** ** ** 13 14 15 >>> >>> The columns that I marked with "**" are identical between "good" and >>> "bad". (These are columns 0-7, 10-12.) >>> >>> Column 8 is overwritten by zeros (every 16th byte). >>> >>> Column 9 is overwritten by 0x3c (every 16th byte). >>> >>> Column 13 is super interesting. The most significant nibble in that >>> column is not disturbed. And, in the least significant nibble, the least >>> significant three bits are turned on. Basically, the corruption could be >>> described, for this column (i.e., every 16th byte), as >>> >>> bad = good | 0x7 >>> >>> Column 14 is overwritten by zeros (every 16th byte). >>> >>> Column 15 is overwritten by zeros (every 16th byte). >>> >>> My take is that your host machine has faulty RAM. Please run memtest86+ >>> or something similar. >> >> I will do so, but for me very unlikely: >> - it never happens with BIOS, only with OVMF >> - for each test I start q new QEMU process, which should use a different >> memory region >> - it repeatedly hits e1000 or libata.ko >> >> After updating from OVMF to 0~20161202.7bbe0b3e-1 from >> (0~20160813.de74668f-2 it has not yet happened again. >> >> Anyway, thank you for your help. > > What host CPU are you using? Everything is amd64: > processor : 3 > vendor_id : GenuineIntel > cpu family : 6 > model : 58 > model name : Intel(R) Core(TM) i5-3337U CPU @ 1.80GHz > stepping : 9 > microcode : 0x19 > cpu MHz : 2591.015 > cache size : 3072 KB > physical id : 0 > siblings : 4 > core id : 1 > cpu cores : 2 > apicid : 3 > initial apicid : 3 > fpu : yes > fpu_exception : yes > cpuid level : 13 > wp : yes > flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx est tm2 ssse3 cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase smep erms xsaveopt dtherm ida arat pln pts > bugs : > bogomips : 3592.75 > clflush size : 64 > cache_alignment : 64 > address sizes : 36 bits physical, 48 bits virtual > power management: Philipp
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web