Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #268392 > unrolled thread

Filsystemkorruption i ext4?

Started byJesper Dybdal <jd-debian-user@dybdal.dk>
First post2024-03-19 15:50 +0100
Last post2024-03-28 16:40 +0100
Articles 11 — 4 participants

Back to article view | Back to linux.debian.user


Contents

  Filsystemkorruption i ext4? Jesper Dybdal <jd-debian-user@dybdal.dk> - 2024-03-19 15:50 +0100
    Re: Filsystemkorruption i ext4? Franco Martelli <martellif67@gmail.com> - 2024-03-19 21:50 +0100
      Re: Filsystemkorruption i ext4? Jesper Dybdal <jd-debian-user@dybdal.dk> - 2024-03-20 09:20 +0100
        Re: Filsystemkorruption i ext4? Franco Martelli <martellif67@gmail.com> - 2024-03-20 16:10 +0100
      Re: Filsystemkorruption i ext4? Jesper Dybdal <jd-debian-user@dybdal.dk> - 2024-03-20 14:30 +0100
        Re: Filsystemkorruption i ext4? Nicholas Geovanis <nickgeovanis@gmail.com> - 2024-03-20 23:00 +0100
          Re: Filsystemkorruption i ext4? Jesper Dybdal <jd-debian-user@dybdal.dk> - 2024-03-28 14:40 +0100
          Re: Filsystemkorruption i ext4? Jesper Dybdal <jd-debian-user@dybdal.dk> - 2024-03-28 15:00 +0100
            Re: Filsystemkorruption i ext4? Hans <hans.ullrich@loop.de> - 2024-03-28 15:10 +0100
              Re: Filsystemkorruption i ext4? Jesper Dybdal <jd-debian-user@dybdal.dk> - 2024-03-28 16:10 +0100
                Re: Filsystemkorruption i ext4? Hans <hans.ullrich@loop.de> - 2024-03-28 16:40 +0100

#268392 — Filsystemkorruption i ext4?

FromJesper Dybdal <jd-debian-user@dybdal.dk>
Date2024-03-19 15:50 +0100
SubjectFilsystemkorruption i ext4?
Message-ID<IjIWS-b9R-27@gated-at.bofh.it>
I have just discovered that my Debian Bullseye server thinks that a file 
has allocated 2251799813684984 blocks (!):

root@nuser:/etc/postfix# stat master.cf.bad-size
   File: master.cf.bad-size
   Size: 10782           Blocks: 2251799813684984 IO Block: 4096 regular 
file
Device: 900h/2304d      Inode: 10748971    Links: 1
Access: (0644/-rw-r--r--)  Uid: (    0/    root)   Gid: (    0/ root)
Access: 2024-03-19 15:23:01.002627532 +0100
Modify: 2024-03-11 16:16:22.152186851 +0100
Change: 2024-03-19 13:37:44.647279278 +0100
  Birth: 2023-06-13 01:07:54.542855184 +0200

The file can be read and has the correct contents.  (I've renamed it and 
replaced master.cf by a fresh backup copy, so nothing will be missed if 
fsck deletes the bad file.)

It seems to me that there is a file system corruption and/or a disk 
error and/or a RAM error.
Had anyone else seen something like this?

(I haven't rebooted or run fsck yet, because I want to put a monitor on 
the machine first and have time to solve problemr - and I haven't got 
that time right now.)

My plan is to boot a rescue disk and mount that partition read-only. Then:
* If the file looks ok after reboot, then I'll strongly suspect the RAM 
- and run memtest.
* Otherwise, I'll have to run fsck and see what happens.

kernel version:
root@nuser:~# uname -a
Linux nuser 5.10.0-28-amd64 #1 SMP Debian 5.10.209-2 (2024-01-31) x86_64 
GNU/Linux

The partition in question is a RAID 1 controlled by md.

Thanks,
Jesper


-- 
Jesper Dybdal
https://www.dybdal.dk

[toc] | [next] | [standalone]


#268407

FromFranco Martelli <martellif67@gmail.com>
Date2024-03-19 21:50 +0100
Message-ID<IjOzf-fuL-9@gated-at.bofh.it>
In reply to#268392
On 19/03/24 at 15:43, Jesper Dybdal wrote:

> 
> My plan is to boot a rescue disk and mount that partition read-only. Then:
> * If the file looks ok after reboot, then I'll strongly suspect the RAM 
> - and run memtest.
> * Otherwise, I'll have to run fsck and see what happens.
> 
> kernel version:
> root@nuser:~# uname -a
> Linux nuser 5.10.0-28-amd64 #1 SMP Debian 5.10.209-2 (2024-01-31) x86_64 
> GNU/Linux
> 
> The partition in question is a RAID 1 controlled by md.

Another check you can perform it is on the RAID array, by default it 
runs on the first Sunday of each month at 00:57. You should have this 
file /etc/cron.d/mdadm that takes care to run this check monthly.

Before you reboot, does it look OK /proc/mdstat ?

-- 
Franco Martelli

[toc] | [prev] | [next] | [standalone]


#268416

FromJesper Dybdal <jd-debian-user@dybdal.dk>
Date2024-03-20 09:20 +0100
Message-ID<IjZkZ-mBf-3@gated-at.bofh.it>
In reply to#268407
[Sorry for the accidental Danish-language subject line :-( ]

On 2024-03-19 21:47, Franco Martelli wrote:
> On 19/03/24 at 15:43, Jesper Dybdal wrote:
>
>>
>> My plan is to boot a rescue disk and mount that partition read-only. 
>> Then:
>> * If the file looks ok after reboot, then I'll strongly suspect the 
>> RAM - and run memtest.
>> * Otherwise, I'll have to run fsck and see what happens.
>>
>> kernel version:
>> root@nuser:~# uname -a
>> Linux nuser 5.10.0-28-amd64 #1 SMP Debian 5.10.209-2 (2024-01-31) 
>> x86_64 GNU/Linux
>>
>> The partition in question is a RAID 1 controlled by md.
>
> Another check you can perform it is on the RAID array, by default it 
> runs on the first Sunday of each month at 00:57. You should have this 
> file /etc/cron.d/mdadm that takes care to run this check monthly.
Good idea!  That should of course be done first.  It's running now.

> Before you reboot, does it look OK /proc/mdstat ?
Yes, it seems ok.

Thanks,
Jesper

-- 
Jesper Dybdal
https://www.dybdal.dk

[toc] | [prev] | [next] | [standalone]


#268431

FromFranco Martelli <martellif67@gmail.com>
Date2024-03-20 16:10 +0100
Message-ID<Ik5JM-qsY-9@gated-at.bofh.it>
In reply to#268416
On 20/03/24 at 09:15, Jesper Dybdal wrote:
> [Sorry for the accidental Danish-language subject line :-( ]
> 
> On 2024-03-19 21:47, Franco Martelli wrote:
>> On 19/03/24 at 15:43, Jesper Dybdal wrote:
>>
>>>
>>> My plan is to boot a rescue disk and mount that partition read-only. 
>>> Then:
>>> * If the file looks ok after reboot, then I'll strongly suspect the 
>>> RAM - and run memtest.
>>> * Otherwise, I'll have to run fsck and see what happens.
>>>
>>> kernel version:
>>> root@nuser:~# uname -a
>>> Linux nuser 5.10.0-28-amd64 #1 SMP Debian 5.10.209-2 (2024-01-31) 
>>> x86_64 GNU/Linux
>>>
>>> The partition in question is a RAID 1 controlled by md.
>>
>> Another check you can perform it is on the RAID array, by default it 
>> runs on the first Sunday of each month at 00:57. You should have this 
>> file /etc/cron.d/mdadm that takes care to run this check monthly.
> Good idea!  That should of course be done first.  It's running now.
> 
>> Before you reboot, does it look OK /proc/mdstat ?
> Yes, it seems ok.

I would suggest you to mount the filesystem yes read-only but also with 
the noload option ( … -o ro,noload … ) see "man mount" for a brief 
explanation.

Cheers

-- 
Franco Martelli

[toc] | [prev] | [next] | [standalone]


#268427

FromJesper Dybdal <jd-debian-user@dybdal.dk>
Date2024-03-20 14:30 +0100
Message-ID<Ik4aZ-psP-3@gated-at.bofh.it>
In reply to#268407
I have now done the following:
* Checked the RAID array - no problems found.
* Run fsck.  It found three cases of the block count being incorrect.  I 
don't know which the other two affected files are.
* Run one pass of memtest86+.  Nothing found.

So it seems not to be a problem with the disks.
A bug in ext4?  Well, ext4 has always done its job for me wihtout problems.
A RAM error that memtest86+ did not find?  Possible.  Once upon a time, 
when you bought an ordinary pc, its RAM had ECC as a matter of course; 
unfortunately, that is not the case nowadays.

I think I'll let memtest86+ run overnight one of the coming nights.

Unless it is simply a RAM error, then it is a bit scary...

Regards,
Jesper

On 2024-03-19 21:47, Franco Martelli wrote:
> On 19/03/24 at 15:43, Jesper Dybdal wrote:
>
>>
>> My plan is to boot a rescue disk and mount that partition read-only. 
>> Then:
>> * If the file looks ok after reboot, then I'll strongly suspect the 
>> RAM - and run memtest.
>> * Otherwise, I'll have to run fsck and see what happens.
>>
>> kernel version:
>> root@nuser:~# uname -a
>> Linux nuser 5.10.0-28-amd64 #1 SMP Debian 5.10.209-2 (2024-01-31) 
>> x86_64 GNU/Linux
>>
>> The partition in question is a RAID 1 controlled by md.
>
> Another check you can perform it is on the RAID array, by default it 
> runs on the first Sunday of each month at 00:57. You should have this 
> file /etc/cron.d/mdadm that takes care to run this check monthly.
>
> Before you reboot, does it look OK /proc/mdstat ?
>

-- 
Jesper Dybdal
https://www.dybdal.dk

[toc] | [prev] | [next] | [standalone]


#268465

FromNicholas Geovanis <nickgeovanis@gmail.com>
Date2024-03-20 23:00 +0100
Message-ID<Ikc8y-umf-7@gated-at.bofh.it>
In reply to#268427

[Multipart message — attachments visible in raw view] — view raw

On Wed, Mar 20, 2024, 11:28 AM Jesper Dybdal <jd-debian-user@dybdal.dk>
wrote:

> I have now done the following:
> * Checked the RAID array - no problems found.
> * Run fsck.  It found three cases of the block count being incorrect.  I
> don't know which the other two affected files are.
> * Run one pass of memtest86+.  Nothing found.
>
> So it seems not to be a problem with the disks.
> A bug in ext4?  Well, ext4 has always done its job for me wihtout problems.
> A RAM error that memtest86+ did not find?  Possible.  Once upon a time,
> when you bought an ordinary pc, its RAM had ECC as a matter of course;
> unfortunately, that is not the case nowadays.
>
> I think I'll let memtest86+ run overnight one of the coming nights.
>
> Unless it is simply a RAM error, then it is a bit scary...
>

I have seen that a couple times, unlikely but possible. Maybe review your
RAM configuration too, ensure that the sticks are on the same supported
refresh rate and distributed across the slots in an approved way.

Regards,
> Jesper
>
> On 2024-03-19 21:47, Franco Martelli wrote:
> > On 19/03/24 at 15:43, Jesper Dybdal wrote:
> >
> >>
> >> My plan is to boot a rescue disk and mount that partition read-only.
> >> Then:
> >> * If the file looks ok after reboot, then I'll strongly suspect the
> >> RAM - and run memtest.
> >> * Otherwise, I'll have to run fsck and see what happens.
> >>
> >> kernel version:
> >> root@nuser:~# uname -a
> >> Linux nuser 5.10.0-28-amd64 #1 SMP Debian 5.10.209-2 (2024-01-31)
> >> x86_64 GNU/Linux
> >>
> >> The partition in question is a RAID 1 controlled by md.
> >
> > Another check you can perform it is on the RAID array, by default it
> > runs on the first Sunday of each month at 00:57. You should have this
> > file /etc/cron.d/mdadm that takes care to run this check monthly.
> >
> > Before you reboot, does it look OK /proc/mdstat ?
> >
>
> --
> Jesper Dybdal
> https://www.dybdal.dk
>
>
>
>

[toc] | [prev] | [next] | [standalone]


#268577

FromJesper Dybdal <jd-debian-user@dybdal.dk>
Date2024-03-28 14:40 +0100
Message-ID<ImY94-2heG-19@gated-at.bofh.it>
In reply to#268465

[Multipart message — attachments visible in raw view] — view raw

On 2024-03-20 22:58, Nicholas Geovanis wrote:
>
> On Wed, Mar 20, 2024, 11:28 AM Jesper Dybdal 
> <jd-debian-user@dybdal.dk> wrote:
>
>     I have now done the following:
>     * Checked the RAID array - no problems found.
>     * Run fsck.  It found three cases of the block count being
>     incorrect.  I
>     don't know which the other two affected files are.
>     * Run one pass of memtest86+.  Nothing found.
>
>     So it seems not to be a problem with the disks.
>     A bug in ext4?  Well, ext4 has always done its job for me wihtout
>     problems.
>     A RAM error that memtest86+ did not find?  Possible.  Once upon a
>     time,
>     when you bought an ordinary pc, its RAM had ECC as a matter of
>     course;
>     unfortunately, that is not the case nowadays.
>
>     I think I'll let memtest86+ run overnight one of the coming nights.
>
>     Unless it is simply a RAM error, then it is a bit scary...
>
>
I've now let memtest86+ run for 8 hours, during which i did 14 passes of 
all its tests.  It found nothing wrong.
> I have seen that a couple times, unlikely but possible. Maybe review 
> your RAM configuration too, ensure that the sticks are on the same 
> supported refresh rate and distributed across the slots in an approved 
> way.
>
>     Regards,
>     Jesper
>
>     On 2024-03-19 21:47, Franco Martelli wrote:
>     > On 19/03/24 at 15:43, Jesper Dybdal wrote:
>     >
>     >>
>     >> My plan is to boot a rescue disk and mount that partition
>     read-only.
>     >> Then:
>     >> * If the file looks ok after reboot, then I'll strongly suspect
>     the
>     >> RAM - and run memtest.
>     >> * Otherwise, I'll have to run fsck and see what happens.
>     >>
>     >> kernel version:
>     >> root@nuser:~# uname -a
>     >> Linux nuser 5.10.0-28-amd64 #1 SMP Debian 5.10.209-2 (2024-01-31)
>     >> x86_64 GNU/Linux
>     >>
>     >> The partition in question is a RAID 1 controlled by md.
>     >
>     > Another check you can perform it is on the RAID array, by
>     default it
>     > runs on the first Sunday of each month at 00:57. You should have
>     this
>     > file /etc/cron.d/mdadm that takes care to run this check monthly.
>     >
>     > Before you reboot, does it look OK /proc/mdstat ?
>     >
>
>     -- 
>     Jesper Dybdal
>     https://www.dybdal.dk
>
>
>

-- 
Jesper Dybdal
https://www.dybdal.dk

[toc] | [prev] | [next] | [standalone]


#268578

FromJesper Dybdal <jd-debian-user@dybdal.dk>
Date2024-03-28 15:00 +0100
Message-ID<ImYsp-2hoP-7@gated-at.bofh.it>
In reply to#268465

[Multipart message — attachments visible in raw view] — view raw

[Sorry - I accidentally sent this too quickly in an incomplete state.  
Second try here:]

> On Wed, Mar 20, 2024, 11:28 AM Jesper Dybdal 
> <jd-debian-user@dybdal.dk> wrote:
>
>
>     I think I'll let memtest86+ run overnight one of the coming nights.
>
>     Unless it is simply a RAM error, then it is a bit scary...
>

I've now let memtest86+ run for 9 hours, during which it did 14 passes 
of all its tests.  It found nothing wrong.

On 2024-03-20 22:58, Nicholas Geovanis wrote:
> I have seen that a couple times, unlikely but possible. Maybe review 
> your RAM configuration too, ensure that the sticks are on the same 
> supported refresh rate and distributed across the slots in an approved 
> way.

There is only one RAM stick (of 16 GB), so there should be no problems 
of that kind.

I'm afraid I won't find an explanation of that file system corruption :-(

Thanks to Franco and Nicholas for your responses,
Jesper

-- 
Jesper Dybdal
https://www.dybdal.dk

[toc] | [prev] | [next] | [standalone]


#268579

FromHans <hans.ullrich@loop.de>
Date2024-03-28 15:10 +0100
Message-ID<ImYC6-2hLu-7@gated-at.bofh.it>
In reply to#268578
Am Donnerstag, 28. März 2024, 14:49:37 CET schrieb Jesper Dybdal:
Hello,

memtest86+ is for testing RAM, but do you not want to test ext4 filesystem?

If so, I suggest to boot a live system like Knoppix or similar, then run your 
test by using

e2fsck -y /dev/sda1

or wherever your filesystem resides. 

Please pay attention: If you have encrypted filesystems, then first open the 
encryption, do NOT mount the filesystem and then check it, for example:

cryptsetup luksOpen /dev/sda1 data1

then enter the password and now you can run

e2fsck -y /dev/mapper/data1

Note: the word "data1" is only an example, you can name it, whatever you want 
like "space", "soap", "bullet", "henry" or whatever.

Hope this helps.

Best

Hans



> [Sorry - I accidentally sent this too quickly in an incomplete state. 
> Second try here:]
> 
> > On Wed, Mar 20, 2024, 11:28 AM Jesper Dybdal
> > 
> > <jd-debian-user@dybdal.dk> wrote:
> >     I think I'll let memtest86+ run overnight one of the coming nights.
> >     
> >     Unless it is simply a RAM error, then it is a bit scary...
> 
> I've now let memtest86+ run for 9 hours, during which it did 14 passes
> of all its tests.  It found nothing wrong.
> 
> On 2024-03-20 22:58, Nicholas Geovanis wrote:
> > I have seen that a couple times, unlikely but possible. Maybe review
> > your RAM configuration too, ensure that the sticks are on the same
> > supported refresh rate and distributed across the slots in an approved
> > way.
> 
> There is only one RAM stick (of 16 GB), so there should be no problems
> of that kind.
> 
> I'm afraid I won't find an explanation of that file system corruption :-(
> 
> Thanks to Franco and Nicholas for your responses,
> Jesper

[toc] | [prev] | [next] | [standalone]


#268581

FromJesper Dybdal <jd-debian-user@dybdal.dk>
Date2024-03-28 16:10 +0100
Message-ID<ImZya-2iyZ-25@gated-at.bofh.it>
In reply to#268579

[Multipart message — attachments visible in raw view] — view raw

On 2024-03-28 15:02, Hans wrote:
> Am Donnerstag, 28. März 2024, 14:49:37 CET schrieb Jesper Dybdal:
> Hello,
>
> memtest86+ is for testing RAM, but do you not want to test ext4 filesystem?

Sorry - I should have left more of the previous mails quoted.  I have 
previously tested the RAID1 consistency (ok), fixed the file system 
(found 3 files with incorrect block count), and now also tested the 
RAM.And since it seems unlikely that it is a bug in ext4 (in Debian 
Bullseye), I don't quite understand how such an inconsistency can occur. 
Thanks for your response, Jesper
> If so, I suggest to boot a live system like Knoppix or similar, then run your
> test by using
>
> e2fsck -y /dev/sda1
>
> or wherever your filesystem resides.
>
> Please pay attention: If you have encrypted filesystems, then first open the
> encryption, do NOT mount the filesystem and then check it, for example:
>
> cryptsetup luksOpen /dev/sda1 data1
>
> then enter the password and now you can run
>
> e2fsck -y /dev/mapper/data1
>
> Note: the word "data1" is only an example, you can name it, whatever you want
> like "space", "soap", "bullet", "henry" or whatever.
>
> Hope this helps.
>
> Best
>
> Hans
>
>
>
>> [Sorry - I accidentally sent this too quickly in an incomplete state.
>> Second try here:]
>>
>>> On Wed, Mar 20, 2024, 11:28 AM Jesper Dybdal
>>>
>>> <jd-debian-user@dybdal.dk>  wrote:
>>>      I think I'll let memtest86+ run overnight one of the coming nights.
>>>      
>>>      Unless it is simply a RAM error, then it is a bit scary...
>> I've now let memtest86+ run for 9 hours, during which it did 14 passes
>> of all its tests.  It found nothing wrong.
>>
>> On 2024-03-20 22:58, Nicholas Geovanis wrote:
>>> I have seen that a couple times, unlikely but possible. Maybe review
>>> your RAM configuration too, ensure that the sticks are on the same
>>> supported refresh rate and distributed across the slots in an approved
>>> way.
>> There is only one RAM stick (of 16 GB), so there should be no problems
>> of that kind.
>>
>> I'm afraid I won't find an explanation of that file system corruption :-(
>>
>> Thanks to Franco and Nicholas for your responses,
>> Jesper
>
>
>

-- 
Jesper Dybdal
https://www.dybdal.dk

[toc] | [prev] | [next] | [standalone]


#268583

FromHans <hans.ullrich@loop.de>
Date2024-03-28 16:40 +0100
Message-ID<In01b-2iOb-5@gated-at.bofh.it>
In reply to#268581
Hi Jesper,

RAID 1 is mirroring. I suppose, a reason for the failure might be a timing 
problem. I do not know for sure, if yous system has got a real RAID-controller 
or if it is made by software.

The real controller should not produce write errors, however maybe at heavy 
load it might happen. I never used RAID 1 myself, as I am a fancy guy and am 
no friend of RAID 1. It is just, when there is an error on one drive, it is on 
the other, too.

My fancy solution was, using one drive and mirror this frequently every 30 
minutes using rsync. IMO doing so, I have several options:

1. If the harddrive is defective, I can boot the other one.

2. If the software is defective, I have 30 minutes, to discover the failure 
(every good logging system should alarm this in time)

3. I have a running backup available.

4. I can exchange the defective harddrive during the running system.

5. After exchange, i can examine, what happened (hardware failure, malware, 
whatever).

Many people will now laugh at me, but doing so, worked for me at best. So I 
reached an uptime of more than 700 days, but this might not be based on my 
work, but the work of all the debian developers!

As I said before, i am not very experienced with RAID 1, other people might 
know much more.

Personally I believe, RAID is mostly used with Windows, as Windows does not 
have these nice tools like rsync or syslog and all the things, that make linux 
and debian so great.

Have a nice eastern!

Best

Hans 

> Sorry - I should have left more of the previous mails quoted.  I have
> previously tested the RAID1 consistency (ok), fixed the file system
> (found 3 files with incorrect block count), and now also tested the
> RAM.And since it seems unlikely that it is a bug in ext4 (in Debian
> Bullseye), I don't quite understand how such an inconsistency can occur.
> Thanks for your response, Jesper
> 
> > If so, I suggest to boot a live system like Knoppix or similar, then run
> > your test by using
> > 
> > e2fsck -y /dev/sda1
> > 
> > or wherever your filesystem resides.
> > 
> > Please pay attention: If you have encrypted filesystems, then first open
> > the encryption, do NOT mount the filesystem and then check it, for
> > example:
> > 
> > cryptsetup luksOpen /dev/sda1 data1
> > 
> > then enter the password and now you can run
> > 
> > e2fsck -y /dev/mapper/data1
> > 
> > Note: the word "data1" is only an example, you can name it, whatever you
> > want like "space", "soap", "bullet", "henry" or whatever.
> > 
> > Hope this helps.
> > 
> > Best
> > 
> > Hans
> > 
> >> [Sorry - I accidentally sent this too quickly in an incomplete state.
> >> Second try here:]
> >> 
> >>> On Wed, Mar 20, 2024, 11:28 AM Jesper Dybdal
> >>> 
> >>> <jd-debian-user@dybdal.dk>  wrote:
> >>>      I think I'll let memtest86+ run overnight one of the coming nights.
> >>>      
> >>>      Unless it is simply a RAM error, then it is a bit scary...
> >> 
> >> I've now let memtest86+ run for 9 hours, during which it did 14 passes
> >> of all its tests.  It found nothing wrong.
> >> 
> >> On 2024-03-20 22:58, Nicholas Geovanis wrote:
> >>> I have seen that a couple times, unlikely but possible. Maybe review
> >>> your RAM configuration too, ensure that the sticks are on the same
> >>> supported refresh rate and distributed across the slots in an approved
> >>> way.
> >> 
> >> There is only one RAM stick (of 16 GB), so there should be no problems
> >> of that kind.
> >> 
> >> I'm afraid I won't find an explanation of that file system corruption :-(
> >> 
> >> Thanks to Franco and Nicholas for your responses,
> >> Jesper

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web