Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #67833

Bug#833170: linux-image-3.16.0-4-amd64: Reproducable XFS filesystem corruption, possibly connected with ACLs

Path csiph.com!aioe.org!bofh.it!news.nic.it!robomod
From Salvatore Bonaccorso <carnil@debian.org>
Newsgroups linux.debian.bugs.dist, linux.debian.kernel
Subject Bug#833170: linux-image-3.16.0-4-amd64: Reproducable XFS filesystem corruption, possibly connected with ACLs
Date Sun, 23 Aug 2020 17:50:01 +0200
Message-ID <AH0pP-8a5-1@gated-at.bofh.it> (permalink)
References <s1pm2-Vl-7@gated-at.bofh.it> <s1pm2-Vl-7@gated-at.bofh.it>
X-Original-To Will Aoki <waoki@umnh.utah.edu>, 833170@bugs.debian.org
X-Mailbox-Line From debian-bugs-dist-request@lists.debian.org Sun Aug 23 15:45:09 2020
Old-Return-Path <debbugs@buxtehude.debian.org>
X-Spam-Flag NO
X-Spam-Score -2.4
Reply-To Salvatore Bonaccorso <carnil@debian.org>, 833170@bugs.debian.org
Original-Sender Salvatore Bonaccorso <salvatore.bonaccorso@gmail.com>
Resent-To debian-bugs-dist@lists.debian.org
Resent-Cc Debian Kernel Team <debian-kernel@lists.debian.org>
X-Debian-Pr-Message followup 833170
X-Debian-Pr-Package src:linux
X-Debian-Pr-Source linux
X-Spam-Bayes score:0.0000 Tokens: new, 86; hammy, 150; neutral, 280; spammy, 0. spammytokens: hammytokens:0.000-+--UD:xz, 0.000-+--H*F:U*carnil, 0.000-+--H*r:TLS1_3, 0.000-+--Hx-spam-relays-external:eldamar, 0.000-+--Hx-spam-relays-external:sk:80-218-
Dkim-Signature v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20161025; h=sender:date:from:to:subject:message-id:references:mime-version :content-disposition:content-transfer-encoding:in-reply-to; bh=El2m+wnBXTS6YEpmT5ZhWULM6xOSyrmvYq6KoHY4igE=; b=hD9ZkOXeN/QCCf7SAL2/WfRO9Y42I4HsmyKmDqFSG+mQ+GxX4MRVE3x/7LWbPpzS3o iB1GYk/VHqbRaOnQ0flzn/ZdluM/2V/dZ7pEoDRVUurdPBhBEjEuYbOVjT2aSK+h+PAl C3kB4Ki5h4fMka8IpXxaPHHtqvnGs52uRSiVDRwcFAH27HaPMdAS6f5THPFvgNUVSSnh pZ6RZ4E5SYrhp3YqQQlxu+UciY4A977GoVO8J9z1QilZ8AzTlmeKitUfDUhZBiOnTL1U HbmS64lDPiRq2+1PaV61wg+hM0ldmaEW30I6j1ncoPp1BnbOPm1FQyFZtuyJz2qhohDj x5XQ==
X-Google-Dkim-Signature v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20161025; h=x-gm-message-state:sender:date:from:to:subject:message-id :references:mime-version:content-disposition :content-transfer-encoding:in-reply-to; bh=El2m+wnBXTS6YEpmT5ZhWULM6xOSyrmvYq6KoHY4igE=; b=m5UKHdLet8wQ/LFKR++dh1KLbsSuIEqqAWV7DTedjNuwAaLWuMb+0BJvVIwOjg07aP oK+GLYqvA+xHE3C3P+Ch5wSiZPNQ4Wrr7mRsv851eazgA1r9c/823Fcp8rhxKevuAAzg ZhEq1Yeg6OU0c6C9U5QH53wo6cFsH2+xf4gRZvo5k5DACR9mRSaXEIGKoikGEt5mu/m6 j/WbvPa8G6dxYiD3ZOrAnUbL+peAjqn8MyCPOQ22BnSQr+h0j7ymisVh11MKBNN1RygP jEE3hFhEVdfVaAMqO5omy9VKqtCzvvjk6ftyJiqzKc1yzTxpNGUWJ1cxD0HNt4Hq3Roj yHVQ==
X-Gm-Message-State AOAM531bkkmTY1N9kr04yh109WbWaKcQXNDYZU11sI6aKMTZ/MYt4GR+ m/A+kiqxfU/aO78IgcAXXmYtsQEeD1OK1A==
X-Google-SMTP-Source ABdhPJwdytsZVe5bxNx/zCPLNuLzPTxh6daAtybkITvNetXOZs0oyDaUEEtKG5uTqmjj4MPpi9OoHw==
X-Received by 2002:a5d:6944:: with SMTP id r4mr2175709wrw.132.1598197412808; Sun, 23 Aug 2020 08:43:32 -0700 (PDT)
Sender robomod@news.nic.it
MIME-Version 1.0
Content-Type text/plain; charset=utf-8
Content-Disposition inline
Content-Transfer-Encoding 8bit
X-Debian-Message from BTS
X-Mailing-List <debian-bugs-dist@lists.debian.org> archive/latest/1619785
List-ID <debian-bugs-dist.lists.debian.org>
List-URL <https://lists.debian.org/debian-bugs-dist/>
Approved robomod@news.nic.it
Lines 165
Organization linux.* mail to news gateway
X-Original-Date Sun, 23 Aug 2020 17:43:31 +0200
X-Original-Message-ID <20200823154331.GA1273666@eldamar.local>
X-Original-References <20160801172528.GF10297@umnh.utah.edu> <20160801172528.GF10297@umnh.utah.edu>
X-Original-Sender Salvatore Bonaccorso <salvatore.bonaccorso@gmail.com>
Xref csiph.com linux.debian.bugs.dist:1022522 linux.debian.kernel:67833

Cross-posted to 2 groups.

Show key headers only | View raw


Hi,

On Mon, Aug 01, 2016 at 11:25:28AM -0600, Will Aoki wrote:
> Package: src:linux
> Version: 3.16.7-ckt25-2+deb8u3
> Severity: important
> 
> I hit a nasty filesystem corruption bug while restoring some backups. I'm able
> to reliably reproduce this on every system where I tried to restore to XFS and
> on a brand-new VM created for testing (which I can share as an OVF, although
> it's pretty big).
> 
> Everything's running on hardware with ECC RAM, so memory errors are unlikely.
> I've reproduced it on VMs spread across three different storage arrays from two
> different vendors. Everything has at least two virtual CPUs.
> 
> All systems tested use XFS on LVM.
> 
> In my test VM, the only packages not from Debian stable are custom stow and tar
> packages to fix bugs in the versions in the stable release and a backport of
> burp because the verson in Debian was very old. These are all userland tools
> and none should be able to cause filessytem corruption.
> 
> I'm using xfsprogs from Debian jessie. On one affected system, I tried version
> 4.3.0+nmu1 and saw no difference in what xfs_repair found.
> 
> 
> At this point, my test case uses the burp backup software to create the I/O
> activity which triggers this bug. I have not been able to make tar trigger this
> problem.
> 
> 
> Steps to reproduce:
> 
> 1: Create new XFS filesystem & mount it on /srv/src
> 
> 2: Create some directories in /srv/src & set ACLs (including default ACLs) on
>    them
> 
> 3: Generate deep tree of files in each of the directories from step #2. For
>    testing, I used a script which created random files & subdirectories. Total
>    bulk was about 2.5 gigabytes.
> 
> 4: Take backup of /srv/src with burp:
> 
>    # burp -a b
> 
> 5: Unmount /srv/src
> 
> 6: Create new XFS filesystem & mount it on /srv/src
> 
> 7: Run restore to /srv/src:
> 
>    # burp -a r -r ^/srv/src
> 
>    Do not suspend the restore process: the bug appears to require sustained I/O
>    to trigger. In trials where I suspended it multiple times during a restore,
>    corruption did not surface.
> 
> 
> Expected outcome (observed when restoring to e.g. ext4):
> 
> 1: Can create files (permissions notwithstanding) in every directory under
>    /srv/src
> 
> 2: Default ACL on every directory is the same as the backup utility wrote
> 
> 3: If filesystem is unmounted and xfs_repair is run on it, no errors will be
>    found
> 
> 
> Actual outcome (observed when restoring to XFS):
> 
> 1: Some files & directories cannot be written. The easiest way to find problem
>    directories them is:
> 
>    # find . -type d -exec touch {}/asdf \;
>    touch: cannot touch ‘./aaaaa/-BIz/asdf’: Cannot allocate memory
>    touch: cannot touch ‘./aaaaa/-BIz/Zp.NyvX0guz./asdf’: Cannot allocate memory
>    touch: cannot touch ‘./aaaaa/-BIz/Zp.NyvX0guz./TWDU/asdf’: Cannot allocate memory
>    [etc]
> 
>    Giving VMs more RAM has no effect on this. Clearing the ACL on the directory
>    has no effect.
> 
>    Affected directories are not always the same between different runs.
> 
> 2: The default ACL has not been restored to problem directories. Directories
>    which I can write to have had the default ACL restored.
> 
> 3: If filesystem is unmounted and xfs_repair is run on it, many errors are
>    reported:
> 
>    # xfs_repair -n /dev/mapper/xfsbugtest--vg-dst 2>&1 | head -90
>    Phase 1 - find and verify superblock...
>    Phase 2 - using internal log
>            - scan filesystem freespace and inode maps...
>            - found root inode chunk
>    Phase 3 - for each AG...
>            - scan (but don't clear) agi unlinked lists...
>            - process known inodes and perform inode discovery...
>            - agno = 0
>    Too many ACL entries, count -2010719080
>    entry contains illegal value in attribute named SGI_ACL_FILE or SGI_ACL_DEFAULT
>    bad security value for attribute entry 1 in attr block 0, inode 133
>    problem with attribute contents in inode 133
>    would clear attr fork
>    bad nblocks 2 for inode 133, would reset to 1
>    bad anextents 1 for inode 133, would reset to 0
>    Too many ACL entries, count -2010719080
>    entry contains illegal value in attribute named SGI_ACL_FILE or SGI_ACL_DEFAULT
>    bad security value for attribute entry 1 in attr block 0, inode 134
>    problem with attribute contents in inode 134
>    would clear attr fork
>    bad nblocks 2 for inode 134, would reset to 1
>    bad anextents 1 for inode 134, would reset to 0
>    [...]
>    bad nblocks 1 for inode 52741928, would reset to 0
>    bad anextents 1 for inode 52741928, would reset to 0
>            - process newly discovered inodes...
>    Phase 4 - check for duplicate blocks...
>            - setting up duplicate extent list...
>            - check for inodes claiming duplicate blocks...
>            - agno = 0
>            - agno = 1
>            - agno = 2
>            - agno = 3
>    No modify flag set, skipping phase 5
>    Phase 6 - check inode connectivity...
>            - traversing filesystem ...
>            - traversal finished ...
>            - moving disconnected inodes to lost+found ...
>    Phase 7 - verify link counts...
>    No modify flag set, skipping filesystem flush and exiting.
> 
>    On my production VMs, running xfs_repair without '-n' typically left many
>    files (the highest was 148k) in /lost+found and left many directories
>    without ACLs.
> 
> 
> 
> xfs_info output on a corrupted filesystem on the test VM:
> 
> meta-data=/dev/mapper/xfsbugtest--vg-dst isize=256    agcount=4, agsize=655360 blks
>          =                       sectsz=512   attr=2, projid32bit=1
>          =                       crc=0        finobt=0
> data     =                       bsize=4096   blocks=2621440, imaxpct=25
>          =                       sunit=0      swidth=0 blks
> naming   =version 2              bsize=4096   ascii-ci=0 ftype=0
> log      =internal               bsize=4096   blocks=2560, version=2
>          =                       sectsz=512   sunit=0 blks, lazy-count=1
> realtime =none                   extsz=4096   blocks=0, rtextents=0
> 
> 
> xfs_metadump of filesystem is at ftp://ftp.umnh.utah.edu/general-temporary/xfs/corrupted.metadump
> 
> 
> Giant (5.9 GB uncompressed) trace-cmd output is at ftp://ftp.umnh.utah.edu/general-temporary/xfs/trace_report.xz

Is this issue reproducible with current supported Debian versions? If
not we might want to close this bug as Jessie respectively v3.16.y is
EOL'ed.

Regards,
Salvatore

Back to linux.debian.kernel | Previous | Next | Find similar | Unroll thread


Thread

Bug#833170: linux-image-3.16.0-4-amd64: Reproducable XFS filesystem corruption, possibly connected with ACLs Salvatore Bonaccorso <carnil@debian.org> - 2020-08-23 17:50 +0200

csiph-web