Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #62271 > unrolled thread

Bug#910074: linux-image-4.9.0-8-amd64: BTRFS data loss - kernel BUG at .../linux-4.9.110/fs/btrfs/ctree.c:3178

Started byNicholas D Steeves <nsteeves@gmail.com>
First post2018-10-02 22:20 +0200
Last post2018-10-09 03:10 +0200
Articles 4 — 2 participants

Back to article view | Back to linux.debian.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Bug#910074: linux-image-4.9.0-8-amd64: BTRFS data loss - kernel BUG at .../linux-4.9.110/fs/btrfs/ctree.c:3178 Nicholas D Steeves <nsteeves@gmail.com> - 2018-10-02 22:20 +0200
    Bug#910074: linux-image-4.9.0-8-amd64: BTRFS data loss - kernel BUG at .../linux-4.9.110/fs/btrfs/ctree.c:3178 Michael Firth <MFirth@nevion.com> - 2018-10-03 12:00 +0200
      Bug#910074: linux-image-4.9.0-8-amd64: BTRFS data loss - kernel BUG at .../linux-4.9.110/fs/btrfs/ctree.c:3178 Nicholas D Steeves <nsteeves@gmail.com> - 2018-10-09 03:30 +0200
    Bug#910074: linux-image-4.9.0-8-amd64: BTRFS data loss - kernel BUG at .../linux-4.9.110/fs/btrfs/ctree.c:3178 Nicholas D Steeves <nsteeves@gmail.com> - 2018-10-09 03:10 +0200

#62271 — Bug#910074: linux-image-4.9.0-8-amd64: BTRFS data loss - kernel BUG at .../linux-4.9.110/fs/btrfs/ctree.c:3178

FromNicholas D Steeves <nsteeves@gmail.com>
Date2018-10-02 22:20 +0200
SubjectBug#910074: linux-image-4.9.0-8-amd64: BTRFS data loss - kernel BUG at .../linux-4.9.110/fs/btrfs/ctree.c:3178
Message-ID<wEzjb-8nB-3@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Hi Michael,

On Tue, Oct 02, 2018 at 11:37:45AM +0100, Michael Firth wrote:
> 
> After this, there was a file that was errored on the filesystem (as
> reported by 'btrfs check'), and it seems BTRFS doesn't have any tools to
> resolve the error. Deleting the file at the reported inode has cleared
> the error from 'btrfs check', but I am not 100% sure that will have
> fixed all the corruption.
>

As far as I know, if 'btrfs check' is clean then you're in the clear
for any known issues involving the fs structure.  Of course, a 'btrfs
scrub' is necessary to check for data and metadata corruption...  BTW,
if you're using an ssd, make sure you're mounting with -o nossd,
because as far as I know linux-4.9.x still hasn't been patched.
P.S. that requires a full rebalance to take effect.  Make up-to-date
backups before running that rebalance...

> This issue looks very like the bug described at:
> 
> https://www.spinics.net/lists/linux-btrfs/msg60984.html
> 
> And in bug report #708509 for a much older kernel (from 2013)
> 
> According to the BTRFS mailing list post above, there are patches
> submitted to fix this issue (or one with the same symptoms).
> Is there any way to easily determine if these patches are in the Debian
> version of the V4.9.110 kernel?

The last time I checked I couldn't find any btrfs-specific ones in
Debian; I used apt-get source and expected to find a quilt series.

> If not, what is the route to get these patches incorporated? Do I need
> to talk to the BTRFS people about getting the patches in to the stock
> V4.9 kernel, or is this something that the Debian team would apply
> directly?

The first of the two patches from that 29 Nov 2016 linux-btrfs email
appears to be queued for linux-4.9.119:
  https://lore.kernel.org/patchwork/patch/972419/

I wasn't able to find status of the second one wrt linux-4.9.x.

> Kernel bug report output included below, in case it is useful.
> 
> BTRFS may not be a filesystem that everyone uses, but I feel if it is in
> the Debian kernel then bugs that can cause data loss should be fixed if
> a patch already exists.

In principle I agree; although I think it would be safer to coordinate
with Greg Kroah-Harman about getting them applied upstream before
importing them into Debian, since (afaik) we don't have any btrfs
specialists working on our kernel...people who would know if importing
one of these patches will introduce unintended side-effects or a
rabbit hole of patches.  Maybe it would be safer to look at the delta
between btrfs in 4.9.x and 4.14.x and ask for backported fixes from
4.14.x to 4.9.x? (eg: more than six months of testing in 4.14.x, like
the -o ssd bug that is still present in 4.9.x)

Cheers,
Nicholas

[toc] | [next] | [standalone]


#62276

FromMichael Firth <MFirth@nevion.com>
Date2018-10-03 12:00 +0200
Message-ID<wEM6K-7m0-5@gated-at.bofh.it>
In reply to#62271
Hi,

> -----Original Message-----
> 
> Hi Michael,
> 
> On Tue, Oct 02, 2018 at 11:37:45AM +0100, Michael Firth wrote:
> >
> > After this, there was a file that was errored on the filesystem (as
> > reported by 'btrfs check'), and it seems BTRFS doesn't have any tools
> > to resolve the error. Deleting the file at the reported inode has
> > cleared the error from 'btrfs check', but I am not 100% sure that will
> > have fixed all the corruption.
> >
> 
> As far as I know, if 'btrfs check' is clean then you're in the clear for any known
> issues involving the fs structure.  Of course, a 'btrfs scrub' is necessary to
> check for data and metadata corruption...  BTW, if you're using an ssd, make
> sure you're mounting with -o nossd, because as far as I know linux-4.9.x still
> hasn't been patched.
> P.S. that requires a full rebalance to take effect.  Make up-to-date backups
> before running that rebalance...

That is good to know. I will run a "scrub" on the partition soon to check for any
other issues. I am running on a VM on top of a hardware RAID array of
spinning disks, so hopefully the SSD issue doesn't apply.

> 
> > This issue looks very like the bug described at:
> >
> > https://www.spinics.net/lists/linux-btrfs/msg60984.html
> >
> > And in bug report #708509 for a much older kernel (from 2013)
> >
> > According to the BTRFS mailing list post above, there are patches
> > submitted to fix this issue (or one with the same symptoms).
> > Is there any way to easily determine if these patches are in the
> > Debian version of the V4.9.110 kernel?
> 
> The last time I checked I couldn't find any btrfs-specific ones in Debian; I used
> apt-get source and expected to find a quilt series.

There was a BTRFS patch in the update that became available an hour after my
crash:

linux (4.9.110-3+deb9u5) stretch-security; urgency=high
.
.
.
  * btrfs: relocation: Only remove reloc rb_trees if reloc control has been
    initialized (CVE-2018-14609)

I'm not sure if that is a Debian specific patch or whether it is a Debian specific 
merge from another version

> 
> > If not, what is the route to get these patches incorporated? Do I need
> > to talk to the BTRFS people about getting the patches in to the stock
> > V4.9 kernel, or is this something that the Debian team would apply
> > directly?
> 
> The first of the two patches from that 29 Nov 2016 linux-btrfs email appears
> to be queued for linux-4.9.119:
>   https://lore.kernel.org/patchwork/patch/972419/
> 

So I guess the related question that I should have asked is whether there is
information on how upstream changes are merged into the Debian kernel, and
what the likely delay between (for example) the 4.9.119 mainline kernel being
released, and the Debian version following it?

> I wasn't able to find status of the second one wrt linux-4.9.x.

Though the description is similar, I don't think patch 972419 is actually either of
the two patches referenced from that mail. I think I will ask on the BTRFS mailing
list what the current status of all of these patches is.

> 
> > Kernel bug report output included below, in case it is useful.
> >
> > BTRFS may not be a filesystem that everyone uses, but I feel if it is
> > in the Debian kernel then bugs that can cause data loss should be
> > fixed if a patch already exists.
> 
> In principle I agree; although I think it would be safer to coordinate with Greg
> Kroah-Harman about getting them applied upstream before importing them
> into Debian, since (afaik) we don't have any btrfs specialists working on our
> kernel...people who would know if importing one of these patches will
> introduce unintended side-effects or a rabbit hole of patches.  Maybe it
> would be safer to look at the delta between btrfs in 4.9.x and 4.14.x and ask
> for backported fixes from 4.14.x to 4.9.x? (eg: more than six months of
> testing in 4.14.x, like the -o ssd bug that is still present in 4.9.x)
> 
I agree with this, and with Hans's comment that because it isn't a Debian specific
issue it should be handled upstream. I guess the question comes partly from not
knowing if/how/when upstream V4.9.X kernel changes are merged into the Debian
Stretch kernel.

> Cheers,
> Nicholas


Regards

Michael

[toc] | [prev] | [next] | [standalone]


#62319

FromNicholas D Steeves <nsteeves@gmail.com>
Date2018-10-09 03:30 +0200
Message-ID<wGP0t-5Uh-1@gated-at.bofh.it>
In reply to#62276
On Wed, Oct 03, 2018 at 09:56:13AM +0000, Michael Firth wrote:
> > 
> > As far as I know, if 'btrfs check' is clean then you're in the clear for any known
> > issues involving the fs structure.  Of course, a 'btrfs scrub' is necessary to
> > check for data and metadata corruption...  BTW, if you're using an ssd, make
> > sure you're mounting with -o nossd, because as far as I know linux-4.9.x still
> > hasn't been patched.
> > P.S. that requires a full rebalance to take effect.  Make up-to-date backups
> > before running that rebalance...
> 
> That is good to know. I will run a "scrub" on the partition soon to check for any
> other issues. I am running on a VM on top of a hardware RAID array of
> spinning disks, so hopefully the SSD issue doesn't apply.

To find out:

  cat /sys/block/your_hardware_raid_block_device/queue/rotational

If it returns "0" then you're affected by the -o ssd bug, but if it
returns "1" then there is nothing to worry about. :-)  While this
issue will become obsolete when the fix is backported I'm curious to
learn if hardware RAID registers as nonrotational, so please let me
know.  I suspect it will register as nonrotational, because then the
kernel will let the RAID controller merge and reorder IO as it sees
fit.

> 
> There was a BTRFS patch in the update that became available an hour after my
> crash:
> 
> linux (4.9.110-3+deb9u5) stretch-security; urgency=high
> .
> .
> .
>   * btrfs: relocation: Only remove reloc rb_trees if reloc control has been
>     initialized (CVE-2018-14609)
> 
> I'm not sure if that is a Debian specific patch or whether it is a Debian specific 
> merge from another version

Oh Nice!  I'm really happy to see this.  Thank you Ben and kernel team!

> > I wasn't able to find status of the second one wrt linux-4.9.x.
> 
> Though the description is similar, I don't think patch 972419 is actually either of
> the two patches referenced from that mail. I think I will ask on the BTRFS mailing
> list what the current status of all of these patches is.

Thanks.

> > In principle I agree; although I think it would be safer to coordinate with Greg
> > Kroah-Harman about getting them applied upstream before importing them
> > into Debian, since (afaik) we don't have any btrfs specialists working on our
> > kernel...people who would know if importing one of these patches will
> > introduce unintended side-effects or a rabbit hole of patches.  Maybe it
> > would be safer to look at the delta between btrfs in 4.9.x and 4.14.x and ask
> > for backported fixes from 4.14.x to 4.9.x? (eg: more than six months of
> > testing in 4.14.x, like the -o ssd bug that is still present in 4.9.x)
> > 
> I agree with this, and with Hans's comment that because it isn't a Debian specific
> issue it should be handled upstream. I guess the question comes partly from not
> knowing if/how/when upstream V4.9.X kernel changes are merged into the Debian
> Stretch kernel.

The version number is a hint.  Looking at the changelog for the
package, you'll see stretch released with 4.9.30-2, was updated with
security fixes in 4.9.30-2+deb9u1, was updated from upstream LTS in
4.9.47-1, etc, and is now at 4.9.110-3+deb9u6.


Cheers,
Nicholas

[toc] | [prev] | [next] | [standalone]


#62318

FromNicholas D Steeves <nsteeves@gmail.com>
Date2018-10-09 03:10 +0200
Message-ID<wGOH7-5Ot-3@gated-at.bofh.it>
In reply to#62271

[Multipart message — attachments visible in raw view] — view raw

Hi Hans,

On Wed, Oct 03, 2018 at 12:05:31AM +0200, Hans van Kranenburg wrote:
> Hi,
> 
> On 10/02/2018 10:08 PM, Nicholas D Steeves wrote:
> > Hi Michael,
> > 
> > On Tue, Oct 02, 2018 at 11:37:45AM +0100, Michael Firth wrote:
> >>
> >> BTRFS may not be a filesystem that everyone uses, but I feel if it is in
> >> the Debian kernel then bugs that can cause data loss should be fixed if
> >> a patch already exists.
> > 
> > In principle I agree; although I think it would be safer to coordinate
> > with Greg Kroah-Harman about getting them applied upstream before
> > importing them into Debian, since (afaik)
> 
> > we don't have any btrfs
> > specialists working on our kernel...people who would know if importing
> > one of these patches will introduce unintended side-effects or a
> > rabbit hole of patches.
> 
> This is not a debian specific issue. The upstream btrfs team does not
> have enough work capacity to do this, and mainly focuses on going
> forward instead of looking back. And I don't think there's really
> someone who would know the things mentioned above except for the authors
> of the patches themselves (who tag them for stable if it's data
> corruption and if they know it will work (tm)), or the btrfs maintainer
> who knows which ones to put together in which order to prepare the next
> kernel release.

Agreed!  Also, acknowledging when issues aren't Debian-specific and
then working with upstream so everyone can benefit is one reason we
have a great reputation for giving back to the larger community :-)

> 
> >  Maybe it would be safer to look at the delta
> > between btrfs in 4.9.x and 4.14.x and ask for backported fixes from
> > 4.14.x to 4.9.x? (eg: more than six months of testing in 4.14.x, like
> > the -o ssd bug that is still present in 4.9.x)
> 
> For the -o ssd issue, in hindsight, it was a mistake to not get that
> into 4.9 earlier.
> 
> Every user who wants to try out btrfs on his/her computer with Stretch
> and uses it as the root filesystem on a disk which is not too large is
> still affected by this sub-optimal behaviour.
> 
> So I guess that's a TODO for me, to still get it done now. It's 951e7966
> and 583b723151 with a few small changes to make it apply. At least it
> has had enough testing, and the amount of users with out-of-space
> filesystems has decreased notably in the last year in #btrfs IRC. :)

Thank you!  I updated our wiki page within a week of learning about
the patch, but in the future would you prefer if I file a bug?  I
don't imagine it will be more than two bugs a year ;-)  Btw, would you
please forward the bug to me for the 951e7966 and 583b723151 backport?

Sincerely,
Nicholas

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web