Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.kernel > #87523 > unrolled thread

Do no harm: Data loss from new (bookworm2trixie) discard=async default

Started byNicholas D Steeves <sten@debian.org>
First post2025-05-10 20:10 +0200
Last post2025-05-11 15:00 +0200
Articles 2 — 2 participants

Back to article view | Back to linux.debian.kernel


Contents

  Do no harm: Data loss from new (bookworm2trixie) discard=async default Nicholas D Steeves <sten@debian.org> - 2025-05-10 20:10 +0200
    Re: Do no harm: Data loss from new (bookworm2trixie) discard=async  default Pascal Hambourg <pascal@plouf.fr.eu.org> - 2025-05-11 15:00 +0200

#87523 — Do no harm: Data loss from new (bookworm2trixie) discard=async default

FromNicholas D Steeves <sten@debian.org>
Date2025-05-10 20:10 +0200
SubjectDo no harm: Data loss from new (bookworm2trixie) discard=async default
Message-ID<KKWO5-28Qt-1@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Hi,

Sorry for not preempting the backlog sooner.

An ultra-brief history: many SSDs including various Samsungs, and if I
remember correctly many drives with old SandForce controllers have
broken discard=async.  This was a big issue back in 2011-2014, and in
some (many?) cases it was a data loss risk.

Linux-6.2 started enabling discard=async by default (at least for
btrfs), and deductively this appears to necessarily harm many users of
at least pre2011-to-2014 SSDs.  Does Linux-6.12.x, for trixie, have
sufficient quirk coverage to make the new default safe, and fall to back
to discard=sync for affected hardware?  Alternatively, has our kernel
been patched to maintain bookworm's 6.2.x behaviour of discard=sync?

Security conscious users maintain that it presents a security risk when
a filesystem issues discards to the underlying LUKS layer.  Are we going
to start doing this by default for trixie, or are we still going to
block it at the dm-crypt layer?

I'm most concerned about the btrfs-specific case, where "mount -o
discard" is significantly riskier than running fstrim; it's a major
contributing factor to those old "btrfs ate my data" stories, and the
primary motivation for my Debian involvement is safe defaults for btrfs.

Or do we ship a default configuration that provides the best performance
for recent (five years) systems, and that is probably safe most
mainstream systems?  It looks like that's where we are now.  In this
case, are release notes really enough for what sounds like a data loss
risk?

Best,
Nicholas

[toc] | [next] | [standalone]


#87526 — Re: Do no harm: Data loss from new (bookworm2trixie) discard=async default

FromPascal Hambourg <pascal@plouf.fr.eu.org>
Date2025-05-11 15:00 +0200
SubjectRe: Do no harm: Data loss from new (bookworm2trixie) discard=async default
Message-ID<KLerD-2jVZ-3@gated-at.bofh.it>
In reply to#87523
On 10/05/2025 at 20:07, Nicholas D Steeves wrote:
> 
> An ultra-brief history: many SSDs including various Samsungs, and if I
> remember correctly many drives with old SandForce controllers have
> broken discard=async.

Correct me if I am wrong, but my understanding is that these SSDs have a 
broken *queued TRIM* feature. discard=sync and discard=async are btrfs 
features and both use queued TRIM if supported by the SSD (and not 
blackisted by the kernel) or non-queued TRIM otherwise.

> Linux-6.2 started enabling discard=async by default (at least for
> btrfs),

To be clear for everyone: without an explicit discard mount option, the 
default for btrfs is now to enable discard in asynchronous mode 
(discard=async) instead of disabling discard (nodiscard). Explicit 
"discard" still enables discard in synchronous mode, (discard=sync). 
Explicit "nodiscard" is now needed to disable discard.

> and deductively this appears to necessarily harm many users of
> at least pre2011-to-2014 SSDs.  Does Linux-6.12.x, for trixie, have
> sufficient quirk coverage to make the new default safe, and fall to back
> to discard=sync for affected hardware?  Alternatively, has our kernel
> been patched to maintain bookworm's 6.2.x behaviour of discard=sync?

As I understand it, it is not the asynchronous mode which may cause data 
loss with non-blacklisted broken TRIM but rather the discard option as a 
whole.

> Security conscious users maintain that it presents a security risk when
> a filesystem issues discards to the underlying LUKS layer.  Are we going
> to start doing this by default for trixie, or are we still going to
> block it at the dm-crypt layer?

The Debian installer already enables discard by default on encrypted 
devices since buster. The "discard" option is available for most 
filesystem types which support it (btrfs, ext4, FAT, HFS+, XFS) but is 
not enabled by default. JFS supports it since Linux 3.7 but this is not 
mentioned in mount(8). It is not available for swap.

The Calamares installer may have different defaults. The package 
calamares-settings-debian has a file /etc/calamares/modules/fstab.conf 
which contains:

ssdExtraMountOptions:
     ext4: discard
     jfs: discard
     xfs: discard
     swap: discard
     btrfs: discard,compress=lzo

Also the fstrim systemd service is enabled and triggered once a week by 
default. Is it safer than online discard with broken TRIM ? If yes, can 
anyone explain why ?

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.kernel


csiph-web