Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #272376 > unrolled thread

wait until swapoff is *actually* finished (it returns too early)?

Started byThorsten Glaser <tg@debian.org>
First post2024-08-21 02:00 +0200
Last post2024-08-23 22:20 +0200
Articles 20 — 10 participants

Back to article view | Back to linux.debian.user


Contents

  wait until swapoff is *actually* finished (it returns too early)? Thorsten Glaser <tg@debian.org> - 2024-08-21 02:00 +0200
    Re: wait until swapoff is *actually* finished (it returns too early)? Roberto C. Sánchez <roberto@debian.org> - 2024-08-21 02:10 +0200
      Re: wait until swapoff is *actually* finished (it returns too early)? Greg Wooledge <greg@wooledge.org> - 2024-08-21 02:20 +0200
        Re: wait until swapoff is *actually* finished (it returns too early)? Erwan David <erwan@rail.eu.org> - 2024-08-21 08:50 +0200
          Re: wait until swapoff is *actually* finished (it returns too early)? Greg Wooledge <greg@wooledge.org> - 2024-08-21 13:40 +0200
            Re: wait until swapoff is *actually* finished (it returns too early)? Alain D D Williams <addw@phcomp.co.uk> - 2024-08-21 14:20 +0200
          Re: wait until swapoff is *actually* finished (it returns too early)? Alain D D Williams <addw@phcomp.co.uk> - 2024-08-21 13:50 +0200
        Re: wait until swapoff is *actually* finished (it returns too early)? Roberto C. Sánchez <roberto@debian.org> - 2024-08-21 13:50 +0200
    Re: wait until swapoff is *actually* finished (it returns too early)? Stefan Monnier <monnier@iro.umontreal.ca> - 2024-08-21 15:40 +0200
    Re: wait until swapoff is *actually* finished (it returns too early)? Franco Martelli <martellif67@gmail.com> - 2024-08-21 16:00 +0200
      Re: wait until swapoff is *actually* finished (it returns too early)? Thorsten Glaser <tg@debian.org> - 2024-08-21 23:20 +0200
        Re: wait until swapoff is *actually* finished (it returns too early)? Stefan Monnier <monnier@iro.umontreal.ca> - 2024-08-22 14:50 +0200
          Re: wait until swapoff is *actually* finished (it returns too early)? Mike Castle <dalgoda+debian@gmail.com> - 2024-08-22 17:50 +0200
            Re: wait until swapoff is *actually* finished (it returns too early)? Thorsten Glaser <tg@debian.org> - 2024-08-22 19:30 +0200
              Re: wait until swapoff is *actually* finished (it returns too early)? <tomas@tuxteam.de> - 2024-08-22 21:40 +0200
                Re: wait until swapoff is *actually* finished (it returns too early)? Mike Castle <dalgoda+debian@gmail.com> - 2024-08-23 05:40 +0200
              Re: wait until swapoff is *actually* finished (it returns too early)? Mike Castle <dalgoda+debian@gmail.com> - 2024-08-23 04:50 +0200
                Re: wait until swapoff is *actually* finished (it returns too early)? Thorsten Glaser <tg@debian.org> - 2024-08-23 05:20 +0200
    Re: wait until swapoff is *actually* finished (it returns too early)? Pierre-Elliott Bécue <peb@debian.org> - 2024-08-23 13:10 +0200
      Re: wait until swapoff is *actually* finished (it returns too early)? Thorsten Glaser <tg@debian.org> - 2024-08-23 22:20 +0200

#272376 — wait until swapoff is *actually* finished (it returns too early)?

FromThorsten Glaser <tg@debian.org>
Date2024-08-21 02:00 +0200
Subjectwait until swapoff is *actually* finished (it returns too early)?
Message-ID<JdGVz-6qT2-1@gated-at.bofh.it>
(Please d̲o̲ Cc me on replies, I don’t subscribe to this list. Thanks!)

Hi,

this is a bit curious problem:

I have a setup with swap devices on dmcrypt:

$ cat /etc/crypttab
# <target name> <source device>         <key file>      <options>
crtpv           LABEL=fooclvm           none            discard,luks,initramfs
cswp1           /dev/vg-foo/lv-swp1     /dev/random     discard,cipher=aes-xts-plain64,size=256,plain,swap
cswp2           /dev/vg-foo/lv-swp2     /dev/random     discard,cipher=aes-xts-plain64,size=256,plain,swap

In a cronjob, I basically do swapoff && cryptdisks_stop && \
cryptdisks_start && swapon for both swaps individually to
throw away the old encryption key regularily (but not too
frequently).

I immediately ran into the problem, when trying this for the
first time, that a “swapoff /dev/mapper/cswp1” returns before
the device is released, so the subsequent cryptdisks_stop fails.

I found that inserting a “cat /proc/swaps”, funnily enough,
makes those failures less frequent but still present; adding
a “sleep 3” as well made it work for months.

Until tonight when it didn’t.

Just adding a “sleep” is no proper fix anyway, so the question
is, how to wait in a shell script until the swap device is
*really* swapoff’d when the syscall returns too early, and
(someone from the Linux kernel maintainers reading this?) should
I report the latter as a bug against the kernel?

This is on bullseye/amd64, on VMs and bare metal both. Using
direct partitions like /dev/sda3 (or via LABEL= to avoid trouble)
makes no difference from using LVs.

Thanks in advance,
//mirabilos
-- 
16:47⎜«mika:#grml» .oO(mira ist einfach gut....)      23:22⎜«mikap:#grml»
mirabilos: und dein bootloader ist geil :)    23:29⎜«mikap:#grml» und ich
finds saugeil dass ich ein bsd zum booten mit grml hab, das muss ich dann
gleich mal auf usb-stick installieren	-- Michael Prokop über MirOS bsd4grml

[toc] | [next] | [standalone]


#272378

FromRoberto C. Sánchez <roberto@debian.org>
Date2024-08-21 02:10 +0200
Message-ID<JdH5f-6rbQ-21@gated-at.bofh.it>
In reply to#272376
On Tue, Aug 20, 2024 at 11:34:32PM +0000, Thorsten Glaser wrote:
> 
> Just adding a “sleep” is no proper fix anyway, so the question
> is, how to wait in a shell script until the swap device is
> *really* swapoff’d when the syscall returns too early, and
> (someone from the Linux kernel maintainers reading this?) should
> I report the latter as a bug against the kernel?
> 
I forget where and when (a long time ago?) but I recall having learned
that prior to swapoff it is necessary to call sync and in my history I
have it like this:

sync && sync && sync && swapoff

I couldn't tell why I have sync 3 times, but I know that it's how I've
called swapoff since as far back as I can remember.

Regards,

-Roberto

-- 
Roberto C. Sánchez

[toc] | [prev] | [next] | [standalone]


#272379

FromGreg Wooledge <greg@wooledge.org>
Date2024-08-21 02:20 +0200
Message-ID<JdHeV-6rgA-5@gated-at.bofh.it>
In reply to#272378
On Tue, Aug 20, 2024 at 20:04:11 -0400, Roberto C. Sánchez wrote:
> sync && sync && sync && swapoff
> 
> I couldn't tell why I have sync 3 times, but I know that it's how I've
> called swapoff since as far back as I can remember.

Cargo cult.  It was never useful to the best of my knowledge.

[toc] | [prev] | [next] | [standalone]


#272383

FromErwan David <erwan@rail.eu.org>
Date2024-08-21 08:50 +0200
Message-ID<JdNkl-6voj-5@gated-at.bofh.it>
In reply to#272379
On Wed, Aug 21, 2024 at 02:18:44AM CEST, Greg Wooledge <greg@wooledge.org> said:
> On Tue, Aug 20, 2024 at 20:04:11 -0400, Roberto C. Sánchez wrote:
> > sync && sync && sync && swapoff
> > 
> > I couldn't tell why I have sync 3 times, but I know that it's how I've
> > called swapoff since as far back as I can remember.
> 
> Cargo cult.  It was never useful to the best of my knowledge.

Once upon a time, the sync command would return before the actual
syscall where completed. Doing 3 times sync gave you a very high
probability that the first one indeed completed all its writes.

But it was already false in SunOS 4 (~1990)

-- 
Erwan David

[toc] | [prev] | [next] | [standalone]


#272390

FromGreg Wooledge <greg@wooledge.org>
Date2024-08-21 13:40 +0200
Message-ID<JdRQZ-6y9S-5@gated-at.bofh.it>
In reply to#272383
On Wed, Aug 21, 2024 at 08:39:37 +0200, Erwan David wrote:
> On Wed, Aug 21, 2024 at 02:18:44AM CEST, Greg Wooledge <greg@wooledge.org> said:
> > On Tue, Aug 20, 2024 at 20:04:11 -0400, Roberto C. Sánchez wrote:
> > > sync && sync && sync && swapoff
> > > 
> > > I couldn't tell why I have sync 3 times, but I know that it's how I've
> > > called swapoff since as far back as I can remember.
> > 
> > Cargo cult.  It was never useful to the best of my knowledge.
> 
> Once upon a time, the sync command would return before the actual
> syscall where completed. Doing 3 times sync gave you a very high
> probability that the first one indeed completed all its writes.
> 
> But it was already false in SunOS 4 (~1990)

Even if that's true, running them all in the same command as Roberto
shows would not give you any benefit.

You'd need to physically *type* the command and press Enter three times
to get any "protection".  And even then, it's really just the extra
time that it takes to type those commands out.  You'd get the same
"protection" by simply waiting 10 seconds (or whatever's appropriate)
after running sync once.

[toc] | [prev] | [next] | [standalone]


#272404

FromAlain D D Williams <addw@phcomp.co.uk>
Date2024-08-21 14:20 +0200
Message-ID<JdStH-6yCo-1@gated-at.bofh.it>
In reply to#272390
On Wed, Aug 21, 2024 at 07:38:29AM -0400, Greg Wooledge wrote:

> Even if that's true, running them all in the same command as Roberto
> shows would not give you any benefit.

In early Unix sync *did* return immediately after scheduling a buffer flush.

> You'd need to physically *type* the command and press Enter three times
> to get any "protection".  And even then, it's really just the extra
> time that it takes to type those commands out.

This is exactly what I was taught to do in the 1980s and the reason was to
cause delay before typing ^p.

> You'd get the same "protection" by simply waiting 10 seconds (or whatever's
> appropriate) after running sync once.

Remembering to wait is much harder than remembering to type sync on 3 lines -
especially late at night.

-- 
Alain Williams
Linux/GNU Consultant - Mail systems, Web sites, Networking, Programmer, IT Lecturer.
+44 (0) 787 668 0256  https://www.phcomp.co.uk/
Parliament Hill Computers. Registration Information: https://www.phcomp.co.uk/Contact.html
#include <std_disclaimer.h>

[toc] | [prev] | [next] | [standalone]


#272396

FromAlain D D Williams <addw@phcomp.co.uk>
Date2024-08-21 13:50 +0200
Message-ID<JdS0F-6ydm-9@gated-at.bofh.it>
In reply to#272383
On Wed, Aug 21, 2024 at 08:39:37AM +0200, Erwan David wrote:
> On Wed, Aug 21, 2024 at 02:18:44AM CEST, Greg Wooledge <greg@wooledge.org> said:
> > On Tue, Aug 20, 2024 at 20:04:11 -0400, Roberto C. Sánchez wrote:
> > > sync && sync && sync && swapoff
> > > 
> > > I couldn't tell why I have sync 3 times, but I know that it's how I've
> > > called swapoff since as far back as I can remember.
> > 
> > Cargo cult.  It was never useful to the best of my knowledge.
> 
> Once upon a time, the sync command would return before the actual
> syscall where completed. Doing 3 times sync gave you a very high
> probability that the first one indeed completed all its writes.

I do remember one smart alec typing the following, he did not realise that
typing on separate lines was to slow him down:

    sync;sync;sync
    ^P
    H

^P on a PDP-11 console made it enter console state and 'H' then halted the
processor.

-- 
Alain Williams
Linux/GNU Consultant - Mail systems, Web sites, Networking, Programmer, IT Lecturer.
+44 (0) 787 668 0256  https://www.phcomp.co.uk/
Parliament Hill Computers. Registration Information: https://www.phcomp.co.uk/Contact.html
#include <std_disclaimer.h>

[toc] | [prev] | [next] | [standalone]


#272395

FromRoberto C. Sánchez <roberto@debian.org>
Date2024-08-21 13:50 +0200
Message-ID<JdS0F-6ydm-5@gated-at.bofh.it>
In reply to#272379
On Tue, Aug 20, 2024 at 08:18:44PM -0400, Greg Wooledge wrote:
> On Tue, Aug 20, 2024 at 20:04:11 -0400, Roberto C. Sánchez wrote:
> > sync && sync && sync && swapoff
> > 
> > I couldn't tell why I have sync 3 times, but I know that it's how I've
> > called swapoff since as far back as I can remember.
> 
> Cargo cult.  It was never useful to the best of my knowledge.
> 
Yeah, that is not at all unspurising to me.

It seemed somewhat odd and since I couldn't, even after wracking my
brain, come up with a source or even a vaguely plausible reason for that
particular incantation, I figured it either wasn't doing what I thought
I what was doing or (more likley) it was doing something but essentially
as a side-effect.

Regards,

-Roberto

-- 
Roberto C. Sánchez

[toc] | [prev] | [next] | [standalone]


#272415

FromStefan Monnier <monnier@iro.umontreal.ca>
Date2024-08-21 15:40 +0200
Message-ID<JdTJ7-6zhv-9@gated-at.bofh.it>
In reply to#272376
> Just adding a “sleep” is no proper fix anyway, so the question
> is, how to wait in a shell script until the swap device is
> *really* swapoff’d when the syscall returns too early, and
> (someone from the Linux kernel maintainers reading this?) should
> I report the latter as a bug against the kernel?

I'd file a bug report against the `mount` package (the one that
provides `swapoff`).


        Stefan

[toc] | [prev] | [next] | [standalone]


#272420

FromFranco Martelli <martellif67@gmail.com>
Date2024-08-21 16:00 +0200
Message-ID<JdU2t-6zom-11@gated-at.bofh.it>
In reply to#272376
On 21/08/24 at 01:34, Thorsten Glaser wrote:
> (Please d̲o̲ Cc me on replies, I don’t subscribe to this list. Thanks!)

<snip>
> 
> Just adding a “sleep” is no proper fix anyway, so the question
> is, how to wait in a shell script until the swap device is
> *really* swapoff’d when the syscall returns too early, and
> (someone from the Linux kernel maintainers reading this?) should
> I report the latter as a bug against the kernel?

I don't think this as a kernel bug, many stuffs have timeout on all 
OSes, however you can stop the execution flow of a script using an 
endless loop then interrupt it when a condition is satisfied, e.g.:

while true
do
         /usr/bin/grep lv-swp1 /proc/swaps >/dev/null 2>&1
         [ $? -ne 0 ] && break
	/usr/bin/sleep 1
done

HTH

P.S.
To other readers, the OP asked to Cc to him when replying, see above
-- 
Franco Martelli

[toc] | [prev] | [next] | [standalone]


#272441

FromThorsten Glaser <tg@debian.org>
Date2024-08-21 23:20 +0200
Message-ID<Je0Uh-6DJp-1@gated-at.bofh.it>
In reply to#272420
(Please Cc me on replies.)

Franco Martelli dixit:

> interrupt it when a condition is satisfied, e.g.:
>
> while true
> do
>        /usr/bin/grep lv-swp1 /proc/swaps >/dev/null 2>&1

Not the right condition though… it’s absent there but still in use.
I am looking for the right thing to check…

bye,
//mirabilos
-- 
15:41⎜<Lo-lan-do:#fusionforge> Somebody write a testsuite for helloworld :-)

[toc] | [prev] | [next] | [standalone]


#272456

FromStefan Monnier <monnier@iro.umontreal.ca>
Date2024-08-22 14:50 +0200
Message-ID<Jefqh-6MWk-1@gated-at.bofh.it>
In reply to#272441
> Not the right condition though… it’s absent there but still in use.
> I am looking for the right thing to check…

How 'bout checking the success of `cryptdisks_stop`?


        Stefan

[toc] | [prev] | [next] | [standalone]


#272462

FromMike Castle <dalgoda+debian@gmail.com>
Date2024-08-22 17:50 +0200
Message-ID<Jeieu-6OGL-23@gated-at.bofh.it>
In reply to#272456
On Thu, Aug 22, 2024 at 5:45 AM Stefan Monnier <monnier@iro.umontreal.ca> wrote:
> How 'bout checking the success of `cryptdisks_stop`?

Does cryptdisks have the ability to display what is in use at the
moment?  Maybe polling that before executing the stop?

I suspect that the race is that, when the the swapoff() syscall
returns, the kernel has indeed moved all of the content off, so that
part is fine... but it has not yet released whatever kind of resources
is has on the backing store (akin to an open file handle).  Or the
cryptdisks stack itself hasn't fully processed the notification.

Ideally the stop command would have a flag that says 'wait for any
pending changes to happen', but short of that, some sort of status
than can be polled with a sleep between it might be a bit more formal.

You could then control the timeout by looping no more than N times,
then failing a bit more gracefully than it is now.

mrc

[toc] | [prev] | [next] | [standalone]


#272464

FromThorsten Glaser <tg@debian.org>
Date2024-08-22 19:30 +0200
Message-ID<JejNf-6PJA-3@gated-at.bofh.it>
In reply to#272462
Mike Castle dixit:

>Does cryptdisks have the ability to display what is in use at the
>moment?  Maybe polling that before executing the stop?

That’s what I would like to ask and why I sent this eMail.

>I suspect that the race is that, when the the swapoff() syscall
>returns, the kernel has indeed moved all of the content off, so that
>part is fine... but it has not yet released whatever kind of resources
>is has on the backing store (akin to an open file handle).

My guess as well.

>You could then control the timeout by looping no more than N times,
>then failing a bit more gracefully than it is now.

That’d be a method of last resort, same category as the sleep,
except even worse. I’d like to find a way to prevent that.

Thanks,
//mirabilos
-- 
<cnuke> den AGP stecker anfeilen, damit er in den slot aufm 440BX board passt…
oder netzteile, an die man auch den monitor angeschlossen hat und die dann für
ein elektrisch aufgeladenes gehäuse gesorgt haben […] für lacher gut auf jeder
LAN party │ <nvb> damals, als der pizzateig noch auf dem monior "gegangen" ist

[toc] | [prev] | [next] | [standalone]


#272466

From<tomas@tuxteam.de>
Date2024-08-22 21:40 +0200
Message-ID<JelP3-6QTL-1@gated-at.bofh.it>
In reply to#272464

[Multipart message — attachments visible in raw view] — view raw

On Thu, Aug 22, 2024 at 01:45:06PM -0500, David Wright wrote:
> On Thu 22 Aug 2024 at 17:21:04 (+0000), Thorsten Glaser wrote:
> > Mike Castle dixit:

[...]

> > >I suspect that the race is that, when the the swapoff() syscall
> > >returns, the kernel has indeed moved all of the content off, so that
> > >part is fine... but it has not yet released whatever kind of resources
> > >is has on the backing store (akin to an open file handle).
> > 
> > My guess as well.
> 
> I'm not convinced. Finding out what needs copying back and locating
> somewhere to put it is AIUI a slow process.

Actually, thinking about it: if the system hasn't enough discardable
RAM, the process might take arbitrarily long, no?

Cheers
-- 
t

[toc] | [prev] | [next] | [standalone]


#272473

FromMike Castle <dalgoda+debian@gmail.com>
Date2024-08-23 05:40 +0200
Message-ID<Jetjz-6Wkr-1@gated-at.bofh.it>
In reply to#272466
On Thu, Aug 22, 2024 at 4:31 PM David Wright <deblis@lionunicorn.co.uk> wrote:
> Irrespective of the time taken, that could trigger the OOM killer,
> couldn't it. Very risky, unless you're using two swaps as mentioned.

I was actually surprised to see this happen in a test right now.  I
*thought* that swapoff() would fail if reduce the available memory to
below current usage.

But indeed, the OOM Killer not only killed my test program, it took
out the swapoff command for good measure!

mrc

[toc] | [prev] | [next] | [standalone]


#272469

FromMike Castle <dalgoda+debian@gmail.com>
Date2024-08-23 04:50 +0200
Message-ID<Jesxb-6VOX-1@gated-at.bofh.it>
In reply to#272464
On Thu, Aug 22, 2024 at 11:45 AM David Wright <deblis@lionunicorn.co.uk> wrote:

> I'm not convinced. Finding out what needs copying back and locating
> somewhere to put it is AIUI a slow process. What's much faster is
> when processes themselves demand something be paged back in from
> swap. I think there are "tricks" available to cause that to occur,
> thus speeding up swapoff.

Exactly.  In my experience, running swapoff(8) _will_ take a long time
if the swap area has a lot of content.  It will block until everything
is moved out.

What I think we are seeing here is that phase has finished, but the
kernel has not yet notified the backing store that it is no longer
used.  Though that is just a WAG.

mrc

[toc] | [prev] | [next] | [standalone]


#272472

FromThorsten Glaser <tg@debian.org>
Date2024-08-23 05:20 +0200
Message-ID<Jet0d-6WdY-3@gated-at.bofh.it>
In reply to#272469
Mike Castle dixit:

>Exactly.  In my experience, running swapoff(8) _will_ take a long time
>if the swap area has a lot of content.

Yes.

>It will block until everything is moved out.

Unfortunately not. It will block until *almost* everything is
moved out. I think what we’re seeing is that the request to move
out the last bits was sent but is processed async, or something.

I think at this point, perhaps the kernel team has an idea.

bye,
//mirabilos
PS: please keep Cc’ing me on replies, thanks!
-- 
<hecker> cool ein Ada Lovelace Google-Doodle. aber zum 197. Geburtstag? Hätten
die nicht noch 3 Jahre warten können? <mirabilos> bis dahin gibts google nicht
mehr <hecker> ja, könnte man meinen. wahrscheinlich ist der angekündigte welt-
untergang aus dem maya-kalender die globale abschaltung von google ☺ und darum
müssen die die doodles vorher noch raushauen

[toc] | [prev] | [next] | [standalone]


#272482

FromPierre-Elliott Bécue <peb@debian.org>
Date2024-08-23 13:10 +0200
Message-ID<JeAl4-713h-1@gated-at.bofh.it>
In reply to#272376

[Multipart message — attachments visible in raw view] — view raw

Hey,

Thorsten Glaser <tg@debian.org> wrote on 21/08/2024 at 01:34:32+0200:

> (Please d̲o̲ Cc me on replies, I don’t subscribe to this list. Thanks!)
>
> Hi,
>
> this is a bit curious problem:
>
> I have a setup with swap devices on dmcrypt:
>
> $ cat /etc/crypttab
> # <target name> <source device>         <key file>      <options>
> crtpv           LABEL=fooclvm           none            discard,luks,initramfs
> cswp1           /dev/vg-foo/lv-swp1     /dev/random     discard,cipher=aes-xts-plain64,size=256,plain,swap
> cswp2           /dev/vg-foo/lv-swp2     /dev/random     discard,cipher=aes-xts-plain64,size=256,plain,swap
>
> In a cronjob, I basically do swapoff && cryptdisks_stop && \
> cryptdisks_start && swapon for both swaps individually to throw away
> the old encryption key regularily (but not too frequently).

Ooc, what do you expect to actually gain from this setup?

Apart from that, I had read that discard does a bad job and prople
should use a fstrim timer and drop discard options from mount points.

Bests,

-- 
PEB

[toc] | [prev] | [next] | [standalone]


#272502

FromThorsten Glaser <tg@debian.org>
Date2024-08-23 22:20 +0200
Message-ID<JeIVj-76bl-5@gated-at.bofh.it>
In reply to#272482
Pierre-Elliott Bécue dixit:

>> In a cronjob, I basically do swapoff && cryptdisks_stop && \
>> cryptdisks_start && swapon for both swaps individually to throw away
>> the old encryption key regularily (but not too frequently).
>
>Ooc, what do you expect to actually gain from this setup?

Encryption key rotation. Pages encrypted with the old key
are no longer readable afterwards. This is for long-running
VMs, on hoster infra, mostly (so the hoster could snapshot
the storage any time (ok, they could also snapshot the RAM,
but…)).

This is to get a bit closer to swapencrypt on BSD, which
uses separate keys for each page or set of pages, AIUI.

bye,
//mirabilos
-- 
Solange man keine schmutzigen Tricks macht, und ich meine *wirklich*
schmutzige Tricks, wie bei einer doppelt verketteten Liste beide
Pointer XORen und in nur einem Word speichern, funktioniert Boehm ganz
hervorragend.		-- Andreas Bogk über boehm-gc in d.a.s.r

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web