Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1704173 > unrolled thread

[PATCH 0/2] mm,fork,security: introduce MADV_WIPEONFORK

Started byriel@redhat.com
First post2017-08-04 21:10 +0200
Last post2017-08-05 17:30 +0200
Articles 3 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH 0/2] mm,fork,security: introduce MADV_WIPEONFORK riel@redhat.com - 2017-08-04 21:10 +0200
    Re: [PATCH 0/2] mm,fork,security: introduce MADV_WIPEONFORK "Kirill A. Shutemov" <kirill@shutemov.name> - 2017-08-05 01:50 +0200
      Re: [PATCH 0/2] mm,fork,security: introduce MADV_WIPEONFORK Rik van Riel <riel@redhat.com> - 2017-08-05 17:30 +0200

#1704173 — [PATCH 0/2] mm,fork,security: introduce MADV_WIPEONFORK

Fromriel@redhat.com
Date2017-08-04 21:10 +0200
Subject[PATCH 0/2] mm,fork,security: introduce MADV_WIPEONFORK
Message-ID<uaQ8W-2CK-17@gated-at.bofh.it>
[resend because half the recipients got dropped due to IPv6 firewall issues]

Introduce MADV_WIPEONFORK semantics, which result in a VMA being
empty in the child process after fork. This differs from MADV_DONTFORK
in one important way.

If a child process accesses memory that was MADV_WIPEONFORK, it
will get zeroes. The address ranges are still valid, they are just empty.

If a child process accesses memory that was MADV_DONTFORK, it will
get a segmentation fault, since those address ranges are no longer
valid in the child after fork.

Since MADV_DONTFORK also seems to be used to allow very large
programs to fork in systems with strict memory overcommit restrictions,
changing the semantics of MADV_DONTFORK might break existing programs.

The use case is libraries that store or cache information, and
want to know that they need to regenerate it in the child process
after fork.

Examples of this would be:
- systemd/pulseaudio API checks (fail after fork)
  (replacing a getpid check, which is too slow without a PID cache)
- PKCS#11 API reinitialization check (mandated by specification)
- glibc's upcoming PRNG (reseed after fork)
- OpenSSL PRNG (reseed after fork)

The security benefits of a forking server having a re-inialized
PRNG in every child process are pretty obvious. However, due to
libraries having all kinds of internal state, and programs getting
compiled with many different versions of each library, it is
unreasonable to expect calling programs to re-initialize everything
manually after fork.

A further complication is the proliferation of clone flags,
programs bypassing glibc's functions to call clone directly,
and programs calling unshare, causing the glibc pthread_atfork
hook to not get called.

It would be better to have the kernel take care of this automatically.

This is similar to the OpenBSD minherit syscall with MAP_INHERIT_ZERO:

    https://man.openbsd.org/minherit.2

[toc] | [next] | [standalone]


#1704413

From"Kirill A. Shutemov" <kirill@shutemov.name>
Date2017-08-05 01:50 +0200
Message-ID<uaUvW-5hz-63@gated-at.bofh.it>
In reply to#1704173
On Fri, Aug 04, 2017 at 03:07:28PM -0400, riel@redhat.com wrote:
> [resend because half the recipients got dropped due to IPv6 firewall issues]
> 
> Introduce MADV_WIPEONFORK semantics, which result in a VMA being
> empty in the child process after fork. This differs from MADV_DONTFORK
> in one important way.
> 
> If a child process accesses memory that was MADV_WIPEONFORK, it
> will get zeroes. The address ranges are still valid, they are just empty.

I feel like we are repeating mistake we made with MADV_DONTNEED.

MADV_WIPEONFORK would require a specific action from kernel, ignoring
the /advise/ would likely lead to application misbehaviour.

Is it something we really want to see from madvise()?

-- 
 Kirill A. Shutemov

[toc] | [prev] | [next] | [standalone]


#1704659

FromRik van Riel <riel@redhat.com>
Date2017-08-05 17:30 +0200
Message-ID<ub9bA-6KA-7@gated-at.bofh.it>
In reply to#1704413
On Sat, 2017-08-05 at 02:44 +0300, Kirill A. Shutemov wrote:
> On Fri, Aug 04, 2017 at 03:07:28PM -0400, riel@redhat.com wrote:
> > [resend because half the recipients got dropped due to IPv6
> > firewall issues]
> > 
> > Introduce MADV_WIPEONFORK semantics, which result in a VMA being
> > empty in the child process after fork. This differs from
> > MADV_DONTFORK
> > in one important way.
> > 
> > If a child process accesses memory that was MADV_WIPEONFORK, it
> > will get zeroes. The address ranges are still valid, they are just
> > empty.
> 
> I feel like we are repeating mistake we made with MADV_DONTNEED.
> 
> MADV_WIPEONFORK would require a specific action from kernel, ignoring
> the /advise/ would likely lead to application misbehaviour.
> 
> Is it something we really want to see from madvise()?

We already have various mandatory madvise behaviors in Linux,
including MADV_REMOVE, MADV_DONTFORK, and MADV_DONTDUMP.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web