Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1262009 > unrolled thread

[PATCH v2 00/13] MADV_FREE support

Started byMinchan Kim <minchan@kernel.org>
First post2015-11-04 02:30 +0100
Last post2015-11-05 02:40 +0100
Articles 15 on this page of 35 — 5 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH v2 00/13] MADV_FREE support Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
    [PATCH v2 11/13] arm: add pmd_mkclean for THP Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
    [PATCH v2 13/13] mm: don't split THP page when syscall is called Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
    [PATCH v2 04/13] mm: free swp_entry in madvise_free Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
    [PATCH v2 02/13] mm: define MADV_FREE for some arches Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
    [PATCH v2 12/13] arm64: add pmd_mkclean for THP Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
    [PATCH v2 08/13] x86: add pmd_[dirty|mkclean] for THP Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
    [PATCH v2 07/13] mm: mark stable page dirty in KSM Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
    [PATCH v2 05/13] mm: move lazily freed pages to inactive list Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
    [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-04 02:40 +0100
      Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2015-11-04 03:20 +0100
        Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 00:40 +0100
          Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2015-11-05 04:50 +0100
      Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2015-11-04 03:40 +0100
        Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 00:50 +0100
      Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Andy Lutomirski <luto@amacapital.net> - 2015-11-04 04:50 +0100
        Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 07:00 +0100
          Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 07:00 +0100
            Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 07:10 +0100
          Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Andy Lutomirski <luto@amacapital.net> - 2015-11-04 19:30 +0100
            Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 23:10 +0100
              Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Shaohua Li <shli@kernel.org> - 2015-11-05 19:20 +0100
                Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-05 21:20 +0100
                  Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-05 21:20 +0100
        Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 01:20 +0100
          Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Andy Lutomirski <luto@amacapital.net> - 2015-11-05 01:50 +0100
            Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 02:00 +0100
              Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Andy Lutomirski <luto@amacapital.net> - 2015-11-05 02:40 +0100
                Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 02:50 +0100
      Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Shaohua Li <shli@kernel.org> - 2015-11-04 21:10 +0100
        Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 22:20 +0100
          Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 22:30 +0100
        Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Andy Lutomirski <luto@amacapital.net> - 2015-11-04 22:50 +0100
        Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 02:40 +0100
          Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 02:40 +0100

Page 2 of 2 — ← Prev page 1 [2]


#1262632 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromDaniel Micay <danielmicay@gmail.com>
Date2015-11-04 23:10 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrepJ-6lw-29@gated-at.bofh.it>
In reply to#1262487

[Multipart message — attachments visible in raw view] — view raw

> With enough pages at once, though, munmap would be fine, too.

That implies lots of page faults and zeroing though. The zeroing alone
is a major performance issue.

There are separate issues with munmap since it ends up resulting in a
lot more virtual memory fragmentation. It would help if the kernel used
first-best-fit for mmap instead of the current naive algorithm (bonus:
O(log n) worst-case, not O(n)). Since allocators like jemalloc and
PartitionAlloc want 2M aligned spans, mixing them with other allocators
can also accelerate the VM fragmentation caused by the dumb mmap
algorithm (i.e. they make a 2M aligned mapping, some other mmap user
does 4k, now there's a nearly 2M gap when the next 2M region is made and
the kernel keeps going rather than reusing it). Anyway, that's a totally
separate issue from this. Just felt like complaining :).

> Maybe what's really needed is a MADV_FREE variant that takes an iovec.
> On an all-cores multithreaded mm, the TLB shootdown broadcast takes
> thousands of cycles on each core more or less regardless of how much
> of the TLB gets zapped.

That would work very well. The allocator ends up having a sequence of
dirty spans that it needs to purge in one go. As long as purging is
fairly spread out, the cost of a single TLB shootdown isn't that bad. It
is extremely bad if it needs to do it over and over to purge a bunch of
ranges, which can happen if the memory has ended up being very, very
fragmentated despite the efforts to compact it (depends on what the
application ends up doing).

[toc] | [prev] | [next] | [standalone]


#1263465 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromShaohua Li <shli@kernel.org>
Date2015-11-05 19:20 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrxiF-1JR-1@gated-at.bofh.it>
In reply to#1262632
On Wed, Nov 04, 2015 at 05:05:47PM -0500, Daniel Micay wrote:
> > With enough pages at once, though, munmap would be fine, too.
> 
> That implies lots of page faults and zeroing though. The zeroing alone
> is a major performance issue.
> 
> There are separate issues with munmap since it ends up resulting in a
> lot more virtual memory fragmentation. It would help if the kernel used
> first-best-fit for mmap instead of the current naive algorithm (bonus:
> O(log n) worst-case, not O(n)). Since allocators like jemalloc and
> PartitionAlloc want 2M aligned spans, mixing them with other allocators
> can also accelerate the VM fragmentation caused by the dumb mmap
> algorithm (i.e. they make a 2M aligned mapping, some other mmap user
> does 4k, now there's a nearly 2M gap when the next 2M region is made and
> the kernel keeps going rather than reusing it). Anyway, that's a totally
> separate issue from this. Just felt like complaining :).
> 
> > Maybe what's really needed is a MADV_FREE variant that takes an iovec.
> > On an all-cores multithreaded mm, the TLB shootdown broadcast takes
> > thousands of cycles on each core more or less regardless of how much
> > of the TLB gets zapped.
> 
> That would work very well. The allocator ends up having a sequence of
> dirty spans that it needs to purge in one go. As long as purging is
> fairly spread out, the cost of a single TLB shootdown isn't that bad. It
> is extremely bad if it needs to do it over and over to purge a bunch of
> ranges, which can happen if the memory has ended up being very, very
> fragmentated despite the efforts to compact it (depends on what the
> application ends up doing).

I posted a patch doing exactly iovec madvise. Doesn't support MADV_FREE yet
though, but should be easy to do it.

http://marc.info/?l=linux-mm&m=144615663522661&w=2
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1263528 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromDaniel Micay <danielmicay@gmail.com>
Date2015-11-05 21:20 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrzaN-2TZ-3@gated-at.bofh.it>
In reply to#1263465

[Multipart message — attachments visible in raw view] — view raw

> I posted a patch doing exactly iovec madvise. Doesn't support MADV_FREE yet
> though, but should be easy to do it.
> 
> http://marc.info/?l=linux-mm&m=144615663522661&w=2

I think that would be a great way to deal with this. It keeps the nice
property of still being able to drop pages in allocations that have been
handed out but not yet touched. The allocator just needs to be designed
to do lots of purging in one go (i.e. something like an 8:1 active:clean
ratio triggers purging and it goes all the way to 16:1).

[toc] | [prev] | [next] | [standalone]


#1263529 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromDaniel Micay <danielmicay@gmail.com>
Date2015-11-05 21:20 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrzaN-2TZ-1@gated-at.bofh.it>
In reply to#1263528

[Multipart message — attachments visible in raw view] — view raw

> active:clean

active:dirty*, sigh.

[toc] | [prev] | [next] | [standalone]


#1262759 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromMinchan Kim <minchan@kernel.org>
Date2015-11-05 01:20 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrgrv-7Cx-1@gated-at.bofh.it>
In reply to#1262060
On Tue, Nov 03, 2015 at 07:41:35PM -0800, Andy Lutomirski wrote:
> On Nov 3, 2015 5:30 PM, "Minchan Kim" <minchan@kernel.org> wrote:
> >
> > Linux doesn't have an ability to free pages lazy while other OS already
> > have been supported that named by madvise(MADV_FREE).
> >
> > The gain is clear that kernel can discard freed pages rather than swapping
> > out or OOM if memory pressure happens.
> >
> > Without memory pressure, freed pages would be reused by userspace without
> > another additional overhead(ex, page fault + allocation + zeroing).
> >
> 
> [...]
> 
> >
> > How it works:
> >
> > When madvise syscall is called, VM clears dirty bit of ptes of the range.
> > If memory pressure happens, VM checks dirty bit of page table and if it
> > found still "clean", it means it's a "lazyfree pages" so VM could discard
> > the page instead of swapping out.  Once there was store operation for the
> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out
> > the page instead of discarding.
> 
> What happens if you MADV_FREE something that's MAP_SHARED or isn't
> ordinary anonymous memory?  There's a long history of MADV_DONTNEED on
> such mappings causing exploitable problems, and I think it would be
> nice if MADV_FREE were obviously safe.

It filter out VM_LOCKED|VM_HUGETLB|VM_PFNMAP and file-backed vma and MAP_SHARED
with vma_is_anonymous.

> 
> Does this set the write protect bit?

No.

> 
> What happens on architectures without hardware dirty tracking?  For
> that matter, even on architecture with hardware dirty tracking, what
> happens in multithreaded processes that have the dirty TLB state
> cached in a different CPU's TLB?
> 
> Using the dirty bit for these semantics scares me.  This API creates a
> page that can have visible nonzero contents and then can
> asynchronously and magically zero itself thereafter.  That makes me
> nervous.  Could we use the accessed bit instead?  Then the observable

Access bit is used by aging algorithm for reclaim. In addition,
we have supported clear_refs feacture.
IOW, it could be reset anytime so it's hard to use marker for
lazy freeing at the moment.

> semantics would be equivalent to having MADV_FREE either zero the page
> or do nothing, except that it doesn't make up its mind until the next
> read.
> 
> > +                       ptent = pte_mkold(ptent);
> > +                       ptent = pte_mkclean(ptent);
> > +                       set_pte_at(mm, addr, pte, ptent);
> > +                       tlb_remove_tlb_entry(tlb, pte, addr);
> 
> It looks like you are flushing the TLB.  In a multithreaded program,
> that's rather expensive.  Potentially silly question: would it be
> better to just zero the page immediately in a multithreaded program
> and then, when swapping out, check the page is zeroed and, if so, skip
> swapping it out?  That could be done without forcing an IPI.

So, we should monitor all of pages in reclaim patch whether they are
zero or not? It is fatster for allocation side but much slower in
reclaim side. For avoiding that, we should mark something for lazy
freeing page out of page table.

Anyway, it depends on the TLB flush overehead vs memset overhead.
If the hinted range is pretty big and small system(ie, not many core),
memset overhead would't not trivial compared to TLB flush.
Even, some of ARM arches doesn't do IPI to TLB flush so the overhead
would be cheaper.

I don't want to push more optimization in new syscall from the beginning.
It's an optimization and might come better idea once we hear from the
voice of userland folks. Then, it's not too late.
Let's do step by step.

> 
> > +static int madvise_free_single_vma(struct vm_area_struct *vma,
> > +                       unsigned long start_addr, unsigned long end_addr)
> > +{
> > +       unsigned long start, end;
> > +       struct mm_struct *mm = vma->vm_mm;
> > +       struct mmu_gather tlb;
> > +
> > +       if (vma->vm_flags & (VM_LOCKED|VM_HUGETLB|VM_PFNMAP))
> > +               return -EINVAL;
> > +
> > +       /* MADV_FREE works for only anon vma at the moment */
> > +       if (!vma_is_anonymous(vma))
> > +               return -EINVAL;
> 
> Does anything weird happen if it's shared?

Hmm, you mean MAP_SHARED|MAP_ANONYMOUS?
In that case, vma->vm_ops = &shmem_vm_ops so vma_is anonymous should filter it out.

> 
> > +               if (!PageDirty(page) && (flags & TTU_FREE)) {
> > +                       /* It's a freeable page by MADV_FREE */
> > +                       dec_mm_counter(mm, MM_ANONPAGES);
> > +                       goto discard;
> > +               }
> 
> Does something clear TTU_FREE the next time the page gets marked clean?

Sorry, I don't understand. Could you elaborate it more?

> 
> --Andy
> 
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org.  For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1262764 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromAndy Lutomirski <luto@amacapital.net>
Date2015-11-05 01:50 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrgUx-7M5-11@gated-at.bofh.it>
In reply to#1262759
On Wed, Nov 4, 2015 at 4:13 PM, Minchan Kim <minchan@kernel.org> wrote:
> On Tue, Nov 03, 2015 at 07:41:35PM -0800, Andy Lutomirski wrote:
>> On Nov 3, 2015 5:30 PM, "Minchan Kim" <minchan@kernel.org> wrote:
>> >
>> > Linux doesn't have an ability to free pages lazy while other OS already
>> > have been supported that named by madvise(MADV_FREE).
>> >
>> > The gain is clear that kernel can discard freed pages rather than swapping
>> > out or OOM if memory pressure happens.
>> >
>> > Without memory pressure, freed pages would be reused by userspace without
>> > another additional overhead(ex, page fault + allocation + zeroing).
>> >
>>
>> [...]
>>
>> >
>> > How it works:
>> >
>> > When madvise syscall is called, VM clears dirty bit of ptes of the range.
>> > If memory pressure happens, VM checks dirty bit of page table and if it
>> > found still "clean", it means it's a "lazyfree pages" so VM could discard
>> > the page instead of swapping out.  Once there was store operation for the
>> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out
>> > the page instead of discarding.
>>
>> What happens if you MADV_FREE something that's MAP_SHARED or isn't
>> ordinary anonymous memory?  There's a long history of MADV_DONTNEED on
>> such mappings causing exploitable problems, and I think it would be
>> nice if MADV_FREE were obviously safe.
>
> It filter out VM_LOCKED|VM_HUGETLB|VM_PFNMAP and file-backed vma and MAP_SHARED
> with vma_is_anonymous.
>
>>
>> Does this set the write protect bit?
>
> No.
>
>>
>> What happens on architectures without hardware dirty tracking?  For
>> that matter, even on architecture with hardware dirty tracking, what
>> happens in multithreaded processes that have the dirty TLB state
>> cached in a different CPU's TLB?
>>
>> Using the dirty bit for these semantics scares me.  This API creates a
>> page that can have visible nonzero contents and then can
>> asynchronously and magically zero itself thereafter.  That makes me
>> nervous.  Could we use the accessed bit instead?  Then the observable
>
> Access bit is used by aging algorithm for reclaim. In addition,
> we have supported clear_refs feacture.
> IOW, it could be reset anytime so it's hard to use marker for
> lazy freeing at the moment.
>

That's unfortunate.  I think that the ABI would be much nicer if it
used the accessed bit.

In any case, shouldn't the aging algorithm be irrelevant here?  A
MADV_FREE page that isn't accessed can be discarded, whereas we could
hopefully just say that a MADV_FREE page that is accessed gets moved
to whatever list holds recently accessed pages and also stops being a
candidate for discarding due to MADV_FREE?

>>
>> > +               if (!PageDirty(page) && (flags & TTU_FREE)) {
>> > +                       /* It's a freeable page by MADV_FREE */
>> > +                       dec_mm_counter(mm, MM_ANONPAGES);
>> > +                       goto discard;
>> > +               }
>>
>> Does something clear TTU_FREE the next time the page gets marked clean?
>
> Sorry, I don't understand. Could you elaborate it more?

I don't fully understand how TTU_FREE ends up being set here, but, if
the page is dirtied by user code and then cleaned later by the kernel,
what prevents TTU_FREE from being incorrectly set here?


--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1262785 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromMinchan Kim <minchan@kernel.org>
Date2015-11-05 02:00 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrh4g-7PW-43@gated-at.bofh.it>
In reply to#1262764
On Wed, Nov 04, 2015 at 04:42:37PM -0800, Andy Lutomirski wrote:
> On Wed, Nov 4, 2015 at 4:13 PM, Minchan Kim <minchan@kernel.org> wrote:
> > On Tue, Nov 03, 2015 at 07:41:35PM -0800, Andy Lutomirski wrote:
> >> On Nov 3, 2015 5:30 PM, "Minchan Kim" <minchan@kernel.org> wrote:
> >> >
> >> > Linux doesn't have an ability to free pages lazy while other OS already
> >> > have been supported that named by madvise(MADV_FREE).
> >> >
> >> > The gain is clear that kernel can discard freed pages rather than swapping
> >> > out or OOM if memory pressure happens.
> >> >
> >> > Without memory pressure, freed pages would be reused by userspace without
> >> > another additional overhead(ex, page fault + allocation + zeroing).
> >> >
> >>
> >> [...]
> >>
> >> >
> >> > How it works:
> >> >
> >> > When madvise syscall is called, VM clears dirty bit of ptes of the range.
> >> > If memory pressure happens, VM checks dirty bit of page table and if it
> >> > found still "clean", it means it's a "lazyfree pages" so VM could discard
> >> > the page instead of swapping out.  Once there was store operation for the
> >> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out
> >> > the page instead of discarding.
> >>
> >> What happens if you MADV_FREE something that's MAP_SHARED or isn't
> >> ordinary anonymous memory?  There's a long history of MADV_DONTNEED on
> >> such mappings causing exploitable problems, and I think it would be
> >> nice if MADV_FREE were obviously safe.
> >
> > It filter out VM_LOCKED|VM_HUGETLB|VM_PFNMAP and file-backed vma and MAP_SHARED
> > with vma_is_anonymous.
> >
> >>
> >> Does this set the write protect bit?
> >
> > No.
> >
> >>
> >> What happens on architectures without hardware dirty tracking?  For
> >> that matter, even on architecture with hardware dirty tracking, what
> >> happens in multithreaded processes that have the dirty TLB state
> >> cached in a different CPU's TLB?
> >>
> >> Using the dirty bit for these semantics scares me.  This API creates a
> >> page that can have visible nonzero contents and then can
> >> asynchronously and magically zero itself thereafter.  That makes me
> >> nervous.  Could we use the accessed bit instead?  Then the observable
> >
> > Access bit is used by aging algorithm for reclaim. In addition,
> > we have supported clear_refs feacture.
> > IOW, it could be reset anytime so it's hard to use marker for
> > lazy freeing at the moment.
> >
> 
> That's unfortunate.  I think that the ABI would be much nicer if it
> used the accessed bit.
> 
> In any case, shouldn't the aging algorithm be irrelevant here?  A
> MADV_FREE page that isn't accessed can be discarded, whereas we could
> hopefully just say that a MADV_FREE page that is accessed gets moved
> to whatever list holds recently accessed pages and also stops being a
> candidate for discarding due to MADV_FREE?

I meant if we use access bit as indicator for lazy-freeing page,
we could discard valid page which is never hinted by MADV_FREE but
just doesn't mark access bit in page table by aging algorithm.

> 
> >>
> >> > +               if (!PageDirty(page) && (flags & TTU_FREE)) {
> >> > +                       /* It's a freeable page by MADV_FREE */
> >> > +                       dec_mm_counter(mm, MM_ANONPAGES);
> >> > +                       goto discard;
> >> > +               }
> >>
> >> Does something clear TTU_FREE the next time the page gets marked clean?
> >
> > Sorry, I don't understand. Could you elaborate it more?
> 
> I don't fully understand how TTU_FREE ends up being set here, but, if
> the page is dirtied by user code and then cleaned later by the kernel,
> what prevents TTU_FREE from being incorrectly set here?

Kernel shouldn't make the page clean without writeback(ie, swapout)
if the page has valid data.

> 
> 
> --Andy
> 
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org.  For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1262800 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromAndy Lutomirski <luto@amacapital.net>
Date2015-11-05 02:40 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrhGV-8jJ-13@gated-at.bofh.it>
In reply to#1262785
On Wed, Nov 4, 2015 at 4:56 PM, Minchan Kim <minchan@kernel.org> wrote:
> On Wed, Nov 04, 2015 at 04:42:37PM -0800, Andy Lutomirski wrote:
>> On Wed, Nov 4, 2015 at 4:13 PM, Minchan Kim <minchan@kernel.org> wrote:
>> > On Tue, Nov 03, 2015 at 07:41:35PM -0800, Andy Lutomirski wrote:
>> >> On Nov 3, 2015 5:30 PM, "Minchan Kim" <minchan@kernel.org> wrote:
>> >> >
>> >> > Linux doesn't have an ability to free pages lazy while other OS already
>> >> > have been supported that named by madvise(MADV_FREE).
>> >> >
>> >> > The gain is clear that kernel can discard freed pages rather than swapping
>> >> > out or OOM if memory pressure happens.
>> >> >
>> >> > Without memory pressure, freed pages would be reused by userspace without
>> >> > another additional overhead(ex, page fault + allocation + zeroing).
>> >> >
>> >>
>> >> [...]
>> >>
>> >> >
>> >> > How it works:
>> >> >
>> >> > When madvise syscall is called, VM clears dirty bit of ptes of the range.
>> >> > If memory pressure happens, VM checks dirty bit of page table and if it
>> >> > found still "clean", it means it's a "lazyfree pages" so VM could discard
>> >> > the page instead of swapping out.  Once there was store operation for the
>> >> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out
>> >> > the page instead of discarding.
>> >>
>> >> What happens if you MADV_FREE something that's MAP_SHARED or isn't
>> >> ordinary anonymous memory?  There's a long history of MADV_DONTNEED on
>> >> such mappings causing exploitable problems, and I think it would be
>> >> nice if MADV_FREE were obviously safe.
>> >
>> > It filter out VM_LOCKED|VM_HUGETLB|VM_PFNMAP and file-backed vma and MAP_SHARED
>> > with vma_is_anonymous.
>> >
>> >>
>> >> Does this set the write protect bit?
>> >
>> > No.
>> >
>> >>
>> >> What happens on architectures without hardware dirty tracking?  For
>> >> that matter, even on architecture with hardware dirty tracking, what
>> >> happens in multithreaded processes that have the dirty TLB state
>> >> cached in a different CPU's TLB?
>> >>
>> >> Using the dirty bit for these semantics scares me.  This API creates a
>> >> page that can have visible nonzero contents and then can
>> >> asynchronously and magically zero itself thereafter.  That makes me
>> >> nervous.  Could we use the accessed bit instead?  Then the observable
>> >
>> > Access bit is used by aging algorithm for reclaim. In addition,
>> > we have supported clear_refs feacture.
>> > IOW, it could be reset anytime so it's hard to use marker for
>> > lazy freeing at the moment.
>> >
>>
>> That's unfortunate.  I think that the ABI would be much nicer if it
>> used the accessed bit.
>>
>> In any case, shouldn't the aging algorithm be irrelevant here?  A
>> MADV_FREE page that isn't accessed can be discarded, whereas we could
>> hopefully just say that a MADV_FREE page that is accessed gets moved
>> to whatever list holds recently accessed pages and also stops being a
>> candidate for discarding due to MADV_FREE?
>
> I meant if we use access bit as indicator for lazy-freeing page,
> we could discard valid page which is never hinted by MADV_FREE but
> just doesn't mark access bit in page table by aging algorithm.

Oh, is the rule that the anonymous pages that are clean are discarded
instead of swapped out?  That is, does your patch set detect that an
anonymous page can be discarded if it's clean and that the lack of a
dirty bit is the only indication that the page has been hit with
MADV_FREE?

If so, that seems potentially error prone -- I had assumed that pages
that were swapped in but not written since swap-in would also be
clean, and I don't see how you distinguish them.

--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1262801 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromMinchan Kim <minchan@kernel.org>
Date2015-11-05 02:50 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrhQB-8ni-3@gated-at.bofh.it>
In reply to#1262800
On Wed, Nov 04, 2015 at 05:29:57PM -0800, Andy Lutomirski wrote:
> On Wed, Nov 4, 2015 at 4:56 PM, Minchan Kim <minchan@kernel.org> wrote:
> > On Wed, Nov 04, 2015 at 04:42:37PM -0800, Andy Lutomirski wrote:
> >> On Wed, Nov 4, 2015 at 4:13 PM, Minchan Kim <minchan@kernel.org> wrote:
> >> > On Tue, Nov 03, 2015 at 07:41:35PM -0800, Andy Lutomirski wrote:
> >> >> On Nov 3, 2015 5:30 PM, "Minchan Kim" <minchan@kernel.org> wrote:
> >> >> >
> >> >> > Linux doesn't have an ability to free pages lazy while other OS already
> >> >> > have been supported that named by madvise(MADV_FREE).
> >> >> >
> >> >> > The gain is clear that kernel can discard freed pages rather than swapping
> >> >> > out or OOM if memory pressure happens.
> >> >> >
> >> >> > Without memory pressure, freed pages would be reused by userspace without
> >> >> > another additional overhead(ex, page fault + allocation + zeroing).
> >> >> >
> >> >>
> >> >> [...]
> >> >>
> >> >> >
> >> >> > How it works:
> >> >> >
> >> >> > When madvise syscall is called, VM clears dirty bit of ptes of the range.
> >> >> > If memory pressure happens, VM checks dirty bit of page table and if it
> >> >> > found still "clean", it means it's a "lazyfree pages" so VM could discard
> >> >> > the page instead of swapping out.  Once there was store operation for the
> >> >> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out
> >> >> > the page instead of discarding.
> >> >>
> >> >> What happens if you MADV_FREE something that's MAP_SHARED or isn't
> >> >> ordinary anonymous memory?  There's a long history of MADV_DONTNEED on
> >> >> such mappings causing exploitable problems, and I think it would be
> >> >> nice if MADV_FREE were obviously safe.
> >> >
> >> > It filter out VM_LOCKED|VM_HUGETLB|VM_PFNMAP and file-backed vma and MAP_SHARED
> >> > with vma_is_anonymous.
> >> >
> >> >>
> >> >> Does this set the write protect bit?
> >> >
> >> > No.
> >> >
> >> >>
> >> >> What happens on architectures without hardware dirty tracking?  For
> >> >> that matter, even on architecture with hardware dirty tracking, what
> >> >> happens in multithreaded processes that have the dirty TLB state
> >> >> cached in a different CPU's TLB?
> >> >>
> >> >> Using the dirty bit for these semantics scares me.  This API creates a
> >> >> page that can have visible nonzero contents and then can
> >> >> asynchronously and magically zero itself thereafter.  That makes me
> >> >> nervous.  Could we use the accessed bit instead?  Then the observable
> >> >
> >> > Access bit is used by aging algorithm for reclaim. In addition,
> >> > we have supported clear_refs feacture.
> >> > IOW, it could be reset anytime so it's hard to use marker for
> >> > lazy freeing at the moment.
> >> >
> >>
> >> That's unfortunate.  I think that the ABI would be much nicer if it
> >> used the accessed bit.
> >>
> >> In any case, shouldn't the aging algorithm be irrelevant here?  A
> >> MADV_FREE page that isn't accessed can be discarded, whereas we could
> >> hopefully just say that a MADV_FREE page that is accessed gets moved
> >> to whatever list holds recently accessed pages and also stops being a
> >> candidate for discarding due to MADV_FREE?
> >
> > I meant if we use access bit as indicator for lazy-freeing page,
> > we could discard valid page which is never hinted by MADV_FREE but
> > just doesn't mark access bit in page table by aging algorithm.
> 
> Oh, is the rule that the anonymous pages that are clean are discarded
> instead of swapped out?  That is, does your patch set detect that an

The page swapped-in after swapped-out has clean pte and swap device
has valid data if the page isn't touch so VM discards the page rather
than swapout. Of course, pte should point out the swap slot.
If VM decide to remove the page from swap slot, it should be marked
PG_dirty.

> anonymous page can be discarded if it's clean and that the lack of a
> dirty bit is the only indication that the page has been hit with
> MADV_FREE?

No dirty bit, exactly speaking, PG_Dirty
because the page I mentioned above has clean pte but will have PG_dirty.

> 
> If so, that seems potentially error prone -- I had assumed that pages
> that were swapped in but not written since swap-in would also be
> clean, and I don't see how you distinguish them.

I hope above will answer.
> 
> --Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1262558 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromShaohua Li <shli@kernel.org>
Date2015-11-04 21:10 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrcxA-59d-7@gated-at.bofh.it>
In reply to#1262021
On Wed, Nov 04, 2015 at 10:25:55AM +0900, Minchan Kim wrote:
> Linux doesn't have an ability to free pages lazy while other OS already
> have been supported that named by madvise(MADV_FREE).
> 
> The gain is clear that kernel can discard freed pages rather than swapping
> out or OOM if memory pressure happens.
> 
> Without memory pressure, freed pages would be reused by userspace without
> another additional overhead(ex, page fault + allocation + zeroing).
> 
> Jason Evans said:
> 
> : Facebook has been using MAP_UNINITIALIZED
> : (https://lkml.org/lkml/2012/1/18/308) in some of its applications for
> : several years, but there are operational costs to maintaining this
> : out-of-tree in our kernel and in jemalloc, and we are anxious to retire it
> : in favor of MADV_FREE.  When we first enabled MAP_UNINITIALIZED it
> : increased throughput for much of our workload by ~5%, and although the
> : benefit has decreased using newer hardware and kernels, there is still
> : enough benefit that we cannot reasonably retire it without a replacement.
> :
> : Aside from Facebook operations, there are numerous broadly used
> : applications that would benefit from MADV_FREE.  The ones that immediately
> : come to mind are redis, varnish, and MariaDB.  I don't have much insight
> : into Android internals and development process, but I would hope to see
> : MADV_FREE support eventually end up there as well to benefit applications
> : linked with the integrated jemalloc.
> :
> : jemalloc will use MADV_FREE once it becomes available in the Linux kernel.
> : In fact, jemalloc already uses MADV_FREE or equivalent everywhere it's
> : available: *BSD, OS X, Windows, and Solaris -- every platform except Linux
> : (and AIX, but I'm not sure it even compiles on AIX).  The lack of
> : MADV_FREE on Linux forced me down a long series of increasingly
> : sophisticated heuristics for madvise() volume reduction, and even so this
> : remains a common performance issue for people using jemalloc on Linux.
> : Please integrate MADV_FREE; many people will benefit substantially.
> 
> How it works:
> 
> When madvise syscall is called, VM clears dirty bit of ptes of the range.
> If memory pressure happens, VM checks dirty bit of page table and if it
> found still "clean", it means it's a "lazyfree pages" so VM could discard
> the page instead of swapping out.  Once there was store operation for the
> page before VM peek a page to reclaim, dirty bit is set so VM can swap out
> the page instead of discarding.
> 
> Firstly, heavy users would be general allocators(ex, jemalloc, tcmalloc
> and hope glibc supports it) and jemalloc/tcmalloc already have supported
> the feature for other OS(ex, FreeBSD)
> 
> barrios@blaptop:~/benchmark/ebizzy$ lscpu
> Architecture:          x86_64
> CPU op-mode(s):        32-bit, 64-bit
> Byte Order:            Little Endian
> CPU(s):                12
> On-line CPU(s) list:   0-11
> Thread(s) per core:    1
> Core(s) per socket:    1
> Socket(s):             12
> NUMA node(s):          1
> Vendor ID:             GenuineIntel
> CPU family:            6
> Model:                 2
> Stepping:              3
> CPU MHz:               3200.185
> BogoMIPS:              6400.53
> Virtualization:        VT-x
> Hypervisor vendor:     KVM
> Virtualization type:   full
> L1d cache:             32K
> L1i cache:             32K
> L2 cache:              4096K
> NUMA node0 CPU(s):     0-11
> ebizzy benchmark(./ebizzy -S 10 -n 512)
> 
> Higher avg is better.
> 
>  vanilla-jemalloc		MADV_free-jemalloc
> 
> 1 thread
> records: 10			    records: 10
> avg:	2961.90			    avg:   12069.70
> std:	  71.96(2.43%)		    std:     186.68(1.55%)
> max:	3070.00			    max:   12385.00
> min:	2796.00			    min:   11746.00
> 
> 2 thread
> records: 10			    records: 10
> avg:	5020.00			    avg:   17827.00
> std:	 264.87(5.28%)		    std:     358.52(2.01%)
> max:	5244.00			    max:   18760.00
> min:	4251.00			    min:   17382.00
> 
> 4 thread
> records: 10			    records: 10
> avg:	8988.80			    avg:   27930.80
> std:	1175.33(13.08%)		    std:    3317.33(11.88%)
> max:	9508.00			    max:   30879.00
> min:	5477.00			    min:   21024.00
> 
> 8 thread
> records: 10			    records: 10
> avg:   13036.50			    avg:   33739.40
> std:	 170.67(1.31%)		    std:    5146.22(15.25%)
> max:   13371.00			    max:   40572.00
> min:   12785.00			    min:   24088.00
> 
> 16 thread
> records: 10			    records: 10
> avg:   11092.40			    avg:   31424.20
> std:	 710.60(6.41%)		    std:    3763.89(11.98%)
> max:   12446.00			    max:   36635.00
> min:	9949.00			    min:   25669.00
> 
> 32 thread
> records: 10			    records: 10
> avg:   11067.00			    avg:   34495.80
> std:	 971.06(8.77%)		    std:    2721.36(7.89%)
> max:   12010.00			    max:   38598.00
> min:	9002.00			    min:   30636.00
> 
> In summary, MADV_FREE is about much faster than MADV_DONTNEED.

The MADV_FREE is discussed for a while, it probably is too late to propose
something new, but we had the new idea (from Ben Maurer, CCed) recently and
think it's better. Our target is still jemalloc.

Compared to MADV_DONTNEED, MADV_FREE's lazy memory free is a huge win to reduce
page fault. But there is one issue remaining, the TLB flush. Both MADV_DONTNEED
and MADV_FREE do TLB flush. TLB flush overhead is quite big in contemporary
multi-thread applications. In our production workload, we observed 80% CPU
spending on TLB flush triggered by jemalloc madvise(MADV_DONTNEED) sometimes.
We haven't tested MADV_FREE yet, but the result should be similar. It's hard to
avoid the TLB flush issue with MADV_FREE, because it helps avoid data
corruption.

The new proposal tries to fix the TLB issue. We introduce two madvise verbs:

MARK_FREE. Userspace notifies kernel the memory range can be discarded. Kernel
just records the range in current stage. Should memory pressure happen, page
reclaim can free the memory directly regardless the pte state.

MARK_NOFREE. Userspace notifies kernel the memory range will be reused soon.
Kernel deletes the record and prevents page reclaim discards the memory. If the
memory isn't reclaimed, userspace will access the old memory, otherwise do
normal page fault handling.

The point is to let userspace notify kernel if memory can be discarded, instead
of depending on pte dirty bit used by MADV_FREE. With these, no TLB flush is
required till page reclaim actually frees the memory (page reclaim need do the
TLB flush for MADV_FREE too). It still preserves the lazy memory free merit of
MADV_FREE.

Compared to MADV_FREE, reusing memory with the new proposal isn't transparent,
eg must call MARK_NOFREE. But it's easy to utilize the new API in jemalloc.

We don't have code to backup this yet, sorry. We'd like to discuss it if it
makes sense.

Thanks,
Shaohua
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1262604 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromDaniel Micay <danielmicay@gmail.com>
Date2015-11-04 22:20 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrdDk-5P5-13@gated-at.bofh.it>
In reply to#1262558

[Multipart message — attachments visible in raw view] — view raw

> Compared to MADV_DONTNEED, MADV_FREE's lazy memory free is a huge win to reduce
> page fault. But there is one issue remaining, the TLB flush. Both MADV_DONTNEED
> and MADV_FREE do TLB flush. TLB flush overhead is quite big in contemporary
> multi-thread applications. In our production workload, we observed 80% CPU
> spending on TLB flush triggered by jemalloc madvise(MADV_DONTNEED) sometimes.
> We haven't tested MADV_FREE yet, but the result should be similar. It's hard to
> avoid the TLB flush issue with MADV_FREE, because it helps avoid data
> corruption.
> 
> The new proposal tries to fix the TLB issue. We introduce two madvise verbs:
> 
> MARK_FREE. Userspace notifies kernel the memory range can be discarded. Kernel
> just records the range in current stage. Should memory pressure happen, page
> reclaim can free the memory directly regardless the pte state.
> 
> MARK_NOFREE. Userspace notifies kernel the memory range will be reused soon.
> Kernel deletes the record and prevents page reclaim discards the memory. If the
> memory isn't reclaimed, userspace will access the old memory, otherwise do
> normal page fault handling.
> 
> The point is to let userspace notify kernel if memory can be discarded, instead
> of depending on pte dirty bit used by MADV_FREE. With these, no TLB flush is
> required till page reclaim actually frees the memory (page reclaim need do the
> TLB flush for MADV_FREE too). It still preserves the lazy memory free merit of
> MADV_FREE.
> 
> Compared to MADV_FREE, reusing memory with the new proposal isn't transparent,
> eg must call MARK_NOFREE. But it's easy to utilize the new API in jemalloc.
> 
> We don't have code to backup this yet, sorry. We'd like to discuss it if it
> makes sense.

That's comparable to Android's pinning / unpinning API for ashmem and I
think it makes sense if it's faster. It's different than the MADV_FREE
API though, because the new allocations that are handed out won't have
the usual lazy commit which MADV_FREE provides. Pages in an allocation
that's handed out can still be dropped until they are actually written
to. It's considered active by jemalloc either way, but only a subset of
the active pages are actually committed. There's probably a use case for
both of these systems.

[toc] | [prev] | [next] | [standalone]


#1262608 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromDaniel Micay <danielmicay@gmail.com>
Date2015-11-04 22:30 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrdN0-5Sz-13@gated-at.bofh.it>
In reply to#1262604

[Multipart message — attachments visible in raw view] — view raw

> That's comparable to Android's pinning / unpinning API for ashmem and I
> think it makes sense if it's faster. It's different than the MADV_FREE
> API though, because the new allocations that are handed out won't have
> the usual lazy commit which MADV_FREE provides. Pages in an allocation
> that's handed out can still be dropped until they are actually written
> to. It's considered active by jemalloc either way, but only a subset of
> the active pages are actually committed. There's probably a use case for
> both of these systems.

Also, consider that MADV_FREE would allow jemalloc to be extremely
aggressive with purging when it actually has to do it. It can start with
the largest span of memory and it can mark more than strictly necessary
to drop below the ratio as there's no cost to using the memory again
(not even a system call).

Since the main cost is using the system call at all, there's going to be
pressure to mark the largest possible spans in one go. It will mean
concentration on memory compaction will improve performance. I think
that's the right direction for the kernel to be guiding userspace. It
will play better with THP than the allocator trying to be very precise
with purging based on aging.

[toc] | [prev] | [next] | [standalone]


#1262621 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromAndy Lutomirski <luto@amacapital.net>
Date2015-11-04 22:50 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qre6n-5Zh-37@gated-at.bofh.it>
In reply to#1262558
On Wed, Nov 4, 2015 at 12:00 PM, Shaohua Li <shli@kernel.org> wrote:
>
> The new proposal tries to fix the TLB issue. We introduce two madvise verbs:
>
> MARK_FREE. Userspace notifies kernel the memory range can be discarded. Kernel
> just records the range in current stage. Should memory pressure happen, page
> reclaim can free the memory directly regardless the pte state.
>
> MARK_NOFREE. Userspace notifies kernel the memory range will be reused soon.
> Kernel deletes the record and prevents page reclaim discards the memory. If the
> memory isn't reclaimed, userspace will access the old memory, otherwise do
> normal page fault handling.
>
> The point is to let userspace notify kernel if memory can be discarded, instead
> of depending on pte dirty bit used by MADV_FREE. With these, no TLB flush is
> required till page reclaim actually frees the memory (page reclaim need do the
> TLB flush for MADV_FREE too). It still preserves the lazy memory free merit of
> MADV_FREE.
>
> Compared to MADV_FREE, reusing memory with the new proposal isn't transparent,
> eg must call MARK_NOFREE. But it's easy to utilize the new API in jemalloc.
>

I can't speak to the usefulness of this or to other arches, but on x86
(unless you have nohz_full or similar enabled), a pair of syscalls
should be *much* faster than an IPI or a page fault.

I don't know how expensive it is to write to a clean page or to access
an unaccessed page on x86.  I'm sure it's not free (there's memory
bandwidth if nothing else), but it could be very cheap.

--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1262796 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromMinchan Kim <minchan@kernel.org>
Date2015-11-05 02:40 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrhGV-8jJ-1@gated-at.bofh.it>
In reply to#1262558
On Wed, Nov 04, 2015 at 12:00:06PM -0800, Shaohua Li wrote:
> On Wed, Nov 04, 2015 at 10:25:55AM +0900, Minchan Kim wrote:
> > Linux doesn't have an ability to free pages lazy while other OS already
> > have been supported that named by madvise(MADV_FREE).
> > 
> > The gain is clear that kernel can discard freed pages rather than swapping
> > out or OOM if memory pressure happens.
> > 
> > Without memory pressure, freed pages would be reused by userspace without
> > another additional overhead(ex, page fault + allocation + zeroing).
> > 
> > Jason Evans said:
> > 
> > : Facebook has been using MAP_UNINITIALIZED
> > : (https://lkml.org/lkml/2012/1/18/308) in some of its applications for
> > : several years, but there are operational costs to maintaining this
> > : out-of-tree in our kernel and in jemalloc, and we are anxious to retire it
> > : in favor of MADV_FREE.  When we first enabled MAP_UNINITIALIZED it
> > : increased throughput for much of our workload by ~5%, and although the
> > : benefit has decreased using newer hardware and kernels, there is still
> > : enough benefit that we cannot reasonably retire it without a replacement.
> > :
> > : Aside from Facebook operations, there are numerous broadly used
> > : applications that would benefit from MADV_FREE.  The ones that immediately
> > : come to mind are redis, varnish, and MariaDB.  I don't have much insight
> > : into Android internals and development process, but I would hope to see
> > : MADV_FREE support eventually end up there as well to benefit applications
> > : linked with the integrated jemalloc.
> > :
> > : jemalloc will use MADV_FREE once it becomes available in the Linux kernel.
> > : In fact, jemalloc already uses MADV_FREE or equivalent everywhere it's
> > : available: *BSD, OS X, Windows, and Solaris -- every platform except Linux
> > : (and AIX, but I'm not sure it even compiles on AIX).  The lack of
> > : MADV_FREE on Linux forced me down a long series of increasingly
> > : sophisticated heuristics for madvise() volume reduction, and even so this
> > : remains a common performance issue for people using jemalloc on Linux.
> > : Please integrate MADV_FREE; many people will benefit substantially.
> > 
> > How it works:
> > 
> > When madvise syscall is called, VM clears dirty bit of ptes of the range.
> > If memory pressure happens, VM checks dirty bit of page table and if it
> > found still "clean", it means it's a "lazyfree pages" so VM could discard
> > the page instead of swapping out.  Once there was store operation for the
> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out
> > the page instead of discarding.
> > 
> > Firstly, heavy users would be general allocators(ex, jemalloc, tcmalloc
> > and hope glibc supports it) and jemalloc/tcmalloc already have supported
> > the feature for other OS(ex, FreeBSD)
> > 
> > barrios@blaptop:~/benchmark/ebizzy$ lscpu
> > Architecture:          x86_64
> > CPU op-mode(s):        32-bit, 64-bit
> > Byte Order:            Little Endian
> > CPU(s):                12
> > On-line CPU(s) list:   0-11
> > Thread(s) per core:    1
> > Core(s) per socket:    1
> > Socket(s):             12
> > NUMA node(s):          1
> > Vendor ID:             GenuineIntel
> > CPU family:            6
> > Model:                 2
> > Stepping:              3
> > CPU MHz:               3200.185
> > BogoMIPS:              6400.53
> > Virtualization:        VT-x
> > Hypervisor vendor:     KVM
> > Virtualization type:   full
> > L1d cache:             32K
> > L1i cache:             32K
> > L2 cache:              4096K
> > NUMA node0 CPU(s):     0-11
> > ebizzy benchmark(./ebizzy -S 10 -n 512)
> > 
> > Higher avg is better.
> > 
> >  vanilla-jemalloc		MADV_free-jemalloc
> > 
> > 1 thread
> > records: 10			    records: 10
> > avg:	2961.90			    avg:   12069.70
> > std:	  71.96(2.43%)		    std:     186.68(1.55%)
> > max:	3070.00			    max:   12385.00
> > min:	2796.00			    min:   11746.00
> > 
> > 2 thread
> > records: 10			    records: 10
> > avg:	5020.00			    avg:   17827.00
> > std:	 264.87(5.28%)		    std:     358.52(2.01%)
> > max:	5244.00			    max:   18760.00
> > min:	4251.00			    min:   17382.00
> > 
> > 4 thread
> > records: 10			    records: 10
> > avg:	8988.80			    avg:   27930.80
> > std:	1175.33(13.08%)		    std:    3317.33(11.88%)
> > max:	9508.00			    max:   30879.00
> > min:	5477.00			    min:   21024.00
> > 
> > 8 thread
> > records: 10			    records: 10
> > avg:   13036.50			    avg:   33739.40
> > std:	 170.67(1.31%)		    std:    5146.22(15.25%)
> > max:   13371.00			    max:   40572.00
> > min:   12785.00			    min:   24088.00
> > 
> > 16 thread
> > records: 10			    records: 10
> > avg:   11092.40			    avg:   31424.20
> > std:	 710.60(6.41%)		    std:    3763.89(11.98%)
> > max:   12446.00			    max:   36635.00
> > min:	9949.00			    min:   25669.00
> > 
> > 32 thread
> > records: 10			    records: 10
> > avg:   11067.00			    avg:   34495.80
> > std:	 971.06(8.77%)		    std:    2721.36(7.89%)
> > max:   12010.00			    max:   38598.00
> > min:	9002.00			    min:   30636.00
> > 
> > In summary, MADV_FREE is about much faster than MADV_DONTNEED.
> 
> The MADV_FREE is discussed for a while, it probably is too late to propose
> something new, but we had the new idea (from Ben Maurer, CCed) recently and
> think it's better. Our target is still jemalloc.
> 
> Compared to MADV_DONTNEED, MADV_FREE's lazy memory free is a huge win to reduce
> page fault. But there is one issue remaining, the TLB flush. Both MADV_DONTNEED
> and MADV_FREE do TLB flush. TLB flush overhead is quite big in contemporary
> multi-thread applications. In our production workload, we observed 80% CPU
> spending on TLB flush triggered by jemalloc madvise(MADV_DONTNEED) sometimes.
> We haven't tested MADV_FREE yet, but the result should be similar. It's hard to
> avoid the TLB flush issue with MADV_FREE, because it helps avoid data
> corruption.
> 
> The new proposal tries to fix the TLB issue. We introduce two madvise verbs:
> 
> MARK_FREE. Userspace notifies kernel the memory range can be discarded. Kernel
> just records the range in current stage. Should memory pressure happen, page
> reclaim can free the memory directly regardless the pte state.
> 
> MARK_NOFREE. Userspace notifies kernel the memory range will be reused soon.
> Kernel deletes the record and prevents page reclaim discards the memory. If the
> memory isn't reclaimed, userspace will access the old memory, otherwise do
> normal page fault handling.
> 
> The point is to let userspace notify kernel if memory can be discarded, instead
> of depending on pte dirty bit used by MADV_FREE. With these, no TLB flush is
> required till page reclaim actually frees the memory (page reclaim need do the
> TLB flush for MADV_FREE too). It still preserves the lazy memory free merit of
> MADV_FREE.
> 
> Compared to MADV_FREE, reusing memory with the new proposal isn't transparent,
> eg must call MARK_NOFREE. But it's easy to utilize the new API in jemalloc.
> 
> We don't have code to backup this yet, sorry. We'd like to discuss it if it
> makes sense.

It's really what volatile range did.
John Stultz and me tried it for a *long* time but it had lots of troubles.
It's really hard to write it down in my time due to really long history
and even I forgot lots of detail(ie, dead brain).
Please search volatile ranges in google.
Finally, people in LSF/MM suggested MADV_FREE to help anonymous page side
rather than stucking hich prevent useful feature. :(

> 
> Thanks,
> Shaohua
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [next] | [standalone]


#1262798 — Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE)

FromMinchan Kim <minchan@kernel.org>
Date2015-11-05 02:40 +0100
SubjectRe: [PATCH v2 01/13] mm: support madvise(MADV_FREE)
Message-ID<qrhGV-8jJ-9@gated-at.bofh.it>
In reply to#1262796
On Thu, Nov 05, 2015 at 10:33:50AM +0900, Minchan Kim wrote:
> On Wed, Nov 04, 2015 at 12:00:06PM -0800, Shaohua Li wrote:
> > On Wed, Nov 04, 2015 at 10:25:55AM +0900, Minchan Kim wrote:
> > > Linux doesn't have an ability to free pages lazy while other OS already
> > > have been supported that named by madvise(MADV_FREE).
> > > 
> > > The gain is clear that kernel can discard freed pages rather than swapping
> > > out or OOM if memory pressure happens.
> > > 
> > > Without memory pressure, freed pages would be reused by userspace without
> > > another additional overhead(ex, page fault + allocation + zeroing).
> > > 
> > > Jason Evans said:
> > > 
> > > : Facebook has been using MAP_UNINITIALIZED
> > > : (https://lkml.org/lkml/2012/1/18/308) in some of its applications for
> > > : several years, but there are operational costs to maintaining this
> > > : out-of-tree in our kernel and in jemalloc, and we are anxious to retire it
> > > : in favor of MADV_FREE.  When we first enabled MAP_UNINITIALIZED it
> > > : increased throughput for much of our workload by ~5%, and although the
> > > : benefit has decreased using newer hardware and kernels, there is still
> > > : enough benefit that we cannot reasonably retire it without a replacement.
> > > :
> > > : Aside from Facebook operations, there are numerous broadly used
> > > : applications that would benefit from MADV_FREE.  The ones that immediately
> > > : come to mind are redis, varnish, and MariaDB.  I don't have much insight
> > > : into Android internals and development process, but I would hope to see
> > > : MADV_FREE support eventually end up there as well to benefit applications
> > > : linked with the integrated jemalloc.
> > > :
> > > : jemalloc will use MADV_FREE once it becomes available in the Linux kernel.
> > > : In fact, jemalloc already uses MADV_FREE or equivalent everywhere it's
> > > : available: *BSD, OS X, Windows, and Solaris -- every platform except Linux
> > > : (and AIX, but I'm not sure it even compiles on AIX).  The lack of
> > > : MADV_FREE on Linux forced me down a long series of increasingly
> > > : sophisticated heuristics for madvise() volume reduction, and even so this
> > > : remains a common performance issue for people using jemalloc on Linux.
> > > : Please integrate MADV_FREE; many people will benefit substantially.
> > > 
> > > How it works:
> > > 
> > > When madvise syscall is called, VM clears dirty bit of ptes of the range.
> > > If memory pressure happens, VM checks dirty bit of page table and if it
> > > found still "clean", it means it's a "lazyfree pages" so VM could discard
> > > the page instead of swapping out.  Once there was store operation for the
> > > page before VM peek a page to reclaim, dirty bit is set so VM can swap out
> > > the page instead of discarding.
> > > 
> > > Firstly, heavy users would be general allocators(ex, jemalloc, tcmalloc
> > > and hope glibc supports it) and jemalloc/tcmalloc already have supported
> > > the feature for other OS(ex, FreeBSD)
> > > 
> > > barrios@blaptop:~/benchmark/ebizzy$ lscpu
> > > Architecture:          x86_64
> > > CPU op-mode(s):        32-bit, 64-bit
> > > Byte Order:            Little Endian
> > > CPU(s):                12
> > > On-line CPU(s) list:   0-11
> > > Thread(s) per core:    1
> > > Core(s) per socket:    1
> > > Socket(s):             12
> > > NUMA node(s):          1
> > > Vendor ID:             GenuineIntel
> > > CPU family:            6
> > > Model:                 2
> > > Stepping:              3
> > > CPU MHz:               3200.185
> > > BogoMIPS:              6400.53
> > > Virtualization:        VT-x
> > > Hypervisor vendor:     KVM
> > > Virtualization type:   full
> > > L1d cache:             32K
> > > L1i cache:             32K
> > > L2 cache:              4096K
> > > NUMA node0 CPU(s):     0-11
> > > ebizzy benchmark(./ebizzy -S 10 -n 512)
> > > 
> > > Higher avg is better.
> > > 
> > >  vanilla-jemalloc		MADV_free-jemalloc
> > > 
> > > 1 thread
> > > records: 10			    records: 10
> > > avg:	2961.90			    avg:   12069.70
> > > std:	  71.96(2.43%)		    std:     186.68(1.55%)
> > > max:	3070.00			    max:   12385.00
> > > min:	2796.00			    min:   11746.00
> > > 
> > > 2 thread
> > > records: 10			    records: 10
> > > avg:	5020.00			    avg:   17827.00
> > > std:	 264.87(5.28%)		    std:     358.52(2.01%)
> > > max:	5244.00			    max:   18760.00
> > > min:	4251.00			    min:   17382.00
> > > 
> > > 4 thread
> > > records: 10			    records: 10
> > > avg:	8988.80			    avg:   27930.80
> > > std:	1175.33(13.08%)		    std:    3317.33(11.88%)
> > > max:	9508.00			    max:   30879.00
> > > min:	5477.00			    min:   21024.00
> > > 
> > > 8 thread
> > > records: 10			    records: 10
> > > avg:   13036.50			    avg:   33739.40
> > > std:	 170.67(1.31%)		    std:    5146.22(15.25%)
> > > max:   13371.00			    max:   40572.00
> > > min:   12785.00			    min:   24088.00
> > > 
> > > 16 thread
> > > records: 10			    records: 10
> > > avg:   11092.40			    avg:   31424.20
> > > std:	 710.60(6.41%)		    std:    3763.89(11.98%)
> > > max:   12446.00			    max:   36635.00
> > > min:	9949.00			    min:   25669.00
> > > 
> > > 32 thread
> > > records: 10			    records: 10
> > > avg:   11067.00			    avg:   34495.80
> > > std:	 971.06(8.77%)		    std:    2721.36(7.89%)
> > > max:   12010.00			    max:   38598.00
> > > min:	9002.00			    min:   30636.00
> > > 
> > > In summary, MADV_FREE is about much faster than MADV_DONTNEED.
> > 
> > The MADV_FREE is discussed for a while, it probably is too late to propose
> > something new, but we had the new idea (from Ben Maurer, CCed) recently and
> > think it's better. Our target is still jemalloc.
> > 
> > Compared to MADV_DONTNEED, MADV_FREE's lazy memory free is a huge win to reduce
> > page fault. But there is one issue remaining, the TLB flush. Both MADV_DONTNEED
> > and MADV_FREE do TLB flush. TLB flush overhead is quite big in contemporary
> > multi-thread applications. In our production workload, we observed 80% CPU
> > spending on TLB flush triggered by jemalloc madvise(MADV_DONTNEED) sometimes.
> > We haven't tested MADV_FREE yet, but the result should be similar. It's hard to
> > avoid the TLB flush issue with MADV_FREE, because it helps avoid data
> > corruption.
> > 
> > The new proposal tries to fix the TLB issue. We introduce two madvise verbs:
> > 
> > MARK_FREE. Userspace notifies kernel the memory range can be discarded. Kernel
> > just records the range in current stage. Should memory pressure happen, page
> > reclaim can free the memory directly regardless the pte state.
> > 
> > MARK_NOFREE. Userspace notifies kernel the memory range will be reused soon.
> > Kernel deletes the record and prevents page reclaim discards the memory. If the
> > memory isn't reclaimed, userspace will access the old memory, otherwise do
> > normal page fault handling.
> > 
> > The point is to let userspace notify kernel if memory can be discarded, instead
> > of depending on pte dirty bit used by MADV_FREE. With these, no TLB flush is
> > required till page reclaim actually frees the memory (page reclaim need do the
> > TLB flush for MADV_FREE too). It still preserves the lazy memory free merit of
> > MADV_FREE.
> > 
> > Compared to MADV_FREE, reusing memory with the new proposal isn't transparent,
> > eg must call MARK_NOFREE. But it's easy to utilize the new API in jemalloc.
> > 
> > We don't have code to backup this yet, sorry. We'd like to discuss it if it
> > makes sense.
> 
> It's really what volatile range did.
> John Stultz and me tried it for a *long* time but it had lots of troubles.
> It's really hard to write it down in my time due to really long history
> and even I forgot lots of detail(ie, dead brain).
> Please search volatile ranges in google.
> Finally, people in LSF/MM suggested MADV_FREE to help anonymous page side
> rather than stucking hich prevent useful feature. :(

I should have Cced John Stutlz.

He would have good memory than me so he would help but I'm not sure
he has a interest on volatile ranges, still.
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at  http://vger.kernel.org/majordomo-info.html
Please read the FAQ at  http://www.tux.org/lkml/

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web