Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1262009 > unrolled thread
| Started by | Minchan Kim <minchan@kernel.org> |
|---|---|
| First post | 2015-11-04 02:30 +0100 |
| Last post | 2015-11-05 02:40 +0100 |
| Articles | 15 on this page of 35 — 5 participants |
Back to article view | Back to linux.kernel
[PATCH v2 00/13] MADV_FREE support Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
[PATCH v2 11/13] arm: add pmd_mkclean for THP Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
[PATCH v2 13/13] mm: don't split THP page when syscall is called Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
[PATCH v2 04/13] mm: free swp_entry in madvise_free Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
[PATCH v2 02/13] mm: define MADV_FREE for some arches Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
[PATCH v2 12/13] arm64: add pmd_mkclean for THP Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
[PATCH v2 08/13] x86: add pmd_[dirty|mkclean] for THP Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
[PATCH v2 07/13] mm: mark stable page dirty in KSM Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
[PATCH v2 05/13] mm: move lazily freed pages to inactive list Minchan Kim <minchan@kernel.org> - 2015-11-04 02:30 +0100
[PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-04 02:40 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2015-11-04 03:20 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 00:40 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2015-11-05 04:50 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Sergey Senozhatsky <sergey.senozhatsky.work@gmail.com> - 2015-11-04 03:40 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 00:50 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Andy Lutomirski <luto@amacapital.net> - 2015-11-04 04:50 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 07:00 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 07:00 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 07:10 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Andy Lutomirski <luto@amacapital.net> - 2015-11-04 19:30 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 23:10 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Shaohua Li <shli@kernel.org> - 2015-11-05 19:20 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-05 21:20 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-05 21:20 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 01:20 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Andy Lutomirski <luto@amacapital.net> - 2015-11-05 01:50 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 02:00 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Andy Lutomirski <luto@amacapital.net> - 2015-11-05 02:40 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 02:50 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Shaohua Li <shli@kernel.org> - 2015-11-04 21:10 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 22:20 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Daniel Micay <danielmicay@gmail.com> - 2015-11-04 22:30 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Andy Lutomirski <luto@amacapital.net> - 2015-11-04 22:50 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 02:40 +0100
Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) Minchan Kim <minchan@kernel.org> - 2015-11-05 02:40 +0100
Page 2 of 2 — ← Prev page 1 [2]
| From | Daniel Micay <danielmicay@gmail.com> |
|---|---|
| Date | 2015-11-04 23:10 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrepJ-6lw-29@gated-at.bofh.it> |
| In reply to | #1262487 |
[Multipart message — attachments visible in raw view] — view raw
> With enough pages at once, though, munmap would be fine, too. That implies lots of page faults and zeroing though. The zeroing alone is a major performance issue. There are separate issues with munmap since it ends up resulting in a lot more virtual memory fragmentation. It would help if the kernel used first-best-fit for mmap instead of the current naive algorithm (bonus: O(log n) worst-case, not O(n)). Since allocators like jemalloc and PartitionAlloc want 2M aligned spans, mixing them with other allocators can also accelerate the VM fragmentation caused by the dumb mmap algorithm (i.e. they make a 2M aligned mapping, some other mmap user does 4k, now there's a nearly 2M gap when the next 2M region is made and the kernel keeps going rather than reusing it). Anyway, that's a totally separate issue from this. Just felt like complaining :). > Maybe what's really needed is a MADV_FREE variant that takes an iovec. > On an all-cores multithreaded mm, the TLB shootdown broadcast takes > thousands of cycles on each core more or less regardless of how much > of the TLB gets zapped. That would work very well. The allocator ends up having a sequence of dirty spans that it needs to purge in one go. As long as purging is fairly spread out, the cost of a single TLB shootdown isn't that bad. It is extremely bad if it needs to do it over and over to purge a bunch of ranges, which can happen if the memory has ended up being very, very fragmentated despite the efforts to compact it (depends on what the application ends up doing).
[toc] | [prev] | [next] | [standalone]
| From | Shaohua Li <shli@kernel.org> |
|---|---|
| Date | 2015-11-05 19:20 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrxiF-1JR-1@gated-at.bofh.it> |
| In reply to | #1262632 |
On Wed, Nov 04, 2015 at 05:05:47PM -0500, Daniel Micay wrote: > > With enough pages at once, though, munmap would be fine, too. > > That implies lots of page faults and zeroing though. The zeroing alone > is a major performance issue. > > There are separate issues with munmap since it ends up resulting in a > lot more virtual memory fragmentation. It would help if the kernel used > first-best-fit for mmap instead of the current naive algorithm (bonus: > O(log n) worst-case, not O(n)). Since allocators like jemalloc and > PartitionAlloc want 2M aligned spans, mixing them with other allocators > can also accelerate the VM fragmentation caused by the dumb mmap > algorithm (i.e. they make a 2M aligned mapping, some other mmap user > does 4k, now there's a nearly 2M gap when the next 2M region is made and > the kernel keeps going rather than reusing it). Anyway, that's a totally > separate issue from this. Just felt like complaining :). > > > Maybe what's really needed is a MADV_FREE variant that takes an iovec. > > On an all-cores multithreaded mm, the TLB shootdown broadcast takes > > thousands of cycles on each core more or less regardless of how much > > of the TLB gets zapped. > > That would work very well. The allocator ends up having a sequence of > dirty spans that it needs to purge in one go. As long as purging is > fairly spread out, the cost of a single TLB shootdown isn't that bad. It > is extremely bad if it needs to do it over and over to purge a bunch of > ranges, which can happen if the memory has ended up being very, very > fragmentated despite the efforts to compact it (depends on what the > application ends up doing). I posted a patch doing exactly iovec madvise. Doesn't support MADV_FREE yet though, but should be easy to do it. http://marc.info/?l=linux-mm&m=144615663522661&w=2 -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Daniel Micay <danielmicay@gmail.com> |
|---|---|
| Date | 2015-11-05 21:20 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrzaN-2TZ-3@gated-at.bofh.it> |
| In reply to | #1263465 |
[Multipart message — attachments visible in raw view] — view raw
> I posted a patch doing exactly iovec madvise. Doesn't support MADV_FREE yet > though, but should be easy to do it. > > http://marc.info/?l=linux-mm&m=144615663522661&w=2 I think that would be a great way to deal with this. It keeps the nice property of still being able to drop pages in allocations that have been handed out but not yet touched. The allocator just needs to be designed to do lots of purging in one go (i.e. something like an 8:1 active:clean ratio triggers purging and it goes all the way to 16:1).
[toc] | [prev] | [next] | [standalone]
| From | Daniel Micay <danielmicay@gmail.com> |
|---|---|
| Date | 2015-11-05 21:20 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrzaN-2TZ-1@gated-at.bofh.it> |
| In reply to | #1263528 |
[Multipart message — attachments visible in raw view] — view raw
> active:clean active:dirty*, sigh.
[toc] | [prev] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2015-11-05 01:20 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrgrv-7Cx-1@gated-at.bofh.it> |
| In reply to | #1262060 |
On Tue, Nov 03, 2015 at 07:41:35PM -0800, Andy Lutomirski wrote:
> On Nov 3, 2015 5:30 PM, "Minchan Kim" <minchan@kernel.org> wrote:
> >
> > Linux doesn't have an ability to free pages lazy while other OS already
> > have been supported that named by madvise(MADV_FREE).
> >
> > The gain is clear that kernel can discard freed pages rather than swapping
> > out or OOM if memory pressure happens.
> >
> > Without memory pressure, freed pages would be reused by userspace without
> > another additional overhead(ex, page fault + allocation + zeroing).
> >
>
> [...]
>
> >
> > How it works:
> >
> > When madvise syscall is called, VM clears dirty bit of ptes of the range.
> > If memory pressure happens, VM checks dirty bit of page table and if it
> > found still "clean", it means it's a "lazyfree pages" so VM could discard
> > the page instead of swapping out. Once there was store operation for the
> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out
> > the page instead of discarding.
>
> What happens if you MADV_FREE something that's MAP_SHARED or isn't
> ordinary anonymous memory? There's a long history of MADV_DONTNEED on
> such mappings causing exploitable problems, and I think it would be
> nice if MADV_FREE were obviously safe.
It filter out VM_LOCKED|VM_HUGETLB|VM_PFNMAP and file-backed vma and MAP_SHARED
with vma_is_anonymous.
>
> Does this set the write protect bit?
No.
>
> What happens on architectures without hardware dirty tracking? For
> that matter, even on architecture with hardware dirty tracking, what
> happens in multithreaded processes that have the dirty TLB state
> cached in a different CPU's TLB?
>
> Using the dirty bit for these semantics scares me. This API creates a
> page that can have visible nonzero contents and then can
> asynchronously and magically zero itself thereafter. That makes me
> nervous. Could we use the accessed bit instead? Then the observable
Access bit is used by aging algorithm for reclaim. In addition,
we have supported clear_refs feacture.
IOW, it could be reset anytime so it's hard to use marker for
lazy freeing at the moment.
> semantics would be equivalent to having MADV_FREE either zero the page
> or do nothing, except that it doesn't make up its mind until the next
> read.
>
> > + ptent = pte_mkold(ptent);
> > + ptent = pte_mkclean(ptent);
> > + set_pte_at(mm, addr, pte, ptent);
> > + tlb_remove_tlb_entry(tlb, pte, addr);
>
> It looks like you are flushing the TLB. In a multithreaded program,
> that's rather expensive. Potentially silly question: would it be
> better to just zero the page immediately in a multithreaded program
> and then, when swapping out, check the page is zeroed and, if so, skip
> swapping it out? That could be done without forcing an IPI.
So, we should monitor all of pages in reclaim patch whether they are
zero or not? It is fatster for allocation side but much slower in
reclaim side. For avoiding that, we should mark something for lazy
freeing page out of page table.
Anyway, it depends on the TLB flush overehead vs memset overhead.
If the hinted range is pretty big and small system(ie, not many core),
memset overhead would't not trivial compared to TLB flush.
Even, some of ARM arches doesn't do IPI to TLB flush so the overhead
would be cheaper.
I don't want to push more optimization in new syscall from the beginning.
It's an optimization and might come better idea once we hear from the
voice of userland folks. Then, it's not too late.
Let's do step by step.
>
> > +static int madvise_free_single_vma(struct vm_area_struct *vma,
> > + unsigned long start_addr, unsigned long end_addr)
> > +{
> > + unsigned long start, end;
> > + struct mm_struct *mm = vma->vm_mm;
> > + struct mmu_gather tlb;
> > +
> > + if (vma->vm_flags & (VM_LOCKED|VM_HUGETLB|VM_PFNMAP))
> > + return -EINVAL;
> > +
> > + /* MADV_FREE works for only anon vma at the moment */
> > + if (!vma_is_anonymous(vma))
> > + return -EINVAL;
>
> Does anything weird happen if it's shared?
Hmm, you mean MAP_SHARED|MAP_ANONYMOUS?
In that case, vma->vm_ops = &shmem_vm_ops so vma_is anonymous should filter it out.
>
> > + if (!PageDirty(page) && (flags & TTU_FREE)) {
> > + /* It's a freeable page by MADV_FREE */
> > + dec_mm_counter(mm, MM_ANONPAGES);
> > + goto discard;
> > + }
>
> Does something clear TTU_FREE the next time the page gets marked clean?
Sorry, I don't understand. Could you elaborate it more?
>
> --Andy
>
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org. For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2015-11-05 01:50 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrgUx-7M5-11@gated-at.bofh.it> |
| In reply to | #1262759 |
On Wed, Nov 4, 2015 at 4:13 PM, Minchan Kim <minchan@kernel.org> wrote:
> On Tue, Nov 03, 2015 at 07:41:35PM -0800, Andy Lutomirski wrote:
>> On Nov 3, 2015 5:30 PM, "Minchan Kim" <minchan@kernel.org> wrote:
>> >
>> > Linux doesn't have an ability to free pages lazy while other OS already
>> > have been supported that named by madvise(MADV_FREE).
>> >
>> > The gain is clear that kernel can discard freed pages rather than swapping
>> > out or OOM if memory pressure happens.
>> >
>> > Without memory pressure, freed pages would be reused by userspace without
>> > another additional overhead(ex, page fault + allocation + zeroing).
>> >
>>
>> [...]
>>
>> >
>> > How it works:
>> >
>> > When madvise syscall is called, VM clears dirty bit of ptes of the range.
>> > If memory pressure happens, VM checks dirty bit of page table and if it
>> > found still "clean", it means it's a "lazyfree pages" so VM could discard
>> > the page instead of swapping out. Once there was store operation for the
>> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out
>> > the page instead of discarding.
>>
>> What happens if you MADV_FREE something that's MAP_SHARED or isn't
>> ordinary anonymous memory? There's a long history of MADV_DONTNEED on
>> such mappings causing exploitable problems, and I think it would be
>> nice if MADV_FREE were obviously safe.
>
> It filter out VM_LOCKED|VM_HUGETLB|VM_PFNMAP and file-backed vma and MAP_SHARED
> with vma_is_anonymous.
>
>>
>> Does this set the write protect bit?
>
> No.
>
>>
>> What happens on architectures without hardware dirty tracking? For
>> that matter, even on architecture with hardware dirty tracking, what
>> happens in multithreaded processes that have the dirty TLB state
>> cached in a different CPU's TLB?
>>
>> Using the dirty bit for these semantics scares me. This API creates a
>> page that can have visible nonzero contents and then can
>> asynchronously and magically zero itself thereafter. That makes me
>> nervous. Could we use the accessed bit instead? Then the observable
>
> Access bit is used by aging algorithm for reclaim. In addition,
> we have supported clear_refs feacture.
> IOW, it could be reset anytime so it's hard to use marker for
> lazy freeing at the moment.
>
That's unfortunate. I think that the ABI would be much nicer if it
used the accessed bit.
In any case, shouldn't the aging algorithm be irrelevant here? A
MADV_FREE page that isn't accessed can be discarded, whereas we could
hopefully just say that a MADV_FREE page that is accessed gets moved
to whatever list holds recently accessed pages and also stops being a
candidate for discarding due to MADV_FREE?
>>
>> > + if (!PageDirty(page) && (flags & TTU_FREE)) {
>> > + /* It's a freeable page by MADV_FREE */
>> > + dec_mm_counter(mm, MM_ANONPAGES);
>> > + goto discard;
>> > + }
>>
>> Does something clear TTU_FREE the next time the page gets marked clean?
>
> Sorry, I don't understand. Could you elaborate it more?
I don't fully understand how TTU_FREE ends up being set here, but, if
the page is dirtied by user code and then cleaned later by the kernel,
what prevents TTU_FREE from being incorrectly set here?
--Andy
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2015-11-05 02:00 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrh4g-7PW-43@gated-at.bofh.it> |
| In reply to | #1262764 |
On Wed, Nov 04, 2015 at 04:42:37PM -0800, Andy Lutomirski wrote:
> On Wed, Nov 4, 2015 at 4:13 PM, Minchan Kim <minchan@kernel.org> wrote:
> > On Tue, Nov 03, 2015 at 07:41:35PM -0800, Andy Lutomirski wrote:
> >> On Nov 3, 2015 5:30 PM, "Minchan Kim" <minchan@kernel.org> wrote:
> >> >
> >> > Linux doesn't have an ability to free pages lazy while other OS already
> >> > have been supported that named by madvise(MADV_FREE).
> >> >
> >> > The gain is clear that kernel can discard freed pages rather than swapping
> >> > out or OOM if memory pressure happens.
> >> >
> >> > Without memory pressure, freed pages would be reused by userspace without
> >> > another additional overhead(ex, page fault + allocation + zeroing).
> >> >
> >>
> >> [...]
> >>
> >> >
> >> > How it works:
> >> >
> >> > When madvise syscall is called, VM clears dirty bit of ptes of the range.
> >> > If memory pressure happens, VM checks dirty bit of page table and if it
> >> > found still "clean", it means it's a "lazyfree pages" so VM could discard
> >> > the page instead of swapping out. Once there was store operation for the
> >> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out
> >> > the page instead of discarding.
> >>
> >> What happens if you MADV_FREE something that's MAP_SHARED or isn't
> >> ordinary anonymous memory? There's a long history of MADV_DONTNEED on
> >> such mappings causing exploitable problems, and I think it would be
> >> nice if MADV_FREE were obviously safe.
> >
> > It filter out VM_LOCKED|VM_HUGETLB|VM_PFNMAP and file-backed vma and MAP_SHARED
> > with vma_is_anonymous.
> >
> >>
> >> Does this set the write protect bit?
> >
> > No.
> >
> >>
> >> What happens on architectures without hardware dirty tracking? For
> >> that matter, even on architecture with hardware dirty tracking, what
> >> happens in multithreaded processes that have the dirty TLB state
> >> cached in a different CPU's TLB?
> >>
> >> Using the dirty bit for these semantics scares me. This API creates a
> >> page that can have visible nonzero contents and then can
> >> asynchronously and magically zero itself thereafter. That makes me
> >> nervous. Could we use the accessed bit instead? Then the observable
> >
> > Access bit is used by aging algorithm for reclaim. In addition,
> > we have supported clear_refs feacture.
> > IOW, it could be reset anytime so it's hard to use marker for
> > lazy freeing at the moment.
> >
>
> That's unfortunate. I think that the ABI would be much nicer if it
> used the accessed bit.
>
> In any case, shouldn't the aging algorithm be irrelevant here? A
> MADV_FREE page that isn't accessed can be discarded, whereas we could
> hopefully just say that a MADV_FREE page that is accessed gets moved
> to whatever list holds recently accessed pages and also stops being a
> candidate for discarding due to MADV_FREE?
I meant if we use access bit as indicator for lazy-freeing page,
we could discard valid page which is never hinted by MADV_FREE but
just doesn't mark access bit in page table by aging algorithm.
>
> >>
> >> > + if (!PageDirty(page) && (flags & TTU_FREE)) {
> >> > + /* It's a freeable page by MADV_FREE */
> >> > + dec_mm_counter(mm, MM_ANONPAGES);
> >> > + goto discard;
> >> > + }
> >>
> >> Does something clear TTU_FREE the next time the page gets marked clean?
> >
> > Sorry, I don't understand. Could you elaborate it more?
>
> I don't fully understand how TTU_FREE ends up being set here, but, if
> the page is dirtied by user code and then cleaned later by the kernel,
> what prevents TTU_FREE from being incorrectly set here?
Kernel shouldn't make the page clean without writeback(ie, swapout)
if the page has valid data.
>
>
> --Andy
>
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org. For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2015-11-05 02:40 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrhGV-8jJ-13@gated-at.bofh.it> |
| In reply to | #1262785 |
On Wed, Nov 4, 2015 at 4:56 PM, Minchan Kim <minchan@kernel.org> wrote: > On Wed, Nov 04, 2015 at 04:42:37PM -0800, Andy Lutomirski wrote: >> On Wed, Nov 4, 2015 at 4:13 PM, Minchan Kim <minchan@kernel.org> wrote: >> > On Tue, Nov 03, 2015 at 07:41:35PM -0800, Andy Lutomirski wrote: >> >> On Nov 3, 2015 5:30 PM, "Minchan Kim" <minchan@kernel.org> wrote: >> >> > >> >> > Linux doesn't have an ability to free pages lazy while other OS already >> >> > have been supported that named by madvise(MADV_FREE). >> >> > >> >> > The gain is clear that kernel can discard freed pages rather than swapping >> >> > out or OOM if memory pressure happens. >> >> > >> >> > Without memory pressure, freed pages would be reused by userspace without >> >> > another additional overhead(ex, page fault + allocation + zeroing). >> >> > >> >> >> >> [...] >> >> >> >> > >> >> > How it works: >> >> > >> >> > When madvise syscall is called, VM clears dirty bit of ptes of the range. >> >> > If memory pressure happens, VM checks dirty bit of page table and if it >> >> > found still "clean", it means it's a "lazyfree pages" so VM could discard >> >> > the page instead of swapping out. Once there was store operation for the >> >> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out >> >> > the page instead of discarding. >> >> >> >> What happens if you MADV_FREE something that's MAP_SHARED or isn't >> >> ordinary anonymous memory? There's a long history of MADV_DONTNEED on >> >> such mappings causing exploitable problems, and I think it would be >> >> nice if MADV_FREE were obviously safe. >> > >> > It filter out VM_LOCKED|VM_HUGETLB|VM_PFNMAP and file-backed vma and MAP_SHARED >> > with vma_is_anonymous. >> > >> >> >> >> Does this set the write protect bit? >> > >> > No. >> > >> >> >> >> What happens on architectures without hardware dirty tracking? For >> >> that matter, even on architecture with hardware dirty tracking, what >> >> happens in multithreaded processes that have the dirty TLB state >> >> cached in a different CPU's TLB? >> >> >> >> Using the dirty bit for these semantics scares me. This API creates a >> >> page that can have visible nonzero contents and then can >> >> asynchronously and magically zero itself thereafter. That makes me >> >> nervous. Could we use the accessed bit instead? Then the observable >> > >> > Access bit is used by aging algorithm for reclaim. In addition, >> > we have supported clear_refs feacture. >> > IOW, it could be reset anytime so it's hard to use marker for >> > lazy freeing at the moment. >> > >> >> That's unfortunate. I think that the ABI would be much nicer if it >> used the accessed bit. >> >> In any case, shouldn't the aging algorithm be irrelevant here? A >> MADV_FREE page that isn't accessed can be discarded, whereas we could >> hopefully just say that a MADV_FREE page that is accessed gets moved >> to whatever list holds recently accessed pages and also stops being a >> candidate for discarding due to MADV_FREE? > > I meant if we use access bit as indicator for lazy-freeing page, > we could discard valid page which is never hinted by MADV_FREE but > just doesn't mark access bit in page table by aging algorithm. Oh, is the rule that the anonymous pages that are clean are discarded instead of swapped out? That is, does your patch set detect that an anonymous page can be discarded if it's clean and that the lack of a dirty bit is the only indication that the page has been hit with MADV_FREE? If so, that seems potentially error prone -- I had assumed that pages that were swapped in but not written since swap-in would also be clean, and I don't see how you distinguish them. --Andy -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2015-11-05 02:50 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrhQB-8ni-3@gated-at.bofh.it> |
| In reply to | #1262800 |
On Wed, Nov 04, 2015 at 05:29:57PM -0800, Andy Lutomirski wrote: > On Wed, Nov 4, 2015 at 4:56 PM, Minchan Kim <minchan@kernel.org> wrote: > > On Wed, Nov 04, 2015 at 04:42:37PM -0800, Andy Lutomirski wrote: > >> On Wed, Nov 4, 2015 at 4:13 PM, Minchan Kim <minchan@kernel.org> wrote: > >> > On Tue, Nov 03, 2015 at 07:41:35PM -0800, Andy Lutomirski wrote: > >> >> On Nov 3, 2015 5:30 PM, "Minchan Kim" <minchan@kernel.org> wrote: > >> >> > > >> >> > Linux doesn't have an ability to free pages lazy while other OS already > >> >> > have been supported that named by madvise(MADV_FREE). > >> >> > > >> >> > The gain is clear that kernel can discard freed pages rather than swapping > >> >> > out or OOM if memory pressure happens. > >> >> > > >> >> > Without memory pressure, freed pages would be reused by userspace without > >> >> > another additional overhead(ex, page fault + allocation + zeroing). > >> >> > > >> >> > >> >> [...] > >> >> > >> >> > > >> >> > How it works: > >> >> > > >> >> > When madvise syscall is called, VM clears dirty bit of ptes of the range. > >> >> > If memory pressure happens, VM checks dirty bit of page table and if it > >> >> > found still "clean", it means it's a "lazyfree pages" so VM could discard > >> >> > the page instead of swapping out. Once there was store operation for the > >> >> > page before VM peek a page to reclaim, dirty bit is set so VM can swap out > >> >> > the page instead of discarding. > >> >> > >> >> What happens if you MADV_FREE something that's MAP_SHARED or isn't > >> >> ordinary anonymous memory? There's a long history of MADV_DONTNEED on > >> >> such mappings causing exploitable problems, and I think it would be > >> >> nice if MADV_FREE were obviously safe. > >> > > >> > It filter out VM_LOCKED|VM_HUGETLB|VM_PFNMAP and file-backed vma and MAP_SHARED > >> > with vma_is_anonymous. > >> > > >> >> > >> >> Does this set the write protect bit? > >> > > >> > No. > >> > > >> >> > >> >> What happens on architectures without hardware dirty tracking? For > >> >> that matter, even on architecture with hardware dirty tracking, what > >> >> happens in multithreaded processes that have the dirty TLB state > >> >> cached in a different CPU's TLB? > >> >> > >> >> Using the dirty bit for these semantics scares me. This API creates a > >> >> page that can have visible nonzero contents and then can > >> >> asynchronously and magically zero itself thereafter. That makes me > >> >> nervous. Could we use the accessed bit instead? Then the observable > >> > > >> > Access bit is used by aging algorithm for reclaim. In addition, > >> > we have supported clear_refs feacture. > >> > IOW, it could be reset anytime so it's hard to use marker for > >> > lazy freeing at the moment. > >> > > >> > >> That's unfortunate. I think that the ABI would be much nicer if it > >> used the accessed bit. > >> > >> In any case, shouldn't the aging algorithm be irrelevant here? A > >> MADV_FREE page that isn't accessed can be discarded, whereas we could > >> hopefully just say that a MADV_FREE page that is accessed gets moved > >> to whatever list holds recently accessed pages and also stops being a > >> candidate for discarding due to MADV_FREE? > > > > I meant if we use access bit as indicator for lazy-freeing page, > > we could discard valid page which is never hinted by MADV_FREE but > > just doesn't mark access bit in page table by aging algorithm. > > Oh, is the rule that the anonymous pages that are clean are discarded > instead of swapped out? That is, does your patch set detect that an The page swapped-in after swapped-out has clean pte and swap device has valid data if the page isn't touch so VM discards the page rather than swapout. Of course, pte should point out the swap slot. If VM decide to remove the page from swap slot, it should be marked PG_dirty. > anonymous page can be discarded if it's clean and that the lack of a > dirty bit is the only indication that the page has been hit with > MADV_FREE? No dirty bit, exactly speaking, PG_Dirty because the page I mentioned above has clean pte but will have PG_dirty. > > If so, that seems potentially error prone -- I had assumed that pages > that were swapped in but not written since swap-in would also be > clean, and I don't see how you distinguish them. I hope above will answer. > > --Andy -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Shaohua Li <shli@kernel.org> |
|---|---|
| Date | 2015-11-04 21:10 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrcxA-59d-7@gated-at.bofh.it> |
| In reply to | #1262021 |
On Wed, Nov 04, 2015 at 10:25:55AM +0900, Minchan Kim wrote: > Linux doesn't have an ability to free pages lazy while other OS already > have been supported that named by madvise(MADV_FREE). > > The gain is clear that kernel can discard freed pages rather than swapping > out or OOM if memory pressure happens. > > Without memory pressure, freed pages would be reused by userspace without > another additional overhead(ex, page fault + allocation + zeroing). > > Jason Evans said: > > : Facebook has been using MAP_UNINITIALIZED > : (https://lkml.org/lkml/2012/1/18/308) in some of its applications for > : several years, but there are operational costs to maintaining this > : out-of-tree in our kernel and in jemalloc, and we are anxious to retire it > : in favor of MADV_FREE. When we first enabled MAP_UNINITIALIZED it > : increased throughput for much of our workload by ~5%, and although the > : benefit has decreased using newer hardware and kernels, there is still > : enough benefit that we cannot reasonably retire it without a replacement. > : > : Aside from Facebook operations, there are numerous broadly used > : applications that would benefit from MADV_FREE. The ones that immediately > : come to mind are redis, varnish, and MariaDB. I don't have much insight > : into Android internals and development process, but I would hope to see > : MADV_FREE support eventually end up there as well to benefit applications > : linked with the integrated jemalloc. > : > : jemalloc will use MADV_FREE once it becomes available in the Linux kernel. > : In fact, jemalloc already uses MADV_FREE or equivalent everywhere it's > : available: *BSD, OS X, Windows, and Solaris -- every platform except Linux > : (and AIX, but I'm not sure it even compiles on AIX). The lack of > : MADV_FREE on Linux forced me down a long series of increasingly > : sophisticated heuristics for madvise() volume reduction, and even so this > : remains a common performance issue for people using jemalloc on Linux. > : Please integrate MADV_FREE; many people will benefit substantially. > > How it works: > > When madvise syscall is called, VM clears dirty bit of ptes of the range. > If memory pressure happens, VM checks dirty bit of page table and if it > found still "clean", it means it's a "lazyfree pages" so VM could discard > the page instead of swapping out. Once there was store operation for the > page before VM peek a page to reclaim, dirty bit is set so VM can swap out > the page instead of discarding. > > Firstly, heavy users would be general allocators(ex, jemalloc, tcmalloc > and hope glibc supports it) and jemalloc/tcmalloc already have supported > the feature for other OS(ex, FreeBSD) > > barrios@blaptop:~/benchmark/ebizzy$ lscpu > Architecture: x86_64 > CPU op-mode(s): 32-bit, 64-bit > Byte Order: Little Endian > CPU(s): 12 > On-line CPU(s) list: 0-11 > Thread(s) per core: 1 > Core(s) per socket: 1 > Socket(s): 12 > NUMA node(s): 1 > Vendor ID: GenuineIntel > CPU family: 6 > Model: 2 > Stepping: 3 > CPU MHz: 3200.185 > BogoMIPS: 6400.53 > Virtualization: VT-x > Hypervisor vendor: KVM > Virtualization type: full > L1d cache: 32K > L1i cache: 32K > L2 cache: 4096K > NUMA node0 CPU(s): 0-11 > ebizzy benchmark(./ebizzy -S 10 -n 512) > > Higher avg is better. > > vanilla-jemalloc MADV_free-jemalloc > > 1 thread > records: 10 records: 10 > avg: 2961.90 avg: 12069.70 > std: 71.96(2.43%) std: 186.68(1.55%) > max: 3070.00 max: 12385.00 > min: 2796.00 min: 11746.00 > > 2 thread > records: 10 records: 10 > avg: 5020.00 avg: 17827.00 > std: 264.87(5.28%) std: 358.52(2.01%) > max: 5244.00 max: 18760.00 > min: 4251.00 min: 17382.00 > > 4 thread > records: 10 records: 10 > avg: 8988.80 avg: 27930.80 > std: 1175.33(13.08%) std: 3317.33(11.88%) > max: 9508.00 max: 30879.00 > min: 5477.00 min: 21024.00 > > 8 thread > records: 10 records: 10 > avg: 13036.50 avg: 33739.40 > std: 170.67(1.31%) std: 5146.22(15.25%) > max: 13371.00 max: 40572.00 > min: 12785.00 min: 24088.00 > > 16 thread > records: 10 records: 10 > avg: 11092.40 avg: 31424.20 > std: 710.60(6.41%) std: 3763.89(11.98%) > max: 12446.00 max: 36635.00 > min: 9949.00 min: 25669.00 > > 32 thread > records: 10 records: 10 > avg: 11067.00 avg: 34495.80 > std: 971.06(8.77%) std: 2721.36(7.89%) > max: 12010.00 max: 38598.00 > min: 9002.00 min: 30636.00 > > In summary, MADV_FREE is about much faster than MADV_DONTNEED. The MADV_FREE is discussed for a while, it probably is too late to propose something new, but we had the new idea (from Ben Maurer, CCed) recently and think it's better. Our target is still jemalloc. Compared to MADV_DONTNEED, MADV_FREE's lazy memory free is a huge win to reduce page fault. But there is one issue remaining, the TLB flush. Both MADV_DONTNEED and MADV_FREE do TLB flush. TLB flush overhead is quite big in contemporary multi-thread applications. In our production workload, we observed 80% CPU spending on TLB flush triggered by jemalloc madvise(MADV_DONTNEED) sometimes. We haven't tested MADV_FREE yet, but the result should be similar. It's hard to avoid the TLB flush issue with MADV_FREE, because it helps avoid data corruption. The new proposal tries to fix the TLB issue. We introduce two madvise verbs: MARK_FREE. Userspace notifies kernel the memory range can be discarded. Kernel just records the range in current stage. Should memory pressure happen, page reclaim can free the memory directly regardless the pte state. MARK_NOFREE. Userspace notifies kernel the memory range will be reused soon. Kernel deletes the record and prevents page reclaim discards the memory. If the memory isn't reclaimed, userspace will access the old memory, otherwise do normal page fault handling. The point is to let userspace notify kernel if memory can be discarded, instead of depending on pte dirty bit used by MADV_FREE. With these, no TLB flush is required till page reclaim actually frees the memory (page reclaim need do the TLB flush for MADV_FREE too). It still preserves the lazy memory free merit of MADV_FREE. Compared to MADV_FREE, reusing memory with the new proposal isn't transparent, eg must call MARK_NOFREE. But it's easy to utilize the new API in jemalloc. We don't have code to backup this yet, sorry. We'd like to discuss it if it makes sense. Thanks, Shaohua -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Daniel Micay <danielmicay@gmail.com> |
|---|---|
| Date | 2015-11-04 22:20 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrdDk-5P5-13@gated-at.bofh.it> |
| In reply to | #1262558 |
[Multipart message — attachments visible in raw view] — view raw
> Compared to MADV_DONTNEED, MADV_FREE's lazy memory free is a huge win to reduce > page fault. But there is one issue remaining, the TLB flush. Both MADV_DONTNEED > and MADV_FREE do TLB flush. TLB flush overhead is quite big in contemporary > multi-thread applications. In our production workload, we observed 80% CPU > spending on TLB flush triggered by jemalloc madvise(MADV_DONTNEED) sometimes. > We haven't tested MADV_FREE yet, but the result should be similar. It's hard to > avoid the TLB flush issue with MADV_FREE, because it helps avoid data > corruption. > > The new proposal tries to fix the TLB issue. We introduce two madvise verbs: > > MARK_FREE. Userspace notifies kernel the memory range can be discarded. Kernel > just records the range in current stage. Should memory pressure happen, page > reclaim can free the memory directly regardless the pte state. > > MARK_NOFREE. Userspace notifies kernel the memory range will be reused soon. > Kernel deletes the record and prevents page reclaim discards the memory. If the > memory isn't reclaimed, userspace will access the old memory, otherwise do > normal page fault handling. > > The point is to let userspace notify kernel if memory can be discarded, instead > of depending on pte dirty bit used by MADV_FREE. With these, no TLB flush is > required till page reclaim actually frees the memory (page reclaim need do the > TLB flush for MADV_FREE too). It still preserves the lazy memory free merit of > MADV_FREE. > > Compared to MADV_FREE, reusing memory with the new proposal isn't transparent, > eg must call MARK_NOFREE. But it's easy to utilize the new API in jemalloc. > > We don't have code to backup this yet, sorry. We'd like to discuss it if it > makes sense. That's comparable to Android's pinning / unpinning API for ashmem and I think it makes sense if it's faster. It's different than the MADV_FREE API though, because the new allocations that are handed out won't have the usual lazy commit which MADV_FREE provides. Pages in an allocation that's handed out can still be dropped until they are actually written to. It's considered active by jemalloc either way, but only a subset of the active pages are actually committed. There's probably a use case for both of these systems.
[toc] | [prev] | [next] | [standalone]
| From | Daniel Micay <danielmicay@gmail.com> |
|---|---|
| Date | 2015-11-04 22:30 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrdN0-5Sz-13@gated-at.bofh.it> |
| In reply to | #1262604 |
[Multipart message — attachments visible in raw view] — view raw
> That's comparable to Android's pinning / unpinning API for ashmem and I > think it makes sense if it's faster. It's different than the MADV_FREE > API though, because the new allocations that are handed out won't have > the usual lazy commit which MADV_FREE provides. Pages in an allocation > that's handed out can still be dropped until they are actually written > to. It's considered active by jemalloc either way, but only a subset of > the active pages are actually committed. There's probably a use case for > both of these systems. Also, consider that MADV_FREE would allow jemalloc to be extremely aggressive with purging when it actually has to do it. It can start with the largest span of memory and it can mark more than strictly necessary to drop below the ratio as there's no cost to using the memory again (not even a system call). Since the main cost is using the system call at all, there's going to be pressure to mark the largest possible spans in one go. It will mean concentration on memory compaction will improve performance. I think that's the right direction for the kernel to be guiding userspace. It will play better with THP than the allocator trying to be very precise with purging based on aging.
[toc] | [prev] | [next] | [standalone]
| From | Andy Lutomirski <luto@amacapital.net> |
|---|---|
| Date | 2015-11-04 22:50 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qre6n-5Zh-37@gated-at.bofh.it> |
| In reply to | #1262558 |
On Wed, Nov 4, 2015 at 12:00 PM, Shaohua Li <shli@kernel.org> wrote: > > The new proposal tries to fix the TLB issue. We introduce two madvise verbs: > > MARK_FREE. Userspace notifies kernel the memory range can be discarded. Kernel > just records the range in current stage. Should memory pressure happen, page > reclaim can free the memory directly regardless the pte state. > > MARK_NOFREE. Userspace notifies kernel the memory range will be reused soon. > Kernel deletes the record and prevents page reclaim discards the memory. If the > memory isn't reclaimed, userspace will access the old memory, otherwise do > normal page fault handling. > > The point is to let userspace notify kernel if memory can be discarded, instead > of depending on pte dirty bit used by MADV_FREE. With these, no TLB flush is > required till page reclaim actually frees the memory (page reclaim need do the > TLB flush for MADV_FREE too). It still preserves the lazy memory free merit of > MADV_FREE. > > Compared to MADV_FREE, reusing memory with the new proposal isn't transparent, > eg must call MARK_NOFREE. But it's easy to utilize the new API in jemalloc. > I can't speak to the usefulness of this or to other arches, but on x86 (unless you have nohz_full or similar enabled), a pair of syscalls should be *much* faster than an IPI or a page fault. I don't know how expensive it is to write to a clean page or to access an unaccessed page on x86. I'm sure it's not free (there's memory bandwidth if nothing else), but it could be very cheap. --Andy -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2015-11-05 02:40 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrhGV-8jJ-1@gated-at.bofh.it> |
| In reply to | #1262558 |
On Wed, Nov 04, 2015 at 12:00:06PM -0800, Shaohua Li wrote: > On Wed, Nov 04, 2015 at 10:25:55AM +0900, Minchan Kim wrote: > > Linux doesn't have an ability to free pages lazy while other OS already > > have been supported that named by madvise(MADV_FREE). > > > > The gain is clear that kernel can discard freed pages rather than swapping > > out or OOM if memory pressure happens. > > > > Without memory pressure, freed pages would be reused by userspace without > > another additional overhead(ex, page fault + allocation + zeroing). > > > > Jason Evans said: > > > > : Facebook has been using MAP_UNINITIALIZED > > : (https://lkml.org/lkml/2012/1/18/308) in some of its applications for > > : several years, but there are operational costs to maintaining this > > : out-of-tree in our kernel and in jemalloc, and we are anxious to retire it > > : in favor of MADV_FREE. When we first enabled MAP_UNINITIALIZED it > > : increased throughput for much of our workload by ~5%, and although the > > : benefit has decreased using newer hardware and kernels, there is still > > : enough benefit that we cannot reasonably retire it without a replacement. > > : > > : Aside from Facebook operations, there are numerous broadly used > > : applications that would benefit from MADV_FREE. The ones that immediately > > : come to mind are redis, varnish, and MariaDB. I don't have much insight > > : into Android internals and development process, but I would hope to see > > : MADV_FREE support eventually end up there as well to benefit applications > > : linked with the integrated jemalloc. > > : > > : jemalloc will use MADV_FREE once it becomes available in the Linux kernel. > > : In fact, jemalloc already uses MADV_FREE or equivalent everywhere it's > > : available: *BSD, OS X, Windows, and Solaris -- every platform except Linux > > : (and AIX, but I'm not sure it even compiles on AIX). The lack of > > : MADV_FREE on Linux forced me down a long series of increasingly > > : sophisticated heuristics for madvise() volume reduction, and even so this > > : remains a common performance issue for people using jemalloc on Linux. > > : Please integrate MADV_FREE; many people will benefit substantially. > > > > How it works: > > > > When madvise syscall is called, VM clears dirty bit of ptes of the range. > > If memory pressure happens, VM checks dirty bit of page table and if it > > found still "clean", it means it's a "lazyfree pages" so VM could discard > > the page instead of swapping out. Once there was store operation for the > > page before VM peek a page to reclaim, dirty bit is set so VM can swap out > > the page instead of discarding. > > > > Firstly, heavy users would be general allocators(ex, jemalloc, tcmalloc > > and hope glibc supports it) and jemalloc/tcmalloc already have supported > > the feature for other OS(ex, FreeBSD) > > > > barrios@blaptop:~/benchmark/ebizzy$ lscpu > > Architecture: x86_64 > > CPU op-mode(s): 32-bit, 64-bit > > Byte Order: Little Endian > > CPU(s): 12 > > On-line CPU(s) list: 0-11 > > Thread(s) per core: 1 > > Core(s) per socket: 1 > > Socket(s): 12 > > NUMA node(s): 1 > > Vendor ID: GenuineIntel > > CPU family: 6 > > Model: 2 > > Stepping: 3 > > CPU MHz: 3200.185 > > BogoMIPS: 6400.53 > > Virtualization: VT-x > > Hypervisor vendor: KVM > > Virtualization type: full > > L1d cache: 32K > > L1i cache: 32K > > L2 cache: 4096K > > NUMA node0 CPU(s): 0-11 > > ebizzy benchmark(./ebizzy -S 10 -n 512) > > > > Higher avg is better. > > > > vanilla-jemalloc MADV_free-jemalloc > > > > 1 thread > > records: 10 records: 10 > > avg: 2961.90 avg: 12069.70 > > std: 71.96(2.43%) std: 186.68(1.55%) > > max: 3070.00 max: 12385.00 > > min: 2796.00 min: 11746.00 > > > > 2 thread > > records: 10 records: 10 > > avg: 5020.00 avg: 17827.00 > > std: 264.87(5.28%) std: 358.52(2.01%) > > max: 5244.00 max: 18760.00 > > min: 4251.00 min: 17382.00 > > > > 4 thread > > records: 10 records: 10 > > avg: 8988.80 avg: 27930.80 > > std: 1175.33(13.08%) std: 3317.33(11.88%) > > max: 9508.00 max: 30879.00 > > min: 5477.00 min: 21024.00 > > > > 8 thread > > records: 10 records: 10 > > avg: 13036.50 avg: 33739.40 > > std: 170.67(1.31%) std: 5146.22(15.25%) > > max: 13371.00 max: 40572.00 > > min: 12785.00 min: 24088.00 > > > > 16 thread > > records: 10 records: 10 > > avg: 11092.40 avg: 31424.20 > > std: 710.60(6.41%) std: 3763.89(11.98%) > > max: 12446.00 max: 36635.00 > > min: 9949.00 min: 25669.00 > > > > 32 thread > > records: 10 records: 10 > > avg: 11067.00 avg: 34495.80 > > std: 971.06(8.77%) std: 2721.36(7.89%) > > max: 12010.00 max: 38598.00 > > min: 9002.00 min: 30636.00 > > > > In summary, MADV_FREE is about much faster than MADV_DONTNEED. > > The MADV_FREE is discussed for a while, it probably is too late to propose > something new, but we had the new idea (from Ben Maurer, CCed) recently and > think it's better. Our target is still jemalloc. > > Compared to MADV_DONTNEED, MADV_FREE's lazy memory free is a huge win to reduce > page fault. But there is one issue remaining, the TLB flush. Both MADV_DONTNEED > and MADV_FREE do TLB flush. TLB flush overhead is quite big in contemporary > multi-thread applications. In our production workload, we observed 80% CPU > spending on TLB flush triggered by jemalloc madvise(MADV_DONTNEED) sometimes. > We haven't tested MADV_FREE yet, but the result should be similar. It's hard to > avoid the TLB flush issue with MADV_FREE, because it helps avoid data > corruption. > > The new proposal tries to fix the TLB issue. We introduce two madvise verbs: > > MARK_FREE. Userspace notifies kernel the memory range can be discarded. Kernel > just records the range in current stage. Should memory pressure happen, page > reclaim can free the memory directly regardless the pte state. > > MARK_NOFREE. Userspace notifies kernel the memory range will be reused soon. > Kernel deletes the record and prevents page reclaim discards the memory. If the > memory isn't reclaimed, userspace will access the old memory, otherwise do > normal page fault handling. > > The point is to let userspace notify kernel if memory can be discarded, instead > of depending on pte dirty bit used by MADV_FREE. With these, no TLB flush is > required till page reclaim actually frees the memory (page reclaim need do the > TLB flush for MADV_FREE too). It still preserves the lazy memory free merit of > MADV_FREE. > > Compared to MADV_FREE, reusing memory with the new proposal isn't transparent, > eg must call MARK_NOFREE. But it's easy to utilize the new API in jemalloc. > > We don't have code to backup this yet, sorry. We'd like to discuss it if it > makes sense. It's really what volatile range did. John Stultz and me tried it for a *long* time but it had lots of troubles. It's really hard to write it down in my time due to really long history and even I forgot lots of detail(ie, dead brain). Please search volatile ranges in google. Finally, people in LSF/MM suggested MADV_FREE to help anonymous page side rather than stucking hich prevent useful feature. :( > > Thanks, > Shaohua -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Minchan Kim <minchan@kernel.org> |
|---|---|
| Date | 2015-11-05 02:40 +0100 |
| Subject | Re: [PATCH v2 01/13] mm: support madvise(MADV_FREE) |
| Message-ID | <qrhGV-8jJ-9@gated-at.bofh.it> |
| In reply to | #1262796 |
On Thu, Nov 05, 2015 at 10:33:50AM +0900, Minchan Kim wrote: > On Wed, Nov 04, 2015 at 12:00:06PM -0800, Shaohua Li wrote: > > On Wed, Nov 04, 2015 at 10:25:55AM +0900, Minchan Kim wrote: > > > Linux doesn't have an ability to free pages lazy while other OS already > > > have been supported that named by madvise(MADV_FREE). > > > > > > The gain is clear that kernel can discard freed pages rather than swapping > > > out or OOM if memory pressure happens. > > > > > > Without memory pressure, freed pages would be reused by userspace without > > > another additional overhead(ex, page fault + allocation + zeroing). > > > > > > Jason Evans said: > > > > > > : Facebook has been using MAP_UNINITIALIZED > > > : (https://lkml.org/lkml/2012/1/18/308) in some of its applications for > > > : several years, but there are operational costs to maintaining this > > > : out-of-tree in our kernel and in jemalloc, and we are anxious to retire it > > > : in favor of MADV_FREE. When we first enabled MAP_UNINITIALIZED it > > > : increased throughput for much of our workload by ~5%, and although the > > > : benefit has decreased using newer hardware and kernels, there is still > > > : enough benefit that we cannot reasonably retire it without a replacement. > > > : > > > : Aside from Facebook operations, there are numerous broadly used > > > : applications that would benefit from MADV_FREE. The ones that immediately > > > : come to mind are redis, varnish, and MariaDB. I don't have much insight > > > : into Android internals and development process, but I would hope to see > > > : MADV_FREE support eventually end up there as well to benefit applications > > > : linked with the integrated jemalloc. > > > : > > > : jemalloc will use MADV_FREE once it becomes available in the Linux kernel. > > > : In fact, jemalloc already uses MADV_FREE or equivalent everywhere it's > > > : available: *BSD, OS X, Windows, and Solaris -- every platform except Linux > > > : (and AIX, but I'm not sure it even compiles on AIX). The lack of > > > : MADV_FREE on Linux forced me down a long series of increasingly > > > : sophisticated heuristics for madvise() volume reduction, and even so this > > > : remains a common performance issue for people using jemalloc on Linux. > > > : Please integrate MADV_FREE; many people will benefit substantially. > > > > > > How it works: > > > > > > When madvise syscall is called, VM clears dirty bit of ptes of the range. > > > If memory pressure happens, VM checks dirty bit of page table and if it > > > found still "clean", it means it's a "lazyfree pages" so VM could discard > > > the page instead of swapping out. Once there was store operation for the > > > page before VM peek a page to reclaim, dirty bit is set so VM can swap out > > > the page instead of discarding. > > > > > > Firstly, heavy users would be general allocators(ex, jemalloc, tcmalloc > > > and hope glibc supports it) and jemalloc/tcmalloc already have supported > > > the feature for other OS(ex, FreeBSD) > > > > > > barrios@blaptop:~/benchmark/ebizzy$ lscpu > > > Architecture: x86_64 > > > CPU op-mode(s): 32-bit, 64-bit > > > Byte Order: Little Endian > > > CPU(s): 12 > > > On-line CPU(s) list: 0-11 > > > Thread(s) per core: 1 > > > Core(s) per socket: 1 > > > Socket(s): 12 > > > NUMA node(s): 1 > > > Vendor ID: GenuineIntel > > > CPU family: 6 > > > Model: 2 > > > Stepping: 3 > > > CPU MHz: 3200.185 > > > BogoMIPS: 6400.53 > > > Virtualization: VT-x > > > Hypervisor vendor: KVM > > > Virtualization type: full > > > L1d cache: 32K > > > L1i cache: 32K > > > L2 cache: 4096K > > > NUMA node0 CPU(s): 0-11 > > > ebizzy benchmark(./ebizzy -S 10 -n 512) > > > > > > Higher avg is better. > > > > > > vanilla-jemalloc MADV_free-jemalloc > > > > > > 1 thread > > > records: 10 records: 10 > > > avg: 2961.90 avg: 12069.70 > > > std: 71.96(2.43%) std: 186.68(1.55%) > > > max: 3070.00 max: 12385.00 > > > min: 2796.00 min: 11746.00 > > > > > > 2 thread > > > records: 10 records: 10 > > > avg: 5020.00 avg: 17827.00 > > > std: 264.87(5.28%) std: 358.52(2.01%) > > > max: 5244.00 max: 18760.00 > > > min: 4251.00 min: 17382.00 > > > > > > 4 thread > > > records: 10 records: 10 > > > avg: 8988.80 avg: 27930.80 > > > std: 1175.33(13.08%) std: 3317.33(11.88%) > > > max: 9508.00 max: 30879.00 > > > min: 5477.00 min: 21024.00 > > > > > > 8 thread > > > records: 10 records: 10 > > > avg: 13036.50 avg: 33739.40 > > > std: 170.67(1.31%) std: 5146.22(15.25%) > > > max: 13371.00 max: 40572.00 > > > min: 12785.00 min: 24088.00 > > > > > > 16 thread > > > records: 10 records: 10 > > > avg: 11092.40 avg: 31424.20 > > > std: 710.60(6.41%) std: 3763.89(11.98%) > > > max: 12446.00 max: 36635.00 > > > min: 9949.00 min: 25669.00 > > > > > > 32 thread > > > records: 10 records: 10 > > > avg: 11067.00 avg: 34495.80 > > > std: 971.06(8.77%) std: 2721.36(7.89%) > > > max: 12010.00 max: 38598.00 > > > min: 9002.00 min: 30636.00 > > > > > > In summary, MADV_FREE is about much faster than MADV_DONTNEED. > > > > The MADV_FREE is discussed for a while, it probably is too late to propose > > something new, but we had the new idea (from Ben Maurer, CCed) recently and > > think it's better. Our target is still jemalloc. > > > > Compared to MADV_DONTNEED, MADV_FREE's lazy memory free is a huge win to reduce > > page fault. But there is one issue remaining, the TLB flush. Both MADV_DONTNEED > > and MADV_FREE do TLB flush. TLB flush overhead is quite big in contemporary > > multi-thread applications. In our production workload, we observed 80% CPU > > spending on TLB flush triggered by jemalloc madvise(MADV_DONTNEED) sometimes. > > We haven't tested MADV_FREE yet, but the result should be similar. It's hard to > > avoid the TLB flush issue with MADV_FREE, because it helps avoid data > > corruption. > > > > The new proposal tries to fix the TLB issue. We introduce two madvise verbs: > > > > MARK_FREE. Userspace notifies kernel the memory range can be discarded. Kernel > > just records the range in current stage. Should memory pressure happen, page > > reclaim can free the memory directly regardless the pte state. > > > > MARK_NOFREE. Userspace notifies kernel the memory range will be reused soon. > > Kernel deletes the record and prevents page reclaim discards the memory. If the > > memory isn't reclaimed, userspace will access the old memory, otherwise do > > normal page fault handling. > > > > The point is to let userspace notify kernel if memory can be discarded, instead > > of depending on pte dirty bit used by MADV_FREE. With these, no TLB flush is > > required till page reclaim actually frees the memory (page reclaim need do the > > TLB flush for MADV_FREE too). It still preserves the lazy memory free merit of > > MADV_FREE. > > > > Compared to MADV_FREE, reusing memory with the new proposal isn't transparent, > > eg must call MARK_NOFREE. But it's easy to utilize the new API in jemalloc. > > > > We don't have code to backup this yet, sorry. We'd like to discuss it if it > > makes sense. > > It's really what volatile range did. > John Stultz and me tried it for a *long* time but it had lots of troubles. > It's really hard to write it down in my time due to really long history > and even I forgot lots of detail(ie, dead brain). > Please search volatile ranges in google. > Finally, people in LSF/MM suggested MADV_FREE to help anonymous page side > rather than stucking hich prevent useful feature. :( I should have Cced John Stutlz. He would have good memory than me so he would help but I'm not sure he has a interest on volatile ranges, still. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [standalone]
Page 2 of 2 — ← Prev page 1 [2]
Back to top | Article view | linux.kernel
csiph-web