Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1642158 > unrolled thread

[PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

Started byjs1304@gmail.com
First post2017-05-16 03:20 +0200
Last post2017-05-24 08:20 +0200
Articles 20 on this page of 40 — 4 participants

Back to article view | Back to linux.kernel


Contents

  [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption js1304@gmail.com - 2017-05-16 03:20 +0200
    [PATCH v1 09/11] x86/kasan: support on-demand shadow mapping js1304@gmail.com - 2017-05-16 03:20 +0200
    [PATCH v1 11/11] mm/kasan: change the order of shadow memory check js1304@gmail.com - 2017-05-16 03:20 +0200
    [PATCH v1 10/11] mm/kasan: support dynamic shadow memory free js1304@gmail.com - 2017-05-16 03:20 +0200
    [PATCH v1 05/11] mm/kasan: introduce per-page shadow memory infrastructure js1304@gmail.com - 2017-05-16 03:20 +0200
    [PATCH v1 08/11] mm/kasan: support on-demand shadow allocation/mapping js1304@gmail.com - 2017-05-16 03:20 +0200
    [PATCH v1 02/11] mm/kasan: don't fetch the next shadow value speculartively js1304@gmail.com - 2017-05-16 03:20 +0200
    [PATCH v1 07/11] x86/kasan: use per-page shadow memory js1304@gmail.com - 2017-05-16 03:20 +0200
    [PATCH(RE-RESEND) v1 01/11] mm/kasan: rename _is_zero to _is_nonzero Joonsoo Kim <js1304@gmail.com> - 2017-05-16 03:30 +0200
    Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-16 06:40 +0200
      Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-16 06:50 +0200
      Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-16 08:30 +0200
        Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-16 22:50 +0200
          Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-17 09:30 +0200
          Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-17 09:30 +0200
          Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-24 09:00 +0200
            Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-24 09:50 +0200
              Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-24 19:30 +0200
                Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-25 02:50 +0200
                  Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-29 17:10 +0200
                    Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-29 17:20 +0200
                      Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-29 17:40 +0200
                        Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 10:00 +0200
                          Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-30 10:20 +0200
                            Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 10:40 +0200
                              Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 10:50 +0200
                                Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-30 10:50 +0200
                                  Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 11:10 +0200
                                    Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-30 11:30 +0200
                                      Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 11:40 +0200
                                        Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-30 11:50 +0200
                                          Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 12:00 +0200
                        Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-31 08:00 +0200
                        Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-06-01 20:10 +0200
    Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-19 04:00 +0200
      Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-22 08:10 +0200
        Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-24 08:10 +0200
          Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-24 18:40 +0200
            Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-25 02:50 +0200
      Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-24 08:20 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1652583 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-29 17:20 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tMuCB-1GC-9@gated-at.bofh.it>
In reply to#1652579
On Mon, May 29, 2017 at 5:07 PM, Dmitry Vyukov <dvyukov@google.com> wrote:
> On Thu, May 25, 2017 at 2:41 AM, Joonsoo Kim <js1304@gmail.com> wrote:
>> On Wed, May 24, 2017 at 07:19:50PM +0200, Dmitry Vyukov wrote:
>>> On Wed, May 24, 2017 at 9:45 AM, Joonsoo Kim <js1304@gmail.com> wrote:
>>> >> > What does make your current patch work then?
>>> >> > Say we map a new shadow page, update the page shadow to say that there
>>> >> > is mapped shadow. Then another CPU loads the page shadow and then
>>> >> > loads from the newly mapped shadow. If we don't flush TLB, what makes
>>> >> > the second CPU see the newly mapped shadow?
>>> >>
>>> >> /\/\/\/\/\/\
>>> >>
>>> >> Joonsoo, please answer this question above.
>>> >
>>> > Hello, I've answered it in another e-mail however it would not be
>>> > sufficient. I try again.
>>> >
>>> > If the page isn't used for kernel stack, slab, and global variable
>>> > (aka. kernel memory), black shadow is mapped for the page. We map a
>>> > new shadow page if the page will be used for kernel memory. We need to
>>> > flush TLB in all cpus when mapping a new shadow however it's not
>>> > possible in some cases. So, this patch does just flushing local cpu's
>>> > TLB. Another cpu could have stale TLB that points black shadow for
>>> > this page. If that cpu with stale TLB try to check vailidity of the
>>> > object on this page, result would be invalid since stale TLB points
>>> > the black shadow and it's shadow value is non-zero. We need a magic
>>> > here. At this moment, we cannot make sure if invalid is correct result
>>> > or not since we didn't do full TLB flush. So fixup processing is
>>> > started. It is implemented in check_memory_region_slow(). Flushing
>>> > local TLB and re-checking the shadow value. With flushing local TLB,
>>> > we will use fresh TLB at this time. Therefore, we can pass the
>>> > validity check as usual.
>>> >
>>> >> I am trying to understand if there is any chance to make mapping a
>>> >> single page for all non-interesting shadow ranges work. That would be
>>> >
>>> > This is what this patchset does. Mapping a single (zero/black) shadow
>>> > page for all non-interesting (non-kernel memory) shadow ranges.
>>> > There is only single instance of zero/black shadow page. On v1,
>>> > I used black shadow page only so fail to get enough performance. On
>>> > v2 mentioned in another thread, I use zero shadow for some region. I
>>> > guess that performance problem would be gone.
>>>
>>>
>>> I can't say I understand everything here, but after staring at the
>>> patch I don't understand why we need pshadow at all now. Especially
>>> with this commit
>>> https://github.com/JoonsooKim/linux/commit/be36ee65f185e3c4026fe93b633056ea811120fb.
>>> It seems that the current shadow is enough.
>>
>> pshadow exists for non-kernel memory like as page cache or anonymous page.
>> This patch doesn't map a new shadow (per-byte shadow) for those pages
>> to reduce memory consumption. However, we need to know if those page
>> are allocated or not in order to check the validity of access to those
>> page. We cannot utilize zero/black shadow page here since mapping
>> single zero/black shadow page represents eight real page's shadow
>> value. Instead, we use per-page shadow here and mark/unmark it when
>> allocation and free happens. With it, we can know the state of the
>> page and we can determine the validity of access to them.
>
> I see the problem with 8 kernel pages mapped to a single shadow page.
>
>
>>> If we see bad shadow when the actual shadow value is good, we fall
>>> onto slow path, flush tlb, reload shadow, see that it is good and
>>> return. Pshadow is not needed in this case.
>>
>> For the kernel memory, if we see bad shadow due to *stale TLB*, we
>> fall onto slow path (check_memory_region_slow()) and flush tlb and
>> reload shadow.
>>
>> For the non-kernel memory, if we see bad shadow, we fall onto
>> pshadow_val() check and we can see actual state of the page.
>>
>>> If we see good shadow when the actual shadow value is bad, we return
>>> immediately and get false negative. Pshadow is not involved as well.
>>> What am I missing?
>>
>> In this patchset, there is no case that we see good shadow when the
>> actual (p)shadow value is bad. This case should not happen since we
>> can miss actual error.
>
> But why is not it possible?
> Let's say we have a real shadow page allocated for range of kernel
> memory. Then we unmap the shadow page and map the back page (maybe
> even unmap the black page and map another real shadow page). Then
> another CPU reads shadow for this range. What prevents it from seeing
> the old shadow page?


Re the async processing in kasan_unmap_shadow_workfn. Can't it lead to
shadow corruption? It seems that it can cause unsynchronized state of
shadow pages and corresponding kernel pages in page alloc.
Consider that we schedule unmap of some pages in kasan_unmap_shadow.
Then the range is reallocated in page_alloc and we get into
kasan_map_shadow, which tries to map shadow for these pages again, but
since they are already mapped it bails out. Then
kasan_unmap_shadow_workfn starts and unmaps shadow for the range.

[toc] | [prev] | [next] | [standalone]


#1652599 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-29 17:40 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tMuVY-1P8-35@gated-at.bofh.it>
In reply to#1652583
On Mon, May 29, 2017 at 5:12 PM, Dmitry Vyukov <dvyukov@google.com> wrote:
>>>> >> > What does make your current patch work then?
>>>> >> > Say we map a new shadow page, update the page shadow to say that there
>>>> >> > is mapped shadow. Then another CPU loads the page shadow and then
>>>> >> > loads from the newly mapped shadow. If we don't flush TLB, what makes
>>>> >> > the second CPU see the newly mapped shadow?
>>>> >>
>>>> >> /\/\/\/\/\/\
>>>> >>
>>>> >> Joonsoo, please answer this question above.
>>>> >
>>>> > Hello, I've answered it in another e-mail however it would not be
>>>> > sufficient. I try again.
>>>> >
>>>> > If the page isn't used for kernel stack, slab, and global variable
>>>> > (aka. kernel memory), black shadow is mapped for the page. We map a
>>>> > new shadow page if the page will be used for kernel memory. We need to
>>>> > flush TLB in all cpus when mapping a new shadow however it's not
>>>> > possible in some cases. So, this patch does just flushing local cpu's
>>>> > TLB. Another cpu could have stale TLB that points black shadow for
>>>> > this page. If that cpu with stale TLB try to check vailidity of the
>>>> > object on this page, result would be invalid since stale TLB points
>>>> > the black shadow and it's shadow value is non-zero. We need a magic
>>>> > here. At this moment, we cannot make sure if invalid is correct result
>>>> > or not since we didn't do full TLB flush. So fixup processing is
>>>> > started. It is implemented in check_memory_region_slow(). Flushing
>>>> > local TLB and re-checking the shadow value. With flushing local TLB,
>>>> > we will use fresh TLB at this time. Therefore, we can pass the
>>>> > validity check as usual.
>>>> >
>>>> >> I am trying to understand if there is any chance to make mapping a
>>>> >> single page for all non-interesting shadow ranges work. That would be
>>>> >
>>>> > This is what this patchset does. Mapping a single (zero/black) shadow
>>>> > page for all non-interesting (non-kernel memory) shadow ranges.
>>>> > There is only single instance of zero/black shadow page. On v1,
>>>> > I used black shadow page only so fail to get enough performance. On
>>>> > v2 mentioned in another thread, I use zero shadow for some region. I
>>>> > guess that performance problem would be gone.
>>>>
>>>>
>>>> I can't say I understand everything here, but after staring at the
>>>> patch I don't understand why we need pshadow at all now. Especially
>>>> with this commit
>>>> https://github.com/JoonsooKim/linux/commit/be36ee65f185e3c4026fe93b633056ea811120fb.
>>>> It seems that the current shadow is enough.
>>>
>>> pshadow exists for non-kernel memory like as page cache or anonymous page.
>>> This patch doesn't map a new shadow (per-byte shadow) for those pages
>>> to reduce memory consumption. However, we need to know if those page
>>> are allocated or not in order to check the validity of access to those
>>> page. We cannot utilize zero/black shadow page here since mapping
>>> single zero/black shadow page represents eight real page's shadow
>>> value. Instead, we use per-page shadow here and mark/unmark it when
>>> allocation and free happens. With it, we can know the state of the
>>> page and we can determine the validity of access to them.
>>
>> I see the problem with 8 kernel pages mapped to a single shadow page.
>>
>>
>>>> If we see bad shadow when the actual shadow value is good, we fall
>>>> onto slow path, flush tlb, reload shadow, see that it is good and
>>>> return. Pshadow is not needed in this case.
>>>
>>> For the kernel memory, if we see bad shadow due to *stale TLB*, we
>>> fall onto slow path (check_memory_region_slow()) and flush tlb and
>>> reload shadow.
>>>
>>> For the non-kernel memory, if we see bad shadow, we fall onto
>>> pshadow_val() check and we can see actual state of the page.
>>>
>>>> If we see good shadow when the actual shadow value is bad, we return
>>>> immediately and get false negative. Pshadow is not involved as well.
>>>> What am I missing?
>>>
>>> In this patchset, there is no case that we see good shadow when the
>>> actual (p)shadow value is bad. This case should not happen since we
>>> can miss actual error.
>>
>> But why is not it possible?
>> Let's say we have a real shadow page allocated for range of kernel
>> memory. Then we unmap the shadow page and map the back page (maybe
>> even unmap the black page and map another real shadow page). Then
>> another CPU reads shadow for this range. What prevents it from seeing
>> the old shadow page?
>
>
> Re the async processing in kasan_unmap_shadow_workfn. Can't it lead to
> shadow corruption? It seems that it can cause unsynchronized state of
> shadow pages and corresponding kernel pages in page alloc.
> Consider that we schedule unmap of some pages in kasan_unmap_shadow.
> Then the range is reallocated in page_alloc and we get into
> kasan_map_shadow, which tries to map shadow for these pages again, but
> since they are already mapped it bails out. Then
> kasan_unmap_shadow_workfn starts and unmaps shadow for the range.


Joonsoo,

I guess mine (and Andrey's) main concern is the amount of additional
complexity (I am still struggling to understand how it all works) and
more arch-dependent code in exchange for moderate memory win.

Joonsoo, Andrey,

I have an alternative proposal. It should be conceptually simpler and
also less arch-dependent. But I don't know if I miss something
important that will render it non working.
Namely, we add a pointer to shadow to the page struct. Then, create a
slab allocator for 512B shadow blocks. Then, attach/detach these
shadow blocks to page structs as necessary. It should lead to even
smaller memory consumption because we won't need a whole shadow page
when only 1 out of 8 corresponding kernel pages are used (we will need
just a single 512B block). I guess with some fragmentation we need
lots of excessive shadow with the current proposed patch.
This does not depend on TLB in any way and does not require hooking
into buddy allocator.
The main downside is that we will need to be careful to not assume
that shadow is continuous. In particular this means that this mode
will work only with outline instrumentation and will need some ifdefs.
Also it will be slower due to the additional indirection when
accessing shadow, but that's meant as "small but slow" mode as far as
I understand.

But the main win as I see it is that that's basically complete support
for 32-bit arches. People do ask about arm32 support:
https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
and probably mips32 is relevant as well.
Such mode does not require a huge continuous address space range, has
minimal memory consumption and requires minimal arch-dependent code.
Works only with outline instrumentation, but I think that's a
reasonable compromise.

What do you think?

[toc] | [prev] | [next] | [standalone]


#1652912 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 10:00 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tMKem-4q1-9@gated-at.bofh.it>
In reply to#1652599
On 29/05/17 16:29, Dmitry Vyukov wrote:
> I have an alternative proposal. It should be conceptually simpler and
> also less arch-dependent. But I don't know if I miss something
> important that will render it non working.
> Namely, we add a pointer to shadow to the page struct. Then, create a
> slab allocator for 512B shadow blocks. Then, attach/detach these
> shadow blocks to page structs as necessary. It should lead to even
> smaller memory consumption because we won't need a whole shadow page
> when only 1 out of 8 corresponding kernel pages are used (we will need
> just a single 512B block). I guess with some fragmentation we need
> lots of excessive shadow with the current proposed patch.
> This does not depend on TLB in any way and does not require hooking
> into buddy allocator.
> The main downside is that we will need to be careful to not assume
> that shadow is continuous. In particular this means that this mode
> will work only with outline instrumentation and will need some ifdefs.
> Also it will be slower due to the additional indirection when
> accessing shadow, but that's meant as "small but slow" mode as far as
> I understand.
> 
> But the main win as I see it is that that's basically complete support
> for 32-bit arches. People do ask about arm32 support:
> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
> and probably mips32 is relevant as well.
> Such mode does not require a huge continuous address space range, has
> minimal memory consumption and requires minimal arch-dependent code.
> Works only with outline instrumentation, but I think that's a
> reasonable compromise.

.. or you can just keep shadow in page extension. It was suggested back in
2015 [1], but seems that lack of stack instrumentation was "no-way"... 

[1] https://lkml.org/lkml/2015/8/24/573 

Cheers
Vladimir

[toc] | [prev] | [next] | [standalone]


#1652928 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-30 10:20 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tMKxI-4OJ-13@gated-at.bofh.it>
In reply to#1652912
On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
<vladimir.murzin@arm.com> wrote:
> On 29/05/17 16:29, Dmitry Vyukov wrote:
>> I have an alternative proposal. It should be conceptually simpler and
>> also less arch-dependent. But I don't know if I miss something
>> important that will render it non working.
>> Namely, we add a pointer to shadow to the page struct. Then, create a
>> slab allocator for 512B shadow blocks. Then, attach/detach these
>> shadow blocks to page structs as necessary. It should lead to even
>> smaller memory consumption because we won't need a whole shadow page
>> when only 1 out of 8 corresponding kernel pages are used (we will need
>> just a single 512B block). I guess with some fragmentation we need
>> lots of excessive shadow with the current proposed patch.
>> This does not depend on TLB in any way and does not require hooking
>> into buddy allocator.
>> The main downside is that we will need to be careful to not assume
>> that shadow is continuous. In particular this means that this mode
>> will work only with outline instrumentation and will need some ifdefs.
>> Also it will be slower due to the additional indirection when
>> accessing shadow, but that's meant as "small but slow" mode as far as
>> I understand.
>>
>> But the main win as I see it is that that's basically complete support
>> for 32-bit arches. People do ask about arm32 support:
>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>> and probably mips32 is relevant as well.
>> Such mode does not require a huge continuous address space range, has
>> minimal memory consumption and requires minimal arch-dependent code.
>> Works only with outline instrumentation, but I think that's a
>> reasonable compromise.
>
> .. or you can just keep shadow in page extension. It was suggested back in
> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>
> [1] https://lkml.org/lkml/2015/8/24/573

Right. It describes basically the same idea.

How is page_ext better than adding data page struct?
It seems that memory for all page_ext is preallocated along with page
structs; but just the lookup is slower.

[toc] | [prev] | [next] | [standalone]


#1652950 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 10:40 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tMKR4-4Wf-29@gated-at.bofh.it>
In reply to#1652928
On 30/05/17 09:15, Dmitry Vyukov wrote:
> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
> <vladimir.murzin@arm.com> wrote:
>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>> I have an alternative proposal. It should be conceptually simpler and
>>> also less arch-dependent. But I don't know if I miss something
>>> important that will render it non working.
>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>> shadow blocks to page structs as necessary. It should lead to even
>>> smaller memory consumption because we won't need a whole shadow page
>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>> just a single 512B block). I guess with some fragmentation we need
>>> lots of excessive shadow with the current proposed patch.
>>> This does not depend on TLB in any way and does not require hooking
>>> into buddy allocator.
>>> The main downside is that we will need to be careful to not assume
>>> that shadow is continuous. In particular this means that this mode
>>> will work only with outline instrumentation and will need some ifdefs.
>>> Also it will be slower due to the additional indirection when
>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>> I understand.
>>>
>>> But the main win as I see it is that that's basically complete support
>>> for 32-bit arches. People do ask about arm32 support:
>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>> and probably mips32 is relevant as well.
>>> Such mode does not require a huge continuous address space range, has
>>> minimal memory consumption and requires minimal arch-dependent code.
>>> Works only with outline instrumentation, but I think that's a
>>> reasonable compromise.
>>
>> .. or you can just keep shadow in page extension. It was suggested back in
>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>
>> [1] https://lkml.org/lkml/2015/8/24/573
> 
> Right. It describes basically the same idea.
> 
> How is page_ext better than adding data page struct?

page_ext is already here along with some other debug options ;)

> It seems that memory for all page_ext is preallocated along with page
> structs; but just the lookup is slower.
> 

Yup. Lookup would look like (based on v4.0):

...
page_ext = lookup_page_ext_begin(virt_to_page(start));

do {
	page_ext->shadow[idx++] = value;
} while (idx < bound);

lookup_page_ext_end((void *)page_ext);

...

Cheers
Vladimir

[toc] | [prev] | [next] | [standalone]


#1652953 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 10:50 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tML0J-504-1@gated-at.bofh.it>
In reply to#1652950
On 30/05/17 09:31, Vladimir Murzin wrote:
> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
> 
> On 30/05/17 09:15, Dmitry Vyukov wrote:
>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>> <vladimir.murzin@arm.com> wrote:
>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>> I have an alternative proposal. It should be conceptually simpler and
>>>> also less arch-dependent. But I don't know if I miss something
>>>> important that will render it non working.
>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>> shadow blocks to page structs as necessary. It should lead to even
>>>> smaller memory consumption because we won't need a whole shadow page
>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>> just a single 512B block). I guess with some fragmentation we need
>>>> lots of excessive shadow with the current proposed patch.
>>>> This does not depend on TLB in any way and does not require hooking
>>>> into buddy allocator.
>>>> The main downside is that we will need to be careful to not assume
>>>> that shadow is continuous. In particular this means that this mode
>>>> will work only with outline instrumentation and will need some ifdefs.
>>>> Also it will be slower due to the additional indirection when
>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>> I understand.
>>>>
>>>> But the main win as I see it is that that's basically complete support
>>>> for 32-bit arches. People do ask about arm32 support:
>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>> and probably mips32 is relevant as well.
>>>> Such mode does not require a huge continuous address space range, has
>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>> Works only with outline instrumentation, but I think that's a
>>>> reasonable compromise.
>>>
>>> .. or you can just keep shadow in page extension. It was suggested back in
>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>
>>> [1] https://lkml.org/lkml/2015/8/24/573
>>
>> Right. It describes basically the same idea.
>>
>> How is page_ext better than adding data page struct?
> 
> page_ext is already here along with some other debug options ;)
> 
>> It seems that memory for all page_ext is preallocated along with page
>> structs; but just the lookup is slower.
>>
> 
> Yup. Lookup would look like (based on v4.0):
> 
> ...
> page_ext = lookup_page_ext_begin(virt_to_page(start));
> 
> do {
>         page_ext->shadow[idx++] = value;
> } while (idx < bound);
> 
> lookup_page_ext_end((void *)page_ext);
> 
> ...

Correction: please, ignore that *_{begin,end} stuff - mainline only
lookup_page_ext() is only used.

Cheers
Vladimir

> 
> Cheers
> Vladimir
> 
> 
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org.  For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
> 

[toc] | [prev] | [next] | [standalone]


#1652958 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-30 10:50 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tML0K-504-25@gated-at.bofh.it>
In reply to#1652953
On Tue, May 30, 2017 at 10:40 AM, Vladimir Murzin
<vladimir.murzin@arm.com> wrote:
> On 30/05/17 09:31, Vladimir Murzin wrote:
>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
>>
>> On 30/05/17 09:15, Dmitry Vyukov wrote:
>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>>> <vladimir.murzin@arm.com> wrote:
>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>>> I have an alternative proposal. It should be conceptually simpler and
>>>>> also less arch-dependent. But I don't know if I miss something
>>>>> important that will render it non working.
>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>>> shadow blocks to page structs as necessary. It should lead to even
>>>>> smaller memory consumption because we won't need a whole shadow page
>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>>> just a single 512B block). I guess with some fragmentation we need
>>>>> lots of excessive shadow with the current proposed patch.
>>>>> This does not depend on TLB in any way and does not require hooking
>>>>> into buddy allocator.
>>>>> The main downside is that we will need to be careful to not assume
>>>>> that shadow is continuous. In particular this means that this mode
>>>>> will work only with outline instrumentation and will need some ifdefs.
>>>>> Also it will be slower due to the additional indirection when
>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>>> I understand.
>>>>>
>>>>> But the main win as I see it is that that's basically complete support
>>>>> for 32-bit arches. People do ask about arm32 support:
>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>>> and probably mips32 is relevant as well.
>>>>> Such mode does not require a huge continuous address space range, has
>>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>>> Works only with outline instrumentation, but I think that's a
>>>>> reasonable compromise.
>>>>
>>>> .. or you can just keep shadow in page extension. It was suggested back in
>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>>
>>>> [1] https://lkml.org/lkml/2015/8/24/573
>>>
>>> Right. It describes basically the same idea.
>>>
>>> How is page_ext better than adding data page struct?
>>
>> page_ext is already here along with some other debug options ;)


But page struct is also here. What am I missing?


>>> It seems that memory for all page_ext is preallocated along with page
>>> structs; but just the lookup is slower.
>>>
>>
>> Yup. Lookup would look like (based on v4.0):
>>
>> ...
>> page_ext = lookup_page_ext_begin(virt_to_page(start));
>>
>> do {
>>         page_ext->shadow[idx++] = value;
>> } while (idx < bound);
>>
>> lookup_page_ext_end((void *)page_ext);
>>
>> ...
>
> Correction: please, ignore that *_{begin,end} stuff - mainline only
> lookup_page_ext() is only used.


Note that this added code will be executed during handling of each and
every memory access in kernel. Every instruction matters on that path.
The additional indirection via page struct will also slow down it, but
that's the cost for lower memory consumption and potentially 32-bit
support. For page_ext it looks like even more overhead for no gain.

[toc] | [prev] | [next] | [standalone]


#1652984 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 11:10 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tMLk6-5mP-17@gated-at.bofh.it>
In reply to#1652958
On 30/05/17 09:49, Dmitry Vyukov wrote:
> On Tue, May 30, 2017 at 10:40 AM, Vladimir Murzin
> <vladimir.murzin@arm.com> wrote:
>> On 30/05/17 09:31, Vladimir Murzin wrote:
>>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
>>>
>>> On 30/05/17 09:15, Dmitry Vyukov wrote:
>>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>>>> <vladimir.murzin@arm.com> wrote:
>>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>>>> I have an alternative proposal. It should be conceptually simpler and
>>>>>> also less arch-dependent. But I don't know if I miss something
>>>>>> important that will render it non working.
>>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>>>> shadow blocks to page structs as necessary. It should lead to even
>>>>>> smaller memory consumption because we won't need a whole shadow page
>>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>>>> just a single 512B block). I guess with some fragmentation we need
>>>>>> lots of excessive shadow with the current proposed patch.
>>>>>> This does not depend on TLB in any way and does not require hooking
>>>>>> into buddy allocator.
>>>>>> The main downside is that we will need to be careful to not assume
>>>>>> that shadow is continuous. In particular this means that this mode
>>>>>> will work only with outline instrumentation and will need some ifdefs.
>>>>>> Also it will be slower due to the additional indirection when
>>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>>>> I understand.
>>>>>>
>>>>>> But the main win as I see it is that that's basically complete support
>>>>>> for 32-bit arches. People do ask about arm32 support:
>>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>>>> and probably mips32 is relevant as well.
>>>>>> Such mode does not require a huge continuous address space range, has
>>>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>>>> Works only with outline instrumentation, but I think that's a
>>>>>> reasonable compromise.
>>>>>
>>>>> .. or you can just keep shadow in page extension. It was suggested back in
>>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>>>
>>>>> [1] https://lkml.org/lkml/2015/8/24/573
>>>>
>>>> Right. It describes basically the same idea.
>>>>
>>>> How is page_ext better than adding data page struct?
>>>
>>> page_ext is already here along with some other debug options ;)
> 
> 
> But page struct is also here. What am I missing?
> 

Probably, free room in page struct? I guess most of the page_ext stuff would
love to live in page struct, but... for instance, look at page idle tracking
which has to live in page_ext only for 32-bit.

> 
>>>> It seems that memory for all page_ext is preallocated along with page
>>>> structs; but just the lookup is slower.
>>>>
>>>
>>> Yup. Lookup would look like (based on v4.0):
>>>
>>> ...
>>> page_ext = lookup_page_ext_begin(virt_to_page(start));
>>>
>>> do {
>>>         page_ext->shadow[idx++] = value;
>>> } while (idx < bound);
>>>
>>> lookup_page_ext_end((void *)page_ext);
>>>
>>> ...
>>
>> Correction: please, ignore that *_{begin,end} stuff - mainline only
>> lookup_page_ext() is only used.
> 
> 
> Note that this added code will be executed during handling of each and
> every memory access in kernel. Every instruction matters on that path.

I know, I know... still better than nothing.

> The additional indirection via page struct will also slow down it, but
> that's the cost for lower memory consumption and potentially 32-bit
> support. For page_ext it looks like even more overhead for no gain.
> 

eefa864 (mm/page_ext: resurrect struct page extending code for debugging)
express some cases where keeping data in page_ext has benefit.

Cheers
Vladimir

[toc] | [prev] | [next] | [standalone]


#1653031 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-30 11:30 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tMLDt-5ux-49@gated-at.bofh.it>
In reply to#1652984
On Tue, May 30, 2017 at 11:08 AM, Vladimir Murzin
<vladimir.murzin@arm.com> wrote:
>> <vladimir.murzin@arm.com> wrote:
>>> On 30/05/17 09:31, Vladimir Murzin wrote:
>>>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
>>>>
>>>> On 30/05/17 09:15, Dmitry Vyukov wrote:
>>>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>>>>> <vladimir.murzin@arm.com> wrote:
>>>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>>>>> I have an alternative proposal. It should be conceptually simpler and
>>>>>>> also less arch-dependent. But I don't know if I miss something
>>>>>>> important that will render it non working.
>>>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>>>>> shadow blocks to page structs as necessary. It should lead to even
>>>>>>> smaller memory consumption because we won't need a whole shadow page
>>>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>>>>> just a single 512B block). I guess with some fragmentation we need
>>>>>>> lots of excessive shadow with the current proposed patch.
>>>>>>> This does not depend on TLB in any way and does not require hooking
>>>>>>> into buddy allocator.
>>>>>>> The main downside is that we will need to be careful to not assume
>>>>>>> that shadow is continuous. In particular this means that this mode
>>>>>>> will work only with outline instrumentation and will need some ifdefs.
>>>>>>> Also it will be slower due to the additional indirection when
>>>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>>>>> I understand.
>>>>>>>
>>>>>>> But the main win as I see it is that that's basically complete support
>>>>>>> for 32-bit arches. People do ask about arm32 support:
>>>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>>>>> and probably mips32 is relevant as well.
>>>>>>> Such mode does not require a huge continuous address space range, has
>>>>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>>>>> Works only with outline instrumentation, but I think that's a
>>>>>>> reasonable compromise.
>>>>>>
>>>>>> .. or you can just keep shadow in page extension. It was suggested back in
>>>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>>>>
>>>>>> [1] https://lkml.org/lkml/2015/8/24/573
>>>>>
>>>>> Right. It describes basically the same idea.
>>>>>
>>>>> How is page_ext better than adding data page struct?
>>>>
>>>> page_ext is already here along with some other debug options ;)
>>
>>
>> But page struct is also here. What am I missing?
>>
>
> Probably, free room in page struct? I guess most of the page_ext stuff would
> love to live in page struct, but... for instance, look at page idle tracking
> which has to live in page_ext only for 32-bit.


Sorry for my ignorance. What's the fundamental problem with just
pushing everything into page struct?

I don't see anything relevant in page struct comment. Nor I see "idle"
nor "tracking" page struct. I see only 2 mentions of CONFIG_64BIT, but
both declare the same fields just with different types (int vs short).



>>>>> It seems that memory for all page_ext is preallocated along with page
>>>>> structs; but just the lookup is slower.
>>>>>
>>>>
>>>> Yup. Lookup would look like (based on v4.0):
>>>>
>>>> ...
>>>> page_ext = lookup_page_ext_begin(virt_to_page(start));
>>>>
>>>> do {
>>>>         page_ext->shadow[idx++] = value;
>>>> } while (idx < bound);
>>>>
>>>> lookup_page_ext_end((void *)page_ext);
>>>>
>>>> ...
>>>
>>> Correction: please, ignore that *_{begin,end} stuff - mainline only
>>> lookup_page_ext() is only used.
>>
>>
>> Note that this added code will be executed during handling of each and
>> every memory access in kernel. Every instruction matters on that path.
>
> I know, I know... still better than nothing.
>
>> The additional indirection via page struct will also slow down it, but
>> that's the cost for lower memory consumption and potentially 32-bit
>> support. For page_ext it looks like even more overhead for no gain.
>>
>
> eefa864 (mm/page_ext: resurrect struct page extending code for debugging)
> express some cases where keeping data in page_ext has benefit.
>
> Cheers
> Vladimir

[toc] | [prev] | [next] | [standalone]


#1653041 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 11:40 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tMLN9-5ya-31@gated-at.bofh.it>
In reply to#1653031
On 30/05/17 10:26, Dmitry Vyukov wrote:
> On Tue, May 30, 2017 at 11:08 AM, Vladimir Murzin
> <vladimir.murzin@arm.com> wrote:
>>> <vladimir.murzin@arm.com> wrote:
>>>> On 30/05/17 09:31, Vladimir Murzin wrote:
>>>>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
>>>>>
>>>>> On 30/05/17 09:15, Dmitry Vyukov wrote:
>>>>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>>>>>> <vladimir.murzin@arm.com> wrote:
>>>>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>>>>>> I have an alternative proposal. It should be conceptually simpler and
>>>>>>>> also less arch-dependent. But I don't know if I miss something
>>>>>>>> important that will render it non working.
>>>>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>>>>>> shadow blocks to page structs as necessary. It should lead to even
>>>>>>>> smaller memory consumption because we won't need a whole shadow page
>>>>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>>>>>> just a single 512B block). I guess with some fragmentation we need
>>>>>>>> lots of excessive shadow with the current proposed patch.
>>>>>>>> This does not depend on TLB in any way and does not require hooking
>>>>>>>> into buddy allocator.
>>>>>>>> The main downside is that we will need to be careful to not assume
>>>>>>>> that shadow is continuous. In particular this means that this mode
>>>>>>>> will work only with outline instrumentation and will need some ifdefs.
>>>>>>>> Also it will be slower due to the additional indirection when
>>>>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>>>>>> I understand.
>>>>>>>>
>>>>>>>> But the main win as I see it is that that's basically complete support
>>>>>>>> for 32-bit arches. People do ask about arm32 support:
>>>>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>>>>>> and probably mips32 is relevant as well.
>>>>>>>> Such mode does not require a huge continuous address space range, has
>>>>>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>>>>>> Works only with outline instrumentation, but I think that's a
>>>>>>>> reasonable compromise.
>>>>>>>
>>>>>>> .. or you can just keep shadow in page extension. It was suggested back in
>>>>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>>>>>
>>>>>>> [1] https://lkml.org/lkml/2015/8/24/573
>>>>>>
>>>>>> Right. It describes basically the same idea.
>>>>>>
>>>>>> How is page_ext better than adding data page struct?
>>>>>
>>>>> page_ext is already here along with some other debug options ;)
>>>
>>>
>>> But page struct is also here. What am I missing?
>>>
>>
>> Probably, free room in page struct? I guess most of the page_ext stuff would
>> love to live in page struct, but... for instance, look at page idle tracking
>> which has to live in page_ext only for 32-bit.
> 
> 
> Sorry for my ignorance. What's the fundamental problem with just
> pushing everything into page struct?

I think [1] has an answer for your question ;)

> 
> I don't see anything relevant in page struct comment. Nor I see "idle"
> nor "tracking" page struct. I see only 2 mentions of CONFIG_64BIT, but
> both declare the same fields just with different types (int vs short).

Right, it is because implementation is based on page flags [1]:

Note, since there is no room for extra page flags on 32 bit, this feature
uses extended page flags when compiled on 32 bit.


[1] https://lwn.net/Articles/565097/
[2] 33c3fc7 ("mm: introduce idle page tracking")

Cheers
Vladimir

> 
> 
> 
>>>>>> It seems that memory for all page_ext is preallocated along with page
>>>>>> structs; but just the lookup is slower.
>>>>>>
>>>>>
>>>>> Yup. Lookup would look like (based on v4.0):
>>>>>
>>>>> ...
>>>>> page_ext = lookup_page_ext_begin(virt_to_page(start));
>>>>>
>>>>> do {
>>>>>         page_ext->shadow[idx++] = value;
>>>>> } while (idx < bound);
>>>>>
>>>>> lookup_page_ext_end((void *)page_ext);
>>>>>
>>>>> ...
>>>>
>>>> Correction: please, ignore that *_{begin,end} stuff - mainline only
>>>> lookup_page_ext() is only used.
>>>
>>>
>>> Note that this added code will be executed during handling of each and
>>> every memory access in kernel. Every instruction matters on that path.
>>
>> I know, I know... still better than nothing.
>>
>>> The additional indirection via page struct will also slow down it, but
>>> that's the cost for lower memory consumption and potentially 32-bit
>>> support. For page_ext it looks like even more overhead for no gain.
>>>
>>
>> eefa864 (mm/page_ext: resurrect struct page extending code for debugging)
>> express some cases where keeping data in page_ext has benefit.
>>
>> Cheers
>> Vladimir
> 

[toc] | [prev] | [next] | [standalone]


#1653050 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-30 11:50 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tMLWO-5BU-31@gated-at.bofh.it>
In reply to#1653041
On Tue, May 30, 2017 at 11:39 AM, Vladimir Murzin
<vladimir.murzin@arm.com> wrote:
>
> On 30/05/17 10:26, Dmitry Vyukov wrote:
> > On Tue, May 30, 2017 at 11:08 AM, Vladimir Murzin
> > <vladimir.murzin@arm.com> wrote:
> >>> <vladimir.murzin@arm.com> wrote:
> >>>> On 30/05/17 09:31, Vladimir Murzin wrote:
> >>>>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
> >>>>>
> >>>>> On 30/05/17 09:15, Dmitry Vyukov wrote:
> >>>>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
> >>>>>> <vladimir.murzin@arm.com> wrote:
> >>>>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
> >>>>>>>> I have an alternative proposal. It should be conceptually simpler and
> >>>>>>>> also less arch-dependent. But I don't know if I miss something
> >>>>>>>> important that will render it non working.
> >>>>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
> >>>>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
> >>>>>>>> shadow blocks to page structs as necessary. It should lead to even
> >>>>>>>> smaller memory consumption because we won't need a whole shadow page
> >>>>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
> >>>>>>>> just a single 512B block). I guess with some fragmentation we need
> >>>>>>>> lots of excessive shadow with the current proposed patch.
> >>>>>>>> This does not depend on TLB in any way and does not require hooking
> >>>>>>>> into buddy allocator.
> >>>>>>>> The main downside is that we will need to be careful to not assume
> >>>>>>>> that shadow is continuous. In particular this means that this mode
> >>>>>>>> will work only with outline instrumentation and will need some ifdefs.
> >>>>>>>> Also it will be slower due to the additional indirection when
> >>>>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
> >>>>>>>> I understand.
> >>>>>>>>
> >>>>>>>> But the main win as I see it is that that's basically complete support
> >>>>>>>> for 32-bit arches. People do ask about arm32 support:
> >>>>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
> >>>>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
> >>>>>>>> and probably mips32 is relevant as well.
> >>>>>>>> Such mode does not require a huge continuous address space range, has
> >>>>>>>> minimal memory consumption and requires minimal arch-dependent code.
> >>>>>>>> Works only with outline instrumentation, but I think that's a
> >>>>>>>> reasonable compromise.
> >>>>>>>
> >>>>>>> .. or you can just keep shadow in page extension. It was suggested back in
> >>>>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
> >>>>>>>
> >>>>>>> [1] https://lkml.org/lkml/2015/8/24/573
> >>>>>>
> >>>>>> Right. It describes basically the same idea.
> >>>>>>
> >>>>>> How is page_ext better than adding data page struct?
> >>>>>
> >>>>> page_ext is already here along with some other debug options ;)
> >>>
> >>>
> >>> But page struct is also here. What am I missing?
> >>>
> >>
> >> Probably, free room in page struct? I guess most of the page_ext stuff would
> >> love to live in page struct, but... for instance, look at page idle tracking
> >> which has to live in page_ext only for 32-bit.
> >
> >
> > Sorry for my ignorance. What's the fundamental problem with just
> > pushing everything into page struct?
>
> I think [1] has an answer for your question ;)

It also has an answer for why we should put it into page struct :)


>
> >
> > I don't see anything relevant in page struct comment. Nor I see "idle"
> > nor "tracking" page struct. I see only 2 mentions of CONFIG_64BIT, but
> > both declare the same fields just with different types (int vs short).
>
> Right, it is because implementation is based on page flags [1]:
>
> Note, since there is no room for extra page flags on 32 bit, this feature
> uses extended page flags when compiled on 32 bit.
>
>
> [1] https://lwn.net/Articles/565097/
> [2] 33c3fc7 ("mm: introduce idle page tracking")

[toc] | [prev] | [next] | [standalone]


#1653062 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 12:00 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tMM6v-5FJ-33@gated-at.bofh.it>
In reply to#1653050
On 30/05/17 10:45, Dmitry Vyukov wrote:
> On Tue, May 30, 2017 at 11:39 AM, Vladimir Murzin
> <vladimir.murzin@arm.com> wrote:
>>
>> On 30/05/17 10:26, Dmitry Vyukov wrote:
>>> On Tue, May 30, 2017 at 11:08 AM, Vladimir Murzin
>>> <vladimir.murzin@arm.com> wrote:
>>>>> <vladimir.murzin@arm.com> wrote:
>>>>>> On 30/05/17 09:31, Vladimir Murzin wrote:
>>>>>>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
>>>>>>>
>>>>>>> On 30/05/17 09:15, Dmitry Vyukov wrote:
>>>>>>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>>>>>>>> <vladimir.murzin@arm.com> wrote:
>>>>>>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>>>>>>>> I have an alternative proposal. It should be conceptually simpler and
>>>>>>>>>> also less arch-dependent. But I don't know if I miss something
>>>>>>>>>> important that will render it non working.
>>>>>>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>>>>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>>>>>>>> shadow blocks to page structs as necessary. It should lead to even
>>>>>>>>>> smaller memory consumption because we won't need a whole shadow page
>>>>>>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>>>>>>>> just a single 512B block). I guess with some fragmentation we need
>>>>>>>>>> lots of excessive shadow with the current proposed patch.
>>>>>>>>>> This does not depend on TLB in any way and does not require hooking
>>>>>>>>>> into buddy allocator.
>>>>>>>>>> The main downside is that we will need to be careful to not assume
>>>>>>>>>> that shadow is continuous. In particular this means that this mode
>>>>>>>>>> will work only with outline instrumentation and will need some ifdefs.
>>>>>>>>>> Also it will be slower due to the additional indirection when
>>>>>>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>>>>>>>> I understand.
>>>>>>>>>>
>>>>>>>>>> But the main win as I see it is that that's basically complete support
>>>>>>>>>> for 32-bit arches. People do ask about arm32 support:
>>>>>>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>>>>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>>>>>>>> and probably mips32 is relevant as well.
>>>>>>>>>> Such mode does not require a huge continuous address space range, has
>>>>>>>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>>>>>>>> Works only with outline instrumentation, but I think that's a
>>>>>>>>>> reasonable compromise.
>>>>>>>>>
>>>>>>>>> .. or you can just keep shadow in page extension. It was suggested back in
>>>>>>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>>>>>>>
>>>>>>>>> [1] https://lkml.org/lkml/2015/8/24/573
>>>>>>>>
>>>>>>>> Right. It describes basically the same idea.
>>>>>>>>
>>>>>>>> How is page_ext better than adding data page struct?
>>>>>>>
>>>>>>> page_ext is already here along with some other debug options ;)
>>>>>
>>>>>
>>>>> But page struct is also here. What am I missing?
>>>>>
>>>>
>>>> Probably, free room in page struct? I guess most of the page_ext stuff would
>>>> love to live in page struct, but... for instance, look at page idle tracking
>>>> which has to live in page_ext only for 32-bit.
>>>
>>>
>>> Sorry for my ignorance. What's the fundamental problem with just
>>> pushing everything into page struct?
>>
>> I think [1] has an answer for your question ;)
> 
> It also has an answer for why we should put it into page struct :)

Glad you find it useful ;) I'd be glad to see it lands into 32-bit world :)

Cheers
Vladimir

> 
> 
>>
>>>
>>> I don't see anything relevant in page struct comment. Nor I see "idle"
>>> nor "tracking" page struct. I see only 2 mentions of CONFIG_64BIT, but
>>> both declare the same fields just with different types (int vs short).
>>
>> Right, it is because implementation is based on page flags [1]:
>>
>> Note, since there is no room for extra page flags on 32 bit, this feature
>> uses extended page flags when compiled on 32 bit.
>>
>>
>> [1] https://lwn.net/Articles/565097/
>> [2] 33c3fc7 ("mm: introduce idle page tracking")
> 

[toc] | [prev] | [next] | [standalone]


#1653850 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromJoonsoo Kim <js1304@gmail.com>
Date2017-05-31 08:00 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tN4PL-FB-1@gated-at.bofh.it>
In reply to#1652599
On Tue, May 30, 2017 at 05:16:56PM +0300, Andrey Ryabinin wrote:
> On 05/29/2017 06:29 PM, Dmitry Vyukov wrote:
> > Joonsoo,
> > 
> > I guess mine (and Andrey's) main concern is the amount of additional
> > complexity (I am still struggling to understand how it all works) and
> > more arch-dependent code in exchange for moderate memory win.
> > 
> > Joonsoo, Andrey,
> > 
> > I have an alternative proposal. It should be conceptually simpler and
> > also less arch-dependent. But I don't know if I miss something
> > important that will render it non working.
> > Namely, we add a pointer to shadow to the page struct. Then, create a
> > slab allocator for 512B shadow blocks. Then, attach/detach these
> > shadow blocks to page structs as necessary. It should lead to even
> > smaller memory consumption because we won't need a whole shadow page
> > when only 1 out of 8 corresponding kernel pages are used (we will need
> > just a single 512B block). I guess with some fragmentation we need
> > lots of excessive shadow with the current proposed patch.
> > This does not depend on TLB in any way and does not require hooking
> > into buddy allocator.
> > The main downside is that we will need to be careful to not assume
> > that shadow is continuous. In particular this means that this mode
> > will work only with outline instrumentation and will need some ifdefs.
> > Also it will be slower due to the additional indirection when
> > accessing shadow, but that's meant as "small but slow" mode as far as
> > I understand.
> 
> It seems that you are forgetting about stack instrumentation.
> You'll have to disable it completely, at least with current implementation of it in gcc.

Correct. Even if we use OUTLINE build, gcc directly inserts codes to the
function prologue/epilogue to mark/unmakr the shadow. And, I'm not
sure we can change it since it would affect performance greately. In
current situation, alternative proposal loses most of benefit mentioned
above.
> 
> > But the main win as I see it is that that's basically complete support
> > for 32-bit arches. People do ask about arm32 support:
> > https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
> > https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
> > and probably mips32 is relevant as well.
> 
> I don't see how above is relevant for 32-bit arches. Current design
> is perfectly fine for 32-bit arches. I did some POC arm32 port couple years
> ago - https://github.com/aryabinin/linux/commits/kasan/arm_v0_1
> It has some ugly hacks and non-critical bugs. AFAIR it also super-slow because I (mistakenly) 
> made shadow memory uncached. But otherwise it works.

Could you explain that where is the code to map shadow memory uncached?
I don't find anything related to it.

> > Such mode does not require a huge continuous address space range, has
> > minimal memory consumption and requires minimal arch-dependent code.
> > Works only with outline instrumentation, but I think that's a
> > reasonable compromise.
> > 
> > What do you think?
>  
> I don't understand why we trying to invent some hacky/complex schemes when we already have
> a simple one - scaling shadow to 1/32. It's easy to implement and should be more performant comparing
> to suggested schemes.

My approach can co-exist with changing scaling approach. It has it's
own benefit.

And, as Dmitry mentioned before, scaling shadow to 1/32 also has downsides,
expecially for inline instrumentation. And, it requires compiler
modification and user needs to update their compiler to newer version
which is not so simple in terms of the user's usability

Thanks.

[toc] | [prev] | [next] | [standalone]


#1655653 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromDmitry Vyukov <dvyukov@google.com>
Date2017-06-01 20:10 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tNCHM-61B-25@gated-at.bofh.it>
In reply to#1652599
On Tue, May 30, 2017 at 4:16 PM, Andrey Ryabinin
<aryabinin@virtuozzo.com> wrote:
> On 05/29/2017 06:29 PM, Dmitry Vyukov wrote:
>> Joonsoo,
>>
>> I guess mine (and Andrey's) main concern is the amount of additional
>> complexity (I am still struggling to understand how it all works) and
>> more arch-dependent code in exchange for moderate memory win.
>>
>> Joonsoo, Andrey,
>>
>> I have an alternative proposal. It should be conceptually simpler and
>> also less arch-dependent. But I don't know if I miss something
>> important that will render it non working.
>> Namely, we add a pointer to shadow to the page struct. Then, create a
>> slab allocator for 512B shadow blocks. Then, attach/detach these
>> shadow blocks to page structs as necessary. It should lead to even
>> smaller memory consumption because we won't need a whole shadow page
>> when only 1 out of 8 corresponding kernel pages are used (we will need
>> just a single 512B block). I guess with some fragmentation we need
>> lots of excessive shadow with the current proposed patch.
>> This does not depend on TLB in any way and does not require hooking
>> into buddy allocator.
>> The main downside is that we will need to be careful to not assume
>> that shadow is continuous. In particular this means that this mode
>> will work only with outline instrumentation and will need some ifdefs.
>> Also it will be slower due to the additional indirection when
>> accessing shadow, but that's meant as "small but slow" mode as far as
>> I understand.
>
> It seems that you are forgetting about stack instrumentation.
> You'll have to disable it completely, at least with current implementation of it in gcc.
>
>> But the main win as I see it is that that's basically complete support
>> for 32-bit arches. People do ask about arm32 support:
>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>> and probably mips32 is relevant as well.
>
> I don't see how above is relevant for 32-bit arches. Current design
> is perfectly fine for 32-bit arches. I did some POC arm32 port couple years
> ago - https://github.com/aryabinin/linux/commits/kasan/arm_v0_1
> It has some ugly hacks and non-critical bugs. AFAIR it also super-slow because I (mistakenly)
> made shadow memory uncached. But otherwise it works.
>
>> Such mode does not require a huge continuous address space range, has
>> minimal memory consumption and requires minimal arch-dependent code.
>> Works only with outline instrumentation, but I think that's a
>> reasonable compromise.
>>
>> What do you think?
>
> I don't understand why we trying to invent some hacky/complex schemes when we already have
> a simple one - scaling shadow to 1/32. It's easy to implement and should be more performant comparing
> to suggested schemes.


If 32-bits work with the current approach, then I would also prefer to
keep things simpler.
FWIW clang supports settings shadow scale via a command line flag
(-asan-mapping-scale).

[toc] | [prev] | [next] | [standalone]


#1645116 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromJoonsoo Kim <js1304@gmail.com>
Date2017-05-19 04:00 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tIFmW-3k7-15@gated-at.bofh.it>
In reply to#1642158
On Wed, May 17, 2017 at 03:17:13PM +0300, Andrey Ryabinin wrote:
> On 05/16/2017 04:16 AM, js1304@gmail.com wrote:
> > From: Joonsoo Kim <iamjoonsoo.kim@lge.com>
> > 
> > Hello, all.
> > 
> > This is an attempt to recude memory consumption of KASAN. Please see
> > following description to get the more information.
> > 
> > 1. What is per-page shadow memory
> > 
> > This patch introduces infrastructure to support per-page shadow memory.
> > Per-page shadow memory is the same with original shadow memory except
> > the granualarity. It's one byte shows the shadow value for the page.
> > The purpose of introducing this new shadow memory is to save memory
> > consumption.
> > 
> > 2. Problem of current approach
> > 
> > Until now, KASAN needs shadow memory for all the range of the memory
> > so the amount of statically allocated memory is so large. It causes
> > the problem that KASAN cannot run on the system with hard memory
> > constraint. Even if KASAN can run, large memory consumption due to
> > KASAN changes behaviour of the workload so we cannot validate
> > the moment that we want to check.
> > 
> > 3. How does this patch fix the problem
> > 
> > This patch tries to fix the problem by reducing memory consumption for
> > the shadow memory. There are two observations.
> > 
> 
> 
> I think that the best way to deal with your problem is to increase shadow scale size.
> 
> You'll need to add tunable to gcc to control shadow size. I expect that gcc has some
> places where 8-shadow scale size is hardcoded, but it should be fixable.
> 
> The kernel also have some small amount of code written with KASAN_SHADOW_SCALE_SIZE == 8 in mind,
> which should be easy to fix.
> 
> Note that bigger shadow scale size requires bigger alignment of allocated memory and variables.
> However, according to comments in gcc/asan.c gcc already aligns stack and global variables and at
> 32-bytes boundary.
> So we could bump shadow scale up to 32 without increasing current stack consumption.
> 
> On a small machine (1Gb) 1/32 of shadow is just 32Mb which is comparable to yours 30Mb, but I expect it to be
> much faster. More importantly, this will require only small amount of simple changes in code, which will be
> a *lot* more easier to maintain.

I agree that it is also a good option to reduce memory consumption.
Nevertheless, there are two reasons that justifies this patchset.

1) With this patchset, memory consumption isn't increased in
proportional to total memory size. Please consider my 4Gb system
example on the below. With increasing shadow scale size to 32, memory
would be consumed by 128M. However, this patchset consumed 50MB. This
difference can be larger if we run KASAN with bigger machine.

2) These two optimization can be applied simulatenously. It is just an
orthogonal feature. If shadow scale size is increased to 32, memory
consumption will be decreased in case of my patchset, too.

Therefore, I think that this patchset is useful in any case.

Note that increasing shadow scale has it's own trade-off. It requires
that the size of slab object is aligned to shadow scale. It will
increase memory consumption due to slab.

Thanks.

[toc] | [prev] | [next] | [standalone]


#1646538 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-22 08:10 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tJOHw-1Mg-15@gated-at.bofh.it>
In reply to#1645116
On Fri, May 19, 2017 at 3:53 AM, Joonsoo Kim <js1304@gmail.com> wrote:
> On Wed, May 17, 2017 at 03:17:13PM +0300, Andrey Ryabinin wrote:
>> On 05/16/2017 04:16 AM, js1304@gmail.com wrote:
>> > From: Joonsoo Kim <iamjoonsoo.kim@lge.com>
>> >
>> > Hello, all.
>> >
>> > This is an attempt to recude memory consumption of KASAN. Please see
>> > following description to get the more information.
>> >
>> > 1. What is per-page shadow memory
>> >
>> > This patch introduces infrastructure to support per-page shadow memory.
>> > Per-page shadow memory is the same with original shadow memory except
>> > the granualarity. It's one byte shows the shadow value for the page.
>> > The purpose of introducing this new shadow memory is to save memory
>> > consumption.
>> >
>> > 2. Problem of current approach
>> >
>> > Until now, KASAN needs shadow memory for all the range of the memory
>> > so the amount of statically allocated memory is so large. It causes
>> > the problem that KASAN cannot run on the system with hard memory
>> > constraint. Even if KASAN can run, large memory consumption due to
>> > KASAN changes behaviour of the workload so we cannot validate
>> > the moment that we want to check.
>> >
>> > 3. How does this patch fix the problem
>> >
>> > This patch tries to fix the problem by reducing memory consumption for
>> > the shadow memory. There are two observations.
>> >
>>
>>
>> I think that the best way to deal with your problem is to increase shadow scale size.
>>
>> You'll need to add tunable to gcc to control shadow size. I expect that gcc has some
>> places where 8-shadow scale size is hardcoded, but it should be fixable.
>>
>> The kernel also have some small amount of code written with KASAN_SHADOW_SCALE_SIZE == 8 in mind,
>> which should be easy to fix.
>>
>> Note that bigger shadow scale size requires bigger alignment of allocated memory and variables.
>> However, according to comments in gcc/asan.c gcc already aligns stack and global variables and at
>> 32-bytes boundary.
>> So we could bump shadow scale up to 32 without increasing current stack consumption.
>>
>> On a small machine (1Gb) 1/32 of shadow is just 32Mb which is comparable to yours 30Mb, but I expect it to be
>> much faster. More importantly, this will require only small amount of simple changes in code, which will be
>> a *lot* more easier to maintain.


Interesting option. We never considered increasing scale in user space
due to performance implications. But the algorithm always supported up
to 128x scale. Definitely worth considering as an option.


> I agree that it is also a good option to reduce memory consumption.
> Nevertheless, there are two reasons that justifies this patchset.
>
> 1) With this patchset, memory consumption isn't increased in
> proportional to total memory size. Please consider my 4Gb system
> example on the below. With increasing shadow scale size to 32, memory
> would be consumed by 128M. However, this patchset consumed 50MB. This
> difference can be larger if we run KASAN with bigger machine.
>
> 2) These two optimization can be applied simulatenously. It is just an
> orthogonal feature. If shadow scale size is increased to 32, memory
> consumption will be decreased in case of my patchset, too.
>
> Therefore, I think that this patchset is useful in any case.

It is definitely useful all else being equal. But it does considerably
increase code size and complexity, which is an important aspect.

Also note that there is also fixed size quarantine (1/32 of RAM) and
redzones. Reducing shadow overhead beyond some threshold has
diminishing returns, because overall overhead will be just dominated
by quarantine/redzones.

What's your target devices and constraints? We run KASAN on phones
today without any issues.


> Note that increasing shadow scale has it's own trade-off. It requires
> that the size of slab object is aligned to shadow scale. It will
> increase memory consumption due to slab.

I've tried to retest your latest change on top of
http://git.cmpxchg.org/cgit.cgi/linux-mmots.git
d9cd9c95cc3b2fed0f04d233ebf2f7056741858c, but now this version
https://codereview.appspot.com/325780043 always crashes during boot
for me. Report points to zero shadow.

[    0.123434] ==================================================================
[    0.125153] BUG: KASAN: double-free or invalid-free in
cleanup_uevent_env+0x2c/0x40
[    0.126900]
[    0.127318] CPU: 1 PID: 226 Comm: kworker/u8:0 Not tainted
4.12.0-rc1-mm1+ #376
[    0.128995] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996),
BIOS Bochs 01/01/2011
[    0.130896] Call Trace:
[    0.131202] kworker/u8:0 (277) used greatest stack depth: 22976 bytes left
[    0.133129]  dump_stack+0xb0/0x13d
[    0.133958]  ? _atomic_dec_and_lock+0x1e3/0x1e3
[    0.135020]  ? load_image_and_restore+0xf6/0xf6
[    0.136083]  ? kmemdup+0x31/0x40
[    0.136143] kworker/u8:0 (320) used greatest stack depth: 22112 bytes left
[    0.138294]  ? cleanup_uevent_env+0x2c/0x40
[    0.139255]  print_address_description+0x6a/0x270
[    0.140285]  ? cleanup_uevent_env+0x2c/0x40
[    0.141224]  ? cleanup_uevent_env+0x2c/0x40
[    0.142168]  kasan_report_double_free+0x55/0x80
[    0.143162]  kasan_slab_free+0xa4/0xc0
[    0.143934]  ? cleanup_uevent_env+0x2c/0x40
[    0.144882]  kfree+0x8f/0x190
[    0.145561]  cleanup_uevent_env+0x2c/0x40
[    0.146455]  umh_complete+0x3c/0x60
[    0.147180]  call_usermodehelper_exec_async+0x671/0x950
[    0.148334]  ? __asan_report_store_n_noabort+0x12/0x20
[    0.149460]  ? native_load_sp0+0xa3/0xb0
[    0.150213]  ? umh_complete+0x60/0x60
[    0.150990]  ? kasan_end_report+0x20/0x50
[    0.151829]  ? finish_task_switch+0x510/0x7d0
[    0.152760]  ? copy_user_overflow+0x20/0x20
[    0.153565]  ? umh_complete+0x60/0x60
[    0.154341]  ? umh_complete+0x60/0x60
[    0.155125]  ret_from_fork+0x2c/0x40
[    0.155888]
[    0.156190] Allocated by task 1:
[    0.156890]  save_stack_trace+0x16/0x20
[    0.157629]  save_stack+0x43/0xd0
[    0.158299]  kasan_kmalloc+0xad/0xe0
[    0.159068]  kmem_cache_alloc_trace+0x61/0x170
[    0.159920]  kobject_uevent_env+0x1b2/0xa20
[    0.160819]  kobject_uevent+0xb/0x10
[    0.161551]  param_sysfs_init+0x28e/0x2d2
[    0.162375]  do_one_initcall+0x8c/0x290
[    0.163083]  kernel_init_freeable+0x4a2/0x554
[    0.163958]  kernel_init+0xe/0x120
[    0.164669]  ret_from_fork+0x2c/0x40
[    0.165393]
[    0.165685] Freed by task 0:
[    0.166232] (stack is not available)
[    0.166954]
[    0.167247] The buggy address belongs to the object at ffff88007b45e818
[    0.167247]  which belongs to the cache kmalloc-4096 of size 4096
[    0.169709] The buggy address is located 0 bytes inside of
[    0.169709]  4096-byte region [ffff88007b45e818, ffff88007b45f818)
[    0.171897] The buggy address belongs to the page:
[    0.172833] page:ffffea0001ed1600 count:1 mapcount:0 mapping:
   (null) index:0x0 compound_mapcount: 0
[    0.174560] flags: 0x100000000008100(slab|head)
[    0.175410] raw: 0100000000008100 0000000000000000 0000000000000000
0000000100070007
[    0.176819] raw: ffffea0001ed0c20 ffffea0001ed3c20 ffff88007c80ed40
0000000000000000
[    0.178250] page dumped because: kasan: bad access detected
[    0.179312]
[    0.179586] Memory state around the buggy address:
[    0.180488]  ffff88007b45e700: fc fc fc fc fc fc fc fc fc fc fc fc
fc fc fc fc
[    0.181801]  ffff88007b45e780: fc fc fc fc fc fc fc fc fc fc fc fc
fc fc fc fc
[    0.183112] >ffff88007b45e800: fc fc fc 00 00 00 00 00 00 00 00 00
00 00 00 00
[    0.184518]                             ^
[    0.185177]  ffff88007b45e880: 00 00 00 00 00 00 00 00 00 00 00 00
00 00 00 00
[    0.186420]  ffff88007b45e900: 00 00 00 00 00 00 00 00 00 00 00 00
00 00 00 00
[    0.187723] ==================================================================

[toc] | [prev] | [next] | [standalone]


#1649136 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromJoonsoo Kim <js1304@gmail.com>
Date2017-05-24 08:10 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tKxEB-6Td-3@gated-at.bofh.it>
In reply to#1646538
On Mon, May 22, 2017 at 08:02:36AM +0200, Dmitry Vyukov wrote:
> On Fri, May 19, 2017 at 3:53 AM, Joonsoo Kim <js1304@gmail.com> wrote:
> > On Wed, May 17, 2017 at 03:17:13PM +0300, Andrey Ryabinin wrote:
> >> On 05/16/2017 04:16 AM, js1304@gmail.com wrote:
> >> > From: Joonsoo Kim <iamjoonsoo.kim@lge.com>
> >> >
> >> > Hello, all.
> >> >
> >> > This is an attempt to recude memory consumption of KASAN. Please see
> >> > following description to get the more information.
> >> >
> >> > 1. What is per-page shadow memory
> >> >
> >> > This patch introduces infrastructure to support per-page shadow memory.
> >> > Per-page shadow memory is the same with original shadow memory except
> >> > the granualarity. It's one byte shows the shadow value for the page.
> >> > The purpose of introducing this new shadow memory is to save memory
> >> > consumption.
> >> >
> >> > 2. Problem of current approach
> >> >
> >> > Until now, KASAN needs shadow memory for all the range of the memory
> >> > so the amount of statically allocated memory is so large. It causes
> >> > the problem that KASAN cannot run on the system with hard memory
> >> > constraint. Even if KASAN can run, large memory consumption due to
> >> > KASAN changes behaviour of the workload so we cannot validate
> >> > the moment that we want to check.
> >> >
> >> > 3. How does this patch fix the problem
> >> >
> >> > This patch tries to fix the problem by reducing memory consumption for
> >> > the shadow memory. There are two observations.
> >> >
> >>
> >>
> >> I think that the best way to deal with your problem is to increase shadow scale size.
> >>
> >> You'll need to add tunable to gcc to control shadow size. I expect that gcc has some
> >> places where 8-shadow scale size is hardcoded, but it should be fixable.
> >>
> >> The kernel also have some small amount of code written with KASAN_SHADOW_SCALE_SIZE == 8 in mind,
> >> which should be easy to fix.
> >>
> >> Note that bigger shadow scale size requires bigger alignment of allocated memory and variables.
> >> However, according to comments in gcc/asan.c gcc already aligns stack and global variables and at
> >> 32-bytes boundary.
> >> So we could bump shadow scale up to 32 without increasing current stack consumption.
> >>
> >> On a small machine (1Gb) 1/32 of shadow is just 32Mb which is comparable to yours 30Mb, but I expect it to be
> >> much faster. More importantly, this will require only small amount of simple changes in code, which will be
> >> a *lot* more easier to maintain.
> 
> 
> Interesting option. We never considered increasing scale in user space
> due to performance implications. But the algorithm always supported up
> to 128x scale. Definitely worth considering as an option.

Could you explain me how does increasing scale reduce performance? I
tried to guess the reason but failed.

> 
> 
> > I agree that it is also a good option to reduce memory consumption.
> > Nevertheless, there are two reasons that justifies this patchset.
> >
> > 1) With this patchset, memory consumption isn't increased in
> > proportional to total memory size. Please consider my 4Gb system
> > example on the below. With increasing shadow scale size to 32, memory
> > would be consumed by 128M. However, this patchset consumed 50MB. This
> > difference can be larger if we run KASAN with bigger machine.
> >
> > 2) These two optimization can be applied simulatenously. It is just an
> > orthogonal feature. If shadow scale size is increased to 32, memory
> > consumption will be decreased in case of my patchset, too.
> >
> > Therefore, I think that this patchset is useful in any case.
> 
> It is definitely useful all else being equal. But it does considerably
> increase code size and complexity, which is an important aspect.
> 
> Also note that there is also fixed size quarantine (1/32 of RAM) and
> redzones. Reducing shadow overhead beyond some threshold has
> diminishing returns, because overall overhead will be just dominated
> by quarantine/redzones.

My usecase doesn't use quarantine yet since it uses old version kernel
and quarantine isn't back-ported. But, this 1/32 of RAM for quarantine
also could affect the system and I think that we need a switch to
disable it. In our case, making the feature work is more important
than detecting more bugs.

Redzone is also a good target to make selectable since
error pattern could be changed with different object layout. I
sometimes saw that error disappears if KASAN is enabled. I'm not sure
what causes it, but, in some case, it would be helpful that everything
else than something compulsory is the same with non-KASAN build.

> What's your target devices and constraints? We run KASAN on phones
> today without any issues.

My target devices are a smart TV or embedded system on a car. Usually,
these devices have specific use scenario and memory is managed more
tightly than a phone. I have heard that some system with 1GB memory
cannot run if 128MB is used for KASAN. I'm not sure that 1/32 scale
changes the picture, but, yes, I guess that most of problem will disappear.

> 
> > Note that increasing shadow scale has it's own trade-off. It requires
> > that the size of slab object is aligned to shadow scale. It will
> > increase memory consumption due to slab.
> 
> I've tried to retest your latest change on top of
> http://git.cmpxchg.org/cgit.cgi/linux-mmots.git
> d9cd9c95cc3b2fed0f04d233ebf2f7056741858c, but now this version
> https://codereview.appspot.com/325780043 always crashes during boot
> for me. Report points to zero shadow.

Oops... Maybe, it's due to lack of stale TLB handling on double-free
check in kasan_slab_free(). I fixed it on my version 2 patchset.
And, I also fixed performance problem due to memory allocated by early
allocator(memblock or (no)bootmem).

https://github.com/JoonsooKim/linux/tree/kasan-opt-memory-consumption-v2.0-next-20170511

This branch is based on next-20170511.

Thanks.

[toc] | [prev] | [next] | [standalone]


#1649754 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-24 18:40 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tKHui-4G7-7@gated-at.bofh.it>
In reply to#1649136
On Wed, May 24, 2017 at 8:04 AM, Joonsoo Kim <js1304@gmail.com> wrote:
>> >> > From: Joonsoo Kim <iamjoonsoo.kim@lge.com>
>> >> >
>> >> > Hello, all.
>> >> >
>> >> > This is an attempt to recude memory consumption of KASAN. Please see
>> >> > following description to get the more information.
>> >> >
>> >> > 1. What is per-page shadow memory
>> >> >
>> >> > This patch introduces infrastructure to support per-page shadow memory.
>> >> > Per-page shadow memory is the same with original shadow memory except
>> >> > the granualarity. It's one byte shows the shadow value for the page.
>> >> > The purpose of introducing this new shadow memory is to save memory
>> >> > consumption.
>> >> >
>> >> > 2. Problem of current approach
>> >> >
>> >> > Until now, KASAN needs shadow memory for all the range of the memory
>> >> > so the amount of statically allocated memory is so large. It causes
>> >> > the problem that KASAN cannot run on the system with hard memory
>> >> > constraint. Even if KASAN can run, large memory consumption due to
>> >> > KASAN changes behaviour of the workload so we cannot validate
>> >> > the moment that we want to check.
>> >> >
>> >> > 3. How does this patch fix the problem
>> >> >
>> >> > This patch tries to fix the problem by reducing memory consumption for
>> >> > the shadow memory. There are two observations.
>> >> >
>> >>
>> >>
>> >> I think that the best way to deal with your problem is to increase shadow scale size.
>> >>
>> >> You'll need to add tunable to gcc to control shadow size. I expect that gcc has some
>> >> places where 8-shadow scale size is hardcoded, but it should be fixable.
>> >>
>> >> The kernel also have some small amount of code written with KASAN_SHADOW_SCALE_SIZE == 8 in mind,
>> >> which should be easy to fix.
>> >>
>> >> Note that bigger shadow scale size requires bigger alignment of allocated memory and variables.
>> >> However, according to comments in gcc/asan.c gcc already aligns stack and global variables and at
>> >> 32-bytes boundary.
>> >> So we could bump shadow scale up to 32 without increasing current stack consumption.
>> >>
>> >> On a small machine (1Gb) 1/32 of shadow is just 32Mb which is comparable to yours 30Mb, but I expect it to be
>> >> much faster. More importantly, this will require only small amount of simple changes in code, which will be
>> >> a *lot* more easier to maintain.
>>
>>
>> Interesting option. We never considered increasing scale in user space
>> due to performance implications. But the algorithm always supported up
>> to 128x scale. Definitely worth considering as an option.
>
> Could you explain me how does increasing scale reduce performance? I
> tried to guess the reason but failed.


The main reason is inline instrumentation. Inline instrumentation for
a check of 8-byte access (which are very common in 64-bit code) is
just a check of the shadow byte for 0. For smaller accesses we have
more complex instrumentation that first checks shadow for 0 and then
does precise check based on size/offset of the access + shadow value.
That's slower and also increases register pressure and code size
(which can further reduce performance due to icache overflow). If we
increase scale to 16/32, all accesses will need that slow path.
Another thing is stack instrumentation: larger scale will require
larger redzones to ensure proper alignment. That will increase stack
frames and also more instructions to poison/unpoison stack shadow on
function entry/exit.

[toc] | [prev] | [next] | [standalone]


#1650091 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromJoonsoo Kim <js1304@gmail.com>
Date2017-05-25 02:50 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tKP8t-11Y-3@gated-at.bofh.it>
In reply to#1649754
On Wed, May 24, 2017 at 06:31:04PM +0200, Dmitry Vyukov wrote:
> On Wed, May 24, 2017 at 8:04 AM, Joonsoo Kim <js1304@gmail.com> wrote:
> >> >> > From: Joonsoo Kim <iamjoonsoo.kim@lge.com>
> >> >> >
> >> >> > Hello, all.
> >> >> >
> >> >> > This is an attempt to recude memory consumption of KASAN. Please see
> >> >> > following description to get the more information.
> >> >> >
> >> >> > 1. What is per-page shadow memory
> >> >> >
> >> >> > This patch introduces infrastructure to support per-page shadow memory.
> >> >> > Per-page shadow memory is the same with original shadow memory except
> >> >> > the granualarity. It's one byte shows the shadow value for the page.
> >> >> > The purpose of introducing this new shadow memory is to save memory
> >> >> > consumption.
> >> >> >
> >> >> > 2. Problem of current approach
> >> >> >
> >> >> > Until now, KASAN needs shadow memory for all the range of the memory
> >> >> > so the amount of statically allocated memory is so large. It causes
> >> >> > the problem that KASAN cannot run on the system with hard memory
> >> >> > constraint. Even if KASAN can run, large memory consumption due to
> >> >> > KASAN changes behaviour of the workload so we cannot validate
> >> >> > the moment that we want to check.
> >> >> >
> >> >> > 3. How does this patch fix the problem
> >> >> >
> >> >> > This patch tries to fix the problem by reducing memory consumption for
> >> >> > the shadow memory. There are two observations.
> >> >> >
> >> >>
> >> >>
> >> >> I think that the best way to deal with your problem is to increase shadow scale size.
> >> >>
> >> >> You'll need to add tunable to gcc to control shadow size. I expect that gcc has some
> >> >> places where 8-shadow scale size is hardcoded, but it should be fixable.
> >> >>
> >> >> The kernel also have some small amount of code written with KASAN_SHADOW_SCALE_SIZE == 8 in mind,
> >> >> which should be easy to fix.
> >> >>
> >> >> Note that bigger shadow scale size requires bigger alignment of allocated memory and variables.
> >> >> However, according to comments in gcc/asan.c gcc already aligns stack and global variables and at
> >> >> 32-bytes boundary.
> >> >> So we could bump shadow scale up to 32 without increasing current stack consumption.
> >> >>
> >> >> On a small machine (1Gb) 1/32 of shadow is just 32Mb which is comparable to yours 30Mb, but I expect it to be
> >> >> much faster. More importantly, this will require only small amount of simple changes in code, which will be
> >> >> a *lot* more easier to maintain.
> >>
> >>
> >> Interesting option. We never considered increasing scale in user space
> >> due to performance implications. But the algorithm always supported up
> >> to 128x scale. Definitely worth considering as an option.
> >
> > Could you explain me how does increasing scale reduce performance? I
> > tried to guess the reason but failed.
> 
> 
> The main reason is inline instrumentation. Inline instrumentation for
> a check of 8-byte access (which are very common in 64-bit code) is
> just a check of the shadow byte for 0. For smaller accesses we have
> more complex instrumentation that first checks shadow for 0 and then
> does precise check based on size/offset of the access + shadow value.
> That's slower and also increases register pressure and code size
> (which can further reduce performance due to icache overflow). If we
> increase scale to 16/32, all accesses will need that slow path.
> Another thing is stack instrumentation: larger scale will require
> larger redzones to ensure proper alignment. That will increase stack
> frames and also more instructions to poison/unpoison stack shadow on
> function entry/exit.

Now, I see. Thanks for explanation.

Thanks.

[toc] | [prev] | [next] | [standalone]


#1649140 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromJoonsoo Kim <js1304@gmail.com>
Date2017-05-24 08:20 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tKxOh-6Xm-7@gated-at.bofh.it>
In reply to#1645116
On Mon, May 22, 2017 at 05:00:29PM +0300, Andrey Ryabinin wrote:
> 
> 
> On 05/19/2017 04:53 AM, Joonsoo Kim wrote:
> > On Wed, May 17, 2017 at 03:17:13PM +0300, Andrey Ryabinin wrote:
> >> On 05/16/2017 04:16 AM, js1304@gmail.com wrote:
> >>> From: Joonsoo Kim <iamjoonsoo.kim@lge.com>
> >>>
> >>> Hello, all.
> >>>
> >>> This is an attempt to recude memory consumption of KASAN. Please see
> >>> following description to get the more information.
> >>>
> >>> 1. What is per-page shadow memory
> >>>
> >>> This patch introduces infrastructure to support per-page shadow memory.
> >>> Per-page shadow memory is the same with original shadow memory except
> >>> the granualarity. It's one byte shows the shadow value for the page.
> >>> The purpose of introducing this new shadow memory is to save memory
> >>> consumption.
> >>>
> >>> 2. Problem of current approach
> >>>
> >>> Until now, KASAN needs shadow memory for all the range of the memory
> >>> so the amount of statically allocated memory is so large. It causes
> >>> the problem that KASAN cannot run on the system with hard memory
> >>> constraint. Even if KASAN can run, large memory consumption due to
> >>> KASAN changes behaviour of the workload so we cannot validate
> >>> the moment that we want to check.
> >>>
> >>> 3. How does this patch fix the problem
> >>>
> >>> This patch tries to fix the problem by reducing memory consumption for
> >>> the shadow memory. There are two observations.
> >>>
> >>
> >>
> >> I think that the best way to deal with your problem is to increase shadow scale size.
> >>
> >> You'll need to add tunable to gcc to control shadow size. I expect that gcc has some
> >> places where 8-shadow scale size is hardcoded, but it should be fixable.
> >>
> >> The kernel also have some small amount of code written with KASAN_SHADOW_SCALE_SIZE == 8 in mind,
> >> which should be easy to fix.
> >>
> >> Note that bigger shadow scale size requires bigger alignment of allocated memory and variables.
> >> However, according to comments in gcc/asan.c gcc already aligns stack and global variables and at
> >> 32-bytes boundary.
> >> So we could bump shadow scale up to 32 without increasing current stack consumption.
> >>
> >> On a small machine (1Gb) 1/32 of shadow is just 32Mb which is comparable to yours 30Mb, but I expect it to be
> >> much faster. More importantly, this will require only small amount of simple changes in code, which will be
> >> a *lot* more easier to maintain.
> > 
> > I agree that it is also a good option to reduce memory consumption.
> > Nevertheless, there are two reasons that justifies this patchset.
> > 
> > 1) With this patchset, memory consumption isn't increased in
> > proportional to total memory size. Please consider my 4Gb system
> > example on the below. With increasing shadow scale size to 32, memory
> > would be consumed by 128M. However, this patchset consumed 50MB. This
> > difference can be larger if we run KASAN with bigger machine.
> > 
> 
> Well, yes, but I assume that bigger machine implies that we can use more memory without
> causing a significant change in system's behavior.

In common case, yes. But, I guess that there is a system that
statically uses most of memory and just a few memory is left for others.
For example, consider 64GB system and some program (DB?) runs with
using 60GB. Only 4GB left. If KASAN uses 2GB, just 2GB is left and it
would cause the problem. So, I'd like to insist that this merit 1)
should be considered as valuable.

> 
> > 2) These two optimization can be applied simulatenously. It is just an
> > orthogonal feature. If shadow scale size is increased to 32, memory
> > consumption will be decreased in case of my patchset, too.
> > 
> > Therefore, I think that this patchset is useful in any case.
>  
> These are valid points, but IMO it's not enough to justify this patchset.
> Too much of hacky and fragile code.
> 
> If our goal is to make KASAN to eat less memory, the first step definitely would be a 1/32 shadow.
> Simply because it's the best way to achieve that goal.
> And only if it's not enough we could think about something else, like decreasing/turning off quarantine
> and/or smaller redzones.

Please refer the reply to Dmitry. I think that we need an option that
everything else than something compulsory is the same with non-KASAN
build as much as possible. 1/32 scale would change object layout so it
will not work for this option.

Thanks.

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web