Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1650093 > unrolled thread

Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

Started byJoonsoo Kim <js1304@gmail.com>
First post2017-05-25 02:50 +0200
Last post2017-06-01 20:10 +0200
Articles 16 — 3 participants

Back to article view | Back to linux.kernel

This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by below is the oldest one visible, not the original post.


Contents

  Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-25 02:50 +0200
    Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-29 17:10 +0200
      Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-29 17:20 +0200
        Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-29 17:40 +0200
          Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 10:00 +0200
            Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-30 10:20 +0200
              Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 10:40 +0200
                Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 10:50 +0200
                  Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-30 10:50 +0200
                    Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 11:10 +0200
                      Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-30 11:30 +0200
                        Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 11:40 +0200
                          Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-05-30 11:50 +0200
                            Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Vladimir Murzin <vladimir.murzin@arm.com> - 2017-05-30 12:00 +0200
          Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Joonsoo Kim <js1304@gmail.com> - 2017-05-31 08:00 +0200
          Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to  reduce memory consumption Dmitry Vyukov <dvyukov@google.com> - 2017-06-01 20:10 +0200

#1650093 — Re: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption

FromJoonsoo Kim <js1304@gmail.com>
Date2017-05-25 02:50 +0200
SubjectRe: [PATCH v1 00/11] mm/kasan: support per-page shadow memory to reduce memory consumption
Message-ID<tKP8t-11Y-9@gated-at.bofh.it>
On Wed, May 24, 2017 at 07:19:50PM +0200, Dmitry Vyukov wrote:
> On Wed, May 24, 2017 at 9:45 AM, Joonsoo Kim <js1304@gmail.com> wrote:
> >> > What does make your current patch work then?
> >> > Say we map a new shadow page, update the page shadow to say that there
> >> > is mapped shadow. Then another CPU loads the page shadow and then
> >> > loads from the newly mapped shadow. If we don't flush TLB, what makes
> >> > the second CPU see the newly mapped shadow?
> >>
> >> /\/\/\/\/\/\
> >>
> >> Joonsoo, please answer this question above.
> >
> > Hello, I've answered it in another e-mail however it would not be
> > sufficient. I try again.
> >
> > If the page isn't used for kernel stack, slab, and global variable
> > (aka. kernel memory), black shadow is mapped for the page. We map a
> > new shadow page if the page will be used for kernel memory. We need to
> > flush TLB in all cpus when mapping a new shadow however it's not
> > possible in some cases. So, this patch does just flushing local cpu's
> > TLB. Another cpu could have stale TLB that points black shadow for
> > this page. If that cpu with stale TLB try to check vailidity of the
> > object on this page, result would be invalid since stale TLB points
> > the black shadow and it's shadow value is non-zero. We need a magic
> > here. At this moment, we cannot make sure if invalid is correct result
> > or not since we didn't do full TLB flush. So fixup processing is
> > started. It is implemented in check_memory_region_slow(). Flushing
> > local TLB and re-checking the shadow value. With flushing local TLB,
> > we will use fresh TLB at this time. Therefore, we can pass the
> > validity check as usual.
> >
> >> I am trying to understand if there is any chance to make mapping a
> >> single page for all non-interesting shadow ranges work. That would be
> >
> > This is what this patchset does. Mapping a single (zero/black) shadow
> > page for all non-interesting (non-kernel memory) shadow ranges.
> > There is only single instance of zero/black shadow page. On v1,
> > I used black shadow page only so fail to get enough performance. On
> > v2 mentioned in another thread, I use zero shadow for some region. I
> > guess that performance problem would be gone.
> 
> 
> I can't say I understand everything here, but after staring at the
> patch I don't understand why we need pshadow at all now. Especially
> with this commit
> https://github.com/JoonsooKim/linux/commit/be36ee65f185e3c4026fe93b633056ea811120fb.
> It seems that the current shadow is enough.

pshadow exists for non-kernel memory like as page cache or anonymous page.
This patch doesn't map a new shadow (per-byte shadow) for those pages
to reduce memory consumption. However, we need to know if those page
are allocated or not in order to check the validity of access to those
page. We cannot utilize zero/black shadow page here since mapping
single zero/black shadow page represents eight real page's shadow
value. Instead, we use per-page shadow here and mark/unmark it when
allocation and free happens. With it, we can know the state of the
page and we can determine the validity of access to them.

> If we see bad shadow when the actual shadow value is good, we fall
> onto slow path, flush tlb, reload shadow, see that it is good and
> return. Pshadow is not needed in this case.

For the kernel memory, if we see bad shadow due to *stale TLB*, we
fall onto slow path (check_memory_region_slow()) and flush tlb and
reload shadow.

For the non-kernel memory, if we see bad shadow, we fall onto
pshadow_val() check and we can see actual state of the page.

> If we see good shadow when the actual shadow value is bad, we return
> immediately and get false negative. Pshadow is not involved as well.
> What am I missing?

In this patchset, there is no case that we see good shadow when the
actual (p)shadow value is bad. This case should not happen since we
can miss actual error.

Please let me know that these explanation is insufficient. I will try
more. :)

Thanks.

[toc] | [next] | [standalone]


#1652579

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-29 17:10 +0200
Message-ID<tMusW-1Cu-15@gated-at.bofh.it>
In reply to#1650093
On Thu, May 25, 2017 at 2:41 AM, Joonsoo Kim <js1304@gmail.com> wrote:
> On Wed, May 24, 2017 at 07:19:50PM +0200, Dmitry Vyukov wrote:
>> On Wed, May 24, 2017 at 9:45 AM, Joonsoo Kim <js1304@gmail.com> wrote:
>> >> > What does make your current patch work then?
>> >> > Say we map a new shadow page, update the page shadow to say that there
>> >> > is mapped shadow. Then another CPU loads the page shadow and then
>> >> > loads from the newly mapped shadow. If we don't flush TLB, what makes
>> >> > the second CPU see the newly mapped shadow?
>> >>
>> >> /\/\/\/\/\/\
>> >>
>> >> Joonsoo, please answer this question above.
>> >
>> > Hello, I've answered it in another e-mail however it would not be
>> > sufficient. I try again.
>> >
>> > If the page isn't used for kernel stack, slab, and global variable
>> > (aka. kernel memory), black shadow is mapped for the page. We map a
>> > new shadow page if the page will be used for kernel memory. We need to
>> > flush TLB in all cpus when mapping a new shadow however it's not
>> > possible in some cases. So, this patch does just flushing local cpu's
>> > TLB. Another cpu could have stale TLB that points black shadow for
>> > this page. If that cpu with stale TLB try to check vailidity of the
>> > object on this page, result would be invalid since stale TLB points
>> > the black shadow and it's shadow value is non-zero. We need a magic
>> > here. At this moment, we cannot make sure if invalid is correct result
>> > or not since we didn't do full TLB flush. So fixup processing is
>> > started. It is implemented in check_memory_region_slow(). Flushing
>> > local TLB and re-checking the shadow value. With flushing local TLB,
>> > we will use fresh TLB at this time. Therefore, we can pass the
>> > validity check as usual.
>> >
>> >> I am trying to understand if there is any chance to make mapping a
>> >> single page for all non-interesting shadow ranges work. That would be
>> >
>> > This is what this patchset does. Mapping a single (zero/black) shadow
>> > page for all non-interesting (non-kernel memory) shadow ranges.
>> > There is only single instance of zero/black shadow page. On v1,
>> > I used black shadow page only so fail to get enough performance. On
>> > v2 mentioned in another thread, I use zero shadow for some region. I
>> > guess that performance problem would be gone.
>>
>>
>> I can't say I understand everything here, but after staring at the
>> patch I don't understand why we need pshadow at all now. Especially
>> with this commit
>> https://github.com/JoonsooKim/linux/commit/be36ee65f185e3c4026fe93b633056ea811120fb.
>> It seems that the current shadow is enough.
>
> pshadow exists for non-kernel memory like as page cache or anonymous page.
> This patch doesn't map a new shadow (per-byte shadow) for those pages
> to reduce memory consumption. However, we need to know if those page
> are allocated or not in order to check the validity of access to those
> page. We cannot utilize zero/black shadow page here since mapping
> single zero/black shadow page represents eight real page's shadow
> value. Instead, we use per-page shadow here and mark/unmark it when
> allocation and free happens. With it, we can know the state of the
> page and we can determine the validity of access to them.

I see the problem with 8 kernel pages mapped to a single shadow page.


>> If we see bad shadow when the actual shadow value is good, we fall
>> onto slow path, flush tlb, reload shadow, see that it is good and
>> return. Pshadow is not needed in this case.
>
> For the kernel memory, if we see bad shadow due to *stale TLB*, we
> fall onto slow path (check_memory_region_slow()) and flush tlb and
> reload shadow.
>
> For the non-kernel memory, if we see bad shadow, we fall onto
> pshadow_val() check and we can see actual state of the page.
>
>> If we see good shadow when the actual shadow value is bad, we return
>> immediately and get false negative. Pshadow is not involved as well.
>> What am I missing?
>
> In this patchset, there is no case that we see good shadow when the
> actual (p)shadow value is bad. This case should not happen since we
> can miss actual error.

But why is not it possible?
Let's say we have a real shadow page allocated for range of kernel
memory. Then we unmap the shadow page and map the back page (maybe
even unmap the black page and map another real shadow page). Then
another CPU reads shadow for this range. What prevents it from seeing
the old shadow page?

[toc] | [prev] | [next] | [standalone]


#1652583

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-29 17:20 +0200
Message-ID<tMuCB-1GC-9@gated-at.bofh.it>
In reply to#1652579
On Mon, May 29, 2017 at 5:07 PM, Dmitry Vyukov <dvyukov@google.com> wrote:
> On Thu, May 25, 2017 at 2:41 AM, Joonsoo Kim <js1304@gmail.com> wrote:
>> On Wed, May 24, 2017 at 07:19:50PM +0200, Dmitry Vyukov wrote:
>>> On Wed, May 24, 2017 at 9:45 AM, Joonsoo Kim <js1304@gmail.com> wrote:
>>> >> > What does make your current patch work then?
>>> >> > Say we map a new shadow page, update the page shadow to say that there
>>> >> > is mapped shadow. Then another CPU loads the page shadow and then
>>> >> > loads from the newly mapped shadow. If we don't flush TLB, what makes
>>> >> > the second CPU see the newly mapped shadow?
>>> >>
>>> >> /\/\/\/\/\/\
>>> >>
>>> >> Joonsoo, please answer this question above.
>>> >
>>> > Hello, I've answered it in another e-mail however it would not be
>>> > sufficient. I try again.
>>> >
>>> > If the page isn't used for kernel stack, slab, and global variable
>>> > (aka. kernel memory), black shadow is mapped for the page. We map a
>>> > new shadow page if the page will be used for kernel memory. We need to
>>> > flush TLB in all cpus when mapping a new shadow however it's not
>>> > possible in some cases. So, this patch does just flushing local cpu's
>>> > TLB. Another cpu could have stale TLB that points black shadow for
>>> > this page. If that cpu with stale TLB try to check vailidity of the
>>> > object on this page, result would be invalid since stale TLB points
>>> > the black shadow and it's shadow value is non-zero. We need a magic
>>> > here. At this moment, we cannot make sure if invalid is correct result
>>> > or not since we didn't do full TLB flush. So fixup processing is
>>> > started. It is implemented in check_memory_region_slow(). Flushing
>>> > local TLB and re-checking the shadow value. With flushing local TLB,
>>> > we will use fresh TLB at this time. Therefore, we can pass the
>>> > validity check as usual.
>>> >
>>> >> I am trying to understand if there is any chance to make mapping a
>>> >> single page for all non-interesting shadow ranges work. That would be
>>> >
>>> > This is what this patchset does. Mapping a single (zero/black) shadow
>>> > page for all non-interesting (non-kernel memory) shadow ranges.
>>> > There is only single instance of zero/black shadow page. On v1,
>>> > I used black shadow page only so fail to get enough performance. On
>>> > v2 mentioned in another thread, I use zero shadow for some region. I
>>> > guess that performance problem would be gone.
>>>
>>>
>>> I can't say I understand everything here, but after staring at the
>>> patch I don't understand why we need pshadow at all now. Especially
>>> with this commit
>>> https://github.com/JoonsooKim/linux/commit/be36ee65f185e3c4026fe93b633056ea811120fb.
>>> It seems that the current shadow is enough.
>>
>> pshadow exists for non-kernel memory like as page cache or anonymous page.
>> This patch doesn't map a new shadow (per-byte shadow) for those pages
>> to reduce memory consumption. However, we need to know if those page
>> are allocated or not in order to check the validity of access to those
>> page. We cannot utilize zero/black shadow page here since mapping
>> single zero/black shadow page represents eight real page's shadow
>> value. Instead, we use per-page shadow here and mark/unmark it when
>> allocation and free happens. With it, we can know the state of the
>> page and we can determine the validity of access to them.
>
> I see the problem with 8 kernel pages mapped to a single shadow page.
>
>
>>> If we see bad shadow when the actual shadow value is good, we fall
>>> onto slow path, flush tlb, reload shadow, see that it is good and
>>> return. Pshadow is not needed in this case.
>>
>> For the kernel memory, if we see bad shadow due to *stale TLB*, we
>> fall onto slow path (check_memory_region_slow()) and flush tlb and
>> reload shadow.
>>
>> For the non-kernel memory, if we see bad shadow, we fall onto
>> pshadow_val() check and we can see actual state of the page.
>>
>>> If we see good shadow when the actual shadow value is bad, we return
>>> immediately and get false negative. Pshadow is not involved as well.
>>> What am I missing?
>>
>> In this patchset, there is no case that we see good shadow when the
>> actual (p)shadow value is bad. This case should not happen since we
>> can miss actual error.
>
> But why is not it possible?
> Let's say we have a real shadow page allocated for range of kernel
> memory. Then we unmap the shadow page and map the back page (maybe
> even unmap the black page and map another real shadow page). Then
> another CPU reads shadow for this range. What prevents it from seeing
> the old shadow page?


Re the async processing in kasan_unmap_shadow_workfn. Can't it lead to
shadow corruption? It seems that it can cause unsynchronized state of
shadow pages and corresponding kernel pages in page alloc.
Consider that we schedule unmap of some pages in kasan_unmap_shadow.
Then the range is reallocated in page_alloc and we get into
kasan_map_shadow, which tries to map shadow for these pages again, but
since they are already mapped it bails out. Then
kasan_unmap_shadow_workfn starts and unmaps shadow for the range.

[toc] | [prev] | [next] | [standalone]


#1652599

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-29 17:40 +0200
Message-ID<tMuVY-1P8-35@gated-at.bofh.it>
In reply to#1652583
On Mon, May 29, 2017 at 5:12 PM, Dmitry Vyukov <dvyukov@google.com> wrote:
>>>> >> > What does make your current patch work then?
>>>> >> > Say we map a new shadow page, update the page shadow to say that there
>>>> >> > is mapped shadow. Then another CPU loads the page shadow and then
>>>> >> > loads from the newly mapped shadow. If we don't flush TLB, what makes
>>>> >> > the second CPU see the newly mapped shadow?
>>>> >>
>>>> >> /\/\/\/\/\/\
>>>> >>
>>>> >> Joonsoo, please answer this question above.
>>>> >
>>>> > Hello, I've answered it in another e-mail however it would not be
>>>> > sufficient. I try again.
>>>> >
>>>> > If the page isn't used for kernel stack, slab, and global variable
>>>> > (aka. kernel memory), black shadow is mapped for the page. We map a
>>>> > new shadow page if the page will be used for kernel memory. We need to
>>>> > flush TLB in all cpus when mapping a new shadow however it's not
>>>> > possible in some cases. So, this patch does just flushing local cpu's
>>>> > TLB. Another cpu could have stale TLB that points black shadow for
>>>> > this page. If that cpu with stale TLB try to check vailidity of the
>>>> > object on this page, result would be invalid since stale TLB points
>>>> > the black shadow and it's shadow value is non-zero. We need a magic
>>>> > here. At this moment, we cannot make sure if invalid is correct result
>>>> > or not since we didn't do full TLB flush. So fixup processing is
>>>> > started. It is implemented in check_memory_region_slow(). Flushing
>>>> > local TLB and re-checking the shadow value. With flushing local TLB,
>>>> > we will use fresh TLB at this time. Therefore, we can pass the
>>>> > validity check as usual.
>>>> >
>>>> >> I am trying to understand if there is any chance to make mapping a
>>>> >> single page for all non-interesting shadow ranges work. That would be
>>>> >
>>>> > This is what this patchset does. Mapping a single (zero/black) shadow
>>>> > page for all non-interesting (non-kernel memory) shadow ranges.
>>>> > There is only single instance of zero/black shadow page. On v1,
>>>> > I used black shadow page only so fail to get enough performance. On
>>>> > v2 mentioned in another thread, I use zero shadow for some region. I
>>>> > guess that performance problem would be gone.
>>>>
>>>>
>>>> I can't say I understand everything here, but after staring at the
>>>> patch I don't understand why we need pshadow at all now. Especially
>>>> with this commit
>>>> https://github.com/JoonsooKim/linux/commit/be36ee65f185e3c4026fe93b633056ea811120fb.
>>>> It seems that the current shadow is enough.
>>>
>>> pshadow exists for non-kernel memory like as page cache or anonymous page.
>>> This patch doesn't map a new shadow (per-byte shadow) for those pages
>>> to reduce memory consumption. However, we need to know if those page
>>> are allocated or not in order to check the validity of access to those
>>> page. We cannot utilize zero/black shadow page here since mapping
>>> single zero/black shadow page represents eight real page's shadow
>>> value. Instead, we use per-page shadow here and mark/unmark it when
>>> allocation and free happens. With it, we can know the state of the
>>> page and we can determine the validity of access to them.
>>
>> I see the problem with 8 kernel pages mapped to a single shadow page.
>>
>>
>>>> If we see bad shadow when the actual shadow value is good, we fall
>>>> onto slow path, flush tlb, reload shadow, see that it is good and
>>>> return. Pshadow is not needed in this case.
>>>
>>> For the kernel memory, if we see bad shadow due to *stale TLB*, we
>>> fall onto slow path (check_memory_region_slow()) and flush tlb and
>>> reload shadow.
>>>
>>> For the non-kernel memory, if we see bad shadow, we fall onto
>>> pshadow_val() check and we can see actual state of the page.
>>>
>>>> If we see good shadow when the actual shadow value is bad, we return
>>>> immediately and get false negative. Pshadow is not involved as well.
>>>> What am I missing?
>>>
>>> In this patchset, there is no case that we see good shadow when the
>>> actual (p)shadow value is bad. This case should not happen since we
>>> can miss actual error.
>>
>> But why is not it possible?
>> Let's say we have a real shadow page allocated for range of kernel
>> memory. Then we unmap the shadow page and map the back page (maybe
>> even unmap the black page and map another real shadow page). Then
>> another CPU reads shadow for this range. What prevents it from seeing
>> the old shadow page?
>
>
> Re the async processing in kasan_unmap_shadow_workfn. Can't it lead to
> shadow corruption? It seems that it can cause unsynchronized state of
> shadow pages and corresponding kernel pages in page alloc.
> Consider that we schedule unmap of some pages in kasan_unmap_shadow.
> Then the range is reallocated in page_alloc and we get into
> kasan_map_shadow, which tries to map shadow for these pages again, but
> since they are already mapped it bails out. Then
> kasan_unmap_shadow_workfn starts and unmaps shadow for the range.


Joonsoo,

I guess mine (and Andrey's) main concern is the amount of additional
complexity (I am still struggling to understand how it all works) and
more arch-dependent code in exchange for moderate memory win.

Joonsoo, Andrey,

I have an alternative proposal. It should be conceptually simpler and
also less arch-dependent. But I don't know if I miss something
important that will render it non working.
Namely, we add a pointer to shadow to the page struct. Then, create a
slab allocator for 512B shadow blocks. Then, attach/detach these
shadow blocks to page structs as necessary. It should lead to even
smaller memory consumption because we won't need a whole shadow page
when only 1 out of 8 corresponding kernel pages are used (we will need
just a single 512B block). I guess with some fragmentation we need
lots of excessive shadow with the current proposed patch.
This does not depend on TLB in any way and does not require hooking
into buddy allocator.
The main downside is that we will need to be careful to not assume
that shadow is continuous. In particular this means that this mode
will work only with outline instrumentation and will need some ifdefs.
Also it will be slower due to the additional indirection when
accessing shadow, but that's meant as "small but slow" mode as far as
I understand.

But the main win as I see it is that that's basically complete support
for 32-bit arches. People do ask about arm32 support:
https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
and probably mips32 is relevant as well.
Such mode does not require a huge continuous address space range, has
minimal memory consumption and requires minimal arch-dependent code.
Works only with outline instrumentation, but I think that's a
reasonable compromise.

What do you think?

[toc] | [prev] | [next] | [standalone]


#1652912

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 10:00 +0200
Message-ID<tMKem-4q1-9@gated-at.bofh.it>
In reply to#1652599
On 29/05/17 16:29, Dmitry Vyukov wrote:
> I have an alternative proposal. It should be conceptually simpler and
> also less arch-dependent. But I don't know if I miss something
> important that will render it non working.
> Namely, we add a pointer to shadow to the page struct. Then, create a
> slab allocator for 512B shadow blocks. Then, attach/detach these
> shadow blocks to page structs as necessary. It should lead to even
> smaller memory consumption because we won't need a whole shadow page
> when only 1 out of 8 corresponding kernel pages are used (we will need
> just a single 512B block). I guess with some fragmentation we need
> lots of excessive shadow with the current proposed patch.
> This does not depend on TLB in any way and does not require hooking
> into buddy allocator.
> The main downside is that we will need to be careful to not assume
> that shadow is continuous. In particular this means that this mode
> will work only with outline instrumentation and will need some ifdefs.
> Also it will be slower due to the additional indirection when
> accessing shadow, but that's meant as "small but slow" mode as far as
> I understand.
> 
> But the main win as I see it is that that's basically complete support
> for 32-bit arches. People do ask about arm32 support:
> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
> and probably mips32 is relevant as well.
> Such mode does not require a huge continuous address space range, has
> minimal memory consumption and requires minimal arch-dependent code.
> Works only with outline instrumentation, but I think that's a
> reasonable compromise.

.. or you can just keep shadow in page extension. It was suggested back in
2015 [1], but seems that lack of stack instrumentation was "no-way"... 

[1] https://lkml.org/lkml/2015/8/24/573 

Cheers
Vladimir

[toc] | [prev] | [next] | [standalone]


#1652928

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-30 10:20 +0200
Message-ID<tMKxI-4OJ-13@gated-at.bofh.it>
In reply to#1652912
On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
<vladimir.murzin@arm.com> wrote:
> On 29/05/17 16:29, Dmitry Vyukov wrote:
>> I have an alternative proposal. It should be conceptually simpler and
>> also less arch-dependent. But I don't know if I miss something
>> important that will render it non working.
>> Namely, we add a pointer to shadow to the page struct. Then, create a
>> slab allocator for 512B shadow blocks. Then, attach/detach these
>> shadow blocks to page structs as necessary. It should lead to even
>> smaller memory consumption because we won't need a whole shadow page
>> when only 1 out of 8 corresponding kernel pages are used (we will need
>> just a single 512B block). I guess with some fragmentation we need
>> lots of excessive shadow with the current proposed patch.
>> This does not depend on TLB in any way and does not require hooking
>> into buddy allocator.
>> The main downside is that we will need to be careful to not assume
>> that shadow is continuous. In particular this means that this mode
>> will work only with outline instrumentation and will need some ifdefs.
>> Also it will be slower due to the additional indirection when
>> accessing shadow, but that's meant as "small but slow" mode as far as
>> I understand.
>>
>> But the main win as I see it is that that's basically complete support
>> for 32-bit arches. People do ask about arm32 support:
>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>> and probably mips32 is relevant as well.
>> Such mode does not require a huge continuous address space range, has
>> minimal memory consumption and requires minimal arch-dependent code.
>> Works only with outline instrumentation, but I think that's a
>> reasonable compromise.
>
> .. or you can just keep shadow in page extension. It was suggested back in
> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>
> [1] https://lkml.org/lkml/2015/8/24/573

Right. It describes basically the same idea.

How is page_ext better than adding data page struct?
It seems that memory for all page_ext is preallocated along with page
structs; but just the lookup is slower.

[toc] | [prev] | [next] | [standalone]


#1652950

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 10:40 +0200
Message-ID<tMKR4-4Wf-29@gated-at.bofh.it>
In reply to#1652928
On 30/05/17 09:15, Dmitry Vyukov wrote:
> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
> <vladimir.murzin@arm.com> wrote:
>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>> I have an alternative proposal. It should be conceptually simpler and
>>> also less arch-dependent. But I don't know if I miss something
>>> important that will render it non working.
>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>> shadow blocks to page structs as necessary. It should lead to even
>>> smaller memory consumption because we won't need a whole shadow page
>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>> just a single 512B block). I guess with some fragmentation we need
>>> lots of excessive shadow with the current proposed patch.
>>> This does not depend on TLB in any way and does not require hooking
>>> into buddy allocator.
>>> The main downside is that we will need to be careful to not assume
>>> that shadow is continuous. In particular this means that this mode
>>> will work only with outline instrumentation and will need some ifdefs.
>>> Also it will be slower due to the additional indirection when
>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>> I understand.
>>>
>>> But the main win as I see it is that that's basically complete support
>>> for 32-bit arches. People do ask about arm32 support:
>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>> and probably mips32 is relevant as well.
>>> Such mode does not require a huge continuous address space range, has
>>> minimal memory consumption and requires minimal arch-dependent code.
>>> Works only with outline instrumentation, but I think that's a
>>> reasonable compromise.
>>
>> .. or you can just keep shadow in page extension. It was suggested back in
>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>
>> [1] https://lkml.org/lkml/2015/8/24/573
> 
> Right. It describes basically the same idea.
> 
> How is page_ext better than adding data page struct?

page_ext is already here along with some other debug options ;)

> It seems that memory for all page_ext is preallocated along with page
> structs; but just the lookup is slower.
> 

Yup. Lookup would look like (based on v4.0):

...
page_ext = lookup_page_ext_begin(virt_to_page(start));

do {
	page_ext->shadow[idx++] = value;
} while (idx < bound);

lookup_page_ext_end((void *)page_ext);

...

Cheers
Vladimir

[toc] | [prev] | [next] | [standalone]


#1652953

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 10:50 +0200
Message-ID<tML0J-504-1@gated-at.bofh.it>
In reply to#1652950
On 30/05/17 09:31, Vladimir Murzin wrote:
> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
> 
> On 30/05/17 09:15, Dmitry Vyukov wrote:
>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>> <vladimir.murzin@arm.com> wrote:
>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>> I have an alternative proposal. It should be conceptually simpler and
>>>> also less arch-dependent. But I don't know if I miss something
>>>> important that will render it non working.
>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>> shadow blocks to page structs as necessary. It should lead to even
>>>> smaller memory consumption because we won't need a whole shadow page
>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>> just a single 512B block). I guess with some fragmentation we need
>>>> lots of excessive shadow with the current proposed patch.
>>>> This does not depend on TLB in any way and does not require hooking
>>>> into buddy allocator.
>>>> The main downside is that we will need to be careful to not assume
>>>> that shadow is continuous. In particular this means that this mode
>>>> will work only with outline instrumentation and will need some ifdefs.
>>>> Also it will be slower due to the additional indirection when
>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>> I understand.
>>>>
>>>> But the main win as I see it is that that's basically complete support
>>>> for 32-bit arches. People do ask about arm32 support:
>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>> and probably mips32 is relevant as well.
>>>> Such mode does not require a huge continuous address space range, has
>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>> Works only with outline instrumentation, but I think that's a
>>>> reasonable compromise.
>>>
>>> .. or you can just keep shadow in page extension. It was suggested back in
>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>
>>> [1] https://lkml.org/lkml/2015/8/24/573
>>
>> Right. It describes basically the same idea.
>>
>> How is page_ext better than adding data page struct?
> 
> page_ext is already here along with some other debug options ;)
> 
>> It seems that memory for all page_ext is preallocated along with page
>> structs; but just the lookup is slower.
>>
> 
> Yup. Lookup would look like (based on v4.0):
> 
> ...
> page_ext = lookup_page_ext_begin(virt_to_page(start));
> 
> do {
>         page_ext->shadow[idx++] = value;
> } while (idx < bound);
> 
> lookup_page_ext_end((void *)page_ext);
> 
> ...

Correction: please, ignore that *_{begin,end} stuff - mainline only
lookup_page_ext() is only used.

Cheers
Vladimir

> 
> Cheers
> Vladimir
> 
> 
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org.  For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
> 

[toc] | [prev] | [next] | [standalone]


#1652958

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-30 10:50 +0200
Message-ID<tML0K-504-25@gated-at.bofh.it>
In reply to#1652953
On Tue, May 30, 2017 at 10:40 AM, Vladimir Murzin
<vladimir.murzin@arm.com> wrote:
> On 30/05/17 09:31, Vladimir Murzin wrote:
>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
>>
>> On 30/05/17 09:15, Dmitry Vyukov wrote:
>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>>> <vladimir.murzin@arm.com> wrote:
>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>>> I have an alternative proposal. It should be conceptually simpler and
>>>>> also less arch-dependent. But I don't know if I miss something
>>>>> important that will render it non working.
>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>>> shadow blocks to page structs as necessary. It should lead to even
>>>>> smaller memory consumption because we won't need a whole shadow page
>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>>> just a single 512B block). I guess with some fragmentation we need
>>>>> lots of excessive shadow with the current proposed patch.
>>>>> This does not depend on TLB in any way and does not require hooking
>>>>> into buddy allocator.
>>>>> The main downside is that we will need to be careful to not assume
>>>>> that shadow is continuous. In particular this means that this mode
>>>>> will work only with outline instrumentation and will need some ifdefs.
>>>>> Also it will be slower due to the additional indirection when
>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>>> I understand.
>>>>>
>>>>> But the main win as I see it is that that's basically complete support
>>>>> for 32-bit arches. People do ask about arm32 support:
>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>>> and probably mips32 is relevant as well.
>>>>> Such mode does not require a huge continuous address space range, has
>>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>>> Works only with outline instrumentation, but I think that's a
>>>>> reasonable compromise.
>>>>
>>>> .. or you can just keep shadow in page extension. It was suggested back in
>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>>
>>>> [1] https://lkml.org/lkml/2015/8/24/573
>>>
>>> Right. It describes basically the same idea.
>>>
>>> How is page_ext better than adding data page struct?
>>
>> page_ext is already here along with some other debug options ;)


But page struct is also here. What am I missing?


>>> It seems that memory for all page_ext is preallocated along with page
>>> structs; but just the lookup is slower.
>>>
>>
>> Yup. Lookup would look like (based on v4.0):
>>
>> ...
>> page_ext = lookup_page_ext_begin(virt_to_page(start));
>>
>> do {
>>         page_ext->shadow[idx++] = value;
>> } while (idx < bound);
>>
>> lookup_page_ext_end((void *)page_ext);
>>
>> ...
>
> Correction: please, ignore that *_{begin,end} stuff - mainline only
> lookup_page_ext() is only used.


Note that this added code will be executed during handling of each and
every memory access in kernel. Every instruction matters on that path.
The additional indirection via page struct will also slow down it, but
that's the cost for lower memory consumption and potentially 32-bit
support. For page_ext it looks like even more overhead for no gain.

[toc] | [prev] | [next] | [standalone]


#1652984

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 11:10 +0200
Message-ID<tMLk6-5mP-17@gated-at.bofh.it>
In reply to#1652958
On 30/05/17 09:49, Dmitry Vyukov wrote:
> On Tue, May 30, 2017 at 10:40 AM, Vladimir Murzin
> <vladimir.murzin@arm.com> wrote:
>> On 30/05/17 09:31, Vladimir Murzin wrote:
>>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
>>>
>>> On 30/05/17 09:15, Dmitry Vyukov wrote:
>>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>>>> <vladimir.murzin@arm.com> wrote:
>>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>>>> I have an alternative proposal. It should be conceptually simpler and
>>>>>> also less arch-dependent. But I don't know if I miss something
>>>>>> important that will render it non working.
>>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>>>> shadow blocks to page structs as necessary. It should lead to even
>>>>>> smaller memory consumption because we won't need a whole shadow page
>>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>>>> just a single 512B block). I guess with some fragmentation we need
>>>>>> lots of excessive shadow with the current proposed patch.
>>>>>> This does not depend on TLB in any way and does not require hooking
>>>>>> into buddy allocator.
>>>>>> The main downside is that we will need to be careful to not assume
>>>>>> that shadow is continuous. In particular this means that this mode
>>>>>> will work only with outline instrumentation and will need some ifdefs.
>>>>>> Also it will be slower due to the additional indirection when
>>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>>>> I understand.
>>>>>>
>>>>>> But the main win as I see it is that that's basically complete support
>>>>>> for 32-bit arches. People do ask about arm32 support:
>>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>>>> and probably mips32 is relevant as well.
>>>>>> Such mode does not require a huge continuous address space range, has
>>>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>>>> Works only with outline instrumentation, but I think that's a
>>>>>> reasonable compromise.
>>>>>
>>>>> .. or you can just keep shadow in page extension. It was suggested back in
>>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>>>
>>>>> [1] https://lkml.org/lkml/2015/8/24/573
>>>>
>>>> Right. It describes basically the same idea.
>>>>
>>>> How is page_ext better than adding data page struct?
>>>
>>> page_ext is already here along with some other debug options ;)
> 
> 
> But page struct is also here. What am I missing?
> 

Probably, free room in page struct? I guess most of the page_ext stuff would
love to live in page struct, but... for instance, look at page idle tracking
which has to live in page_ext only for 32-bit.

> 
>>>> It seems that memory for all page_ext is preallocated along with page
>>>> structs; but just the lookup is slower.
>>>>
>>>
>>> Yup. Lookup would look like (based on v4.0):
>>>
>>> ...
>>> page_ext = lookup_page_ext_begin(virt_to_page(start));
>>>
>>> do {
>>>         page_ext->shadow[idx++] = value;
>>> } while (idx < bound);
>>>
>>> lookup_page_ext_end((void *)page_ext);
>>>
>>> ...
>>
>> Correction: please, ignore that *_{begin,end} stuff - mainline only
>> lookup_page_ext() is only used.
> 
> 
> Note that this added code will be executed during handling of each and
> every memory access in kernel. Every instruction matters on that path.

I know, I know... still better than nothing.

> The additional indirection via page struct will also slow down it, but
> that's the cost for lower memory consumption and potentially 32-bit
> support. For page_ext it looks like even more overhead for no gain.
> 

eefa864 (mm/page_ext: resurrect struct page extending code for debugging)
express some cases where keeping data in page_ext has benefit.

Cheers
Vladimir

[toc] | [prev] | [next] | [standalone]


#1653031

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-30 11:30 +0200
Message-ID<tMLDt-5ux-49@gated-at.bofh.it>
In reply to#1652984
On Tue, May 30, 2017 at 11:08 AM, Vladimir Murzin
<vladimir.murzin@arm.com> wrote:
>> <vladimir.murzin@arm.com> wrote:
>>> On 30/05/17 09:31, Vladimir Murzin wrote:
>>>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
>>>>
>>>> On 30/05/17 09:15, Dmitry Vyukov wrote:
>>>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>>>>> <vladimir.murzin@arm.com> wrote:
>>>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>>>>> I have an alternative proposal. It should be conceptually simpler and
>>>>>>> also less arch-dependent. But I don't know if I miss something
>>>>>>> important that will render it non working.
>>>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>>>>> shadow blocks to page structs as necessary. It should lead to even
>>>>>>> smaller memory consumption because we won't need a whole shadow page
>>>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>>>>> just a single 512B block). I guess with some fragmentation we need
>>>>>>> lots of excessive shadow with the current proposed patch.
>>>>>>> This does not depend on TLB in any way and does not require hooking
>>>>>>> into buddy allocator.
>>>>>>> The main downside is that we will need to be careful to not assume
>>>>>>> that shadow is continuous. In particular this means that this mode
>>>>>>> will work only with outline instrumentation and will need some ifdefs.
>>>>>>> Also it will be slower due to the additional indirection when
>>>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>>>>> I understand.
>>>>>>>
>>>>>>> But the main win as I see it is that that's basically complete support
>>>>>>> for 32-bit arches. People do ask about arm32 support:
>>>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>>>>> and probably mips32 is relevant as well.
>>>>>>> Such mode does not require a huge continuous address space range, has
>>>>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>>>>> Works only with outline instrumentation, but I think that's a
>>>>>>> reasonable compromise.
>>>>>>
>>>>>> .. or you can just keep shadow in page extension. It was suggested back in
>>>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>>>>
>>>>>> [1] https://lkml.org/lkml/2015/8/24/573
>>>>>
>>>>> Right. It describes basically the same idea.
>>>>>
>>>>> How is page_ext better than adding data page struct?
>>>>
>>>> page_ext is already here along with some other debug options ;)
>>
>>
>> But page struct is also here. What am I missing?
>>
>
> Probably, free room in page struct? I guess most of the page_ext stuff would
> love to live in page struct, but... for instance, look at page idle tracking
> which has to live in page_ext only for 32-bit.


Sorry for my ignorance. What's the fundamental problem with just
pushing everything into page struct?

I don't see anything relevant in page struct comment. Nor I see "idle"
nor "tracking" page struct. I see only 2 mentions of CONFIG_64BIT, but
both declare the same fields just with different types (int vs short).



>>>>> It seems that memory for all page_ext is preallocated along with page
>>>>> structs; but just the lookup is slower.
>>>>>
>>>>
>>>> Yup. Lookup would look like (based on v4.0):
>>>>
>>>> ...
>>>> page_ext = lookup_page_ext_begin(virt_to_page(start));
>>>>
>>>> do {
>>>>         page_ext->shadow[idx++] = value;
>>>> } while (idx < bound);
>>>>
>>>> lookup_page_ext_end((void *)page_ext);
>>>>
>>>> ...
>>>
>>> Correction: please, ignore that *_{begin,end} stuff - mainline only
>>> lookup_page_ext() is only used.
>>
>>
>> Note that this added code will be executed during handling of each and
>> every memory access in kernel. Every instruction matters on that path.
>
> I know, I know... still better than nothing.
>
>> The additional indirection via page struct will also slow down it, but
>> that's the cost for lower memory consumption and potentially 32-bit
>> support. For page_ext it looks like even more overhead for no gain.
>>
>
> eefa864 (mm/page_ext: resurrect struct page extending code for debugging)
> express some cases where keeping data in page_ext has benefit.
>
> Cheers
> Vladimir

[toc] | [prev] | [next] | [standalone]


#1653041

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 11:40 +0200
Message-ID<tMLN9-5ya-31@gated-at.bofh.it>
In reply to#1653031
On 30/05/17 10:26, Dmitry Vyukov wrote:
> On Tue, May 30, 2017 at 11:08 AM, Vladimir Murzin
> <vladimir.murzin@arm.com> wrote:
>>> <vladimir.murzin@arm.com> wrote:
>>>> On 30/05/17 09:31, Vladimir Murzin wrote:
>>>>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
>>>>>
>>>>> On 30/05/17 09:15, Dmitry Vyukov wrote:
>>>>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>>>>>> <vladimir.murzin@arm.com> wrote:
>>>>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>>>>>> I have an alternative proposal. It should be conceptually simpler and
>>>>>>>> also less arch-dependent. But I don't know if I miss something
>>>>>>>> important that will render it non working.
>>>>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>>>>>> shadow blocks to page structs as necessary. It should lead to even
>>>>>>>> smaller memory consumption because we won't need a whole shadow page
>>>>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>>>>>> just a single 512B block). I guess with some fragmentation we need
>>>>>>>> lots of excessive shadow with the current proposed patch.
>>>>>>>> This does not depend on TLB in any way and does not require hooking
>>>>>>>> into buddy allocator.
>>>>>>>> The main downside is that we will need to be careful to not assume
>>>>>>>> that shadow is continuous. In particular this means that this mode
>>>>>>>> will work only with outline instrumentation and will need some ifdefs.
>>>>>>>> Also it will be slower due to the additional indirection when
>>>>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>>>>>> I understand.
>>>>>>>>
>>>>>>>> But the main win as I see it is that that's basically complete support
>>>>>>>> for 32-bit arches. People do ask about arm32 support:
>>>>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>>>>>> and probably mips32 is relevant as well.
>>>>>>>> Such mode does not require a huge continuous address space range, has
>>>>>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>>>>>> Works only with outline instrumentation, but I think that's a
>>>>>>>> reasonable compromise.
>>>>>>>
>>>>>>> .. or you can just keep shadow in page extension. It was suggested back in
>>>>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>>>>>
>>>>>>> [1] https://lkml.org/lkml/2015/8/24/573
>>>>>>
>>>>>> Right. It describes basically the same idea.
>>>>>>
>>>>>> How is page_ext better than adding data page struct?
>>>>>
>>>>> page_ext is already here along with some other debug options ;)
>>>
>>>
>>> But page struct is also here. What am I missing?
>>>
>>
>> Probably, free room in page struct? I guess most of the page_ext stuff would
>> love to live in page struct, but... for instance, look at page idle tracking
>> which has to live in page_ext only for 32-bit.
> 
> 
> Sorry for my ignorance. What's the fundamental problem with just
> pushing everything into page struct?

I think [1] has an answer for your question ;)

> 
> I don't see anything relevant in page struct comment. Nor I see "idle"
> nor "tracking" page struct. I see only 2 mentions of CONFIG_64BIT, but
> both declare the same fields just with different types (int vs short).

Right, it is because implementation is based on page flags [1]:

Note, since there is no room for extra page flags on 32 bit, this feature
uses extended page flags when compiled on 32 bit.


[1] https://lwn.net/Articles/565097/
[2] 33c3fc7 ("mm: introduce idle page tracking")

Cheers
Vladimir

> 
> 
> 
>>>>>> It seems that memory for all page_ext is preallocated along with page
>>>>>> structs; but just the lookup is slower.
>>>>>>
>>>>>
>>>>> Yup. Lookup would look like (based on v4.0):
>>>>>
>>>>> ...
>>>>> page_ext = lookup_page_ext_begin(virt_to_page(start));
>>>>>
>>>>> do {
>>>>>         page_ext->shadow[idx++] = value;
>>>>> } while (idx < bound);
>>>>>
>>>>> lookup_page_ext_end((void *)page_ext);
>>>>>
>>>>> ...
>>>>
>>>> Correction: please, ignore that *_{begin,end} stuff - mainline only
>>>> lookup_page_ext() is only used.
>>>
>>>
>>> Note that this added code will be executed during handling of each and
>>> every memory access in kernel. Every instruction matters on that path.
>>
>> I know, I know... still better than nothing.
>>
>>> The additional indirection via page struct will also slow down it, but
>>> that's the cost for lower memory consumption and potentially 32-bit
>>> support. For page_ext it looks like even more overhead for no gain.
>>>
>>
>> eefa864 (mm/page_ext: resurrect struct page extending code for debugging)
>> express some cases where keeping data in page_ext has benefit.
>>
>> Cheers
>> Vladimir
> 

[toc] | [prev] | [next] | [standalone]


#1653050

FromDmitry Vyukov <dvyukov@google.com>
Date2017-05-30 11:50 +0200
Message-ID<tMLWO-5BU-31@gated-at.bofh.it>
In reply to#1653041
On Tue, May 30, 2017 at 11:39 AM, Vladimir Murzin
<vladimir.murzin@arm.com> wrote:
>
> On 30/05/17 10:26, Dmitry Vyukov wrote:
> > On Tue, May 30, 2017 at 11:08 AM, Vladimir Murzin
> > <vladimir.murzin@arm.com> wrote:
> >>> <vladimir.murzin@arm.com> wrote:
> >>>> On 30/05/17 09:31, Vladimir Murzin wrote:
> >>>>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
> >>>>>
> >>>>> On 30/05/17 09:15, Dmitry Vyukov wrote:
> >>>>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
> >>>>>> <vladimir.murzin@arm.com> wrote:
> >>>>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
> >>>>>>>> I have an alternative proposal. It should be conceptually simpler and
> >>>>>>>> also less arch-dependent. But I don't know if I miss something
> >>>>>>>> important that will render it non working.
> >>>>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
> >>>>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
> >>>>>>>> shadow blocks to page structs as necessary. It should lead to even
> >>>>>>>> smaller memory consumption because we won't need a whole shadow page
> >>>>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
> >>>>>>>> just a single 512B block). I guess with some fragmentation we need
> >>>>>>>> lots of excessive shadow with the current proposed patch.
> >>>>>>>> This does not depend on TLB in any way and does not require hooking
> >>>>>>>> into buddy allocator.
> >>>>>>>> The main downside is that we will need to be careful to not assume
> >>>>>>>> that shadow is continuous. In particular this means that this mode
> >>>>>>>> will work only with outline instrumentation and will need some ifdefs.
> >>>>>>>> Also it will be slower due to the additional indirection when
> >>>>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
> >>>>>>>> I understand.
> >>>>>>>>
> >>>>>>>> But the main win as I see it is that that's basically complete support
> >>>>>>>> for 32-bit arches. People do ask about arm32 support:
> >>>>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
> >>>>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
> >>>>>>>> and probably mips32 is relevant as well.
> >>>>>>>> Such mode does not require a huge continuous address space range, has
> >>>>>>>> minimal memory consumption and requires minimal arch-dependent code.
> >>>>>>>> Works only with outline instrumentation, but I think that's a
> >>>>>>>> reasonable compromise.
> >>>>>>>
> >>>>>>> .. or you can just keep shadow in page extension. It was suggested back in
> >>>>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
> >>>>>>>
> >>>>>>> [1] https://lkml.org/lkml/2015/8/24/573
> >>>>>>
> >>>>>> Right. It describes basically the same idea.
> >>>>>>
> >>>>>> How is page_ext better than adding data page struct?
> >>>>>
> >>>>> page_ext is already here along with some other debug options ;)
> >>>
> >>>
> >>> But page struct is also here. What am I missing?
> >>>
> >>
> >> Probably, free room in page struct? I guess most of the page_ext stuff would
> >> love to live in page struct, but... for instance, look at page idle tracking
> >> which has to live in page_ext only for 32-bit.
> >
> >
> > Sorry for my ignorance. What's the fundamental problem with just
> > pushing everything into page struct?
>
> I think [1] has an answer for your question ;)

It also has an answer for why we should put it into page struct :)


>
> >
> > I don't see anything relevant in page struct comment. Nor I see "idle"
> > nor "tracking" page struct. I see only 2 mentions of CONFIG_64BIT, but
> > both declare the same fields just with different types (int vs short).
>
> Right, it is because implementation is based on page flags [1]:
>
> Note, since there is no room for extra page flags on 32 bit, this feature
> uses extended page flags when compiled on 32 bit.
>
>
> [1] https://lwn.net/Articles/565097/
> [2] 33c3fc7 ("mm: introduce idle page tracking")

[toc] | [prev] | [next] | [standalone]


#1653062

FromVladimir Murzin <vladimir.murzin@arm.com>
Date2017-05-30 12:00 +0200
Message-ID<tMM6v-5FJ-33@gated-at.bofh.it>
In reply to#1653050
On 30/05/17 10:45, Dmitry Vyukov wrote:
> On Tue, May 30, 2017 at 11:39 AM, Vladimir Murzin
> <vladimir.murzin@arm.com> wrote:
>>
>> On 30/05/17 10:26, Dmitry Vyukov wrote:
>>> On Tue, May 30, 2017 at 11:08 AM, Vladimir Murzin
>>> <vladimir.murzin@arm.com> wrote:
>>>>> <vladimir.murzin@arm.com> wrote:
>>>>>> On 30/05/17 09:31, Vladimir Murzin wrote:
>>>>>>> [This sender failed our fraud detection checks and may not be who they appear to be. Learn about spoofing at http://aka.ms/LearnAboutSpoofing]
>>>>>>>
>>>>>>> On 30/05/17 09:15, Dmitry Vyukov wrote:
>>>>>>>> On Tue, May 30, 2017 at 9:58 AM, Vladimir Murzin
>>>>>>>> <vladimir.murzin@arm.com> wrote:
>>>>>>>>> On 29/05/17 16:29, Dmitry Vyukov wrote:
>>>>>>>>>> I have an alternative proposal. It should be conceptually simpler and
>>>>>>>>>> also less arch-dependent. But I don't know if I miss something
>>>>>>>>>> important that will render it non working.
>>>>>>>>>> Namely, we add a pointer to shadow to the page struct. Then, create a
>>>>>>>>>> slab allocator for 512B shadow blocks. Then, attach/detach these
>>>>>>>>>> shadow blocks to page structs as necessary. It should lead to even
>>>>>>>>>> smaller memory consumption because we won't need a whole shadow page
>>>>>>>>>> when only 1 out of 8 corresponding kernel pages are used (we will need
>>>>>>>>>> just a single 512B block). I guess with some fragmentation we need
>>>>>>>>>> lots of excessive shadow with the current proposed patch.
>>>>>>>>>> This does not depend on TLB in any way and does not require hooking
>>>>>>>>>> into buddy allocator.
>>>>>>>>>> The main downside is that we will need to be careful to not assume
>>>>>>>>>> that shadow is continuous. In particular this means that this mode
>>>>>>>>>> will work only with outline instrumentation and will need some ifdefs.
>>>>>>>>>> Also it will be slower due to the additional indirection when
>>>>>>>>>> accessing shadow, but that's meant as "small but slow" mode as far as
>>>>>>>>>> I understand.
>>>>>>>>>>
>>>>>>>>>> But the main win as I see it is that that's basically complete support
>>>>>>>>>> for 32-bit arches. People do ask about arm32 support:
>>>>>>>>>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>>>>>>>>>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>>>>>>>>>> and probably mips32 is relevant as well.
>>>>>>>>>> Such mode does not require a huge continuous address space range, has
>>>>>>>>>> minimal memory consumption and requires minimal arch-dependent code.
>>>>>>>>>> Works only with outline instrumentation, but I think that's a
>>>>>>>>>> reasonable compromise.
>>>>>>>>>
>>>>>>>>> .. or you can just keep shadow in page extension. It was suggested back in
>>>>>>>>> 2015 [1], but seems that lack of stack instrumentation was "no-way"...
>>>>>>>>>
>>>>>>>>> [1] https://lkml.org/lkml/2015/8/24/573
>>>>>>>>
>>>>>>>> Right. It describes basically the same idea.
>>>>>>>>
>>>>>>>> How is page_ext better than adding data page struct?
>>>>>>>
>>>>>>> page_ext is already here along with some other debug options ;)
>>>>>
>>>>>
>>>>> But page struct is also here. What am I missing?
>>>>>
>>>>
>>>> Probably, free room in page struct? I guess most of the page_ext stuff would
>>>> love to live in page struct, but... for instance, look at page idle tracking
>>>> which has to live in page_ext only for 32-bit.
>>>
>>>
>>> Sorry for my ignorance. What's the fundamental problem with just
>>> pushing everything into page struct?
>>
>> I think [1] has an answer for your question ;)
> 
> It also has an answer for why we should put it into page struct :)

Glad you find it useful ;) I'd be glad to see it lands into 32-bit world :)

Cheers
Vladimir

> 
> 
>>
>>>
>>> I don't see anything relevant in page struct comment. Nor I see "idle"
>>> nor "tracking" page struct. I see only 2 mentions of CONFIG_64BIT, but
>>> both declare the same fields just with different types (int vs short).
>>
>> Right, it is because implementation is based on page flags [1]:
>>
>> Note, since there is no room for extra page flags on 32 bit, this feature
>> uses extended page flags when compiled on 32 bit.
>>
>>
>> [1] https://lwn.net/Articles/565097/
>> [2] 33c3fc7 ("mm: introduce idle page tracking")
> 

[toc] | [prev] | [next] | [standalone]


#1653850

FromJoonsoo Kim <js1304@gmail.com>
Date2017-05-31 08:00 +0200
Message-ID<tN4PL-FB-1@gated-at.bofh.it>
In reply to#1652599
On Tue, May 30, 2017 at 05:16:56PM +0300, Andrey Ryabinin wrote:
> On 05/29/2017 06:29 PM, Dmitry Vyukov wrote:
> > Joonsoo,
> > 
> > I guess mine (and Andrey's) main concern is the amount of additional
> > complexity (I am still struggling to understand how it all works) and
> > more arch-dependent code in exchange for moderate memory win.
> > 
> > Joonsoo, Andrey,
> > 
> > I have an alternative proposal. It should be conceptually simpler and
> > also less arch-dependent. But I don't know if I miss something
> > important that will render it non working.
> > Namely, we add a pointer to shadow to the page struct. Then, create a
> > slab allocator for 512B shadow blocks. Then, attach/detach these
> > shadow blocks to page structs as necessary. It should lead to even
> > smaller memory consumption because we won't need a whole shadow page
> > when only 1 out of 8 corresponding kernel pages are used (we will need
> > just a single 512B block). I guess with some fragmentation we need
> > lots of excessive shadow with the current proposed patch.
> > This does not depend on TLB in any way and does not require hooking
> > into buddy allocator.
> > The main downside is that we will need to be careful to not assume
> > that shadow is continuous. In particular this means that this mode
> > will work only with outline instrumentation and will need some ifdefs.
> > Also it will be slower due to the additional indirection when
> > accessing shadow, but that's meant as "small but slow" mode as far as
> > I understand.
> 
> It seems that you are forgetting about stack instrumentation.
> You'll have to disable it completely, at least with current implementation of it in gcc.

Correct. Even if we use OUTLINE build, gcc directly inserts codes to the
function prologue/epilogue to mark/unmakr the shadow. And, I'm not
sure we can change it since it would affect performance greately. In
current situation, alternative proposal loses most of benefit mentioned
above.
> 
> > But the main win as I see it is that that's basically complete support
> > for 32-bit arches. People do ask about arm32 support:
> > https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
> > https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
> > and probably mips32 is relevant as well.
> 
> I don't see how above is relevant for 32-bit arches. Current design
> is perfectly fine for 32-bit arches. I did some POC arm32 port couple years
> ago - https://github.com/aryabinin/linux/commits/kasan/arm_v0_1
> It has some ugly hacks and non-critical bugs. AFAIR it also super-slow because I (mistakenly) 
> made shadow memory uncached. But otherwise it works.

Could you explain that where is the code to map shadow memory uncached?
I don't find anything related to it.

> > Such mode does not require a huge continuous address space range, has
> > minimal memory consumption and requires minimal arch-dependent code.
> > Works only with outline instrumentation, but I think that's a
> > reasonable compromise.
> > 
> > What do you think?
>  
> I don't understand why we trying to invent some hacky/complex schemes when we already have
> a simple one - scaling shadow to 1/32. It's easy to implement and should be more performant comparing
> to suggested schemes.

My approach can co-exist with changing scaling approach. It has it's
own benefit.

And, as Dmitry mentioned before, scaling shadow to 1/32 also has downsides,
expecially for inline instrumentation. And, it requires compiler
modification and user needs to update their compiler to newer version
which is not so simple in terms of the user's usability

Thanks.

[toc] | [prev] | [next] | [standalone]


#1655653

FromDmitry Vyukov <dvyukov@google.com>
Date2017-06-01 20:10 +0200
Message-ID<tNCHM-61B-25@gated-at.bofh.it>
In reply to#1652599
On Tue, May 30, 2017 at 4:16 PM, Andrey Ryabinin
<aryabinin@virtuozzo.com> wrote:
> On 05/29/2017 06:29 PM, Dmitry Vyukov wrote:
>> Joonsoo,
>>
>> I guess mine (and Andrey's) main concern is the amount of additional
>> complexity (I am still struggling to understand how it all works) and
>> more arch-dependent code in exchange for moderate memory win.
>>
>> Joonsoo, Andrey,
>>
>> I have an alternative proposal. It should be conceptually simpler and
>> also less arch-dependent. But I don't know if I miss something
>> important that will render it non working.
>> Namely, we add a pointer to shadow to the page struct. Then, create a
>> slab allocator for 512B shadow blocks. Then, attach/detach these
>> shadow blocks to page structs as necessary. It should lead to even
>> smaller memory consumption because we won't need a whole shadow page
>> when only 1 out of 8 corresponding kernel pages are used (we will need
>> just a single 512B block). I guess with some fragmentation we need
>> lots of excessive shadow with the current proposed patch.
>> This does not depend on TLB in any way and does not require hooking
>> into buddy allocator.
>> The main downside is that we will need to be careful to not assume
>> that shadow is continuous. In particular this means that this mode
>> will work only with outline instrumentation and will need some ifdefs.
>> Also it will be slower due to the additional indirection when
>> accessing shadow, but that's meant as "small but slow" mode as far as
>> I understand.
>
> It seems that you are forgetting about stack instrumentation.
> You'll have to disable it completely, at least with current implementation of it in gcc.
>
>> But the main win as I see it is that that's basically complete support
>> for 32-bit arches. People do ask about arm32 support:
>> https://groups.google.com/d/msg/kasan-dev/Sk6BsSPMRRc/Gqh4oD_wAAAJ
>> https://groups.google.com/d/msg/kasan-dev/B22vOFp-QWg/EVJPbrsgAgAJ
>> and probably mips32 is relevant as well.
>
> I don't see how above is relevant for 32-bit arches. Current design
> is perfectly fine for 32-bit arches. I did some POC arm32 port couple years
> ago - https://github.com/aryabinin/linux/commits/kasan/arm_v0_1
> It has some ugly hacks and non-critical bugs. AFAIR it also super-slow because I (mistakenly)
> made shadow memory uncached. But otherwise it works.
>
>> Such mode does not require a huge continuous address space range, has
>> minimal memory consumption and requires minimal arch-dependent code.
>> Works only with outline instrumentation, but I think that's a
>> reasonable compromise.
>>
>> What do you think?
>
> I don't understand why we trying to invent some hacky/complex schemes when we already have
> a simple one - scaling shadow to 1/32. It's easy to implement and should be more performant comparing
> to suggested schemes.


If 32-bits work with the current approach, then I would also prefer to
keep things simpler.
FWIW clang supports settings shadow scale via a command line flag
(-asan-mapping-scale).

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web