Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1739652 > unrolled thread

[RFC] a question about mlockall() and mprotect()

Started byXishi Qiu <qiuxishi@huawei.com>
First post2017-09-26 10:00 +0200
Last post2017-09-27 08:00 +0200
Articles 10 — 3 participants

Back to article view | Back to linux.kernel


Contents

  [RFC] a question about mlockall() and mprotect() Xishi Qiu <qiuxishi@huawei.com> - 2017-09-26 10:00 +0200
    Re: [RFC] a question about mlockall() and mprotect() Michal Hocko <mhocko@kernel.org> - 2017-09-26 10:20 +0200
      Re: [RFC] a question about mlockall() and mprotect() Xishi Qiu <qiuxishi@huawei.com> - 2017-09-26 10:50 +0200
        Re: [RFC] a question about mlockall() and mprotect() Michal Hocko <mhocko@kernel.org> - 2017-09-26 11:10 +0200
          Re: [RFC] a question about mlockall() and mprotect() Xishi Qiu <qiuxishi@huawei.com> - 2017-09-26 11:20 +0200
            Re: [RFC] a question about mlockall() and mprotect() Michal Hocko <mhocko@kernel.org> - 2017-09-26 11:30 +0200
            Re: [RFC] a question about mlockall() and mprotect() Xishi Qiu <qiuxishi@huawei.com> - 2017-09-26 11:30 +0200
              Re: [RFC] a question about mlockall() and mprotect() Vlastimil Babka <vbabka@suse.cz> - 2017-09-26 11:50 +0200
                Re: [RFC] a question about mlockall() and mprotect() Michal Hocko <mhocko@kernel.org> - 2017-09-26 13:10 +0200
                  Re: [RFC] a question about mlockall() and mprotect() Xishi Qiu <qiuxishi@huawei.com> - 2017-09-27 08:00 +0200

#1739652 — [RFC] a question about mlockall() and mprotect()

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-09-26 10:00 +0200
Subject[RFC] a question about mlockall() and mprotect()
Message-ID<utSWD-1N5-43@gated-at.bofh.it>
When we call mlockall(), we will add VM_LOCKED to the vma,
if the vma prot is ---p, then mm_populate -> get_user_pages
will not alloc memory.

I find it said "ignore errors" in mm_populate()
static inline void mm_populate(unsigned long addr, unsigned long len)
{
	/* Ignore errors */
	(void) __mm_populate(addr, len, 1);
}

And later we call mprotect() to change the prot, then it is
still not alloc memory for the mlocked vma.

My question is that, shall we alloc memory if the prot changed,
and who(kernel, glibc, user) should alloc the memory?

Thanks,
Xishi Qiu

[toc] | [next] | [standalone]


#1739662

FromMichal Hocko <mhocko@kernel.org>
Date2017-09-26 10:20 +0200
Message-ID<utTfY-29X-23@gated-at.bofh.it>
In reply to#1739652
On Tue 26-09-17 15:56:55, Xishi Qiu wrote:
> When we call mlockall(), we will add VM_LOCKED to the vma,
> if the vma prot is ---p,

not sure what you mean here. apply_mlockall_flags will set the flag on
all vmas except for special mappings (mlock_fixup). This phase will
cause that memory reclaim will not free already mapped pages in those
vmas (see page_check_references and the lazy mlock pages move to
unevictable LRUs).

> then mm_populate -> get_user_pages will not alloc memory.

mm_populate all the vmas with pages. Well there are certainly some
constrains - e.g. memory cgroup hard limit might be hit and so the
faulting might fail.

> I find it said "ignore errors" in mm_populate()
> static inline void mm_populate(unsigned long addr, unsigned long len)
> {
> 	/* Ignore errors */
> 	(void) __mm_populate(addr, len, 1);
> }

But we do not report the failure because any failure past
apply_mlockall_flags would be tricky to handle. We have already dropped
the mmap_sem lock so some other address space operations could have
interfered.
 
> And later we call mprotect() to change the prot, then it is
> still not alloc memory for the mlocked vma.
> 
> My question is that, shall we alloc memory if the prot changed,
> and who(kernel, glibc, user) should alloc the memory?

I do not understand your question but if you are asking how to get pages
to map your vmas then touching that area will fault the memory in.
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1739708

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-09-26 10:50 +0200
Message-ID<utTJ0-2nL-27@gated-at.bofh.it>
In reply to#1739662
On 2017/9/26 16:17, Michal Hocko wrote:

> On Tue 26-09-17 15:56:55, Xishi Qiu wrote:
>> When we call mlockall(), we will add VM_LOCKED to the vma,
>> if the vma prot is ---p,
> 
> not sure what you mean here. apply_mlockall_flags will set the flag on
> all vmas except for special mappings (mlock_fixup). This phase will
> cause that memory reclaim will not free already mapped pages in those
> vmas (see page_check_references and the lazy mlock pages move to
> unevictable LRUs).
> 
>> then mm_populate -> get_user_pages will not alloc memory.
> 
> mm_populate all the vmas with pages. Well there are certainly some
> constrains - e.g. memory cgroup hard limit might be hit and so the
> faulting might fail.
> 
>> I find it said "ignore errors" in mm_populate()
>> static inline void mm_populate(unsigned long addr, unsigned long len)
>> {
>> 	/* Ignore errors */
>> 	(void) __mm_populate(addr, len, 1);
>> }
> 
> But we do not report the failure because any failure past
> apply_mlockall_flags would be tricky to handle. We have already dropped
> the mmap_sem lock so some other address space operations could have
> interfered.
>  
>> And later we call mprotect() to change the prot, then it is
>> still not alloc memory for the mlocked vma.
>>
>> My question is that, shall we alloc memory if the prot changed,
>> and who(kernel, glibc, user) should alloc the memory?
> 
> I do not understand your question but if you are asking how to get pages
> to map your vmas then touching that area will fault the memory in.

Hi Michal,

syscall mlockall() will first apply the VM_LOCKED to the vma, then
call mm_populate() to map the vmas.

mm_populate
	populate_vma_page_range
		__get_user_pages
			check_vma_flags
And the above path maybe return -EFAULT in some case, right?

If we call mprotect() to change the prot of vma, just let
check_vma_flags() return 0, then we will get the mlocked pages
in following page-fault, right?

My question is that, shall we map the vmas immediately when
the prot changed? If we should map it immediately, who(kernel, glibc, user)
do this step?

Thanks,
Xishi Qiu

[toc] | [prev] | [next] | [standalone]


#1739727

FromMichal Hocko <mhocko@kernel.org>
Date2017-09-26 11:10 +0200
Message-ID<utU2m-2KN-31@gated-at.bofh.it>
In reply to#1739708
On Tue 26-09-17 16:39:56, Xishi Qiu wrote:
> On 2017/9/26 16:17, Michal Hocko wrote:
> 
> > On Tue 26-09-17 15:56:55, Xishi Qiu wrote:
> >> When we call mlockall(), we will add VM_LOCKED to the vma,
> >> if the vma prot is ---p,
> > 
> > not sure what you mean here. apply_mlockall_flags will set the flag on
> > all vmas except for special mappings (mlock_fixup). This phase will
> > cause that memory reclaim will not free already mapped pages in those
> > vmas (see page_check_references and the lazy mlock pages move to
> > unevictable LRUs).
> > 
> >> then mm_populate -> get_user_pages will not alloc memory.
> > 
> > mm_populate all the vmas with pages. Well there are certainly some
> > constrains - e.g. memory cgroup hard limit might be hit and so the
> > faulting might fail.
> > 
> >> I find it said "ignore errors" in mm_populate()
> >> static inline void mm_populate(unsigned long addr, unsigned long len)
> >> {
> >> 	/* Ignore errors */
> >> 	(void) __mm_populate(addr, len, 1);
> >> }
> > 
> > But we do not report the failure because any failure past
> > apply_mlockall_flags would be tricky to handle. We have already dropped
> > the mmap_sem lock so some other address space operations could have
> > interfered.
> >  
> >> And later we call mprotect() to change the prot, then it is
> >> still not alloc memory for the mlocked vma.
> >>
> >> My question is that, shall we alloc memory if the prot changed,
> >> and who(kernel, glibc, user) should alloc the memory?
> > 
> > I do not understand your question but if you are asking how to get pages
> > to map your vmas then touching that area will fault the memory in.
> 
> Hi Michal,
> 
> syscall mlockall() will first apply the VM_LOCKED to the vma, then
> call mm_populate() to map the vmas.
> 
> mm_populate
> 	populate_vma_page_range
> 		__get_user_pages
> 			check_vma_flags
> And the above path maybe return -EFAULT in some case, right?
> 
> If we call mprotect() to change the prot of vma, just let
> check_vma_flags() return 0, then we will get the mlocked pages
> in following page-fault, right?

Any future page fault to the existing vma will result in the mlocked
page. That is what VM_LOCKED guarantess.

> My question is that, shall we map the vmas immediately when
> the prot changed? If we should map it immediately, who(kernel, glibc, user)
> do this step?

This is still very fuzzy. What are you actually trying to achieve?
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1739735

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-09-26 11:20 +0200
Message-ID<utUc3-2OD-31@gated-at.bofh.it>
In reply to#1739727
On 2017/9/26 17:02, Michal Hocko wrote:

> On Tue 26-09-17 16:39:56, Xishi Qiu wrote:
>> On 2017/9/26 16:17, Michal Hocko wrote:
>>
>>> On Tue 26-09-17 15:56:55, Xishi Qiu wrote:
>>>> When we call mlockall(), we will add VM_LOCKED to the vma,
>>>> if the vma prot is ---p,
>>>
>>> not sure what you mean here. apply_mlockall_flags will set the flag on
>>> all vmas except for special mappings (mlock_fixup). This phase will
>>> cause that memory reclaim will not free already mapped pages in those
>>> vmas (see page_check_references and the lazy mlock pages move to
>>> unevictable LRUs).
>>>
>>>> then mm_populate -> get_user_pages will not alloc memory.
>>>
>>> mm_populate all the vmas with pages. Well there are certainly some
>>> constrains - e.g. memory cgroup hard limit might be hit and so the
>>> faulting might fail.
>>>
>>>> I find it said "ignore errors" in mm_populate()
>>>> static inline void mm_populate(unsigned long addr, unsigned long len)
>>>> {
>>>> 	/* Ignore errors */
>>>> 	(void) __mm_populate(addr, len, 1);
>>>> }
>>>
>>> But we do not report the failure because any failure past
>>> apply_mlockall_flags would be tricky to handle. We have already dropped
>>> the mmap_sem lock so some other address space operations could have
>>> interfered.
>>>  
>>>> And later we call mprotect() to change the prot, then it is
>>>> still not alloc memory for the mlocked vma.
>>>>
>>>> My question is that, shall we alloc memory if the prot changed,
>>>> and who(kernel, glibc, user) should alloc the memory?
>>>
>>> I do not understand your question but if you are asking how to get pages
>>> to map your vmas then touching that area will fault the memory in.
>>
>> Hi Michal,
>>
>> syscall mlockall() will first apply the VM_LOCKED to the vma, then
>> call mm_populate() to map the vmas.
>>
>> mm_populate
>> 	populate_vma_page_range
>> 		__get_user_pages
>> 			check_vma_flags
>> And the above path maybe return -EFAULT in some case, right?
>>
>> If we call mprotect() to change the prot of vma, just let
>> check_vma_flags() return 0, then we will get the mlocked pages
>> in following page-fault, right?
> 
> Any future page fault to the existing vma will result in the mlocked
> page. That is what VM_LOCKED guarantess.
> 
>> My question is that, shall we map the vmas immediately when
>> the prot changed? If we should map it immediately, who(kernel, glibc, user)
>> do this step?
> 
> This is still very fuzzy. What are you actually trying to achieve?

I don't expect page fault any more after mlock.

[toc] | [prev] | [next] | [standalone]


#1739737

FromMichal Hocko <mhocko@kernel.org>
Date2017-09-26 11:30 +0200
Message-ID<utUlH-2Sr-13@gated-at.bofh.it>
In reply to#1739735
On Tue 26-09-17 17:13:59, Xishi Qiu wrote:
> On 2017/9/26 17:02, Michal Hocko wrote:
[...]
> > This is still very fuzzy. What are you actually trying to achieve?
> 
> I don't expect page fault any more after mlock.

This should be the case normally. Except when mm_populate fails which
can happen e.g. when running inside a memcg with the hard limit
configured. Is there any other unexpected failure scenario you are
seeing?

-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1739738

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-09-26 11:30 +0200
Message-ID<utUlI-2Sr-17@gated-at.bofh.it>
In reply to#1739735
On 2017/9/26 17:13, Xishi Qiu wrote:

> On 2017/9/26 17:02, Michal Hocko wrote:
> 
>> On Tue 26-09-17 16:39:56, Xishi Qiu wrote:
>>> On 2017/9/26 16:17, Michal Hocko wrote:
>>>
>>>> On Tue 26-09-17 15:56:55, Xishi Qiu wrote:
>>>>> When we call mlockall(), we will add VM_LOCKED to the vma,
>>>>> if the vma prot is ---p,
>>>>
>>>> not sure what you mean here. apply_mlockall_flags will set the flag on
>>>> all vmas except for special mappings (mlock_fixup). This phase will
>>>> cause that memory reclaim will not free already mapped pages in those
>>>> vmas (see page_check_references and the lazy mlock pages move to
>>>> unevictable LRUs).
>>>>
>>>>> then mm_populate -> get_user_pages will not alloc memory.
>>>>
>>>> mm_populate all the vmas with pages. Well there are certainly some
>>>> constrains - e.g. memory cgroup hard limit might be hit and so the
>>>> faulting might fail.
>>>>
>>>>> I find it said "ignore errors" in mm_populate()
>>>>> static inline void mm_populate(unsigned long addr, unsigned long len)
>>>>> {
>>>>> 	/* Ignore errors */
>>>>> 	(void) __mm_populate(addr, len, 1);
>>>>> }
>>>>
>>>> But we do not report the failure because any failure past
>>>> apply_mlockall_flags would be tricky to handle. We have already dropped
>>>> the mmap_sem lock so some other address space operations could have
>>>> interfered.
>>>>  
>>>>> And later we call mprotect() to change the prot, then it is
>>>>> still not alloc memory for the mlocked vma.
>>>>>
>>>>> My question is that, shall we alloc memory if the prot changed,
>>>>> and who(kernel, glibc, user) should alloc the memory?
>>>>
>>>> I do not understand your question but if you are asking how to get pages
>>>> to map your vmas then touching that area will fault the memory in.
>>>
>>> Hi Michal,
>>>
>>> syscall mlockall() will first apply the VM_LOCKED to the vma, then
>>> call mm_populate() to map the vmas.
>>>
>>> mm_populate
>>> 	populate_vma_page_range
>>> 		__get_user_pages
>>> 			check_vma_flags
>>> And the above path maybe return -EFAULT in some case, right?
>>>
>>> If we call mprotect() to change the prot of vma, just let
>>> check_vma_flags() return 0, then we will get the mlocked pages
>>> in following page-fault, right?
>>
>> Any future page fault to the existing vma will result in the mlocked
>> page. That is what VM_LOCKED guarantess.
>>
>>> My question is that, shall we map the vmas immediately when
>>> the prot changed? If we should map it immediately, who(kernel, glibc, user)
>>> do this step?
>>
>> This is still very fuzzy. What are you actually trying to achieve?
> 
> I don't expect page fault any more after mlock.
> 

Our apps is some thing like RT, and page-fault maybe cause a lot of time,
e.g. lock, mem reclaim ..., so I use mlock and don't want page fault
any more.

Thanks,
Xishi Qiu

> 
> .
> 

[toc] | [prev] | [next] | [standalone]


#1739745

FromVlastimil Babka <vbabka@suse.cz>
Date2017-09-26 11:50 +0200
Message-ID<utUF4-2ZV-11@gated-at.bofh.it>
In reply to#1739738
On 09/26/2017 11:22 AM, Xishi Qiu wrote:
> On 2017/9/26 17:13, Xishi Qiu wrote:
>>> This is still very fuzzy. What are you actually trying to achieve?
>>
>> I don't expect page fault any more after mlock.
>>
> 
> Our apps is some thing like RT, and page-fault maybe cause a lot of time,
> e.g. lock, mem reclaim ..., so I use mlock and don't want page fault
> any more.

Why does your app then have restricted mprotect when calling mlockall()
and only later adjusts the mprotect?

Vlastimil

[toc] | [prev] | [next] | [standalone]


#1739780

FromMichal Hocko <mhocko@kernel.org>
Date2017-09-26 13:10 +0200
Message-ID<utVUu-3YA-13@gated-at.bofh.it>
In reply to#1739745
On Tue 26-09-17 11:45:16, Vlastimil Babka wrote:
> On 09/26/2017 11:22 AM, Xishi Qiu wrote:
> > On 2017/9/26 17:13, Xishi Qiu wrote:
> >>> This is still very fuzzy. What are you actually trying to achieve?
> >>
> >> I don't expect page fault any more after mlock.
> >>
> > 
> > Our apps is some thing like RT, and page-fault maybe cause a lot of time,
> > e.g. lock, mem reclaim ..., so I use mlock and don't want page fault
> > any more.
> 
> Why does your app then have restricted mprotect when calling mlockall()
> and only later adjusts the mprotect?

Ahh, OK I see what is goging on. So you have PROT_NONE vma at the time
mlockall and then later mprotect it something else and want to fault all
that memory at the mprotect time?

So basically to do
---
diff --git a/mm/mprotect.c b/mm/mprotect.c
index 6d3e2f082290..b665b5d1c544 100644
--- a/mm/mprotect.c
+++ b/mm/mprotect.c
@@ -369,7 +369,7 @@ mprotect_fixup(struct vm_area_struct *vma, struct vm_area_struct **pprev,
 	 * Private VM_LOCKED VMA becoming writable: trigger COW to avoid major
 	 * fault on access.
 	 */
-	if ((oldflags & (VM_WRITE | VM_SHARED | VM_LOCKED)) == VM_LOCKED &&
+	if ((oldflags & (VM_WRITE | VM_LOCKED)) == VM_LOCKED &&
 			(newflags & VM_WRITE)) {
 		populate_vma_page_range(vma, start, end, NULL);
 	}

-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1740400

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-09-27 08:00 +0200
Message-ID<uudy1-6E8-3@gated-at.bofh.it>
In reply to#1739780
On 2017/9/26 19:00, Michal Hocko wrote:

> On Tue 26-09-17 11:45:16, Vlastimil Babka wrote:
>> On 09/26/2017 11:22 AM, Xishi Qiu wrote:
>>> On 2017/9/26 17:13, Xishi Qiu wrote:
>>>>> This is still very fuzzy. What are you actually trying to achieve?
>>>>
>>>> I don't expect page fault any more after mlock.
>>>>
>>>
>>> Our apps is some thing like RT, and page-fault maybe cause a lot of time,
>>> e.g. lock, mem reclaim ..., so I use mlock and don't want page fault
>>> any more.
>>
>> Why does your app then have restricted mprotect when calling mlockall()
>> and only later adjusts the mprotect?
> 
> Ahh, OK I see what is goging on. So you have PROT_NONE vma at the time
> mlockall and then later mprotect it something else and want to fault all
> that memory at the mprotect time?
> 
> So basically to do
> ---
> diff --git a/mm/mprotect.c b/mm/mprotect.c
> index 6d3e2f082290..b665b5d1c544 100644
> --- a/mm/mprotect.c
> +++ b/mm/mprotect.c
> @@ -369,7 +369,7 @@ mprotect_fixup(struct vm_area_struct *vma, struct vm_area_struct **pprev,
>  	 * Private VM_LOCKED VMA becoming writable: trigger COW to avoid major
>  	 * fault on access.
>  	 */
> -	if ((oldflags & (VM_WRITE | VM_SHARED | VM_LOCKED)) == VM_LOCKED &&
> +	if ((oldflags & (VM_WRITE | VM_LOCKED)) == VM_LOCKED &&
>  			(newflags & VM_WRITE)) {
>  		populate_vma_page_range(vma, start, end, NULL);
>  	}
> 

Hi Michal,

My kernel is v3.10, and I missed this code, thank you reminding me.

Thanks,
Xishi Qiu

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web