Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1424208 > unrolled thread
| Started by | Michal Hocko <mhocko@kernel.org> |
|---|---|
| First post | 2016-06-16 17:50 +0200 |
| Last post | 2016-06-17 13:20 +0200 |
| Articles | 11 — 4 participants |
Back to article view | Back to linux.kernel
This discussion starts older than the indexed window; earlier articles aren't shown. The article labeled Started by
below is the oldest one visible, not the original post.
Re: [PATCH] mm: fix account pmd page to the process Michal Hocko <mhocko@kernel.org> - 2016-06-16 17:50 +0200
Re: [PATCH] mm: fix account pmd page to the process Michal Hocko <mhocko@kernel.org> - 2016-06-16 17:50 +0200
Re: [PATCH] mm: fix account pmd page to the process Mike Kravetz <mike.kravetz@oracle.com> - 2016-06-16 18:10 +0200
Re: [PATCH] mm: fix account pmd page to the process Michal Hocko <mhocko@kernel.org> - 2016-06-16 18:40 +0200
Re: [PATCH] mm: fix account pmd page to the process Mike Kravetz <mike.kravetz@oracle.com> - 2016-06-16 18:50 +0200
Re: [PATCH] mm: fix account pmd page to the process "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-06-17 14:30 +0200
Re: [PATCH] mm: fix account pmd page to the process Michal Hocko <mhocko@kernel.org> - 2016-06-17 15:10 +0200
Re: [PATCH] mm: fix account pmd page to the process "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-06-17 16:30 +0200
Re: [PATCH] mm: fix account pmd page to the process Mike Kravetz <mike.kravetz@oracle.com> - 2016-06-17 17:40 +0200
Re: [PATCH] mm: fix account pmd page to the process zhong jiang <zhongjiang@huawei.com> - 2016-06-18 07:10 +0200
Re: [PATCH] mm: fix account pmd page to the process zhong jiang <zhongjiang@huawei.com> - 2016-06-17 13:20 +0200
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-06-16 17:50 +0200 |
| Subject | Re: [PATCH] mm: fix account pmd page to the process |
| Message-ID | <rKHIl-1K6-11@gated-at.bofh.it> |
On Thu 16-06-16 19:36:11, zhongjiang wrote:
> From: zhong jiang <zhongjiang@huawei.com>
>
> when a process acquire a pmd table shared by other process, we
> increase the account to current process. otherwise, a race result
> in other tasks have set the pud entry. so it no need to increase it.
>
> Signed-off-by: zhong jiang <zhongjiang@huawei.com>
> ---
> mm/hugetlb.c | 5 ++---
> 1 file changed, 2 insertions(+), 3 deletions(-)
>
> diff --git a/mm/hugetlb.c b/mm/hugetlb.c
> index 19d0d08..3b025c5 100644
> --- a/mm/hugetlb.c
> +++ b/mm/hugetlb.c
> @@ -4189,10 +4189,9 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
> if (pud_none(*pud)) {
> pud_populate(mm, pud,
> (pmd_t *)((unsigned long)spte & PAGE_MASK));
> - } else {
> + } else
> put_page(virt_to_page(spte));
> - mm_inc_nr_pmds(mm);
> - }
The code is quite puzzling but is this correct? Shouldn't we rather do
mm_dec_nr_pmds(mm) in that path to undo the previous inc?
> +
> spin_unlock(ptl);
> out:
> pte = (pte_t *)pmd_alloc(mm, pud, addr);
> --
> 1.8.3.1
--
Michal Hocko
SUSE Labs
[toc] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-06-16 17:50 +0200 |
| Message-ID | <rKHIm-1K6-35@gated-at.bofh.it> |
| In reply to | #1424208 |
[It seems that this patch has been sent several times and this
particular copy didn't add Kirill who has added this code CC him now]
On Thu 16-06-16 17:42:14, Michal Hocko wrote:
> On Thu 16-06-16 19:36:11, zhongjiang wrote:
> > From: zhong jiang <zhongjiang@huawei.com>
> >
> > when a process acquire a pmd table shared by other process, we
> > increase the account to current process. otherwise, a race result
> > in other tasks have set the pud entry. so it no need to increase it.
> >
> > Signed-off-by: zhong jiang <zhongjiang@huawei.com>
> > ---
> > mm/hugetlb.c | 5 ++---
> > 1 file changed, 2 insertions(+), 3 deletions(-)
> >
> > diff --git a/mm/hugetlb.c b/mm/hugetlb.c
> > index 19d0d08..3b025c5 100644
> > --- a/mm/hugetlb.c
> > +++ b/mm/hugetlb.c
> > @@ -4189,10 +4189,9 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
> > if (pud_none(*pud)) {
> > pud_populate(mm, pud,
> > (pmd_t *)((unsigned long)spte & PAGE_MASK));
> > - } else {
> > + } else
> > put_page(virt_to_page(spte));
> > - mm_inc_nr_pmds(mm);
> > - }
>
> The code is quite puzzling but is this correct? Shouldn't we rather do
> mm_dec_nr_pmds(mm) in that path to undo the previous inc?
>
> > +
> > spin_unlock(ptl);
> > out:
> > pte = (pte_t *)pmd_alloc(mm, pud, addr);
> > --
> > 1.8.3.1
>
> --
> Michal Hocko
> SUSE Labs
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Mike Kravetz <mike.kravetz@oracle.com> |
|---|---|
| Date | 2016-06-16 18:10 +0200 |
| Message-ID | <rKI1H-25D-19@gated-at.bofh.it> |
| In reply to | #1424215 |
On 06/16/2016 08:43 AM, Michal Hocko wrote:
> [It seems that this patch has been sent several times and this
> particular copy didn't add Kirill who has added this code CC him now]
>
> On Thu 16-06-16 17:42:14, Michal Hocko wrote:
>> On Thu 16-06-16 19:36:11, zhongjiang wrote:
>>> From: zhong jiang <zhongjiang@huawei.com>
>>>
>>> when a process acquire a pmd table shared by other process, we
>>> increase the account to current process. otherwise, a race result
>>> in other tasks have set the pud entry. so it no need to increase it.
>>>
>>> Signed-off-by: zhong jiang <zhongjiang@huawei.com>
>>> ---
>>> mm/hugetlb.c | 5 ++---
>>> 1 file changed, 2 insertions(+), 3 deletions(-)
>>>
>>> diff --git a/mm/hugetlb.c b/mm/hugetlb.c
>>> index 19d0d08..3b025c5 100644
>>> --- a/mm/hugetlb.c
>>> +++ b/mm/hugetlb.c
>>> @@ -4189,10 +4189,9 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
>>> if (pud_none(*pud)) {
>>> pud_populate(mm, pud,
>>> (pmd_t *)((unsigned long)spte & PAGE_MASK));
>>> - } else {
>>> + } else
>>> put_page(virt_to_page(spte));
>>> - mm_inc_nr_pmds(mm);
>>> - }
>>
>> The code is quite puzzling but is this correct? Shouldn't we rather do
>> mm_dec_nr_pmds(mm) in that path to undo the previous inc?
I agree that the code is quite puzzling. :(
However, if this were an issue I would have expected to see some reports.
Oracle DB makes use of this feature (shared page tables) and if the pmd
count is wrong we would catch it in check_mm() at exit time.
Upon closer examination, I believe the code in question is never executed.
Note the callers of huge_pmd_share. The calling code looks like:
if (want_pmd_share() && pud_none(*pud))
pte = huge_pmd_share(mm, addr, pud);
else
pte = (pte_t *)pmd_alloc(mm, pud, addr);
Therefore, we do not call huge_pmd_share unless pud_none(*pud). The
code in question is only executed when !pud_none(*pud).
I think that entire if/else statement can be removed. We know
pud_none(*pud), so just do pud_populate().
--
Mike Kravetz
>>
>>> +
>>> spin_unlock(ptl);
>>> out:
>>> pte = (pte_t *)pmd_alloc(mm, pud, addr);
>>> --
>>> 1.8.3.1
>>
>> --
>> Michal Hocko
>> SUSE Labs
>
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-06-16 18:40 +0200 |
| Message-ID | <rKIuK-2fo-21@gated-at.bofh.it> |
| In reply to | #1424229 |
On Thu 16-06-16 09:05:23, Mike Kravetz wrote:
> On 06/16/2016 08:43 AM, Michal Hocko wrote:
> > [It seems that this patch has been sent several times and this
> > particular copy didn't add Kirill who has added this code CC him now]
> >
> > On Thu 16-06-16 17:42:14, Michal Hocko wrote:
> >> On Thu 16-06-16 19:36:11, zhongjiang wrote:
> >>> From: zhong jiang <zhongjiang@huawei.com>
> >>>
> >>> when a process acquire a pmd table shared by other process, we
> >>> increase the account to current process. otherwise, a race result
> >>> in other tasks have set the pud entry. so it no need to increase it.
> >>>
> >>> Signed-off-by: zhong jiang <zhongjiang@huawei.com>
> >>> ---
> >>> mm/hugetlb.c | 5 ++---
> >>> 1 file changed, 2 insertions(+), 3 deletions(-)
> >>>
> >>> diff --git a/mm/hugetlb.c b/mm/hugetlb.c
> >>> index 19d0d08..3b025c5 100644
> >>> --- a/mm/hugetlb.c
> >>> +++ b/mm/hugetlb.c
> >>> @@ -4189,10 +4189,9 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
> >>> if (pud_none(*pud)) {
> >>> pud_populate(mm, pud,
> >>> (pmd_t *)((unsigned long)spte & PAGE_MASK));
> >>> - } else {
> >>> + } else
> >>> put_page(virt_to_page(spte));
> >>> - mm_inc_nr_pmds(mm);
> >>> - }
> >>
> >> The code is quite puzzling but is this correct? Shouldn't we rather do
> >> mm_dec_nr_pmds(mm) in that path to undo the previous inc?
>
> I agree that the code is quite puzzling. :(
>
> However, if this were an issue I would have expected to see some reports.
> Oracle DB makes use of this feature (shared page tables) and if the pmd
> count is wrong we would catch it in check_mm() at exit time.
>
> Upon closer examination, I believe the code in question is never executed.
> Note the callers of huge_pmd_share. The calling code looks like:
>
> if (want_pmd_share() && pud_none(*pud))
> pte = huge_pmd_share(mm, addr, pud);
> else
> pte = (pte_t *)pmd_alloc(mm, pud, addr);
>
> Therefore, we do not call huge_pmd_share unless pud_none(*pud). The
> code in question is only executed when !pud_none(*pud).
My understanding is that the check is needed after we retake page lock
because we might have raced with other thread. But it's been quite some
time since I've looked at hugetlb locking and page table sharing code.
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Mike Kravetz <mike.kravetz@oracle.com> |
|---|---|
| Date | 2016-06-16 18:50 +0200 |
| Message-ID | <rKIEp-2jH-27@gated-at.bofh.it> |
| In reply to | #1424256 |
On 06/16/2016 09:31 AM, Michal Hocko wrote:
> On Thu 16-06-16 09:05:23, Mike Kravetz wrote:
>> On 06/16/2016 08:43 AM, Michal Hocko wrote:
>>> [It seems that this patch has been sent several times and this
>>> particular copy didn't add Kirill who has added this code CC him now]
>>>
>>> On Thu 16-06-16 17:42:14, Michal Hocko wrote:
>>>> On Thu 16-06-16 19:36:11, zhongjiang wrote:
>>>>> From: zhong jiang <zhongjiang@huawei.com>
>>>>>
>>>>> when a process acquire a pmd table shared by other process, we
>>>>> increase the account to current process. otherwise, a race result
>>>>> in other tasks have set the pud entry. so it no need to increase it.
>>>>>
>>>>> Signed-off-by: zhong jiang <zhongjiang@huawei.com>
>>>>> ---
>>>>> mm/hugetlb.c | 5 ++---
>>>>> 1 file changed, 2 insertions(+), 3 deletions(-)
>>>>>
>>>>> diff --git a/mm/hugetlb.c b/mm/hugetlb.c
>>>>> index 19d0d08..3b025c5 100644
>>>>> --- a/mm/hugetlb.c
>>>>> +++ b/mm/hugetlb.c
>>>>> @@ -4189,10 +4189,9 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
>>>>> if (pud_none(*pud)) {
>>>>> pud_populate(mm, pud,
>>>>> (pmd_t *)((unsigned long)spte & PAGE_MASK));
>>>>> - } else {
>>>>> + } else
>>>>> put_page(virt_to_page(spte));
>>>>> - mm_inc_nr_pmds(mm);
>>>>> - }
>>>>
>>>> The code is quite puzzling but is this correct? Shouldn't we rather do
>>>> mm_dec_nr_pmds(mm) in that path to undo the previous inc?
>>
>> I agree that the code is quite puzzling. :(
>>
>> However, if this were an issue I would have expected to see some reports.
>> Oracle DB makes use of this feature (shared page tables) and if the pmd
>> count is wrong we would catch it in check_mm() at exit time.
>>
>> Upon closer examination, I believe the code in question is never executed.
>> Note the callers of huge_pmd_share. The calling code looks like:
>>
>> if (want_pmd_share() && pud_none(*pud))
>> pte = huge_pmd_share(mm, addr, pud);
>> else
>> pte = (pte_t *)pmd_alloc(mm, pud, addr);
>>
>> Therefore, we do not call huge_pmd_share unless pud_none(*pud). The
>> code in question is only executed when !pud_none(*pud).
>
> My understanding is that the check is needed after we retake page lock
> because we might have raced with other thread. But it's been quite some
> time since I've looked at hugetlb locking and page table sharing code.
That is correct, we could have raced. Duh!
In the case of a race, the other thread would have incremented the
PMD count already. Your suggestion of decrementing pmd count in
this case seems to be the correct approach. But, I need to think
about this some more.
--
Mike Kravetz
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2016-06-17 14:30 +0200 |
| Message-ID | <rL14m-6sY-21@gated-at.bofh.it> |
| In reply to | #1424269 |
On Thu, Jun 16, 2016 at 09:47:46AM -0700, Mike Kravetz wrote:
> On 06/16/2016 09:31 AM, Michal Hocko wrote:
> > On Thu 16-06-16 09:05:23, Mike Kravetz wrote:
> >> On 06/16/2016 08:43 AM, Michal Hocko wrote:
> >>> [It seems that this patch has been sent several times and this
> >>> particular copy didn't add Kirill who has added this code CC him now]
> >>>
> >>> On Thu 16-06-16 17:42:14, Michal Hocko wrote:
> >>>> On Thu 16-06-16 19:36:11, zhongjiang wrote:
> >>>>> From: zhong jiang <zhongjiang@huawei.com>
> >>>>>
> >>>>> when a process acquire a pmd table shared by other process, we
> >>>>> increase the account to current process. otherwise, a race result
> >>>>> in other tasks have set the pud entry. so it no need to increase it.
> >>>>>
> >>>>> Signed-off-by: zhong jiang <zhongjiang@huawei.com>
> >>>>> ---
> >>>>> mm/hugetlb.c | 5 ++---
> >>>>> 1 file changed, 2 insertions(+), 3 deletions(-)
> >>>>>
> >>>>> diff --git a/mm/hugetlb.c b/mm/hugetlb.c
> >>>>> index 19d0d08..3b025c5 100644
> >>>>> --- a/mm/hugetlb.c
> >>>>> +++ b/mm/hugetlb.c
> >>>>> @@ -4189,10 +4189,9 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
> >>>>> if (pud_none(*pud)) {
> >>>>> pud_populate(mm, pud,
> >>>>> (pmd_t *)((unsigned long)spte & PAGE_MASK));
> >>>>> - } else {
> >>>>> + } else
> >>>>> put_page(virt_to_page(spte));
> >>>>> - mm_inc_nr_pmds(mm);
> >>>>> - }
> >>>>
> >>>> The code is quite puzzling but is this correct? Shouldn't we rather do
> >>>> mm_dec_nr_pmds(mm) in that path to undo the previous inc?
> >>
> >> I agree that the code is quite puzzling. :(
> >>
> >> However, if this were an issue I would have expected to see some reports.
> >> Oracle DB makes use of this feature (shared page tables) and if the pmd
> >> count is wrong we would catch it in check_mm() at exit time.
> >>
> >> Upon closer examination, I believe the code in question is never executed.
> >> Note the callers of huge_pmd_share. The calling code looks like:
> >>
> >> if (want_pmd_share() && pud_none(*pud))
> >> pte = huge_pmd_share(mm, addr, pud);
> >> else
> >> pte = (pte_t *)pmd_alloc(mm, pud, addr);
> >>
> >> Therefore, we do not call huge_pmd_share unless pud_none(*pud). The
> >> code in question is only executed when !pud_none(*pud).
> >
> > My understanding is that the check is needed after we retake page lock
> > because we might have raced with other thread. But it's been quite some
> > time since I've looked at hugetlb locking and page table sharing code.
>
> That is correct, we could have raced. Duh!
>
> In the case of a race, the other thread would have incremented the
> PMD count already. Your suggestion of decrementing pmd count in
> this case seems to be the correct approach. But, I need to think
> about this some more.
Yes, I made mistake by increasing nr_pmds, not descreasing here.
Testcase:
#include <errno.h>
#include <stdio.h>
#include <stdint.h>
#include <stdlib.h>
#include <unistd.h>
#include <sys/mman.h>
#include <sys/syscall.h>
#include <sys/time.h>
#define HPGSZ 2097152UL
int main(int argc, char **argv) {
char *addr;
system("echo 1024 > /proc/sys/vm/nr_hugepages");
addr = mmap(NULL, 1024*HPGSZ, PROT_WRITE | PROT_READ,
MAP_SHARED | MAP_ANONYMOUS | MAP_HUGETLB | MAP_POPULATE, -1, 0);
if (addr == MAP_FAILED) {
fprintf(stderr, "Failed to alloc hugepage\n");
return -1;
}
addr[0] = 1;
fork();
printf("addr[0]: %d\n", addr[0]);
sleep(1);
return 0;
}
You can simulate race by replacing 'if (pud_none(*pud))' with "if (0)". It
would produce "BUG: non-zero nr_pmds on freeing mm: 2" on the test-case.
Fix:
From fd22922e7b4664e83653a84331f0a95b985bff0c Mon Sep 17 00:00:00 2001
From: "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com>
Date: Fri, 17 Jun 2016 15:07:03 +0300
Subject: [PATCH] hugetlb: fix nr_pmds accounting with shared page tables
We account HugeTLB's shared page table to all processes who share it.
The accounting happens during huge_pmd_share().
If somebody populates pud entry under us, we should decrease pagetable's
refcount and decrease nr_pmds of the process.
By mistake, I increase nr_pmds again in this case. :-/
It will lead to "BUG: non-zero nr_pmds on freeing mm: 2" on process'
exit.
Let's fix this by increasing nr_pmds only when we're sure that the page
table will be used.
Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Reported-by: zhongjiang <zhongjiang@huawei.com>
Fixes: dc6c9a35b66b ("mm: account pmd page tables to the process")
Cc: <stable@vger.kernel.org> [4.0+]
---
mm/hugetlb.c | 3 +--
1 file changed, 1 insertion(+), 2 deletions(-)
diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index e197cd7080e6..ed6a537f0878 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -4216,7 +4216,6 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
if (saddr) {
spte = huge_pte_offset(svma->vm_mm, saddr);
if (spte) {
- mm_inc_nr_pmds(mm);
get_page(virt_to_page(spte));
break;
}
@@ -4231,9 +4230,9 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
if (pud_none(*pud)) {
pud_populate(mm, pud,
(pmd_t *)((unsigned long)spte & PAGE_MASK));
+ mm_inc_nr_pmds(mm);
} else {
put_page(virt_to_page(spte));
- mm_inc_nr_pmds(mm);
}
spin_unlock(ptl);
out:
--
Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Michal Hocko <mhocko@kernel.org> |
|---|---|
| Date | 2016-06-17 15:10 +0200 |
| Message-ID | <rL1H4-6Xf-17@gated-at.bofh.it> |
| In reply to | #1425008 |
On Fri 17-06-16 15:25:06, Kirill A. Shutemov wrote:
[...]
> >From fd22922e7b4664e83653a84331f0a95b985bff0c Mon Sep 17 00:00:00 2001
> From: "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com>
> Date: Fri, 17 Jun 2016 15:07:03 +0300
> Subject: [PATCH] hugetlb: fix nr_pmds accounting with shared page tables
>
> We account HugeTLB's shared page table to all processes who share it.
> The accounting happens during huge_pmd_share().
>
> If somebody populates pud entry under us, we should decrease pagetable's
> refcount and decrease nr_pmds of the process.
>
> By mistake, I increase nr_pmds again in this case. :-/
> It will lead to "BUG: non-zero nr_pmds on freeing mm: 2" on process'
> exit.
>
> Let's fix this by increasing nr_pmds only when we're sure that the page
> table will be used.
>
> Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
> Reported-by: zhongjiang <zhongjiang@huawei.com>
> Fixes: dc6c9a35b66b ("mm: account pmd page tables to the process")
> Cc: <stable@vger.kernel.org> [4.0+]
Yes this patch is better. Is it worth backporting to stable though?
BUG message sounds scary but it is not a real BUG().
Acked-by: Michal Hocko <mhocko@suse.com>
> ---
> mm/hugetlb.c | 3 +--
> 1 file changed, 1 insertion(+), 2 deletions(-)
>
> diff --git a/mm/hugetlb.c b/mm/hugetlb.c
> index e197cd7080e6..ed6a537f0878 100644
> --- a/mm/hugetlb.c
> +++ b/mm/hugetlb.c
> @@ -4216,7 +4216,6 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
> if (saddr) {
> spte = huge_pte_offset(svma->vm_mm, saddr);
> if (spte) {
> - mm_inc_nr_pmds(mm);
> get_page(virt_to_page(spte));
> break;
> }
> @@ -4231,9 +4230,9 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
> if (pud_none(*pud)) {
> pud_populate(mm, pud,
> (pmd_t *)((unsigned long)spte & PAGE_MASK));
> + mm_inc_nr_pmds(mm);
> } else {
> put_page(virt_to_page(spte));
> - mm_inc_nr_pmds(mm);
> }
> spin_unlock(ptl);
> out:
> --
> Kirill A. Shutemov
--
Michal Hocko
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2016-06-17 16:30 +0200 |
| Message-ID | <rL2Wt-7ET-1@gated-at.bofh.it> |
| In reply to | #1425055 |
On Fri, Jun 17, 2016 at 03:00:00PM +0200, Michal Hocko wrote:
> On Fri 17-06-16 15:25:06, Kirill A. Shutemov wrote:
> [...]
> > >From fd22922e7b4664e83653a84331f0a95b985bff0c Mon Sep 17 00:00:00 2001
> > From: "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com>
> > Date: Fri, 17 Jun 2016 15:07:03 +0300
> > Subject: [PATCH] hugetlb: fix nr_pmds accounting with shared page tables
> >
> > We account HugeTLB's shared page table to all processes who share it.
> > The accounting happens during huge_pmd_share().
> >
> > If somebody populates pud entry under us, we should decrease pagetable's
> > refcount and decrease nr_pmds of the process.
> >
> > By mistake, I increase nr_pmds again in this case. :-/
> > It will lead to "BUG: non-zero nr_pmds on freeing mm: 2" on process'
> > exit.
> >
> > Let's fix this by increasing nr_pmds only when we're sure that the page
> > table will be used.
> >
> > Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
> > Reported-by: zhongjiang <zhongjiang@huawei.com>
> > Fixes: dc6c9a35b66b ("mm: account pmd page tables to the process")
> > Cc: <stable@vger.kernel.org> [4.0+]
>
> Yes this patch is better. Is it worth backporting to stable though?
> BUG message sounds scary but it is not a real BUG().
I guess, we can live without stable backport.
>
> Acked-by: Michal Hocko <mhocko@suse.com>
Thanks.
--
Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Mike Kravetz <mike.kravetz@oracle.com> |
|---|---|
| Date | 2016-06-17 17:40 +0200 |
| Message-ID | <rL42d-8iY-9@gated-at.bofh.it> |
| In reply to | #1425008 |
On 06/17/2016 05:25 AM, Kirill A. Shutemov wrote:
>
> From fd22922e7b4664e83653a84331f0a95b985bff0c Mon Sep 17 00:00:00 2001
> From: "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com>
> Date: Fri, 17 Jun 2016 15:07:03 +0300
> Subject: [PATCH] hugetlb: fix nr_pmds accounting with shared page tables
>
> We account HugeTLB's shared page table to all processes who share it.
> The accounting happens during huge_pmd_share().
>
> If somebody populates pud entry under us, we should decrease pagetable's
> refcount and decrease nr_pmds of the process.
>
> By mistake, I increase nr_pmds again in this case. :-/
> It will lead to "BUG: non-zero nr_pmds on freeing mm: 2" on process'
> exit.
>
> Let's fix this by increasing nr_pmds only when we're sure that the page
> table will be used.
>
> Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Nice,
Reviewed-by: Mike Kravetz <mike.kravetz@oracle.com>
I agree that we do not necessarily need a back port. I have not seen
reports of people experiencing this race and seeing the BUG (on mm
tear-down).
zhongjiang, did someone actually hit the BUG? Or, did you find it by
code examination?
--
Mike Kravetz
> Reported-by: zhongjiang <zhongjiang@huawei.com>
> Fixes: dc6c9a35b66b ("mm: account pmd page tables to the process")
> Cc: <stable@vger.kernel.org> [4.0+]
> ---
> mm/hugetlb.c | 3 +--
> 1 file changed, 1 insertion(+), 2 deletions(-)
>
> diff --git a/mm/hugetlb.c b/mm/hugetlb.c
> index e197cd7080e6..ed6a537f0878 100644
> --- a/mm/hugetlb.c
> +++ b/mm/hugetlb.c
> @@ -4216,7 +4216,6 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
> if (saddr) {
> spte = huge_pte_offset(svma->vm_mm, saddr);
> if (spte) {
> - mm_inc_nr_pmds(mm);
> get_page(virt_to_page(spte));
> break;
> }
> @@ -4231,9 +4230,9 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
> if (pud_none(*pud)) {
> pud_populate(mm, pud,
> (pmd_t *)((unsigned long)spte & PAGE_MASK));
> + mm_inc_nr_pmds(mm);
> } else {
> put_page(virt_to_page(spte));
> - mm_inc_nr_pmds(mm);
> }
> spin_unlock(ptl);
> out:
>
[toc] | [prev] | [next] | [standalone]
| From | zhong jiang <zhongjiang@huawei.com> |
|---|---|
| Date | 2016-06-18 07:10 +0200 |
| Message-ID | <rLgG5-890-5@gated-at.bofh.it> |
| In reply to | #1425231 |
On 2016/6/17 23:39, Mike Kravetz wrote: > On 06/17/2016 05:25 AM, Kirill A. Shutemov wrote: >> From fd22922e7b4664e83653a84331f0a95b985bff0c Mon Sep 17 00:00:00 2001 >> From: "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> >> Date: Fri, 17 Jun 2016 15:07:03 +0300 >> Subject: [PATCH] hugetlb: fix nr_pmds accounting with shared page tables >> >> We account HugeTLB's shared page table to all processes who share it. >> The accounting happens during huge_pmd_share(). >> >> If somebody populates pud entry under us, we should decrease pagetable's >> refcount and decrease nr_pmds of the process. >> >> By mistake, I increase nr_pmds again in this case. :-/ >> It will lead to "BUG: non-zero nr_pmds on freeing mm: 2" on process' >> exit. >> >> Let's fix this by increasing nr_pmds only when we're sure that the page >> table will be used. >> >> Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com> > Nice, > Reviewed-by: Mike Kravetz <mike.kravetz@oracle.com> > > I agree that we do not necessarily need a back port. I have not seen > reports of people experiencing this race and seeing the BUG (on mm > tear-down). > > zhongjiang, did someone actually hit the BUG? Or, did you find it by > code examination? > just code examination.
[toc] | [prev] | [next] | [standalone]
| From | zhong jiang <zhongjiang@huawei.com> |
|---|---|
| Date | 2016-06-17 13:20 +0200 |
| Message-ID | <rKZYB-5P6-7@gated-at.bofh.it> |
| In reply to | #1424208 |
On 2016/6/16 23:42, Michal Hocko wrote:
> On Thu 16-06-16 19:36:11, zhongjiang wrote:
>> From: zhong jiang <zhongjiang@huawei.com>
>>
>> when a process acquire a pmd table shared by other process, we
>> increase the account to current process. otherwise, a race result
>> in other tasks have set the pud entry. so it no need to increase it.
>>
>> Signed-off-by: zhong jiang <zhongjiang@huawei.com>
>> ---
>> mm/hugetlb.c | 5 ++---
>> 1 file changed, 2 insertions(+), 3 deletions(-)
>>
>> diff --git a/mm/hugetlb.c b/mm/hugetlb.c
>> index 19d0d08..3b025c5 100644
>> --- a/mm/hugetlb.c
>> +++ b/mm/hugetlb.c
>> @@ -4189,10 +4189,9 @@ pte_t *huge_pmd_share(struct mm_struct *mm, unsigned long addr, pud_t *pud)
>> if (pud_none(*pud)) {
>> pud_populate(mm, pud,
>> (pmd_t *)((unsigned long)spte & PAGE_MASK));
>> - } else {
>> + } else
>> put_page(virt_to_page(spte));
>> - mm_inc_nr_pmds(mm);
>> - }
> The code is quite puzzling but is this correct? Shouldn't we rather do
> mm_dec_nr_pmds(mm) in that path to undo the previous inc?
Yes, you are right. I will modify it in V2.
Thanks
zhongjiang
>
>> +
>> spin_unlock(ptl);
>> out:
>> pte = (pte_t *)pmd_alloc(mm, pud, addr);
>> --
>> 1.8.3.1
[toc] | [prev] | [standalone]
Back to top | Article view | linux.kernel
csiph-web