Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1332281 > unrolled thread
| Started by | Gerald Schaefer <gerald.schaefer@de.ibm.com> |
|---|---|
| First post | 2016-02-11 19:30 +0100 |
| Last post | 2016-02-15 17:50 +0100 |
| Articles | 20 on this page of 30 — 7 participants |
Back to article view | Back to linux.kernel
[BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Gerald Schaefer <gerald.schaefer@de.ibm.com> - 2016-02-11 19:30 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-02-11 20:10 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2016-02-11 20:20 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Sebastian Ott <sebott@linux.vnet.ibm.com> - 2016-02-12 13:30 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Gerald Schaefer <gerald.schaefer@de.ibm.com> - 2016-02-11 21:00 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-02-12 05:10 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Gerald Schaefer <gerald.schaefer@de.ibm.com> - 2016-02-12 13:10 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> - 2016-02-12 17:20 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Will Deacon <will.deacon@arm.com> - 2016-02-12 11:10 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Sebastian Ott <sebott@linux.vnet.ibm.com> - 2016-02-12 11:20 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Will Deacon <will.deacon@arm.com> - 2016-02-12 17:00 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-02-12 16:50 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Christian Borntraeger <borntraeger@de.ibm.com> - 2016-02-12 17:00 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Gerald Schaefer <gerald.schaefer@de.ibm.com> - 2016-02-12 18:20 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-02-13 00:20 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Sebastian Ott <sebott@linux.vnet.ibm.com> - 2016-02-13 13:00 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-02-15 16:50 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Sebastian Ott <sebott@linux.vnet.ibm.com> - 2016-02-15 17:40 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Gerald Schaefer <gerald.schaefer@de.ibm.com> - 2016-02-15 19:40 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-02-16 00:30 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Sebastian Ott <sebott@linux.vnet.ibm.com> - 2016-02-16 11:00 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Gerald Schaefer <gerald.schaefer@de.ibm.com> - 2016-02-16 17:30 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-02-17 16:30 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Christian Borntraeger <borntraeger@de.ibm.com> - 2016-02-16 19:50 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Gerald Schaefer <gerald.schaefer@de.ibm.com> - 2016-02-17 20:20 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-02-18 08:00 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Gerald Schaefer <gerald.schaefer@de.ibm.com> - 2016-02-18 16:10 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) "Kirill A. Shutemov" <kirill@shutemov.name> - 2016-02-18 18:10 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Sebastian Ott <sebott@linux.vnet.ibm.com> - 2016-02-19 15:20 +0100
Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) Gerald Schaefer <gerald.schaefer@de.ibm.com> - 2016-02-15 17:50 +0100
Page 1 of 2 [1] 2 Next page →
| From | Gerald Schaefer <gerald.schaefer@de.ibm.com> |
|---|---|
| Date | 2016-02-11 19:30 +0100 |
| Subject | [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r14a7-44q-17@gated-at.bofh.it> |
Hi,
Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and
he also bisected this to commit 61f5d698 "mm: re-enable THP". Further
review of the THP rework patches, which cannot be bisected, revealed
commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs"
(and also similar commits for other archs).
This commit removes the THP splitting bit and also the architecture
implementation of pmdp_splitting_flush(), which took care of the IPI for
fast_gup serialization. The commit message says
pmdp_splitting_flush() is not needed too: on splitting PMD we will do
pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as
needed for fast_gup
The assumption that a TLB flush will also produce an IPI is wrong on s390,
and maybe also on other architectures, and I thought that this was actually
the main reason for having an arch-specific pmdp_splitting_flush().
At least PowerPC and ARM also had an individual implementation of
pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB
flush to send the IPI, and those were also removed. Putting the arch
maintainers and mailing lists on cc to verify.
On s390 this will break the IPI serialization against fast_gup, which
would certainly explain the random kernel crashes, please revert or fix
the pmdp_splitting_flush() removal.
Regards,
Gerald
[toc] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2016-02-11 20:10 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r14MN-4Cs-9@gated-at.bofh.it> |
| In reply to | #1332281 |
On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote: > Hi, > > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further > review of the THP rework patches, which cannot be bisected, revealed > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs" > (and also similar commits for other archs). > > This commit removes the THP splitting bit and also the architecture > implementation of pmdp_splitting_flush(), which took care of the IPI for > fast_gup serialization. The commit message says > > pmdp_splitting_flush() is not needed too: on splitting PMD we will do > pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as > needed for fast_gup > > The assumption that a TLB flush will also produce an IPI is wrong on s390, > and maybe also on other architectures, and I thought that this was actually > the main reason for having an arch-specific pmdp_splitting_flush(). > > At least PowerPC and ARM also had an individual implementation of > pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB > flush to send the IPI, and those were also removed. Putting the arch > maintainers and mailing lists on cc to verify. > > On s390 this will break the IPI serialization against fast_gup, which > would certainly explain the random kernel crashes, please revert or fix > the pmdp_splitting_flush() removal. Sorry for that. I believe, the problem was already addressed for PowerPC: http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do the trick, right? If yes, I'll prepare patch tomorrow (some sleep required). -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> |
|---|---|
| Date | 2016-02-11 20:20 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r14Wu-4FV-5@gated-at.bofh.it> |
| In reply to | #1332324 |
On Thu, Feb 11, 2016 at 09:09:42PM +0200, Kirill A. Shutemov wrote: > On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote: > > Hi, > > > > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and > > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further > > review of the THP rework patches, which cannot be bisected, revealed > > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs" > > (and also similar commits for other archs). > > > > This commit removes the THP splitting bit and also the architecture > > implementation of pmdp_splitting_flush(), which took care of the IPI for > > fast_gup serialization. The commit message says > > > > pmdp_splitting_flush() is not needed too: on splitting PMD we will do > > pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as > > needed for fast_gup > > > > The assumption that a TLB flush will also produce an IPI is wrong on s390, > > and maybe also on other architectures, and I thought that this was actually > > the main reason for having an arch-specific pmdp_splitting_flush(). > > > > At least PowerPC and ARM also had an individual implementation of > > pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB > > flush to send the IPI, and those were also removed. Putting the arch > > maintainers and mailing lists on cc to verify. > > > > On s390 this will break the IPI serialization against fast_gup, which > > would certainly explain the random kernel crashes, please revert or fix > > the pmdp_splitting_flush() removal. > > Sorry for that. > > I believe, the problem was already addressed for PowerPC: > > http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com Correct link is http://lkml.kernel.org/g/1454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Sebastian Ott <sebott@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-02-12 13:30 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1l1h-6TD-25@gated-at.bofh.it> |
| In reply to | #1332326 |
On Thu, 11 Feb 2016, Kirill A. Shutemov wrote:
> On Thu, Feb 11, 2016 at 09:09:42PM +0200, Kirill A. Shutemov wrote:
> > On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote:
> > > Hi,
> > >
> > > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and
> > > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further
> > > review of the THP rework patches, which cannot be bisected, revealed
> > > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs"
> > > (and also similar commits for other archs).
> > >
> > > This commit removes the THP splitting bit and also the architecture
> > > implementation of pmdp_splitting_flush(), which took care of the IPI for
> > > fast_gup serialization. The commit message says
> > >
> > > pmdp_splitting_flush() is not needed too: on splitting PMD we will do
> > > pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as
> > > needed for fast_gup
> > >
> > > The assumption that a TLB flush will also produce an IPI is wrong on s390,
> > > and maybe also on other architectures, and I thought that this was actually
> > > the main reason for having an arch-specific pmdp_splitting_flush().
> > >
> > > At least PowerPC and ARM also had an individual implementation of
> > > pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB
> > > flush to send the IPI, and those were also removed. Putting the arch
> > > maintainers and mailing lists on cc to verify.
> > >
> > > On s390 this will break the IPI serialization against fast_gup, which
> > > would certainly explain the random kernel crashes, please revert or fix
> > > the pmdp_splitting_flush() removal.
> >
> > Sorry for that.
> >
> > I believe, the problem was already addressed for PowerPC:
> >
> > http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com
>
> Correct link is
>
> http://lkml.kernel.org/g/1454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com
>
Based on your suggestion Gerald provided the following patch but sadly it
didn't fix the problem.
Sebastian
---
arch/s390/include/asm/pgtable.h | 2 ++
1 file changed, 2 insertions(+)
--- a/arch/s390/include/asm/pgtable.h
+++ b/arch/s390/include/asm/pgtable.h
@@ -1587,6 +1587,8 @@ static inline void pmdp_invalidate(struc
unsigned long address, pmd_t *pmdp)
{
pmdp_flush_direct(vma->vm_mm, address, pmdp);
+ /* Serialize against fast_gup with IPI */
+ kick_all_cpus_sync();
}
#define __HAVE_ARCH_PMDP_SET_WRPROTECT
[toc] | [prev] | [next] | [standalone]
| From | Gerald Schaefer <gerald.schaefer@de.ibm.com> |
|---|---|
| Date | 2016-02-11 21:00 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r15zb-4W1-9@gated-at.bofh.it> |
| In reply to | #1332324 |
On Thu, 11 Feb 2016 21:09:42 +0200 "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote: > > Hi, > > > > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and > > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further > > review of the THP rework patches, which cannot be bisected, revealed > > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs" > > (and also similar commits for other archs). > > > > This commit removes the THP splitting bit and also the architecture > > implementation of pmdp_splitting_flush(), which took care of the IPI for > > fast_gup serialization. The commit message says > > > > pmdp_splitting_flush() is not needed too: on splitting PMD we will do > > pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as > > needed for fast_gup > > > > The assumption that a TLB flush will also produce an IPI is wrong on s390, > > and maybe also on other architectures, and I thought that this was actually > > the main reason for having an arch-specific pmdp_splitting_flush(). > > > > At least PowerPC and ARM also had an individual implementation of > > pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB > > flush to send the IPI, and those were also removed. Putting the arch > > maintainers and mailing lists on cc to verify. > > > > On s390 this will break the IPI serialization against fast_gup, which > > would certainly explain the random kernel crashes, please revert or fix > > the pmdp_splitting_flush() removal. > > Sorry for that. > > I believe, the problem was already addressed for PowerPC: > > http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com > > I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do > the trick, right? Hmm, not sure about that. After pmdp_invalidate(), a pmd_none() check in fast_gup will still return false, because the pmd is not empty (at least on s390). So I don't see spontaneously how it will help fast_gup to break out to the slow path in case of THP splitting. > > If yes, I'll prepare patch tomorrow (some sleep required). > We'll check if adding kick_all_cpus_sync() to pmdp_invalidate() helps. It would also be good if Martin has a look at this, he'll return on Monday.
[toc] | [prev] | [next] | [standalone]
| From | "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-02-12 05:10 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1ddo-1Mt-3@gated-at.bofh.it> |
| In reply to | #1332342 |
Gerald Schaefer <gerald.schaefer@de.ibm.com> writes:
> On Thu, 11 Feb 2016 21:09:42 +0200
> "Kirill A. Shutemov" <kirill@shutemov.name> wrote:
>
>> On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote:
>> > Hi,
>> >
>> > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and
>> > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further
>> > review of the THP rework patches, which cannot be bisected, revealed
>> > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs"
>> > (and also similar commits for other archs).
>> >
>> > This commit removes the THP splitting bit and also the architecture
>> > implementation of pmdp_splitting_flush(), which took care of the IPI for
>> > fast_gup serialization. The commit message says
>> >
>> > pmdp_splitting_flush() is not needed too: on splitting PMD we will do
>> > pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as
>> > needed for fast_gup
>> >
>> > The assumption that a TLB flush will also produce an IPI is wrong on s390,
>> > and maybe also on other architectures, and I thought that this was actually
>> > the main reason for having an arch-specific pmdp_splitting_flush().
>> >
>> > At least PowerPC and ARM also had an individual implementation of
>> > pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB
>> > flush to send the IPI, and those were also removed. Putting the arch
>> > maintainers and mailing lists on cc to verify.
>> >
>> > On s390 this will break the IPI serialization against fast_gup, which
>> > would certainly explain the random kernel crashes, please revert or fix
>> > the pmdp_splitting_flush() removal.
>>
>> Sorry for that.
>>
>> I believe, the problem was already addressed for PowerPC:
>>
>> http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com
>>
>> I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do
>> the trick, right?
>
> Hmm, not sure about that. After pmdp_invalidate(), a pmd_none() check in
> fast_gup will still return false, because the pmd is not empty (at least
> on s390).
Why can't we do this ? I did this for ppc64.
void pmdp_invalidate(struct vm_area_struct *vma, unsigned long address,
pmd_t *pmdp)
{
- pmd_hugepage_update(vma->vm_mm, address, pmdp, _PAGE_PRESENT, 0);
+ pmd_hugepage_update(vma->vm_mm, address, pmdp, ~0UL, 0);
>So I don't see spontaneously how it will help fast_gup to break
> out to the slow path in case of THP splitting.
>
>>
>> If yes, I'll prepare patch tomorrow (some sleep required).
>>
>
> We'll check if adding kick_all_cpus_sync() to pmdp_invalidate() helps.
> It would also be good if Martin has a look at this, he'll return on
> Monday.
-aneesh
[toc] | [prev] | [next] | [standalone]
| From | Gerald Schaefer <gerald.schaefer@de.ibm.com> |
|---|---|
| Date | 2016-02-12 13:10 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1kHU-6LW-19@gated-at.bofh.it> |
| In reply to | #1332510 |
On Fri, 12 Feb 2016 09:34:33 +0530
"Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> wrote:
> Gerald Schaefer <gerald.schaefer@de.ibm.com> writes:
>
> > On Thu, 11 Feb 2016 21:09:42 +0200
> > "Kirill A. Shutemov" <kirill@shutemov.name> wrote:
> >
> >> On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote:
> >> > Hi,
> >> >
> >> > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and
> >> > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further
> >> > review of the THP rework patches, which cannot be bisected, revealed
> >> > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs"
> >> > (and also similar commits for other archs).
> >> >
> >> > This commit removes the THP splitting bit and also the architecture
> >> > implementation of pmdp_splitting_flush(), which took care of the IPI for
> >> > fast_gup serialization. The commit message says
> >> >
> >> > pmdp_splitting_flush() is not needed too: on splitting PMD we will do
> >> > pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as
> >> > needed for fast_gup
> >> >
> >> > The assumption that a TLB flush will also produce an IPI is wrong on s390,
> >> > and maybe also on other architectures, and I thought that this was actually
> >> > the main reason for having an arch-specific pmdp_splitting_flush().
> >> >
> >> > At least PowerPC and ARM also had an individual implementation of
> >> > pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB
> >> > flush to send the IPI, and those were also removed. Putting the arch
> >> > maintainers and mailing lists on cc to verify.
> >> >
> >> > On s390 this will break the IPI serialization against fast_gup, which
> >> > would certainly explain the random kernel crashes, please revert or fix
> >> > the pmdp_splitting_flush() removal.
> >>
> >> Sorry for that.
> >>
> >> I believe, the problem was already addressed for PowerPC:
> >>
> >> http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com
> >>
> >> I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do
> >> the trick, right?
> >
> > Hmm, not sure about that. After pmdp_invalidate(), a pmd_none() check in
> > fast_gup will still return false, because the pmd is not empty (at least
> > on s390).
>
> Why can't we do this ? I did this for ppc64.
>
> void pmdp_invalidate(struct vm_area_struct *vma, unsigned long address,
> pmd_t *pmdp)
> {
> - pmd_hugepage_update(vma->vm_mm, address, pmdp, _PAGE_PRESENT, 0);
> + pmd_hugepage_update(vma->vm_mm, address, pmdp, ~0UL, 0);
>
Wouldn't that semantically change what pmdp_invalidate() was supposed to
do? The comment before the call says "the pmd_trans_huge and
pmd_trans_splitting must remain set at all times on the pmd". So, after
removing pmd_trans_splitting, it seems to be necessary to at least keep
pmd_trans_huge set.
In your case, the pmd would be completely cleared, which may help to find
it in fast_gup with pmd_none(), but I'm not sure if this would open up
other problems, e.g. with concurrent page faults. But I must also admit that
my THP overview got a little rusty.
> >So I don't see spontaneously how it will help fast_gup to break
> > out to the slow path in case of THP splitting.
> >
> >>
> >> If yes, I'll prepare patch tomorrow (some sleep required).
> >>
> >
> > We'll check if adding kick_all_cpus_sync() to pmdp_invalidate() helps.
> > It would also be good if Martin has a look at this, he'll return on
> > Monday.
>
> -aneesh
>
> --
> To unsubscribe from this list: send the line "unsubscribe linux-s390" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at http://vger.kernel.org/majordomo-info.html
>
[toc] | [prev] | [next] | [standalone]
| From | "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-02-12 17:20 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1oBR-SL-33@gated-at.bofh.it> |
| In reply to | #1332696 |
Gerald Schaefer <gerald.schaefer@de.ibm.com> writes:
> On Fri, 12 Feb 2016 09:34:33 +0530
> "Aneesh Kumar K.V" <aneesh.kumar@linux.vnet.ibm.com> wrote:
>
>> Gerald Schaefer <gerald.schaefer@de.ibm.com> writes:
>>
>> > On Thu, 11 Feb 2016 21:09:42 +0200
>> > "Kirill A. Shutemov" <kirill@shutemov.name> wrote:
>> >
>> >> On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote:
>> >> > Hi,
>> >> >
>> >> > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and
>> >> > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further
>> >> > review of the THP rework patches, which cannot be bisected, revealed
>> >> > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs"
>> >> > (and also similar commits for other archs).
>> >> >
>> >> > This commit removes the THP splitting bit and also the architecture
>> >> > implementation of pmdp_splitting_flush(), which took care of the IPI for
>> >> > fast_gup serialization. The commit message says
>> >> >
>> >> > pmdp_splitting_flush() is not needed too: on splitting PMD we will do
>> >> > pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as
>> >> > needed for fast_gup
>> >> >
>> >> > The assumption that a TLB flush will also produce an IPI is wrong on s390,
>> >> > and maybe also on other architectures, and I thought that this was actually
>> >> > the main reason for having an arch-specific pmdp_splitting_flush().
>> >> >
>> >> > At least PowerPC and ARM also had an individual implementation of
>> >> > pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB
>> >> > flush to send the IPI, and those were also removed. Putting the arch
>> >> > maintainers and mailing lists on cc to verify.
>> >> >
>> >> > On s390 this will break the IPI serialization against fast_gup, which
>> >> > would certainly explain the random kernel crashes, please revert or fix
>> >> > the pmdp_splitting_flush() removal.
>> >>
>> >> Sorry for that.
>> >>
>> >> I believe, the problem was already addressed for PowerPC:
>> >>
>> >> http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com
>> >>
>> >> I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do
>> >> the trick, right?
>> >
>> > Hmm, not sure about that. After pmdp_invalidate(), a pmd_none() check in
>> > fast_gup will still return false, because the pmd is not empty (at least
>> > on s390).
>>
>> Why can't we do this ? I did this for ppc64.
>>
>> void pmdp_invalidate(struct vm_area_struct *vma, unsigned long address,
>> pmd_t *pmdp)
>> {
>> - pmd_hugepage_update(vma->vm_mm, address, pmdp, _PAGE_PRESENT, 0);
>> + pmd_hugepage_update(vma->vm_mm, address, pmdp, ~0UL, 0);
>>
>
> Wouldn't that semantically change what pmdp_invalidate() was supposed to
> do? The comment before the call says "the pmd_trans_huge and
> pmd_trans_splitting must remain set at all times on the pmd". So, after
> removing pmd_trans_splitting, it seems to be necessary to at least keep
> pmd_trans_huge set.
>
> In your case, the pmd would be completely cleared, which may help to find
> it in fast_gup with pmd_none(), but I'm not sure if this would open up
> other problems, e.g. with concurrent page faults. But I must also admit that
> my THP overview got a little rusty.
Thinking about this more, I guess, I should not be doing this. Because
this bring in the exit_mmap race that I outlined in the patch even
though the window now is small.
I guess we should fix this in the gup path by checking for what ever
trick we are using to mark the pmd splitting. For ppc64 we clear the
_PAGE_USER. We are ok as long as autonuma is enabled because
pmd_protnone() check will check against _PAGE_USER. But that may not be
sufficient.
-aneesh
[toc] | [prev] | [next] | [standalone]
| From | Will Deacon <will.deacon@arm.com> |
|---|---|
| Date | 2016-02-12 11:10 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1iPM-5vL-19@gated-at.bofh.it> |
| In reply to | #1332342 |
On Thu, Feb 11, 2016 at 08:57:02PM +0100, Gerald Schaefer wrote: > On Thu, 11 Feb 2016 21:09:42 +0200 > "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > > On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote: > > > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and > > > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further > > > review of the THP rework patches, which cannot be bisected, revealed > > > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs" > > > (and also similar commits for other archs). > > > > > > This commit removes the THP splitting bit and also the architecture > > > implementation of pmdp_splitting_flush(), which took care of the IPI for > > > fast_gup serialization. The commit message says > > > > > > pmdp_splitting_flush() is not needed too: on splitting PMD we will do > > > pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as > > > needed for fast_gup > > > > > > The assumption that a TLB flush will also produce an IPI is wrong on s390, > > > and maybe also on other architectures, and I thought that this was actually > > > the main reason for having an arch-specific pmdp_splitting_flush(). > > > > > > At least PowerPC and ARM also had an individual implementation of > > > pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB > > > flush to send the IPI, and those were also removed. Putting the arch > > > maintainers and mailing lists on cc to verify. > > > > > > On s390 this will break the IPI serialization against fast_gup, which > > > would certainly explain the random kernel crashes, please revert or fix > > > the pmdp_splitting_flush() removal. > > > > Sorry for that. > > > > I believe, the problem was already addressed for PowerPC: > > > > http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com > > > > I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do > > the trick, right? > > Hmm, not sure about that. After pmdp_invalidate(), a pmd_none() check in > fast_gup will still return false, because the pmd is not empty (at least > on s390). So I don't see spontaneously how it will help fast_gup to break > out to the slow path in case of THP splitting. > > > > > If yes, I'll prepare patch tomorrow (some sleep required). > > > > We'll check if adding kick_all_cpus_sync() to pmdp_invalidate() helps. > It would also be good if Martin has a look at this, he'll return on > Monday. Do you have a reliable way to trigger the "random kernel crashes"? We've not seen anything reported on arm64, but I don't see why we wouldn't be affected by the same bug and it would be good to confirm and validate a fix. Cheers, Will
[toc] | [prev] | [next] | [standalone]
| From | Sebastian Ott <sebott@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-02-12 11:20 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1iZr-5zo-3@gated-at.bofh.it> |
| In reply to | #1332657 |
On Fri, 12 Feb 2016, Will Deacon wrote: > On Thu, Feb 11, 2016 at 08:57:02PM +0100, Gerald Schaefer wrote: > > On Thu, 11 Feb 2016 21:09:42 +0200 > > "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > > > On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote: > > > > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and > > > > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further > > > > review of the THP rework patches, which cannot be bisected, revealed > > > > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs" > > > > (and also similar commits for other archs). > > > > > > > > This commit removes the THP splitting bit and also the architecture > > > > implementation of pmdp_splitting_flush(), which took care of the IPI for > > > > fast_gup serialization. The commit message says > > > > > > > > pmdp_splitting_flush() is not needed too: on splitting PMD we will do > > > > pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as > > > > needed for fast_gup > > > > > > > > The assumption that a TLB flush will also produce an IPI is wrong on s390, > > > > and maybe also on other architectures, and I thought that this was actually > > > > the main reason for having an arch-specific pmdp_splitting_flush(). > > > > > > > > At least PowerPC and ARM also had an individual implementation of > > > > pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB > > > > flush to send the IPI, and those were also removed. Putting the arch > > > > maintainers and mailing lists on cc to verify. > > > > > > > > On s390 this will break the IPI serialization against fast_gup, which > > > > would certainly explain the random kernel crashes, please revert or fix > > > > the pmdp_splitting_flush() removal. > > > > > > Sorry for that. > > > > > > I believe, the problem was already addressed for PowerPC: > > > > > > http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com > > > > > > I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do > > > the trick, right? > > > > Hmm, not sure about that. After pmdp_invalidate(), a pmd_none() check in > > fast_gup will still return false, because the pmd is not empty (at least > > on s390). So I don't see spontaneously how it will help fast_gup to break > > out to the slow path in case of THP splitting. > > > > > > > > If yes, I'll prepare patch tomorrow (some sleep required). > > > > > > > We'll check if adding kick_all_cpus_sync() to pmdp_invalidate() helps. > > It would also be good if Martin has a look at this, he'll return on > > Monday. > > Do you have a reliable way to trigger the "random kernel crashes"? We've not > seen anything reported on arm64, but I don't see why we wouldn't be affected > by the same bug and it would be good to confirm and validate a fix. My testcase was compiling the kernel. Most of the time my test system didn't survive a single compile run. During bisect I did at least 20 compile runs to flag a commit as good. Sebastian
[toc] | [prev] | [next] | [standalone]
| From | Will Deacon <will.deacon@arm.com> |
|---|---|
| Date | 2016-02-12 17:00 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1oiu-ws-23@gated-at.bofh.it> |
| In reply to | #1332659 |
On Fri, Feb 12, 2016 at 11:12:34AM +0100, Sebastian Ott wrote: > On Fri, 12 Feb 2016, Will Deacon wrote: > > On Thu, Feb 11, 2016 at 08:57:02PM +0100, Gerald Schaefer wrote: > > > On Thu, 11 Feb 2016 21:09:42 +0200 > > > "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > > > > On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote: > > > > > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and > > > > > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further > > > > > review of the THP rework patches, which cannot be bisected, revealed > > > > > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs" > > > > > (and also similar commits for other archs). [...] > > Do you have a reliable way to trigger the "random kernel crashes"? We've not > > seen anything reported on arm64, but I don't see why we wouldn't be affected > > by the same bug and it would be good to confirm and validate a fix. > > My testcase was compiling the kernel. Most of the time my test system > didn't survive a single compile run. During bisect I did at least 20 > compile runs to flag a commit as good. I've been building kernels all day with -rc3 on my arm64 box and haven't seen any problems yet.. :/. I'll leave it going over the weekend. Will
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2016-02-12 16:50 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1o8O-tk-33@gated-at.bofh.it> |
| In reply to | #1332342 |
On Thu, Feb 11, 2016 at 08:57:02PM +0100, Gerald Schaefer wrote: > On Thu, 11 Feb 2016 21:09:42 +0200 > "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > > > On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote: > > > Hi, > > > > > > Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and > > > he also bisected this to commit 61f5d698 "mm: re-enable THP". Further > > > review of the THP rework patches, which cannot be bisected, revealed > > > commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs" > > > (and also similar commits for other archs). > > > > > > This commit removes the THP splitting bit and also the architecture > > > implementation of pmdp_splitting_flush(), which took care of the IPI for > > > fast_gup serialization. The commit message says > > > > > > pmdp_splitting_flush() is not needed too: on splitting PMD we will do > > > pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as > > > needed for fast_gup > > > > > > The assumption that a TLB flush will also produce an IPI is wrong on s390, > > > and maybe also on other architectures, and I thought that this was actually > > > the main reason for having an arch-specific pmdp_splitting_flush(). > > > > > > At least PowerPC and ARM also had an individual implementation of > > > pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB > > > flush to send the IPI, and those were also removed. Putting the arch > > > maintainers and mailing lists on cc to verify. > > > > > > On s390 this will break the IPI serialization against fast_gup, which > > > would certainly explain the random kernel crashes, please revert or fix > > > the pmdp_splitting_flush() removal. > > > > Sorry for that. > > > > I believe, the problem was already addressed for PowerPC: > > > > http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com > > > > I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do > > the trick, right? > > Hmm, not sure about that. After pmdp_invalidate(), a pmd_none() check in > fast_gup will still return false, because the pmd is not empty (at least > on s390). So I don't see spontaneously how it will help fast_gup to break > out to the slow path in case of THP splitting. What pmdp_flush_direct() does in pmdp_invalidate()? It's hard to unwrap for me :-/ Does it make the pmd !pmd_present()? I'm also confused by pmd_none() is equal to !pmd_present() on s390. Hm? -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Christian Borntraeger <borntraeger@de.ibm.com> |
|---|---|
| Date | 2016-02-12 17:00 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1oiu-ws-7@gated-at.bofh.it> |
| In reply to | #1332824 |
On 02/12/2016 04:41 PM, Kirill A. Shutemov wrote: > On Thu, Feb 11, 2016 at 08:57:02PM +0100, Gerald Schaefer wrote: >> On Thu, 11 Feb 2016 21:09:42 +0200 >> "Kirill A. Shutemov" <kirill@shutemov.name> wrote: >> >>> On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote: >>>> Hi, >>>> >>>> Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and >>>> he also bisected this to commit 61f5d698 "mm: re-enable THP". Further >>>> review of the THP rework patches, which cannot be bisected, revealed >>>> commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs" >>>> (and also similar commits for other archs). >>>> >>>> This commit removes the THP splitting bit and also the architecture >>>> implementation of pmdp_splitting_flush(), which took care of the IPI for >>>> fast_gup serialization. The commit message says >>>> >>>> pmdp_splitting_flush() is not needed too: on splitting PMD we will do >>>> pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as >>>> needed for fast_gup >>>> >>>> The assumption that a TLB flush will also produce an IPI is wrong on s390, >>>> and maybe also on other architectures, and I thought that this was actually >>>> the main reason for having an arch-specific pmdp_splitting_flush(). >>>> >>>> At least PowerPC and ARM also had an individual implementation of >>>> pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB >>>> flush to send the IPI, and those were also removed. Putting the arch >>>> maintainers and mailing lists on cc to verify. >>>> >>>> On s390 this will break the IPI serialization against fast_gup, which >>>> would certainly explain the random kernel crashes, please revert or fix >>>> the pmdp_splitting_flush() removal. >>> >>> Sorry for that. >>> >>> I believe, the problem was already addressed for PowerPC: >>> >>> http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com >>> >>> I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do >>> the trick, right? >> >> Hmm, not sure about that. After pmdp_invalidate(), a pmd_none() check in >> fast_gup will still return false, because the pmd is not empty (at least >> on s390). So I don't see spontaneously how it will help fast_gup to break >> out to the slow path in case of THP splitting. > > What pmdp_flush_direct() does in pmdp_invalidate()? It's hard to unwrap for me :-/ > Does it make the pmd !pmd_present()? It uses the idte instruction, which in an atomic fashion flushes the associated TLB entry and changes the value of the pmd entry to invalid. This comes from the HW requirement to not change a PTE/PMD that might be still in use, other than with special instructions that does the tlb handling and the invalidation together. (It also does some some other magic to the attach_count, which might hold off finish_arch_post_lock_switch while some flushing is happening, but this should be unrelated here) > I'm also confused by pmd_none() is equal to !pmd_present() on s390. Hm? Don't know, Gerald or Martin?
[toc] | [prev] | [next] | [standalone]
| From | Gerald Schaefer <gerald.schaefer@de.ibm.com> |
|---|---|
| Date | 2016-02-12 18:20 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1pxW-1vA-55@gated-at.bofh.it> |
| In reply to | #1332830 |
On Fri, 12 Feb 2016 16:57:27 +0100 Christian Borntraeger <borntraeger@de.ibm.com> wrote: > On 02/12/2016 04:41 PM, Kirill A. Shutemov wrote: > > On Thu, Feb 11, 2016 at 08:57:02PM +0100, Gerald Schaefer wrote: > >> On Thu, 11 Feb 2016 21:09:42 +0200 > >> "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > >> > >>> On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote: > >>>> Hi, > >>>> > >>>> Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and > >>>> he also bisected this to commit 61f5d698 "mm: re-enable THP". Further > >>>> review of the THP rework patches, which cannot be bisected, revealed > >>>> commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs" > >>>> (and also similar commits for other archs). > >>>> > >>>> This commit removes the THP splitting bit and also the architecture > >>>> implementation of pmdp_splitting_flush(), which took care of the IPI for > >>>> fast_gup serialization. The commit message says > >>>> > >>>> pmdp_splitting_flush() is not needed too: on splitting PMD we will do > >>>> pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as > >>>> needed for fast_gup > >>>> > >>>> The assumption that a TLB flush will also produce an IPI is wrong on s390, > >>>> and maybe also on other architectures, and I thought that this was actually > >>>> the main reason for having an arch-specific pmdp_splitting_flush(). > >>>> > >>>> At least PowerPC and ARM also had an individual implementation of > >>>> pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB > >>>> flush to send the IPI, and those were also removed. Putting the arch > >>>> maintainers and mailing lists on cc to verify. > >>>> > >>>> On s390 this will break the IPI serialization against fast_gup, which > >>>> would certainly explain the random kernel crashes, please revert or fix > >>>> the pmdp_splitting_flush() removal. > >>> > >>> Sorry for that. > >>> > >>> I believe, the problem was already addressed for PowerPC: > >>> > >>> http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com > >>> > >>> I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do > >>> the trick, right? > >> > >> Hmm, not sure about that. After pmdp_invalidate(), a pmd_none() check in > >> fast_gup will still return false, because the pmd is not empty (at least > >> on s390). So I don't see spontaneously how it will help fast_gup to break > >> out to the slow path in case of THP splitting. > > > > What pmdp_flush_direct() does in pmdp_invalidate()? It's hard to unwrap for me :-/ > > Does it make the pmd !pmd_present()? > > It uses the idte instruction, which in an atomic fashion flushes the associated > TLB entry and changes the value of the pmd entry to invalid. This comes from the > HW requirement to not change a PTE/PMD that might be still in use, other than > with special instructions that does the tlb handling and the invalidation together. Correct, and it does _not_ make the pmd !pmd_present(), that would only be the case after a _clear_flush(). It only marks the pmd as invalid and flushes, so that it cannot generate a new TLB entry before the following pmd_populate(), but it keeps its other content. This is to fulfill the requirements outlined in the comment in mm/huge_memory.c before the call to pmdp_invalidate(). And independent from that comment, we would need such an _invalidate() or _clear_flush() on s390 before the pmd_populate() because of the HW details that Christian described. Reading the comment again, I do now notice that it also says "mark the current pmd notpresent", which we cannot do w/o losing the huge and (formerly) splitting bits, but it also shouldn't be needed to provide the "single TLB guarantee" that is required from the comment. So, a pmd_present() check on s390 in this state would still return true. Not sure yet if this is a problem, need more thinking, this behavior was already present before the THP rework but maybe it was OK before and is not OK now. At least for fast_gup this should not be a problem though. > (It also does some some other magic to the attach_count, which might hold off > finish_arch_post_lock_switch while some flushing is happening, but this should > be unrelated here) > > > > I'm also confused by pmd_none() is equal to !pmd_present() on s390. Hm? > > Don't know, Gerald or Martin? The implementation frequently changes depending on how many new bits Martin needs to squeeze out :-) We don't have a _PAGE_PRESENT bit for pmds, so pmd_present() just checks if the entry is not empty. pmd_none() of course does the opposite, it checks if it is empty. > > -- > To unsubscribe from this list: send the line "unsubscribe linux-s390" in > the body of a message to majordomo@vger.kernel.org > More majordomo info at http://vger.kernel.org/majordomo-info.html >
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2016-02-13 00:20 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1vai-5dV-13@gated-at.bofh.it> |
| In reply to | #1332947 |
On Fri, Feb 12, 2016 at 06:16:40PM +0100, Gerald Schaefer wrote: > On Fri, 12 Feb 2016 16:57:27 +0100 > Christian Borntraeger <borntraeger@de.ibm.com> wrote: > > > On 02/12/2016 04:41 PM, Kirill A. Shutemov wrote: > > > On Thu, Feb 11, 2016 at 08:57:02PM +0100, Gerald Schaefer wrote: > > >> On Thu, 11 Feb 2016 21:09:42 +0200 > > >> "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > > >> > > >>> On Thu, Feb 11, 2016 at 07:22:23PM +0100, Gerald Schaefer wrote: > > >>>> Hi, > > >>>> > > >>>> Sebastian Ott reported random kernel crashes beginning with v4.5-rc1 and > > >>>> he also bisected this to commit 61f5d698 "mm: re-enable THP". Further > > >>>> review of the THP rework patches, which cannot be bisected, revealed > > >>>> commit fecffad "s390, thp: remove infrastructure for handling splitting PMDs" > > >>>> (and also similar commits for other archs). > > >>>> > > >>>> This commit removes the THP splitting bit and also the architecture > > >>>> implementation of pmdp_splitting_flush(), which took care of the IPI for > > >>>> fast_gup serialization. The commit message says > > >>>> > > >>>> pmdp_splitting_flush() is not needed too: on splitting PMD we will do > > >>>> pmdp_clear_flush() + set_pte_at(). pmdp_clear_flush() will do IPI as > > >>>> needed for fast_gup > > >>>> > > >>>> The assumption that a TLB flush will also produce an IPI is wrong on s390, > > >>>> and maybe also on other architectures, and I thought that this was actually > > >>>> the main reason for having an arch-specific pmdp_splitting_flush(). > > >>>> > > >>>> At least PowerPC and ARM also had an individual implementation of > > >>>> pmdp_splitting_flush() that used kick_all_cpus_sync() instead of a TLB > > >>>> flush to send the IPI, and those were also removed. Putting the arch > > >>>> maintainers and mailing lists on cc to verify. > > >>>> > > >>>> On s390 this will break the IPI serialization against fast_gup, which > > >>>> would certainly explain the random kernel crashes, please revert or fix > > >>>> the pmdp_splitting_flush() removal. > > >>> > > >>> Sorry for that. > > >>> > > >>> I believe, the problem was already addressed for PowerPC: > > >>> > > >>> http://lkml.kernel.org/g/454980831-16631-1-git-send-email-aneesh.kumar@linux.vnet.ibm.com > > >>> > > >>> I think kick_all_cpus_sync() in arch-specific pmdp_invalidate() would do > > >>> the trick, right? > > >> > > >> Hmm, not sure about that. After pmdp_invalidate(), a pmd_none() check in > > >> fast_gup will still return false, because the pmd is not empty (at least > > >> on s390). So I don't see spontaneously how it will help fast_gup to break > > >> out to the slow path in case of THP splitting. > > > > > > What pmdp_flush_direct() does in pmdp_invalidate()? It's hard to unwrap for me :-/ > > > Does it make the pmd !pmd_present()? > > > > It uses the idte instruction, which in an atomic fashion flushes the associated > > TLB entry and changes the value of the pmd entry to invalid. This comes from the > > HW requirement to not change a PTE/PMD that might be still in use, other than > > with special instructions that does the tlb handling and the invalidation together. > > Correct, and it does _not_ make the pmd !pmd_present(), that would only be the > case after a _clear_flush(). It only marks the pmd as invalid and flushes, > so that it cannot generate a new TLB entry before the following pmd_populate(), > but it keeps its other content. This is to fulfill the requirements outlined in > the comment in mm/huge_memory.c before the call to pmdp_invalidate(). And > independent from that comment, we would need such an _invalidate() or > _clear_flush() on s390 before the pmd_populate() because of the HW details > that Christian described. > > Reading the comment again, I do now notice that it also says "mark the current > pmd notpresent", which we cannot do w/o losing the huge and (formerly) splitting > bits, but it also shouldn't be needed to provide the "single TLB guarantee" that > is required from the comment. So, a pmd_present() check on s390 in this state > would still return true. Not sure yet if this is a problem, need more thinking, > this behavior was already present before the THP rework but maybe it was OK > before and is not OK now. > > At least for fast_gup this should not be a problem though. I'm trying to wrap my head around the issue and I don't think missing serialization with gup_fast is the cause -- we just don't need it anymore. Previously, __split_huge_page_splitting() required serialization against gup_fast to make sure nobody can obtain new reference to the page after __split_huge_page_splitting() returns. This was a way to stabilize page references before starting to distribute them from head page to tail pages. With new refcounting, we don't care about this. Splitting PMD is now decoupled from splitting underlying compound page. It's okay to get new pins after split_huge_pmd(). To stabilize page references during split_huge_page() we rely on setting up migration entries once all pmds are split into page table entries. The theory that serialization against gup_fast is not a root cause of the crashes is consistent no crashes on arm64. Problem is somewhere else. > > (It also does some some other magic to the attach_count, which might hold off > > finish_arch_post_lock_switch while some flushing is happening, but this should > > be unrelated here) > > > > > > > I'm also confused by pmd_none() is equal to !pmd_present() on s390. Hm? > > > > Don't know, Gerald or Martin? > > The implementation frequently changes depending on how many new bits Martin > needs to squeeze out :-) One bit was freed up by the commit you've pointed to as a cause. I wounder If it's possible that screw up something while removing it? I don't see it, but who knows. Could you check if revert of fecffad25458 helps? And could you share how crashes looks like? I haven't seen backtraces yet. > We don't have a _PAGE_PRESENT bit for pmds, so pmd_present() just checks if the > entry is not empty. pmd_none() of course does the opposite, it checks if it is > empty. -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Sebastian Ott <sebott@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-02-13 13:00 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r1H1L-4f9-9@gated-at.bofh.it> |
| In reply to | #1333202 |
[Multipart message — attachments visible in raw view] — view raw
On Sat, 13 Feb 2016, Kirill A. Shutemov wrote:
> Could you check if revert of fecffad25458 helps?
I reverted fecffad25458 on top of 721675fcf277cf - it oopsed with:
¢ 1851.721062! Unable to handle kernel pointer dereference in virtual kernel address space
¢ 1851.721075! failing address: 0000000000000000 TEID: 0000000000000483
¢ 1851.721078! Fault in home space mode while using kernel ASCE.
¢ 1851.721085! AS:0000000000d5c007 R3:00000000ffff0007 S:00000000ffffa800 P:000000000000003d
¢ 1851.721128! Oops: 0004 ilc:3 ¢#1! PREEMPT SMP DEBUG_PAGEALLOC
¢ 1851.721135! Modules linked in: bridge stp llc btrfs mlx4_ib mlx4_en ib_sa ib_mad vxlan xor ip6_udp_tunnel ib_core udp_tunnel ptp pps_core ib_addr ghash_s390raid6_pq prng ecb aes_s390 mlx4_core des_s390 des_generic genwqe_card sha512_s390 sha256_s390 sha1_s390 sha_common crc_itu_t dm_mod scm_block vhost_net tun vhost eadm_sch macvtap macvlan kvm autofs4
¢ 1851.721183! CPU: 7 PID: 256422 Comm: bash Not tainted 4.5.0-rc3-00058-g07923d7-dirty #178
¢ 1851.721186! task: 000000007fbfd290 ti: 000000008c604000 task.ti: 000000008c604000
¢ 1851.721189! Krnl PSW : 0704d00180000000 000000000045d3b8 (__rb_erase_color+0x280/0x308)
¢ 1851.721200! R:0 T:1 IO:1 EX:1 Key:0 M:1 W:0 P:0 AS:3 CC:1 PM:0 EA:3
Krnl GPRS: 0000000000000001 0000000000000020 0000000000000000 00000000bd07eff1
¢ 1851.721205! 000000000027ca10 0000000000000000 0000000083e45898 0000000077b61198
¢ 1851.721207! 000000007ce1a490 00000000bd07eff0 000000007ce1a548 000000000027ca10
¢ 1851.721210! 00000000bd07c350 00000000bd07eff0 000000008c607aa8 000000008c607a68
¢ 1851.721221! Krnl Code: 000000000045d3aa: e3c0d0080024 stg %%r12,8(%%r13)
000000000045d3b0: b9040039 lgr %%r3,%%r9
#000000000045d3b4: a53b0001 oill %%r3,1
>000000000045d3b8: e33010000024 stg %%r3,0(%%r1)
000000000045d3be: ec28000e007c cgij %%r2,0,8,45d3da
000000000045d3c4: e34020000004 lg %%r4,0(%%r2)
000000000045d3ca: b904001c lgr %%r1,%%r12
000000000045d3ce: ec143f3f0056 rosbg %%r1,%%r4,63,63,0
¢ 1851.721269! Call Trace:
¢ 1851.721273! (¢<0000000083e45898>! 0x83e45898)
¢ 1851.721279! ¢<000000000029342a>! unlink_anon_vmas+0x9a/0x1d8
¢ 1851.721282! ¢<0000000000283f34>! free_pgtables+0xcc/0x148
¢ 1851.721285! ¢<000000000028c376>! exit_mmap+0xd6/0x300
¢ 1851.721289! ¢<0000000000134db8>! mmput+0x90/0x118
¢ 1851.721294! ¢<00000000002d76bc>! flush_old_exec+0x5d4/0x700
¢ 1851.721298! ¢<00000000003369f4>! load_elf_binary+0x2f4/0x13e8
¢ 1851.721301! ¢<00000000002d6e4a>! search_binary_handler+0x9a/0x1f8
¢ 1851.721304! ¢<00000000002d8970>! do_execveat_common.isra.32+0x668/0x9a0
¢ 1851.721307! ¢<00000000002d8cec>! do_execve+0x44/0x58
¢ 1851.721310! ¢<00000000002d8f92>! SyS_execve+0x3a/0x48
¢ 1851.721315! ¢<00000000006fb096>! system_call+0xd6/0x258
¢ 1851.721317! ¢<000003ff997436d6>! 0x3ff997436d6
¢ 1851.721319! INFO: lockdep is turned off.
¢ 1851.721321! Last Breaking-Event-Address:
¢ 1851.721323! ¢<000000000045d31a>! __rb_erase_color+0x1e2/0x308
¢ 1851.721327!
¢ 1851.721329! ---¢ end trace 0d80041ac00cfae2 !---
>
> And could you share how crashes looks like? I haven't seen backtraces yet.
>
Sure. I didn't because they really looked random to me. Most of the time
in rcu or list debugging but I thought these have just been the messenger
observing a corruption first. Anyhow, here is an older one that might look
interesting:
[ 59.851421] list_del corruption. next->prev should be 000000006e1eb000, but was 0000000000000400
[ 59.851469] ------------[ cut here ]------------
[ 59.851472] WARNING: at lib/list_debug.c:71
[ 59.851475] Modules linked in: bridge stp llc btrfs xor mlx4_en vxlan ip6_udp_tunnel udp_tunnel mlx4_ib ptp pps_core ib_sa ib_mad ib_core ib_addr ghash_s390 prng raid6_pq ecb aes_s390 des_s390 des_generic sha512_s390 sha256_s390 sha1_s390 mlx4_core sha_common genwqe_card scm_block crc_itu_t vhost_net tun vhost dm_mod macvtap eadm_sch macvlan kvm autofs4
[ 59.851532] CPU: 0 PID: 5400 Comm: git Not tainted 4.4.0-07794-ga4eff16-dirty #77
[ 59.851535] task: 00000000d2310000 ti: 00000000d6610000 task.ti: 00000000d6610000
[ 59.851539] Krnl PSW : 0704c00180000000 0000000000487434 (__list_del_entry+0xa4/0xe0)
[ 59.851548] R:0 T:1 IO:1 EX:1 Key:0 M:1 W:0 P:0 AS:3 CC:0 PM:0 EA:3
Krnl GPRS: 0000000001a7a1cf 00000000d2310000 0000000000000054 0000000000000001
[ 59.851554] 0000000000487430 0000000000000000 0000000000000000 00000000774e6900
[ 59.851557] 000003ff53000000 000000006d4017a0 000003ff52f00000 000003ff52f00000
[ 59.851560] 000003d101780000 000000006e1eb000 0000000000487430 00000000d6613b00
[ 59.851571] Krnl Code: 0000000000487424: c02000219e3a larl %%r2,8bb098
000000000048742a: c0e5ffee05db brasl %%r14,247fe0
#0000000000487430: a7f40001 brc 15,487432
>0000000000487434: a7f40017 brc 15,487462
0000000000487438: a7390200 lghi %%r3,512
000000000048743c: ec13ffd28064 cgrj %%r1,%%r3,8,4873e0
0000000000487442: e32010000020 cg %%r2,0(%%r1)
0000000000487448: a774ffda brc 7,4873fc
[ 59.851615] Call Trace:
[ 59.851618] ([<0000000000487430>] __list_del_entry+0xa0/0xe0)
[ 59.851621] [<0000000000487498>] list_del+0x28/0x40
[ 59.851627] [<00000000001259ec>] pgtable_trans_huge_withdraw+0x74/0x90
[ 59.851632] [<00000000002bf234>] __split_huge_pmd_locked+0x3ec/0xa10
[ 59.851635] [<00000000002c4310>] __split_huge_pmd+0x118/0x218
[ 59.851639] [<00000000002810e8>] unmap_single_vma+0x2d8/0xb40
[ 59.851643] [<0000000000282d66>] zap_page_range+0x116/0x318
[ 59.851646] [<000000000029b834>] SyS_madvise+0x23c/0x5e8
[ 59.851652] [<00000000006f9f56>] system_call+0xd6/0x258
[ 59.851656] [<000003ff9bbfd282>] 0x3ff9bbfd282
[ 59.851658] 2 locks held by git/5400:
[ 59.851660] #0: (&mm->mmap_sem){++++++}, at: [<000000000029bb5a>] SyS_madvise+0x562/0x5e8
[ 59.851670] #1: (&(ptlock_ptr(page))->rlock){+.+...}, at: [<00000000002c4268>] __split_huge_pmd+0x70/0x218
[ 59.851679] Last Breaking-Event-Address:
[ 59.851682] [<0000000000487430>] __list_del_entry+0xa0/0xe0
[ 59.851686] ---[ end trace 7bce9a4f571985b6 ]---
[ 59.875754] list_del corruption. prev->next should be 000000006e1eb820, but was (null)
[ 59.875768] ------------[ cut here ]------------
[ 59.875771] WARNING: at lib/list_debug.c:68
[ 59.875774] Modules linked in: bridge stp llc btrfs xor mlx4_en vxlan ip6_udp_tunnel udp_tunnel mlx4_ib ptp pps_core ib_sa ib_mad ib_core ib_addr ghash_s390 prng raid6_pq ecb aes_s390 des_s390 des_generic sha512_s390 sha256_s390 sha1_s390 mlx4_core sha_common genwqe_card scm_block crc_itu_t vhost_net tun vhost dm_mod macvtap eadm_sch macvlan kvm autofs4
[ 59.875820] CPU: 2 PID: 5402 Comm: git Tainted: G W 4.4.0-07794-ga4eff16-dirty #77
[ 59.875823] task: 00000000d2312948 ti: 00000000cfecc000 task.ti: 00000000cfecc000
[ 59.875826] Krnl PSW : 0704c00180000000 0000000000487416 (__list_del_entry+0x86/0xe0)
[ 59.875832] R:0 T:1 IO:1 EX:1 Key:0 M:1 W:0 P:0 AS:3 CC:0 PM:0 EA:3
Krnl GPRS: 0000000001a7a1cf 00000000d2312948 0000000000000054 0000000000000001
[ 59.875838] 0000000000487412 0000000000000000 0000000000000000 00000000774e6900
[ 59.875841] 000003ff52000000 000000006d403b10 000003ff51f00000 000003ff51f00000
[ 59.875843] 000003d10177c000 000000006e1eb820 0000000000487412 00000000cfecfb00
[ 59.875851] Krnl Code: 0000000000487406: c02000219e2c larl %%r2,8bb05e
000000000048740c: c0e5ffee05ea brasl %%r14,247fe0
#0000000000487412: a7f40001 brc 15,487414
>0000000000487416: a7f40026 brc 15,487462
000000000048741a: b9040032 lgr %%r3,%%r2
000000000048741e: e34040080004 lg %%r4,8(%%r4)
0000000000487424: c02000219e3a larl %%r2,8bb098
000000000048742a: c0e5ffee05db brasl %%r14,247fe0
[ 59.875874] Call Trace:
[ 59.875876] ([<0000000000487412>] __list_del_entry+0x82/0xe0)
[ 59.875879] [<0000000000487498>] list_del+0x28/0x40
[ 59.875882] [<00000000001259ec>] pgtable_trans_huge_withdraw+0x74/0x90
[ 59.875885] [<00000000002bf234>] __split_huge_pmd_locked+0x3ec/0xa10
[ 59.875888] [<00000000002c4310>] __split_huge_pmd+0x118/0x218
[ 59.875891] [<00000000002810e8>] unmap_single_vma+0x2d8/0xb40
[ 59.875894] [<0000000000282d66>] zap_page_range+0x116/0x318
[ 59.875896] [<000000000029b834>] SyS_madvise+0x23c/0x5e8
[ 59.875899] [<00000000006f9f56>] system_call+0xd6/0x258
[ 59.875902] [<000003ff9bbfd282>] 0x3ff9bbfd282
[ 59.875904] 2 locks held by git/5402:
[ 59.875906] #0: (&mm->mmap_sem){++++++}, at: [<000000000029bb5a>] SyS_madvise+0x562/0x5e8
[ 59.875914] #1: (&(ptlock_ptr(page))->rlock){+.+...}, at: [<00000000002c4268>] __split_huge_pmd+0x70/0x218
[ 59.875922] Last Breaking-Event-Address:
[ 59.875925] [<0000000000487412>] __list_del_entry+0x82/0xe0
[ 59.875927] ---[ end trace 7bce9a4f571985b7 ]---
[ 59.875935] ------------[ cut here ]------------
[ 59.875937] kernel BUG at mm/huge_memory.c:2884!
[ 59.875979] illegal operation: 0001 ilc:1 [#1] PREEMPT SMP DEBUG_PAGEALLOC
[ 59.875986] Modules linked in: bridge stp llc btrfs xor mlx4_en vxlan ip6_udp_tunnel udp_tunnel mlx4_ib ptp pps_core ib_sa ib_mad ib_core ib_addr ghash_s390 prng raid6_pq ecb aes_s390 des_s390 des_generic sha512_s390 sha256_s390 sha1_s390 mlx4_core sha_common genwqe_card scm_block crc_itu_t vhost_net tun vhost dm_mod macvtap eadm_sch macvlan kvm autofs4
[ 59.876033] CPU: 2 PID: 5402 Comm: git Tainted: G W 4.4.0-07794-ga4eff16-dirty #77
[ 59.876036] task: 00000000d2312948 ti: 00000000cfecc000 task.ti: 00000000cfecc000
[ 59.876039] Krnl PSW : 0704d00180000000 00000000002bf3aa (__split_huge_pmd_locked+0x562/0xa10)
[ 59.876045] R:0 T:1 IO:1 EX:1 Key:0 M:1 W:0 P:0 AS:3 CC:1 PM:0 EA:3
Krnl GPRS: 0000000001a7a1cf 000003d10177c000 0000000000044068 000000005df00215
[ 59.876051] 0000000000000001 0000000000000001 0000000000000000 00000000774e6900
[ 59.876054] 000003ff52000000 000000006d403b10 000000006e1eb800 000003ff51f00000
[ 59.876058] 000003d10177c000 0000000000715190 00000000002bf234 00000000cfecfb58
[ 59.876068] Krnl Code: 00000000002bf39c: d507d010a000 clc 16(8,%%r13),0(%%r10)
00000000002bf3a2: a7840004 brc 8,2bf3aa
#00000000002bf3a6: a7f40001 brc 15,2bf3a8
>00000000002bf3aa: 91407440 tm 1088(%%r7),64
00000000002bf3ae: a7840208 brc 8,2bf7be
00000000002bf3b2: a7f401e9 brc 15,2bf784
00000000002bf3b6: 9104a006 tm 6(%%r10),4
00000000002bf3ba: a7740004 brc 7,2bf3c2
[ 59.876089] Call Trace:
[ 59.876092] ([<00000000002bf234>] __split_huge_pmd_locked+0x3ec/0xa10)
[ 59.876095] [<00000000002c4310>] __split_huge_pmd+0x118/0x218
[ 59.876099] [<00000000002810e8>] unmap_single_vma+0x2d8/0xb40
[ 59.876102] [<0000000000282d66>] zap_page_range+0x116/0x318
[ 59.876105] [<000000000029b834>] SyS_madvise+0x23c/0x5e8
[ 59.876108] [<00000000006f9f56>] system_call+0xd6/0x258
[ 59.876111] [<000003ff9bbfd282>] 0x3ff9bbfd282
[ 59.876113] INFO: lockdep is turned off.
[ 59.876115] Last Breaking-Event-Address:
[ 59.876118] [<00000000002bf3a6>] __split_huge_pmd_locked+0x55e/0xa10
[ 59.876122]
[ 59.876124] ---[ end trace 7bce9a4f571985b8 ]---
[ 59.876128] BUG: sleeping function called from invalid context at include/linux/sched.h:2791
[ 59.876130] in_atomic(): 1, irqs_disabled(): 0, pid: 5402, name: git
[ 59.876132] INFO: lockdep is turned off.
[ 59.876134] Preemption disabled at:[<00000000002c4268>] __split_huge_pmd+0x70/0x218
[ 59.876138]
[ 59.876141] CPU: 2 PID: 5402 Comm: git Tainted: G D W 4.4.0-07794-ga4eff16-dirty #77
[ 59.876144] 00000000cfecf610 00000000cfecf6a0 0000000000000002 0000000000000000
00000000cfecf740 00000000cfecf6b8 00000000cfecf6b8 0000000000113402
0000000000000000 000000000089ab4e 00000000008b0a84 0704d0010000000b
00000000cfecf700 00000000cfecf6a0 0000000000000000 0000000000000000
0000000000000000 0000000000113402 00000000cfecf6a0 00000000cfecf700
[ 59.876176] Call Trace:
[ 59.876182] ([<000000000011330e>] show_trace+0x126/0x148)
[ 59.876185] [<00000000001133b8>] show_stack+0x88/0xe8
[ 59.876189] [<000000000045549a>] dump_stack+0x7a/0xd8
[ 59.876193] [<00000000001666c6>] ___might_sleep+0x236/0x248
[ 59.876198] [<000000000014a314>] exit_signals+0x3c/0x158
[ 59.876202] [<000000000013a4e0>] do_exit+0x140/0xd18
[ 59.876206] [<00000000001137c4>] die+0x164/0x170
[ 59.876209] [<0000000000100ac6>] do_report_trap+0x14e/0x160
[ 59.876211] [<0000000000100c94>] illegal_op+0x134/0x148
[ 59.876214] [<00000000006fa26c>] pgm_check_handler+0x15c/0x1b4
[ 59.876217] [<00000000002bf3aa>] __split_huge_pmd_locked+0x562/0xa10
[ 59.876221] ([<00000000002bf234>] __split_huge_pmd_locked+0x3ec/0xa10)
[ 59.876223] [<00000000002c4310>] __split_huge_pmd+0x118/0x218
[ 59.876226] [<00000000002810e8>] unmap_single_vma+0x2d8/0xb40
[ 59.876229] [<0000000000282d66>] zap_page_range+0x116/0x318
[ 59.876232] [<000000000029b834>] SyS_madvise+0x23c/0x5e8
[ 59.876235] [<00000000006f9f56>] system_call+0xd6/0x258
[ 59.876238] [<000003ff9bbfd282>] 0x3ff9bbfd282
[ 59.876240] INFO: lockdep is turned off.
[ 59.876243] note: git[5402] exited with preempt_count 1
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2016-02-15 16:50 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r2tzs-2WJ-35@gated-at.bofh.it> |
| In reply to | #1333309 |
On Sat, Feb 13, 2016 at 12:58:31PM +0100, Sebastian Ott wrote: > > On Sat, 13 Feb 2016, Kirill A. Shutemov wrote: > > Could you check if revert of fecffad25458 helps? > > I reverted fecffad25458 on top of 721675fcf277cf - it oopsed with: > > ¢ 1851.721062! Unable to handle kernel pointer dereference in virtual kernel address space > ¢ 1851.721075! failing address: 0000000000000000 TEID: 0000000000000483 > ¢ 1851.721078! Fault in home space mode while using kernel ASCE. > ¢ 1851.721085! AS:0000000000d5c007 R3:00000000ffff0007 S:00000000ffffa800 P:000000000000003d > ¢ 1851.721128! Oops: 0004 ilc:3 ¢#1! PREEMPT SMP DEBUG_PAGEALLOC > ¢ 1851.721135! Modules linked in: bridge stp llc btrfs mlx4_ib mlx4_en ib_sa ib_mad vxlan xor ip6_udp_tunnel ib_core udp_tunnel ptp pps_core ib_addr ghash_s390raid6_pq prng ecb aes_s390 mlx4_core des_s390 des_generic genwqe_card sha512_s390 sha256_s390 sha1_s390 sha_common crc_itu_t dm_mod scm_block vhost_net tun vhost eadm_sch macvtap macvlan kvm autofs4 > ¢ 1851.721183! CPU: 7 PID: 256422 Comm: bash Not tainted 4.5.0-rc3-00058-g07923d7-dirty #178 > ¢ 1851.721186! task: 000000007fbfd290 ti: 000000008c604000 task.ti: 000000008c604000 > ¢ 1851.721189! Krnl PSW : 0704d00180000000 000000000045d3b8 (__rb_erase_color+0x280/0x308) > ¢ 1851.721200! R:0 T:1 IO:1 EX:1 Key:0 M:1 W:0 P:0 AS:3 CC:1 PM:0 EA:3 > Krnl GPRS: 0000000000000001 0000000000000020 0000000000000000 00000000bd07eff1 > ¢ 1851.721205! 000000000027ca10 0000000000000000 0000000083e45898 0000000077b61198 > ¢ 1851.721207! 000000007ce1a490 00000000bd07eff0 000000007ce1a548 000000000027ca10 > ¢ 1851.721210! 00000000bd07c350 00000000bd07eff0 000000008c607aa8 000000008c607a68 > ¢ 1851.721221! Krnl Code: 000000000045d3aa: e3c0d0080024 stg %%r12,8(%%r13) > 000000000045d3b0: b9040039 lgr %%r3,%%r9 > #000000000045d3b4: a53b0001 oill %%r3,1 > >000000000045d3b8: e33010000024 stg %%r3,0(%%r1) > 000000000045d3be: ec28000e007c cgij %%r2,0,8,45d3da > 000000000045d3c4: e34020000004 lg %%r4,0(%%r2) > 000000000045d3ca: b904001c lgr %%r1,%%r12 > 000000000045d3ce: ec143f3f0056 rosbg %%r1,%%r4,63,63,0 > ¢ 1851.721269! Call Trace: > ¢ 1851.721273! (¢<0000000083e45898>! 0x83e45898) > ¢ 1851.721279! ¢<000000000029342a>! unlink_anon_vmas+0x9a/0x1d8 > ¢ 1851.721282! ¢<0000000000283f34>! free_pgtables+0xcc/0x148 > ¢ 1851.721285! ¢<000000000028c376>! exit_mmap+0xd6/0x300 > ¢ 1851.721289! ¢<0000000000134db8>! mmput+0x90/0x118 > ¢ 1851.721294! ¢<00000000002d76bc>! flush_old_exec+0x5d4/0x700 > ¢ 1851.721298! ¢<00000000003369f4>! load_elf_binary+0x2f4/0x13e8 > ¢ 1851.721301! ¢<00000000002d6e4a>! search_binary_handler+0x9a/0x1f8 > ¢ 1851.721304! ¢<00000000002d8970>! do_execveat_common.isra.32+0x668/0x9a0 > ¢ 1851.721307! ¢<00000000002d8cec>! do_execve+0x44/0x58 > ¢ 1851.721310! ¢<00000000002d8f92>! SyS_execve+0x3a/0x48 > ¢ 1851.721315! ¢<00000000006fb096>! system_call+0xd6/0x258 > ¢ 1851.721317! ¢<000003ff997436d6>! 0x3ff997436d6 > ¢ 1851.721319! INFO: lockdep is turned off. > ¢ 1851.721321! Last Breaking-Event-Address: > ¢ 1851.721323! ¢<000000000045d31a>! __rb_erase_color+0x1e2/0x308 > ¢ 1851.721327! > ¢ 1851.721329! ---¢ end trace 0d80041ac00cfae2 !--- > > > > > > And could you share how crashes looks like? I haven't seen backtraces yet. > > > > Sure. I didn't because they really looked random to me. Most of the time > in rcu or list debugging but I thought these have just been the messenger > observing a corruption first. Anyhow, here is an older one that might look > interesting: > > [ 59.851421] list_del corruption. next->prev should be 000000006e1eb000, but was 0000000000000400 This kinda interesting: 0x400 is TAIL_MAPPING.. Hm.. Could you check if you see the problem on commit 1c290f642101 and its immediate parent? -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
| From | Sebastian Ott <sebott@linux.vnet.ibm.com> |
|---|---|
| Date | 2016-02-15 17:40 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r2ulR-3v3-37@gated-at.bofh.it> |
| In reply to | #1334553 |
On Mon, 15 Feb 2016, Kirill A. Shutemov wrote: > > [ 59.851421] list_del corruption. next->prev should be 000000006e1eb000, but was 0000000000000400 > > This kinda interesting: 0x400 is TAIL_MAPPING.. Hm.. > > Could you check if you see the problem on commit 1c290f642101 and its > immediate parent? Both 1c290f642101 and 1c290f642101^ survived 20 compile runs each. Sebastian
[toc] | [prev] | [next] | [standalone]
| From | Gerald Schaefer <gerald.schaefer@de.ibm.com> |
|---|---|
| Date | 2016-02-15 19:40 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r2wdY-4Rn-7@gated-at.bofh.it> |
| In reply to | #1334553 |
On Mon, 15 Feb 2016 13:31:59 +0200 "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > On Sat, Feb 13, 2016 at 12:58:31PM +0100, Sebastian Ott wrote: > > > > On Sat, 13 Feb 2016, Kirill A. Shutemov wrote: > > > Could you check if revert of fecffad25458 helps? > > > > I reverted fecffad25458 on top of 721675fcf277cf - it oopsed with: > > > > ¢ 1851.721062! Unable to handle kernel pointer dereference in virtual kernel address space > > ¢ 1851.721075! failing address: 0000000000000000 TEID: 0000000000000483 > > ¢ 1851.721078! Fault in home space mode while using kernel ASCE. > > ¢ 1851.721085! AS:0000000000d5c007 R3:00000000ffff0007 S:00000000ffffa800 P:000000000000003d > > ¢ 1851.721128! Oops: 0004 ilc:3 ¢#1! PREEMPT SMP DEBUG_PAGEALLOC > > ¢ 1851.721135! Modules linked in: bridge stp llc btrfs mlx4_ib mlx4_en ib_sa ib_mad vxlan xor ip6_udp_tunnel ib_core udp_tunnel ptp pps_core ib_addr ghash_s390raid6_pq prng ecb aes_s390 mlx4_core des_s390 des_generic genwqe_card sha512_s390 sha256_s390 sha1_s390 sha_common crc_itu_t dm_mod scm_block vhost_net tun vhost eadm_sch macvtap macvlan kvm autofs4 > > ¢ 1851.721183! CPU: 7 PID: 256422 Comm: bash Not tainted 4.5.0-rc3-00058-g07923d7-dirty #178 > > ¢ 1851.721186! task: 000000007fbfd290 ti: 000000008c604000 task.ti: 000000008c604000 > > ¢ 1851.721189! Krnl PSW : 0704d00180000000 000000000045d3b8 (__rb_erase_color+0x280/0x308) > > ¢ 1851.721200! R:0 T:1 IO:1 EX:1 Key:0 M:1 W:0 P:0 AS:3 CC:1 PM:0 EA:3 > > Krnl GPRS: 0000000000000001 0000000000000020 0000000000000000 00000000bd07eff1 > > ¢ 1851.721205! 000000000027ca10 0000000000000000 0000000083e45898 0000000077b61198 > > ¢ 1851.721207! 000000007ce1a490 00000000bd07eff0 000000007ce1a548 000000000027ca10 > > ¢ 1851.721210! 00000000bd07c350 00000000bd07eff0 000000008c607aa8 000000008c607a68 > > ¢ 1851.721221! Krnl Code: 000000000045d3aa: e3c0d0080024 stg %%r12,8(%%r13) > > 000000000045d3b0: b9040039 lgr %%r3,%%r9 > > #000000000045d3b4: a53b0001 oill %%r3,1 > > >000000000045d3b8: e33010000024 stg %%r3,0(%%r1) > > 000000000045d3be: ec28000e007c cgij %%r2,0,8,45d3da > > 000000000045d3c4: e34020000004 lg %%r4,0(%%r2) > > 000000000045d3ca: b904001c lgr %%r1,%%r12 > > 000000000045d3ce: ec143f3f0056 rosbg %%r1,%%r4,63,63,0 > > ¢ 1851.721269! Call Trace: > > ¢ 1851.721273! (¢<0000000083e45898>! 0x83e45898) > > ¢ 1851.721279! ¢<000000000029342a>! unlink_anon_vmas+0x9a/0x1d8 > > ¢ 1851.721282! ¢<0000000000283f34>! free_pgtables+0xcc/0x148 > > ¢ 1851.721285! ¢<000000000028c376>! exit_mmap+0xd6/0x300 > > ¢ 1851.721289! ¢<0000000000134db8>! mmput+0x90/0x118 > > ¢ 1851.721294! ¢<00000000002d76bc>! flush_old_exec+0x5d4/0x700 > > ¢ 1851.721298! ¢<00000000003369f4>! load_elf_binary+0x2f4/0x13e8 > > ¢ 1851.721301! ¢<00000000002d6e4a>! search_binary_handler+0x9a/0x1f8 > > ¢ 1851.721304! ¢<00000000002d8970>! do_execveat_common.isra.32+0x668/0x9a0 > > ¢ 1851.721307! ¢<00000000002d8cec>! do_execve+0x44/0x58 > > ¢ 1851.721310! ¢<00000000002d8f92>! SyS_execve+0x3a/0x48 > > ¢ 1851.721315! ¢<00000000006fb096>! system_call+0xd6/0x258 > > ¢ 1851.721317! ¢<000003ff997436d6>! 0x3ff997436d6 > > ¢ 1851.721319! INFO: lockdep is turned off. > > ¢ 1851.721321! Last Breaking-Event-Address: > > ¢ 1851.721323! ¢<000000000045d31a>! __rb_erase_color+0x1e2/0x308 > > ¢ 1851.721327! > > ¢ 1851.721329! ---¢ end trace 0d80041ac00cfae2 !--- > > > > > > > > > > And could you share how crashes looks like? I haven't seen backtraces yet. > > > > > > > Sure. I didn't because they really looked random to me. Most of the time > > in rcu or list debugging but I thought these have just been the messenger > > observing a corruption first. Anyhow, here is an older one that might look > > interesting: > > > > [ 59.851421] list_del corruption. next->prev should be 000000006e1eb000, but was 0000000000000400 > > This kinda interesting: 0x400 is TAIL_MAPPING.. Hm.. > > Could you check if you see the problem on commit 1c290f642101 and its > immediate parent? > How should the page->mapping poison end up as next->prev in the list of pre-allocated THP splitting page tables? Also, commit 1c290f642101 is before the THP rework, at least the non-bisectable part, so we should expect not to see the problem there. 0x400 is also the value of an empty pte on s390, and the thp_deposit/withdraw listheads are placed inside the pre-allocated pagetables instead of page->lru, because we have 2K pagetables on s390 and cannot use struct page == pgtable_t. So, for example, two concurrent withdraws could produce such a list corruption, because the first withdraw will overwrite the listhead at the beginning of the pagetable with 2 empty ptes. Has anything changed regarding the general THP deposit/withdraw logic?
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2016-02-16 00:30 +0100 |
| Subject | Re: [BUG] random kernel crashes after THP rework on s390 (maybe also on PowerPC and ARM) |
| Message-ID | <r2AKC-7Tu-5@gated-at.bofh.it> |
| In reply to | #1334710 |
On Mon, Feb 15, 2016 at 07:37:02PM +0100, Gerald Schaefer wrote: > On Mon, 15 Feb 2016 13:31:59 +0200 > "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > > > On Sat, Feb 13, 2016 at 12:58:31PM +0100, Sebastian Ott wrote: > > > > > > On Sat, 13 Feb 2016, Kirill A. Shutemov wrote: > > > > Could you check if revert of fecffad25458 helps? > > > > > > I reverted fecffad25458 on top of 721675fcf277cf - it oopsed with: > > > > > > ¢ 1851.721062! Unable to handle kernel pointer dereference in virtual kernel address space > > > ¢ 1851.721075! failing address: 0000000000000000 TEID: 0000000000000483 > > > ¢ 1851.721078! Fault in home space mode while using kernel ASCE. > > > ¢ 1851.721085! AS:0000000000d5c007 R3:00000000ffff0007 S:00000000ffffa800 P:000000000000003d > > > ¢ 1851.721128! Oops: 0004 ilc:3 ¢#1! PREEMPT SMP DEBUG_PAGEALLOC > > > ¢ 1851.721135! Modules linked in: bridge stp llc btrfs mlx4_ib mlx4_en ib_sa ib_mad vxlan xor ip6_udp_tunnel ib_core udp_tunnel ptp pps_core ib_addr ghash_s390raid6_pq prng ecb aes_s390 mlx4_core des_s390 des_generic genwqe_card sha512_s390 sha256_s390 sha1_s390 sha_common crc_itu_t dm_mod scm_block vhost_net tun vhost eadm_sch macvtap macvlan kvm autofs4 > > > ¢ 1851.721183! CPU: 7 PID: 256422 Comm: bash Not tainted 4.5.0-rc3-00058-g07923d7-dirty #178 > > > ¢ 1851.721186! task: 000000007fbfd290 ti: 000000008c604000 task.ti: 000000008c604000 > > > ¢ 1851.721189! Krnl PSW : 0704d00180000000 000000000045d3b8 (__rb_erase_color+0x280/0x308) > > > ¢ 1851.721200! R:0 T:1 IO:1 EX:1 Key:0 M:1 W:0 P:0 AS:3 CC:1 PM:0 EA:3 > > > Krnl GPRS: 0000000000000001 0000000000000020 0000000000000000 00000000bd07eff1 > > > ¢ 1851.721205! 000000000027ca10 0000000000000000 0000000083e45898 0000000077b61198 > > > ¢ 1851.721207! 000000007ce1a490 00000000bd07eff0 000000007ce1a548 000000000027ca10 > > > ¢ 1851.721210! 00000000bd07c350 00000000bd07eff0 000000008c607aa8 000000008c607a68 > > > ¢ 1851.721221! Krnl Code: 000000000045d3aa: e3c0d0080024 stg %%r12,8(%%r13) > > > 000000000045d3b0: b9040039 lgr %%r3,%%r9 > > > #000000000045d3b4: a53b0001 oill %%r3,1 > > > >000000000045d3b8: e33010000024 stg %%r3,0(%%r1) > > > 000000000045d3be: ec28000e007c cgij %%r2,0,8,45d3da > > > 000000000045d3c4: e34020000004 lg %%r4,0(%%r2) > > > 000000000045d3ca: b904001c lgr %%r1,%%r12 > > > 000000000045d3ce: ec143f3f0056 rosbg %%r1,%%r4,63,63,0 > > > ¢ 1851.721269! Call Trace: > > > ¢ 1851.721273! (¢<0000000083e45898>! 0x83e45898) > > > ¢ 1851.721279! ¢<000000000029342a>! unlink_anon_vmas+0x9a/0x1d8 > > > ¢ 1851.721282! ¢<0000000000283f34>! free_pgtables+0xcc/0x148 > > > ¢ 1851.721285! ¢<000000000028c376>! exit_mmap+0xd6/0x300 > > > ¢ 1851.721289! ¢<0000000000134db8>! mmput+0x90/0x118 > > > ¢ 1851.721294! ¢<00000000002d76bc>! flush_old_exec+0x5d4/0x700 > > > ¢ 1851.721298! ¢<00000000003369f4>! load_elf_binary+0x2f4/0x13e8 > > > ¢ 1851.721301! ¢<00000000002d6e4a>! search_binary_handler+0x9a/0x1f8 > > > ¢ 1851.721304! ¢<00000000002d8970>! do_execveat_common.isra.32+0x668/0x9a0 > > > ¢ 1851.721307! ¢<00000000002d8cec>! do_execve+0x44/0x58 > > > ¢ 1851.721310! ¢<00000000002d8f92>! SyS_execve+0x3a/0x48 > > > ¢ 1851.721315! ¢<00000000006fb096>! system_call+0xd6/0x258 > > > ¢ 1851.721317! ¢<000003ff997436d6>! 0x3ff997436d6 > > > ¢ 1851.721319! INFO: lockdep is turned off. > > > ¢ 1851.721321! Last Breaking-Event-Address: > > > ¢ 1851.721323! ¢<000000000045d31a>! __rb_erase_color+0x1e2/0x308 > > > ¢ 1851.721327! > > > ¢ 1851.721329! ---¢ end trace 0d80041ac00cfae2 !--- > > > > > > > > > > > > > > And could you share how crashes looks like? I haven't seen backtraces yet. > > > > > > > > > > Sure. I didn't because they really looked random to me. Most of the time > > > in rcu or list debugging but I thought these have just been the messenger > > > observing a corruption first. Anyhow, here is an older one that might look > > > interesting: > > > > > > [ 59.851421] list_del corruption. next->prev should be 000000006e1eb000, but was 0000000000000400 > > > > This kinda interesting: 0x400 is TAIL_MAPPING.. Hm.. > > > > Could you check if you see the problem on commit 1c290f642101 and its > > immediate parent? > > > > How should the page->mapping poison end up as next->prev in the list of > pre-allocated THP splitting page tables? May be pgtable was casted to struct page or something. I don't know. > Also, commit 1c290f642101 is before the THP rework, at least the > non-bisectable part, so we should expect not to see the problem there. Just to make sure: commit 122afea9626a is fine, commit 61f5d698cc97 crashes. Correct? > 0x400 is also the value of an empty pte on s390, and the thp_deposit/withdraw > listheads are placed inside the pre-allocated pagetables instead of page->lru, > because we have 2K pagetables on s390 and cannot use struct page == pgtable_t. 0x400 from empty pte makes more sense than TAIL_MAPPING. But I guess it worth changing TAIL_MAPPING to some other value to make sure. > So, for example, two concurrent withdraws could produce such a list > corruption, because the first withdraw will overwrite the listhead at the > beginning of the pagetable with 2 empty ptes. > > Has anything changed regarding the general THP deposit/withdraw logic? I don't see any changes in this area. To eliminate one more variable, I would propose to disable split pmd lock for testing and check if it makes difference. Is there any chance that I'll be able to trigger the bug using QEMU? Does anybody have an QEMU image I can use? -- Kirill A. Shutemov
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web