Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1209737 > unrolled thread
| Started by | "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> |
|---|---|
| First post | 2015-08-19 11:30 +0200 |
| Last post | 2015-08-24 11:40 +0200 |
| Articles | 20 on this page of 37 — 9 participants |
Back to article view | Back to linux.kernel
[PATCHv3 0/5] Fix compound_head() race "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2015-08-19 11:30 +0200
[PATCHv3 4/5] mm: make compound_head() robust "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2015-08-19 11:30 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Andrew Morton <akpm@linux-foundation.org> - 2015-08-21 01:40 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Kirill A. Shutemov" <kirill@shutemov.name> - 2015-08-21 14:20 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Christoph Lameter <cl@linux.com> - 2015-08-21 18:20 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Kirill A. Shutemov" <kirill@shutemov.name> - 2015-08-21 21:40 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Andrew Morton <akpm@linux-foundation.org> - 2015-08-21 21:40 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Christoph Lameter <cl@linux.com> - 2015-08-21 23:20 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Vlastimil Babka <vbabka@suse.cz> - 2015-08-24 17:50 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Vlastimil Babka <vbabka@suse.cz> - 2015-08-25 13:50 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Kirill A. Shutemov" <kirill@shutemov.name> - 2015-08-25 20:40 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2015-08-25 22:20 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Vlastimil Babka <vbabka@suse.cz> - 2015-08-25 22:50 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2015-08-25 23:30 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Kirill A. Shutemov" <kirill@shutemov.name> - 2015-08-26 17:10 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Vlastimil Babka <vbabka@suse.cz> - 2015-08-26 17:40 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2015-08-26 18:40 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Hugh Dickins <hughd@google.com> - 2015-08-26 20:20 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2015-08-26 23:30 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Hugh Dickins <hughd@google.com> - 2015-08-27 00:30 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2015-08-27 01:40 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Michal Hocko <mhocko@kernel.org> - 2015-08-27 17:10 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Michal Hocko <mhocko@kernel.org> - 2015-08-27 18:10 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Hugh Dickins <hughd@google.com> - 2015-08-27 19:30 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Michal Hocko <mhocko@kernel.org> - 2015-08-27 20:10 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2015-08-27 18:40 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Michal Hocko <mhocko@kernel.org> - 2015-08-27 20:20 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> - 2015-08-27 21:20 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust Jesper Dangaard Brouer <netdev@brouer.com> - 2015-08-24 02:20 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Kirill A. Shutemov" <kirill@shutemov.name> - 2015-08-24 11:40 +0200
Re: [PATCHv3 4/5] mm: make compound_head() robust "Kirill A. Shutemov" <kirill@shutemov.name> - 2015-08-24 12:20 +0200
[PATCHv3 5/5] mm: use 'unsigned int' for page order "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> - 2015-08-19 11:30 +0200
Re: [PATCHv3 5/5] mm: use 'unsigned int' for page order Michal Hocko <mhocko@kernel.org> - 2015-08-20 10:40 +0200
Re: [PATCHv3 0/5] Fix compound_head() race "Kirill A. Shutemov" <kirill@shutemov.name> - 2015-08-20 14:40 +0200
Re: [PATCHv3 0/5] Fix compound_head() race Andrew Morton <akpm@linux-foundation.org> - 2015-08-21 01:40 +0200
Re: [PATCHv3 0/5] Fix compound_head() race Hugh Dickins <hughd@google.com> - 2015-08-22 22:20 +0200
Re: [PATCHv3 0/5] Fix compound_head() race "Kirill A. Shutemov" <kirill@shutemov.name> - 2015-08-24 11:40 +0200
Page 1 of 2 [1] 2 Next page →
| From | "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> |
|---|---|
| Date | 2015-08-19 11:30 +0200 |
| Subject | [PATCHv3 0/5] Fix compound_head() race |
| Message-ID | <pZ7QZ-1Mc-13@gated-at.bofh.it> |
Here's my attempt on fixing recently discovered race in compound_head().
It should make compound_head() reliable in all contexts.
The patchset is against Linus' tree. Let me know if it need to be rebased
onto different baseline.
It's expected to have conflicts with my page-flags patchset and probably
should be applied before it.
v3:
- Fix build without hugetlb;
- Drop page->first_page;
- Update comment for free_compound_page();
- Use 'unsigned int' for page order;
v2: Per Hugh's suggestion page->compound_head is moved into third double
word. This way we can avoid memory overhead which v1 had in some
cases.
This place in struct page is rather overloaded. More testing is
required to make sure we don't collide with anyone.
Kirill A. Shutemov (5):
mm: drop page->slab_page
zsmalloc: use page->private instead of page->first_page
mm: pack compound_dtor and compound_order into one word in struct page
mm: make compound_head() robust
mm: use 'unsigned int' for page order
Documentation/vm/split_page_table_lock | 4 +-
arch/xtensa/configs/iss_defconfig | 1 -
include/linux/mm.h | 82 +++++++++++-----------------------
include/linux/mm_types.h | 21 ++++++---
include/linux/page-flags.h | 80 ++++++++-------------------------
mm/Kconfig | 12 -----
mm/debug.c | 5 ---
mm/huge_memory.c | 3 +-
mm/hugetlb.c | 35 +++++++--------
mm/internal.h | 8 ++--
mm/memory-failure.c | 7 ---
mm/page_alloc.c | 76 ++++++++++++++++++-------------
mm/swap.c | 4 +-
mm/zsmalloc.c | 11 +++--
14 files changed, 133 insertions(+), 216 deletions(-)
--
2.5.0
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> |
|---|---|
| Date | 2015-08-19 11:30 +0200 |
| Subject | [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <pZ7R0-1Mc-27@gated-at.bofh.it> |
| In reply to | #1209737 |
Hugh has pointed that compound_head() call can be unsafe in some
context. There's one example:
CPU0 CPU1
isolate_migratepages_block()
page_count()
compound_head()
!!PageTail() == true
put_page()
tail->first_page = NULL
head = tail->first_page
alloc_pages(__GFP_COMP)
prep_compound_page()
tail->first_page = head
__SetPageTail(p);
!!PageTail() == true
<head == NULL dereferencing>
The race is pure theoretical. I don't it's possible to trigger it in
practice. But who knows.
We can fix the race by changing how encode PageTail() and compound_head()
within struct page to be able to update them in one shot.
The patch introduces page->compound_head into third double word block in
front of compound_dtor and compound_order. That means it shares storage
space with:
- page->lru.next;
- page->next;
- page->rcu_head.next;
- page->pmd_huge_pte;
That's too long list to be absolutely sure, but looks like nobody uses
bit 0 of the word. It can be used to encode PageTail(). And if the bit
set, rest of the word is pointer to head page.
Signed-off-by: Kirill A. Shutemov <kirill.shutemov@linux.intel.com>
Acked-by: Michal Hocko <mhocko@suse.com>
Cc: Hugh Dickins <hughd@google.com>
Cc: David Rientjes <rientjes@google.com>
Cc: Vlastimil Babka <vbabka@suse.cz>
---
Documentation/vm/split_page_table_lock | 4 +-
arch/xtensa/configs/iss_defconfig | 1 -
include/linux/mm.h | 53 ++--------------------
include/linux/mm_types.h | 9 +++-
include/linux/page-flags.h | 80 ++++++++--------------------------
mm/Kconfig | 12 -----
mm/debug.c | 5 ---
mm/huge_memory.c | 3 +-
mm/hugetlb.c | 8 +---
mm/internal.h | 4 +-
mm/memory-failure.c | 7 ---
mm/page_alloc.c | 38 ++++++++--------
mm/swap.c | 4 +-
13 files changed, 58 insertions(+), 170 deletions(-)
diff --git a/Documentation/vm/split_page_table_lock b/Documentation/vm/split_page_table_lock
index 6dea4fd5c961..62842a857dab 100644
--- a/Documentation/vm/split_page_table_lock
+++ b/Documentation/vm/split_page_table_lock
@@ -54,8 +54,8 @@ everything required is done by pgtable_page_ctor() and pgtable_page_dtor(),
which must be called on PTE table allocation / freeing.
Make sure the architecture doesn't use slab allocator for page table
-allocation: slab uses page->slab_cache and page->first_page for its pages.
-These fields share storage with page->ptl.
+allocation: slab uses page->slab_cache for its pages.
+This field shares storage with page->ptl.
PMD split lock only makes sense if you have more than two page table
levels.
diff --git a/arch/xtensa/configs/iss_defconfig b/arch/xtensa/configs/iss_defconfig
index e4d193e7a300..5c7c385f21c4 100644
--- a/arch/xtensa/configs/iss_defconfig
+++ b/arch/xtensa/configs/iss_defconfig
@@ -169,7 +169,6 @@ CONFIG_FLATMEM_MANUAL=y
# CONFIG_SPARSEMEM_MANUAL is not set
CONFIG_FLATMEM=y
CONFIG_FLAT_NODE_MEM_MAP=y
-CONFIG_PAGEFLAGS_EXTENDED=y
CONFIG_SPLIT_PTLOCK_CPUS=4
# CONFIG_PHYS_ADDR_T_64BIT is not set
CONFIG_ZONE_DMA_FLAG=1
diff --git a/include/linux/mm.h b/include/linux/mm.h
index 0735bc0a351a..a4c4b7d07473 100644
--- a/include/linux/mm.h
+++ b/include/linux/mm.h
@@ -437,46 +437,6 @@ static inline void compound_unlock_irqrestore(struct page *page,
#endif
}
-static inline struct page *compound_head_by_tail(struct page *tail)
-{
- struct page *head = tail->first_page;
-
- /*
- * page->first_page may be a dangling pointer to an old
- * compound page, so recheck that it is still a tail
- * page before returning.
- */
- smp_rmb();
- if (likely(PageTail(tail)))
- return head;
- return tail;
-}
-
-/*
- * Since either compound page could be dismantled asynchronously in THP
- * or we access asynchronously arbitrary positioned struct page, there
- * would be tail flag race. To handle this race, we should call
- * smp_rmb() before checking tail flag. compound_head_by_tail() did it.
- */
-static inline struct page *compound_head(struct page *page)
-{
- if (unlikely(PageTail(page)))
- return compound_head_by_tail(page);
- return page;
-}
-
-/*
- * If we access compound page synchronously such as access to
- * allocated page, there is no need to handle tail flag race, so we can
- * check tail flag directly without any synchronization primitive.
- */
-static inline struct page *compound_head_fast(struct page *page)
-{
- if (unlikely(PageTail(page)))
- return page->first_page;
- return page;
-}
-
/*
* The atomic page->_mapcount, starts from -1: so that transitions
* both from it and to it can be tracked, using atomic_inc_and_test
@@ -525,7 +485,7 @@ static inline void get_huge_page_tail(struct page *page)
VM_BUG_ON_PAGE(!PageTail(page), page);
VM_BUG_ON_PAGE(page_mapcount(page) < 0, page);
VM_BUG_ON_PAGE(atomic_read(&page->_count) != 0, page);
- if (compound_tail_refcounted(page->first_page))
+ if (compound_tail_refcounted(compound_head(page)))
atomic_inc(&page->_mapcount);
}
@@ -548,13 +508,7 @@ static inline struct page *virt_to_head_page(const void *x)
{
struct page *page = virt_to_page(x);
- /*
- * We don't need to worry about synchronization of tail flag
- * when we call virt_to_head_page() since it is only called for
- * already allocated page and this page won't be freed until
- * this virt_to_head_page() is finished. So use _fast variant.
- */
- return compound_head_fast(page);
+ return compound_head(page);
}
/*
@@ -1496,8 +1450,7 @@ static inline bool ptlock_init(struct page *page)
* with 0. Make sure nobody took it in use in between.
*
* It can happen if arch try to use slab for page table allocation:
- * slab code uses page->slab_cache and page->first_page (for tail
- * pages), which share storage with page->ptl.
+ * slab code uses page->slab_cache, which share storage with page->ptl.
*/
VM_BUG_ON_PAGE(*(unsigned long *)&page->ptl, page);
if (!ptlock_alloc(page))
diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h
index 63cdfe7ec336..e324768b6cc7 100644
--- a/include/linux/mm_types.h
+++ b/include/linux/mm_types.h
@@ -120,7 +120,12 @@ struct page {
};
};
- /* Third double word block */
+ /*
+ * Third double word block
+ *
+ * WARNING: bit 0 of the first word encode PageTail and *must* be 0
+ * for non-tail pages.
+ */
union {
struct list_head lru; /* Pageout list, eg. active_list
* protected by zone->lru_lock !
@@ -143,6 +148,7 @@ struct page {
*/
/* First tail page of compound page */
struct {
+ unsigned long compound_head; /* If bit zero is set */
#ifdef CONFIG_64BIT
unsigned int compound_dtor;
unsigned int compound_order;
@@ -174,7 +180,6 @@ struct page {
#endif
#endif
struct kmem_cache *slab_cache; /* SL[AU]B: Pointer to slab */
- struct page *first_page; /* Compound tail pages */
};
#ifdef CONFIG_MEMCG
diff --git a/include/linux/page-flags.h b/include/linux/page-flags.h
index 41c93844fb1d..9b865158e452 100644
--- a/include/linux/page-flags.h
+++ b/include/linux/page-flags.h
@@ -86,12 +86,7 @@ enum pageflags {
PG_private, /* If pagecache, has fs-private data */
PG_private_2, /* If pagecache, has fs aux data */
PG_writeback, /* Page is under writeback */
-#ifdef CONFIG_PAGEFLAGS_EXTENDED
PG_head, /* A head page */
- PG_tail, /* A tail page */
-#else
- PG_compound, /* A compound page */
-#endif
PG_swapcache, /* Swap page: swp_entry_t in private */
PG_mappedtodisk, /* Has blocks allocated on-disk */
PG_reclaim, /* To be reclaimed asap */
@@ -387,85 +382,46 @@ static inline void set_page_writeback_keepwrite(struct page *page)
test_set_page_writeback_keepwrite(page);
}
-#ifdef CONFIG_PAGEFLAGS_EXTENDED
-/*
- * System with lots of page flags available. This allows separate
- * flags for PageHead() and PageTail() checks of compound pages so that bit
- * tests can be used in performance sensitive paths. PageCompound is
- * generally not used in hot code paths except arch/powerpc/mm/init_64.c
- * and arch/powerpc/kvm/book3s_64_vio_hv.c which use it to detect huge pages
- * and avoid handling those in real mode.
- */
__PAGEFLAG(Head, head) CLEARPAGEFLAG(Head, head)
-__PAGEFLAG(Tail, tail)
-static inline int PageCompound(struct page *page)
-{
- return page->flags & ((1L << PG_head) | (1L << PG_tail));
-
-}
-#ifdef CONFIG_TRANSPARENT_HUGEPAGE
-static inline void ClearPageCompound(struct page *page)
+static inline int PageTail(struct page *page)
{
- BUG_ON(!PageHead(page));
- ClearPageHead(page);
+ return READ_ONCE(page->compound_head) & 1;
}
-#endif
-
-#define PG_head_mask ((1L << PG_head))
-#else
-/*
- * Reduce page flag use as much as possible by overlapping
- * compound page flags with the flags used for page cache pages. Possible
- * because PageCompound is always set for compound pages and not for
- * pages on the LRU and/or pagecache.
- */
-TESTPAGEFLAG(Compound, compound)
-__SETPAGEFLAG(Head, compound) __CLEARPAGEFLAG(Head, compound)
-
-/*
- * PG_reclaim is used in combination with PG_compound to mark the
- * head and tail of a compound page. This saves one page flag
- * but makes it impossible to use compound pages for the page cache.
- * The PG_reclaim bit would have to be used for reclaim or readahead
- * if compound pages enter the page cache.
- *
- * PG_compound & PG_reclaim => Tail page
- * PG_compound & ~PG_reclaim => Head page
- */
-#define PG_head_mask ((1L << PG_compound))
-#define PG_head_tail_mask ((1L << PG_compound) | (1L << PG_reclaim))
-
-static inline int PageHead(struct page *page)
+static inline void set_compound_head(struct page *page, struct page *head)
{
- return ((page->flags & PG_head_tail_mask) == PG_head_mask);
+ WRITE_ONCE(page->compound_head, (unsigned long)head + 1);
}
-static inline int PageTail(struct page *page)
+static inline void clear_compound_head(struct page *page)
{
- return ((page->flags & PG_head_tail_mask) == PG_head_tail_mask);
+ WRITE_ONCE(page->compound_head, 0);
}
-static inline void __SetPageTail(struct page *page)
+static inline struct page *compound_head(struct page *page)
{
- page->flags |= PG_head_tail_mask;
+ unsigned long head = READ_ONCE(page->compound_head);
+
+ if (unlikely(head & 1))
+ return (struct page *) (head - 1);
+ return page;
}
-static inline void __ClearPageTail(struct page *page)
+static inline int PageCompound(struct page *page)
{
- page->flags &= ~PG_head_tail_mask;
-}
+ return PageHead(page) || PageTail(page);
+}
#ifdef CONFIG_TRANSPARENT_HUGEPAGE
static inline void ClearPageCompound(struct page *page)
{
- BUG_ON((page->flags & PG_head_tail_mask) != (1 << PG_compound));
- clear_bit(PG_compound, &page->flags);
+ BUG_ON(!PageHead(page));
+ ClearPageHead(page);
}
#endif
-#endif /* !PAGEFLAGS_EXTENDED */
+#define PG_head_mask ((1L << PG_head))
#ifdef CONFIG_HUGETLB_PAGE
int PageHuge(struct page *page);
diff --git a/mm/Kconfig b/mm/Kconfig
index e79de2bd12cd..454579d31081 100644
--- a/mm/Kconfig
+++ b/mm/Kconfig
@@ -200,18 +200,6 @@ config MEMORY_HOTREMOVE
depends on MEMORY_HOTPLUG && ARCH_ENABLE_MEMORY_HOTREMOVE
depends on MIGRATION
-#
-# If we have space for more page flags then we can enable additional
-# optimizations and functionality.
-#
-# Regular Sparsemem takes page flag bits for the sectionid if it does not
-# use a virtual memmap. Disable extended page flags for 32 bit platforms
-# that require the use of a sectionid in the page flags.
-#
-config PAGEFLAGS_EXTENDED
- def_bool y
- depends on 64BIT || SPARSEMEM_VMEMMAP || !SPARSEMEM
-
# Heavily threaded applications may benefit from splitting the mm-wide
# page_table_lock, so that faults on different parts of the user address
# space can be handled with less contention: split it at this NR_CPUS.
diff --git a/mm/debug.c b/mm/debug.c
index 76089ddf99ea..205e5ef957ab 100644
--- a/mm/debug.c
+++ b/mm/debug.c
@@ -25,12 +25,7 @@ static const struct trace_print_flags pageflag_names[] = {
{1UL << PG_private, "private" },
{1UL << PG_private_2, "private_2" },
{1UL << PG_writeback, "writeback" },
-#ifdef CONFIG_PAGEFLAGS_EXTENDED
{1UL << PG_head, "head" },
- {1UL << PG_tail, "tail" },
-#else
- {1UL << PG_compound, "compound" },
-#endif
{1UL << PG_swapcache, "swapcache" },
{1UL << PG_mappedtodisk, "mappedtodisk" },
{1UL << PG_reclaim, "reclaim" },
diff --git a/mm/huge_memory.c b/mm/huge_memory.c
index 097c7a4bfbd9..330377f83ac7 100644
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -1686,8 +1686,7 @@ static void __split_huge_page_refcount(struct page *page,
(1L << PG_unevictable)));
page_tail->flags |= (1L << PG_dirty);
- /* clear PageTail before overwriting first_page */
- smp_wmb();
+ clear_compound_head(page_tail);
/*
* __split_huge_page_splitting() already set the
diff --git a/mm/hugetlb.c b/mm/hugetlb.c
index 8ea74caa1fa8..53c0709fd87b 100644
--- a/mm/hugetlb.c
+++ b/mm/hugetlb.c
@@ -824,9 +824,8 @@ static void destroy_compound_gigantic_page(struct page *page,
struct page *p = page + 1;
for (i = 1; i < nr_pages; i++, p = mem_map_next(p, page, i)) {
- __ClearPageTail(p);
+ clear_compound_head(p);
set_page_refcounted(p);
- p->first_page = NULL;
}
set_compound_order(page, 0);
@@ -1099,10 +1098,7 @@ static void prep_compound_gigantic_page(struct page *page, unsigned long order)
*/
__ClearPageReserved(p);
set_page_count(p, 0);
- p->first_page = page;
- /* Make sure p->first_page is always valid for PageTail() */
- smp_wmb();
- __SetPageTail(p);
+ set_compound_head(p, page);
}
}
diff --git a/mm/internal.h b/mm/internal.h
index 36b23f1e2ca6..89e21a07080a 100644
--- a/mm/internal.h
+++ b/mm/internal.h
@@ -61,9 +61,9 @@ static inline void __get_page_tail_foll(struct page *page,
* speculative page access (like in
* page_cache_get_speculative()) on tail pages.
*/
- VM_BUG_ON_PAGE(atomic_read(&page->first_page->_count) <= 0, page);
+ VM_BUG_ON_PAGE(atomic_read(&compound_head(page)->_count) <= 0, page);
if (get_page_head)
- atomic_inc(&page->first_page->_count);
+ atomic_inc(&compound_head(page)->_count);
get_huge_page_tail(page);
}
diff --git a/mm/memory-failure.c b/mm/memory-failure.c
index 1f4446a90cef..4d1a5de9653d 100644
--- a/mm/memory-failure.c
+++ b/mm/memory-failure.c
@@ -787,8 +787,6 @@ static int me_huge_page(struct page *p, unsigned long pfn)
#define lru (1UL << PG_lru)
#define swapbacked (1UL << PG_swapbacked)
#define head (1UL << PG_head)
-#define tail (1UL << PG_tail)
-#define compound (1UL << PG_compound)
#define slab (1UL << PG_slab)
#define reserved (1UL << PG_reserved)
@@ -811,12 +809,7 @@ static struct page_state {
*/
{ slab, slab, MF_MSG_SLAB, me_kernel },
-#ifdef CONFIG_PAGEFLAGS_EXTENDED
{ head, head, MF_MSG_HUGE, me_huge_page },
- { tail, tail, MF_MSG_HUGE, me_huge_page },
-#else
- { compound, compound, MF_MSG_HUGE, me_huge_page },
-#endif
{ sc|dirty, sc|dirty, MF_MSG_DIRTY_SWAPCACHE, me_swapcache_dirty },
{ sc|dirty, sc, MF_MSG_CLEAN_SWAPCACHE, me_swapcache_clean },
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index c6733cc3cbce..78859d47aaf4 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -424,15 +424,15 @@ out:
/*
* Higher-order pages are called "compound pages". They are structured thusly:
*
- * The first PAGE_SIZE page is called the "head page".
+ * The first PAGE_SIZE page is called the "head page" and have PG_head set.
*
- * The remaining PAGE_SIZE pages are called "tail pages".
+ * The remaining PAGE_SIZE pages are called "tail pages". PageTail() is encoded
+ * in bit 0 of page->compound_head. The rest of bits is pointer to head page.
*
- * All pages have PG_compound set. All tail pages have their ->first_page
- * pointing at the head page.
+ * The first tail page's ->compound_dtor holds the offset in array of compound
+ * page destructors. See compound_page_dtors.
*
- * The first tail page's ->lru.next holds the address of the compound page's
- * put_page() function. Its ->lru.prev holds the order of allocation.
+ * The first tail page's ->compound_order holds the order of allocation.
* This usage means that zero-order pages may not be compound.
*/
@@ -452,10 +452,7 @@ void prep_compound_page(struct page *page, unsigned long order)
for (i = 1; i < nr_pages; i++) {
struct page *p = page + i;
set_page_count(p, 0);
- p->first_page = page;
- /* Make sure p->first_page is always valid for PageTail() */
- smp_wmb();
- __SetPageTail(p);
+ set_compound_head(p, page);
}
}
@@ -830,17 +827,24 @@ static void free_one_page(struct zone *zone,
static int free_tail_pages_check(struct page *head_page, struct page *page)
{
- if (!IS_ENABLED(CONFIG_DEBUG_VM))
- return 0;
+ int ret = 1;
+
+ if (!IS_ENABLED(CONFIG_DEBUG_VM)) {
+ ret = 0;
+ goto out;
+ }
if (unlikely(!PageTail(page))) {
bad_page(page, "PageTail not set", 0);
- return 1;
+ goto out;
}
- if (unlikely(page->first_page != head_page)) {
- bad_page(page, "first_page not consistent", 0);
- return 1;
+ if (unlikely(compound_head(page) != head_page)) {
+ bad_page(page, "compound_head not consistent", 0);
+ goto out;
}
- return 0;
+ ret = 0;
+out:
+ clear_compound_head(page);
+ return ret;
}
static void __meminit __init_single_page(struct page *page, unsigned long pfn,
diff --git a/mm/swap.c b/mm/swap.c
index a3a0a2f1f7c3..faa9e1687dea 100644
--- a/mm/swap.c
+++ b/mm/swap.c
@@ -200,7 +200,7 @@ out_put_single:
__put_single_page(page);
return;
}
- VM_BUG_ON_PAGE(page_head != page->first_page, page);
+ VM_BUG_ON_PAGE(page_head != compound_head(page), page);
/*
* We can release the refcount taken by
* get_page_unless_zero() now that
@@ -261,7 +261,7 @@ static void put_compound_page(struct page *page)
* Case 3 is possible, as we may race with
* __split_huge_page_refcount tearing down a THP page.
*/
- page_head = compound_head_by_tail(page);
+ page_head = compound_head(page);
if (!__compound_tail_refcounted(page_head))
put_unrefcounted_compound_page(page_head, page);
else
--
2.5.0
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Andrew Morton <akpm@linux-foundation.org> |
|---|---|
| Date | 2015-08-21 01:40 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <pZHB8-33E-19@gated-at.bofh.it> |
| In reply to | #1209739 |
On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote:
> Hugh has pointed that compound_head() call can be unsafe in some
> context. There's one example:
>
> CPU0 CPU1
>
> isolate_migratepages_block()
> page_count()
> compound_head()
> !!PageTail() == true
> put_page()
> tail->first_page = NULL
> head = tail->first_page
> alloc_pages(__GFP_COMP)
> prep_compound_page()
> tail->first_page = head
> __SetPageTail(p);
> !!PageTail() == true
> <head == NULL dereferencing>
>
> The race is pure theoretical. I don't it's possible to trigger it in
> practice. But who knows.
>
> We can fix the race by changing how encode PageTail() and compound_head()
> within struct page to be able to update them in one shot.
>
> The patch introduces page->compound_head into third double word block in
> front of compound_dtor and compound_order. That means it shares storage
> space with:
>
> - page->lru.next;
> - page->next;
> - page->rcu_head.next;
> - page->pmd_huge_pte;
>
> That's too long list to be absolutely sure, but looks like nobody uses
> bit 0 of the word. It can be used to encode PageTail(). And if the bit
> set, rest of the word is pointer to head page.
So nothing else which participates in the union in the "Third double
word block" is allowed to use bit zero of the first word.
Is this really true? For example if it's a slab page, will that page
ever be inspected by code which is looking for the PageTail bit?
Anyway, this is quite subtle and there's a risk that people will
accidentally break it later on. I don't think the patch puts
sufficient documentation in place to prevent this. And even
documentation might not be enough to prevent accidents.
>
> ...
>
> --- a/include/linux/mm_types.h
> +++ b/include/linux/mm_types.h
> @@ -120,7 +120,12 @@ struct page {
> };
> };
>
> - /* Third double word block */
> + /*
> + * Third double word block
> + *
> + * WARNING: bit 0 of the first word encode PageTail and *must* be 0
> + * for non-tail pages.
> + */
> union {
> struct list_head lru; /* Pageout list, eg. active_list
> * protected by zone->lru_lock !
> @@ -143,6 +148,7 @@ struct page {
> */
> /* First tail page of compound page */
> struct {
> + unsigned long compound_head; /* If bit zero is set */
I think the comments around here should have more details and should
be louder!
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2015-08-21 14:20 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <pZTsC-3k5-17@gated-at.bofh.it> |
| In reply to | #1210792 |
On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote:
> On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote:
>
> > Hugh has pointed that compound_head() call can be unsafe in some
> > context. There's one example:
> >
> > CPU0 CPU1
> >
> > isolate_migratepages_block()
> > page_count()
> > compound_head()
> > !!PageTail() == true
> > put_page()
> > tail->first_page = NULL
> > head = tail->first_page
> > alloc_pages(__GFP_COMP)
> > prep_compound_page()
> > tail->first_page = head
> > __SetPageTail(p);
> > !!PageTail() == true
> > <head == NULL dereferencing>
> >
> > The race is pure theoretical. I don't it's possible to trigger it in
> > practice. But who knows.
> >
> > We can fix the race by changing how encode PageTail() and compound_head()
> > within struct page to be able to update them in one shot.
> >
> > The patch introduces page->compound_head into third double word block in
> > front of compound_dtor and compound_order. That means it shares storage
> > space with:
> >
> > - page->lru.next;
> > - page->next;
> > - page->rcu_head.next;
> > - page->pmd_huge_pte;
> >
> > That's too long list to be absolutely sure, but looks like nobody uses
> > bit 0 of the word. It can be used to encode PageTail(). And if the bit
> > set, rest of the word is pointer to head page.
>
> So nothing else which participates in the union in the "Third double
> word block" is allowed to use bit zero of the first word.
Correct.
> Is this really true? For example if it's a slab page, will that page
> ever be inspected by code which is looking for the PageTail bit?
+Christoph.
What we know for sure is that space is not used in tail pages, otherwise
it would collide with current compound_dtor.
For head/small pages it gets trickier. I convinced myself that it should
be safe this way:
All fields it shares space with are pointers (with possible exception of
pmd_huge_pte, see below) to objects with sizeof() > 1. I think it's
reasonable to expect that the bit 0 in such pointers would be clear due
alignment. We do the same for page->mapping.
On pmd_huge_pte: it's pgtable_t which on most architectures is typedef to
struct page *. That should not create any conflicts. On some architectures
it's pte_t *, which is fine too. On arc it's virtual address of the page
in form of unsigned long. It should work.
The worry I have about pmd_huge_pte is that some new architecture may
choose to implement pgtable_t as pfn and that will collide on bit 0. :-/
We can address this worry by shifting pmd_huge_pte to the second word in
the double word block. But I'm not sure if we should.
And of course there's chance that these field are used not according to
its type. I didn't find such cases, but I can't guarantee that they don't
exist.
I tested patched kernel with all three SLAB allocator and was not able to
crash it under trinity. More testing is required.
> Anyway, this is quite subtle and there's a risk that people will
> accidentally break it later on. I don't think the patch puts
> sufficient documentation in place to prevent this.
I would appreciate for suggestion on place and form of documentation.
> And even documentation might not be enough to prevent accidents.
The only think I can propose is VM_BUG_ON() in PageTail() and
compound_head() which would ensure that page->compound_page points to
place within MAX_ORDER_NR_PAGES before the current page if bit 0 is set.
Do you consider this helpful?
> >
> > ...
> >
> > --- a/include/linux/mm_types.h
> > +++ b/include/linux/mm_types.h
> > @@ -120,7 +120,12 @@ struct page {
> > };
> > };
> >
> > - /* Third double word block */
> > + /*
> > + * Third double word block
> > + *
> > + * WARNING: bit 0 of the first word encode PageTail and *must* be 0
> > + * for non-tail pages.
> > + */
> > union {
> > struct list_head lru; /* Pageout list, eg. active_list
> > * protected by zone->lru_lock !
> > @@ -143,6 +148,7 @@ struct page {
> > */
> > /* First tail page of compound page */
> > struct {
> > + unsigned long compound_head; /* If bit zero is set */
>
> I think the comments around here should have more details and should
> be louder!
I'm always bad when it comes to documentation. Is it enough?
/*
* Third double word block
*
* WARNING: bit 0 of the first word encode PageTail(). That means
* the rest users of the storage space MUST NOT use the bit to
* avoid collision and false-positive PageTail().
*/
--
Kirill A. Shutemov
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2015-08-21 18:20 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <pZXcS-kI-37@gated-at.bofh.it> |
| In reply to | #1211128 |
On Fri, 21 Aug 2015, Kirill A. Shutemov wrote: > > Is this really true? For example if it's a slab page, will that page > > ever be inspected by code which is looking for the PageTail bit? > > +Christoph. > > What we know for sure is that space is not used in tail pages, otherwise > it would collide with current compound_dtor. Sl*b allocators only do a virt_to_head_page on tail pages. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2015-08-21 21:40 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q00kr-4Le-25@gated-at.bofh.it> |
| In reply to | #1211233 |
On Fri, Aug 21, 2015 at 11:11:27AM -0500, Christoph Lameter wrote: > On Fri, 21 Aug 2015, Kirill A. Shutemov wrote: > > > > Is this really true? For example if it's a slab page, will that page > > > ever be inspected by code which is looking for the PageTail bit? > > > > +Christoph. > > > > What we know for sure is that space is not used in tail pages, otherwise > > it would collide with current compound_dtor. > > Sl*b allocators only do a virt_to_head_page on tail pages. The question was whether it's safe to assume that the bit 0 is always zero in the word as this bit will encode PageTail(). -- Kirill A. Shutemov -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Andrew Morton <akpm@linux-foundation.org> |
|---|---|
| Date | 2015-08-21 21:40 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q00kt-4Le-71@gated-at.bofh.it> |
| In reply to | #1211306 |
On Fri, 21 Aug 2015 22:31:09 +0300 "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > On Fri, Aug 21, 2015 at 11:11:27AM -0500, Christoph Lameter wrote: > > On Fri, 21 Aug 2015, Kirill A. Shutemov wrote: > > > > > > Is this really true? For example if it's a slab page, will that page > > > > ever be inspected by code which is looking for the PageTail bit? > > > > > > +Christoph. > > > > > > What we know for sure is that space is not used in tail pages, otherwise > > > it would collide with current compound_dtor. > > > > Sl*b allocators only do a virt_to_head_page on tail pages. > > The question was whether it's safe to assume that the bit 0 is always zero > in the word as this bit will encode PageTail(). That wasn't my question actually... What I'm wondering is: if this page is being used for slab, will any code path ever run PageTail() against it? If not, we don't need to be concerned about that bit. And slab was just the example I chose. The same question petains to all other uses of that union. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Christoph Lameter <cl@linux.com> |
|---|---|
| Date | 2015-08-21 23:20 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q01Tc-74X-13@gated-at.bofh.it> |
| In reply to | #1211307 |
On Fri, 21 Aug 2015, Andrew Morton wrote: > On Fri, 21 Aug 2015 22:31:09 +0300 "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > > > On Fri, Aug 21, 2015 at 11:11:27AM -0500, Christoph Lameter wrote: > > > On Fri, 21 Aug 2015, Kirill A. Shutemov wrote: > > > > > > > > Is this really true? For example if it's a slab page, will that page > > > > > ever be inspected by code which is looking for the PageTail bit? > > > > > > > > +Christoph. > > > > > > > > What we know for sure is that space is not used in tail pages, otherwise > > > > it would collide with current compound_dtor. > > > > > > Sl*b allocators only do a virt_to_head_page on tail pages. > > > > The question was whether it's safe to assume that the bit 0 is always zero > > in the word as this bit will encode PageTail(). > > That wasn't my question actually... > > What I'm wondering is: if this page is being used for slab, will any > code path ever run PageTail() against it? If not, we don't need to be > concerned about that bit. virt_to_head_page will run PageTail because it uses compound_head(). And compound_head needs to use the first_page pointer if its a tail page. -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vlastimil Babka <vbabka@suse.cz> |
|---|---|
| Date | 2015-08-24 17:50 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q12au-3Av-7@gated-at.bofh.it> |
| In reply to | #1211307 |
On 08/21/2015 09:34 PM, Andrew Morton wrote: > On Fri, 21 Aug 2015 22:31:09 +0300 "Kirill A. Shutemov" <kirill@shutemov.name> wrote: > >> On Fri, Aug 21, 2015 at 11:11:27AM -0500, Christoph Lameter wrote: >>> On Fri, 21 Aug 2015, Kirill A. Shutemov wrote: >>> >>>>> Is this really true? For example if it's a slab page, will that page >>>>> ever be inspected by code which is looking for the PageTail bit? >>>> >>>> +Christoph. >>>> >>>> What we know for sure is that space is not used in tail pages, otherwise >>>> it would collide with current compound_dtor. >>> >>> Sl*b allocators only do a virt_to_head_page on tail pages. >> >> The question was whether it's safe to assume that the bit 0 is always zero >> in the word as this bit will encode PageTail(). > > That wasn't my question actually... > > What I'm wondering is: if this page is being used for slab, will any > code path ever run PageTail() against it? If not, we don't need to be > concerned about that bit. Pfn scanners such as compaction might inspect such pages and run compound_head() (and thus PageTail) on them. I think no kind of page within a zone (slab or otherwise) is "protected" from this, which is why it needs to be robust. > And slab was just the example I chose. The same question petains to > all other uses of that union. > -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vlastimil Babka <vbabka@suse.cz> |
|---|---|
| Date | 2015-08-25 13:50 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1kTL-5od-9@gated-at.bofh.it> |
| In reply to | #1211128 |
On 08/21/2015 02:10 PM, Kirill A. Shutemov wrote:
> On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote:
>> On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote:
>>
>>> The patch introduces page->compound_head into third double word block in
>>> front of compound_dtor and compound_order. That means it shares storage
>>> space with:
>>>
>>> - page->lru.next;
>>> - page->next;
>>> - page->rcu_head.next;
>>> - page->pmd_huge_pte;
>>>
We should probably ask Paul about the chances that rcu_head.next would
like to use the bit too one day?
For pgtable_t I can't think of anything better than a warning in the
generic definition in include/asm-generic/page.h and hope that anyone
reimplementing it for a new arch will look there first.
The lru part is probably the hardest to prevent danger. It can be used
for any private purposes. Hopefully everyone currently uses only
standard list operations here, and the list poison values don't set bit
0. But I see there can be some arbitrary CONFIG_ILLEGAL_POINTER_VALUE
added to the poisons, so maybe that's worth some build error check?
Anyway we would be imposing restrictions on types that are not ours, so
there might be some resistance...
>
>> Anyway, this is quite subtle and there's a risk that people will
>> accidentally break it later on. I don't think the patch puts
>> sufficient documentation in place to prevent this.
>
> I would appreciate for suggestion on place and form of documentation.
>
>> And even documentation might not be enough to prevent accidents.
>
> The only think I can propose is VM_BUG_ON() in PageTail() and
> compound_head() which would ensure that page->compound_page points to
> place within MAX_ORDER_NR_PAGES before the current page if bit 0 is set.
That should probably catch some bad stuff, but probably only moments
before it would crash anyway if the pointer was bogus. But I also don't
see better way, because we can't proactively put checks in those who
would "misbehave", as we don't know who they are. Putting more debug
checks in e.g. page freeing might help, but probably not much.
> Do you consider this helpful?
>
>>>
>>> ...
>>>
>>> --- a/include/linux/mm_types.h
>>> +++ b/include/linux/mm_types.h
>>> @@ -120,7 +120,12 @@ struct page {
>>> };
>>> };
>>>
>>> - /* Third double word block */
>>> + /*
>>> + * Third double word block
>>> + *
>>> + * WARNING: bit 0 of the first word encode PageTail and *must* be 0
>>> + * for non-tail pages.
>>> + */
>>> union {
>>> struct list_head lru; /* Pageout list, eg. active_list
>>> * protected by zone->lru_lock !
>>> @@ -143,6 +148,7 @@ struct page {
>>> */
>>> /* First tail page of compound page */
Note that compound_head is not just in the *first* tail page. Only the
rest is.
>>> struct {
>>> + unsigned long compound_head; /* If bit zero is set */
>>
>> I think the comments around here should have more details and should
>> be louder!
>
> I'm always bad when it comes to documentation. Is it enough?
>
> /*
> * Third double word block
> *
> * WARNING: bit 0 of the first word encode PageTail(). That means
> * the rest users of the storage space MUST NOT use the bit to
> * avoid collision and false-positive PageTail().
> */
>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2015-08-25 20:40 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1riz-6ju-51@gated-at.bofh.it> |
| In reply to | #1212985 |
On Tue, Aug 25, 2015 at 01:44:13PM +0200, Vlastimil Babka wrote:
> On 08/21/2015 02:10 PM, Kirill A. Shutemov wrote:
> >On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote:
> >>On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote:
> >>
> >>>The patch introduces page->compound_head into third double word block in
> >>>front of compound_dtor and compound_order. That means it shares storage
> >>>space with:
> >>>
> >>> - page->lru.next;
> >>> - page->next;
> >>> - page->rcu_head.next;
> >>> - page->pmd_huge_pte;
> >>>
>
> We should probably ask Paul about the chances that rcu_head.next would like
> to use the bit too one day?
+Paul.
> For pgtable_t I can't think of anything better than a warning in the generic
> definition in include/asm-generic/page.h and hope that anyone reimplementing
> it for a new arch will look there first.
I will move it to other word, just in case.
> The lru part is probably the hardest to prevent danger. It can be used for
> any private purposes. Hopefully everyone currently uses only standard list
> operations here, and the list poison values don't set bit 0. But I see there
> can be some arbitrary CONFIG_ILLEGAL_POINTER_VALUE added to the poisons, so
> maybe that's worth some build error check? Anyway we would be imposing
> restrictions on types that are not ours, so there might be some
> resistance...
I will add BUILD_BUG_ON((unsigned long)LIST_POISON1 & 1);
> >>Anyway, this is quite subtle and there's a risk that people will
> >>accidentally break it later on. I don't think the patch puts
> >>sufficient documentation in place to prevent this.
> >
> >I would appreciate for suggestion on place and form of documentation.
> >
> >>And even documentation might not be enough to prevent accidents.
> >
> >The only think I can propose is VM_BUG_ON() in PageTail() and
> >compound_head() which would ensure that page->compound_page points to
> >place within MAX_ORDER_NR_PAGES before the current page if bit 0 is set.
>
> That should probably catch some bad stuff, but probably only moments before
> it would crash anyway if the pointer was bogus. But I also don't see better
> way, because we can't proactively put checks in those who would "misbehave",
> as we don't know who they are. Putting more debug checks in e.g. page
> freeing might help, but probably not much.
So, do you think it worth it or not after all?
>
> >Do you consider this helpful?
> >
> >>>
> >>>...
> >>>
> >>>--- a/include/linux/mm_types.h
> >>>+++ b/include/linux/mm_types.h
> >>>@@ -120,7 +120,12 @@ struct page {
> >>> };
> >>> };
> >>>
> >>>- /* Third double word block */
> >>>+ /*
> >>>+ * Third double word block
> >>>+ *
> >>>+ * WARNING: bit 0 of the first word encode PageTail and *must* be 0
> >>>+ * for non-tail pages.
> >>>+ */
> >>> union {
> >>> struct list_head lru; /* Pageout list, eg. active_list
> >>> * protected by zone->lru_lock !
> >>>@@ -143,6 +148,7 @@ struct page {
> >>> */
> >>> /* First tail page of compound page */
>
> Note that compound_head is not just in the *first* tail page. Only the rest
> is.
Right.
--
Kirill A. Shutemov
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2015-08-25 22:20 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1sRk-ky-23@gated-at.bofh.it> |
| In reply to | #1213238 |
On Tue, Aug 25, 2015 at 09:33:54PM +0300, Kirill A. Shutemov wrote:
> On Tue, Aug 25, 2015 at 01:44:13PM +0200, Vlastimil Babka wrote:
> > On 08/21/2015 02:10 PM, Kirill A. Shutemov wrote:
> > >On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote:
> > >>On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote:
> > >>
> > >>>The patch introduces page->compound_head into third double word block in
> > >>>front of compound_dtor and compound_order. That means it shares storage
> > >>>space with:
> > >>>
> > >>> - page->lru.next;
> > >>> - page->next;
> > >>> - page->rcu_head.next;
> > >>> - page->pmd_huge_pte;
> > >>>
> >
> > We should probably ask Paul about the chances that rcu_head.next would like
> > to use the bit too one day?
>
> +Paul.
The call_rcu() function does stomp that bit, but if you stop using that
bit before you invoke call_rcu(), no problem.
Thanx, Paul
> > For pgtable_t I can't think of anything better than a warning in the generic
> > definition in include/asm-generic/page.h and hope that anyone reimplementing
> > it for a new arch will look there first.
>
> I will move it to other word, just in case.
>
> > The lru part is probably the hardest to prevent danger. It can be used for
> > any private purposes. Hopefully everyone currently uses only standard list
> > operations here, and the list poison values don't set bit 0. But I see there
> > can be some arbitrary CONFIG_ILLEGAL_POINTER_VALUE added to the poisons, so
> > maybe that's worth some build error check? Anyway we would be imposing
> > restrictions on types that are not ours, so there might be some
> > resistance...
>
> I will add BUILD_BUG_ON((unsigned long)LIST_POISON1 & 1);
>
> > >>Anyway, this is quite subtle and there's a risk that people will
> > >>accidentally break it later on. I don't think the patch puts
> > >>sufficient documentation in place to prevent this.
> > >
> > >I would appreciate for suggestion on place and form of documentation.
> > >
> > >>And even documentation might not be enough to prevent accidents.
> > >
> > >The only think I can propose is VM_BUG_ON() in PageTail() and
> > >compound_head() which would ensure that page->compound_page points to
> > >place within MAX_ORDER_NR_PAGES before the current page if bit 0 is set.
> >
> > That should probably catch some bad stuff, but probably only moments before
> > it would crash anyway if the pointer was bogus. But I also don't see better
> > way, because we can't proactively put checks in those who would "misbehave",
> > as we don't know who they are. Putting more debug checks in e.g. page
> > freeing might help, but probably not much.
>
> So, do you think it worth it or not after all?
> >
> > >Do you consider this helpful?
> > >
> > >>>
> > >>>...
> > >>>
> > >>>--- a/include/linux/mm_types.h
> > >>>+++ b/include/linux/mm_types.h
> > >>>@@ -120,7 +120,12 @@ struct page {
> > >>> };
> > >>> };
> > >>>
> > >>>- /* Third double word block */
> > >>>+ /*
> > >>>+ * Third double word block
> > >>>+ *
> > >>>+ * WARNING: bit 0 of the first word encode PageTail and *must* be 0
> > >>>+ * for non-tail pages.
> > >>>+ */
> > >>> union {
> > >>> struct list_head lru; /* Pageout list, eg. active_list
> > >>> * protected by zone->lru_lock !
> > >>>@@ -143,6 +148,7 @@ struct page {
> > >>> */
> > >>> /* First tail page of compound page */
> >
> > Note that compound_head is not just in the *first* tail page. Only the rest
> > is.
>
> Right.
>
> --
> Kirill A. Shutemov
>
--
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vlastimil Babka <vbabka@suse.cz> |
|---|---|
| Date | 2015-08-25 22:50 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1tkm-SD-15@gated-at.bofh.it> |
| In reply to | #1213303 |
On 25.8.2015 22:11, Paul E. McKenney wrote: > On Tue, Aug 25, 2015 at 09:33:54PM +0300, Kirill A. Shutemov wrote: >> On Tue, Aug 25, 2015 at 01:44:13PM +0200, Vlastimil Babka wrote: >>> On 08/21/2015 02:10 PM, Kirill A. Shutemov wrote: >>>> On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote: >>>>> On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote: >>>>> >>>>>> The patch introduces page->compound_head into third double word block in >>>>>> front of compound_dtor and compound_order. That means it shares storage >>>>>> space with: >>>>>> >>>>>> - page->lru.next; >>>>>> - page->next; >>>>>> - page->rcu_head.next; >>>>>> - page->pmd_huge_pte; >>>>>> >>> >>> We should probably ask Paul about the chances that rcu_head.next would like >>> to use the bit too one day? >> >> +Paul. > > The call_rcu() function does stomp that bit, but if you stop using that > bit before you invoke call_rcu(), no problem. You mean that it sets the bit 0 of rcu_head.next during its processing? That's bad news then. It's not that we would trigger that bit when the rcu_head part of the union is "active". It's that pfn scanners could inspect such page at arbitrary time, see the bit 0 set (due to RCU processing) and think that it's a tail page of a compound page, and interpret the rest of the pointer as a pointer to the head page (to test it for flags etc). -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2015-08-25 23:30 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1tX4-1RJ-27@gated-at.bofh.it> |
| In reply to | #1213333 |
On Tue, Aug 25, 2015 at 10:46:44PM +0200, Vlastimil Babka wrote: > On 25.8.2015 22:11, Paul E. McKenney wrote: > > On Tue, Aug 25, 2015 at 09:33:54PM +0300, Kirill A. Shutemov wrote: > >> On Tue, Aug 25, 2015 at 01:44:13PM +0200, Vlastimil Babka wrote: > >>> On 08/21/2015 02:10 PM, Kirill A. Shutemov wrote: > >>>> On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote: > >>>>> On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote: > >>>>> > >>>>>> The patch introduces page->compound_head into third double word block in > >>>>>> front of compound_dtor and compound_order. That means it shares storage > >>>>>> space with: > >>>>>> > >>>>>> - page->lru.next; > >>>>>> - page->next; > >>>>>> - page->rcu_head.next; > >>>>>> - page->pmd_huge_pte; > >>>>>> > >>> > >>> We should probably ask Paul about the chances that rcu_head.next would like > >>> to use the bit too one day? > >> > >> +Paul. > > > > The call_rcu() function does stomp that bit, but if you stop using that > > bit before you invoke call_rcu(), no problem. > > You mean that it sets the bit 0 of rcu_head.next during its processing? Not at the moment, though RCU will splat if given a misaligned rcu_head structure because of the possibility to use that bit to flag callbacks that do nothing but free memory. If RCU needs to do that (e.g., to promote energy efficiency), then that bit might well be set during RCU grace-period processing. > That's > bad news then. It's not that we would trigger that bit when the rcu_head part of > the union is "active". It's that pfn scanners could inspect such page at > arbitrary time, see the bit 0 set (due to RCU processing) and think that it's a > tail page of a compound page, and interpret the rest of the pointer as a pointer > to the head page (to test it for flags etc). On the other hand, if you avoid scanning rcu_head structures for pages that are currently waiting for a grace period, no problem. RCU does not use the rcu_head structure at all except for during the time between when call_rcu() is invoked on that rcu_head structure and the time that the callback is invoked. Is there some other page state that indicates that the page is waiting for a grace period? If so, you could simply avoid testing that bit in that case. Thanx, Paul -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Kirill A. Shutemov" <kirill@shutemov.name> |
|---|---|
| Date | 2015-08-26 17:10 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1KuR-Rx-5@gated-at.bofh.it> |
| In reply to | #1213370 |
On Tue, Aug 25, 2015 at 02:19:54PM -0700, Paul E. McKenney wrote: > On Tue, Aug 25, 2015 at 10:46:44PM +0200, Vlastimil Babka wrote: > > On 25.8.2015 22:11, Paul E. McKenney wrote: > > > On Tue, Aug 25, 2015 at 09:33:54PM +0300, Kirill A. Shutemov wrote: > > >> On Tue, Aug 25, 2015 at 01:44:13PM +0200, Vlastimil Babka wrote: > > >>> On 08/21/2015 02:10 PM, Kirill A. Shutemov wrote: > > >>>> On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote: > > >>>>> On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote: > > >>>>> > > >>>>>> The patch introduces page->compound_head into third double word block in > > >>>>>> front of compound_dtor and compound_order. That means it shares storage > > >>>>>> space with: > > >>>>>> > > >>>>>> - page->lru.next; > > >>>>>> - page->next; > > >>>>>> - page->rcu_head.next; > > >>>>>> - page->pmd_huge_pte; > > >>>>>> > > >>> > > >>> We should probably ask Paul about the chances that rcu_head.next would like > > >>> to use the bit too one day? > > >> > > >> +Paul. > > > > > > The call_rcu() function does stomp that bit, but if you stop using that > > > bit before you invoke call_rcu(), no problem. > > > > You mean that it sets the bit 0 of rcu_head.next during its processing? > > Not at the moment, though RCU will splat if given a misaligned rcu_head > structure because of the possibility to use that bit to flag callbacks > that do nothing but free memory. If RCU needs to do that (e.g., to > promote energy efficiency), then that bit might well be set during > RCU grace-period processing. Ugh.. :-/ > > That's > > bad news then. It's not that we would trigger that bit when the rcu_head part of > > the union is "active". It's that pfn scanners could inspect such page at > > arbitrary time, see the bit 0 set (due to RCU processing) and think that it's a > > tail page of a compound page, and interpret the rest of the pointer as a pointer > > to the head page (to test it for flags etc). > > On the other hand, if you avoid scanning rcu_head structures for pages > that are currently waiting for a grace period, no problem. RCU does > not use the rcu_head structure at all except for during the time between > when call_rcu() is invoked on that rcu_head structure and the time that > the callback is invoked. > > Is there some other page state that indicates that the page is waiting > for a grace period? If so, you could simply avoid testing that bit in > that case. No, I don't think so. For compound pages most of info of its state is stored in head page (e.g. page_count(), flags, etc). So if we examine random page (pfn scanner case) the very first thing we want to know if we stepped on tail page. PageTail() is what I wanted to encode in the bit... What if we change order of fields within rcu_head and put ->func first? Can we expect this pointer to have bit 0 always clear? -- Kirill A. Shutemov -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Vlastimil Babka <vbabka@suse.cz> |
|---|---|
| Date | 2015-08-26 17:40 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1KXT-1pg-19@gated-at.bofh.it> |
| In reply to | #1213949 |
On 08/26/2015 05:04 PM, Kirill A. Shutemov wrote: >>> That's >>> bad news then. It's not that we would trigger that bit when the rcu_head part of >>> the union is "active". It's that pfn scanners could inspect such page at >>> arbitrary time, see the bit 0 set (due to RCU processing) and think that it's a >>> tail page of a compound page, and interpret the rest of the pointer as a pointer >>> to the head page (to test it for flags etc). >> >> On the other hand, if you avoid scanning rcu_head structures for pages >> that are currently waiting for a grace period, no problem. RCU does >> not use the rcu_head structure at all except for during the time between >> when call_rcu() is invoked on that rcu_head structure and the time that >> the callback is invoked. >> >> Is there some other page state that indicates that the page is waiting >> for a grace period? If so, you could simply avoid testing that bit in >> that case. > > No, I don't think so. > > For compound pages most of info of its state is stored in head page (e.g. > page_count(), flags, etc). So if we examine random page (pfn scanner case) > the very first thing we want to know if we stepped on tail page. > PageTail() is what I wanted to encode in the bit... > > What if we change order of fields within rcu_head and put ->func first? Or change the order of compound_head wrt the rest? > Can we expect this pointer to have bit 0 always clear? That's probably a question whether $compiler is guaranteed to align functions on all architectures... -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2015-08-26 18:40 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1LTY-2LN-23@gated-at.bofh.it> |
| In reply to | #1213949 |
On Wed, Aug 26, 2015 at 06:04:12PM +0300, Kirill A. Shutemov wrote: > On Tue, Aug 25, 2015 at 02:19:54PM -0700, Paul E. McKenney wrote: > > On Tue, Aug 25, 2015 at 10:46:44PM +0200, Vlastimil Babka wrote: > > > On 25.8.2015 22:11, Paul E. McKenney wrote: > > > > On Tue, Aug 25, 2015 at 09:33:54PM +0300, Kirill A. Shutemov wrote: > > > >> On Tue, Aug 25, 2015 at 01:44:13PM +0200, Vlastimil Babka wrote: > > > >>> On 08/21/2015 02:10 PM, Kirill A. Shutemov wrote: > > > >>>> On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote: > > > >>>>> On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote: > > > >>>>> > > > >>>>>> The patch introduces page->compound_head into third double word block in > > > >>>>>> front of compound_dtor and compound_order. That means it shares storage > > > >>>>>> space with: > > > >>>>>> > > > >>>>>> - page->lru.next; > > > >>>>>> - page->next; > > > >>>>>> - page->rcu_head.next; > > > >>>>>> - page->pmd_huge_pte; > > > >>>>>> > > > >>> > > > >>> We should probably ask Paul about the chances that rcu_head.next would like > > > >>> to use the bit too one day? > > > >> > > > >> +Paul. > > > > > > > > The call_rcu() function does stomp that bit, but if you stop using that > > > > bit before you invoke call_rcu(), no problem. > > > > > > You mean that it sets the bit 0 of rcu_head.next during its processing? > > > > Not at the moment, though RCU will splat if given a misaligned rcu_head > > structure because of the possibility to use that bit to flag callbacks > > that do nothing but free memory. If RCU needs to do that (e.g., to > > promote energy efficiency), then that bit might well be set during > > RCU grace-period processing. > > Ugh.. :-/ > > > > bad news then. It's not that we would trigger that bit when the rcu_head part of > > > the union is "active". It's that pfn scanners could inspect such page at > > > arbitrary time, see the bit 0 set (due to RCU processing) and think that it's a > > > tail page of a compound page, and interpret the rest of the pointer as a pointer > > > to the head page (to test it for flags etc). > > > > On the other hand, if you avoid scanning rcu_head structures for pages > > that are currently waiting for a grace period, no problem. RCU does > > not use the rcu_head structure at all except for during the time between > > when call_rcu() is invoked on that rcu_head structure and the time that > > the callback is invoked. > > > > Is there some other page state that indicates that the page is waiting > > for a grace period? If so, you could simply avoid testing that bit in > > that case. > > No, I don't think so. OK, I'll bite... How do you know that it is safe to invoke call_rcu(), given that you are not allowed to invoke call_rcu() until the previous callback has been invoked? > For compound pages most of info of its state is stored in head page (e.g. > page_count(), flags, etc). So if we examine random page (pfn scanner case) > the very first thing we want to know if we stepped on tail page. > PageTail() is what I wanted to encode in the bit... Ah, so that would require the page scanner to do reverse mapping or some such, then. Which is perhaps what you are trying to avoid. > What if we change order of fields within rcu_head and put ->func first? > Can we expect this pointer to have bit 0 always clear? I asked that question some time back, and the answer was "no". You can apparently have functions that start at odd addresses on some architectures. That said, there are likely to be reserved bits somewhere in the function address, perhaps varying depending on architecture and/or boot, in the case of address-space randomization. Perhaps some way of identifying those bits with architecture-independent ways of querying and setting them? Thanx, Paul -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Hugh Dickins <hughd@google.com> |
|---|---|
| Date | 2015-08-26 20:20 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1NsK-55Y-29@gated-at.bofh.it> |
| In reply to | #1213370 |
On Tue, 25 Aug 2015, Paul E. McKenney wrote: > On Tue, Aug 25, 2015 at 10:46:44PM +0200, Vlastimil Babka wrote: > > On 25.8.2015 22:11, Paul E. McKenney wrote: > > > On Tue, Aug 25, 2015 at 09:33:54PM +0300, Kirill A. Shutemov wrote: > > >> On Tue, Aug 25, 2015 at 01:44:13PM +0200, Vlastimil Babka wrote: > > >>> On 08/21/2015 02:10 PM, Kirill A. Shutemov wrote: > > >>>> On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote: > > >>>>> On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote: > > >>>>> > > >>>>>> The patch introduces page->compound_head into third double word block in > > >>>>>> front of compound_dtor and compound_order. That means it shares storage > > >>>>>> space with: > > >>>>>> > > >>>>>> - page->lru.next; > > >>>>>> - page->next; > > >>>>>> - page->rcu_head.next; > > >>>>>> - page->pmd_huge_pte; > > >>>>>> > > >>> > > >>> We should probably ask Paul about the chances that rcu_head.next would like > > >>> to use the bit too one day? > > >> > > >> +Paul. > > > > > > The call_rcu() function does stomp that bit, but if you stop using that > > > bit before you invoke call_rcu(), no problem. > > > > You mean that it sets the bit 0 of rcu_head.next during its processing? > > Not at the moment, though RCU will splat if given a misaligned rcu_head > structure because of the possibility to use that bit to flag callbacks > that do nothing but free memory. If RCU needs to do that (e.g., to > promote energy efficiency), then that bit might well be set during > RCU grace-period processing. But if you do one day implement that, wouldn't sl?b.c have to use call_rcu_with_added_meaning() instead of call_rcu(), to be in danger of getting that bit set? (No rcu_head is placed in a PageTail page.) So although it might be a little strange not to use a variant intended for freeing memory when indeed that's what it's doing, it would not be the end of the world for SLAB_DESTROY_BY_RCU to carry on using straight call_rcu(), in defence of the struct page safety Kirill is proposing. hUgh -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | "Paul E. McKenney" <paulmck@linux.vnet.ibm.com> |
|---|---|
| Date | 2015-08-26 23:30 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1QqB-RP-7@gated-at.bofh.it> |
| In reply to | #1214096 |
On Wed, Aug 26, 2015 at 11:18:45AM -0700, Hugh Dickins wrote: > On Tue, 25 Aug 2015, Paul E. McKenney wrote: > > On Tue, Aug 25, 2015 at 10:46:44PM +0200, Vlastimil Babka wrote: > > > On 25.8.2015 22:11, Paul E. McKenney wrote: > > > > On Tue, Aug 25, 2015 at 09:33:54PM +0300, Kirill A. Shutemov wrote: > > > >> On Tue, Aug 25, 2015 at 01:44:13PM +0200, Vlastimil Babka wrote: > > > >>> On 08/21/2015 02:10 PM, Kirill A. Shutemov wrote: > > > >>>> On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote: > > > >>>>> On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote: > > > >>>>> > > > >>>>>> The patch introduces page->compound_head into third double word block in > > > >>>>>> front of compound_dtor and compound_order. That means it shares storage > > > >>>>>> space with: > > > >>>>>> > > > >>>>>> - page->lru.next; > > > >>>>>> - page->next; > > > >>>>>> - page->rcu_head.next; > > > >>>>>> - page->pmd_huge_pte; > > > >>>>>> > > > >>> > > > >>> We should probably ask Paul about the chances that rcu_head.next would like > > > >>> to use the bit too one day? > > > >> > > > >> +Paul. > > > > > > > > The call_rcu() function does stomp that bit, but if you stop using that > > > > bit before you invoke call_rcu(), no problem. > > > > > > You mean that it sets the bit 0 of rcu_head.next during its processing? > > > > Not at the moment, though RCU will splat if given a misaligned rcu_head > > structure because of the possibility to use that bit to flag callbacks > > that do nothing but free memory. If RCU needs to do that (e.g., to > > promote energy efficiency), then that bit might well be set during > > RCU grace-period processing. > > But if you do one day implement that, wouldn't sl?b.c have to use > call_rcu_with_added_meaning() instead of call_rcu(), to be in danger > of getting that bit set? (No rcu_head is placed in a PageTail page.) Good point, call_rcu_lazy(), but yes. > So although it might be a little strange not to use a variant intended > for freeing memory when indeed that's what it's doing, it would not be > the end of the world for SLAB_DESTROY_BY_RCU to carry on using straight > call_rcu(), in defence of the struct page safety Kirill is proposing. As long as you are OK with the bottom bit being zero throughout the RCU processing, yes. Thanx, Paul -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
| From | Hugh Dickins <hughd@google.com> |
|---|---|
| Date | 2015-08-27 00:30 +0200 |
| Subject | Re: [PATCHv3 4/5] mm: make compound_head() robust |
| Message-ID | <q1RmG-2ei-7@gated-at.bofh.it> |
| In reply to | #1214202 |
On Wed, 26 Aug 2015, Paul E. McKenney wrote: > On Wed, Aug 26, 2015 at 11:18:45AM -0700, Hugh Dickins wrote: > > On Tue, 25 Aug 2015, Paul E. McKenney wrote: > > > On Tue, Aug 25, 2015 at 10:46:44PM +0200, Vlastimil Babka wrote: > > > > On 25.8.2015 22:11, Paul E. McKenney wrote: > > > > > On Tue, Aug 25, 2015 at 09:33:54PM +0300, Kirill A. Shutemov wrote: > > > > >> On Tue, Aug 25, 2015 at 01:44:13PM +0200, Vlastimil Babka wrote: > > > > >>> On 08/21/2015 02:10 PM, Kirill A. Shutemov wrote: > > > > >>>> On Thu, Aug 20, 2015 at 04:36:43PM -0700, Andrew Morton wrote: > > > > >>>>> On Wed, 19 Aug 2015 12:21:45 +0300 "Kirill A. Shutemov" <kirill.shutemov@linux.intel.com> wrote: > > > > >>>>> > > > > >>>>>> The patch introduces page->compound_head into third double word block in > > > > >>>>>> front of compound_dtor and compound_order. That means it shares storage > > > > >>>>>> space with: > > > > >>>>>> > > > > >>>>>> - page->lru.next; > > > > >>>>>> - page->next; > > > > >>>>>> - page->rcu_head.next; > > > > >>>>>> - page->pmd_huge_pte; > > > > >>>>>> > > > > >>> > > > > >>> We should probably ask Paul about the chances that rcu_head.next would like > > > > >>> to use the bit too one day? > > > > >> > > > > >> +Paul. > > > > > > > > > > The call_rcu() function does stomp that bit, but if you stop using that > > > > > bit before you invoke call_rcu(), no problem. > > > > > > > > You mean that it sets the bit 0 of rcu_head.next during its processing? > > > > > > Not at the moment, though RCU will splat if given a misaligned rcu_head > > > structure because of the possibility to use that bit to flag callbacks > > > that do nothing but free memory. If RCU needs to do that (e.g., to > > > promote energy efficiency), then that bit might well be set during > > > RCU grace-period processing. > > > > But if you do one day implement that, wouldn't sl?b.c have to use > > call_rcu_with_added_meaning() instead of call_rcu(), to be in danger > > of getting that bit set? (No rcu_head is placed in a PageTail page.) > > Good point, call_rcu_lazy(), but yes. > > > So although it might be a little strange not to use a variant intended > > for freeing memory when indeed that's what it's doing, it would not be > > the end of the world for SLAB_DESTROY_BY_RCU to carry on using straight > > call_rcu(), in defence of the struct page safety Kirill is proposing. > > As long as you are OK with the bottom bit being zero throughout the RCU > processing, yes. That's exactly what we want: sounds like we have no problem, thanks Paul. Hugh -- To unsubscribe from this list: send the line "unsubscribe linux-kernel" in the body of a message to majordomo@vger.kernel.org More majordomo info at http://vger.kernel.org/majordomo-info.html Please read the FAQ at http://www.tux.org/lkml/
[toc] | [prev] | [next] | [standalone]
Page 1 of 2 [1] 2 Next page →
Back to top | Article view | linux.kernel
csiph-web