Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.kernel > #1379561 > unrolled thread
| Started by | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| First post | 2016-04-15 11:00 +0200 |
| Last post | 2016-04-26 13:50 +0200 |
| Articles | 10 on this page of 70 — 4 participants |
Back to article view | Back to linux.kernel
[PATCH 00/28] Optimise page alloc/free fast paths v3 Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:00 +0200
[PATCH 01/28] mm, page_alloc: Only check PageCompound for high-order pages Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:00 +0200
Re: [PATCH 01/28] mm, page_alloc: Only check PageCompound for high-order pages Vlastimil Babka <vbabka@suse.cz> - 2016-04-25 11:40 +0200
Re: [PATCH 01/28] mm, page_alloc: Only check PageCompound for high-order pages Mel Gorman <mgorman@techsingularity.net> - 2016-04-26 12:40 +0200
Re: [PATCH 01/28] mm, page_alloc: Only check PageCompound for high-order pages Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 13:30 +0200
[PATCH 21/28] mm, page_alloc: Avoid looking up the first zone in a zonelist twice Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 21/28] mm, page_alloc: Avoid looking up the first zone in a zonelist twice Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 19:50 +0200
[PATCH 04/28] mm, page_alloc: Inline zone_statistics Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 04/28] mm, page_alloc: Inline zone_statistics Vlastimil Babka <vbabka@suse.cz> - 2016-04-25 13:20 +0200
[PATCH 15/28] mm, page_alloc: Move might_sleep_if check to the allocator slowpath Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 15/28] mm, page_alloc: Move might_sleep_if check to the allocator slowpath Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 15:50 +0200
Re: [PATCH 15/28] mm, page_alloc: Move might_sleep_if check to the allocator slowpath Mel Gorman <mgorman@techsingularity.net> - 2016-04-26 17:00 +0200
Re: [PATCH 15/28] mm, page_alloc: Move might_sleep_if check to the allocator slowpath Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 17:20 +0200
Re: [PATCH 15/28] mm, page_alloc: Move might_sleep_if check to the allocator slowpath Mel Gorman <mgorman@techsingularity.net> - 2016-04-26 18:30 +0200
[PATCH 19/28] mm, page_alloc: Reduce cost of fair zone allocation policy retry Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
[PATCH 13/28] mm, page_alloc: Remove redundant check for empty zonelist Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
[PATCH 22/28] mm, page_alloc: Remove field from alloc_context Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
[PATCH 14/28] mm, page_alloc: Simplify last cpupid reset Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 14/28] mm, page_alloc: Simplify last cpupid reset Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 15:40 +0200
[PATCH 20/28] mm, page_alloc: Shortcut watermark checks for order-0 pages Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
[PATCH 16/28] mm, page_alloc: Move __GFP_HARDWALL modifications out of the fastpath Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 16/28] mm, page_alloc: Move __GFP_HARDWALL modifications out of the fastpath Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 16:20 +0200
[PATCH 17/28] mm, page_alloc: Check once if a zone has isolated pageblocks Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 17/28] mm, page_alloc: Check once if a zone has isolated pageblocks Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 16:30 +0200
[PATCH 18/28] mm, page_alloc: Shorten the page allocator fast path Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 18/28] mm, page_alloc: Shorten the page allocator fast path Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 17:30 +0200
[PATCH 26/28] cpuset: use static key better and convert to new API Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:20 +0200
Re: [PATCH 26/28] cpuset: use static key better and convert to new API Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 22:00 +0200
[PATCH 27/28] mm, page_alloc: Defer debugging checks of freed pages until a PCP drain Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:20 +0200
Re: [PATCH 27/28] mm, page_alloc: Defer debugging checks of freed pages until a PCP drain Vlastimil Babka <vbabka@suse.cz> - 2016-04-27 14:00 +0200
[PATCH 2/3] mm, page_alloc: pull out side effects from free_pages_check Vlastimil Babka <vbabka@suse.cz> - 2016-04-27 14:10 +0200
Re: [PATCH 2/3] mm, page_alloc: pull out side effects from free_pages_check Mel Gorman <mgorman@techsingularity.net> - 2016-04-27 14:50 +0200
Re: [PATCH 2/3] mm, page_alloc: pull out side effects from free_pages_check Vlastimil Babka <vbabka@suse.cz> - 2016-04-27 15:10 +0200
[PATCH 3/3] mm, page_alloc: don't duplicate code in free_pcp_prepare Vlastimil Babka <vbabka@suse.cz> - 2016-04-27 14:10 +0200
[PATCH 1/3] mm, page_alloc: un-inline the bad part of free_pages_check Vlastimil Babka <vbabka@suse.cz> - 2016-04-27 14:10 +0200
Re: [PATCH 1/3] mm, page_alloc: un-inline the bad part of free_pages_check Mel Gorman <mgorman@techsingularity.net> - 2016-04-27 14:40 +0200
Re: [PATCH 1/3] mm, page_alloc: un-inline the bad part of free_pages_check Vlastimil Babka <vbabka@suse.cz> - 2016-04-27 15:00 +0200
[PATCH 25/28] mm, page_alloc: Inline pageblock lookup in page free fast paths Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:20 +0200
[PATCH 24/28] mm, page_alloc: Remove unnecessary variable from free_pcppages_bulk Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:20 +0200
[PATCH 23/28] mm, page_alloc: Check multiple page fields with a single branch Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:20 +0200
Re: [PATCH 23/28] mm, page_alloc: Check multiple page fields with a single branch Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 20:50 +0200
Re: [PATCH 23/28] mm, page_alloc: Check multiple page fields with a single branch Mel Gorman <mgorman@techsingularity.net> - 2016-04-27 12:10 +0200
[PATCH 28/28] mm, page_alloc: Defer debugging checks of pages allocated from the PCP Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:20 +0200
Re: [PATCH 28/28] mm, page_alloc: Defer debugging checks of pages allocated from the PCP Vlastimil Babka <vbabka@suse.cz> - 2016-04-27 16:10 +0200
Re: [PATCH 28/28] mm, page_alloc: Defer debugging checks of pages allocated from the PCP Mel Gorman <mgorman@techsingularity.net> - 2016-04-27 17:40 +0200
Re: [PATCH 13/28] mm, page_alloc: Remove redundant check for empty zonelist Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 14:10 +0200
Re: [PATCH 13/28] mm, page_alloc: Remove redundant check for empty zonelist Mel Gorman <mgorman@techsingularity.net> - 2016-04-26 15:10 +0200
Re: [PATCH 13/28] mm, page_alloc: Remove redundant check for empty zonelist Andrew Morton <akpm@linux-foundation.org> - 2016-04-26 21:20 +0200
[PATCH 07/28] mm, page_alloc: Avoid unnecessary zone lookups during pageblock operations Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 07/28] mm, page_alloc: Avoid unnecessary zone lookups during pageblock operations Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 13:30 +0200
[PATCH 02/28] mm, page_alloc: Use new PageAnonHead helper in the free page fast path Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 02/28] mm, page_alloc: Use new PageAnonHead helper in the free page fast path Vlastimil Babka <vbabka@suse.cz> - 2016-04-25 12:00 +0200
[PATCH 11/28] mm, page_alloc: Remove unnecessary initialisation in get_page_from_freelist Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
[PATCH 06/28] mm, page_alloc: Use __dec_zone_state for order-0 page allocation Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 06/28] mm, page_alloc: Use __dec_zone_state for order-0 page allocation Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 13:30 +0200
[PATCH 09/28] mm, page_alloc: Convert nr_fair_skipped to bool Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 09/28] mm, page_alloc: Convert nr_fair_skipped to bool Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 13:40 +0200
[PATCH 08/28] mm, page_alloc: Convert alloc_flags to unsigned Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
[PATCH 03/28] mm, page_alloc: Reduce branches in zone_statistics Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 03/28] mm, page_alloc: Reduce branches in zone_statistics Vlastimil Babka <vbabka@suse.cz> - 2016-04-25 13:20 +0200
[PATCH 05/28] mm, page_alloc: Inline the fast path of the zonelist iterator Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 05/28] mm, page_alloc: Inline the fast path of the zonelist iterator Vlastimil Babka <vbabka@suse.cz> - 2016-04-25 17:00 +0200
Re: [PATCH 05/28] mm, page_alloc: Inline the fast path of the zonelist iterator Mel Gorman <mgorman@techsingularity.net> - 2016-04-26 12:40 +0200
Re: [PATCH 05/28] mm, page_alloc: Inline the fast path of the zonelist iterator Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 13:10 +0200
[PATCH 10/28] mm, page_alloc: Remove unnecessary local variable in get_page_from_freelist Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 11:10 +0200
Re: [PATCH 10/28] mm, page_alloc: Remove unnecessary local variable in get_page_from_freelist Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 13:40 +0200
Re: [PATCH 00/28] Optimise page alloc/free fast paths v3 Jesper Dangaard Brouer <brouer@redhat.com> - 2016-04-15 14:50 +0200
Re: [PATCH 00/28] Optimise page alloc/free fast paths v3 Mel Gorman <mgorman@techsingularity.net> - 2016-04-15 15:10 +0200
[PATCH 12/28] mm, page_alloc: Remove unnecessary initialisation from __alloc_pages_nodemask() Mel Gorman <mgorman@techsingularity.net> - 2016-04-16 09:30 +0200
Re: [PATCH 12/28] mm, page_alloc: Remove unnecessary initialisation from __alloc_pages_nodemask() Vlastimil Babka <vbabka@suse.cz> - 2016-04-26 13:50 +0200
Page 4 of 4 — ← Prev page 1 2 3 [4]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-04-15 11:10 +0200 |
| Subject | [PATCH 05/28] mm, page_alloc: Inline the fast path of the zonelist iterator |
| Message-ID | <ro7Vi-Er-55@gated-at.bofh.it> |
| In reply to | #1379561 |
The page allocator iterates through a zonelist for zones that match
the addressing limitations and nodemask of the caller but many allocations
will not be restricted. Despite this, there is always functional call
overhead which builds up.
This patch inlines the optimistic basic case and only calls the
iterator function for the complex case. A hindrance was the fact that
cpuset_current_mems_allowed is used in the fastpath as the allowed nodemask
even though all nodes are allowed on most systems. The patch handles this
by only considering cpuset_current_mems_allowed if a cpuset exists. As well
as being faster in the fast-path, this removes some junk in the slowpath.
The performance difference on a page allocator microbenchmark is;
4.6.0-rc2 4.6.0-rc2
statinline-v1r20 optiter-v1r20
Min alloc-odr0-1 412.00 ( 0.00%) 382.00 ( 7.28%)
Min alloc-odr0-2 301.00 ( 0.00%) 282.00 ( 6.31%)
Min alloc-odr0-4 247.00 ( 0.00%) 233.00 ( 5.67%)
Min alloc-odr0-8 215.00 ( 0.00%) 203.00 ( 5.58%)
Min alloc-odr0-16 199.00 ( 0.00%) 188.00 ( 5.53%)
Min alloc-odr0-32 191.00 ( 0.00%) 182.00 ( 4.71%)
Min alloc-odr0-64 187.00 ( 0.00%) 177.00 ( 5.35%)
Min alloc-odr0-128 185.00 ( 0.00%) 175.00 ( 5.41%)
Min alloc-odr0-256 193.00 ( 0.00%) 184.00 ( 4.66%)
Min alloc-odr0-512 207.00 ( 0.00%) 197.00 ( 4.83%)
Min alloc-odr0-1024 213.00 ( 0.00%) 203.00 ( 4.69%)
Min alloc-odr0-2048 220.00 ( 0.00%) 209.00 ( 5.00%)
Min alloc-odr0-4096 226.00 ( 0.00%) 214.00 ( 5.31%)
Min alloc-odr0-8192 229.00 ( 0.00%) 218.00 ( 4.80%)
Min alloc-odr0-16384 229.00 ( 0.00%) 219.00 ( 4.37%)
perf indicated that next_zones_zonelist disappeared in the profile and
__next_zones_zonelist did not appear. This is expected as the micro-benchmark
would hit the inlined fast-path every time.
Signed-off-by: Mel Gorman <mgorman@techsingularity.net>
---
include/linux/mmzone.h | 13 +++++++++++--
mm/mmzone.c | 2 +-
mm/page_alloc.c | 26 +++++++++-----------------
3 files changed, 21 insertions(+), 20 deletions(-)
diff --git a/include/linux/mmzone.h b/include/linux/mmzone.h
index c60df9257cc7..0c4d5ebb3849 100644
--- a/include/linux/mmzone.h
+++ b/include/linux/mmzone.h
@@ -922,6 +922,10 @@ static inline int zonelist_node_idx(struct zoneref *zoneref)
#endif /* CONFIG_NUMA */
}
+struct zoneref *__next_zones_zonelist(struct zoneref *z,
+ enum zone_type highest_zoneidx,
+ nodemask_t *nodes);
+
/**
* next_zones_zonelist - Returns the next zone at or below highest_zoneidx within the allowed nodemask using a cursor within a zonelist as a starting point
* @z - The cursor used as a starting point for the search
@@ -934,9 +938,14 @@ static inline int zonelist_node_idx(struct zoneref *zoneref)
* being examined. It should be advanced by one before calling
* next_zones_zonelist again.
*/
-struct zoneref *next_zones_zonelist(struct zoneref *z,
+static __always_inline struct zoneref *next_zones_zonelist(struct zoneref *z,
enum zone_type highest_zoneidx,
- nodemask_t *nodes);
+ nodemask_t *nodes)
+{
+ if (likely(!nodes && zonelist_zone_idx(z) <= highest_zoneidx))
+ return z;
+ return __next_zones_zonelist(z, highest_zoneidx, nodes);
+}
/**
* first_zones_zonelist - Returns the first zone at or below highest_zoneidx within the allowed nodemask in a zonelist
diff --git a/mm/mmzone.c b/mm/mmzone.c
index 52687fb4de6f..5652be858e5e 100644
--- a/mm/mmzone.c
+++ b/mm/mmzone.c
@@ -52,7 +52,7 @@ static inline int zref_in_nodemask(struct zoneref *zref, nodemask_t *nodes)
}
/* Returns the next zone at or below highest_zoneidx in a zonelist */
-struct zoneref *next_zones_zonelist(struct zoneref *z,
+struct zoneref *__next_zones_zonelist(struct zoneref *z,
enum zone_type highest_zoneidx,
nodemask_t *nodes)
{
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index b56c2b2911a2..e9acc0b0f787 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -3193,17 +3193,6 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
*/
alloc_flags = gfp_to_alloc_flags(gfp_mask);
- /*
- * Find the true preferred zone if the allocation is unconstrained by
- * cpusets.
- */
- if (!(alloc_flags & ALLOC_CPUSET) && !ac->nodemask) {
- struct zoneref *preferred_zoneref;
- preferred_zoneref = first_zones_zonelist(ac->zonelist,
- ac->high_zoneidx, NULL, &ac->preferred_zone);
- ac->classzone_idx = zonelist_zone_idx(preferred_zoneref);
- }
-
/* This is the last chance, in general, before the goto nopage. */
page = get_page_from_freelist(gfp_mask, order,
alloc_flags & ~ALLOC_NO_WATERMARKS, ac);
@@ -3359,14 +3348,21 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
struct zoneref *preferred_zoneref;
struct page *page = NULL;
unsigned int cpuset_mems_cookie;
- int alloc_flags = ALLOC_WMARK_LOW|ALLOC_CPUSET|ALLOC_FAIR;
+ int alloc_flags = ALLOC_WMARK_LOW|ALLOC_FAIR;
gfp_t alloc_mask; /* The gfp_t that was actually used for allocation */
struct alloc_context ac = {
.high_zoneidx = gfp_zone(gfp_mask),
+ .zonelist = zonelist,
.nodemask = nodemask,
.migratetype = gfpflags_to_migratetype(gfp_mask),
};
+ if (cpusets_enabled()) {
+ alloc_flags |= ALLOC_CPUSET;
+ if (!ac.nodemask)
+ ac.nodemask = &cpuset_current_mems_allowed;
+ }
+
gfp_mask &= gfp_allowed_mask;
lockdep_trace_alloc(gfp_mask);
@@ -3390,16 +3386,12 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
retry_cpuset:
cpuset_mems_cookie = read_mems_allowed_begin();
- /* We set it here, as __alloc_pages_slowpath might have changed it */
- ac.zonelist = zonelist;
-
/* Dirty zone balancing only done in the fast path */
ac.spread_dirty_pages = (gfp_mask & __GFP_WRITE);
/* The preferred zone is used for statistics later */
preferred_zoneref = first_zones_zonelist(ac.zonelist, ac.high_zoneidx,
- ac.nodemask ? : &cpuset_current_mems_allowed,
- &ac.preferred_zone);
+ ac.nodemask, &ac.preferred_zone);
if (!ac.preferred_zone)
goto out;
ac.classzone_idx = zonelist_zone_idx(preferred_zoneref);
--
2.6.4
[toc] | [prev] | [next] | [standalone]
| From | Vlastimil Babka <vbabka@suse.cz> |
|---|---|
| Date | 2016-04-25 17:00 +0200 |
| Subject | Re: [PATCH 05/28] mm, page_alloc: Inline the fast path of the zonelist iterator |
| Message-ID | <rrQ9s-7Bm-17@gated-at.bofh.it> |
| In reply to | #1379586 |
On 04/15/2016 10:58 AM, Mel Gorman wrote:
> The page allocator iterates through a zonelist for zones that match
> the addressing limitations and nodemask of the caller but many allocations
> will not be restricted. Despite this, there is always functional call
> overhead which builds up.
>
> This patch inlines the optimistic basic case and only calls the
> iterator function for the complex case. A hindrance was the fact that
> cpuset_current_mems_allowed is used in the fastpath as the allowed nodemask
> even though all nodes are allowed on most systems. The patch handles this
> by only considering cpuset_current_mems_allowed if a cpuset exists. As well
> as being faster in the fast-path, this removes some junk in the slowpath.
I don't think this part is entirely correct (or at least argued as being
correct above), see below.
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -3193,17 +3193,6 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
> */
> alloc_flags = gfp_to_alloc_flags(gfp_mask);
>
> - /*
> - * Find the true preferred zone if the allocation is unconstrained by
> - * cpusets.
> - */
> - if (!(alloc_flags & ALLOC_CPUSET) && !ac->nodemask) {
> - struct zoneref *preferred_zoneref;
> - preferred_zoneref = first_zones_zonelist(ac->zonelist,
> - ac->high_zoneidx, NULL, &ac->preferred_zone);
> - ac->classzone_idx = zonelist_zone_idx(preferred_zoneref);
> - }
> -
> /* This is the last chance, in general, before the goto nopage. */
> page = get_page_from_freelist(gfp_mask, order,
> alloc_flags & ~ALLOC_NO_WATERMARKS, ac);
> @@ -3359,14 +3348,21 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
> struct zoneref *preferred_zoneref;
> struct page *page = NULL;
> unsigned int cpuset_mems_cookie;
> - int alloc_flags = ALLOC_WMARK_LOW|ALLOC_CPUSET|ALLOC_FAIR;
> + int alloc_flags = ALLOC_WMARK_LOW|ALLOC_FAIR;
> gfp_t alloc_mask; /* The gfp_t that was actually used for allocation */
> struct alloc_context ac = {
> .high_zoneidx = gfp_zone(gfp_mask),
> + .zonelist = zonelist,
> .nodemask = nodemask,
> .migratetype = gfpflags_to_migratetype(gfp_mask),
> };
>
> + if (cpusets_enabled()) {
> + alloc_flags |= ALLOC_CPUSET;
> + if (!ac.nodemask)
> + ac.nodemask = &cpuset_current_mems_allowed;
> + }
My initial reaction is that this is setting ac.nodemask in stone outside
of cpuset_mems_cookie, but I guess it's ok since we're taking a pointer
into current's task_struct, not the contents of the current's nodemask.
It's however setting a non-NULL nodemask into stone, which means no
zonelist iterator fasthpaths... but only in the slowpath. I guess it's
not an issue then.
> +
> gfp_mask &= gfp_allowed_mask;
>
> lockdep_trace_alloc(gfp_mask);
> @@ -3390,16 +3386,12 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
> retry_cpuset:
> cpuset_mems_cookie = read_mems_allowed_begin();
>
> - /* We set it here, as __alloc_pages_slowpath might have changed it */
> - ac.zonelist = zonelist;
This doesn't seem relevant to the preferred_zoneref changes in
__alloc_pages_slowpath, so why it became ok? Maybe it is, but it's not
clear from the changelog.
Anyway, thinking about it made me realize that maybe we could move the
whole mems_cookie thing into slowpath? As soon as the optimistic
fastpath succeeds, we don't check the cookie anyway, so what about
something like this on top?
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index d18061535c8b..07bf1065e7c9 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -3183,6 +3183,7 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
bool can_direct_reclaim = gfp_mask & __GFP_DIRECT_RECLAIM;
struct page *page = NULL;
int alloc_flags;
+ unsigned int cpuset_mems_cookie;
unsigned long pages_reclaimed = 0;
unsigned long did_some_progress;
enum migrate_mode migration_mode = MIGRATE_ASYNC;
@@ -3209,6 +3210,8 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
gfp_mask &= ~__GFP_ATOMIC;
retry:
+ cpuset_mems_cookie = read_mems_allowed_begin();
+
if (gfp_mask & __GFP_KSWAPD_RECLAIM)
wake_all_kswapds(order, ac);
@@ -3219,17 +3222,6 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
*/
alloc_flags = gfp_to_alloc_flags(gfp_mask);
- /*
- * Find the true preferred zone if the allocation is unconstrained by
- * cpusets.
- */
- if (!(alloc_flags & ALLOC_CPUSET) && !ac->nodemask) {
- struct zoneref *preferred_zoneref;
- preferred_zoneref = first_zones_zonelist(ac->zonelist,
- ac->high_zoneidx, NULL, &ac->preferred_zone);
- ac->classzone_idx = zonelist_zone_idx(preferred_zoneref);
- }
-
/* This is the last chance, in general, before the goto nopage. */
page = get_page_from_freelist(gfp_mask, order,
alloc_flags & ~ALLOC_NO_WATERMARKS, ac);
@@ -3370,7 +3362,17 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
if (page)
goto got_pg;
nopage:
+ /*
+ * When updating a task's mems_allowed, it is possible to race with
+ * parallel threads in such a way that an allocation can fail while
+ * the mask is being updated. If a page allocation is about to fail,
+ * check if the cpuset changed during allocation and if so, retry.
+ */
+ if (read_mems_allowed_retry(cpuset_mems_cookie))
+ goto retry;
+
warn_alloc_failed(gfp_mask, order, NULL);
+
got_pg:
return page;
}
@@ -3384,7 +3386,6 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
{
struct zoneref *preferred_zoneref;
struct page *page = NULL;
- unsigned int cpuset_mems_cookie;
int alloc_flags = ALLOC_WMARK_LOW|ALLOC_FAIR;
gfp_t alloc_mask; /* The gfp_t that was actually used for allocation */
struct alloc_context ac = {
@@ -3420,9 +3421,6 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
if (IS_ENABLED(CONFIG_CMA) && ac.migratetype == MIGRATE_MOVABLE)
alloc_flags |= ALLOC_CMA;
-retry_cpuset:
- cpuset_mems_cookie = read_mems_allowed_begin();
-
/* Dirty zone balancing only done in the fast path */
ac.spread_dirty_pages = (gfp_mask & __GFP_WRITE);
@@ -3430,13 +3428,15 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
preferred_zoneref = first_zones_zonelist(ac.zonelist, ac.high_zoneidx,
ac.nodemask, &ac.preferred_zone);
if (!ac.preferred_zone)
- goto out;
+ goto slowpath;
ac.classzone_idx = zonelist_zone_idx(preferred_zoneref);
/* First allocation attempt */
alloc_mask = gfp_mask|__GFP_HARDWALL;
page = get_page_from_freelist(alloc_mask, order, alloc_flags, &ac);
+
if (unlikely(!page)) {
+slowpath:
/*
* Runtime PM, block IO and its error handling path
* can deadlock because I/O on the device might not
@@ -3453,16 +3453,6 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
trace_mm_page_alloc(page, order, alloc_mask, ac.migratetype);
-out:
- /*
- * When updating a task's mems_allowed, it is possible to race with
- * parallel threads in such a way that an allocation can fail while
- * the mask is being updated. If a page allocation is about to fail,
- * check if the cpuset changed during allocation and if so, retry.
- */
- if (unlikely(!page && read_mems_allowed_retry(cpuset_mems_cookie)))
- goto retry_cpuset;
-
return page;
}
EXPORT_SYMBOL(__alloc_pages_nodemask);
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-04-26 12:40 +0200 |
| Subject | Re: [PATCH 05/28] mm, page_alloc: Inline the fast path of the zonelist iterator |
| Message-ID | <rs8zo-652-27@gated-at.bofh.it> |
| In reply to | #1386554 |
On Mon, Apr 25, 2016 at 04:50:18PM +0200, Vlastimil Babka wrote:
> > @@ -3193,17 +3193,6 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
> > */
> > alloc_flags = gfp_to_alloc_flags(gfp_mask);
> >
> > - /*
> > - * Find the true preferred zone if the allocation is unconstrained by
> > - * cpusets.
> > - */
> > - if (!(alloc_flags & ALLOC_CPUSET) && !ac->nodemask) {
> > - struct zoneref *preferred_zoneref;
> > - preferred_zoneref = first_zones_zonelist(ac->zonelist,
> > - ac->high_zoneidx, NULL, &ac->preferred_zone);
> > - ac->classzone_idx = zonelist_zone_idx(preferred_zoneref);
> > - }
> > -
> > /* This is the last chance, in general, before the goto nopage. */
> > page = get_page_from_freelist(gfp_mask, order,
> > alloc_flags & ~ALLOC_NO_WATERMARKS, ac);
> > @@ -3359,14 +3348,21 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
> > struct zoneref *preferred_zoneref;
> > struct page *page = NULL;
> > unsigned int cpuset_mems_cookie;
> > - int alloc_flags = ALLOC_WMARK_LOW|ALLOC_CPUSET|ALLOC_FAIR;
> > + int alloc_flags = ALLOC_WMARK_LOW|ALLOC_FAIR;
> > gfp_t alloc_mask; /* The gfp_t that was actually used for allocation */
> > struct alloc_context ac = {
> > .high_zoneidx = gfp_zone(gfp_mask),
> > + .zonelist = zonelist,
> > .nodemask = nodemask,
> > .migratetype = gfpflags_to_migratetype(gfp_mask),
> > };
> >
> > + if (cpusets_enabled()) {
> > + alloc_flags |= ALLOC_CPUSET;
> > + if (!ac.nodemask)
> > + ac.nodemask = &cpuset_current_mems_allowed;
> > + }
>
> My initial reaction is that this is setting ac.nodemask in stone outside
> of cpuset_mems_cookie, but I guess it's ok since we're taking a pointer
> into current's task_struct, not the contents of the current's nodemask.
> It's however setting a non-NULL nodemask into stone, which means no
> zonelist iterator fasthpaths... but only in the slowpath. I guess it's
> not an issue then.
>
You're right in that setting it in stone is problematic if the cpuset
nodemask changes duration allocation. The retry loop knows there is a
change but does not look it up which would loop once then potentially fail
unnecessarily. I should have moved the retry_cpuset label above the point
where cpuset_current_mems_allowed gets set. That's option 1 as a fixlet
to this patch.
> > +
> > gfp_mask &= gfp_allowed_mask;
> >
> > lockdep_trace_alloc(gfp_mask);
> > @@ -3390,16 +3386,12 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
> > retry_cpuset:
> > cpuset_mems_cookie = read_mems_allowed_begin();
> >
> > - /* We set it here, as __alloc_pages_slowpath might have changed it */
> > - ac.zonelist = zonelist;
>
> This doesn't seem relevant to the preferred_zoneref changes in
> __alloc_pages_slowpath, so why it became ok? Maybe it is, but it's not
> clear from the changelog.
>
The slowpath is no longer altering the preferred_zoneref.
> Anyway, thinking about it made me realize that maybe we could move the
> whole mems_cookie thing into slowpath? As soon as the optimistic
> fastpath succeeds, we don't check the cookie anyway, so what about
> something like this on top?
>
That in general would seem reasonable although I don't think it applies
to the series properly. Do you want to do this as a patch on top of the
series or will I use the fixlet for now and probably follow up with the
cookie move in a week or so when I've caught up after LSF/MM?
--
Mel Gorman
SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Vlastimil Babka <vbabka@suse.cz> |
|---|---|
| Date | 2016-04-26 13:10 +0200 |
| Subject | Re: [PATCH 05/28] mm, page_alloc: Inline the fast path of the zonelist iterator |
| Message-ID | <rs92p-6zI-17@gated-at.bofh.it> |
| In reply to | #1387346 |
On 04/26/2016 12:30 PM, Mel Gorman wrote:
> On Mon, Apr 25, 2016 at 04:50:18PM +0200, Vlastimil Babka wrote:
>> > @@ -3193,17 +3193,6 @@ __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
>> > */
>> > alloc_flags = gfp_to_alloc_flags(gfp_mask);
>> >
>> > - /*
>> > - * Find the true preferred zone if the allocation is unconstrained by
>> > - * cpusets.
>> > - */
>> > - if (!(alloc_flags & ALLOC_CPUSET) && !ac->nodemask) {
>> > - struct zoneref *preferred_zoneref;
>> > - preferred_zoneref = first_zones_zonelist(ac->zonelist,
>> > - ac->high_zoneidx, NULL, &ac->preferred_zone);
>> > - ac->classzone_idx = zonelist_zone_idx(preferred_zoneref);
>> > - }
>> > -
>> > /* This is the last chance, in general, before the goto nopage. */
>> > page = get_page_from_freelist(gfp_mask, order,
>> > alloc_flags & ~ALLOC_NO_WATERMARKS, ac);
>> > @@ -3359,14 +3348,21 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
>> > struct zoneref *preferred_zoneref;
>> > struct page *page = NULL;
>> > unsigned int cpuset_mems_cookie;
>> > - int alloc_flags = ALLOC_WMARK_LOW|ALLOC_CPUSET|ALLOC_FAIR;
>> > + int alloc_flags = ALLOC_WMARK_LOW|ALLOC_FAIR;
>> > gfp_t alloc_mask; /* The gfp_t that was actually used for allocation */
>> > struct alloc_context ac = {
>> > .high_zoneidx = gfp_zone(gfp_mask),
>> > + .zonelist = zonelist,
>> > .nodemask = nodemask,
>> > .migratetype = gfpflags_to_migratetype(gfp_mask),
>> > };
>> >
>> > + if (cpusets_enabled()) {
>> > + alloc_flags |= ALLOC_CPUSET;
>> > + if (!ac.nodemask)
>> > + ac.nodemask = &cpuset_current_mems_allowed;
>> > + }
>>
>> My initial reaction is that this is setting ac.nodemask in stone outside
>> of cpuset_mems_cookie, but I guess it's ok since we're taking a pointer
>> into current's task_struct, not the contents of the current's nodemask.
>> It's however setting a non-NULL nodemask into stone, which means no
>> zonelist iterator fasthpaths... but only in the slowpath. I guess it's
>> not an issue then.
>>
>
> You're right in that setting it in stone is problematic if the cpuset
> nodemask changes duration allocation. The retry loop knows there is a
> change but does not look it up which would loop once then potentially fail
> unnecessarily.
That's what I thought first, but I think the *pointer*
cpuset_current_mems_allowed itself doesn't change when cookie changes, only the
bitmask it points to, so changes in that bitmask should be seen. But it deserves
a comment maybe so people reading the code in future won't get the same suspicion.
> I should have moved the retry_cpuset label above the point
> where cpuset_current_mems_allowed gets set. That's option 1 as a fixlet
> to this patch.
>
>> > +
>> > gfp_mask &= gfp_allowed_mask;
>> >
>> > lockdep_trace_alloc(gfp_mask);
>> > @@ -3390,16 +3386,12 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
>> > retry_cpuset:
>> > cpuset_mems_cookie = read_mems_allowed_begin();
>> >
>> > - /* We set it here, as __alloc_pages_slowpath might have changed it */
>> > - ac.zonelist = zonelist;
>>
>> This doesn't seem relevant to the preferred_zoneref changes in
>> __alloc_pages_slowpath, so why it became ok? Maybe it is, but it's not
>> clear from the changelog.
>>
>
> The slowpath is no longer altering the preferred_zoneref.
But the hunk above is about ac.zonelist, not preferred_zoneref?
>
>> Anyway, thinking about it made me realize that maybe we could move the
>> whole mems_cookie thing into slowpath? As soon as the optimistic
>> fastpath succeeds, we don't check the cookie anyway, so what about
>> something like this on top?
>>
>
> That in general would seem reasonable although I don't think it applies
> to the series properly. Do you want to do this as a patch on top of the
> series or will I use the fixlet for now and probably follow up with the
> cookie move in a week or so when I've caught up after LSF/MM?
I guess fixlet is fine for now and you have better setup to test the effect (if
any) of the cookie move.
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-04-15 11:10 +0200 |
| Subject | [PATCH 10/28] mm, page_alloc: Remove unnecessary local variable in get_page_from_freelist |
| Message-ID | <ro7Vi-Er-57@gated-at.bofh.it> |
| In reply to | #1379561 |
zonelist here is a copy of a struct field that is used once. Ditch it.
Signed-off-by: Mel Gorman <mgorman@techsingularity.net>
---
mm/page_alloc.c | 3 +--
1 file changed, 1 insertion(+), 2 deletions(-)
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index e778485a64c1..313db1c43839 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -2673,7 +2673,6 @@ static struct page *
get_page_from_freelist(gfp_t gfp_mask, unsigned int order, int alloc_flags,
const struct alloc_context *ac)
{
- struct zonelist *zonelist = ac->zonelist;
struct zoneref *z;
struct page *page = NULL;
struct zone *zone;
@@ -2687,7 +2686,7 @@ get_page_from_freelist(gfp_t gfp_mask, unsigned int order, int alloc_flags,
* Scan zonelist, looking for a zone with enough free.
* See also __cpuset_node_allowed() comment in kernel/cpuset.c.
*/
- for_each_zone_zonelist_nodemask(zone, z, zonelist, ac->high_zoneidx,
+ for_each_zone_zonelist_nodemask(zone, z, ac->zonelist, ac->high_zoneidx,
ac->nodemask) {
unsigned long mark;
--
2.6.4
[toc] | [prev] | [next] | [standalone]
| From | Vlastimil Babka <vbabka@suse.cz> |
|---|---|
| Date | 2016-04-26 13:40 +0200 |
| Subject | Re: [PATCH 10/28] mm, page_alloc: Remove unnecessary local variable in get_page_from_freelist |
| Message-ID | <rs9vs-6Of-11@gated-at.bofh.it> |
| In reply to | #1379587 |
On 04/15/2016 10:59 AM, Mel Gorman wrote: > zonelist here is a copy of a struct field that is used once. Ditch it. > > Signed-off-by: Mel Gorman <mgorman@techsingularity.net> Acked-by: Vlastimil Babka <vbabka@suse.cz>
[toc] | [prev] | [next] | [standalone]
| From | Jesper Dangaard Brouer <brouer@redhat.com> |
|---|---|
| Date | 2016-04-15 14:50 +0200 |
| Message-ID | <robma-36m-21@gated-at.bofh.it> |
| In reply to | #1379561 |
On Fri, 15 Apr 2016 09:58:52 +0100
Mel Gorman <mgorman@techsingularity.net> wrote:
> There were no further responses to the last series but I kept going and
> added a few more small bits. Most are basic micro-optimisations. The last
> two patches weaken debugging checks to improve performance at the cost of
> delayed detection of some use-after-free and memory corruption bugs. If
> they make people uncomfortable, they can be dropped and the rest of the
> series stands on its own.
>
> Changelog since v2
> o Add more micro-optimisations
> o Weak debugging checks in favor of speed
>
[...]
>
> The overall impact on a page allocator microbenchmark for a range of orders
I also micro benchmarked this patchset. Avail via Mel Gorman's kernel tree:
http://git.kernel.org/cgit/linux/kernel/git/mel/linux.git
tested branch mm-vmscan-node-lru-v5r9 which also contain the node-lru series.
Tool:
https://github.com/netoptimizer/prototype-kernel/blob/master/kernel/mm/bench/page_bench01.c
Run as:
modprobe page_bench01; rmmod page_bench01 ; dmesg | tail -n40 | grep 'alloc_pages order'
Results kernel 4.6.0-rc1 :
alloc_pages order:0(4096B/x1) 272 cycles per-4096B 272 cycles
alloc_pages order:1(8192B/x2) 395 cycles per-4096B 197 cycles
alloc_pages order:2(16384B/x4) 433 cycles per-4096B 108 cycles
alloc_pages order:3(32768B/x8) 503 cycles per-4096B 62 cycles
alloc_pages order:4(65536B/x16) 682 cycles per-4096B 42 cycles
alloc_pages order:5(131072B/x32) 910 cycles per-4096B 28 cycles
alloc_pages order:6(262144B/x64) 1384 cycles per-4096B 21 cycles
alloc_pages order:7(524288B/x128) 2335 cycles per-4096B 18 cycles
alloc_pages order:8(1048576B/x256) 4108 cycles per-4096B 16 cycles
alloc_pages order:9(2097152B/x512) 8398 cycles per-4096B 16 cycles
After Mel Gorman's optimizations, results from mm-vmscan-node-lru-v5r::
alloc_pages order:0(4096B/x1) 231 cycles per-4096B 231 cycles
alloc_pages order:1(8192B/x2) 351 cycles per-4096B 175 cycles
alloc_pages order:2(16384B/x4) 357 cycles per-4096B 89 cycles
alloc_pages order:3(32768B/x8) 397 cycles per-4096B 49 cycles
alloc_pages order:4(65536B/x16) 481 cycles per-4096B 30 cycles
alloc_pages order:5(131072B/x32) 652 cycles per-4096B 20 cycles
alloc_pages order:6(262144B/x64) 1054 cycles per-4096B 16 cycles
alloc_pages order:7(524288B/x128) 1852 cycles per-4096B 14 cycles
alloc_pages order:8(1048576B/x256) 3156 cycles per-4096B 12 cycles
alloc_pages order:9(2097152B/x512) 6790 cycles per-4096B 13 cycles
I've also started doing some parallel concurrency testing workloads[1]
[1] https://github.com/netoptimizer/prototype-kernel/blob/master/kernel/mm/bench/page_bench03.c
Order-0 pages scale nicely:
Results kernel 4.6.0-rc1 :
Parallel-CPUs:1 page order:0(4096B/x1) ave 274 cycles per-4096B 274 cycles
Parallel-CPUs:2 page order:0(4096B/x1) ave 283 cycles per-4096B 283 cycles
Parallel-CPUs:3 page order:0(4096B/x1) ave 284 cycles per-4096B 284 cycles
Parallel-CPUs:4 page order:0(4096B/x1) ave 288 cycles per-4096B 288 cycles
Parallel-CPUs:5 page order:0(4096B/x1) ave 417 cycles per-4096B 417 cycles
Parallel-CPUs:6 page order:0(4096B/x1) ave 503 cycles per-4096B 503 cycles
Parallel-CPUs:7 page order:0(4096B/x1) ave 567 cycles per-4096B 567 cycles
Parallel-CPUs:8 page order:0(4096B/x1) ave 620 cycles per-4096B 620 cycles
And even better with you changes! :-))) This is great work!
Results from mm-vmscan-node-lru-v5r:
Parallel-CPUs:1 page order:0(4096B/x1) ave 246 cycles per-4096B 246 cycles
Parallel-CPUs:2 page order:0(4096B/x1) ave 251 cycles per-4096B 251 cycles
Parallel-CPUs:3 page order:0(4096B/x1) ave 254 cycles per-4096B 254 cycles
Parallel-CPUs:4 page order:0(4096B/x1) ave 258 cycles per-4096B 258 cycles
Parallel-CPUs:5 page order:0(4096B/x1) ave 313 cycles per-4096B 313 cycles
Parallel-CPUs:6 page order:0(4096B/x1) ave 369 cycles per-4096B 369 cycles
Parallel-CPUs:7 page order:0(4096B/x1) ave 379 cycles per-4096B 379 cycles
Parallel-CPUs:8 page order:0(4096B/x1) ave 399 cycles per-4096B 399 cycles
It does not seem that higher order page scale... and your patches does
not change this pattern.
Example order-3 pages, which is often used in the network stack:
Results kernel 4.6.0-rc1 ::
Parallel-CPUs:1 page order:3(32768B/x8) ave 524 cycles per-4096B 65 cycles
Parallel-CPUs:2 page order:3(32768B/x8) ave 2131 cycles per-4096B 266 cycles
Parallel-CPUs:3 page order:3(32768B/x8) ave 3885 cycles per-4096B 485 cycles
Parallel-CPUs:4 page order:3(32768B/x8) ave 4520 cycles per-4096B 565 cycles
Parallel-CPUs:5 page order:3(32768B/x8) ave 5604 cycles per-4096B 700 cycles
Parallel-CPUs:6 page order:3(32768B/x8) ave 7125 cycles per-4096B 890 cycles
Parallel-CPUs:7 page order:3(32768B/x8) ave 7883 cycles per-4096B 985 cycles
Parallel-CPUs:8 page order:3(32768B/x8) ave 9364 cycles per-4096B 1170 cycles
Results from mm-vmscan-node-lru-v5r:
Parallel-CPUs:1 page order:3(32768B/x8) ave 421 cycles per-4096B 52 cycles
Parallel-CPUs:2 page order:3(32768B/x8) ave 2236 cycles per-4096B 279 cycles
Parallel-CPUs:3 page order:3(32768B/x8) ave 3408 cycles per-4096B 426 cycles
Parallel-CPUs:4 page order:3(32768B/x8) ave 4687 cycles per-4096B 585 cycles
Parallel-CPUs:5 page order:3(32768B/x8) ave 5972 cycles per-4096B 746 cycles
Parallel-CPUs:6 page order:3(32768B/x8) ave 7349 cycles per-4096B 918 cycles
Parallel-CPUs:7 page order:3(32768B/x8) ave 8436 cycles per-4096B 1054 cycles
Parallel-CPUs:8 page order:3(32768B/x8) ave 9589 cycles per-4096B 1198 cycles
--
Best regards,
Jesper Dangaard Brouer
MSc.CS, Principal Kernel Engineer at Red Hat
Author of http://www.iptv-analyzer.org
LinkedIn: http://www.linkedin.com/in/brouer
https://github.com/netoptimizer/prototype-kernel/blob/master/kernel/mm/bench/page_bench03.c
for ORDER in $(seq 0 5) ; do \
for X in $(seq 1 8) ; do \
modprobe page_bench03 page_order=$ORDER parallel_cpus=$X run_flags=$((2#100)); \
rmmod page_bench03 ; dmesg | tail -n 3 | grep Parallel-CPUs ; \
done; \
done
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-04-15 15:10 +0200 |
| Message-ID | <robFv-3uA-5@gated-at.bofh.it> |
| In reply to | #1379795 |
On Fri, Apr 15, 2016 at 02:44:02PM +0200, Jesper Dangaard Brouer wrote: > On Fri, 15 Apr 2016 09:58:52 +0100 > Mel Gorman <mgorman@techsingularity.net> wrote: > > > There were no further responses to the last series but I kept going and > > added a few more small bits. Most are basic micro-optimisations. The last > > two patches weaken debugging checks to improve performance at the cost of > > delayed detection of some use-after-free and memory corruption bugs. If > > they make people uncomfortable, they can be dropped and the rest of the > > series stands on its own. > > > > Changelog since v2 > > o Add more micro-optimisations > > o Weak debugging checks in favor of speed > > > [...] > > > > The overall impact on a page allocator microbenchmark for a range of orders > > I also micro benchmarked this patchset. Avail via Mel Gorman's kernel tree: > http://git.kernel.org/cgit/linux/kernel/git/mel/linux.git > tested branch mm-vmscan-node-lru-v5r9 which also contain the node-lru series. > > Tool: > https://github.com/netoptimizer/prototype-kernel/blob/master/kernel/mm/bench/page_bench01.c > Run as: > modprobe page_bench01; rmmod page_bench01 ; dmesg | tail -n40 | grep 'alloc_pages order' > Thanks Jesper. > Results kernel 4.6.0-rc1 : > > alloc_pages order:0(4096B/x1) 272 cycles per-4096B 272 cycles > alloc_pages order:1(8192B/x2) 395 cycles per-4096B 197 cycles > alloc_pages order:2(16384B/x4) 433 cycles per-4096B 108 cycles > alloc_pages order:3(32768B/x8) 503 cycles per-4096B 62 cycles > alloc_pages order:4(65536B/x16) 682 cycles per-4096B 42 cycles > alloc_pages order:5(131072B/x32) 910 cycles per-4096B 28 cycles > alloc_pages order:6(262144B/x64) 1384 cycles per-4096B 21 cycles > alloc_pages order:7(524288B/x128) 2335 cycles per-4096B 18 cycles > alloc_pages order:8(1048576B/x256) 4108 cycles per-4096B 16 cycles > alloc_pages order:9(2097152B/x512) 8398 cycles per-4096B 16 cycles > > After Mel Gorman's optimizations, results from mm-vmscan-node-lru-v5r:: > > alloc_pages order:0(4096B/x1) 231 cycles per-4096B 231 cycles > alloc_pages order:1(8192B/x2) 351 cycles per-4096B 175 cycles > alloc_pages order:2(16384B/x4) 357 cycles per-4096B 89 cycles > alloc_pages order:3(32768B/x8) 397 cycles per-4096B 49 cycles > alloc_pages order:4(65536B/x16) 481 cycles per-4096B 30 cycles > alloc_pages order:5(131072B/x32) 652 cycles per-4096B 20 cycles > alloc_pages order:6(262144B/x64) 1054 cycles per-4096B 16 cycles > alloc_pages order:7(524288B/x128) 1852 cycles per-4096B 14 cycles > alloc_pages order:8(1048576B/x256) 3156 cycles per-4096B 12 cycles > alloc_pages order:9(2097152B/x512) 6790 cycles per-4096B 13 cycles > This is broadly in line with expectations. order-0 sees the biggest boost because that's what the series focused on. High-order allocations see some benefits but they're still going through the slower paths of the allocator so it's less obvious. I'm glad to see this independently verified. > > I've also started doing some parallel concurrency testing workloads[1] > [1] https://github.com/netoptimizer/prototype-kernel/blob/master/kernel/mm/bench/page_bench03.c > > Order-0 pages scale nicely: > > Results kernel 4.6.0-rc1 : > Parallel-CPUs:1 page order:0(4096B/x1) ave 274 cycles per-4096B 274 cycles > Parallel-CPUs:2 page order:0(4096B/x1) ave 283 cycles per-4096B 283 cycles > Parallel-CPUs:3 page order:0(4096B/x1) ave 284 cycles per-4096B 284 cycles > Parallel-CPUs:4 page order:0(4096B/x1) ave 288 cycles per-4096B 288 cycles > Parallel-CPUs:5 page order:0(4096B/x1) ave 417 cycles per-4096B 417 cycles > Parallel-CPUs:6 page order:0(4096B/x1) ave 503 cycles per-4096B 503 cycles > Parallel-CPUs:7 page order:0(4096B/x1) ave 567 cycles per-4096B 567 cycles > Parallel-CPUs:8 page order:0(4096B/x1) ave 620 cycles per-4096B 620 cycles > > And even better with you changes! :-))) This is great work! > > Results from mm-vmscan-node-lru-v5r: > Parallel-CPUs:1 page order:0(4096B/x1) ave 246 cycles per-4096B 246 cycles > Parallel-CPUs:2 page order:0(4096B/x1) ave 251 cycles per-4096B 251 cycles > Parallel-CPUs:3 page order:0(4096B/x1) ave 254 cycles per-4096B 254 cycles > Parallel-CPUs:4 page order:0(4096B/x1) ave 258 cycles per-4096B 258 cycles > Parallel-CPUs:5 page order:0(4096B/x1) ave 313 cycles per-4096B 313 cycles > Parallel-CPUs:6 page order:0(4096B/x1) ave 369 cycles per-4096B 369 cycles > Parallel-CPUs:7 page order:0(4096B/x1) ave 379 cycles per-4096B 379 cycles > Parallel-CPUs:8 page order:0(4096B/x1) ave 399 cycles per-4096B 399 cycles > Excellent, thanks! > > It does not seem that higher order page scale... and your patches does > not change this pattern. > > Example order-3 pages, which is often used in the network stack: > Unfortunately, this lack of scaling is expected. All the high-order allocations bypass the per-cpu allocator so multiple parallel requests will contend on the zone->lock. Technically, the per-cpu allocator could handle high-order pages but failures would require IPIs to drain the remote lists and the memory footprint would be high. Whatever about the memory footprint, sending IPIs on every allocation failure is going to cause undesirable latency spikes. The original design of the per-cpu allocator assumed that high-order allocations were rare. This expectation is partially violated by SLUB using high-order pages, the network layer using compound pages and also by the test case unfortunately. I'll put some thought into how it could be improved on the flight over to LSF/MM but right now, I'm not very optimistic that a solution will be simple. -- Mel Gorman SUSE Labs
[toc] | [prev] | [next] | [standalone]
| From | Mel Gorman <mgorman@techsingularity.net> |
|---|---|
| Date | 2016-04-16 09:30 +0200 |
| Subject | [PATCH 12/28] mm, page_alloc: Remove unnecessary initialisation from __alloc_pages_nodemask() |
| Message-ID | <rosQ1-8hF-3@gated-at.bofh.it> |
| In reply to | #1379561 |
page is guaranteed to be set before it is read with or without the
initialisation.
Signed-off-by: Mel Gorman <mgorman@techsingularity.net>
---
mm/page_alloc.c | 2 +-
1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index f5ddb342c967..df03ccc7f07c 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -3348,7 +3348,7 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
struct zonelist *zonelist, nodemask_t *nodemask)
{
struct zoneref *preferred_zoneref;
- struct page *page = NULL;
+ struct page *page;
unsigned int cpuset_mems_cookie;
unsigned int alloc_flags = ALLOC_WMARK_LOW|ALLOC_FAIR;
gfp_t alloc_mask; /* The gfp_t that was actually used for allocation */
--
2.6.4
[toc] | [prev] | [next] | [standalone]
| From | Vlastimil Babka <vbabka@suse.cz> |
|---|---|
| Date | 2016-04-26 13:50 +0200 |
| Subject | Re: [PATCH 12/28] mm, page_alloc: Remove unnecessary initialisation from __alloc_pages_nodemask() |
| Message-ID | <rs9F8-6Su-9@gated-at.bofh.it> |
| In reply to | #1380494 |
On 04/16/2016 09:21 AM, Mel Gorman wrote: > page is guaranteed to be set before it is read with or without the > initialisation. > > Signed-off-by: Mel Gorman <mgorman@techsingularity.net> Acked-by: Vlastimil Babka <vbabka@suse.cz>
[toc] | [prev] | [standalone]
Page 4 of 4 — ← Prev page 1 2 3 [4]
Back to top | Article view | linux.kernel
csiph-web