Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1439233

[PATCH 12/34] mm: vmscan: do not reclaim from kswapd if there is any eligible zone

From Mel Gorman <mgorman@techsingularity.net>
Newsgroups linux.kernel
Subject [PATCH 12/34] mm: vmscan: do not reclaim from kswapd if there is any eligible zone
Date 2016-07-08 11:40 +0200
Message-ID <rSAqn-2Tv-51@gated-at.bofh.it> (permalink)
References <rSAqm-2Tv-3@gated-at.bofh.it>
Organization linux.* mail to news gateway

Show all headers | View raw


kswapd scans from highest to lowest for a zone that requires balancing.
This was necessary when reclaim was per-zone to fairly age pages on lower
zones.  Now that we are reclaiming on a per-node basis, any eligible zone
can be used and pages will still be aged fairly.  This patch avoids
reclaiming excessively unless buffer_heads are over the limit and it's
necessary to reclaim from a higher zone than requested by the waker of
kswapd to relieve low memory pressure.

[hillf.zj@alibaba-inc.com: Force kswapd reclaim no more than needed]
Link: http://lkml.kernel.org/r/1466518566-30034-12-git-send-email-mgorman@techsingularity.net
Signed-off-by: Mel Gorman <mgorman@techsingularity.net>
Signed-off-by: Hillf Danton <hillf.zj@alibaba-inc.com>
Acked-by: Vlastimil Babka <vbabka@suse.cz>
---
 mm/vmscan.c | 59 +++++++++++++++++++++++++++--------------------------------
 1 file changed, 27 insertions(+), 32 deletions(-)

diff --git a/mm/vmscan.c b/mm/vmscan.c
index 8b39b903bd14..b7a276f4b1b0 100644
--- a/mm/vmscan.c
+++ b/mm/vmscan.c
@@ -3144,31 +3144,39 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int classzone_idx)
 
 		sc.nr_reclaimed = 0;
 
-		/* Scan from the highest requested zone to dma */
-		for (i = classzone_idx; i >= 0; i--) {
-			zone = pgdat->node_zones + i;
-			if (!populated_zone(zone))
-				continue;
-
-			/*
-			 * If the number of buffer_heads in the machine
-			 * exceeds the maximum allowed level and this node
-			 * has a highmem zone, force kswapd to reclaim from
-			 * it to relieve lowmem pressure.
-			 */
-			if (buffer_heads_over_limit && is_highmem_idx(i)) {
-				classzone_idx = i;
-				break;
-			}
+		/*
+		 * If the number of buffer_heads in the machine exceeds the
+		 * maximum allowed level then reclaim from all zones. This is
+		 * not specific to highmem as highmem may not exist but it is
+		 * it is expected that buffer_heads are stripped in writeback.
+		 */
+		if (buffer_heads_over_limit) {
+			for (i = MAX_NR_ZONES - 1; i >= 0; i--) {
+				zone = pgdat->node_zones + i;
+				if (!populated_zone(zone))
+					continue;
 
-			if (!zone_balanced(zone, order, 0)) {
 				classzone_idx = i;
 				break;
 			}
 		}
 
-		if (i < 0)
-			goto out;
+		/*
+		 * Only reclaim if there are no eligible zones. Check from
+		 * high to low zone as allocations prefer higher zones.
+		 * Scanning from low to high zone would allow congestion to be
+		 * cleared during a very small window when a small low
+		 * zone was balanced even under extreme pressure when the
+		 * overall node may be congested.
+		 */
+		for (i = classzone_idx; i >= 0; i--) {
+			zone = pgdat->node_zones + i;
+			if (!populated_zone(zone))
+				continue;
+
+			if (zone_balanced(zone, sc.order, classzone_idx))
+				goto out;
+		}
 
 		/*
 		 * Do some background aging of the anon list, to give
@@ -3214,19 +3222,6 @@ static int balance_pgdat(pg_data_t *pgdat, int order, int classzone_idx)
 			break;
 
 		/*
-		 * Stop reclaiming if any eligible zone is balanced and clear
-		 * node writeback or congested.
-		 */
-		for (i = 0; i <= classzone_idx; i++) {
-			zone = pgdat->node_zones + i;
-			if (!populated_zone(zone))
-				continue;
-
-			if (zone_balanced(zone, sc.order, classzone_idx))
-				goto out;
-		}
-
-		/*
 		 * Raise priority if scanning rate is too low or there was no
 		 * progress in reclaiming pages
 		 */
-- 
2.6.4

Back to linux.kernel | Previous | NextPrevious in thread | Next in thread | Find similar | Unroll thread


Thread

[PATCH 00/34] Move LRU page reclaim from zones to nodes v9 Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 09/34] mm, vmscan: simplify the logic deciding whether kswapd sleeps Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 01/34] mm, vmstat: add infrastructure for per-node vmstats Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 13/34] mm, vmscan: make shrink_node decisions more node-centric Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 16/34] mm, page_alloc: consider dirtyable memory in terms of nodes Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 20/34] mm: move vmscan writes and file write accounting to the node Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 12/34] mm: vmscan: do not reclaim from kswapd if there is any eligible zone Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 24/34] mm, vmscan: avoid passing in classzone_idx unnecessarily to shrink_node Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 14/34] mm, memcg: move memcg limit enforcement from zones to nodes Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 04/34] mm, mmzone: clarify the usage of zone padding Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 02/34] mm, vmscan: move lru_lock to the node Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 15/34] mm, workingset: make working set detection node-aware Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 23/34] mm: convert zone_reclaim to node_reclaim Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:40 +0200
  [PATCH 07/34] mm, vmscan: make kswapd reclaim in terms of nodes Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 30/34] mm: page_alloc: cache the last node whose dirty limit is reached Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 31/34] mm: vmstat: replace __count_zone_vm_events with a zone id equivalent Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 34/34] mm, vmstat: remove zone and node double accounting by approximating retries Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 10/34] mm, vmscan: by default have direct reclaim only shrink once per node Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 32/34] mm: vmstat: account per-zone stalls and pages skipped during reclaim Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 25/34] mm, vmscan: avoid passing in classzone_idx unnecessarily to compaction_ready Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 28/34] mm, vmscan: add classzone information to tracepoints Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 27/34] mm, vmscan: Have kswapd reclaim from all zones if reclaiming and buffer_heads_over_limit Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 26/34] mm, vmscan: avoid passing in remaining unnecessarily to prepare_kswapd_sleep Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 29/34] mm, page_alloc: remove fair zone allocation policy Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 22/34] mm, page_alloc: wake kswapd based on the highest eligible zone Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 33/34] mm, vmstat: print node-based stats in zoneinfo file Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200
  [PATCH 19/34] mm: move most file-based accounting to the node Mel Gorman <mgorman@techsingularity.net> - 2016-07-08 11:50 +0200

csiph-web