mm: sl[au]b: add knowledge of PFMEMALLOC reserve pages

When a user or administrator requires swap for their application, they create a swap partition and file, format it with mkswap and activate it with swapon. Swap over the network is considered as an option in diskless systems. The two likely scenarios are when blade servers are used as part of a cluster where the form factor or maintenance costs do not allow the use of disks and thin clients. The Linux Terminal Server Project recommends the use of the Network Block Device (NBD) for swap according to the manual at https://sourceforge.net/projects/ltsp/files/Docs-Admin-Guide/LTSPManual.pdf/download There is also documentation and tutorials on how to setup swap over NBD at places like https://help.ubuntu.com/community/UbuntuLTSP/EnableNBDSWAP The nbd-client also documents the use of NBD as swap. Despite this, the fact is that a machine using NBD for swap can deadlock within minutes if swap is used intensively. This patch series addresses the problem. The core issue is that network block devices do not use mempools like normal block devices do. As the host cannot control where they receive packets from, they cannot reliably work out in advance how much memory they might need. Some years ago, Peter Zijlstra developed a series of patches that supported swap over an NFS that at least one distribution is carrying within their kernels. This patch series borrows very heavily from Peter's work to support swapping over NBD as a pre-requisite to supporting swap-over-NFS. The bulk of the complexity is concerned with preserving memory that is allocated from the PFMEMALLOC reserves for use by the network layer which is needed for both NBD and NFS. Patch 1 adds knowledge of the PFMEMALLOC reserves to SLAB and SLUB to preserve access to pages allocated under low memory situations to callers that are freeing memory. Patch 2 optimises the SLUB fast path to avoid pfmemalloc checks Patch 3 introduces __GFP_MEMALLOC to allow access to the PFMEMALLOC reserves without setting PFMEMALLOC. Patch 4 opens the possibility for softirqs to use PFMEMALLOC reserves for later use by network packet processing. Patch 5 only sets page->pfmemalloc when ALLOC_NO_WATERMARKS was required Patch 6 ignores memory policies when ALLOC_NO_WATERMARKS is set. Patches 7-12 allows network processing to use PFMEMALLOC reserves when the socket has been marked as being used by the VM to clean pages. If packets are received and stored in pages that were allocated under low-memory situations and are unrelated to the VM, the packets are dropped. Patch 11 reintroduces __skb_alloc_page which the networking folk may object to but is needed in some cases to propogate pfmemalloc from a newly allocated page to an skb. If there is a strong objection, this patch can be dropped with the impact being that swap-over-network will be slower in some cases but it should not fail. Patch 13 is a micro-optimisation to avoid a function call in the common case. Patch 14 tags NBD sockets as being SOCK_MEMALLOC so they can use PFMEMALLOC if necessary. Patch 15 notes that it is still possible for the PFMEMALLOC reserve to be depleted. To prevent this, direct reclaimers get throttled on a waitqueue if 50% of the PFMEMALLOC reserves are depleted. It is expected that kswapd and the direct reclaimers already running will clean enough pages for the low watermark to be reached and the throttled processes are woken up. Patch 16 adds a statistic to track how often processes get throttled Some basic performance testing was run using kernel builds, netperf on loopback for UDP and TCP, hackbench (pipes and sockets), iozone and sysbench. Each of them were expected to use the sl*b allocators reasonably heavily but there did not appear to be significant performance variances. For testing swap-over-NBD, a machine was booted with 2G of RAM with a swapfile backed by NBD. 8*NUM_CPU processes were started that create anonymous memory mappings and read them linearly in a loop. The total size of the mappings were 4*PHYSICAL_MEMORY to use swap heavily under memory pressure. Without the patches and using SLUB, the machine locks up within minutes and runs to completion with them applied. With SLAB, the story is different as an unpatched kernel run to completion. However, the patched kernel completed the test 45% faster. MICRO 3.5.0-rc2 3.5.0-rc2 vanilla swapnbd Unrecognised test vmscan-anon-mmap-write MMTests Statistics: duration Sys Time Running Test (seconds) 197.80 173.07 User+Sys Time Running Test (seconds) 206.96 182.03 Total Elapsed Time (seconds) 3240.70 1762.09 This patch: mm: sl[au]b: add knowledge of PFMEMALLOC reserve pages Allocations of pages below the min watermark run a risk of the machine hanging due to a lack of memory. To prevent this, only callers who have PF_MEMALLOC or TIF_MEMDIE set and are not processing an interrupt are allowed to allocate with ALLOC_NO_WATERMARKS. Once they are allocated to a slab though, nothing prevents other callers consuming free objects within those slabs. This patch limits access to slab pages that were alloced from the PFMEMALLOC reserves. When this patch is applied, pages allocated from below the low watermark are returned with page->pfmemalloc set and it is up to the caller to determine how the page should be protected. SLAB restricts access to any page with page->pfmemalloc set to callers which are known to able to access the PFMEMALLOC reserve. If one is not available, an attempt is made to allocate a new page rather than use a reserve. SLUB is a bit more relaxed in that it only records if the current per-CPU page was allocated from PFMEMALLOC reserve and uses another partial slab if the caller does not have the necessary GFP or process flags. This was found to be sufficient in tests to avoid hangs due to SLUB generally maintaining smaller lists than SLAB. In low-memory conditions it does mean that !PFMEMALLOC allocators can fail a slab allocation even though free objects are available because they are being preserved for callers that are freeing pages. [a.p.zijlstra@chello.nl: Original implementation] [sebastian@breakpoint.cc: Correct order of page flag clearing] Signed-off-by: Mel Gorman <mgorman@suse.de> Cc: David Miller <davem@davemloft.net> Cc: Neil Brown <neilb@suse.de> Cc: Peter Zijlstra <a.p.zijlstra@chello.nl> Cc: Mike Christie <michaelc@cs.wisc.edu> Cc: Eric B Munson <emunson@mgebm.net> Cc: Eric Dumazet <eric.dumazet@gmail.com> Cc: Sebastian Andrzej Siewior <sebastian@breakpoint.cc> Cc: Mel Gorman <mgorman@suse.de> Cc: Christoph Lameter <cl@linux.com> Signed-off-by: Andrew Morton <akpm@linux-foundation.org> Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
author: Mel Gorman <mgorman@suse.de> 2012-07-31 19:43:58 -0400
committer: Linus Torvalds <torvalds@linux-foundation.org> 2012-07-31 21:42:45 -0400
commit: 072bb0aa5e062902968c5c1007bba332c7820cf4 (patch)
tree: 1b4a602c16b07a41484c0664d1936848387f0916 /mm/page_alloc.c
parent: 702d1a6e0766d45642c934444fd41f658d251305 (diff)
1 files changed, 22 insertions, 5 deletions
diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 6a29ed8e6e60..38e5be65f24e 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -1513,6 +1513,7 @@ failed:
 #define ALLOC_HARDER            0x10 /* try to alloc harder */
 #define ALLOC_HIGH              0x20 /* __GFP_HIGH set */
 #define ALLOC_CPUSET            0x40 /* check for correct cpuset */
+#define ALLOC_PFMEMALLOC        0x80 /* Caller has PF_MEMALLOC set */
 #ifdef CONFIG_FAIL_PAGE_ALLOC
@@ -2293,16 +2294,22 @@ gfp_to_alloc_flags(gfp_t gfp_mask)
        } else if (unlikely(rt_task(current)) && !in_interrupt())
                alloc_flags |= ALLOC_HARDER;
-        if (likely(!(gfp_mask & __GFP_NOMEMALLOC))) {
+        if ((current->flags & PF_MEMALLOC) ||
-                if (!in_interrupt() &&
+                        unlikely(test_thread_flag(TIF_MEMDIE))) {
-                    ((current->flags & PF_MEMALLOC) ||
+                alloc_flags |= ALLOC_PFMEMALLOC;
-                     unlikely(test_thread_flag(TIF_MEMDIE))))
+                if (likely(!(gfp_mask & __GFP_NOMEMALLOC)) && !in_interrupt())
                        alloc_flags |= ALLOC_NO_WATERMARKS;
        }
        return alloc_flags;
 }
+bool gfp_pfmemalloc_allowed(gfp_t gfp_mask)
+{
+        return !!(gfp_to_alloc_flags(gfp_mask) & ALLOC_PFMEMALLOC);
+}
 static inline struct page *
 __alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
        struct zonelist *zonelist, enum zone_type high_zoneidx,
@@ -2490,10 +2497,18 @@ nopage:
        warn_alloc_failed(gfp_mask, order, NULL);
        return page;
 got_pg:
+        /*
+         * page->pfmemalloc is set when the caller had PFMEMALLOC set or is
+         * been OOM killed. The expectation is that the caller is taking
+         * steps that will free more memory. The caller should avoid the
+         * page being used for !PFMEMALLOC purposes.
+         */
+        page->pfmemalloc = !!(alloc_flags & ALLOC_PFMEMALLOC);
        if (kmemcheck_enabled)
                kmemcheck_pagealloc_alloc(page, order, gfp_mask);
-        return page;
+        return page;
 }
 /*
@@ -2544,6 +2559,8 @@ retry_cpuset:
                page = __alloc_pages_slowpath(gfp_mask, order,
                                zonelist, high_zoneidx, nodemask,
                                preferred_zone, migratetype);
+        else
+                page->pfmemalloc = false;
        trace_mm_page_alloc(page, order, gfp_mask, migratetype);
author	Mel Gorman <mgorman@suse.de>	2012-07-31 19:43:58 -0400
committer	Linus Torvalds <torvalds@linux-foundation.org>	2012-07-31 21:42:45 -0400
commit	072bb0aa5e062902968c5c1007bba332c7820cf4 (patch)
tree	1b4a602c16b07a41484c0664d1936848387f0916 /mm/page_alloc.c
parent	702d1a6e0766d45642c934444fd41f658d251305 (diff)

diff --git a/mm/page_alloc.c b/mm/page_alloc.c index 6a29ed8e6e60..38e5be65f24e 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c
@@ -1513,6 +1513,7 @@ failed:
1513	#define ALLOC_HARDER 0x10 /* try to alloc harder */	1513	#define ALLOC_HARDER 0x10 /* try to alloc harder */
1514	#define ALLOC_HIGH 0x20 /* __GFP_HIGH set */	1514	#define ALLOC_HIGH 0x20 /* __GFP_HIGH set */
1515	#define ALLOC_CPUSET 0x40 /* check for correct cpuset */	1515	#define ALLOC_CPUSET 0x40 /* check for correct cpuset */
		1516	#define ALLOC_PFMEMALLOC 0x80 /* Caller has PF_MEMALLOC set */
1516		1517
1517	#ifdef CONFIG_FAIL_PAGE_ALLOC	1518	#ifdef CONFIG_FAIL_PAGE_ALLOC
1518		1519
@@ -2293,16 +2294,22 @@ gfp_to_alloc_flags(gfp_t gfp_mask)
2293	} else if (unlikely(rt_task(current)) && !in_interrupt())	2294	} else if (unlikely(rt_task(current)) && !in_interrupt())
2294	alloc_flags \|= ALLOC_HARDER;	2295	alloc_flags \|= ALLOC_HARDER;
2295		2296
2296	if (likely(!(gfp_mask & __GFP_NOMEMALLOC))) {	2297	if ((current->flags & PF_MEMALLOC) \|\|
2297	if (!in_interrupt() &&	2298	unlikely(test_thread_flag(TIF_MEMDIE))) {
2298	((current->flags & PF_MEMALLOC) \|\|	2299	alloc_flags \|= ALLOC_PFMEMALLOC;
2299	unlikely(test_thread_flag(TIF_MEMDIE))))	2300
		2301	if (likely(!(gfp_mask & __GFP_NOMEMALLOC)) && !in_interrupt())
2300	alloc_flags \|= ALLOC_NO_WATERMARKS;	2302	alloc_flags \|= ALLOC_NO_WATERMARKS;
2301	}	2303	}
2302		2304
2303	return alloc_flags;	2305	return alloc_flags;
2304	}	2306	}
2305		2307
		2308	bool gfp_pfmemalloc_allowed(gfp_t gfp_mask)
		2309	{
		2310	return !!(gfp_to_alloc_flags(gfp_mask) & ALLOC_PFMEMALLOC);
		2311	}
		2312
2306	static inline struct page *	2313	static inline struct page *
2307	__alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,	2314	__alloc_pages_slowpath(gfp_t gfp_mask, unsigned int order,
2308	struct zonelist *zonelist, enum zone_type high_zoneidx,	2315	struct zonelist *zonelist, enum zone_type high_zoneidx,
@@ -2490,10 +2497,18 @@ nopage:
2490	warn_alloc_failed(gfp_mask, order, NULL);	2497	warn_alloc_failed(gfp_mask, order, NULL);
2491	return page;	2498	return page;
2492	got_pg:	2499	got_pg:
		2500	/*
		2501	* page->pfmemalloc is set when the caller had PFMEMALLOC set or is
		2502	* been OOM killed. The expectation is that the caller is taking
		2503	* steps that will free more memory. The caller should avoid the
		2504	* page being used for !PFMEMALLOC purposes.
		2505	*/
		2506	page->pfmemalloc = !!(alloc_flags & ALLOC_PFMEMALLOC);
		2507
2493	if (kmemcheck_enabled)	2508	if (kmemcheck_enabled)
2494	kmemcheck_pagealloc_alloc(page, order, gfp_mask);	2509	kmemcheck_pagealloc_alloc(page, order, gfp_mask);
2495	return page;
2496		2510
		2511	return page;
2497	}	2512	}
2498		2513
2499	/*	2514	/*
@@ -2544,6 +2559,8 @@ retry_cpuset:
2544	page = __alloc_pages_slowpath(gfp_mask, order,	2559	page = __alloc_pages_slowpath(gfp_mask, order,
2545	zonelist, high_zoneidx, nodemask,	2560	zonelist, high_zoneidx, nodemask,
2546	preferred_zone, migratetype);	2561	preferred_zone, migratetype);
		2562	else
		2563	page->pfmemalloc = false;
2547		2564
2548	trace_mm_page_alloc(page, order, gfp_mask, migratetype);	2565	trace_mm_page_alloc(page, order, gfp_mask, migratetype);
2549		2566