Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1556394 > unrolled thread

getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel

Started byGanapatrao Kulkarni <gpkulkarni@gmail.com>
First post2017-01-11 12:00 +0100
Last post2017-01-11 17:40 +0100
Articles 12 — 3 participants

Back to article view | Back to linux.kernel


Contents

  getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Ganapatrao Kulkarni <gpkulkarni@gmail.com> - 2017-01-11 12:00 +0100
    Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Vlastimil Babka <vbabka@suse.cz> - 2017-01-11 12:10 +0100
      Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Vlastimil Babka <vbabka@suse.cz> - 2017-01-11 13:40 +0100
        Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Michal Hocko <mhocko@kernel.org> - 2017-01-11 17:50 +0100
          Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Vlastimil Babka <vbabka@suse.cz> - 2017-01-12 12:20 +0100
            Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Ganapatrao Kulkarni <gpkulkarni@gmail.com> - 2017-01-13 05:40 +0100
              Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Vlastimil Babka <vbabka@suse.cz> - 2017-01-13 10:10 +0100
                Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Michal Hocko <mhocko@kernel.org> - 2017-01-13 17:00 +0100
                Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Ganapatrao Kulkarni <gpkulkarni@gmail.com> - 2017-01-16 11:50 +0100
                  Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Vlastimil Babka <vbabka@suse.cz> - 2017-01-16 14:30 +0100
      Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Michal Hocko <mhocko@kernel.org> - 2017-01-11 17:40 +0100
    Re: getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel Michal Hocko <mhocko@kernel.org> - 2017-01-11 17:40 +0100

#1556394 — getting oom/stalls for ltp test cpuset01 with latest/4.9 kernel

FromGanapatrao Kulkarni <gpkulkarni@gmail.com>
Date2017-01-11 12:00 +0100
Subjectgetting oom/stalls for ltp test cpuset01 with latest/4.9 kernel
Message-ID<sYoNk-5Rd-25@gated-at.bofh.it>
Hi,

we are seeing OOM/stalls messages when we run ltp cpuset01(cpuset01 -I
360) test for few minutes, even through the numa system has adequate
memory on both nodes.

this we have observed same on both arm64/thunderx numa and on x86 numa system!

using latest ltp from master branch version 20160920-197-gbc4d3db
and linux kernel version 4.9

is this known bug already?

below is the oops log:
[ 2280.275193] cgroup: new mount options do not match the existing
superblock, will be ignored
[ 2316.565940] cgroup: new mount options do not match the existing
superblock, will be ignored
[ 2393.388361] cpuset01: page allocation stalls for 10051ms, order:0,
mode:0x24280ca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO)
[ 2393.388371] CPU: 9 PID: 18188 Comm: cpuset01 Not tainted 4.9.0 #1
[ 2393.388373] Hardware name: Dell Inc. PowerEdge T630/0W9WXC, BIOS
1.0.4 08/29/2014
[ 2393.388374]  ffffc9000c1afba8 ffffffff813c771e ffffffff81a40be8
0000000000000001
[ 2393.388377]  ffffc9000c1afc30 ffffffff811b8c9a 024280ca00000202
ffffffff81a40be8
[ 2393.388380]  ffffc9000c1afbd0 0000000000000010 ffffc9000c1afc40
ffffc9000c1afbf0
[ 2393.388383] Call Trace:
[ 2393.388392]  [<ffffffff813c771e>] dump_stack+0x63/0x85
[ 2393.388397]  [<ffffffff811b8c9a>] warn_alloc+0x13a/0x170
[ 2393.388399]  [<ffffffff811b95c4>] __alloc_pages_slowpath+0x884/0xac0
[ 2393.388402]  [<ffffffff811b9ac5>] __alloc_pages_nodemask+0x2c5/0x310
[ 2393.388405]  [<ffffffff8120f663>] alloc_pages_vma+0xb3/0x260
[ 2393.388410]  [<ffffffff811e0534>] ? anon_vma_interval_tree_insert+0x84/0x90
[ 2393.388413]  [<ffffffff811ea42c>] handle_mm_fault+0x129c/0x1550
[ 2393.388417]  [<ffffffff813d65bb>] ? call_rwsem_wake+0x1b/0x30
[ 2393.388422]  [<ffffffff8106a362>] __do_page_fault+0x222/0x4b0
[ 2393.388424]  [<ffffffff8106a61f>] do_page_fault+0x2f/0x80
[ 2393.388429]  [<ffffffff817ca588>] page_fault+0x28/0x30
[ 2393.388431] Mem-Info:
[ 2393.388437] active_anon:92316 inactive_anon:21059 isolated_anon:32
 active_file:202031 inactive_file:137088 isolated_file:0
 unevictable:16 dirty:20 writeback:5883 unstable:0
 slab_reclaimable:40274 slab_unreclaimable:21605
 mapped:26819 shmem:28393 pagetables:11375 bounce:0
 free:5494728 free_pcp:549 free_cma:0
[ 2393.388446] Node 0 active_anon:310368kB inactive_anon:25684kB
active_file:807836kB inactive_file:548592kB unevictable:60kB
isolated(anon):0kB isolated(file):0kB mapped:101672kB dirty:80kB
writeback:148kB shmem:0kB shmem_thp: 0kB shmem_pmdmapped: 0kB
anon_thp: 25780kB writeback_tmp:0kB unstable:0kB pages_scanned:0
all_unreclaimable? no
[ 2393.388455] Node 1 active_anon:58896kB inactive_anon:58552kB
active_file:288kB inactive_file:0kB unevictable:4kB
isolated(anon):128kB isolated(file):0kB mapped:5604kB dirty:0kB
writeback:23384kB shmem:0kB shmem_thp: 0kB shmem_pmdmapped: 0kB
anon_thp: 87792kB writeback_tmp:0kB unstable:0kB pages_scanned:0
all_unreclaimable? no
[ 2393.388457] Node 1 Normal free:11937124kB min:45532kB low:62044kB
high:78556kB active_anon:58896kB inactive_anon:58552kB
active_file:288kB inactive_file:0kB unevictable:4kB
writepending:23384kB present:16777216kB managed:16512808kB mlocked:4kB
slab_reclaimable:37876kB slab_unreclaimable:44812kB
kernel_stack:4264kB pagetables:27612kB bounce:0kB free_pcp:2240kB
local_pcp:0kB free_cma:0kB
[ 2393.388462] lowmem_reserve[]: 0 0 0 0
[ 2393.388465] Node 1 Normal: 1179*4kB (UME) 1396*8kB (UME) 1193*16kB
(UME) 910*32kB (UME) 721*64kB (UME) 568*128kB (UME) 444*256kB (UME)
328*512kB (ME) 223*1024kB (UM) 138*2048kB (ME) 2676*4096kB (M) =
11936412kB
[ 2393.388479] Node 0 hugepages_total=4 hugepages_free=4
hugepages_surp=0 hugepages_size=1048576kB
[ 2393.388481] Node 1 hugepages_total=4 hugepages_free=4
hugepages_surp=0 hugepages_size=1048576kB
[ 2393.388481] 374277 total pagecache pages
[ 2393.388483] 6667 pages in swap cache
[ 2393.388484] Swap cache stats: add 101786, delete 95119, find 393/682
[ 2393.388485] Free swap  = 15979384kB
[ 2393.388485] Total swap = 16383996kB
[ 2393.388486] 8331071 pages RAM
[ 2393.388486] 0 pages HighMem/MovableOnly
[ 2393.388487] 152036 pages reserved
[ 2393.388487] 0 pages hwpoisoned
[ 2397.331098] cpuset01 invoked oom-killer:
gfp_mask=0x24280ca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO), nodemask=1,
order=0, oom_score_adj=0


[gkulkarni@xeon-numa ltp]$ numactl --hardware
available: 2 nodes (0-1)
node 0 cpus: 0 2 4 6 8 10 12 14 16 18 20 22
node 0 size: 15823 MB
node 0 free: 10211 MB
node 1 cpus: 1 3 5 7 9 11 13 15 17 19 21 23
node 1 size: 16125 MB
node 1 free: 11628 MB
node distances:
node   0   1
  0:  10  21
  1:  21  10


thanks
Ganapat

[toc] | [next] | [standalone]


#1556399

FromVlastimil Babka <vbabka@suse.cz>
Date2017-01-11 12:10 +0100
Message-ID<sYoX0-69r-11@gated-at.bofh.it>
In reply to#1556394
On 01/11/2017 11:50 AM, Ganapatrao Kulkarni wrote:
> Hi,
> 
> we are seeing OOM/stalls messages when we run ltp cpuset01(cpuset01 -I
> 360) test for few minutes, even through the numa system has adequate
> memory on both nodes.
> 
> this we have observed same on both arm64/thunderx numa and on x86 numa system!
> 
> using latest ltp from master branch version 20160920-197-gbc4d3db
> and linux kernel version 4.9
> 
> is this known bug already?

Probably not.

Is it possible that cpuset limits the process to one node, and numa
mempolicy to the other node?

> below is the oops log:
> [ 2280.275193] cgroup: new mount options do not match the existing
> superblock, will be ignored
> [ 2316.565940] cgroup: new mount options do not match the existing
> superblock, will be ignored
> [ 2393.388361] cpuset01: page allocation stalls for 10051ms, order:0,
> mode:0x24280ca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO)
> [ 2393.388371] CPU: 9 PID: 18188 Comm: cpuset01 Not tainted 4.9.0 #1
> [ 2393.388373] Hardware name: Dell Inc. PowerEdge T630/0W9WXC, BIOS
> 1.0.4 08/29/2014
> [ 2393.388374]  ffffc9000c1afba8 ffffffff813c771e ffffffff81a40be8
> 0000000000000001
> [ 2393.388377]  ffffc9000c1afc30 ffffffff811b8c9a 024280ca00000202
> ffffffff81a40be8
> [ 2393.388380]  ffffc9000c1afbd0 0000000000000010 ffffc9000c1afc40
> ffffc9000c1afbf0
> [ 2393.388383] Call Trace:
> [ 2393.388392]  [<ffffffff813c771e>] dump_stack+0x63/0x85
> [ 2393.388397]  [<ffffffff811b8c9a>] warn_alloc+0x13a/0x170
> [ 2393.388399]  [<ffffffff811b95c4>] __alloc_pages_slowpath+0x884/0xac0
> [ 2393.388402]  [<ffffffff811b9ac5>] __alloc_pages_nodemask+0x2c5/0x310
> [ 2393.388405]  [<ffffffff8120f663>] alloc_pages_vma+0xb3/0x260
> [ 2393.388410]  [<ffffffff811e0534>] ? anon_vma_interval_tree_insert+0x84/0x90
> [ 2393.388413]  [<ffffffff811ea42c>] handle_mm_fault+0x129c/0x1550
> [ 2393.388417]  [<ffffffff813d65bb>] ? call_rwsem_wake+0x1b/0x30
> [ 2393.388422]  [<ffffffff8106a362>] __do_page_fault+0x222/0x4b0
> [ 2393.388424]  [<ffffffff8106a61f>] do_page_fault+0x2f/0x80
> [ 2393.388429]  [<ffffffff817ca588>] page_fault+0x28/0x30
> [ 2393.388431] Mem-Info:
> [ 2393.388437] active_anon:92316 inactive_anon:21059 isolated_anon:32
>  active_file:202031 inactive_file:137088 isolated_file:0
>  unevictable:16 dirty:20 writeback:5883 unstable:0
>  slab_reclaimable:40274 slab_unreclaimable:21605
>  mapped:26819 shmem:28393 pagetables:11375 bounce:0
>  free:5494728 free_pcp:549 free_cma:0
> [ 2393.388446] Node 0 active_anon:310368kB inactive_anon:25684kB
> active_file:807836kB inactive_file:548592kB unevictable:60kB
> isolated(anon):0kB isolated(file):0kB mapped:101672kB dirty:80kB
> writeback:148kB shmem:0kB shmem_thp: 0kB shmem_pmdmapped: 0kB
> anon_thp: 25780kB writeback_tmp:0kB unstable:0kB pages_scanned:0
> all_unreclaimable? no
> [ 2393.388455] Node 1 active_anon:58896kB inactive_anon:58552kB
> active_file:288kB inactive_file:0kB unevictable:4kB
> isolated(anon):128kB isolated(file):0kB mapped:5604kB dirty:0kB
> writeback:23384kB shmem:0kB shmem_thp: 0kB shmem_pmdmapped: 0kB
> anon_thp: 87792kB writeback_tmp:0kB unstable:0kB pages_scanned:0
> all_unreclaimable? no
> [ 2393.388457] Node 1 Normal free:11937124kB min:45532kB low:62044kB
> high:78556kB active_anon:58896kB inactive_anon:58552kB
> active_file:288kB inactive_file:0kB unevictable:4kB
> writepending:23384kB present:16777216kB managed:16512808kB mlocked:4kB
> slab_reclaimable:37876kB slab_unreclaimable:44812kB
> kernel_stack:4264kB pagetables:27612kB bounce:0kB free_pcp:2240kB
> local_pcp:0kB free_cma:0kB
> [ 2393.388462] lowmem_reserve[]: 0 0 0 0
> [ 2393.388465] Node 1 Normal: 1179*4kB (UME) 1396*8kB (UME) 1193*16kB
> (UME) 910*32kB (UME) 721*64kB (UME) 568*128kB (UME) 444*256kB (UME)
> 328*512kB (ME) 223*1024kB (UM) 138*2048kB (ME) 2676*4096kB (M) =
> 11936412kB
> [ 2393.388479] Node 0 hugepages_total=4 hugepages_free=4
> hugepages_surp=0 hugepages_size=1048576kB
> [ 2393.388481] Node 1 hugepages_total=4 hugepages_free=4
> hugepages_surp=0 hugepages_size=1048576kB
> [ 2393.388481] 374277 total pagecache pages
> [ 2393.388483] 6667 pages in swap cache
> [ 2393.388484] Swap cache stats: add 101786, delete 95119, find 393/682
> [ 2393.388485] Free swap  = 15979384kB
> [ 2393.388485] Total swap = 16383996kB
> [ 2393.388486] 8331071 pages RAM
> [ 2393.388486] 0 pages HighMem/MovableOnly
> [ 2393.388487] 152036 pages reserved
> [ 2393.388487] 0 pages hwpoisoned
> [ 2397.331098] cpuset01 invoked oom-killer:
> gfp_mask=0x24280ca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO), nodemask=1,
> order=0, oom_score_adj=0
> 
> 
> [gkulkarni@xeon-numa ltp]$ numactl --hardware
> available: 2 nodes (0-1)
> node 0 cpus: 0 2 4 6 8 10 12 14 16 18 20 22
> node 0 size: 15823 MB
> node 0 free: 10211 MB
> node 1 cpus: 1 3 5 7 9 11 13 15 17 19 21 23
> node 1 size: 16125 MB
> node 1 free: 11628 MB
> node distances:
> node   0   1
>   0:  10  21
>   1:  21  10
> 
> 
> thanks
> Ganapat
> 
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org.  For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
> 

[toc] | [prev] | [next] | [standalone]


#1556451

FromVlastimil Babka <vbabka@suse.cz>
Date2017-01-11 13:40 +0100
Message-ID<sYqm6-6WF-19@gated-at.bofh.it>
In reply to#1556399
On 01/11/2017 12:05 PM, Vlastimil Babka wrote:
> On 01/11/2017 11:50 AM, Ganapatrao Kulkarni wrote:
>> Hi,
>>
>> we are seeing OOM/stalls messages when we run ltp cpuset01(cpuset01 -I
>> 360) test for few minutes, even through the numa system has adequate
>> memory on both nodes.
>>
>> this we have observed same on both arm64/thunderx numa and on x86 numa system!
>>
>> using latest ltp from master branch version 20160920-197-gbc4d3db
>> and linux kernel version 4.9
>>
>> is this known bug already?
> 
> Probably not.
> 
> Is it possible that cpuset limits the process to one node, and numa
> mempolicy to the other node?

Ah, so 4.9 has commit 82e7d3abec86 ("oom: print nodemask in the oom
report"), so such state should be visible in the oom report. Can you
post it whole instead of just the header line (i.e. the last line in
your report)? Thanks.

>> below is the oops log:
>> [ 2280.275193] cgroup: new mount options do not match the existing
>> superblock, will be ignored
>> [ 2316.565940] cgroup: new mount options do not match the existing
>> superblock, will be ignored
>> [ 2393.388361] cpuset01: page allocation stalls for 10051ms, order:0,
>> mode:0x24280ca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO)
>> [ 2393.388371] CPU: 9 PID: 18188 Comm: cpuset01 Not tainted 4.9.0 #1
>> [ 2393.388373] Hardware name: Dell Inc. PowerEdge T630/0W9WXC, BIOS
>> 1.0.4 08/29/2014
>> [ 2393.388374]  ffffc9000c1afba8 ffffffff813c771e ffffffff81a40be8
>> 0000000000000001
>> [ 2393.388377]  ffffc9000c1afc30 ffffffff811b8c9a 024280ca00000202
>> ffffffff81a40be8
>> [ 2393.388380]  ffffc9000c1afbd0 0000000000000010 ffffc9000c1afc40
>> ffffc9000c1afbf0
>> [ 2393.388383] Call Trace:
>> [ 2393.388392]  [<ffffffff813c771e>] dump_stack+0x63/0x85
>> [ 2393.388397]  [<ffffffff811b8c9a>] warn_alloc+0x13a/0x170
>> [ 2393.388399]  [<ffffffff811b95c4>] __alloc_pages_slowpath+0x884/0xac0
>> [ 2393.388402]  [<ffffffff811b9ac5>] __alloc_pages_nodemask+0x2c5/0x310
>> [ 2393.388405]  [<ffffffff8120f663>] alloc_pages_vma+0xb3/0x260
>> [ 2393.388410]  [<ffffffff811e0534>] ? anon_vma_interval_tree_insert+0x84/0x90
>> [ 2393.388413]  [<ffffffff811ea42c>] handle_mm_fault+0x129c/0x1550
>> [ 2393.388417]  [<ffffffff813d65bb>] ? call_rwsem_wake+0x1b/0x30
>> [ 2393.388422]  [<ffffffff8106a362>] __do_page_fault+0x222/0x4b0
>> [ 2393.388424]  [<ffffffff8106a61f>] do_page_fault+0x2f/0x80
>> [ 2393.388429]  [<ffffffff817ca588>] page_fault+0x28/0x30
>> [ 2393.388431] Mem-Info:
>> [ 2393.388437] active_anon:92316 inactive_anon:21059 isolated_anon:32
>>  active_file:202031 inactive_file:137088 isolated_file:0
>>  unevictable:16 dirty:20 writeback:5883 unstable:0
>>  slab_reclaimable:40274 slab_unreclaimable:21605
>>  mapped:26819 shmem:28393 pagetables:11375 bounce:0
>>  free:5494728 free_pcp:549 free_cma:0
>> [ 2393.388446] Node 0 active_anon:310368kB inactive_anon:25684kB
>> active_file:807836kB inactive_file:548592kB unevictable:60kB
>> isolated(anon):0kB isolated(file):0kB mapped:101672kB dirty:80kB
>> writeback:148kB shmem:0kB shmem_thp: 0kB shmem_pmdmapped: 0kB
>> anon_thp: 25780kB writeback_tmp:0kB unstable:0kB pages_scanned:0
>> all_unreclaimable? no
>> [ 2393.388455] Node 1 active_anon:58896kB inactive_anon:58552kB
>> active_file:288kB inactive_file:0kB unevictable:4kB
>> isolated(anon):128kB isolated(file):0kB mapped:5604kB dirty:0kB
>> writeback:23384kB shmem:0kB shmem_thp: 0kB shmem_pmdmapped: 0kB
>> anon_thp: 87792kB writeback_tmp:0kB unstable:0kB pages_scanned:0
>> all_unreclaimable? no
>> [ 2393.388457] Node 1 Normal free:11937124kB min:45532kB low:62044kB
>> high:78556kB active_anon:58896kB inactive_anon:58552kB
>> active_file:288kB inactive_file:0kB unevictable:4kB
>> writepending:23384kB present:16777216kB managed:16512808kB mlocked:4kB
>> slab_reclaimable:37876kB slab_unreclaimable:44812kB
>> kernel_stack:4264kB pagetables:27612kB bounce:0kB free_pcp:2240kB
>> local_pcp:0kB free_cma:0kB
>> [ 2393.388462] lowmem_reserve[]: 0 0 0 0
>> [ 2393.388465] Node 1 Normal: 1179*4kB (UME) 1396*8kB (UME) 1193*16kB
>> (UME) 910*32kB (UME) 721*64kB (UME) 568*128kB (UME) 444*256kB (UME)
>> 328*512kB (ME) 223*1024kB (UM) 138*2048kB (ME) 2676*4096kB (M) =
>> 11936412kB
>> [ 2393.388479] Node 0 hugepages_total=4 hugepages_free=4
>> hugepages_surp=0 hugepages_size=1048576kB
>> [ 2393.388481] Node 1 hugepages_total=4 hugepages_free=4
>> hugepages_surp=0 hugepages_size=1048576kB
>> [ 2393.388481] 374277 total pagecache pages
>> [ 2393.388483] 6667 pages in swap cache
>> [ 2393.388484] Swap cache stats: add 101786, delete 95119, find 393/682
>> [ 2393.388485] Free swap  = 15979384kB
>> [ 2393.388485] Total swap = 16383996kB
>> [ 2393.388486] 8331071 pages RAM
>> [ 2393.388486] 0 pages HighMem/MovableOnly
>> [ 2393.388487] 152036 pages reserved
>> [ 2393.388487] 0 pages hwpoisoned
>> [ 2397.331098] cpuset01 invoked oom-killer:
>> gfp_mask=0x24280ca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO), nodemask=1,
>> order=0, oom_score_adj=0
>>
>>
>> [gkulkarni@xeon-numa ltp]$ numactl --hardware
>> available: 2 nodes (0-1)
>> node 0 cpus: 0 2 4 6 8 10 12 14 16 18 20 22
>> node 0 size: 15823 MB
>> node 0 free: 10211 MB
>> node 1 cpus: 1 3 5 7 9 11 13 15 17 19 21 23
>> node 1 size: 16125 MB
>> node 1 free: 11628 MB
>> node distances:
>> node   0   1
>>   0:  10  21
>>   1:  21  10
>>
>>
>> thanks
>> Ganapat
>>
>> --
>> To unsubscribe, send a message with 'unsubscribe linux-mm' in
>> the body to majordomo@kvack.org.  For more info on Linux MM,
>> see: http://www.linux-mm.org/ .
>> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
>>
> 

[toc] | [prev] | [next] | [standalone]


#1556709

FromMichal Hocko <mhocko@kernel.org>
Date2017-01-11 17:50 +0100
Message-ID<sYug1-OR-15@gated-at.bofh.it>
In reply to#1556451
On Wed 11-01-17 21:52:29, Ganapatrao Kulkarni wrote:
[...]
> [ 2397.331098] cpuset01 invoked oom-killer: gfp_mask=0x24280ca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO), nodemask=1, order=0, oom_score_adj=0
> [ 2397.331100] cpuset01 cpuset=1 mems_allowed=1
[...]
> [ 2397.331206] Node 1 active_anon:5160kB inactive_anon:4968kB
> active_file:260kB inactive_file:0kB unevictable:4kB isolated(anon):0kB
> isolated(file):0kB mapped:1636kB dirty:0kB writeback:5164kB shmem:0kB
> shmem_thp: 0kB shmem_pmdmapped: 0kB anon_thp: 1624kB writeback_tmp:0kB
> unstable:0kB pages_scanned:17440 all_unreclaimable? yes

Hmm, so we consider the whole not unreclaimable...

> [ 2397.331208] Node 1 Normal free:12046572kB min:45532kB low:62044kB

while there is 12G of free memory. That sounds fishy...

> high:78556kB active_anon:5160kB inactive_anon:4968kB active_file:260kB
> inactive_file:0kB unevictable:4kB writepending:5164kB
> present:16777216kB managed:16512808kB mlocked:4kB
> slab_reclaimable:37876kB slab_unreclaimable:42904kB
> kernel_stack:4264kB pagetables:27612kB bounce:0kB free_pcp:1968kB
> local_pcp:0kB free_cma:0kB
[...]
> [ 2397.331236] Free swap  = 15892444kB
> [ 2397.331236] Total swap = 16383996kB

There is a lot of swap space free as well.

[...]
> [ 2398.146123] cpuset01 invoked oom-killer:  gfp_mask=0x24280ca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO), nodemask=1, order=0, oom_score_adj=0
> [ 2398.146124] cpuset01 cpuset=1 mems_allowed=1
[...]
> [ 2398.146217] Node 1 active_anon:3948kB inactive_anon:4736kB
> active_file:528kB inactive_file:204kB unevictable:4kB
> isolated(anon):0kB isolated(file):0kB mapped:1548kB dirty:0kB
> writeback:5100kB shmem:0kB shmem_thp: 0kB shmem_pmdmapped: 0kB
> anon_thp: 1724kB writeback_tmp:0kB unstable:0kB pages_scanned:16433
> all_unreclaimable? yes
> [ 2398.146220] Node 1 Normal free:12047352kB min:45532kB low:62044kB
> high:78556kB active_anon:3948kB inactive_anon:4736kB active_file:528kB
> inactive_file:204kB unevictable:4kB writepending:5100kB
> present:16777216kB managed:16512808kB mlocked:4kB
> slab_reclaimable:37876kB slab_unreclaimable:42856kB
> kernel_stack:4248kB pagetables:26644kB bounce:0kB free_pcp:1900kB
> local_pcp:120kB free_cma:0kB

Hmm, so there is another very similar oom report 1s later with similar
numbers. This doesn't look like a race when somehting else would free a
lot of memory at once. This smells like something different. Maybe we
cannot use any of the available pages for the allocation?

> [ 2398.169391] Node 1 Normal: 951*4kB (UME) 1308*8kB (UME) 1034*16kB (UME) 742*32kB (UME) 581*64kB (UME) 450*128kB (UME) 362*256kB (UME) 275*512kB (ME) 189*1024kB (UM) 117*2048kB (ME) 2742*4096kB (M) = 12047196kB

Most of the memblocks are marked Unmovable (except for the 4MB bloks)
which shouldn't matter because we can fallback to unmovable blocks for
movable allocation AFAIR so we shouldn't really fail the request. I
really fail to see what is going on there but it smells really
suspicious.
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1557372

FromVlastimil Babka <vbabka@suse.cz>
Date2017-01-12 12:20 +0100
Message-ID<sYLAe-3jp-25@gated-at.bofh.it>
In reply to#1556709
On 01/11/2017 05:46 PM, Michal Hocko wrote:
> On Wed 11-01-17 21:52:29, Ganapatrao Kulkarni wrote:
>
>> [ 2398.169391] Node 1 Normal: 951*4kB (UME) 1308*8kB (UME) 1034*16kB (UME) 742*32kB (UME) 581*64kB (UME) 450*128kB (UME) 362*256kB (UME) 275*512kB (ME) 189*1024kB (UM) 117*2048kB (ME) 2742*4096kB (M) = 12047196kB
>
> Most of the memblocks are marked Unmovable (except for the 4MB bloks)

No, UME here means that e.g. 4kB blocks are available on unmovable, movable and 
reclaimable lists.

> which shouldn't matter because we can fallback to unmovable blocks for
> movable allocation AFAIR so we shouldn't really fail the request. I
> really fail to see what is going on there but it smells really
> suspicious.

Perhaps there's something wrong with zonelists and we are skipping the Node 1 
Normal zone. Or there's some race with cpuset operations (but can't see how).

The question is, how reproducible is this? And what exactly the test cpuset01 
does? Is it doing multiple things in a loop that could be reduced to a single 
testcase?

[toc] | [prev] | [next] | [standalone]


#1558018

FromGanapatrao Kulkarni <gpkulkarni@gmail.com>
Date2017-01-13 05:40 +0100
Message-ID<sZ1OF-4L0-3@gated-at.bofh.it>
In reply to#1557372
On Thu, Jan 12, 2017 at 4:40 PM, Vlastimil Babka <vbabka@suse.cz> wrote:
> On 01/11/2017 05:46 PM, Michal Hocko wrote:
>>
>> On Wed 11-01-17 21:52:29, Ganapatrao Kulkarni wrote:
>>
>>> [ 2398.169391] Node 1 Normal: 951*4kB (UME) 1308*8kB (UME) 1034*16kB
>>> (UME) 742*32kB (UME) 581*64kB (UME) 450*128kB (UME) 362*256kB (UME)
>>> 275*512kB (ME) 189*1024kB (UM) 117*2048kB (ME) 2742*4096kB (M) = 12047196kB
>>
>>
>> Most of the memblocks are marked Unmovable (except for the 4MB bloks)
>
>
> No, UME here means that e.g. 4kB blocks are available on unmovable, movable
> and reclaimable lists.
>
>> which shouldn't matter because we can fallback to unmovable blocks for
>> movable allocation AFAIR so we shouldn't really fail the request. I
>> really fail to see what is going on there but it smells really
>> suspicious.
>
>
> Perhaps there's something wrong with zonelists and we are skipping the Node
> 1 Normal zone. Or there's some race with cpuset operations (but can't see
> how).
>
> The question is, how reproducible is this? And what exactly the test
> cpuset01 does? Is it doing multiple things in a loop that could be reduced
> to a single testcase?

IIUC, this test does node change to  cpuset.mems in loop in parent
process in loop and child processes(equal to no of cpus) keeps on
allocation and freeing
10 pages till the execution time is over.
more details at
https://github.com/linux-test-project/ltp/blob/master/testcases/kernel/mem/cpuset/cpuset01.c

thanks
Ganapat

>
>

[toc] | [prev] | [next] | [standalone]


#1558138

FromVlastimil Babka <vbabka@suse.cz>
Date2017-01-13 10:10 +0100
Message-ID<sZ61X-7rG-11@gated-at.bofh.it>
In reply to#1558018
On 01/13/2017 05:35 AM, Ganapatrao Kulkarni wrote:
> On Thu, Jan 12, 2017 at 4:40 PM, Vlastimil Babka <vbabka@suse.cz> wrote:
>> On 01/11/2017 05:46 PM, Michal Hocko wrote:
>>>
>>> On Wed 11-01-17 21:52:29, Ganapatrao Kulkarni wrote:
>>>
>>>> [ 2398.169391] Node 1 Normal: 951*4kB (UME) 1308*8kB (UME) 1034*16kB
>>>> (UME) 742*32kB (UME) 581*64kB (UME) 450*128kB (UME) 362*256kB (UME)
>>>> 275*512kB (ME) 189*1024kB (UM) 117*2048kB (ME) 2742*4096kB (M) = 12047196kB
>>>
>>>
>>> Most of the memblocks are marked Unmovable (except for the 4MB bloks)
>>
>>
>> No, UME here means that e.g. 4kB blocks are available on unmovable, movable
>> and reclaimable lists.
>>
>>> which shouldn't matter because we can fallback to unmovable blocks for
>>> movable allocation AFAIR so we shouldn't really fail the request. I
>>> really fail to see what is going on there but it smells really
>>> suspicious.
>>
>>
>> Perhaps there's something wrong with zonelists and we are skipping the Node
>> 1 Normal zone. Or there's some race with cpuset operations (but can't see
>> how).
>>
>> The question is, how reproducible is this? And what exactly the test
>> cpuset01 does? Is it doing multiple things in a loop that could be reduced
>> to a single testcase?
> 
> IIUC, this test does node change to  cpuset.mems in loop in parent
> process in loop and child processes(equal to no of cpus) keeps on
> allocation and freeing
> 10 pages till the execution time is over.
> more details at
> https://github.com/linux-test-project/ltp/blob/master/testcases/kernel/mem/cpuset/cpuset01.c

Ah, thanks for explaining. Looks like there might be a race where determining
ac.preferred_zone using current_mems_allowed as ac.nodemask skips the only zone
that is allowed after the cpuset.mems update, and we only recalculate
ac.preferred_zone for allocations that are allowed to escape cpusets/watermarks.
Thus we see only part of the zonelist, missing the only allowed zone. This would
be due to commit 682a3385e773 ("mm, page_alloc: inline the fast path of the
zonelist iterator") and/or some others from that series.

Could you try with the following patch please? It also tries to protect from
race with last non-root cpuset removal, which could cause cpusets_enable() to
become false in the middle of the function.

----8<----
From 9f041839401681f2678edf5040c851d11963c5fe Mon Sep 17 00:00:00 2001
From: Vlastimil Babka <vbabka@suse.cz>
Date: Fri, 13 Jan 2017 10:01:26 +0100
Subject: [PATCH] mm, page_alloc: fix race with cpuset update or removal

Changelog and S-O-B TBD.
---
 mm/page_alloc.c | 10 +++++++++-
 1 file changed, 9 insertions(+), 1 deletion(-)

diff --git a/mm/page_alloc.c b/mm/page_alloc.c
index 6de9440e3ae2..c397f146843a 100644
--- a/mm/page_alloc.c
+++ b/mm/page_alloc.c
@@ -3775,9 +3775,17 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
 	/*
 	 * Restore the original nodemask if it was potentially replaced with
 	 * &cpuset_current_mems_allowed to optimize the fast-path attempt.
+	 * Also recalculate the starting point for the zonelist iterator or
+	 * we could end up iterating over non-eligible zones endlessly.
 	 */
-	if (cpusets_enabled())
+	if (unlikely(ac.nodemask != nodemask)) {
 		ac.nodemask = nodemask;
+		ac.preferred_zoneref = first_zones_zonelist(ac.zonelist,
+						ac.high_zoneidx, ac.nodemask);
+		if (!ac.preferred_zoneref)
+			goto no_zone;
+	}
+
 	page = __alloc_pages_slowpath(alloc_mask, order, &ac);
 
 no_zone:
-- 
2.11.0

[toc] | [prev] | [next] | [standalone]


#1558518

FromMichal Hocko <mhocko@kernel.org>
Date2017-01-13 17:00 +0100
Message-ID<sZcqJ-2FK-5@gated-at.bofh.it>
In reply to#1558138
On Fri 13-01-17 10:06:14, Vlastimil Babka wrote:
[...]
> >From 9f041839401681f2678edf5040c851d11963c5fe Mon Sep 17 00:00:00 2001
> From: Vlastimil Babka <vbabka@suse.cz>
> Date: Fri, 13 Jan 2017 10:01:26 +0100
> Subject: [PATCH] mm, page_alloc: fix race with cpuset update or removal
> 
> Changelog and S-O-B TBD.
> ---
>  mm/page_alloc.c | 10 +++++++++-
>  1 file changed, 9 insertions(+), 1 deletion(-)
> 
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index 6de9440e3ae2..c397f146843a 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -3775,9 +3775,17 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
>  	/*
>  	 * Restore the original nodemask if it was potentially replaced with
>  	 * &cpuset_current_mems_allowed to optimize the fast-path attempt.
> +	 * Also recalculate the starting point for the zonelist iterator or
> +	 * we could end up iterating over non-eligible zones endlessly.
>  	 */
> -	if (cpusets_enabled())
> +	if (unlikely(ac.nodemask != nodemask)) {
>  		ac.nodemask = nodemask;
> +		ac.preferred_zoneref = first_zones_zonelist(ac.zonelist,
> +						ac.high_zoneidx, ac.nodemask);
> +		if (!ac.preferred_zoneref)
> +			goto no_zone;
> +	}
> +
>  	page = __alloc_pages_slowpath(alloc_mask, order, &ac);

I think you nailed it. It is really possible that preferred_zoneref is
outside of the cpuset_current_mems_allowed and if we are unlucky there
won't be any other zones on the zonelist...

-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1559631

FromGanapatrao Kulkarni <gpkulkarni@gmail.com>
Date2017-01-16 11:50 +0100
Message-ID<t0d1o-7E3-25@gated-at.bofh.it>
In reply to#1558138
On Fri, Jan 13, 2017 at 2:36 PM, Vlastimil Babka <vbabka@suse.cz> wrote:
> On 01/13/2017 05:35 AM, Ganapatrao Kulkarni wrote:
>> On Thu, Jan 12, 2017 at 4:40 PM, Vlastimil Babka <vbabka@suse.cz> wrote:
>>> On 01/11/2017 05:46 PM, Michal Hocko wrote:
>>>>
>>>> On Wed 11-01-17 21:52:29, Ganapatrao Kulkarni wrote:
>>>>
>>>>> [ 2398.169391] Node 1 Normal: 951*4kB (UME) 1308*8kB (UME) 1034*16kB
>>>>> (UME) 742*32kB (UME) 581*64kB (UME) 450*128kB (UME) 362*256kB (UME)
>>>>> 275*512kB (ME) 189*1024kB (UM) 117*2048kB (ME) 2742*4096kB (M) = 12047196kB
>>>>
>>>>
>>>> Most of the memblocks are marked Unmovable (except for the 4MB bloks)
>>>
>>>
>>> No, UME here means that e.g. 4kB blocks are available on unmovable, movable
>>> and reclaimable lists.
>>>
>>>> which shouldn't matter because we can fallback to unmovable blocks for
>>>> movable allocation AFAIR so we shouldn't really fail the request. I
>>>> really fail to see what is going on there but it smells really
>>>> suspicious.
>>>
>>>
>>> Perhaps there's something wrong with zonelists and we are skipping the Node
>>> 1 Normal zone. Or there's some race with cpuset operations (but can't see
>>> how).
>>>
>>> The question is, how reproducible is this? And what exactly the test
>>> cpuset01 does? Is it doing multiple things in a loop that could be reduced
>>> to a single testcase?
>>
>> IIUC, this test does node change to  cpuset.mems in loop in parent
>> process in loop and child processes(equal to no of cpus) keeps on
>> allocation and freeing
>> 10 pages till the execution time is over.
>> more details at
>> https://github.com/linux-test-project/ltp/blob/master/testcases/kernel/mem/cpuset/cpuset01.c
>
> Ah, thanks for explaining. Looks like there might be a race where determining
> ac.preferred_zone using current_mems_allowed as ac.nodemask skips the only zone
> that is allowed after the cpuset.mems update, and we only recalculate
> ac.preferred_zone for allocations that are allowed to escape cpusets/watermarks.
> Thus we see only part of the zonelist, missing the only allowed zone. This would
> be due to commit 682a3385e773 ("mm, page_alloc: inline the fast path of the
> zonelist iterator") and/or some others from that series.
>
> Could you try with the following patch please? It also tries to protect from
> race with last non-root cpuset removal, which could cause cpusets_enable() to
> become false in the middle of the function.
>
> ----8<----
> From 9f041839401681f2678edf5040c851d11963c5fe Mon Sep 17 00:00:00 2001
> From: Vlastimil Babka <vbabka@suse.cz>
> Date: Fri, 13 Jan 2017 10:01:26 +0100
> Subject: [PATCH] mm, page_alloc: fix race with cpuset update or removal
>
> Changelog and S-O-B TBD.
> ---
>  mm/page_alloc.c | 10 +++++++++-
>  1 file changed, 9 insertions(+), 1 deletion(-)
>
> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
> index 6de9440e3ae2..c397f146843a 100644
> --- a/mm/page_alloc.c
> +++ b/mm/page_alloc.c
> @@ -3775,9 +3775,17 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
>         /*
>          * Restore the original nodemask if it was potentially replaced with
>          * &cpuset_current_mems_allowed to optimize the fast-path attempt.
> +        * Also recalculate the starting point for the zonelist iterator or
> +        * we could end up iterating over non-eligible zones endlessly.
>          */
> -       if (cpusets_enabled())
> +       if (unlikely(ac.nodemask != nodemask)) {
>                 ac.nodemask = nodemask;
> +               ac.preferred_zoneref = first_zones_zonelist(ac.zonelist,
> +                                               ac.high_zoneidx, ac.nodemask);
> +               if (!ac.preferred_zoneref)
> +                       goto no_zone;
> +       }
> +
>         page = __alloc_pages_slowpath(alloc_mask, order, &ac);
>
>  no_zone:
> --
> 2.11.0
>

this patch did not fix the issue.
issue still exists!
i did bisect and this test passes in 4.4,4.5 and 4.6
test failing since 4.7-rc1

thanks
Ganapat
>
>
>

[toc] | [prev] | [next] | [standalone]


#1559731

FromVlastimil Babka <vbabka@suse.cz>
Date2017-01-16 14:30 +0100
Message-ID<t0fwf-WM-53@gated-at.bofh.it>
In reply to#1559631
On 01/16/2017 11:41 AM, Ganapatrao Kulkarni wrote:
> On Fri, Jan 13, 2017 at 2:36 PM, Vlastimil Babka <vbabka@suse.cz> wrote:
>> On 01/13/2017 05:35 AM, Ganapatrao Kulkarni wrote:
>>> On Thu, Jan 12, 2017 at 4:40 PM, Vlastimil Babka <vbabka@suse.cz> wrote:
>>>> On 01/11/2017 05:46 PM, Michal Hocko wrote:
>>>>>
>>>>> On Wed 11-01-17 21:52:29, Ganapatrao Kulkarni wrote:
>>>>>
>>>>>> [ 2398.169391] Node 1 Normal: 951*4kB (UME) 1308*8kB (UME) 1034*16kB
>>>>>> (UME) 742*32kB (UME) 581*64kB (UME) 450*128kB (UME) 362*256kB (UME)
>>>>>> 275*512kB (ME) 189*1024kB (UM) 117*2048kB (ME) 2742*4096kB (M) = 12047196kB
>>>>>
>>>>>
>>>>> Most of the memblocks are marked Unmovable (except for the 4MB bloks)
>>>>
>>>>
>>>> No, UME here means that e.g. 4kB blocks are available on unmovable, movable
>>>> and reclaimable lists.
>>>>
>>>>> which shouldn't matter because we can fallback to unmovable blocks for
>>>>> movable allocation AFAIR so we shouldn't really fail the request. I
>>>>> really fail to see what is going on there but it smells really
>>>>> suspicious.
>>>>
>>>>
>>>> Perhaps there's something wrong with zonelists and we are skipping the Node
>>>> 1 Normal zone. Or there's some race with cpuset operations (but can't see
>>>> how).
>>>>
>>>> The question is, how reproducible is this? And what exactly the test
>>>> cpuset01 does? Is it doing multiple things in a loop that could be reduced
>>>> to a single testcase?
>>>
>>> IIUC, this test does node change to  cpuset.mems in loop in parent
>>> process in loop and child processes(equal to no of cpus) keeps on
>>> allocation and freeing
>>> 10 pages till the execution time is over.
>>> more details at
>>> https://github.com/linux-test-project/ltp/blob/master/testcases/kernel/mem/cpuset/cpuset01.c
>>
>> Ah, thanks for explaining. Looks like there might be a race where determining
>> ac.preferred_zone using current_mems_allowed as ac.nodemask skips the only zone
>> that is allowed after the cpuset.mems update, and we only recalculate
>> ac.preferred_zone for allocations that are allowed to escape cpusets/watermarks.
>> Thus we see only part of the zonelist, missing the only allowed zone. This would
>> be due to commit 682a3385e773 ("mm, page_alloc: inline the fast path of the
>> zonelist iterator") and/or some others from that series.
>>
>> Could you try with the following patch please? It also tries to protect from
>> race with last non-root cpuset removal, which could cause cpusets_enable() to
>> become false in the middle of the function.
>>
>> ----8<----
>> From 9f041839401681f2678edf5040c851d11963c5fe Mon Sep 17 00:00:00 2001
>> From: Vlastimil Babka <vbabka@suse.cz>
>> Date: Fri, 13 Jan 2017 10:01:26 +0100
>> Subject: [PATCH] mm, page_alloc: fix race with cpuset update or removal
>>
>> Changelog and S-O-B TBD.
>> ---
>>  mm/page_alloc.c | 10 +++++++++-
>>  1 file changed, 9 insertions(+), 1 deletion(-)
>>
>> diff --git a/mm/page_alloc.c b/mm/page_alloc.c
>> index 6de9440e3ae2..c397f146843a 100644
>> --- a/mm/page_alloc.c
>> +++ b/mm/page_alloc.c
>> @@ -3775,9 +3775,17 @@ __alloc_pages_nodemask(gfp_t gfp_mask, unsigned int order,
>>         /*
>>          * Restore the original nodemask if it was potentially replaced with
>>          * &cpuset_current_mems_allowed to optimize the fast-path attempt.
>> +        * Also recalculate the starting point for the zonelist iterator or
>> +        * we could end up iterating over non-eligible zones endlessly.
>>          */
>> -       if (cpusets_enabled())
>> +       if (unlikely(ac.nodemask != nodemask)) {
>>                 ac.nodemask = nodemask;
>> +               ac.preferred_zoneref = first_zones_zonelist(ac.zonelist,
>> +                                               ac.high_zoneidx, ac.nodemask);
>> +               if (!ac.preferred_zoneref)
>> +                       goto no_zone;
>> +       }
>> +
>>         page = __alloc_pages_slowpath(alloc_mask, order, &ac);
>>
>>  no_zone:
>> --
>> 2.11.0
>>
> 
> this patch did not fix the issue.
> issue still exists!

Hmm, that's unfortunate.

> i did bisect and this test passes in 4.4,4.5 and 4.6
> test failing since 4.7-rc1

4.7 would match the commit I was trying to fix. But I don't see other
problems now. Could you bisect to a single commit then, to be sure? Thanks.

> thanks
> Ganapat
>>
>>
>>

[toc] | [prev] | [next] | [standalone]


#1556703

FromMichal Hocko <mhocko@kernel.org>
Date2017-01-11 17:40 +0100
Message-ID<sYu6l-Lo-23@gated-at.bofh.it>
In reply to#1556399
On Wed 11-01-17 12:05:44, Vlastimil Babka wrote:
> On 01/11/2017 11:50 AM, Ganapatrao Kulkarni wrote:
> > Hi,
> > 
> > we are seeing OOM/stalls messages when we run ltp cpuset01(cpuset01 -I
> > 360) test for few minutes, even through the numa system has adequate
> > memory on both nodes.
> > 
> > this we have observed same on both arm64/thunderx numa and on x86 numa system!
> > 
> > using latest ltp from master branch version 20160920-197-gbc4d3db
> > and linux kernel version 4.9
> > 
> > is this known bug already?
> 
> Probably not.
> 
> Is it possible that cpuset limits the process to one node, and numa
> mempolicy to the other node?

No this shouldn't happen AFAICS. It is more likely that there is an
unrelated memory pressure happenning at the same time.
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1556705

FromMichal Hocko <mhocko@kernel.org>
Date2017-01-11 17:40 +0100
Message-ID<sYu6m-Lo-31@gated-at.bofh.it>
In reply to#1556394
On Wed 11-01-17 16:20:45, Ganapatrao Kulkarni wrote:
> Hi,
> 
> we are seeing OOM/stalls messages when we run ltp cpuset01(cpuset01 -I
> 360) test for few minutes, even through the numa system has adequate
> memory on both nodes.
> 
> this we have observed same on both arm64/thunderx numa and on x86 numa system!
> 
> using latest ltp from master branch version 20160920-197-gbc4d3db
> and linux kernel version 4.9
> 
> is this known bug already?
> 
> below is the oops log:
> [ 2280.275193] cgroup: new mount options do not match the existing
> superblock, will be ignored
> [ 2316.565940] cgroup: new mount options do not match the existing
> superblock, will be ignored
> [ 2393.388361] cpuset01: page allocation stalls for 10051ms, order:0, mode:0x24280ca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO)

For some reason I thought we are printing the nodemask here. We are
not... Which sucks in situations like this. I will cook up a patch...

[...[
> [ 2393.388457] Node 1 Normal free:11937124kB min:45532kB low:62044kB
> high:78556kB active_anon:58896kB inactive_anon:58552kB
> active_file:288kB inactive_file:0kB unevictable:4kB
> writepending:23384kB present:16777216kB managed:16512808kB mlocked:4kB
> slab_reclaimable:37876kB slab_unreclaimable:44812kB
> kernel_stack:4264kB pagetables:27612kB bounce:0kB free_pcp:2240kB
> local_pcp:0kB free_cma:0kB

It seems that there is a lot of free memory in this node which seems to
be the only eligible one because there are no details about Node 0
zones. So there shouldn't be any real reason to stall this allocation.
Unless there was a huge memory pressure and the relief came only
recently when the current task just managed to get out of the reclaim
and report the stall.

Is there any other workload running on this system?
[...]
> [ 2397.331098] cpuset01 invoked oom-killer:
> gfp_mask=0x24280ca(GFP_HIGHUSER_MOVABLE|__GFP_ZERO), nodemask=1,
> order=0, oom_score_adj=0

Please attach the full oom report.
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web