Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1644028 > unrolled thread

mm, something wring in page_lock_anon_vma_read()?

Started byXishi Qiu <qiuxishi@huawei.com>
First post2017-05-18 12:00 +0200
Last post2017-05-23 05:00 +0200
Articles 17 — 4 participants

Back to article view | Back to linux.kernel


Contents

  mm, something wring in page_lock_anon_vma_read()? Xishi Qiu <qiuxishi@huawei.com> - 2017-05-18 12:00 +0200
    Re: mm, something wring in page_lock_anon_vma_read()? Xishi Qiu <qiuxishi@huawei.com> - 2017-05-19 11:00 +0200
      Re: mm, something wring in page_lock_anon_vma_read()? Xishi Qiu <qiuxishi@huawei.com> - 2017-05-19 11:50 +0200
        Re: mm, something wring in page_lock_anon_vma_read()? Hugh Dickins <hughd@google.com> - 2017-05-20 00:10 +0200
          Re: mm, something wring in page_lock_anon_vma_read()? Xishi Qiu <qiuxishi@huawei.com> - 2017-05-20 03:30 +0200
            Re: mm, something wring in page_lock_anon_vma_read()? Hugh Dickins <hughd@google.com> - 2017-05-20 04:10 +0200
              Re: mm, something wring in page_lock_anon_vma_read()? Xishi Qiu <qiuxishi@huawei.com> - 2017-05-20 04:30 +0200
                Re: mm, something wring in page_lock_anon_vma_read()? Hugh Dickins <hughd@google.com> - 2017-05-20 04:50 +0200
                  Re: mm, something wring in page_lock_anon_vma_read()? zhong jiang <zhongjiang@huawei.com> - 2017-05-20 05:10 +0200
                    Re: mm, something wring in page_lock_anon_vma_read()? Vlastimil Babka <vbabka@suse.cz> - 2017-05-22 19:00 +0200
                      Re: mm, something wring in page_lock_anon_vma_read()? zhong jiang <zhongjiang@huawei.com> - 2017-05-23 11:30 +0200
                        Re: mm, something wring in page_lock_anon_vma_read()? Vlastimil Babka <vbabka@suse.cz> - 2017-05-23 11:40 +0200
                          Re: mm, something wring in page_lock_anon_vma_read()? zhong jiang <zhongjiang@huawei.com> - 2017-05-23 12:40 +0200
                  Re: mm, something wring in page_lock_anon_vma_read()? Xishi Qiu <qiuxishi@huawei.com> - 2017-05-22 12:00 +0200
                    Re: mm, something wring in page_lock_anon_vma_read()? Hugh Dickins <hughd@google.com> - 2017-05-22 21:30 +0200
                      Re: mm, something wring in page_lock_anon_vma_read()? Xishi Qiu <qiuxishi@huawei.com> - 2017-05-23 04:30 +0200
                        Re: mm, something wring in page_lock_anon_vma_read()? Hugh Dickins <hughd@google.com> - 2017-05-23 05:00 +0200

#1644028 — mm, something wring in page_lock_anon_vma_read()?

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-05-18 12:00 +0200
Subjectmm, something wring in page_lock_anon_vma_read()?
Message-ID<tIqnU-qA-19@gated-at.bofh.it>
Hi, my system triggers this bug, and the vmcore shows the anon_vma seems be freed.
The kernel is RHEL 7.2, and the bug is hard to reproduce, so I don't know if it
exists in mainline, any reply is welcome!

[35030.332666] general protection fault: 0000 [#1] SMP
[35030.333016] Modules linked in: veth ipt_MASQUERADE nf_nat_masquerade_ipv4 iptable_nat nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 xt_addrtype iptable_filter xt_conntrack nf_nat nf_conntrack bridge stp llc dm_thin_pool dm_persistent_data dm_bio_prison dm_bufio libcrc32c rtos_kbox_panic(OE) ipmi_devintf ipmi_si ipmi_msghandler signo_catch(O) cirrus syscopyarea sysfillrect sysimgblt ttm crc32_pclmul ghash_clmulni_intel drm_kms_helper aesni_intel ppdev drm lrw gf128mul parport_pc glue_helper ablk_helper serio_raw cryptd i2c_piix4 parport pcspkr sg floppy i2c_core dm_mod sha512_generic ip_tables sd_mod crc_t10dif crct10dif_generic sr_mod cdrom virtio_console virtio_scsi virtio_net ata_generic pata_acpi crct10dif_pclmul crct10dif_common crc32c_intel virtio_pci virtio_ring virtio ata_piix libata ext4 mbcache
[35030.333016]  jbd2
[35030.333016] CPU: 3 PID: 48 Comm: kswapd0 Tainted: G           OE  ---- -------   3.10.0-327.36.58.4.x86_64 #1
[35030.333016] Hardware name: OpenStack Foundation OpenStack Nova, BIOS rel-1.8.1-0-g4adadbd-20160826_044443-hghoulaslx112 04/01/2014
[35030.333016] task: ffff8801b2d20000 ti: ffff8801b4c38000 task.ti: ffff8801b4c38000
[35030.333016] RIP: 0010:[<ffffffff810acac5>]  [<ffffffff810acac5>] down_read_trylock+0x5/0x50
[35030.333016] RSP: 0000:ffff8801b4c3ba90  EFLAGS: 00010282
[35030.333016] RAX: 0000000000000000 RBX: ffff8801b3e2a100 RCX: 0000000000000000
[35030.333016] RDX: 0000000000000000 RSI: 0000000000000000 RDI: deb604d497705c5d
[35030.333016] RBP: ffff8801b4c3bab8 R08: ffffea0002c34460 R09: ffff8801b3d7e8a0
[35030.333016] R10: 0000000000000004 R11: fff00000fe000000 R12: ffff8801b3e2a101
[35030.333016] R13: ffffea0002c34440 R14: deb604d497705c5d R15: ffffea0002c34440
[35030.333016] FS:  0000000000000000(0000) GS:ffff8801bed80000(0000) knlGS:0000000000000000
[35030.333016] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[35030.333016] CR2: 000000c422011080 CR3: 0000000001976000 CR4: 00000000001407e0
[35030.333016] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[35030.333016] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
[35030.333016] Stack:
[35030.333016]  ffffffff811b2795 ffffea0002c34440 0000000000000000 000000000000000f
[35030.333016]  0000000000000001 ffff8801b4c3bb30 ffffffff811b2a17 ffff8800a712d640
[35030.333016]  000000000c4229e2 ffff8801b4c3bb80 0000000100000000 000000000c41fe38
[35030.333016] Call Trace:
[35030.333016]  [<ffffffff811b2795>] ? page_lock_anon_vma_read+0x55/0x110
[35030.333016]  [<ffffffff811b2a17>] page_referenced+0x1c7/0x350
[35030.333016]  [<ffffffff8118d9b4>] shrink_active_list+0x1e4/0x400
[35030.333016]  [<ffffffff8118e08d>] shrink_lruvec+0x4bd/0x770
[35030.333016]  [<ffffffff8118e3b6>] shrink_zone+0x76/0x1a0
[35030.333016]  [<ffffffff8118f6cc>] balance_pgdat+0x49c/0x610
[35030.333016]  [<ffffffff8118f9b3>] kswapd+0x173/0x450
[35030.333016]  [<ffffffff810a8a00>] ? wake_up_atomic_t+0x30/0x30
[35030.333016]  [<ffffffff8118f840>] ? balance_pgdat+0x610/0x610
[35030.333016]  [<ffffffff810a79bf>] kthread+0xcf/0xe0
[35030.333016]  [<ffffffff810a78f0>] ? kthread_create_on_node+0x120/0x120
[35030.333016]  [<ffffffff81665bd8>] ret_from_fork+0x58/0x90
[35030.333016]  [<ffffffff810a78f0>] ? kthread_create_on_node+0x120/0x120
[35030.333016] Code: 00 ba ff ff ff ff 48 89 d8 f0 48 0f c1 10 79 05 e8 31 06 27 00 5b 5d c3 66 66 66 66 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00 <48> 8b 07 48 89 c2 48 83 c2 01 7e 07 f0 48 0f b1 17 75 f0 48 f7
[35030.333016] RIP  [<ffffffff810acac5>] down_read_trylock+0x5/0x50
[35030.333016]  RSP <ffff8801b4c3ba90>
[35030.333016] ------------[ cut here ]------------

struct page {
  flags = 9007194960298056,
  mapping = 0xffff8801b3e2a101,
  {
    {
      index = 34324593617,
      freelist = 0x7fde7bbd1,
      pfmemalloc = 209,
      thp_mmu_gather = {
        counter = -35144751
      },
      pmd_huge_pte = 0x7fde7bbd1
    },
    {
      counters = 8589934592,
      {
        {
          _mapcount = {
            counter = 0
          },
          {
            inuse = 0,
            objects = 0,
            frozen = 0
          },
          units = 0
        },
        _count = {
          counter = 2
        }
      }
    }
  },
  {
    lru = {
      next = 0xdead000000100100,
      prev = 0xdead000000200200
    },
    {
      next = 0xdead000000100100,
      pages = 2097664,
      pobjects = -559087616
    },
    list = {
      next = 0xdead000000100100,
      prev = 0xdead000000200200
    },
    slab_page = 0xdead000000100100
  },
  {
    private = 0,
    ptl = {
      {
        rlock = {
          raw_lock = {
            {
              head_tail = 0,
              tickets = {
                head = 0,
                tail = 0
              }
            }
          }
        }
      }
    },
    slab_cache = 0x0,
    first_page = 0x0
  }
}



crash> struct anon_vma 0xffff8801b3e2a100
struct anon_vma {
  root = 0xdeb604d497705c55,
  rwsem = {
    count = -8192007903225070328,
    wait_lock = {
      raw_lock = {
        {
          head_tail = 2955503940,
          tickets = {
            head = 26948,
            tail = 45097
          }
        }
      }
    },
    wait_list = {
      next = 0x559f9107c1b47439,
      prev = 0x3de13f709bfa043b
    }
  },
  refcount = {
    counter = -13243516
  },
  rb_root = {
    rb_node = 0x11dd18f9ce0bb2e9
  }
}

This address 0xffff8801b3e2a100 can not find in "kmem -S anon_vma"

The page flags is
crash> kmem -g 0x1FFFFF00080048
FLAGS: 1fffff00080048
  PAGE-FLAG        BIT  VALUE
  PG_uptodate        3  0000008
  PG_active          6  0000040
  PG_swapbacked     19  0080000

[toc] | [next] | [standalone]


#1645393

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-05-19 11:00 +0200
Message-ID<tILVn-84l-3@gated-at.bofh.it>
In reply to#1644028
On 2017/5/18 17:46, Xishi Qiu wrote:

> Hi, my system triggers this bug, and the vmcore shows the anon_vma seems be freed.
> The kernel is RHEL 7.2, and the bug is hard to reproduce, so I don't know if it
> exists in mainline, any reply is welcome!
> 

When we alloc anon_vma, we will init the value of anon_vma->root,
so can we set anon_vma->root to NULL when calling
anon_vma_free -> kmem_cache_free(anon_vma_cachep, anon_vma);

anon_vma_free()
	...
	anon_vma->root = NULL;
	kmem_cache_free(anon_vma_cachep, anon_vma);

I find if we do this above, system boot failed, why?

Thanks,
Xishi Qiu

> [35030.332666] general protection fault: 0000 [#1] SMP
> [35030.333016] Modules linked in: veth ipt_MASQUERADE nf_nat_masquerade_ipv4 iptable_nat nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 xt_addrtype iptable_filter xt_conntrack nf_nat nf_conntrack bridge stp llc dm_thin_pool dm_persistent_data dm_bio_prison dm_bufio libcrc32c rtos_kbox_panic(OE) ipmi_devintf ipmi_si ipmi_msghandler signo_catch(O) cirrus syscopyarea sysfillrect sysimgblt ttm crc32_pclmul ghash_clmulni_intel drm_kms_helper aesni_intel ppdev drm lrw gf128mul parport_pc glue_helper ablk_helper serio_raw cryptd i2c_piix4 parport pcspkr sg floppy i2c_core dm_mod sha512_generic ip_tables sd_mod crc_t10dif crct10dif_generic sr_mod cdrom virtio_console virtio_scsi virtio_net ata_generic pata_acpi crct10dif_pclmul crct10dif_common crc32c_intel virtio_pci virtio_ring virtio ata_piix libata ext4 mbcache
> [35030.333016]  jbd2
> [35030.333016] CPU: 3 PID: 48 Comm: kswapd0 Tainted: G           OE  ---- -------   3.10.0-327.36.58.4.x86_64 #1
> [35030.333016] Hardware name: OpenStack Foundation OpenStack Nova, BIOS rel-1.8.1-0-g4adadbd-20160826_044443-hghoulaslx112 04/01/2014
> [35030.333016] task: ffff8801b2d20000 ti: ffff8801b4c38000 task.ti: ffff8801b4c38000
> [35030.333016] RIP: 0010:[<ffffffff810acac5>]  [<ffffffff810acac5>] down_read_trylock+0x5/0x50
> [35030.333016] RSP: 0000:ffff8801b4c3ba90  EFLAGS: 00010282
> [35030.333016] RAX: 0000000000000000 RBX: ffff8801b3e2a100 RCX: 0000000000000000
> [35030.333016] RDX: 0000000000000000 RSI: 0000000000000000 RDI: deb604d497705c5d
> [35030.333016] RBP: ffff8801b4c3bab8 R08: ffffea0002c34460 R09: ffff8801b3d7e8a0
> [35030.333016] R10: 0000000000000004 R11: fff00000fe000000 R12: ffff8801b3e2a101
> [35030.333016] R13: ffffea0002c34440 R14: deb604d497705c5d R15: ffffea0002c34440
> [35030.333016] FS:  0000000000000000(0000) GS:ffff8801bed80000(0000) knlGS:0000000000000000
> [35030.333016] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
> [35030.333016] CR2: 000000c422011080 CR3: 0000000001976000 CR4: 00000000001407e0
> [35030.333016] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
> [35030.333016] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
> [35030.333016] Stack:
> [35030.333016]  ffffffff811b2795 ffffea0002c34440 0000000000000000 000000000000000f
> [35030.333016]  0000000000000001 ffff8801b4c3bb30 ffffffff811b2a17 ffff8800a712d640
> [35030.333016]  000000000c4229e2 ffff8801b4c3bb80 0000000100000000 000000000c41fe38
> [35030.333016] Call Trace:
> [35030.333016]  [<ffffffff811b2795>] ? page_lock_anon_vma_read+0x55/0x110
> [35030.333016]  [<ffffffff811b2a17>] page_referenced+0x1c7/0x350
> [35030.333016]  [<ffffffff8118d9b4>] shrink_active_list+0x1e4/0x400
> [35030.333016]  [<ffffffff8118e08d>] shrink_lruvec+0x4bd/0x770
> [35030.333016]  [<ffffffff8118e3b6>] shrink_zone+0x76/0x1a0
> [35030.333016]  [<ffffffff8118f6cc>] balance_pgdat+0x49c/0x610
> [35030.333016]  [<ffffffff8118f9b3>] kswapd+0x173/0x450
> [35030.333016]  [<ffffffff810a8a00>] ? wake_up_atomic_t+0x30/0x30
> [35030.333016]  [<ffffffff8118f840>] ? balance_pgdat+0x610/0x610
> [35030.333016]  [<ffffffff810a79bf>] kthread+0xcf/0xe0
> [35030.333016]  [<ffffffff810a78f0>] ? kthread_create_on_node+0x120/0x120
> [35030.333016]  [<ffffffff81665bd8>] ret_from_fork+0x58/0x90
> [35030.333016]  [<ffffffff810a78f0>] ? kthread_create_on_node+0x120/0x120
> [35030.333016] Code: 00 ba ff ff ff ff 48 89 d8 f0 48 0f c1 10 79 05 e8 31 06 27 00 5b 5d c3 66 66 66 66 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00 <48> 8b 07 48 89 c2 48 83 c2 01 7e 07 f0 48 0f b1 17 75 f0 48 f7
> [35030.333016] RIP  [<ffffffff810acac5>] down_read_trylock+0x5/0x50
> [35030.333016]  RSP <ffff8801b4c3ba90>
> [35030.333016] ------------[ cut here ]------------
> 
> struct page {
>   flags = 9007194960298056,
>   mapping = 0xffff8801b3e2a101,
>   {
>     {
>       index = 34324593617,
>       freelist = 0x7fde7bbd1,
>       pfmemalloc = 209,
>       thp_mmu_gather = {
>         counter = -35144751
>       },
>       pmd_huge_pte = 0x7fde7bbd1
>     },
>     {
>       counters = 8589934592,
>       {
>         {
>           _mapcount = {
>             counter = 0
>           },
>           {
>             inuse = 0,
>             objects = 0,
>             frozen = 0
>           },
>           units = 0
>         },
>         _count = {
>           counter = 2
>         }
>       }
>     }
>   },
>   {
>     lru = {
>       next = 0xdead000000100100,
>       prev = 0xdead000000200200
>     },
>     {
>       next = 0xdead000000100100,
>       pages = 2097664,
>       pobjects = -559087616
>     },
>     list = {
>       next = 0xdead000000100100,
>       prev = 0xdead000000200200
>     },
>     slab_page = 0xdead000000100100
>   },
>   {
>     private = 0,
>     ptl = {
>       {
>         rlock = {
>           raw_lock = {
>             {
>               head_tail = 0,
>               tickets = {
>                 head = 0,
>                 tail = 0
>               }
>             }
>           }
>         }
>       }
>     },
>     slab_cache = 0x0,
>     first_page = 0x0
>   }
> }
> 
> 
> 
> crash> struct anon_vma 0xffff8801b3e2a100
> struct anon_vma {
>   root = 0xdeb604d497705c55,
>   rwsem = {
>     count = -8192007903225070328,
>     wait_lock = {
>       raw_lock = {
>         {
>           head_tail = 2955503940,
>           tickets = {
>             head = 26948,
>             tail = 45097
>           }
>         }
>       }
>     },
>     wait_list = {
>       next = 0x559f9107c1b47439,
>       prev = 0x3de13f709bfa043b
>     }
>   },
>   refcount = {
>     counter = -13243516
>   },
>   rb_root = {
>     rb_node = 0x11dd18f9ce0bb2e9
>   }
> }
> 
> This address 0xffff8801b3e2a100 can not find in "kmem -S anon_vma"
> 
> The page flags is
> crash> kmem -g 0x1FFFFF00080048
> FLAGS: 1fffff00080048
>   PAGE-FLAG        BIT  VALUE
>   PG_uptodate        3  0000008
>   PG_active          6  0000040
>   PG_swapbacked     19  0080000
> 
> 
> .
> 

[toc] | [prev] | [next] | [standalone]


#1645470

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-05-19 11:50 +0200
Message-ID<tIMHM-hj-13@gated-at.bofh.it>
In reply to#1645393
On 2017/5/19 16:52, Xishi Qiu wrote:

> On 2017/5/18 17:46, Xishi Qiu wrote:
> 
>> Hi, my system triggers this bug, and the vmcore shows the anon_vma seems be freed.
>> The kernel is RHEL 7.2, and the bug is hard to reproduce, so I don't know if it
>> exists in mainline, any reply is welcome!
>>
> 
> When we alloc anon_vma, we will init the value of anon_vma->root,
> so can we set anon_vma->root to NULL when calling
> anon_vma_free -> kmem_cache_free(anon_vma_cachep, anon_vma);
> 
> anon_vma_free()
> 	...
> 	anon_vma->root = NULL;
> 	kmem_cache_free(anon_vma_cachep, anon_vma);
> 
> I find if we do this above, system boot failed, why?
> 

If anon_vma was freed, we should not to access the root_anon_vma, because it maybe also
freed(e.g. anon_vma == root_anon_vma), right?

page_lock_anon_vma_read()
	...
	anon_vma = (struct anon_vma *) (anon_mapping - PAGE_MAPPING_ANON);
	root_anon_vma = ACCESS_ONCE(anon_vma->root);
	if (down_read_trylock(&root_anon_vma->rwsem)) {  // it's not safe
	...
	if (!atomic_inc_not_zero(&anon_vma->refcount)) {  // check anon_vma was not freed
	...
	anon_vma_lock_read(anon_vma);  // it's safe
	...


> Thanks,
> Xishi Qiu
> 
>> [35030.332666] general protection fault: 0000 [#1] SMP
>> [35030.333016] Modules linked in: veth ipt_MASQUERADE nf_nat_masquerade_ipv4 iptable_nat nf_conntrack_ipv4 nf_defrag_ipv4 nf_nat_ipv4 xt_addrtype iptable_filter xt_conntrack nf_nat nf_conntrack bridge stp llc dm_thin_pool dm_persistent_data dm_bio_prison dm_bufio libcrc32c rtos_kbox_panic(OE) ipmi_devintf ipmi_si ipmi_msghandler signo_catch(O) cirrus syscopyarea sysfillrect sysimgblt ttm crc32_pclmul ghash_clmulni_intel drm_kms_helper aesni_intel ppdev drm lrw gf128mul parport_pc glue_helper ablk_helper serio_raw cryptd i2c_piix4 parport pcspkr sg floppy i2c_core dm_mod sha512_generic ip_tables sd_mod crc_t10dif crct10dif_generic sr_mod cdrom virtio_console virtio_scsi virtio_net ata_generic pata_acpi crct10dif_pclmul crct10dif_common crc32c_intel virtio_pci virtio_ring virtio ata_piix libata ext4 mbcache
>> [35030.333016]  jbd2
>> [35030.333016] CPU: 3 PID: 48 Comm: kswapd0 Tainted: G           OE  ---- -------   3.10.0-327.36.58.4.x86_64 #1
>> [35030.333016] Hardware name: OpenStack Foundation OpenStack Nova, BIOS rel-1.8.1-0-g4adadbd-20160826_044443-hghoulaslx112 04/01/2014
>> [35030.333016] task: ffff8801b2d20000 ti: ffff8801b4c38000 task.ti: ffff8801b4c38000
>> [35030.333016] RIP: 0010:[<ffffffff810acac5>]  [<ffffffff810acac5>] down_read_trylock+0x5/0x50
>> [35030.333016] RSP: 0000:ffff8801b4c3ba90  EFLAGS: 00010282
>> [35030.333016] RAX: 0000000000000000 RBX: ffff8801b3e2a100 RCX: 0000000000000000
>> [35030.333016] RDX: 0000000000000000 RSI: 0000000000000000 RDI: deb604d497705c5d
>> [35030.333016] RBP: ffff8801b4c3bab8 R08: ffffea0002c34460 R09: ffff8801b3d7e8a0
>> [35030.333016] R10: 0000000000000004 R11: fff00000fe000000 R12: ffff8801b3e2a101
>> [35030.333016] R13: ffffea0002c34440 R14: deb604d497705c5d R15: ffffea0002c34440
>> [35030.333016] FS:  0000000000000000(0000) GS:ffff8801bed80000(0000) knlGS:0000000000000000
>> [35030.333016] CS:  0010 DS: 0000 ES: 0000 CR0: 0000000080050033
>> [35030.333016] CR2: 000000c422011080 CR3: 0000000001976000 CR4: 00000000001407e0
>> [35030.333016] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
>> [35030.333016] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
>> [35030.333016] Stack:
>> [35030.333016]  ffffffff811b2795 ffffea0002c34440 0000000000000000 000000000000000f
>> [35030.333016]  0000000000000001 ffff8801b4c3bb30 ffffffff811b2a17 ffff8800a712d640
>> [35030.333016]  000000000c4229e2 ffff8801b4c3bb80 0000000100000000 000000000c41fe38
>> [35030.333016] Call Trace:
>> [35030.333016]  [<ffffffff811b2795>] ? page_lock_anon_vma_read+0x55/0x110
>> [35030.333016]  [<ffffffff811b2a17>] page_referenced+0x1c7/0x350
>> [35030.333016]  [<ffffffff8118d9b4>] shrink_active_list+0x1e4/0x400
>> [35030.333016]  [<ffffffff8118e08d>] shrink_lruvec+0x4bd/0x770
>> [35030.333016]  [<ffffffff8118e3b6>] shrink_zone+0x76/0x1a0
>> [35030.333016]  [<ffffffff8118f6cc>] balance_pgdat+0x49c/0x610
>> [35030.333016]  [<ffffffff8118f9b3>] kswapd+0x173/0x450
>> [35030.333016]  [<ffffffff810a8a00>] ? wake_up_atomic_t+0x30/0x30
>> [35030.333016]  [<ffffffff8118f840>] ? balance_pgdat+0x610/0x610
>> [35030.333016]  [<ffffffff810a79bf>] kthread+0xcf/0xe0
>> [35030.333016]  [<ffffffff810a78f0>] ? kthread_create_on_node+0x120/0x120
>> [35030.333016]  [<ffffffff81665bd8>] ret_from_fork+0x58/0x90
>> [35030.333016]  [<ffffffff810a78f0>] ? kthread_create_on_node+0x120/0x120
>> [35030.333016] Code: 00 ba ff ff ff ff 48 89 d8 f0 48 0f c1 10 79 05 e8 31 06 27 00 5b 5d c3 66 66 66 66 66 2e 0f 1f 84 00 00 00 00 00 0f 1f 44 00 00 <48> 8b 07 48 89 c2 48 83 c2 01 7e 07 f0 48 0f b1 17 75 f0 48 f7
>> [35030.333016] RIP  [<ffffffff810acac5>] down_read_trylock+0x5/0x50
>> [35030.333016]  RSP <ffff8801b4c3ba90>
>> [35030.333016] ------------[ cut here ]------------
>>
>> struct page {
>>   flags = 9007194960298056,
>>   mapping = 0xffff8801b3e2a101,
>>   {
>>     {
>>       index = 34324593617,
>>       freelist = 0x7fde7bbd1,
>>       pfmemalloc = 209,
>>       thp_mmu_gather = {
>>         counter = -35144751
>>       },
>>       pmd_huge_pte = 0x7fde7bbd1
>>     },
>>     {
>>       counters = 8589934592,
>>       {
>>         {
>>           _mapcount = {
>>             counter = 0
>>           },
>>           {
>>             inuse = 0,
>>             objects = 0,
>>             frozen = 0
>>           },
>>           units = 0
>>         },
>>         _count = {
>>           counter = 2
>>         }
>>       }
>>     }
>>   },
>>   {
>>     lru = {
>>       next = 0xdead000000100100,
>>       prev = 0xdead000000200200
>>     },
>>     {
>>       next = 0xdead000000100100,
>>       pages = 2097664,
>>       pobjects = -559087616
>>     },
>>     list = {
>>       next = 0xdead000000100100,
>>       prev = 0xdead000000200200
>>     },
>>     slab_page = 0xdead000000100100
>>   },
>>   {
>>     private = 0,
>>     ptl = {
>>       {
>>         rlock = {
>>           raw_lock = {
>>             {
>>               head_tail = 0,
>>               tickets = {
>>                 head = 0,
>>                 tail = 0
>>               }
>>             }
>>           }
>>         }
>>       }
>>     },
>>     slab_cache = 0x0,
>>     first_page = 0x0
>>   }
>> }
>>
>>
>>
>> crash> struct anon_vma 0xffff8801b3e2a100
>> struct anon_vma {
>>   root = 0xdeb604d497705c55,
>>   rwsem = {
>>     count = -8192007903225070328,
>>     wait_lock = {
>>       raw_lock = {
>>         {
>>           head_tail = 2955503940,
>>           tickets = {
>>             head = 26948,
>>             tail = 45097
>>           }
>>         }
>>       }
>>     },
>>     wait_list = {
>>       next = 0x559f9107c1b47439,
>>       prev = 0x3de13f709bfa043b
>>     }
>>   },
>>   refcount = {
>>     counter = -13243516
>>   },
>>   rb_root = {
>>     rb_node = 0x11dd18f9ce0bb2e9
>>   }
>> }
>>
>> This address 0xffff8801b3e2a100 can not find in "kmem -S anon_vma"
>>
>> The page flags is
>> crash> kmem -g 0x1FFFFF00080048
>> FLAGS: 1fffff00080048
>>   PAGE-FLAG        BIT  VALUE
>>   PG_uptodate        3  0000008
>>   PG_active          6  0000040
>>   PG_swapbacked     19  0080000
>>
>>
>> .
>>
> 
> 
> 
> 
> .
> 

[toc] | [prev] | [next] | [standalone]


#1645977

FromHugh Dickins <hughd@google.com>
Date2017-05-20 00:10 +0200
Message-ID<tIYfT-57-1@gated-at.bofh.it>
In reply to#1645470
On Fri, 19 May 2017, Xishi Qiu wrote:
> On 2017/5/19 16:52, Xishi Qiu wrote:
> > On 2017/5/18 17:46, Xishi Qiu wrote:
> > 
> >> Hi, my system triggers this bug, and the vmcore shows the anon_vma seems be freed.
> >> The kernel is RHEL 7.2, and the bug is hard to reproduce, so I don't know if it
> >> exists in mainline, any reply is welcome!
> >>
> > 
> > When we alloc anon_vma, we will init the value of anon_vma->root,
> > so can we set anon_vma->root to NULL when calling
> > anon_vma_free -> kmem_cache_free(anon_vma_cachep, anon_vma);
> > 
> > anon_vma_free()
> > 	...
> > 	anon_vma->root = NULL;
> > 	kmem_cache_free(anon_vma_cachep, anon_vma);
> > 
> > I find if we do this above, system boot failed, why?
> > 
> 
> If anon_vma was freed, we should not to access the root_anon_vma, because it maybe also
> freed(e.g. anon_vma == root_anon_vma), right?
> 
> page_lock_anon_vma_read()
> 	...
> 	anon_vma = (struct anon_vma *) (anon_mapping - PAGE_MAPPING_ANON);
> 	root_anon_vma = ACCESS_ONCE(anon_vma->root);
> 	if (down_read_trylock(&root_anon_vma->rwsem)) {  // it's not safe
> 	...
> 	if (!atomic_inc_not_zero(&anon_vma->refcount)) {  // check anon_vma was not freed
> 	...
> 	anon_vma_lock_read(anon_vma);  // it's safe
> 	...

You're ignoring the rcu_read_lock() on entry to page_lock_anon_vma_read(),
and the SLAB_DESTROY_BY_RCU (recently renamed SLAB_TYPESAFE_BY_RCU) nature
of the anon_vma_cachep kmem cache.  It is not safe to muck with anon_vma->
root in anon_vma_free(), others could still be looking at it.

Hugh

[toc] | [prev] | [next] | [standalone]


#1646038

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-05-20 03:30 +0200
Message-ID<tJ1ns-2mZ-1@gated-at.bofh.it>
In reply to#1645977
On 2017/5/20 6:00, Hugh Dickins wrote:

> On Fri, 19 May 2017, Xishi Qiu wrote:
>> On 2017/5/19 16:52, Xishi Qiu wrote:
>>> On 2017/5/18 17:46, Xishi Qiu wrote:
>>>
>>>> Hi, my system triggers this bug, and the vmcore shows the anon_vma seems be freed.
>>>> The kernel is RHEL 7.2, and the bug is hard to reproduce, so I don't know if it
>>>> exists in mainline, any reply is welcome!
>>>>
>>>
>>> When we alloc anon_vma, we will init the value of anon_vma->root,
>>> so can we set anon_vma->root to NULL when calling
>>> anon_vma_free -> kmem_cache_free(anon_vma_cachep, anon_vma);
>>>
>>> anon_vma_free()
>>> 	...
>>> 	anon_vma->root = NULL;
>>> 	kmem_cache_free(anon_vma_cachep, anon_vma);
>>>
>>> I find if we do this above, system boot failed, why?
>>>
>>
>> If anon_vma was freed, we should not to access the root_anon_vma, because it maybe also
>> freed(e.g. anon_vma == root_anon_vma), right?
>>
>> page_lock_anon_vma_read()
>> 	...
>> 	anon_vma = (struct anon_vma *) (anon_mapping - PAGE_MAPPING_ANON);
>> 	root_anon_vma = ACCESS_ONCE(anon_vma->root);
>> 	if (down_read_trylock(&root_anon_vma->rwsem)) {  // it's not safe
>> 	...
>> 	if (!atomic_inc_not_zero(&anon_vma->refcount)) {  // check anon_vma was not freed
>> 	...
>> 	anon_vma_lock_read(anon_vma);  // it's safe
>> 	...
> 
> You're ignoring the rcu_read_lock() on entry to page_lock_anon_vma_read(),
> and the SLAB_DESTROY_BY_RCU (recently renamed SLAB_TYPESAFE_BY_RCU) nature
> of the anon_vma_cachep kmem cache.  It is not safe to muck with anon_vma->
> root in anon_vma_free(), others could still be looking at it.
> 
> Hugh
> 

Hi Hugh,

Thanks for your reply.

SLAB_DESTROY_BY_RCU will let it call call_rcu() in free_slab(), but if the
anon_vma *reuse* by someone again, access root_anon_vma is not safe, right?

e.g. if I clean the root pointer before free it, then access root_anon_vma
in page_lock_anon_vma_read() is NULL pointer access, right?

anon_vma_free()
	...
	anon_vma->root = NULL;
	kmem_cache_free(anon_vma_cachep, anon_vma);
	...

Thanks,
Xishi Qiu

> .
> 

[toc] | [prev] | [next] | [standalone]


#1646043

FromHugh Dickins <hughd@google.com>
Date2017-05-20 04:10 +0200
Message-ID<tJ209-2WG-3@gated-at.bofh.it>
In reply to#1646038
On Sat, 20 May 2017, Xishi Qiu wrote:
> On 2017/5/20 6:00, Hugh Dickins wrote:
> > 
> > You're ignoring the rcu_read_lock() on entry to page_lock_anon_vma_read(),
> > and the SLAB_DESTROY_BY_RCU (recently renamed SLAB_TYPESAFE_BY_RCU) nature
> > of the anon_vma_cachep kmem cache.  It is not safe to muck with anon_vma->
> > root in anon_vma_free(), others could still be looking at it.
> > 
> > Hugh
> > 
> 
> Hi Hugh,
> 
> Thanks for your reply.
> 
> SLAB_DESTROY_BY_RCU will let it call call_rcu() in free_slab(), but if the
> anon_vma *reuse* by someone again, access root_anon_vma is not safe, right?

That is safe, on reuse it is still a struct anon_vma; then the test for
!page_mapped(page) will show that it's no longer a reliable anon_vma for
this page, so page_lock_anon_vma_read() returns NULL.

But of course, if page->_mapcount has been corrupted or misaccounted,
it may think page_mapped(page) when actually page is not mapped,
and the anon_vma is not good for it.

> 
> e.g. if I clean the root pointer before free it, then access root_anon_vma
> in page_lock_anon_vma_read() is NULL pointer access, right?

Yes, cleaning root pointer before free may result in NULL pointer access.

Hugh

> 
> anon_vma_free()
> 	...
> 	anon_vma->root = NULL;
> 	kmem_cache_free(anon_vma_cachep, anon_vma);
> 	...
> 
> Thanks,
> Xishi Qiu

[toc] | [prev] | [next] | [standalone]


#1646047

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-05-20 04:30 +0200
Message-ID<tJ2jv-39J-3@gated-at.bofh.it>
In reply to#1646043
On 2017/5/20 10:02, Hugh Dickins wrote:

> On Sat, 20 May 2017, Xishi Qiu wrote:
>> On 2017/5/20 6:00, Hugh Dickins wrote:
>>>
>>> You're ignoring the rcu_read_lock() on entry to page_lock_anon_vma_read(),
>>> and the SLAB_DESTROY_BY_RCU (recently renamed SLAB_TYPESAFE_BY_RCU) nature
>>> of the anon_vma_cachep kmem cache.  It is not safe to muck with anon_vma->
>>> root in anon_vma_free(), others could still be looking at it.
>>>
>>> Hugh
>>>
>>
>> Hi Hugh,
>>
>> Thanks for your reply.
>>
>> SLAB_DESTROY_BY_RCU will let it call call_rcu() in free_slab(), but if the
>> anon_vma *reuse* by someone again, access root_anon_vma is not safe, right?
> 
> That is safe, on reuse it is still a struct anon_vma; then the test for
> !page_mapped(page) will show that it's no longer a reliable anon_vma for
> this page, so page_lock_anon_vma_read() returns NULL.
> 
> But of course, if page->_mapcount has been corrupted or misaccounted,
> it may think page_mapped(page) when actually page is not mapped,
> and the anon_vma is not good for it.
> 

Hi Hugh,

Here is a bug report form redhat: https://bugzilla.redhat.com/show_bug.cgi?id=1305620
And I meet the bug too. However it is hard to reproduce, and 
624483f3ea82598("mm: rmap: fix use-after-free in __put_anon_vma") is not help.

From the vmcore, it seems that the page is still mapped(_mapcount=0 and _count=2),
and the value of mapping is a valid address(mapping = 0xffff8801b3e2a101),
but anon_vma has been corrupted.

Any ideas?

Thanks,
Xishi Qiu

>>
>> e.g. if I clean the root pointer before free it, then access root_anon_vma
>> in page_lock_anon_vma_read() is NULL pointer access, right?
> 
> Yes, cleaning root pointer before free may result in NULL pointer access.
> 
> Hugh
> 
>>
>> anon_vma_free()
>> 	...
>> 	anon_vma->root = NULL;
>> 	kmem_cache_free(anon_vma_cachep, anon_vma);
>> 	...
>>
>> Thanks,
>> Xishi Qiu
> 
> .
> 

[toc] | [prev] | [next] | [standalone]


#1646049

FromHugh Dickins <hughd@google.com>
Date2017-05-20 04:50 +0200
Message-ID<tJ2CR-3nU-1@gated-at.bofh.it>
In reply to#1646047
On Sat, 20 May 2017, Xishi Qiu wrote:
> 
> Here is a bug report form redhat: https://bugzilla.redhat.com/show_bug.cgi?id=1305620
> And I meet the bug too. However it is hard to reproduce, and 
> 624483f3ea82598("mm: rmap: fix use-after-free in __put_anon_vma") is not help.
> 
> From the vmcore, it seems that the page is still mapped(_mapcount=0 and _count=2),
> and the value of mapping is a valid address(mapping = 0xffff8801b3e2a101),
> but anon_vma has been corrupted.
> 
> Any ideas?

Sorry, no.  I assume that _mapcount has been misaccounted, for example
a pte mapped in on top of another pte; but cannot begin tell you where
in Red Hat's kernel-3.10.0-229.4.2.el7 that might happen.

Hugh

[toc] | [prev] | [next] | [standalone]


#1646051

Fromzhong jiang <zhongjiang@huawei.com>
Date2017-05-20 05:10 +0200
Message-ID<tJ2Wd-3LQ-1@gated-at.bofh.it>
In reply to#1646049
On 2017/5/20 10:40, Hugh Dickins wrote:
> On Sat, 20 May 2017, Xishi Qiu wrote:
>> Here is a bug report form redhat: https://bugzilla.redhat.com/show_bug.cgi?id=1305620
>> And I meet the bug too. However it is hard to reproduce, and 
>> 624483f3ea82598("mm: rmap: fix use-after-free in __put_anon_vma") is not help.
>>
>> From the vmcore, it seems that the page is still mapped(_mapcount=0 and _count=2),
>> and the value of mapping is a valid address(mapping = 0xffff8801b3e2a101),
>> but anon_vma has been corrupted.
>>
>> Any ideas?
> Sorry, no.  I assume that _mapcount has been misaccounted, for example
> a pte mapped in on top of another pte; but cannot begin tell you where
> in Red Hat's kernel-3.10.0-229.4.2.el7 that might happen.
>
> Hugh
>
> .
>
Hi, Hugh

I find the following message from the dmesg.

[26068.316592] BUG: Bad rss-counter state mm:ffff8800a7de2d80 idx:1 val:1

I can prove that the __mapcount is misaccount.  when task is exited. the rmap
still exist.

Thanks
zhongjiang

[toc] | [prev] | [next] | [standalone]


#1647184

FromVlastimil Babka <vbabka@suse.cz>
Date2017-05-22 19:00 +0200
Message-ID<tJYQA-8az-75@gated-at.bofh.it>
In reply to#1646051
On 05/20/2017 05:01 AM, zhong jiang wrote:
> On 2017/5/20 10:40, Hugh Dickins wrote:
>> On Sat, 20 May 2017, Xishi Qiu wrote:
>>> Here is a bug report form redhat: https://bugzilla.redhat.com/show_bug.cgi?id=1305620
>>> And I meet the bug too. However it is hard to reproduce, and 
>>> 624483f3ea82598("mm: rmap: fix use-after-free in __put_anon_vma") is not help.
>>>
>>> From the vmcore, it seems that the page is still mapped(_mapcount=0 and _count=2),
>>> and the value of mapping is a valid address(mapping = 0xffff8801b3e2a101),
>>> but anon_vma has been corrupted.
>>>
>>> Any ideas?
>> Sorry, no.  I assume that _mapcount has been misaccounted, for example
>> a pte mapped in on top of another pte; but cannot begin tell you where
>> in Red Hat's kernel-3.10.0-229.4.2.el7 that might happen.
>>
>> Hugh
>>
>> .
>>
> Hi, Hugh
> 
> I find the following message from the dmesg.
> 
> [26068.316592] BUG: Bad rss-counter state mm:ffff8800a7de2d80 idx:1 val:1
> 
> I can prove that the __mapcount is misaccount.  when task is exited. the rmap
> still exist.

Check if the kernel in question contains this commit: ad33bb04b2a6 ("mm:
thp: fix SMP race condition between THP page fault and MADV_DONTNEED")

> Thanks
> zhongjiang
> 
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org.  For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
> 

[toc] | [prev] | [next] | [standalone]


#1647872

Fromzhong jiang <zhongjiang@huawei.com>
Date2017-05-23 11:30 +0200
Message-ID<tKeiB-1cT-3@gated-at.bofh.it>
In reply to#1647184
On 2017/5/23 0:51, Vlastimil Babka wrote:
> On 05/20/2017 05:01 AM, zhong jiang wrote:
>> On 2017/5/20 10:40, Hugh Dickins wrote:
>>> On Sat, 20 May 2017, Xishi Qiu wrote:
>>>> Here is a bug report form redhat: https://bugzilla.redhat.com/show_bug.cgi?id=1305620
>>>> And I meet the bug too. However it is hard to reproduce, and 
>>>> 624483f3ea82598("mm: rmap: fix use-after-free in __put_anon_vma") is not help.
>>>>
>>>> From the vmcore, it seems that the page is still mapped(_mapcount=0 and _count=2),
>>>> and the value of mapping is a valid address(mapping = 0xffff8801b3e2a101),
>>>> but anon_vma has been corrupted.
>>>>
>>>> Any ideas?
>>> Sorry, no.  I assume that _mapcount has been misaccounted, for example
>>> a pte mapped in on top of another pte; but cannot begin tell you where
>>> in Red Hat's kernel-3.10.0-229.4.2.el7 that might happen.
>>>
>>> Hugh
>>>
>>> .
>>>
>> Hi, Hugh
>>
>> I find the following message from the dmesg.
>>
>> [26068.316592] BUG: Bad rss-counter state mm:ffff8800a7de2d80 idx:1 val:1
>>
>> I can prove that the __mapcount is misaccount.  when task is exited. the rmap
>> still exist.
> Check if the kernel in question contains this commit: ad33bb04b2a6 ("mm:
> thp: fix SMP race condition between THP page fault and MADV_DONTNEED")
  HI, Vlastimil
 
  I miss the patch.  when I read the patch. I find the following issue. but I am sure it is right.

      if (unlikely(pmd_trans_unstable(pmd)))
        return 0;
    /*
     * A regular pmd is established and it can't morph into a huge pmd
     * from under us anymore at this point because we hold the mmap_sem
     * read mode and khugepaged takes it in write mode. So now it's
     * safe to run pte_offset_map().
     */
    pte = pte_offset_map(pmd, address);

  after pmd_trans_unstable call,  without any protect method.  by the comments,
  it think the pte_offset_map is safe.    before pte_offset_map call, it still may be
  unstable. it is possible?

  Thanks
zhongjiang
>> Thanks
>> zhongjiang
>>
>> --
>> To unsubscribe, send a message with 'unsubscribe linux-mm' in
>> the body to majordomo@kvack.org.  For more info on Linux MM,
>> see: http://www.linux-mm.org/ .
>> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
>>
>
> .
>

[toc] | [prev] | [next] | [standalone]


#1647887

FromVlastimil Babka <vbabka@suse.cz>
Date2017-05-23 11:40 +0200
Message-ID<tKesi-1hY-15@gated-at.bofh.it>
In reply to#1647872
On 05/23/2017 11:21 AM, zhong jiang wrote:
> On 2017/5/23 0:51, Vlastimil Babka wrote:
>> On 05/20/2017 05:01 AM, zhong jiang wrote:
>>> On 2017/5/20 10:40, Hugh Dickins wrote:
>>>> On Sat, 20 May 2017, Xishi Qiu wrote:
>>>>> Here is a bug report form redhat: https://bugzilla.redhat.com/show_bug.cgi?id=1305620
>>>>> And I meet the bug too. However it is hard to reproduce, and 
>>>>> 624483f3ea82598("mm: rmap: fix use-after-free in __put_anon_vma") is not help.
>>>>>
>>>>> From the vmcore, it seems that the page is still mapped(_mapcount=0 and _count=2),
>>>>> and the value of mapping is a valid address(mapping = 0xffff8801b3e2a101),
>>>>> but anon_vma has been corrupted.
>>>>>
>>>>> Any ideas?
>>>> Sorry, no.  I assume that _mapcount has been misaccounted, for example
>>>> a pte mapped in on top of another pte; but cannot begin tell you where
>>>> in Red Hat's kernel-3.10.0-229.4.2.el7 that might happen.
>>>>
>>>> Hugh
>>>>
>>>> .
>>>>
>>> Hi, Hugh
>>>
>>> I find the following message from the dmesg.
>>>
>>> [26068.316592] BUG: Bad rss-counter state mm:ffff8800a7de2d80 idx:1 val:1
>>>
>>> I can prove that the __mapcount is misaccount.  when task is exited. the rmap
>>> still exist.
>> Check if the kernel in question contains this commit: ad33bb04b2a6 ("mm:
>> thp: fix SMP race condition between THP page fault and MADV_DONTNEED")
>   HI, Vlastimil
>  
>   I miss the patch.

Try applying it then, there's good chance the error and crash will go
away. Even if your workload doesn't actually run any madvise(MADV_DONTNEED).

> when I read the patch. I find the following issue. but I am sure it is right.
> 
>       if (unlikely(pmd_trans_unstable(pmd)))
>         return 0;
>     /*
>      * A regular pmd is established and it can't morph into a huge pmd
>      * from under us anymore at this point because we hold the mmap_sem
>      * read mode and khugepaged takes it in write mode. So now it's
>      * safe to run pte_offset_map().
>      */
>     pte = pte_offset_map(pmd, address);
> 
>   after pmd_trans_unstable call,  without any protect method.  by the comments,
>   it think the pte_offset_map is safe.    before pte_offset_map call, it still may be
>   unstable. it is possible?

IIRC it's "unstable" wrt possible none->huge->none transition. But once
we've seen it's a regular pmd via pmd_trans_unstable(), we're safe as a
transition from regular pmd can't happen.

>   Thanks
> zhongjiang
>>> Thanks
>>> zhongjiang
>>>
>>> --
>>> To unsubscribe, send a message with 'unsubscribe linux-mm' in
>>> the body to majordomo@kvack.org.  For more info on Linux MM,
>>> see: http://www.linux-mm.org/ .
>>> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
>>>
>>
>> .
>>
> 
> 
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org.  For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
> 

[toc] | [prev] | [next] | [standalone]


#1647930

Fromzhong jiang <zhongjiang@huawei.com>
Date2017-05-23 12:40 +0200
Message-ID<tKfol-1W6-17@gated-at.bofh.it>
In reply to#1647887
On 2017/5/23 17:33, Vlastimil Babka wrote:
> On 05/23/2017 11:21 AM, zhong jiang wrote:
>> On 2017/5/23 0:51, Vlastimil Babka wrote:
>>> On 05/20/2017 05:01 AM, zhong jiang wrote:
>>>> On 2017/5/20 10:40, Hugh Dickins wrote:
>>>>> On Sat, 20 May 2017, Xishi Qiu wrote:
>>>>>> Here is a bug report form redhat: https://bugzilla.redhat.com/show_bug.cgi?id=1305620
>>>>>> And I meet the bug too. However it is hard to reproduce, and 
>>>>>> 624483f3ea82598("mm: rmap: fix use-after-free in __put_anon_vma") is not help.
>>>>>>
>>>>>> From the vmcore, it seems that the page is still mapped(_mapcount=0 and _count=2),
>>>>>> and the value of mapping is a valid address(mapping = 0xffff8801b3e2a101),
>>>>>> but anon_vma has been corrupted.
>>>>>>
>>>>>> Any ideas?
>>>>> Sorry, no.  I assume that _mapcount has been misaccounted, for example
>>>>> a pte mapped in on top of another pte; but cannot begin tell you where
>>>>> in Red Hat's kernel-3.10.0-229.4.2.el7 that might happen.
>>>>>
>>>>> Hugh
>>>>>
>>>>> .
>>>>>
>>>> Hi, Hugh
>>>>
>>>> I find the following message from the dmesg.
>>>>
>>>> [26068.316592] BUG: Bad rss-counter state mm:ffff8800a7de2d80 idx:1 val:1
>>>>
>>>> I can prove that the __mapcount is misaccount.  when task is exited. the rmap
>>>> still exist.
>>> Check if the kernel in question contains this commit: ad33bb04b2a6 ("mm:
>>> thp: fix SMP race condition between THP page fault and MADV_DONTNEED")
>>   HI, Vlastimil
>>  
>>   I miss the patch.
> Try applying it then, there's good chance the error and crash will go
> away. Even if your workload doesn't actually run any madvise(MADV_DONTNEED).
 ok , I will try.   Thanks
>> when I read the patch. I find the following issue. but I am sure it is right.
>>
>>       if (unlikely(pmd_trans_unstable(pmd)))
>>         return 0;
>>     /*
>>      * A regular pmd is established and it can't morph into a huge pmd
>>      * from under us anymore at this point because we hold the mmap_sem
>>      * read mode and khugepaged takes it in write mode. So now it's
>>      * safe to run pte_offset_map().
>>      */
>>     pte = pte_offset_map(pmd, address);
>>
>>   after pmd_trans_unstable call,  without any protect method.  by the comments,
>>   it think the pte_offset_map is safe.    before pte_offset_map call, it still may be
>>   unstable. it is possible?
> IIRC it's "unstable" wrt possible none->huge->none transition. But once
> we've seen it's a regular pmd via pmd_trans_unstable(), we're safe as a
> transition from regular pmd can't happen.
  Thank you for clarify. 
 
  Regards
 zhongjiang
>>   Thanks
>> zhongjiang
>>>> Thanks
>>>> zhongjiang
>>>>
>>>> --
>>>> To unsubscribe, send a message with 'unsubscribe linux-mm' in
>>>> the body to majordomo@kvack.org.  For more info on Linux MM,
>>>> see: http://www.linux-mm.org/ .
>>>> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
>>>>
>>> .
>>>
>>
>> --
>> To unsubscribe, send a message with 'unsubscribe linux-mm' in
>> the body to majordomo@kvack.org.  For more info on Linux MM,
>> see: http://www.linux-mm.org/ .
>> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
>>
> --
> To unsubscribe, send a message with 'unsubscribe linux-mm' in
> the body to majordomo@kvack.org.  For more info on Linux MM,
> see: http://www.linux-mm.org/ .
> Don't email: <a href=mailto:"dont@kvack.org"> email@kvack.org </a>
>
> .
>

[toc] | [prev] | [next] | [standalone]


#1646690

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-05-22 12:00 +0200
Message-ID<tJSi6-44C-19@gated-at.bofh.it>
In reply to#1646049
On 2017/5/20 10:40, Hugh Dickins wrote:

> On Sat, 20 May 2017, Xishi Qiu wrote:
>>
>> Here is a bug report form redhat: https://bugzilla.redhat.com/show_bug.cgi?id=1305620
>> And I meet the bug too. However it is hard to reproduce, and 
>> 624483f3ea82598("mm: rmap: fix use-after-free in __put_anon_vma") is not help.
>>
>> From the vmcore, it seems that the page is still mapped(_mapcount=0 and _count=2),
>> and the value of mapping is a valid address(mapping = 0xffff8801b3e2a101),
>> but anon_vma has been corrupted.
>>
>> Any ideas?
> 
> Sorry, no.  I assume that _mapcount has been misaccounted, for example
> a pte mapped in on top of another pte; but cannot begin tell you where

Hi Hugh,

What does "a pte mapped in on top of another pte" mean? Could you give more info?

Thanks,
Xishi Qiu

> in Red Hat's kernel-3.10.0-229.4.2.el7 that might happen.
> 
> Hugh
> 
> .
> 

[toc] | [prev] | [next] | [standalone]


#1647298

FromHugh Dickins <hughd@google.com>
Date2017-05-22 21:30 +0200
Message-ID<tK1bJ-1hH-29@gated-at.bofh.it>
In reply to#1646690
On Mon, 22 May 2017, Xishi Qiu wrote:
> On 2017/5/20 10:40, Hugh Dickins wrote:
> > On Sat, 20 May 2017, Xishi Qiu wrote:
> >>
> >> Here is a bug report form redhat: https://bugzilla.redhat.com/show_bug.cgi?id=1305620
> >> And I meet the bug too. However it is hard to reproduce, and 
> >> 624483f3ea82598("mm: rmap: fix use-after-free in __put_anon_vma") is not help.
> >>
> >> From the vmcore, it seems that the page is still mapped(_mapcount=0 and _count=2),
> >> and the value of mapping is a valid address(mapping = 0xffff8801b3e2a101),
> >> but anon_vma has been corrupted.
> >>
> >> Any ideas?
> > 
> > Sorry, no.  I assume that _mapcount has been misaccounted, for example
> > a pte mapped in on top of another pte; but cannot begin tell you where
> 
> Hi Hugh,
> 
> What does "a pte mapped in on top of another pte" mean? Could you give more info?

I mean, there are various places in mm/memory.c which decide what they
intend to do based on orig_pte, then take pte lock, then check that
pte_same(pte, orig_pte) before taking it any further.  If a pte_same()
check were missing (I do not know of any such case), then two racing
tasks might install the same pte, one on top of the other - page
mapcount being incremented twice, but decremented only once when
that pte is finally unmapped later.

Please see similar discussion in the earlier thread at
marc.info/?l=linux-mm&m=148222656211837&w=2

Hugh

> 
> Thanks,
> Xishi Qiu
> 
> > in Red Hat's kernel-3.10.0-229.4.2.el7 that might happen.
> > 
> > Hugh

[toc] | [prev] | [next] | [standalone]


#1647587

FromXishi Qiu <qiuxishi@huawei.com>
Date2017-05-23 04:30 +0200
Message-ID<tK7Ka-5sv-13@gated-at.bofh.it>
In reply to#1647298
On 2017/5/23 3:26, Hugh Dickins wrote:

> On Mon, 22 May 2017, Xishi Qiu wrote:
>> On 2017/5/20 10:40, Hugh Dickins wrote:
>>> On Sat, 20 May 2017, Xishi Qiu wrote:
>>>>
>>>> Here is a bug report form redhat: https://bugzilla.redhat.com/show_bug.cgi?id=1305620
>>>> And I meet the bug too. However it is hard to reproduce, and 
>>>> 624483f3ea82598("mm: rmap: fix use-after-free in __put_anon_vma") is not help.
>>>>
>>>> From the vmcore, it seems that the page is still mapped(_mapcount=0 and _count=2),
>>>> and the value of mapping is a valid address(mapping = 0xffff8801b3e2a101),
>>>> but anon_vma has been corrupted.
>>>>
>>>> Any ideas?
>>>
>>> Sorry, no.  I assume that _mapcount has been misaccounted, for example
>>> a pte mapped in on top of another pte; but cannot begin tell you where
>>
>> Hi Hugh,
>>
>> What does "a pte mapped in on top of another pte" mean? Could you give more info?
> 
> I mean, there are various places in mm/memory.c which decide what they
> intend to do based on orig_pte, then take pte lock, then check that
> pte_same(pte, orig_pte) before taking it any further.  If a pte_same()
> check were missing (I do not know of any such case), then two racing
> tasks might install the same pte, one on top of the other - page
> mapcount being incremented twice, but decremented only once when
> that pte is finally unmapped later.
> 

Hi Hugh,

Do you mean that the ptes from two racing point to the same page?
or the two racing point to two pages, but one covers the other later?
and the first page maybe alone in the lru list, and it will never be freed
when the process exit.

We got this info before crash.
[26068.316592] BUG: Bad rss-counter state mm:ffff8800a7de2d80 idx:1 val:1

Thanks,
Xishi Qiu

> Please see similar discussion in the earlier thread at
> marc.info/?l=linux-mm&m=148222656211837&w=2
> 
> Hugh
> 
>>
>> Thanks,
>> Xishi Qiu
>>
>>> in Red Hat's kernel-3.10.0-229.4.2.el7 that might happen.
>>>
>>> Hugh
> 
> .
> 

[toc] | [prev] | [next] | [standalone]


#1647593

FromHugh Dickins <hughd@google.com>
Date2017-05-23 05:00 +0200
Message-ID<tK8db-5Bc-3@gated-at.bofh.it>
In reply to#1647587
On Tue, 23 May 2017, Xishi Qiu wrote:
> On 2017/5/23 3:26, Hugh Dickins wrote:
> > I mean, there are various places in mm/memory.c which decide what they
> > intend to do based on orig_pte, then take pte lock, then check that
> > pte_same(pte, orig_pte) before taking it any further.  If a pte_same()
> > check were missing (I do not know of any such case), then two racing
> > tasks might install the same pte, one on top of the other - page
> > mapcount being incremented twice, but decremented only once when
> > that pte is finally unmapped later.
> > 
> 
> Hi Hugh,
> 
> Do you mean that the ptes from two racing point to the same page?
> or the two racing point to two pages, but one covers the other later?
> and the first page maybe alone in the lru list, and it will never be freed
> when the process exit.
> 
> We got this info before crash.
> [26068.316592] BUG: Bad rss-counter state mm:ffff8800a7de2d80 idx:1 val:1

I might mean either: you are taking my suggestion too seriously,
it is merely a suggestion of one way in which this could happen.

Another way is ordinary memory corruption (whether by software error
or by flipped DRAM bits) of a page table: that could end up here too.

Hugh

[toc] | [prev] | [standalone]


Back to top | Article view | linux.kernel


csiph-web