Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #207487 > unrolled thread

Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch)

Started byMartin Schwarz <debian-lists@alias.kuroi.de>
First post2019-04-15 14:50 +0200
Last post2019-04-17 01:40 +0200
Articles 16 — 4 participants

Back to article view | Back to linux.debian.user


Contents

  Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Martin Schwarz <debian-lists@alias.kuroi.de> - 2019-04-15 14:50 +0200
    Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Reco <recoverym4n@enotuniq.net> - 2019-04-15 15:40 +0200
      Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Martin Schwarz <debian-lists@alias.kuroi.de> - 2019-04-15 16:50 +0200
        Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Reco <recoverym4n@enotuniq.net> - 2019-04-15 18:10 +0200
          Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Martin Schwarz <debian-lists@alias.kuroi.de> - 2019-04-16 10:30 +0200
            Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Reco <recoverym4n@enotuniq.net> - 2019-04-16 11:00 +0200
              Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Martin Schwarz <debian-lists@alias.kuroi.de> - 2019-04-16 14:40 +0200
                Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Reco <recoverym4n@enotuniq.net> - 2019-04-16 15:10 +0200
    Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Kenneth Parker <sea7kenp@gmail.com> - 2019-04-15 16:50 +0200
      Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Martin Schwarz <debian-lists@alias.kuroi.de> - 2019-04-15 17:10 +0200
    Re: Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch) Peter Wiersig <peter@friesenpeter.de> - 2019-04-16 16:40 +0200
      Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Reco <recoverym4n@enotuniq.net> - 2019-04-16 18:50 +0200
        Re: Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch) Peter Wiersig <peter@friesenpeter.de> - 2019-04-17 01:20 +0200
          Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Reco <recoverym4n@enotuniq.net> - 2019-04-17 07:50 +0200
      Re: Need help analyzing (kernel?) memory usage and reclaiming RAM  (Debian Stretch) Martin Schwarz <debian-lists@alias.kuroi.de> - 2019-04-17 11:40 +0200
    Re: Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch) Peter Wiersig <peter@friesenpeter.de> - 2019-04-17 01:40 +0200

#207487 — Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch)

FromMartin Schwarz <debian-lists@alias.kuroi.de>
Date2019-04-15 14:50 +0200
SubjectNeed help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch)
Message-ID<xN9dD-3LR-3@gated-at.bofh.it>
Hello,

(please let me know if this is more appropriate somewhere else, e.g. on
ebian-kernel)

I need help debugging/solving a weird memory problem. The symptoms are
the usual ones for high memory usage: free/available memory is getting
low, systems start swapping, disk I/O increases, performance drops.

However, from what I can see, the memory is not used up by user space
processes but from the Kernel (NOT caches/buffers), see commands output
at the end.

I'm still puzzled about what exactly eats all the RAM and how to reclaim
it (without rebooting the machine, of course!). Any help would be highly
appreciated!

Some findings so far:

- same problem on many systems, all Debian 9 Stretch, all running stock
  4.9 kernel from the official package, all amd64 virtual machines on
  several (different) VMware ESXi hosts.
- not all Stretch systems seem to be affected, but we haven't yet found
  the common ground.
- problem can occur after some days or some weeks, not at the same time
  on all affected machines. And not at the same time for all VMs on the
  same host
- problem only occurs on Stretch systems, not Jessie, even running on
  the same host.
- we haven't yet seen the problem on real hardware machines, only VMs
  (but since the vast majority of our systems are VMs, this may not be
  relevant)
- problem seems not directly related to the machine's load. it occurs on
  machines that are mostly idle as well as on more heavily-loaded
  systems
- problem occurs the same on single-core VMs as well as on multi-core
  VMs
- problem occurs the same on VMs running on single-socket hosts as well
  as on multi-socket hosts
- problem occurs the same on VMs running on hosts with different
  hypervisor releases, both VMware ESXi 5.5 and 6.5, both standalone and
  in a vSphere cluster.

Here's the output from some commands I hope to be helpful:

The machine in this example is a RADIUS server but has not even gone
productive ... no incoming client requests yet.  (But the problem is not
related to the RADIUS server software - OSC Radiator - since the same
symptoms show on different machines: not only RADIUS servers but also
nameservers, shell servers or jumphosts, etc.)

[values while the problem persists:]
------------------------------------------------------------------------
root@rad-m2m-srv02:~# free -thwl
              total        used        free      shared     buffers       cache   available
Mem:           987M        910M         59M          0B        704K         16M         13M
Low:           987M        927M         59M
High:            0B          0B          0B
Swap:          2,0G        345M        1,7G
Total:         3,0G        1,2G        1,7G
root@rad-m2m-srv02:~# smem -twk
Area                           Used      Cache   Noncache 
firmware/hardware                 0          0          0 
kernel image                      0          0          0 
kernel dynamic memory        914.9M      11.1M     903.8M 
userspace memory              13.0M       5.5M       7.4M 
free memory                   59.4M      59.4M          0 
----------------------------------------------------------
                             987.3M      76.1M     911.2M 
root@rad-m2m-srv02:~# smem -uktr
User     Count     Swap      USS      PSS      RSS 
root        39   332.8M    10.4M    12.4M    44.7M 
msch         6     7.0M        0   607.0K     8.3M 
_chrony      1   360.0K     4.0K    20.0K   572.0K 
messagebus     1   580.0K     4.0K    17.0K   480.0K 
postfix      2     1.6M        0    13.0K   568.0K 
daemon       1   208.0K     4.0K     6.0K    72.0K 
---------------------------------------------------
            50   342.5M    10.4M    13.0M    54.7M 
root@rad-m2m-srv02:~# sort -k2,2nr /proc/meminfo
VmallocTotal:   34359738367 kB
CommitLimit:     2602636 kB
SwapTotal:       2097148 kB
SwapFree:        1741028 kB
MemTotal:        1010976 kB
DirectMap4k:     1007488 kB
Committed_AS:     465128 kB
Slab:              79680 kB
SUnreclaim:        69268 kB
MemFree:           61068 kB
DirectMap2M:       40960 kB
SReclaimable:      10412 kB
Active:             6944 kB
Inactive:           6660 kB
AnonPages:          6608 kB
PageTables:         5804 kB
Cached:             5748 kB
Mapped:             4660 kB
SwapCached:         3988 kB
Active(file):       3920 kB
Inactive(anon):     3828 kB
Active(anon):       3024 kB
KernelStack:        2992 kB
Inactive(file):     2832 kB
Hugepagesize:       2048 kB
Buffers:            1020 kB
Dirty:                 8 kB
AnonHugePages:         0 kB
Bounce:                0 kB
HardwareCorrupted:     0 kB
HugePages_Free:        0
HugePages_Rsvd:        0
HugePages_Surp:        0
HugePages_Total:       0
MemAvailable:          0 kB
Mlocked:               0 kB
NFS_Unstable:          0 kB
Shmem:                 0 kB
ShmemHugePages:        0 kB
ShmemPmdMapped:        0 kB
Unevictable:           0 kB
VmallocChunk:          0 kB
VmallocUsed:           0 kB
Writeback:             0 kB
WritebackTmp:          0 kB
root@rad-m2m-srv02:~# ps aux --sort=-rss | head -15
USER       PID %CPU %MEM    VSZ   RSS TTY      STAT START   TIME COMMAND
root     34718 12.0  0.5  29596  5672 ?        D    09:01   0:00 /usr/bin/python3 -Es /usr/bin/lsb_release --short --description
root     26491  3.1  0.2  79328  2860 ?        D    08:04   1:50 apt-get update -qq
root     32551  6.8  0.2 119036  2800 ?        D    08:51   0:43 /usr/bin/python3 /usr/bin/unattended-upgrade
root     34719  0.0  0.2  41164  2232 pts/1    R+   09:02   0:00 ps aux --sort=-rss
msch     33960  0.1  0.1  23720  1844 pts/0    Ss   08:58   0:00 -bash
root     34492  0.2  0.1  23816  1812 pts/1    S    09:00   0:00 -bash
msch     33996  0.0  0.1  23576  1768 pts/1    Ss   08:58   0:00 bash -i
root     12792  2.2  0.1 159720  1748 ?        D    06:06   3:54 /usr/bin/perl -w /usr/bin/apt-show-versions -i
root     34521  0.7  0.1  95180  1712 ?        Ss   09:01   0:00 sshd: root@notty
root     15502  2.4  0.1 167660  1608 ?        D    06:25   3:51 /usr/bin/perl -w /usr/bin/apt-show-versions -i
root     34527  1.7  0.1  14096  1596 ?        Ss   09:01   0:00 /bin/bash /usr/bin/check_mk_agent
root     33947  0.0  0.1  95180  1564 ?        Ss   08:58   0:00 sshd: msch [priv]
root     26486  0.0  0.1   9600  1436 ?        S    08:04   0:00 /bin/bash 3600/mk_apt
root     26483  0.0  0.1   9588  1424 ?        S    08:04   0:00 /bin/bash
root@rad-m2m-srv02:~# lsof | wc -l
1943
root@rad-m2m-srv02:~# df -Th -t tmpfs
Filesystem     Type   Size  Used Avail Use% Mounted on
tmpfs          tmpfs   99M   12M   87M  12% /run
tmpfs          tmpfs  494M     0  494M   0% /dev/shm
tmpfs          tmpfs  5,0M     0  5,0M   0% /run/lock
tmpfs          tmpfs  494M     0  494M   0% /sys/fs/cgroup
tmpfs          tmpfs  1,0G     0  1,0G   0% /tmp
tmpfs          tmpfs   99M     0   99M   0% /run/user/0
tmpfs          tmpfs   99M     0   99M   0% /run/user/2029
root@rad-m2m-srv02:~# vmware-toolbox-cmd stat balloon
0 MB
root@rad-m2m-srv02:~# cat /sys/kernel/debug/vmmemctl
balloon capabilities:   0x1e
used capabilities:      0x1e
is resetting:           n
target:                    0 pages
current:                   0 pages
rateSleepAlloc:         2048 pages/sec

timer:               3968363
doorbell:                  0
start:                     7 (   0 failed)
guestType:                 7 (   0 failed)
2m-lock:                   0 (   0 failed)
lock:                      0 (   0 failed)
2m-unlock:                 0 (   0 failed)
unlock:                    0 (   0 failed)
target:              3968363 (   6 failed)
prim2mAlloc:               0 (   0 failed)
primNoSleepAlloc:          0 (   0 failed)
primCanSleepAlloc:         0 (   0 failed)
prim2mFree:                0
primFree:                  0
err2mAlloc:                0
errAlloc:                  0
err2mFree:                 0
errFree:                   0
doorbellSet:               6
doorbellUnset:             7
root@rad-m2m-srv02:~# nice vmstat -w 1 10
procs -----------------------memory---------------------- ---swap-- -----io---- -system-- --------cpu--------
 r  b         swpd         free         buff        cache   si   so    bi    bo   in   cs  us  sy  id  wa  st
 0  5       356620        60868         1140        16280   37   19   704    31    3    2   1   2  97   1   0
 1  4       356180        60372          320        16224 3008  624  6180  1236 1109 1915   2  18   0  80   0
 2  5       356632        61476          320        15568 2776 1452  3128  2012 1146 1802   1  14   0  85   0
 1  3       356592        62228          324        15244 2848  952  3784  1564 1029 1780   0  11   0  89   0
 2  4       356732        61492          612        15544 2864 1144  3932  1720 1164 1839   2   9   0  89   0
 1  4       357252        62836          556        15248 4000 1800  4432  3048 1398 2359   1  15   0  84   0
 0  4       356700        61744          448        15248 3368  668  3368  1276 1093 2039   0   9   0  91   0
 2  4       356708        61372          456        16272 1940  868  4744   888  876 1377   0  12   0  88   0
 0  4       356704        61744         1156        14700 2740  660  4828  1940 1123 1768   0  14   0  86   0
 0  4       357556        62240          680        15568 2908 1476  5436  2064 1062 1804   1  15   0  84   0
root@rad-m2m-srv02:~# lsb_release -a
No LSB modules are available.
Distributor ID:	Debian
Description:	Debian GNU/Linux 9.8 (stretch)
Release:	9.8
Codename:	stretch
root@rad-m2m-srv02:~# uname -a
Linux rad-m2m-srv02 4.9.0-8-amd64 #1 SMP Debian 4.9.144-3 (2019-02-02) x86_64 GNU/Linux
root@rad-m2m-srv02:~# w
 09:02:30 up 45 days, 22:20,  1 user,  load average: 5,13, 5,03, 6,58
USER     TTY      FROM             LOGIN@   IDLE   JCPU   PCPU WHAT
msch     pts/0    10.208.105.87    08:58    4.00s  0.26s  0.03s script memdebug
root@rad-m2m-srv02:~# 

[values directly after rebooting:]
------------------------------------------------------------------------
root@rad-m2m-srv02:~# w
 09:23:02 up 4 min,  1 user,  load average: 0,01, 0,08, 0,04
USER     TTY      FROM             LOGIN@   IDLE   JCPU   PCPU WHAT
msch     pts/0    10.208.105.87    09:21    6.00s  0.26s  0.02s sshd: msch [priv]   
root@rad-m2m-srv02:~# free -thwl
              total        used        free      shared     buffers       cache   available
Mem:           987M        112M        610M        4,3M         16M        247M        735M
Low:           987M        377M        610M
High:            0B          0B          0B
Swap:          2,0G          0B        2,0G
Total:         3,0G        112M        2,6G
root@rad-m2m-srv02:~# smem -twk
Area                           Used      Cache   Noncache 
firmware/hardware                 0          0          0 
kernel image                      0          0          0 
kernel dynamic memory        287.1M     226.6M      60.5M 
userspace memory              93.8M      37.8M      56.0M 
free memory                  606.4M     606.4M          0 
----------------------------------------------------------
                             987.3M     870.8M     116.5M 
root@rad-m2m-srv02:~# smem -uktr
User     Count     Swap      USS      PSS      RSS 
root        19        0    62.9M    72.8M   128.7M 
postfix      6        0     7.9M    12.1M    42.9M 
msch         4        0     3.7M     7.3M    19.4M 
messagebus     1        0     1.2M     1.5M     3.8M 
_chrony      1        0   896.0K  1020.0K     2.8M 
daemon       1        0   228.0K   309.0K     2.1M 
---------------------------------------------------
            32        0    76.9M    95.0M   199.7M 
root@rad-m2m-srv02:~# sort -k2,2nr /proc/meminfo
VmallocTotal:   34359738367 kB
CommitLimit:     2602636 kB
SwapFree:        2097148 kB
SwapTotal:       2097148 kB
MemTotal:        1010976 kB
DirectMap2M:      983040 kB
MemAvailable:     753520 kB
MemFree:          624520 kB
Cached:           234508 kB
Active:           161672 kB
Inactive:         142964 kB
Inactive(file):   138936 kB
Committed_AS:     124808 kB
Active(file):     108028 kB
DirectMap4k:       65408 kB
Active(anon):      53644 kB
AnonPages:         53300 kB
Slab:              36968 kB
Mapped:            36760 kB
SReclaimable:      19424 kB
SUnreclaim:        17544 kB
Buffers:           16836 kB
Shmem:              4392 kB
Inactive(anon):     4028 kB
PageTables:         3836 kB
KernelStack:        2748 kB
Hugepagesize:       2048 kB
Dirty:                60 kB
AnonHugePages:         0 kB
Bounce:                0 kB
HardwareCorrupted:     0 kB
HugePages_Free:        0
HugePages_Rsvd:        0
HugePages_Surp:        0
HugePages_Total:       0
Mlocked:               0 kB
NFS_Unstable:          0 kB
ShmemHugePages:        0 kB
ShmemPmdMapped:        0 kB
SwapCached:            0 kB
Unevictable:           0 kB
VmallocChunk:          0 kB
VmallocUsed:           0 kB
Writeback:             0 kB
WritebackTmp:          0 kB
root@rad-m2m-srv02:~# ps aux --sort=-rss | head -15
USER       PID %CPU %MEM    VSZ   RSS TTY      STAT START   TIME COMMAND
root       651  0.1  2.6  78748 26992 ?        S    09:18   0:00 /usr/bin/perl /opt/radiator/bin/radiusd -daemon -pid_file /var/run/radiator.pid -config_file /opt/radiator/etc/radiator.cfg -I /opt/radiator/share/perl/5.24.1/
root       411  0.0  1.7 153488 18144 ?        Ss   09:18   0:00 /usr/bin/VGAuthService
root       221  0.1  1.0 136488 10464 ?        Ss   09:18   0:00 /usr/bin/vmtoolsd
postfix   2033  0.1  0.8  88652  8968 ?        S    09:22   0:00 smtp -t unix -u
postfix   2034  0.0  0.8  87480  8132 ?        S    09:22   0:00 tlsmgr -l -t unix -u
root         1  0.3  0.6  57052  6736 ?        Ss   09:18   0:00 /sbin/init
root      1462  0.0  0.6  95180  6736 ?        Ss   09:21   0:00 sshd: msch [priv]
postfix   2031  0.0  0.6  83352  6700 ?        S    09:22   0:00 cleanup -z -t unix -u
postfix    649  0.0  0.6  83296  6600 ?        S    09:18   0:00 qmgr -l -t unix -u
postfix   2032  0.0  0.6  83260  6600 ?        S    09:22   0:00 trivial-rewrite -n rewrite -t unix -u
postfix    648  0.0  0.6  83248  6284 ?        S    09:18   0:00 pickup -l -t unix -u
root       527  0.0  0.6  69952  6168 ?        Ss   09:18   0:00 /usr/sbin/sshd -D
msch      1464  0.0  0.6  64832  6144 ?        Ss   09:21   0:00 /lib/systemd/systemd --user
root       251  0.0  0.5  47844  5872 ?        Ss   09:18   0:00 /lib/systemd/systemd-udevd
root@rad-m2m-srv02:~# lsof | wc -l
1605
root@rad-m2m-srv02:~# df -Th -t tmpfs
Filesystem     Type   Size  Used Avail Use% Mounted on
tmpfs          tmpfs   99M  4,3M   95M   5% /run
tmpfs          tmpfs  494M     0  494M   0% /dev/shm
tmpfs          tmpfs  5,0M     0  5,0M   0% /run/lock
tmpfs          tmpfs  494M     0  494M   0% /sys/fs/cgroup
tmpfs          tmpfs  1,0G     0  1,0G   0% /tmp
tmpfs          tmpfs   99M     0   99M   0% /run/user/2029
root@rad-m2m-srv02:~# vmware-toolbox-cmd stat balloon
0 MB
root@rad-m2m-srv02:~# cat /sys/kernel/debug/vmmemctl
balloon capabilities:   0x1e
used capabilities:      0x1e
is resetting:           n
target:                    0 pages
current:                   0 pages
rateSleepAlloc:         2048 pages/sec

timer:                   292
doorbell:                  0
start:                     1 (   0 failed)
guestType:                 1 (   0 failed)
2m-lock:                   0 (   0 failed)
lock:                      0 (   0 failed)
2m-unlock:                 0 (   0 failed)
unlock:                    0 (   0 failed)
target:                  292 (   0 failed)
prim2mAlloc:               0 (   0 failed)
primNoSleepAlloc:          0 (   0 failed)
primCanSleepAlloc:         0 (   0 failed)
prim2mFree:                0
primFree:                  0
err2mAlloc:                0
errAlloc:                  0
err2mFree:                 0
errFree:                   0
doorbellSet:               1
doorbellUnset:             1
root@rad-m2m-srv02:~# nice vmstat -w 1 10
procs -----------------------memory---------------------- ---swap-- -----io---- -system-- --------cpu--------
 r  b         swpd         free         buff        cache   si   so    bi    bo   in   cs  us  sy  id  wa  st
 0  0            0       622948        16868       254624    0    0   728   254  104  231   4   2  88   5   0
 0  0            0       622948        16868       254624    0    0     0     0   53   98   0   0 100   0   0
 0  0            0       622948        16876       254600    0    0     0    20   50   96   0   0 100   0   0
 0  0            0       622948        16876       254600    0    0     0     0   50   91   0   0 100   0   0
 0  0            0       622948        16876       254600    0    0     0     0   43   84   0   0 100   0   0
 0  0            0       622948        16876       254604    0    0     0     0   57  105   1   0  99   0   0
 0  0            0       622948        16876       254600    0    0     0     0   53  106   0   1  99   0   0
 0  0            0       622948        16876       254600    0    0     0     0   50   91   1   0  99   0   0
 1  0            0       622948        16876       254600    0    0     0     0   49   96   0   0 100   0   0
 0  0            0       622948        16876       254600    0    0     0    12   50   94   0   1  99   0   0
root@rad-m2m-srv02:~# 

------------------------------------------------------------------------

Anything else I could check to help pinpoint the memory hog?


Thanks in advance!
Martin

-- 
Martin Schwarz * Karlsruhe, Germany * http://kuroi.de/

[toc] | [next] | [standalone]


#207491

FromReco <recoverym4n@enotuniq.net>
Date2019-04-15 15:40 +0200
Message-ID<xNa01-4hU-1@gated-at.bofh.it>
In reply to#207487
	Hi.

On Mon, Apr 15, 2019 at 02:21:16PM +0200, Martin Schwarz wrote:
> I need help debugging/solving a weird memory problem. The symptoms are
> the usual ones for high memory usage: free/available memory is getting
> low, systems start swapping, disk I/O increases, performance drops.

Can you please provide unsorted outputs of /proc/meminfo? It's easier to
compare them if they are unsorted.
And "smem -tm | tail" would be helpful too.

Reco

[toc] | [prev] | [next] | [standalone]


#207502

FromMartin Schwarz <debian-lists@alias.kuroi.de>
Date2019-04-15 16:50 +0200
Message-ID<xNb5N-4Vm-27@gated-at.bofh.it>
In reply to#207491
On Mon, Apr 15, 2019 at 04:35:27PM +0300, Reco wrote:
> Can you please provide unsorted outputs of /proc/meminfo? It's easier to
> compare them if they are unsorted.
> And "smem -tm | tail" would be helpful too.

Thanks for your input!

The system from my previous example has already been rebooted, sorry!
But here's from another system that currently starts showing the same
problem and has an equally small workload:

------------------------------------------------------------------------
root@rad-wgv-srv01:~# free -thwl
              total        used        free      shared     buffers       cache   available
Mem:           987M        843M         72M        1,1M        9,7M         61M         37M
Low:           987M        914M         72M
High:            0B          0B          0B
Swap:          2,0G         75M        1,9G
Total:         3,0G        919M        2,0G
root@rad-wgv-srv01:~# cat /proc/meminfo
MemTotal:        1010976 kB
MemFree:           73980 kB
MemAvailable:      38756 kB
Buffers:            9964 kB
Cached:            50340 kB
SwapCached:         2728 kB
Active:            58776 kB
Inactive:          15164 kB
Active(anon):      11068 kB
Inactive(anon):     3696 kB
Active(file):      47708 kB
Inactive(file):    11468 kB
Unevictable:           0 kB
Mlocked:               0 kB
SwapTotal:       2097148 kB
SwapFree:        2019416 kB
Dirty:               104 kB
Writeback:             0 kB
AnonPages:         13048 kB
Mapped:            19904 kB
Shmem:              1120 kB
Slab:              90744 kB
SReclaimable:      13100 kB
SUnreclaim:        77644 kB
KernelStack:        2700 kB
PageTables:         3764 kB
NFS_Unstable:          0 kB
Bounce:                0 kB
WritebackTmp:          0 kB
CommitLimit:     2602636 kB
Committed_AS:     155208 kB
VmallocTotal:   34359738367 kB
VmallocUsed:           0 kB
VmallocChunk:          0 kB
HardwareCorrupted:     0 kB
AnonHugePages:         0 kB
ShmemHugePages:        0 kB
ShmemPmdMapped:        0 kB
HugePages_Total:       0
HugePages_Free:        0
HugePages_Rsvd:        0
HugePages_Surp:        0
Hugepagesize:       2048 kB
DirectMap4k:      925568 kB
DirectMap2M:      122880 kB
root@rad-wgv-srv01:~# smem -tm | tail
/bin/bash                                    3      358     1076 
/lib/systemd/systemd                         3      386     1158 
/lib/x86_64-linux-gnu/libc-2.24.so          33       54     1783 
/usr/lib/x86_64-linux-gnu/libcrypto.so.1     5      386     1933 
/usr/bin/python2.7                           1     2220     2220 
/lib/systemd/libsystemd-shared-232.so        5      544     2723 
<anonymous>                                 33      146     4848 
[heap]                                      33      304    10060 
-----------------------------------------------------------------
179                                        922    11110    41011 
root@rad-wgv-srv01:~# ps aux --sort=-rss | head -15
USER       PID %CPU %MEM    VSZ   RSS TTY      STAT START   TIME COMMAND
root     61714  0.0  0.6  95180  6676 ?        Ss   16:35   0:00 sshd: msch [priv]
msch     61716  0.0  0.5  64832  5852 ?        Ss   16:35   0:00 /lib/systemd/systemd --user
msch     61723  0.0  0.4  95452  4996 ?        S    16:35   0:00 sshd: msch@pts/0
msch     61724  0.2  0.4  23636  4924 pts/0    Ss   16:35   0:00 -bash
root     62015  1.0  0.4  23740  4892 pts/0    S    16:36   0:00 -bash
root         1  0.0  0.4  57128  4440 ?        Ss   Mär29  14:30 /lib/systemd/systemd --system --deserialize 19
root     62014  0.0  0.3  51964  3836 pts/0    S    16:36   0:00 sudo -i
root     62040  0.0  0.3  41164  3584 pts/0    R+   16:37   0:00 ps aux --sort=-rss
root       228  0.1  0.2 136620  2884 ?        Ss   Mär29  31:20 /usr/bin/vmtoolsd
root       224  0.0  0.2  55556  2560 ?        Ss   Mär29  13:48 /lib/systemd/systemd-journald
root       411  0.0  0.2  46520  2408 ?        Ss   Mär29   6:40 /lib/systemd/systemd-logind
message+   419  0.0  0.1  45200  1332 ?        Ss   Mär29   6:39 /usr/bin/dbus-daemon --system --address=systemd: --nofork --nopidfile --systemd-activation
msch     61717  0.0  0.1  82652  1268 ?        S    16:35   0:00 (sd-pam)
_chrony    531  0.0  0.1  29980  1184 ?        S    Mär29   0:52 /usr/sbin/chronyd
root@rad-wgv-srv01:~# uname -a
Linux rad-wgv-srv01 4.9.0-8-amd64 #1 SMP Debian 4.9.144-3 (2019-02-02) x86_64 GNU/Linux
root@rad-wgv-srv01:~# 

------------------------------------------------------------------------
(any other commands I should repeat on the new system?)

Thanks
Martin

-- 
Martin Schwarz * Karlsruhe, Germany * http://kuroi.de/

[toc] | [prev] | [next] | [standalone]


#207509

FromReco <recoverym4n@enotuniq.net>
Date2019-04-15 18:10 +0200
Message-ID<xNclc-5Rv-13@gated-at.bofh.it>
In reply to#207502
	Hi.

On Mon, Apr 15, 2019 at 04:40:56PM +0200, Martin Schwarz wrote:
> The system from my previous example has already been rebooted, sorry!

Kind of expected. It's useful nevertheless.


> But here's from another system that currently starts showing the same
> problem and has an equally small workload:
> 
> root@rad-wgv-srv01:~# free -thwl

Nothing out of the ordinary here.

> root@rad-wgv-srv01:~# cat /proc/meminfo
> MemTotal:        1010976 kB
> MemFree:           73980 kB
> MemAvailable:      38756 kB
> Buffers:            9964 kB
> Cached:            50340 kB

It's not the file cache who ate the memory.

> SwapCached:         2728 kB

And it's not the swap caching.

> Active(anon):      11068 kB
> Inactive(anon):     3696 kB

Memory consumption cannot be attributed to tmpfs.
I know, you've posted 'df' output earlier, but it does not take mount
namespaces into the account.

> Mapped:            19904 kB

To my biggest disappointment, the problem cannot be explained by
excessive use of mmap(2) syscall. Would be easy otherwise.


> Shmem:              1120 kB

It's not the shared memory segments.


> Slab:              90744 kB
> SReclaimable:      13100 kB
> SUnreclaim:        77644 kB

And it's not dentries cache (saw the thing grown once or twice. was
ugly).


> AnonHugePages:         0 kB
> ShmemHugePages:        0 kB
> ShmemPmdMapped:        0 kB
> HugePages_Total:       0
> HugePages_Free:        0
> HugePages_Rsvd:        0
> HugePages_Surp:        0

And last, but not the least, there are no hugepages in use.


> root@rad-wgv-srv01:~# smem -tm | tail
> /bin/bash                                    3      358     1076 
> /lib/systemd/systemd                         3      386     1158 
> /lib/x86_64-linux-gnu/libc-2.24.so          33       54     1783 
> /usr/lib/x86_64-linux-gnu/libcrypto.so.1     5      386     1933 
> /usr/bin/python2.7                           1     2220     2220 
> /lib/systemd/libsystemd-shared-232.so        5      544     2723 
> <anonymous>                                 33      146     4848 
> [heap]                                      33      304    10060 
> -----------------------------------------------------------------
> 179                                        922    11110    41011 

Moreover, no current running visible process consume the memory.
I suspect that this host does not utilize them anyway.


In short. I do believe that this is happening, but I never seen anything
like this. I cannot imagine the scenario that can lead to this, as long
as we're talking real hardware aka big iron.

What I suspect is happening here is runaway memory allocation by a
kernel module (at least one of them), and said kernel module is likely
to be VMWare-specific.
It could be vmxnet3 (network). It could be that LSI kernel module or
whatever they're using for SCSI these days (vmw_pvscsi?).


And that means - 'perf top', or better yet - 'perf record'.

Reco

[toc] | [prev] | [next] | [standalone]


#207542

FromMartin Schwarz <debian-lists@alias.kuroi.de>
Date2019-04-16 10:30 +0200
Message-ID<xNrDz-6Tn-3@gated-at.bofh.it>
In reply to#207509
Hello,

On Mon, Apr 15, 2019 at 07:03:13PM +0300, Reco wrote:
[...]
> What I suspect is happening here is runaway memory allocation by a
> kernel module (at least one of them), and said kernel module is likely
> to be VMWare-specific.
> It could be vmxnet3 (network). It could be that LSI kernel module or
> whatever they're using for SCSI these days (vmw_pvscsi?).

sounds interesting.

That would explain why I haven't seen this problem on one of my (few)
personal Stretch installations running as Xen DomU. But then, I guess
we're not the only ones who use Debian Stretch on VMware ESXi ;-)
but haven't found any mention of this problem. I wonder what makes our
setup so special ...

Yes, we're using vmw_pvscsi (VMware Paravirtual SCSI controller).

Here's the outout from lsmod:

msch@rad-wgv-srv01:~$ lsmod 
Module                  Size  Used by
tcp_diag               16384  0
inet_diag              20480  1 tcp_diag
ppdev                  20480  0
vmw_balloon            20480  0
joydev                 20480  0
evdev                  24576  1
pcspkr                 16384  0
serio_raw              16384  0
vmwgfx                237568  1
ttm                    98304  1 vmwgfx
drm_kms_helper        155648  1 vmwgfx
sg                     32768  0
drm                   360448  4 vmwgfx,ttm,drm_kms_helper
shpchp                 36864  0
parport_pc             28672  0
parport                49152  2 parport_pc,ppdev
ac                     16384  0
button                 16384  0
vmw_vsock_vmci_transport    28672  0
vsock                  36864  1 vmw_vsock_vmci_transport
vmw_vmci               69632  2 vmw_balloon,vmw_vsock_vmci_transport
ip_tables              24576  0
x_tables               36864  1 ip_tables
autofs4                40960  2
ext4                  585728  3
crc16                  16384  1 ext4
jbd2                  106496  1 ext4
crc32c_generic         16384  0
fscrypto               28672  1 ext4
ecb                    16384  0
glue_helper            16384  0
lrw                    16384  0
gf128mul               16384  1 lrw
ablk_helper            16384  0
cryptd                 24576  1 ablk_helper
aes_x86_64             20480  0
mbcache                16384  4 ext4
dm_mod                118784  16
sr_mod                 24576  0
cdrom                  61440  1 sr_mod
sd_mod                 49152  2
ata_generic            16384  0
crc32c_intel           24576  6
psmouse               135168  0
vmxnet3                61440  0
ata_piix               36864  0
vmw_pvscsi             24576  1
i2c_piix4              24576  0
libata                249856  2 ata_piix,ata_generic
scsi_mod              225280  5 sd_mod,libata,sr_mod,sg,vmw_pvscsi
floppy                 69632  0
msch@rad-wgv-srv01:~$ 

Is there a way to show the memory consumed by each module? (besides the
'perf' tool you recommend below)

Would memory consumed by a module be released when the module is
unloaded? I guess so. Only I can't unload modules that are in use, of
course. Unloading vmw_balloon, vmw_vmci, and vmw_vsock_vmci_transport
didn't help.

> And that means - 'perf top', or better yet - 'perf record'.

I have never used perf before, will look into it.

Thanks a lot for your insight!

Martin

-- 
Martin Schwarz * Karlsruhe, Germany * http://kuroi.de/

[toc] | [prev] | [next] | [standalone]


#207543

FromReco <recoverym4n@enotuniq.net>
Date2019-04-16 11:00 +0200
Message-ID<xNs6B-73g-1@gated-at.bofh.it>
In reply to#207542
	Hi.

On Tue, Apr 16, 2019 at 10:22:16AM +0200, Martin Schwarz wrote:
> Hello,
> 
> On Mon, Apr 15, 2019 at 07:03:13PM +0300, Reco wrote:
> [...]
> > What I suspect is happening here is runaway memory allocation by a
> > kernel module (at least one of them), and said kernel module is likely
> > to be VMWare-specific.
> > It could be vmxnet3 (network). It could be that LSI kernel module or
> > whatever they're using for SCSI these days (vmw_pvscsi?).
> 
> sounds interesting.
> 
> That would explain why I haven't seen this problem on one of my (few)
> personal Stretch installations running as Xen DomU. But then, I guess
> we're not the only ones who use Debian Stretch on VMware ESXi ;-)
> but haven't found any mention of this problem. I wonder what makes our
> setup so special ...

I see nothing unusual short of somewhat low amount of RAM by today's
standards.
It's no excuse for the kernel to behave that way, for obvious reasons.


> Here's the outout from lsmod:

Nothing unexpected here.


> Is there a way to show the memory consumed by each module? (besides the
> 'perf' tool you recommend below)

slabtop from "procps" package, definitely.
Should've thought of it earlier.


> Would memory consumed by a module be released when the module is
> unloaded? I guess so.

Barring kernel memory leaks - yes. Yep, they "invented" them too.
A price to pay if you write a kernel in C.


> Only I can't unload modules that are in use, of course. Unloading
> vmw_balloon, vmw_vmci, and vmw_vsock_vmci_transport didn't help.

I doubt that these are the problem. Unless you're changing VM's RAM at
runtime (and you wrote you don't).
If you can unload it - it's not used, hence no kernel memory allocations
worthy of speaking.


> > And that means - 'perf top', or better yet - 'perf record'.
> 
> I have never used perf before, will look into it.

It's very indirect method. Basically it shows internal kernel functions
used, which may or may not be the source of the leak.

There's this SystemTap thing that can presumably do it better, but last
time I've checked it required a debug version of the kernel (and that's
unsuitable for just about anything short of kernel development).

Reco

[toc] | [prev] | [next] | [standalone]


#207550

FromMartin Schwarz <debian-lists@alias.kuroi.de>
Date2019-04-16 14:40 +0200
Message-ID<xNvxw-PD-9@gated-at.bofh.it>
In reply to#207543
Hello,

On Tue, Apr 16, 2019 at 11:49:57AM +0300, Reco wrote:
> I see nothing unusual short of somewhat low amount of RAM by today's
> standards.
> It's no excuse for the kernel to behave that way, for obvious reasons.

yes, 1 GB is not that much - but should be more than enough for the two
small RADIUSd instances (some 80 MB). Of course, if the workload
requires it, we can assign more RAM. In fact, that was what we first did
when the problem started to show, but this only lasted so long until the
additional RAM was consumed as well - it just takes a bit longer.

> slabtop from "procps" package, definitely.
> Should've thought of it earlier.

I did take a look at slaptop (and at /proc/slabinfo) before. But since the
values for "SReclaimable" and "SUnreclaim" from /proc/meminfo are not high
enough to explain the memory consumption, I didn't investigate any further in
that direction.

Looks unsuspicious to me:

root@rad-wgv-srv01:~# slabtop --once --sort=s | head -15
 Active / Total Objects (% used)    : 144152 / 186846 (77,2%)
 Active / Total Slabs (% used)      : 11205 / 11206 (100,0%)
 Active / Total Caches (% used)     : 66 / 116 (56,9%)
 Active / Total Size (% used)       : 40726,60K / 46301,70K (88,0%)
 Minimum / Average / Maximum Object : 0,02K / 0,25K / 4096,00K

  OBJS ACTIVE  USE OBJ SIZE  SLABS OBJ/SLAB CACHE SIZE NAME                   
     0      0   0% 4096,00K      0        1         0K kmalloc-4194304        
     0      0   0% 4096,00K      0        1         0K dma-kmalloc-4194304    
     0      0   0% 2048,00K      0        1         0K kmalloc-2097152        
     0      0   0% 2048,00K      0        1         0K dma-kmalloc-2097152    
     0      0   0% 1024,00K      0        1         0K kmalloc-1048576        
     0      0   0% 1024,00K      0        1         0K dma-kmalloc-1048576    
     0      0   0%  512,00K      0        1         0K kmalloc-524288         
     0      0   0%  512,00K      0        1         0K dma-kmalloc-524288     
root@rad-wgv-srv01:~# 

> > Only I can't unload modules that are in use, of course. Unloading
> > vmw_balloon, vmw_vmci, and vmw_vsock_vmci_transport didn't help.
> 
> I doubt that these are the problem. Unless you're changing VM's RAM at
> runtime (and you wrote you don't).

We sometimes do increase a VM's RAM while it is running. But most of the VMs
showing the problem still have their initial memory size from the template they
were deployed from.

['perf top' and 'perf record']
> It's very indirect method. Basically it shows internal kernel functions
> used, which may or may not be the source of the leak.

I'm afraid 'perf' is beyond my knowledge. In its default invocation, 'perf_4.9
top' seems to be more focussed on CPU usage?  And 'perf_4.9 record' is used to
profile a certain command? Tried with "sleep 30" as command, but not sure how
to interpret the recording.

So I'm really unsure how to use them to further drill into kernel or module
memory usage. Any hints?

> There's this SystemTap thing that can presumably do it better, but last
> time I've checked it required a debug version of the kernel (and that's
> unsuitable for just about anything short of kernel development).

By then I should probably go to debian-kernel? ;-)

Thanks again for all your help!

Martin

-- 
Martin Schwarz * Karlsruhe, Germany * http://kuroi.de/

[toc] | [prev] | [next] | [standalone]


#207551

FromReco <recoverym4n@enotuniq.net>
Date2019-04-16 15:10 +0200
Message-ID<xNw0y-1fl-5@gated-at.bofh.it>
In reply to#207550
	Hi.

On Tue, Apr 16, 2019 at 02:30:57PM +0200, Martin Schwarz wrote:
> > slabtop from "procps" package, definitely.
> > Should've thought of it earlier.
> 
> I did take a look at slaptop (and at /proc/slabinfo) before. But since the
> values for "SReclaimable" and "SUnreclaim" from /proc/meminfo are not high
> enough to explain the memory consumption, I didn't investigate any further in
> that direction.
> 
> Looks unsuspicious to me:

Agreed. Nothing here that explains the behaviour observed.


> ['perf top' and 'perf record']
> > It's very indirect method. Basically it shows internal kernel functions
> > used, which may or may not be the source of the leak.
> 
> I'm afraid 'perf' is beyond my knowledge. In its default invocation, 'perf_4.9
> top' seems to be more focussed on CPU usage?

Nope, it's deeper than this. What "perf top" should show by default is
the kernel and userspace library functions called by every kernel thread
and every userspace process. For instance,

  81.98%  [kernel]                  [k] load_balance
  18.02%  [kernel]                  [k] __tick_nohz_idle_enter
   0.00%  [kernel]                  [k] module_get_kallsym

This shows mostly idle OS, and the kernel calls load_balance() 81% of
the time, and __tick_nohz_idle_enter() 18% of the time. It eats CPU, of
course, but CPU consumption is not a concern here. The answer to the
question 'what the kernel does' is.


> And 'perf_4.9 record' is used to profile a certain command?

'perf record -a' in this case. All userspace and the kernel.
Terminate it with Ctrl+C after a minute or two.
Sorry if I didn't mention that.


> Tried with "sleep 30" as command, but not sure how
> to interpret the recording.

'perf report' from the same directory where it wrote "pref.data"
earlier.


> So I'm really unsure how to use them to further drill into kernel or module
> memory usage. Any hints?

Send it here, unless it's classified or somehow private. You'll need
those anyway, because...

> 
> > There's this SystemTap thing that can presumably do it better, but last
> > time I've checked it required a debug version of the kernel (and that's
> > unsuitable for just about anything short of kernel development).
> 
> By then I should probably go to debian-kernel? ;-)

I'd file a bug against linux-image package. You're using only in-tree
kernel modules shipped by Debian. The kernel is obviously misbehaving.
And if they won't answer you in a week (it's a small volunteer team
after all) - then you'll go at debian-kernel with a bug number at hands.

Reco

[toc] | [prev] | [next] | [standalone]


#207501

FromKenneth Parker <sea7kenp@gmail.com>
Date2019-04-15 16:50 +0200
Message-ID<xNb5M-4Vm-23@gated-at.bofh.it>
In reply to#207487

[Multipart message — attachments visible in raw view] — view raw

I had a Symptom like this a few years ago, which was tracked to something
called "zram", which tries to use "excess RAM" as Swap Space.

If so, it would show up on /proc/swaps

Verify that.

Best regards,

Kenneth Parker

[toc] | [prev] | [next] | [standalone]


#207503

FromMartin Schwarz <debian-lists@alias.kuroi.de>
Date2019-04-15 17:10 +0200
Message-ID<xNbp8-5hx-29@gated-at.bofh.it>
In reply to#207501
On Mon, Apr 15, 2019 at 10:44:26AM -0400, Kenneth Parker wrote:
> I had a Symptom like this a few years ago, which was tracked to something
> called "zram", which tries to use "excess RAM" as Swap Space.

thanks for your input!

We do not use zram. (I assume that would also show up in `lsmod`?)

/proc/swap shows only one "normal" swap partition (on a LVM logical
volume in this case):

------------------------------------------------------------------------
root@rad-wgv-srv01:~# cat /proc/swaps
Filename				Type		Size	Used	Priority
/dev/dm-3                               partition	2097148	70876	-1
root@rad-wgv-srv01:~# fgrep swap /etc/fstab 
/dev/mapper/vg0-lv_swap none            swap    sw              0       0
root@rad-wgv-srv01:~# ls -l /dev/mapper/vg0-lv_swap
lrwxrwxrwx 1 root root 7 Mär 29 08:33 /dev/mapper/vg0-lv_swap -> ../dm-3
root@rad-wgv-srv01:~# ls -l /dev/dm-3
brw-rw---- 1 root disk 254, 3 Mär 29 08:33 /dev/dm-3
root@rad-wgv-srv01:~# lvs /dev/mapper/vg0-lv_swap
  LV      VG  Attr       LSize Pool Origin Data%  Meta%  Move Log Cpy%Sync Convert
  lv_swap vg0 -wi-ao---- 2,00g                                                    
root@rad-wgv-srv01:~# lsmod | fgrep z
Module                  Size  Used by
root@rad-wgv-srv01:~# 
------------------------------------------------------------------------
-- 
Martin Schwarz * Karlsruhe, Germany * http://kuroi.de/

[toc] | [prev] | [next] | [standalone]


#207559 — Re: Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch)

FromPeter Wiersig <peter@friesenpeter.de>
Date2019-04-16 16:40 +0200
SubjectRe: Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch)
Message-ID<xNxpD-1Z5-9@gated-at.bofh.it>
In reply to#207487
Martin Schwarz <debian-lists@alias.kuroi.de> writes:
> root@rad-m2m-srv02:~# ps aux --sort=-rss | head -15

you're choosing the wrong sort field to debug your problem here:
man ps:
"""
       rss         RSS       resident set size, the non-swapped physical memory that a task has used (in kiloBytes).
                             (alias rssize, rsz).
...
       vsz         VSZ       virtual memory size of the process in KiB (1024-byte units).  Device mappings are currently
                             excluded; this is subject to change. (alias vsize)."""

try the latter field if the problem is repeating.  VSZ can be very
misleading with graphical processes, but that should not occur here.

https://stackoverflow.com/questions/7880784/what-is-rss-and-vsz-in-linux-memory-management

https://stackoverflow.com/a/21049737/2911961
""RSS is the Resident Set Size and is used to show how much memory is
allocated to that process and is in RAM. It does not include memory that
is swapped out. It does include memory from shared libraries as long as
the pages from those libraries are actually in memory. It does include
all stack and heap memory.

VSZ is the Virtual Memory Size. It includes all memory that the process
can access, including memory that is swapped out, memory that is
allocated, but not used, and memory that is from shared libraries.""

Peter

[toc] | [prev] | [next] | [standalone]


#207566

FromReco <recoverym4n@enotuniq.net>
Date2019-04-16 18:50 +0200
Message-ID<xNzrs-3bR-5@gated-at.bofh.it>
In reply to#207559
	Hi.

On Tue, Apr 16, 2019 at 04:39:32PM +0200, Peter Wiersig wrote:
> VSZ is the Virtual Memory Size. It includes all memory that the process
> can access, including memory that is swapped out, memory that is
> allocated, but not used, and memory that is from shared libraries.""

Given this:

== cut ==

#include <stdlib.h>
#include <unistd.h>

void main(void)
{
    char* a = malloc(4l*1024l*1024l*1024l);
    sleep(1200);
}

== cut ==

One can expect the process to "consume" 4Gb of VSZ, and about 700kb of
RSS.

Do you really suggest to treat all this VSZ as a "used memory"? A hint -
compile the sample, run it a hundred times.

Reco

[toc] | [prev] | [next] | [standalone]


#207599 — Re: Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch)

FromPeter Wiersig <peter@friesenpeter.de>
Date2019-04-17 01:20 +0200
SubjectRe: Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch)
Message-ID<xNFwR-76n-3@gated-at.bofh.it>
In reply to#207566
Reco <recoverym4n@enotuniq.net> writes:

Hi Reco,

> On Tue, Apr 16, 2019 at 04:39:32PM +0200, Peter Wiersig wrote:
>> VSZ is the Virtual Memory Size. (...),
>>including memory that is swapped out,
>> memory that is allocated, but not used,
  ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
>> and memory that is from shared libraries.""
>
> Given this:
>
> == cut ==
>
> One can expect the process to "consume" 4Gb of VSZ, and about 700kb of
> RSS.
>
> Do you really suggest to treat all this VSZ as a "used memory"? A hint -
> compile the sample, run it a hundred times.

No, but that is all in the explanation above albeit tersely. I
underlined the part which comes into play with your example code.

If Martin wants to find out what's using his memory during timespans
with degraded performance, it will not be helpful to list the processes
and sort by thenot swapped amount.  If the system is thrashing due to
swapping in and out, I always look at the VSZ values to find out where I
should direct my attention to.

If Martin or if you want, I can explode the the previous sentence about
what VSZ is but you'd probably find better answers if you read about
that.

Feel free to ask, I think I can explain more, but not better than
stackexchange answers or for example LWN kernel article series.

Have fun,
Peter

[toc] | [prev] | [next] | [standalone]


#207615

FromReco <recoverym4n@enotuniq.net>
Date2019-04-17 07:50 +0200
Message-ID<xNLCh-2oO-1@gated-at.bofh.it>
In reply to#207599
On Wed, Apr 17, 2019 at 01:14:36AM +0200, Peter Wiersig wrote:
> Reco <recoverym4n@enotuniq.net> writes:
> 
> Hi Reco,
> 
> > On Tue, Apr 16, 2019 at 04:39:32PM +0200, Peter Wiersig wrote:
> >> VSZ is the Virtual Memory Size. (...),
> >>including memory that is swapped out,
> >> memory that is allocated, but not used,
>   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
> >> and memory that is from shared libraries.""
> >
> > Given this:
> >
> > == cut ==
> >
> > One can expect the process to "consume" 4Gb of VSZ, and about 700kb of
> > RSS.
> >
> > Do you really suggest to treat all this VSZ as a "used memory"? A hint -
> > compile the sample, run it a hundred times.
> 
> No, but that is all in the explanation above albeit tersely. I
> underlined the part which comes into play with your example code.

That was the whole point. If it's allocated, but ain't used - one should
not account it.


> If Martin wants to find out what's using his memory during timespans
> with degraded performance, it will not be helpful to list the processes
> and sort by thenot swapped amount.

To quote the original e-mail:

root@rad-m2m-srv02:~# free -thwl
              total        used        free      shared     buffers cache   available
Mem:           987M        910M         59M          0B        704K 16M         13M
Low:           987M        927M         59M
High:            0B          0B          0B
Swap:          2,0G        345M        1,7G
Total:         3,0G        1,2G        1,7G

root@rad-m2m-srv02:~# smem -uktr
User     Count     Swap      USS      PSS      RSS
root        39   332.8M    10.4M    12.4M    44.7M
msch         6     7.0M        0   607.0K     8.3M
_chrony      1   360.0K     4.0K    20.0K   572.0K
messagebus     1   580.0K     4.0K    17.0K   480.0K
postfix      2     1.6M        0    13.0K   568.0K
daemon       1   208.0K     4.0K     6.0K    72.0K
---------------------------------------------------
            50   342.5M    10.4M    13.0M    54.7M

A host has 1Gb of RAM and 3Gb swap.
Both "free" and "/proc/meminfo" show that there's no free memory left.
"smem" output shows us that this memory ain't used by userspace.
The question was - where all this memory has gone?


> If the system is thrashing due to swapping in and out, I always look
> at the VSZ values to find out where I should direct my attention to.

Thanks for sharing. In this particular case, swapping is a consequence,
not a root cause.
But, just in case, how exactly your method is superior to "top -o SWAP"?

Reco

[toc] | [prev] | [next] | [standalone]


#207632

FromMartin Schwarz <debian-lists@alias.kuroi.de>
Date2019-04-17 11:40 +0200
Message-ID<xNPcS-4DL-11@gated-at.bofh.it>
In reply to#207559
On Tue, Apr 16, 2019 at 04:39:32PM +0200, Peter Wiersig wrote:
>        rss         RSS       resident set size, the non-swapped physical memory that a task has used (in kiloBytes).
>                              (alias rssize, rsz).
> ...
>        vsz         VSZ       virtual memory size of the process in KiB (1024-byte units).  Device mappings are currently
>                              excluded; this is subject to change. (alias vsize)."""
> 

Thanks for pointing out the difference. However, in this case I'm trying
to find out what consumes RAM specifically, not virtual memory in
general. Besides, from all I can see, the high memory usage is NOT
caused user space processes. So the output from ps was just to show that
even though some hundred MB of RAM (!) are used, just some few MB are
consumed by processes.

Kind regards
Martin

-- 
Martin Schwarz * Karlsruhe, Germany * http://kuroi.de/

[toc] | [prev] | [next] | [standalone]


#207601 — Re: Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch)

FromPeter Wiersig <peter@friesenpeter.de>
Date2019-04-17 01:40 +0200
SubjectRe: Need help analyzing (kernel?) memory usage and reclaiming RAM (Debian Stretch)
Message-ID<xNFQe-7d3-13@gated-at.bofh.it>
In reply to#207487
Martin Schwarz <debian-lists@alias.kuroi.de> writes:
>
> Here's the output from some commands I hope to be helpful:
>
> The machine in this example is a RADIUS server but has not even gone
> productive ... no incoming client requests yet.  (But the problem is not
> related to the RADIUS server software - OSC Radiator - since the same
> symptoms show on different machines: not only RADIUS servers but also
> nameservers, shell servers or jumphosts, etc.)
>
> [values while the problem persists:]
...
> USER       PID %CPU %MEM    VSZ   RSS TTY      STAT START   TIME COMMAND
> root     34718 12.0  0.5  29596  5672 ?        D    09:01   0:00 /usr/bin/python3 -Es /usr/bin/lsb_release --short --description
> root     26491  3.1  0.2  79328  2860 ?        D    08:04   1:50 apt-get update -qq
> root     32551  6.8  0.2 119036  2800 ?        D    08:51   0:43 /usr/bin/python3 /usr/bin/unattended-upgrade

Disable this, do your upgrades by some schedule for the duration in
which you're debugging this problem.  Think about system orchestration
tools with push mechanisms if you want to minimize RAM allocated to
VMs.  We're thinking about deploying ansible for patch management.

> root     12792  2.2  0.1 159720  1748 ?        D    06:06   3:54 /usr/bin/perl -w /usr/bin/apt-show-versions -i
> root     15502  2.4  0.1 167660  1608 ?        D    06:25   3:51 /usr/bin/perl -w /usr/bin/apt-show-versions -i

Do they need to run on 6:06 and then parallel at 6:25?  What's their
process tree calling structure, ie. what's starting them?

> root     34527  1.7  0.1  14096  1596 ?        Ss   09:01   0:00 /bin/bash /usr/bin/check_mk_agent

Can you show a zoomed image of the memory graph prior to a problem?  And
a load graph of the same duration?

I had some webservers which were also prone to death spiraling, the only
real solution was to throw RAM at them until they were able to process
the requests and to optimize the database indices to speed up the time
spent fetching and sorting rows.

Peter

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.user


csiph-web