Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.kernel > #1461492 > unrolled thread

[PACTH v2 0/3] Implement /proc/<pid>/totmaps

Started byrobert.foss@collabora.com
First post2016-08-13 00:10 +0200
Last post2016-08-19 07:20 +0200
Articles 18 on this page of 38 — 8 participants

Back to article view | Back to linux.kernel


Contents

  [PACTH v2 0/3] Implement /proc/<pid>/totmaps robert.foss@collabora.com - 2016-08-13 00:10 +0200
    [PACTH v2 3/3] Documentation/filesystems: Added /proc/PID/totmaps documentation robert.foss@collabora.com - 2016-08-13 00:10 +0200
    Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-14 11:10 +0200
      Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Robert Foss <robert.foss@collabora.com> - 2016-08-15 15:10 +0200
        Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-15 15:50 +0200
          Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Robert Foss <robert.foss@collabora.com> - 2016-08-15 18:30 +0200
            Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-16 09:20 +0200
              Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Robert Foss <robert.foss@collabora.com> - 2016-08-16 18:50 +0200
                Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-17 10:30 +0200
                  Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Jann Horn <jann@thejh.net> - 2016-08-17 11:40 +0200
                    Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-17 15:10 +0200
                      Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Robert Foss <robert.foss@collabora.com> - 2016-08-17 18:50 +0200
                      Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Sonny Rao <sonnyrao@chromium.org> - 2016-08-17 21:10 +0200
                        Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-18 09:50 +0200
                          Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-19 03:10 +0200
                            Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Robert Foss <robert.foss@collabora.com> - 2016-08-19 04:00 +0200
                              Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Sonny Rao <sonnyrao@chromium.org> - 2016-08-19 08:30 +0200
                            Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Minchan Kim <minchan@kernel.org> - 2016-08-19 04:30 +0200
                              Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Sonny Rao <sonnyrao@chromium.org> - 2016-08-19 08:50 +0200
                              Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-19 11:10 +0200
                                Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Sonny Rao <sonnyrao@chromium.org> - 2016-08-19 20:30 +0200
                                Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Minchan Kim <minchan@kernel.org> - 2016-08-22 02:10 +0200
                                  Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-22 09:50 +0200
                                    Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Minchan Kim <minchan@kernel.org> - 2016-08-22 16:20 +0200
                                      Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Robert Foss <robert.foss@collabora.com> - 2016-08-22 16:40 +0200
                                      Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-22 18:50 +0200
                                        Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-22 19:30 +0200
                                          Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-22 19:50 +0200
                                            Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-23 10:30 +0200
                                              utime accounting regression since 4.6 (was: Re: [PACTH v2 0/3]  Implement /proc/<pid>/totmaps) Michal Hocko <mhocko@kernel.org> - 2016-08-23 16:40 +0200
                                                Re: utime accounting regression since 4.6 (was: Re: [PACTH v2 0/3]  Implement /proc/<pid>/totmaps) Rik van Riel <riel@redhat.com> - 2016-08-23 23:50 +0200
                            Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Sonny Rao <sonnyrao@chromium.org> - 2016-08-19 08:50 +0200
                              Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-19 10:00 +0200
                                Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Sonny Rao <sonnyrao@chromium.org> - 2016-08-19 20:00 +0200
                                  Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Michal Hocko <mhocko@kernel.org> - 2016-08-22 10:00 +0200
                                    Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Sonny Rao <sonnyrao@chromium.org> - 2016-08-23 00:50 +0200
                                      Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Marcin Jabrzyk <m.jabrzyk@samsung.com> - 2016-08-24 12:20 +0200
                          Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps Sonny Rao <sonnyrao@chromium.org> - 2016-08-19 07:20 +0200

Page 2 of 2 — ← Prev page 1 [2]


#1466606

FromSonny Rao <sonnyrao@chromium.org>
Date2016-08-19 20:30 +0200
Message-ID<s7WIi-Ez-23@gated-at.bofh.it>
In reply to#1466245
On Fri, Aug 19, 2016 at 1:05 AM, Michal Hocko <mhocko@kernel.org> wrote:
> On Fri 19-08-16 11:26:34, Minchan Kim wrote:
>> Hi Michal,
>>
>> On Thu, Aug 18, 2016 at 08:01:04PM +0200, Michal Hocko wrote:
>> > On Thu 18-08-16 10:47:57, Sonny Rao wrote:
>> > > On Thu, Aug 18, 2016 at 12:44 AM, Michal Hocko <mhocko@kernel.org> wrote:
>> > > > On Wed 17-08-16 11:57:56, Sonny Rao wrote:
>> > [...]
>> > > >> 2) User space OOM handling -- we'd rather do a more graceful shutdown
>> > > >> than let the kernel's OOM killer activate and need to gather this
>> > > >> information and we'd like to be able to get this information to make
>> > > >> the decision much faster than 400ms
>> > > >
>> > > > Global OOM handling in userspace is really dubious if you ask me. I
>> > > > understand you want something better than SIGKILL and in fact this is
>> > > > already possible with memory cgroup controller (btw. memcg will give
>> > > > you a cheap access to rss, amount of shared, swapped out memory as
>> > > > well). Anyway if you are getting close to the OOM your system will most
>> > > > probably be really busy and chances are that also reading your new file
>> > > > will take much more time. I am also not quite sure how is pss useful for
>> > > > oom decisions.
>> > >
>> > > I mentioned it before, but based on experience RSS just isn't good
>> > > enough -- there's too much sharing going on in our use case to make
>> > > the correct decision based on RSS.  If RSS were good enough, simply
>> > > put, this patch wouldn't exist.
>> >
>> > But that doesn't answer my question, I am afraid. So how exactly do you
>> > use pss for oom decisions?
>>
>> My case is not for OOM decision but I agree it would be great if we can get
>> *fast* smap summary information.
>>
>> PSS is really great tool to figure out how processes consume memory
>> more exactly rather than RSS. We have been used it for monitoring
>> of memory for per-process. Although it is not used for OOM decision,
>> it would be great if it is speed up because we don't want to spend
>> many CPU time for just monitoring.
>>
>> For our usecase, we don't need AnonHugePages, ShmemPmdMapped, Shared_Hugetlb,
>> Private_Hugetlb, KernelPageSize, MMUPageSize because we never enable THP and
>> hugetlb. Additionally, Locked can be known via vma flags so we don't need it,
>> either. Even, we don't need address range for just monitoring when we don't
>> investigate in detail.
>>
>> Although they are not severe overhead, why does it emit the useless
>> information? Even bloat day by day. :( With that, userspace tools should
>> spend more time to parse which is pointless.
>
> So far it doesn't really seem that the parsing is the biggest problem.
> The major cycles killer is the output formatting and that doesn't sound
> like a problem we are not able to address. And I would even argue that
> we want to address it in a generic way as much as possible.
>
>> Having said that, I'm not fan of creating new stat knob for that, either.
>> How about appending summary information in the end of smap?
>> So, monitoring users can just open the file and lseek to the (end - 1) and
>> read the summary only.
>
> That might confuse existing parsers. Besides that we already have
> /proc/<pid>/statm which gives cumulative numbers already. I am not sure
> how often it is used and whether the pte walk is too expensive for
> existing users but that should be explored and evaluated before a new
> file is created.
>
> The /proc became a dump of everything people found interesting just
> because we were to easy to allow those additions. Do not repeat those
> mistakes, please!

Another thing I noticed was that we lock down smaps on Chromium OS.  I
think this is to avoid exposing more information than necessary via
proc.  The totmaps file gives us just the information we need and
nothing else.   I certainly don't think we need a proc file for this
use case -- do you think a new system call is better or something
else?

> --
> Michal Hocko
> SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1467289

FromMinchan Kim <minchan@kernel.org>
Date2016-08-22 02:10 +0200
Message-ID<s8KYp-7hH-3@gated-at.bofh.it>
In reply to#1466245
On Fri, Aug 19, 2016 at 10:05:32AM +0200, Michal Hocko wrote:
> On Fri 19-08-16 11:26:34, Minchan Kim wrote:
> > Hi Michal,
> > 
> > On Thu, Aug 18, 2016 at 08:01:04PM +0200, Michal Hocko wrote:
> > > On Thu 18-08-16 10:47:57, Sonny Rao wrote:
> > > > On Thu, Aug 18, 2016 at 12:44 AM, Michal Hocko <mhocko@kernel.org> wrote:
> > > > > On Wed 17-08-16 11:57:56, Sonny Rao wrote:
> > > [...]
> > > > >> 2) User space OOM handling -- we'd rather do a more graceful shutdown
> > > > >> than let the kernel's OOM killer activate and need to gather this
> > > > >> information and we'd like to be able to get this information to make
> > > > >> the decision much faster than 400ms
> > > > >
> > > > > Global OOM handling in userspace is really dubious if you ask me. I
> > > > > understand you want something better than SIGKILL and in fact this is
> > > > > already possible with memory cgroup controller (btw. memcg will give
> > > > > you a cheap access to rss, amount of shared, swapped out memory as
> > > > > well). Anyway if you are getting close to the OOM your system will most
> > > > > probably be really busy and chances are that also reading your new file
> > > > > will take much more time. I am also not quite sure how is pss useful for
> > > > > oom decisions.
> > > > 
> > > > I mentioned it before, but based on experience RSS just isn't good
> > > > enough -- there's too much sharing going on in our use case to make
> > > > the correct decision based on RSS.  If RSS were good enough, simply
> > > > put, this patch wouldn't exist.
> > > 
> > > But that doesn't answer my question, I am afraid. So how exactly do you
> > > use pss for oom decisions?
> > 
> > My case is not for OOM decision but I agree it would be great if we can get
> > *fast* smap summary information.
> > 
> > PSS is really great tool to figure out how processes consume memory
> > more exactly rather than RSS. We have been used it for monitoring
> > of memory for per-process. Although it is not used for OOM decision,
> > it would be great if it is speed up because we don't want to spend
> > many CPU time for just monitoring.
> > 
> > For our usecase, we don't need AnonHugePages, ShmemPmdMapped, Shared_Hugetlb,
> > Private_Hugetlb, KernelPageSize, MMUPageSize because we never enable THP and
> > hugetlb. Additionally, Locked can be known via vma flags so we don't need it,
> > either. Even, we don't need address range for just monitoring when we don't
> > investigate in detail.
> > 
> > Although they are not severe overhead, why does it emit the useless
> > information? Even bloat day by day. :( With that, userspace tools should
> > spend more time to parse which is pointless.
> 
> So far it doesn't really seem that the parsing is the biggest problem.
> The major cycles killer is the output formatting and that doesn't sound

I cannot understand how kernel space is more expensive.
Hmm. I tested your test program on my machine.


#!/bin/sh
./smap_test &
pid=$!

for i in $(seq 25)
do
        cat /proc/$pid/smaps > /dev/null
done
kill $pid

root@bbox:/home/barrios/test/smap# time ./s_v.sh
pid:21925
real    0m3.365s
user    0m0.031s
sys     0m3.046s


vs.

#!/bin/sh
./smap_test &
pid=$!

for i in $(seq 25)
do
        awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {}' \
         /proc/$pid/smaps
done
kill $pid

root@bbox:/home/barrios/test/smap# time ./s.sh 
pid:21973

real    0m17.812s
user    0m12.612s
sys     0m5.187s

perf report says

    39.56%  awk        gawk               [.] dfaexec                             
     7.61%  awk        [kernel.kallsyms]  [k] format_decode                       
     6.37%  awk        gawk               [.] avoid_dfa                           
     5.85%  awk        gawk               [.] interpret                           
     5.69%  awk        [kernel.kallsyms]  [k] __memcpy                            
     4.37%  awk        [kernel.kallsyms]  [k] vsnprintf                           
     2.69%  awk        [kernel.kallsyms]  [k] number.isra.13                      
     2.10%  awk        gawk               [.] research                            
     1.91%  awk        gawk               [.] 0x00000000000351d0                  
     1.49%  awk        gawk               [.] free_wstr                           
     1.27%  awk        gawk               [.] unref                               
     1.19%  awk        gawk               [.] reset_record                        
     0.95%  awk        gawk               [.] set_record                          
     0.95%  awk        gawk               [.] get_field                           
     0.94%  awk        [kernel.kallsyms]  [k] show_smap                           

Parsing is much expensive than kernel.
Could you retest your test program?

> like a problem we are not able to address. And I would even argue that
> we want to address it in a generic way as much as possible.

Sure. What solution do you think as generic way?

> 
> > Having said that, I'm not fan of creating new stat knob for that, either.
> > How about appending summary information in the end of smap?
> > So, monitoring users can just open the file and lseek to the (end - 1) and
> > read the summary only.
> 
> That might confuse existing parsers. Besides that we already have
> /proc/<pid>/statm which gives cumulative numbers already. I am not sure
> how often it is used and whether the pte walk is too expensive for
> existing users but that should be explored and evaluated before a new
> file is created.
> 
> The /proc became a dump of everything people found interesting just
> because we were to easy to allow those additions. Do not repeat those
> mistakes, please!
> -- 
> Michal Hocko
> SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1467428

FromMichal Hocko <mhocko@kernel.org>
Date2016-08-22 09:50 +0200
Message-ID<s8S9z-3nT-11@gated-at.bofh.it>
In reply to#1467289
On Mon 22-08-16 09:07:45, Minchan Kim wrote:
[...]
> #!/bin/sh
> ./smap_test &
> pid=$!
> 
> for i in $(seq 25)
> do
>         awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {}' \
>          /proc/$pid/smaps
> done
> kill $pid
> 
> root@bbox:/home/barrios/test/smap# time ./s.sh 
> pid:21973
> 
> real    0m17.812s
> user    0m12.612s
> sys     0m5.187s

retested on the bare metal (x86_64 - 2CPUs)
        Command being timed: "sh s.sh"
        User time (seconds): 0.00
        System time (seconds): 18.08
        Percent of CPU this job got: 98%
        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:18.29

multiple runs are quite consistent in those numbers. I am running with
$ awk --version
GNU Awk 4.1.3, API: 1.1 (GNU MPFR 3.1.4, GNU MP 6.1.0)

> > like a problem we are not able to address. And I would even argue that
> > we want to address it in a generic way as much as possible.
> 
> Sure. What solution do you think as generic way?

either optimize seq_printf or replace it with something faster.

-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1467665

FromMinchan Kim <minchan@kernel.org>
Date2016-08-22 16:20 +0200
Message-ID<s8YeZ-7mP-17@gated-at.bofh.it>
In reply to#1467428

[Multipart message — attachments visible in raw view] — view raw

On Mon, Aug 22, 2016 at 09:40:52AM +0200, Michal Hocko wrote:
> On Mon 22-08-16 09:07:45, Minchan Kim wrote:
> [...]
> > #!/bin/sh
> > ./smap_test &
> > pid=$!
> > 
> > for i in $(seq 25)
> > do
> >         awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {}' \
> >          /proc/$pid/smaps
> > done
> > kill $pid
> > 
> > root@bbox:/home/barrios/test/smap# time ./s.sh 
> > pid:21973
> > 
> > real    0m17.812s
> > user    0m12.612s
> > sys     0m5.187s
> 
> retested on the bare metal (x86_64 - 2CPUs)
>         Command being timed: "sh s.sh"
>         User time (seconds): 0.00
>         System time (seconds): 18.08
>         Percent of CPU this job got: 98%
>         Elapsed (wall clock) time (h:mm:ss or m:ss): 0:18.29
> 
> multiple runs are quite consistent in those numbers. I am running with
> $ awk --version
> GNU Awk 4.1.3, API: 1.1 (GNU MPFR 3.1.4, GNU MP 6.1.0)
> 
> > > like a problem we are not able to address. And I would even argue that
> > > we want to address it in a generic way as much as possible.
> > 
> > Sure. What solution do you think as generic way?
> 
> either optimize seq_printf or replace it with something faster.

If it's real culprit, I agree. However, I tested your test program on
my 2 x86 machines and my friend's machine.

Ubuntu, Fedora, Arch

They have awk 4.0.1 and 4.1.3.

Result are same. Userspace speand more times I mentioned.

[root@blaptop smap_test]# time awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {printf "rss:%d pss:%d\n", rss, pss}' /proc/3552/smaps
rss:263484 pss:262188

real    0m0.770s
user    0m0.574s
sys     0m0.197s

I will attach my test progrma source.
I hope you guys test and repost the result because it's the key for direction
of patchset.

Thanks.

[toc] | [prev] | [next] | [standalone]


#1467676

FromRobert Foss <robert.foss@collabora.com>
Date2016-08-22 16:40 +0200
Message-ID<s8Yyl-7uQ-1@gated-at.bofh.it>
In reply to#1467665

On 2016-08-22 10:12 AM, Minchan Kim wrote:
> On Mon, Aug 22, 2016 at 09:40:52AM +0200, Michal Hocko wrote:
>> On Mon 22-08-16 09:07:45, Minchan Kim wrote:
>> [...]
>>> #!/bin/sh
>>> ./smap_test &
>>> pid=$!
>>>
>>> for i in $(seq 25)
>>> do
>>>         awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {}' \
>>>          /proc/$pid/smaps
>>> done
>>> kill $pid
>>>
>>> root@bbox:/home/barrios/test/smap# time ./s.sh
>>> pid:21973
>>>
>>> real    0m17.812s
>>> user    0m12.612s
>>> sys     0m5.187s
>>
>> retested on the bare metal (x86_64 - 2CPUs)
>>         Command being timed: "sh s.sh"
>>         User time (seconds): 0.00
>>         System time (seconds): 18.08
>>         Percent of CPU this job got: 98%
>>         Elapsed (wall clock) time (h:mm:ss or m:ss): 0:18.29
>>
>> multiple runs are quite consistent in those numbers. I am running with
>> $ awk --version
>> GNU Awk 4.1.3, API: 1.1 (GNU MPFR 3.1.4, GNU MP 6.1.0)
>>

$ ./smap_test &
pid:19658 nr_vma:65514

$ time awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {printf "rss:%d 
pss:%d\n", rss, pss}' /proc/19658/smaps
rss:263452 pss:262151

real	0m0.625s
user	0m0.404s
sys	0m0.216s

$ awk --version
GNU Awk 4.1.3, API: 1.1 (GNU MPFR 3.1.4, GNU MP 6.1.0)

>>>> like a problem we are not able to address. And I would even argue that
>>>> we want to address it in a generic way as much as possible.
>>>
>>> Sure. What solution do you think as generic way?
>>
>> either optimize seq_printf or replace it with something faster.
>
> If it's real culprit, I agree. However, I tested your test program on
> my 2 x86 machines and my friend's machine.
>
> Ubuntu, Fedora, Arch
>
> They have awk 4.0.1 and 4.1.3.
>
> Result are same. Userspace speand more times I mentioned.
>
> [root@blaptop smap_test]# time awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {printf "rss:%d pss:%d\n", rss, pss}' /proc/3552/smaps
> rss:263484 pss:262188
>
> real    0m0.770s
> user    0m0.574s
> sys     0m0.197s
>
> I will attach my test progrma source.
> I hope you guys test and repost the result because it's the key for direction
> of patchset.
>
> Thanks.
>

[toc] | [prev] | [next] | [standalone]


#1467806

FromMichal Hocko <mhocko@kernel.org>
Date2016-08-22 18:50 +0200
Message-ID<s90Aa-iZ-43@gated-at.bofh.it>
In reply to#1467665
On Mon 22-08-16 23:12:41, Minchan Kim wrote:
> On Mon, Aug 22, 2016 at 09:40:52AM +0200, Michal Hocko wrote:
> > On Mon 22-08-16 09:07:45, Minchan Kim wrote:
> > [...]
> > > #!/bin/sh
> > > ./smap_test &
> > > pid=$!
> > > 
> > > for i in $(seq 25)
> > > do
> > >         awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {}' \
> > >          /proc/$pid/smaps
> > > done
> > > kill $pid
> > > 
> > > root@bbox:/home/barrios/test/smap# time ./s.sh 
> > > pid:21973
> > > 
> > > real    0m17.812s
> > > user    0m12.612s
> > > sys     0m5.187s
> > 
> > retested on the bare metal (x86_64 - 2CPUs)
> >         Command being timed: "sh s.sh"
> >         User time (seconds): 0.00
> >         System time (seconds): 18.08
> >         Percent of CPU this job got: 98%
> >         Elapsed (wall clock) time (h:mm:ss or m:ss): 0:18.29
> > 
> > multiple runs are quite consistent in those numbers. I am running with
> > $ awk --version
> > GNU Awk 4.1.3, API: 1.1 (GNU MPFR 3.1.4, GNU MP 6.1.0)
> > 
> > > > like a problem we are not able to address. And I would even argue that
> > > > we want to address it in a generic way as much as possible.
> > > 
> > > Sure. What solution do you think as generic way?
> > 
> > either optimize seq_printf or replace it with something faster.
> 
> If it's real culprit, I agree. However, I tested your test program on
> my 2 x86 machines and my friend's machine.
> 
> Ubuntu, Fedora, Arch
> 
> They have awk 4.0.1 and 4.1.3.
> 
> Result are same. Userspace speand more times I mentioned.
> 
> [root@blaptop smap_test]# time awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {printf "rss:%d pss:%d\n", rss, pss}' /proc/3552/smaps
> rss:263484 pss:262188
> 
> real    0m0.770s
> user    0m0.574s
> sys     0m0.197s
> 
> I will attach my test progrma source.
> I hope you guys test and repost the result because it's the key for direction
> of patchset.

Hmm, this is really interesting. I have checked a different machine and
it shows different results. Same code, slightly different version of awk
(4.1.0) and the results are different
        Command being timed: "awk /^Rss/{rss+=$2} /^Pss/{pss+=$2} END {printf "rss:%d pss:%d\n", rss, pss} /proc/48925/smaps"
        User time (seconds): 0.43
        System time (seconds): 0.27

I have no idea why those numbers are so different on my laptop
yet. It surely looks suspicious. I will try to debug this further
tomorrow. Anyway, the performance is just one side of the problem. I
have tried to express my concerns about a single exported pss value in
other email. Please try to step back and think about how useful is this
information without the knowing which resource we are talking about.
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1467845

FromMichal Hocko <mhocko@kernel.org>
Date2016-08-22 19:30 +0200
Message-ID<s91cS-Nb-25@gated-at.bofh.it>
In reply to#1467806
On Mon 22-08-16 18:45:54, Michal Hocko wrote:
[...]
> I have no idea why those numbers are so different on my laptop
> yet. It surely looks suspicious. I will try to debug this further
> tomorrow.

Hmm, so I've tried to use my version of awk on other machine and vice
versa and it didn't make any difference. So this is independent on the
awk version it seems. So I've tried to strace /usr/bin/time and
wait4(-1, [{WIFEXITED(s) && WEXITSTATUS(s) == 0}], 0, {ru_utime={0, 0}, ru_stime={0, 688438}, ...}) = 9128

so the kernel indeed reports 0 user time for some reason. Note I
was testing with 4.7 and right now with 4.8.0-rc3 kernel (no local
modifications). The other machine which reports non-0 utime is 3.12
SLES kernel. Maybe I am hitting some accounting bug. At first I was
suspecting CONFIG_NO_HZ_FULL because that is the main difference between
my and the other machine but then I've noticed that the tests I was
doing in kvm have this disabled too.. so it must be something else.

Weird...
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1467858

FromMichal Hocko <mhocko@kernel.org>
Date2016-08-22 19:50 +0200
Message-ID<s91we-Tx-27@gated-at.bofh.it>
In reply to#1467845
On Mon 22-08-16 19:29:36, Michal Hocko wrote:
> On Mon 22-08-16 18:45:54, Michal Hocko wrote:
> [...]
> > I have no idea why those numbers are so different on my laptop
> > yet. It surely looks suspicious. I will try to debug this further
> > tomorrow.
> 
> Hmm, so I've tried to use my version of awk on other machine and vice
> versa and it didn't make any difference. So this is independent on the
> awk version it seems. So I've tried to strace /usr/bin/time and
> wait4(-1, [{WIFEXITED(s) && WEXITSTATUS(s) == 0}], 0, {ru_utime={0, 0}, ru_stime={0, 688438}, ...}) = 9128
> 
> so the kernel indeed reports 0 user time for some reason. Note I
> was testing with 4.7 and right now with 4.8.0-rc3 kernel (no local
> modifications). The other machine which reports non-0 utime is 3.12
> SLES kernel. Maybe I am hitting some accounting bug. At first I was
> suspecting CONFIG_NO_HZ_FULL because that is the main difference between
> my and the other machine but then I've noticed that the tests I was
> doing in kvm have this disabled too.. so it must be something else.

4.5 reports non-0 while 4.6 zero utime. NO_HZ configuration is the same
in both kernels.
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1468377

FromMichal Hocko <mhocko@kernel.org>
Date2016-08-23 10:30 +0200
Message-ID<s9ffQ-1oj-15@gated-at.bofh.it>
In reply to#1467858
On Mon 22-08-16 19:47:09, Michal Hocko wrote:
> On Mon 22-08-16 19:29:36, Michal Hocko wrote:
> > On Mon 22-08-16 18:45:54, Michal Hocko wrote:
> > [...]
> > > I have no idea why those numbers are so different on my laptop
> > > yet. It surely looks suspicious. I will try to debug this further
> > > tomorrow.
> > 
> > Hmm, so I've tried to use my version of awk on other machine and vice
> > versa and it didn't make any difference. So this is independent on the
> > awk version it seems. So I've tried to strace /usr/bin/time and
> > wait4(-1, [{WIFEXITED(s) && WEXITSTATUS(s) == 0}], 0, {ru_utime={0, 0}, ru_stime={0, 688438}, ...}) = 9128
> > 
> > so the kernel indeed reports 0 user time for some reason. Note I
> > was testing with 4.7 and right now with 4.8.0-rc3 kernel (no local
> > modifications). The other machine which reports non-0 utime is 3.12
> > SLES kernel. Maybe I am hitting some accounting bug. At first I was
> > suspecting CONFIG_NO_HZ_FULL because that is the main difference between
> > my and the other machine but then I've noticed that the tests I was
> > doing in kvm have this disabled too.. so it must be something else.
> 
> 4.5 reports non-0 while 4.6 zero utime. NO_HZ configuration is the same
> in both kernels.

and one more thing. It is not like utime accounting would be completely
broken and always report 0. Other commands report non-0 values even on
4.6 kernels. I will try to bisect this down later today.
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1468599 — utime accounting regression since 4.6 (was: Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps)

FromMichal Hocko <mhocko@kernel.org>
Date2016-08-23 16:40 +0200
Subjectutime accounting regression since 4.6 (was: Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps)
Message-ID<s9l1U-5eM-17@gated-at.bofh.it>
In reply to#1468377

[Multipart message — attachments visible in raw view] — view raw

On Tue 23-08-16 10:26:03, Michal Hocko wrote:
> On Mon 22-08-16 19:47:09, Michal Hocko wrote:
> > On Mon 22-08-16 19:29:36, Michal Hocko wrote:
> > > On Mon 22-08-16 18:45:54, Michal Hocko wrote:
> > > [...]
> > > > I have no idea why those numbers are so different on my laptop
> > > > yet. It surely looks suspicious. I will try to debug this further
> > > > tomorrow.
> > > 
> > > Hmm, so I've tried to use my version of awk on other machine and vice
> > > versa and it didn't make any difference. So this is independent on the
> > > awk version it seems. So I've tried to strace /usr/bin/time and
> > > wait4(-1, [{WIFEXITED(s) && WEXITSTATUS(s) == 0}], 0, {ru_utime={0, 0}, ru_stime={0, 688438}, ...}) = 9128
> > > 
> > > so the kernel indeed reports 0 user time for some reason. Note I
> > > was testing with 4.7 and right now with 4.8.0-rc3 kernel (no local
> > > modifications). The other machine which reports non-0 utime is 3.12
> > > SLES kernel. Maybe I am hitting some accounting bug. At first I was
> > > suspecting CONFIG_NO_HZ_FULL because that is the main difference between
> > > my and the other machine but then I've noticed that the tests I was
> > > doing in kvm have this disabled too.. so it must be something else.
> > 
> > 4.5 reports non-0 while 4.6 zero utime. NO_HZ configuration is the same
> > in both kernels.
> 
> and one more thing. It is not like utime accounting would be completely
> broken and always report 0. Other commands report non-0 values even on
> 4.6 kernels. I will try to bisect this down later today.

OK, so it seems I found it. I was quite lucky because account_user_time
is not all that popular function and there were basically no changes
besides Riks ff9a9b4c4334 ("sched, time: Switch VIRT_CPU_ACCOUNTING_GEN
to jiffy granularity") and that seems to cause the regression. Reverting
the commit on top of the current mmotm seems to fix the issue for me.

And just to give Rik more context. While debugging overhead of the
/proc/<pid>/smaps I am getting a misleading output from /usr/bin/time -v
(source for ./max_mmap is [1])

root@test1:~# uname -r
4.5.0-rc6-bisect1-00025-gff9a9b4c4334
root@test1:~# ./max_map 
pid:2990 maps:65515
root@test1:~# /usr/bin/time -v awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {printf "rss:%d pss:%d\n", rss, pss}' /proc/2990/smaps
rss:263368 pss:262203
        Command being timed: "awk /^Rss/{rss+=$2} /^Pss/{pss+=$2} END {printf "rss:%d pss:%d\n", rss, pss} /proc/2990/smaps"
        User time (seconds): 0.00
        System time (seconds): 0.45
        Percent of CPU this job got: 98%
        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.46
        Average shared text size (kbytes): 0
        Average unshared data size (kbytes): 0
        Average stack size (kbytes): 0
        Average total size (kbytes): 0
        Maximum resident set size (kbytes): 1796
        Average resident set size (kbytes): 0
        Major (requiring I/O) page faults: 1
        Minor (reclaiming a frame) page faults: 83
        Voluntary context switches: 6
        Involuntary context switches: 6
        Swaps: 0
        File system inputs: 248
        File system outputs: 0
        Socket messages sent: 0
        Socket messages received: 0
        Signals delivered: 0
        Page size (bytes): 4096
        Exit status: 0

See the User time being 0 (as you can see above in the quoted text it
is not a rounding error in userspace or something similar because wait4
really returns 0). Now with the revert
root@test1:~# uname -r
4.5.0-rc6-revert-00026-g7fc86f968bf5
root@test1:~# /usr/bin/time -v awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {printf "rss:%d pss:%d\n", rss, pss}' /proc/3015/smaps
rss:263316 pss:262199
        Command being timed: "awk /^Rss/{rss+=$2} /^Pss/{pss+=$2} END {printf "rss:%d pss:%d\n", rss, pss} /proc/3015/smaps"
        User time (seconds): 0.18
        System time (seconds): 0.29
        Percent of CPU this job got: 97%
        Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.50
        Average shared text size (kbytes): 0
        Average unshared data size (kbytes): 0
        Average stack size (kbytes): 0
        Average total size (kbytes): 0
        Maximum resident set size (kbytes): 1760
        Average resident set size (kbytes): 0
        Major (requiring I/O) page faults: 1
        Minor (reclaiming a frame) page faults: 79
        Voluntary context switches: 5
        Involuntary context switches: 7
        Swaps: 0
        File system inputs: 248
        File system outputs: 0
        Socket messages sent: 0
        Socket messages received: 0
        Signals delivered: 0
        Page size (bytes): 4096
        Exit status: 0

So it looks like the whole user time is accounted as the system time.
My config is attached and yes I do have CONFIG_VIRT_CPU_ACCOUNTING_GEN
enabled. Could you have a look please?

[1] http://lkml.kernel.org/r/20160817082200.GA10547@dhcp22.suse.cz
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1468903 — Re: utime accounting regression since 4.6 (was: Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps)

FromRik van Riel <riel@redhat.com>
Date2016-08-23 23:50 +0200
SubjectRe: utime accounting regression since 4.6 (was: Re: [PACTH v2 0/3] Implement /proc/<pid>/totmaps)
Message-ID<s9rK2-1fM-43@gated-at.bofh.it>
In reply to#1468599
On Tue, 2016-08-23 at 16:33 +0200, Michal Hocko wrote:
> On Tue 23-08-16 10:26:03, Michal Hocko wrote:
> > On Mon 22-08-16 19:47:09, Michal Hocko wrote:
> > > On Mon 22-08-16 19:29:36, Michal Hocko wrote:
> > > > On Mon 22-08-16 18:45:54, Michal Hocko wrote:
> > > > [...]
> > > > > I have no idea why those numbers are so different on my
> > > > > laptop
> > > > > yet. It surely looks suspicious. I will try to debug this
> > > > > further
> > > > > tomorrow.
> > > > 
> > > > Hmm, so I've tried to use my version of awk on other machine
> > > > and vice
> > > > versa and it didn't make any difference. So this is independent
> > > > on the
> > > > awk version it seems. So I've tried to strace /usr/bin/time and
> > > > wait4(-1, [{WIFEXITED(s) && WEXITSTATUS(s) == 0}], 0,
> > > > {ru_utime={0, 0}, ru_stime={0, 688438}, ...}) = 9128
> > > > 
> > > > so the kernel indeed reports 0 user time for some reason. Note
> > > > I
> > > > was testing with 4.7 and right now with 4.8.0-rc3 kernel (no
> > > > local
> > > > modifications). The other machine which reports non-0 utime is
> > > > 3.12
> > > > SLES kernel. Maybe I am hitting some accounting bug. At first I
> > > > was
> > > > suspecting CONFIG_NO_HZ_FULL because that is the main
> > > > difference between
> > > > my and the other machine but then I've noticed that the tests I
> > > > was
> > > > doing in kvm have this disabled too.. so it must be something
> > > > else.
> > > 
> > > 4.5 reports non-0 while 4.6 zero utime. NO_HZ configuration is
> > > the same
> > > in both kernels.
> > 
> > and one more thing. It is not like utime accounting would be
> > completely
> > broken and always report 0. Other commands report non-0 values even
> > on
> > 4.6 kernels. I will try to bisect this down later today.
> 
> OK, so it seems I found it. I was quite lucky because
> account_user_time
> is not all that popular function and there were basically no changes
> besides Riks ff9a9b4c4334 ("sched, time: Switch
> VIRT_CPU_ACCOUNTING_GEN
> to jiffy granularity") and that seems to cause the regression.
> Reverting
> the commit on top of the current mmotm seems to fix the issue for me.
> 
> And just to give Rik more context. While debugging overhead of the
> /proc/<pid>/smaps I am getting a misleading output from /usr/bin/time
> -v
> (source for ./max_mmap is [1])
> 
> root@test1:~# uname -r
> 4.5.0-rc6-bisect1-00025-gff9a9b4c4334
> root@test1:~# ./max_map 
> pid:2990 maps:65515
> root@test1:~# /usr/bin/time -v awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2}
> END {printf "rss:%d pss:%d\n", rss, pss}' /proc/2990/smaps
> rss:263368 pss:262203
>         Command being timed: "awk /^Rss/{rss+=$2} /^Pss/{pss+=$2} END
> {printf "rss:%d pss:%d\n", rss, pss} /proc/2990/smaps"
>         User time (seconds): 0.00
>         System time (seconds): 0.45
>         Percent of CPU this job got: 98%
> 

> root@test1:~# /usr/bin/time -v awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2}
> END {printf "rss:%d pss:%d\n", rss, pss}' /proc/3015/smaps
> rss:263316 pss:262199
>         Command being timed: "awk /^Rss/{rss+=$2} /^Pss/{pss+=$2} END
> {printf "rss:%d pss:%d\n", rss, pss} /proc/3015/smaps"
>         User time (seconds): 0.18
>         System time (seconds): 0.29
>         Percent of CPU this job got: 97%

The patch in question makes user and system
time accounting essentially tick-based. If
jiffies changes while the task is in user
mode, time gets accounted as user time, if
jiffies changes while the task is in system
mode, time gets accounted as system time.

If you get "unlucky", with a job like the
above, it is possible all time gets accounted
to system time.

This would be true both with the system running
with a periodic timer tick (before and after my
patch is applied), and in nohz_idle mode (after
my patch).

However, it does seem quite unlikely that you
get zero user time, since you have 125 timer
ticks in half a second. Furthermore, you do not
even have NO_HZ_FULL enabled...

Does the workload consistently get zero user
time?

If so, we need to dig further to see under
what precise circumstances that happens.

On my laptop, with kernel 4.6.3-300.fc24.x86_64
I get this:

$ /usr/bin/time -v awk '/^Rss/{rss+=$2} /^Pss/{pss+=$2} END {printf
"rss:%d pss:%d\n", rss, pss}' /proc/19825/smaps
rss:263368 pss:262145
	Command being timed: "awk /^Rss/{rss+=$2} /^Pss/{pss+=$2} END
{printf "rss:%d pss:%d\n", rss, pss} /proc/19825/smaps"
	User time (seconds): 0.64
	System time (seconds): 0.19
	Percent of CPU this job got: 99%
	Elapsed (wall clock) time (h:mm:ss or m:ss): 0:00.83

The main difference between your and my
NO_HZ config seems to be that NO_HZ_FULL
is set here. However, it is not enabled
at run time, so both of our systems
should only really get NO_HZ_IDLE
effectively.

Running tasks should get sampled with the
regular timer tick, while they are running.

In other words, vtime accounting should be
disabled in both of our tests, for everything
except the idle task.

Do I need to do anything special to reproduce
your bug, besides running the max mmap program
and the awk script?

[toc] | [prev] | [next] | [standalone]


#1466064

FromSonny Rao <sonnyrao@chromium.org>
Date2016-08-19 08:50 +0200
Message-ID<s7LMR-27O-25@gated-at.bofh.it>
In reply to#1465723
On Thu, Aug 18, 2016 at 11:01 AM, Michal Hocko <mhocko@kernel.org> wrote:
> On Thu 18-08-16 10:47:57, Sonny Rao wrote:
>> On Thu, Aug 18, 2016 at 12:44 AM, Michal Hocko <mhocko@kernel.org> wrote:
>> > On Wed 17-08-16 11:57:56, Sonny Rao wrote:
> [...]
>> >> 2) User space OOM handling -- we'd rather do a more graceful shutdown
>> >> than let the kernel's OOM killer activate and need to gather this
>> >> information and we'd like to be able to get this information to make
>> >> the decision much faster than 400ms
>> >
>> > Global OOM handling in userspace is really dubious if you ask me. I
>> > understand you want something better than SIGKILL and in fact this is
>> > already possible with memory cgroup controller (btw. memcg will give
>> > you a cheap access to rss, amount of shared, swapped out memory as
>> > well). Anyway if you are getting close to the OOM your system will most
>> > probably be really busy and chances are that also reading your new file
>> > will take much more time. I am also not quite sure how is pss useful for
>> > oom decisions.
>>
>> I mentioned it before, but based on experience RSS just isn't good
>> enough -- there's too much sharing going on in our use case to make
>> the correct decision based on RSS.  If RSS were good enough, simply
>> put, this patch wouldn't exist.
>
> But that doesn't answer my question, I am afraid. So how exactly do you
> use pss for oom decisions?

We use PSS to calculate the memory used by a process among all the
processes in the system, in the case of Chrome this tells us how much
each renderer process (which is roughly tied to a particular "tab" in
Chrome) is using and how much it has swapped out, so we know what the
worst offenders are -- I'm not sure what's unclear about that?

Chrome tends to use a lot of shared memory so we found PSS to be
better than RSS, and I can give you examples of the  RSS and PSS on
real systems to illustrate the magnitude of the difference between
those two numbers if that would be useful.

>
>> So even with memcg I think we'd have the same problem?
>
> memcg will give you instant anon, shared counters for all processes in
> the memcg.
>

We want to be able to get per-process granularity quickly.  I'm not
sure if memcg provides that exactly?

>> > Don't take me wrong, /proc/<pid>/totmaps might be suitable for your
>> > specific usecase but so far I haven't heard any sound argument for it to
>> > be generally usable. It is true that smaps is unnecessarily costly but
>> > at least I can see some room for improvements. A simple patch I've
>> > posted cut the formatting overhead by 7%. Maybe we can do more.
>>
>> It seems like a general problem that if you want these values the
>> existing kernel interface can be very expensive, so it would be
>> generally usable by any application which wants a per process PSS,
>> private data, dirty data or swap value.
>
> yes this is really unfortunate. And if at all possible we should address
> that. Precise values require the expensive rmap walk. We can introduce
> some caching to help that. But so far it seems the biggest overhead is
> to simply format the output and that should be addressed before any new
> proc file is added.
>
>> I mentioned two use cases, but I guess I don't understand the comment
>> about why it's not usable by other use cases.
>
> I might be wrong here but a use of pss is quite limited and I do not
> remember anybody asking for large optimizations in that area. I still do
> not understand your use cases properly so I am quite skeptical about a
> general usefulness of a new file.

How do you know that usage of PSS is quite limited?  I can only say
that we've been using it on Chromium OS for at least four years and
have found it very valuable, and I think I've explained the use cases
in this thread. If you have more specific questions then I can try to
clarify.

>
> --
> Michal Hocko
> SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1466208

FromMichal Hocko <mhocko@kernel.org>
Date2016-08-19 10:00 +0200
Message-ID<s7MSC-2Nt-33@gated-at.bofh.it>
In reply to#1466064
On Thu 18-08-16 23:43:39, Sonny Rao wrote:
> On Thu, Aug 18, 2016 at 11:01 AM, Michal Hocko <mhocko@kernel.org> wrote:
> > On Thu 18-08-16 10:47:57, Sonny Rao wrote:
> >> On Thu, Aug 18, 2016 at 12:44 AM, Michal Hocko <mhocko@kernel.org> wrote:
> >> > On Wed 17-08-16 11:57:56, Sonny Rao wrote:
> > [...]
> >> >> 2) User space OOM handling -- we'd rather do a more graceful shutdown
> >> >> than let the kernel's OOM killer activate and need to gather this
> >> >> information and we'd like to be able to get this information to make
> >> >> the decision much faster than 400ms
> >> >
> >> > Global OOM handling in userspace is really dubious if you ask me. I
> >> > understand you want something better than SIGKILL and in fact this is
> >> > already possible with memory cgroup controller (btw. memcg will give
> >> > you a cheap access to rss, amount of shared, swapped out memory as
> >> > well). Anyway if you are getting close to the OOM your system will most
> >> > probably be really busy and chances are that also reading your new file
> >> > will take much more time. I am also not quite sure how is pss useful for
> >> > oom decisions.
> >>
> >> I mentioned it before, but based on experience RSS just isn't good
> >> enough -- there's too much sharing going on in our use case to make
> >> the correct decision based on RSS.  If RSS were good enough, simply
> >> put, this patch wouldn't exist.
> >
> > But that doesn't answer my question, I am afraid. So how exactly do you
> > use pss for oom decisions?
> 
> We use PSS to calculate the memory used by a process among all the
> processes in the system, in the case of Chrome this tells us how much
> each renderer process (which is roughly tied to a particular "tab" in
> Chrome) is using and how much it has swapped out, so we know what the
> worst offenders are -- I'm not sure what's unclear about that?

So let me ask more specifically. How can you make any decision based on
the pss when you do not know _what_ is the shared resource. In other
words if you select a task to terminate based on the pss then you have to
kill others who share the same resource otherwise you do not release
that shared resource. Not to mention that such a shared resource might
be on tmpfs/shmem and it won't get released even after all processes
which map it are gone.

I am sorry for being dense but it is still not clear to me how the
single pss number can be used for oom or, in general, any serious
decisions. The counter might be useful of course for debugging purposes
or to have a general overview but then arguing about 40 vs 20ms sounds a
bit strange to me.

> Chrome tends to use a lot of shared memory so we found PSS to be
> better than RSS, and I can give you examples of the  RSS and PSS on
> real systems to illustrate the magnitude of the difference between
> those two numbers if that would be useful.
> 
> >
> >> So even with memcg I think we'd have the same problem?
> >
> > memcg will give you instant anon, shared counters for all processes in
> > the memcg.
> >
> 
> We want to be able to get per-process granularity quickly.  I'm not
> sure if memcg provides that exactly?

I will give you that information if you do process-per-memcg but that
doesn't sound ideal. I thought those 20-something processes you were
talking about are treated together but it seems I misunderstood.
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1466590

FromSonny Rao <sonnyrao@chromium.org>
Date2016-08-19 20:00 +0200
Message-ID<s7Wff-e9-15@gated-at.bofh.it>
In reply to#1466208
On Fri, Aug 19, 2016 at 12:59 AM, Michal Hocko <mhocko@kernel.org> wrote:
> On Thu 18-08-16 23:43:39, Sonny Rao wrote:
>> On Thu, Aug 18, 2016 at 11:01 AM, Michal Hocko <mhocko@kernel.org> wrote:
>> > On Thu 18-08-16 10:47:57, Sonny Rao wrote:
>> >> On Thu, Aug 18, 2016 at 12:44 AM, Michal Hocko <mhocko@kernel.org> wrote:
>> >> > On Wed 17-08-16 11:57:56, Sonny Rao wrote:
>> > [...]
>> >> >> 2) User space OOM handling -- we'd rather do a more graceful shutdown
>> >> >> than let the kernel's OOM killer activate and need to gather this
>> >> >> information and we'd like to be able to get this information to make
>> >> >> the decision much faster than 400ms
>> >> >
>> >> > Global OOM handling in userspace is really dubious if you ask me. I
>> >> > understand you want something better than SIGKILL and in fact this is
>> >> > already possible with memory cgroup controller (btw. memcg will give
>> >> > you a cheap access to rss, amount of shared, swapped out memory as
>> >> > well). Anyway if you are getting close to the OOM your system will most
>> >> > probably be really busy and chances are that also reading your new file
>> >> > will take much more time. I am also not quite sure how is pss useful for
>> >> > oom decisions.
>> >>
>> >> I mentioned it before, but based on experience RSS just isn't good
>> >> enough -- there's too much sharing going on in our use case to make
>> >> the correct decision based on RSS.  If RSS were good enough, simply
>> >> put, this patch wouldn't exist.
>> >
>> > But that doesn't answer my question, I am afraid. So how exactly do you
>> > use pss for oom decisions?
>>
>> We use PSS to calculate the memory used by a process among all the
>> processes in the system, in the case of Chrome this tells us how much
>> each renderer process (which is roughly tied to a particular "tab" in
>> Chrome) is using and how much it has swapped out, so we know what the
>> worst offenders are -- I'm not sure what's unclear about that?
>
> So let me ask more specifically. How can you make any decision based on
> the pss when you do not know _what_ is the shared resource. In other
> words if you select a task to terminate based on the pss then you have to
> kill others who share the same resource otherwise you do not release
> that shared resource. Not to mention that such a shared resource might
> be on tmpfs/shmem and it won't get released even after all processes
> which map it are gone.

Ok I see why you're confused now, sorry.

In our case that we do know what is being shared in general because
the sharing is mostly between those processes that we're looking at
and not other random processes or tmpfs, so PSS gives us useful data
in the context of these processes which are sharing the data
especially for monitoring between the set of these renderer processes.

We also use the private clean and private dirty and swap fields to
make a few metrics for the processes and charge each process for it's
private, shared, and swap data. Private clean and dirty are used for
estimating a lower bound on how much memory would be freed.  Swap and
PSS also give us some indication of additional memory which might get
freed up.

>
> I am sorry for being dense but it is still not clear to me how the
> single pss number can be used for oom or, in general, any serious
> decisions. The counter might be useful of course for debugging purposes
> or to have a general overview but then arguing about 40 vs 20ms sounds a
> bit strange to me.

Yeah so it's more than just the single PSS number, it's PSS,
Private_Clean, Private_dirty, Swap are all interesting numbers to make
these decisions.

>
>> Chrome tends to use a lot of shared memory so we found PSS to be
>> better than RSS, and I can give you examples of the  RSS and PSS on
>> real systems to illustrate the magnitude of the difference between
>> those two numbers if that would be useful.
>>
>> >
>> >> So even with memcg I think we'd have the same problem?
>> >
>> > memcg will give you instant anon, shared counters for all processes in
>> > the memcg.
>> >
>>
>> We want to be able to get per-process granularity quickly.  I'm not
>> sure if memcg provides that exactly?
>
> I will give you that information if you do process-per-memcg but that
> doesn't sound ideal. I thought those 20-something processes you were
> talking about are treated together but it seems I misunderstood.
> --
> Michal Hocko
> SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1467436

FromMichal Hocko <mhocko@kernel.org>
Date2016-08-22 10:00 +0200
Message-ID<s8Sjg-3ra-29@gated-at.bofh.it>
In reply to#1466590
On Fri 19-08-16 10:57:48, Sonny Rao wrote:
> On Fri, Aug 19, 2016 at 12:59 AM, Michal Hocko <mhocko@kernel.org> wrote:
> > On Thu 18-08-16 23:43:39, Sonny Rao wrote:
> >> On Thu, Aug 18, 2016 at 11:01 AM, Michal Hocko <mhocko@kernel.org> wrote:
> >> > On Thu 18-08-16 10:47:57, Sonny Rao wrote:
> >> >> On Thu, Aug 18, 2016 at 12:44 AM, Michal Hocko <mhocko@kernel.org> wrote:
> >> >> > On Wed 17-08-16 11:57:56, Sonny Rao wrote:
> >> > [...]
> >> >> >> 2) User space OOM handling -- we'd rather do a more graceful shutdown
> >> >> >> than let the kernel's OOM killer activate and need to gather this
> >> >> >> information and we'd like to be able to get this information to make
> >> >> >> the decision much faster than 400ms
> >> >> >
> >> >> > Global OOM handling in userspace is really dubious if you ask me. I
> >> >> > understand you want something better than SIGKILL and in fact this is
> >> >> > already possible with memory cgroup controller (btw. memcg will give
> >> >> > you a cheap access to rss, amount of shared, swapped out memory as
> >> >> > well). Anyway if you are getting close to the OOM your system will most
> >> >> > probably be really busy and chances are that also reading your new file
> >> >> > will take much more time. I am also not quite sure how is pss useful for
> >> >> > oom decisions.
> >> >>
> >> >> I mentioned it before, but based on experience RSS just isn't good
> >> >> enough -- there's too much sharing going on in our use case to make
> >> >> the correct decision based on RSS.  If RSS were good enough, simply
> >> >> put, this patch wouldn't exist.
> >> >
> >> > But that doesn't answer my question, I am afraid. So how exactly do you
> >> > use pss for oom decisions?
> >>
> >> We use PSS to calculate the memory used by a process among all the
> >> processes in the system, in the case of Chrome this tells us how much
> >> each renderer process (which is roughly tied to a particular "tab" in
> >> Chrome) is using and how much it has swapped out, so we know what the
> >> worst offenders are -- I'm not sure what's unclear about that?
> >
> > So let me ask more specifically. How can you make any decision based on
> > the pss when you do not know _what_ is the shared resource. In other
> > words if you select a task to terminate based on the pss then you have to
> > kill others who share the same resource otherwise you do not release
> > that shared resource. Not to mention that such a shared resource might
> > be on tmpfs/shmem and it won't get released even after all processes
> > which map it are gone.
> 
> Ok I see why you're confused now, sorry.
> 
> In our case that we do know what is being shared in general because
> the sharing is mostly between those processes that we're looking at
> and not other random processes or tmpfs, so PSS gives us useful data
> in the context of these processes which are sharing the data
> especially for monitoring between the set of these renderer processes.

OK, I see and agree that pss might be useful when you _know_ what is
shared. But this sounds quite specific to a particular workload. How
many users are in a similar situation? In other words, if we present
a single number without the context, how much useful it will be in
general? Is it possible that presenting such a number could be even
misleading for somebody who doesn't have an idea which resources are
shared? These are all questions which should be answered before we
actually add this number (be it a new/existing proc file or a syscall).
I still believe that the number without wider context is just not all
that useful.

> We also use the private clean and private dirty and swap fields to
> make a few metrics for the processes and charge each process for it's
> private, shared, and swap data. Private clean and dirty are used for
> estimating a lower bound on how much memory would be freed.

I can imagine that this kind of information might be useful and
presented in /proc/<pid>/statm. The question is whether some of the
existing consumers would see the performance impact due to he page table
walk. Anyway even these counters might get quite tricky because even
shareable resources are considered private if the process is the only
one to map them (so again this might be a file on tmpfs...).

> Swap and
> PSS also give us some indication of additional memory which might get
> freed up.
-- 
Michal Hocko
SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1468115

FromSonny Rao <sonnyrao@chromium.org>
Date2016-08-23 00:50 +0200
Message-ID<s96cy-3WR-13@gated-at.bofh.it>
In reply to#1467436
On Mon, Aug 22, 2016 at 12:54 AM, Michal Hocko <mhocko@kernel.org> wrote:
> On Fri 19-08-16 10:57:48, Sonny Rao wrote:
>> On Fri, Aug 19, 2016 at 12:59 AM, Michal Hocko <mhocko@kernel.org> wrote:
>> > On Thu 18-08-16 23:43:39, Sonny Rao wrote:
>> >> On Thu, Aug 18, 2016 at 11:01 AM, Michal Hocko <mhocko@kernel.org> wrote:
>> >> > On Thu 18-08-16 10:47:57, Sonny Rao wrote:
>> >> >> On Thu, Aug 18, 2016 at 12:44 AM, Michal Hocko <mhocko@kernel.org> wrote:
>> >> >> > On Wed 17-08-16 11:57:56, Sonny Rao wrote:
>> >> > [...]
>> >> >> >> 2) User space OOM handling -- we'd rather do a more graceful shutdown
>> >> >> >> than let the kernel's OOM killer activate and need to gather this
>> >> >> >> information and we'd like to be able to get this information to make
>> >> >> >> the decision much faster than 400ms
>> >> >> >
>> >> >> > Global OOM handling in userspace is really dubious if you ask me. I
>> >> >> > understand you want something better than SIGKILL and in fact this is
>> >> >> > already possible with memory cgroup controller (btw. memcg will give
>> >> >> > you a cheap access to rss, amount of shared, swapped out memory as
>> >> >> > well). Anyway if you are getting close to the OOM your system will most
>> >> >> > probably be really busy and chances are that also reading your new file
>> >> >> > will take much more time. I am also not quite sure how is pss useful for
>> >> >> > oom decisions.
>> >> >>
>> >> >> I mentioned it before, but based on experience RSS just isn't good
>> >> >> enough -- there's too much sharing going on in our use case to make
>> >> >> the correct decision based on RSS.  If RSS were good enough, simply
>> >> >> put, this patch wouldn't exist.
>> >> >
>> >> > But that doesn't answer my question, I am afraid. So how exactly do you
>> >> > use pss for oom decisions?
>> >>
>> >> We use PSS to calculate the memory used by a process among all the
>> >> processes in the system, in the case of Chrome this tells us how much
>> >> each renderer process (which is roughly tied to a particular "tab" in
>> >> Chrome) is using and how much it has swapped out, so we know what the
>> >> worst offenders are -- I'm not sure what's unclear about that?
>> >
>> > So let me ask more specifically. How can you make any decision based on
>> > the pss when you do not know _what_ is the shared resource. In other
>> > words if you select a task to terminate based on the pss then you have to
>> > kill others who share the same resource otherwise you do not release
>> > that shared resource. Not to mention that such a shared resource might
>> > be on tmpfs/shmem and it won't get released even after all processes
>> > which map it are gone.
>>
>> Ok I see why you're confused now, sorry.
>>
>> In our case that we do know what is being shared in general because
>> the sharing is mostly between those processes that we're looking at
>> and not other random processes or tmpfs, so PSS gives us useful data
>> in the context of these processes which are sharing the data
>> especially for monitoring between the set of these renderer processes.
>
> OK, I see and agree that pss might be useful when you _know_ what is
> shared. But this sounds quite specific to a particular workload. How
> many users are in a similar situation? In other words, if we present
> a single number without the context, how much useful it will be in
> general? Is it possible that presenting such a number could be even
> misleading for somebody who doesn't have an idea which resources are
> shared? These are all questions which should be answered before we
> actually add this number (be it a new/existing proc file or a syscall).
> I still believe that the number without wider context is just not all
> that useful.


I see the specific point about  PSS -- because you need to know what
is being shared or otherwise use it in a whole system context, but I
still think the whole system context is a valid and generally useful
thing.  But what about the private_clean and private_dirty?  Surely
those are more generally useful for calculating a lower bound on
process memory usage without additional knowledge?

At the end of the day all of these metrics are approximations, and it
comes down to how far off the various approximations are and what
trade offs we are willing to make.
RSS is the cheapest but the most coarse.

PSS (with the correct context) and Private data plus swap are much
better but also more expensive due to the PT walk.
As far as I know, to get anything but RSS we have to go through smaps
or use memcg.  Swap seems to be available in /proc/<pid>/status.

I looked at the "shared" value in /proc/<pid>/statm but it doesn't
seem to correlate well with the shared value in smaps -- not sure why?

It might be useful to show the magnitude of difference of using RSS vs
PSS/Private in the case of the Chrome renderer processes.  On the
system I was looking at there were about 40 of these processes, but I
picked a few to give an idea:

localhost ~ # cat /proc/21550/totmaps
Rss:               98972 kB
Pss:               54717 kB
Shared_Clean:      19020 kB
Shared_Dirty:      26352 kB
Private_Clean:         0 kB
Private_Dirty:     53600 kB
Referenced:        92184 kB
Anonymous:         46524 kB
AnonHugePages:     24576 kB
Swap:              13148 kB


RSS is 80% higher than PSS and 84% higher than private data

localhost ~ # cat /proc/21470/totmaps
Rss:              118420 kB
Pss:               70938 kB
Shared_Clean:      22212 kB
Shared_Dirty:      26520 kB
Private_Clean:         0 kB
Private_Dirty:     69688 kB
Referenced:       111500 kB
Anonymous:         79928 kB
AnonHugePages:     24576 kB
Swap:              12964 kB

RSS is 66% higher than RSS and 69% higher than private data

localhost ~ # cat /proc/21435/totmaps
Rss:               97156 kB
Pss:               50044 kB
Shared_Clean:      21920 kB
Shared_Dirty:      26400 kB
Private_Clean:         0 kB
Private_Dirty:     48836 kB
Referenced:        90012 kB
Anonymous:         75228 kB
AnonHugePages:     24576 kB
Swap:              13064 kB

RSS is 94% higher than PSS and 98% higher than private data.

It looks like there's a set of about 40MB of shared pages which cause
the difference in this case.
Swap was roughly even on these but I don't think it's always going to be true.


>
>> We also use the private clean and private dirty and swap fields to
>> make a few metrics for the processes and charge each process for it's
>> private, shared, and swap data. Private clean and dirty are used for
>> estimating a lower bound on how much memory would be freed.
>
> I can imagine that this kind of information might be useful and
> presented in /proc/<pid>/statm. The question is whether some of the
> existing consumers would see the performance impact due to he page table
> walk. Anyway even these counters might get quite tricky because even
> shareable resources are considered private if the process is the only
> one to map them (so again this might be a file on tmpfs...).
>
>> Swap and
>> PSS also give us some indication of additional memory which might get
>> freed up.
> --
> Michal Hocko
> SUSE Labs

[toc] | [prev] | [next] | [standalone]


#1469310

FromMarcin Jabrzyk <m.jabrzyk@samsung.com>
Date2016-08-24 12:20 +0200
Message-ID<s9DrQ-QH-15@gated-at.bofh.it>
In reply to#1468115

On 23/08/16 00:44, Sonny Rao wrote:
> On Mon, Aug 22, 2016 at 12:54 AM, Michal Hocko <mhocko@kernel.org> wrote:
>> On Fri 19-08-16 10:57:48, Sonny Rao wrote:
>>> On Fri, Aug 19, 2016 at 12:59 AM, Michal Hocko <mhocko@kernel.org> wrote:
>>>> On Thu 18-08-16 23:43:39, Sonny Rao wrote:
>>>>> On Thu, Aug 18, 2016 at 11:01 AM, Michal Hocko <mhocko@kernel.org> wrote:
>>>>>> On Thu 18-08-16 10:47:57, Sonny Rao wrote:
>>>>>>> On Thu, Aug 18, 2016 at 12:44 AM, Michal Hocko <mhocko@kernel.org> wrote:
>>>>>>>> On Wed 17-08-16 11:57:56, Sonny Rao wrote:
>>>>>> [...]
>>>>>>>>> 2) User space OOM handling -- we'd rather do a more graceful shutdown
>>>>>>>>> than let the kernel's OOM killer activate and need to gather this
>>>>>>>>> information and we'd like to be able to get this information to make
>>>>>>>>> the decision much faster than 400ms
>>>>>>>>
>>>>>>>> Global OOM handling in userspace is really dubious if you ask me. I
>>>>>>>> understand you want something better than SIGKILL and in fact this is
>>>>>>>> already possible with memory cgroup controller (btw. memcg will give
>>>>>>>> you a cheap access to rss, amount of shared, swapped out memory as
>>>>>>>> well). Anyway if you are getting close to the OOM your system will most
>>>>>>>> probably be really busy and chances are that also reading your new file
>>>>>>>> will take much more time. I am also not quite sure how is pss useful for
>>>>>>>> oom decisions.
>>>>>>>
>>>>>>> I mentioned it before, but based on experience RSS just isn't good
>>>>>>> enough -- there's too much sharing going on in our use case to make
>>>>>>> the correct decision based on RSS.  If RSS were good enough, simply
>>>>>>> put, this patch wouldn't exist.
>>>>>>
>>>>>> But that doesn't answer my question, I am afraid. So how exactly do you
>>>>>> use pss for oom decisions?
>>>>>
>>>>> We use PSS to calculate the memory used by a process among all the
>>>>> processes in the system, in the case of Chrome this tells us how much
>>>>> each renderer process (which is roughly tied to a particular "tab" in
>>>>> Chrome) is using and how much it has swapped out, so we know what the
>>>>> worst offenders are -- I'm not sure what's unclear about that?
>>>>
>>>> So let me ask more specifically. How can you make any decision based on
>>>> the pss when you do not know _what_ is the shared resource. In other
>>>> words if you select a task to terminate based on the pss then you have to
>>>> kill others who share the same resource otherwise you do not release
>>>> that shared resource. Not to mention that such a shared resource might
>>>> be on tmpfs/shmem and it won't get released even after all processes
>>>> which map it are gone.
>>>
>>> Ok I see why you're confused now, sorry.
>>>
>>> In our case that we do know what is being shared in general because
>>> the sharing is mostly between those processes that we're looking at
>>> and not other random processes or tmpfs, so PSS gives us useful data
>>> in the context of these processes which are sharing the data
>>> especially for monitoring between the set of these renderer processes.
>>
>> OK, I see and agree that pss might be useful when you _know_ what is
>> shared. But this sounds quite specific to a particular workload. How
>> many users are in a similar situation? In other words, if we present
>> a single number without the context, how much useful it will be in
>> general? Is it possible that presenting such a number could be even
>> misleading for somebody who doesn't have an idea which resources are
>> shared? These are all questions which should be answered before we
>> actually add this number (be it a new/existing proc file or a syscall).
>> I still believe that the number without wider context is just not all
>> that useful.
>
>
> I see the specific point about  PSS -- because you need to know what
> is being shared or otherwise use it in a whole system context, but I
> still think the whole system context is a valid and generally useful
> thing.  But what about the private_clean and private_dirty?  Surely
> those are more generally useful for calculating a lower bound on
> process memory usage without additional knowledge?
>
> At the end of the day all of these metrics are approximations, and it
> comes down to how far off the various approximations are and what
> trade offs we are willing to make.
> RSS is the cheapest but the most coarse.
>
> PSS (with the correct context) and Private data plus swap are much
> better but also more expensive due to the PT walk.
> As far as I know, to get anything but RSS we have to go through smaps
> or use memcg.  Swap seems to be available in /proc/<pid>/status.
>
> I looked at the "shared" value in /proc/<pid>/statm but it doesn't
> seem to correlate well with the shared value in smaps -- not sure why?
>
> It might be useful to show the magnitude of difference of using RSS vs
> PSS/Private in the case of the Chrome renderer processes.  On the
> system I was looking at there were about 40 of these processes, but I
> picked a few to give an idea:
>
> localhost ~ # cat /proc/21550/totmaps
> Rss:               98972 kB
> Pss:               54717 kB
> Shared_Clean:      19020 kB
> Shared_Dirty:      26352 kB
> Private_Clean:         0 kB
> Private_Dirty:     53600 kB
> Referenced:        92184 kB
> Anonymous:         46524 kB
> AnonHugePages:     24576 kB
> Swap:              13148 kB
>
>
> RSS is 80% higher than PSS and 84% higher than private data
>
> localhost ~ # cat /proc/21470/totmaps
> Rss:              118420 kB
> Pss:               70938 kB
> Shared_Clean:      22212 kB
> Shared_Dirty:      26520 kB
> Private_Clean:         0 kB
> Private_Dirty:     69688 kB
> Referenced:       111500 kB
> Anonymous:         79928 kB
> AnonHugePages:     24576 kB
> Swap:              12964 kB
>
> RSS is 66% higher than RSS and 69% higher than private data
>
> localhost ~ # cat /proc/21435/totmaps
> Rss:               97156 kB
> Pss:               50044 kB
> Shared_Clean:      21920 kB
> Shared_Dirty:      26400 kB
> Private_Clean:         0 kB
> Private_Dirty:     48836 kB
> Referenced:        90012 kB
> Anonymous:         75228 kB
> AnonHugePages:     24576 kB
> Swap:              13064 kB
>
> RSS is 94% higher than PSS and 98% higher than private data.
>
> It looks like there's a set of about 40MB of shared pages which cause
> the difference in this case.
> Swap was roughly even on these but I don't think it's always going to be true.
>
>

Sorry to hijack the thread, but I've found it recently
and I guess it's the best place to present our point.
We are working at our custom OS based on Linux and we also suffered much
by /proc/<pid>/smaps file. As in Chrome we tried to improve our internal
application memory management polices (Low Memory Killer) using data
provided by smaps but we failed due to very long time needed for reading
and parsing properly the file.

We've also observed that RSS measurement is often highly over PSS which
seems to be more real memory usage for process. Using smaps we would
be able to calculate USS usage and know exact minimum value of memory
that would be freed after terminating some process. Those are very
important sources of information as they give as the possibility to
provide best possible app life-cycle.

We have also tried to use smaps in some application for OS developers
as source of detailed information of memory usage of the system.
For checking possible ways of improvement we tried totmaps from earlier
version. On sample case for our app the CPU usage as presented by 'top'
decreases from ~60% to ~4.5% only by changing source from smpas to tomaps.

So we are also very interested in using interface such as totmaps as it
gives detailed and complete memory usage information for user-space and
in our case much of information provided by smaps is for us not useful
at all.

We are also using or tried using other interfaces like status, statm,
cgroups.memory etc. but still totmaps/smaps are still the best interface
to get all of the informations per process based in single place.

>>
>>> We also use the private clean and private dirty and swap fields to
>>> make a few metrics for the processes and charge each process for it's
>>> private, shared, and swap data. Private clean and dirty are used for
>>> estimating a lower bound on how much memory would be freed.
>>
>> I can imagine that this kind of information might be useful and
>> presented in /proc/<pid>/statm. The question is whether some of the
>> existing consumers would see the performance impact due to he page table
>> walk. Anyway even these counters might get quite tricky because even
>> shareable resources are considered private if the process is the only
>> one to map them (so again this might be a file on tmpfs...).
>>
>>> Swap and
>>> PSS also give us some indication of additional memory which might get
>>> freed up.
>> --
>> Michal Hocko
>> SUSE Labs
>
>

-- 
Marcin Jabrzyk
Samsung R&D Institute Poland
Samsung Electronics

[toc] | [prev] | [next] | [standalone]


#1465996

FromSonny Rao <sonnyrao@chromium.org>
Date2016-08-19 07:20 +0200
Message-ID<s7GtR-7go-77@gated-at.bofh.it>
In reply to#1464982
On Thu, Aug 18, 2016 at 12:44 AM, Michal Hocko <mhocko@kernel.org> wrote:
> On Wed 17-08-16 11:57:56, Sonny Rao wrote:
>> On Wed, Aug 17, 2016 at 6:03 AM, Michal Hocko <mhocko@kernel.org> wrote:
>> > On Wed 17-08-16 11:31:25, Jann Horn wrote:
> [...]
>> >> That's at least 30.43% + 9.12% + 7.66% = 47.21% of the task's kernel
>> >> time spent on evaluating format strings. The new interface
>> >> wouldn't have to spend that much time on format strings because there
>> >> isn't so much text to format.
>> >
>> > well, this is true of course but I would much rather try to reduce the
>> > overhead of smaps file than add a new file. The following should help
>> > already. I've measured ~7% systime cut down. I guess there is still some
>> > room for improvements but I have to say I'm far from being convinced about
>> > a new proc file just because we suck at dumping information to the
>> > userspace.
>> > If this was something like /proc/<pid>/stat which is
>> > essentially read all the time then it would be a different question but
>> > is the rss, pss going to be all that often? If yes why?
>>
>> If the question is why do we need to read RSS, PSS, Private_*, Swap
>> and the other fields so often?
>>
>> I have two use cases so far involving monitoring per-process memory
>> usage, and we usually need to read stats for about 25 processes.
>>
>> Here's a timing example on an fairly recent ARM system 4 core RK3288
>> running at 1.8Ghz
>>
>> localhost ~ # time cat /proc/25946/smaps > /dev/null
>>
>> real    0m0.036s
>> user    0m0.020s
>> sys     0m0.020s
>>
>> localhost ~ # time cat /proc/25946/totmaps > /dev/null
>>
>> real    0m0.027s
>> user    0m0.010s
>> sys     0m0.010s
>> localhost ~ #
>>
>> I'll ignore the user time for now, and we see about 20 ms of system
>> time with smaps and 10 ms with totmaps, with 20 similar processes it
>> would be 400 milliseconds of cpu time for the kernel to get this
>> information from smaps vs 200 milliseconds with totmaps.  Even totmaps
>> is still pretty slow, but much better than smaps.
>>
>> Use cases:
>> 1) Basic task monitoring -- like "top" that shows memory consumption
>> including PSS, Private, Swap
>>     1 second update means about 40% of one CPU is spent in the kernel
>> gathering the data with smaps
>
> I would argue that even 20% is way too much for such a monitoring. What
> is the value to do it so often tha 20 vs 40ms really matters?

Yeah it is too much (I believe I said that) but it's significantly better.

>> 2) User space OOM handling -- we'd rather do a more graceful shutdown
>> than let the kernel's OOM killer activate and need to gather this
>> information and we'd like to be able to get this information to make
>> the decision much faster than 400ms
>
> Global OOM handling in userspace is really dubious if you ask me. I
> understand you want something better than SIGKILL and in fact this is
> already possible with memory cgroup controller (btw. memcg will give
> you a cheap access to rss, amount of shared, swapped out memory as
> well). Anyway if you are getting close to the OOM your system will most
> probably be really busy and chances are that also reading your new file
> will take much more time. I am also not quite sure how is pss useful for
> oom decisions.

I mentioned it before, but based on experience RSS just isn't good
enough -- there's too much sharing going on in our use case to make
the correct decision based on RSS.  If RSS were good enough, simply
put, this patch wouldn't exist.  So even with memcg I think we'd have
the same problem?

>
> Don't take me wrong, /proc/<pid>/totmaps might be suitable for your
> specific usecase but so far I haven't heard any sound argument for it to
> be generally usable. It is true that smaps is unnecessarily costly but
> at least I can see some room for improvements. A simple patch I've
> posted cut the formatting overhead by 7%. Maybe we can do more.

It seems like a general problem that if you want these values the
existing kernel interface can be very expensive, so it would be
generally usable by any application which wants a per process PSS,
private data, dirty data or swap value.   I mentioned two use cases,
but I guess I don't understand the comment about why it's not usable
by other use cases.

> --
> Michal Hocko
> SUSE Labs

[toc] | [prev] | [standalone]


Page 2 of 2 — ← Prev page 1 [2]

Back to top | Article view | linux.kernel


csiph-web