Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.user > #246707 > unrolled thread

Out of memory killer misconfigured?

Started bypiorunz <piorunz@gmx.com>
First post2022-03-29 11:40 +0200
Last post2022-04-20 12:30 +0200
Articles 20 on this page of 26 — 9 participants

Back to article view | Back to linux.debian.user


Contents

  Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-03-29 11:40 +0200
    Re: Out of memory killer misconfigured? Sven Hoexter <sven@stormbind.net> - 2022-03-29 12:10 +0200
      Re: Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-03-29 20:40 +0200
        Re: Out of memory killer misconfigured? Tixy <tixy@yxit.co.uk> - 2022-03-30 10:20 +0200
          Re: Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-03-31 19:00 +0200
            Re: Out of memory killer misconfigured? Jonathan Dowland <jon+debian-user@dow.land> - 2022-04-20 12:30 +0200
              Re: Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-04-20 17:30 +0200
                Re: Out of memory killer misconfigured? Jonathan Dowland <jon+debian-user@dow.land> - 2022-04-20 18:20 +0200
                  Re: Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-04-20 19:50 +0200
                    Re: Out of memory killer misconfigured? David Wright <deblis@lionunicorn.co.uk> - 2022-04-20 20:40 +0200
      Re: Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-03-29 20:50 +0200
        Re: Out of memory killer misconfigured? Greg Wooledge <greg@wooledge.org> - 2022-03-29 21:20 +0200
          Re: Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-03-29 22:10 +0200
            Re: Out of memory killer misconfigured? Greg Wooledge <greg@wooledge.org> - 2022-03-29 22:20 +0200
      Re: Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-03-29 21:00 +0200
        Re: Out of memory killer misconfigured? Nicholas Geovanis <nickgeovanis@gmail.com> - 2022-03-29 21:20 +0200
        Re: Out of memory killer misconfigured? <tomas@tuxteam.de> - 2022-04-01 08:10 +0200
          Re: Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-04-15 03:10 +0200
            Re: Out of memory killer misconfigured? <tomas@tuxteam.de> - 2022-04-15 08:00 +0200
              Re: Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-04-15 12:10 +0200
                Re: Out of memory killer misconfigured? <tomas@tuxteam.de> - 2022-04-15 12:20 +0200
                  Re: Out of memory killer misconfigured? piorunz <piorunz@gmx.com> - 2022-04-18 22:10 +0200
                    Re: Out of memory killer misconfigured? Tim Woodall <debianuser@woodall.me.uk> - 2022-04-19 17:50 +0200
                      Re: Out of memory killer misconfigured? <tomas@tuxteam.de> - 2022-04-19 18:10 +0200
                        Re: Out of memory killer misconfigured? Nicholas Geovanis <nickgeovanis@gmail.com> - 2022-04-19 22:10 +0200
    Re: Out of memory killer misconfigured? Jonathan Dowland <jon+debian-user@dow.land> - 2022-04-20 12:30 +0200

Page 1 of 2  [1] 2  Next page →


#246707 — Out of memory killer misconfigured?

Frompiorunz <piorunz@gmx.com>
Date2022-03-29 11:40 +0200
SubjectOut of memory killer misconfigured?
Message-ID<E6gut-4YqN-5@gated-at.bofh.it>

[Multipart message — attachments visible in raw view] — view raw

Hello,

I use Debian Testing on AMD64, on a workstation with Ryzen 5800X - 16
CPU cores and 64GB of ECC DDR4 RAM.

Today, Windows application I run on Wine for work has decided to eat all
available memory, CPU and HDD I/O. I don't have swapfile, so Linux
kernel must kill something to remain online when all RAM is taken by
rogue application.
That's where problem I noticed comes in - Debian oom-kill has killed
EVERYTHING and actual offending memory hungry application at the end.
Why?! It destroyed working KDE session and I had to hard reset the PC.

Have a look at journalctl results from last boot (I cut timestamps for
easier reading):

kernel: RSP: 002b:00007ffcff9ead98 EFLAGS: 00010246
systemd-journald[411]: Missed 10 kernel messages
kernel: lowmem_reserve[]: 0 3128 64155 64155 64155
kernel: Node 0 DMA32 free:246472kB boost:0kB min:3292kB low:6492kB
high:9692kB reserved_highatomic:0KB active_anon:44kB
inactive_anon:3057032kB active_file:0kB inactive_file:220kB un>
kernel: lowmem_reserve[]: 0 0 61027 61027 61027
kernel: Node 0 Normal free:245324kB boost:283884kB min:348156kB
low:410648kB high:473140kB reserved_highatomic:2048KB
active_anon:270712kB inactive_anon:60254116kB active_file:29564k>
kernel: lowmem_reserve[]: 0 0 0 0 0

(...)

Mar 29 08:58:28 ryzen kernel: 539654 total pagecache pages
Mar 29 08:58:28 ryzen kernel: 0 pages in swap cache
Mar 29 08:58:28 ryzen kernel: Swap cache stats: add 0, delete 0, find 0/0
Mar 29 08:58:28 ryzen kernel: Free swap  = 0kB
Mar 29 08:58:28 ryzen kernel: Total swap = 0kB
Mar 29 08:58:28 ryzen kernel: 16753821 pages RAM
Mar 29 08:58:28 ryzen kernel: 0 pages HighMem/MovableOnly
Mar 29 08:58:28 ryzen kernel: 296896 pages reserved
Mar 29 08:58:28 ryzen kernel: 0 pages hwpoisoned

And here we have all processes running, let me only highlight a few:

kernel: Tasks state (memory values in pages):
kernel: [  pid  ]   uid  tgid total_vm      rss pgtables_bytes swapents
oom_score_adj name
kernel: [   4611]  1000  4611   137776     1767   253952        0
     200 kactivitymanage
kernel: [ 676751]  1000 676751  2356154   115060  2740224        0
        0 terminal64.exe
kernel: [ 702184]  1000 702184  1226983   824654  7540736        0
        0 metatester64.ex
kernel: [ 731468]  1000 731468  1194211   814761  7442432        0
        0 metatester64.ex
kernel: [ 731471]  1000 731471  1245415   835020  7593984        0
        0 metatester64.ex
(and it goes on, at least 16 Wine exe processes like that eating all RAM)

As you can see, my Wine application has spawned a lot of exes, each one
of them uses around 7.5 million pagetables of memory (I am not sure what
is pagetable size in bytes in my Debian), and there are several of such
processes. But instead of killing one, or a few of these processes, OOM
manager has decided to kill everything *but* the offending exes. Killing
of all processes has begun:

kernel:
oom-kill:constraint=CONSTRAINT_NONE,nodemask=(null),cpuset=/,mems_allowed=0,global_oom,task_memcg=/user.slice/user-1000.slice/user@1000.service/background.slice/plasma-kactiv>
kernel: Out of memory: Killed process 4611 (kactivitymanage)
total-vm:551104kB, anon-rss:7068kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:248kB oom_score_adj:200

Instead of killing ONE 7.5 million-worth pagetable process, Linux is
killing everything else! KDE activity manager killed. Then it goes on to
kill EVERYTHING in the system:

kernel: Out of memory: Killed process 4555 (kglobalaccel5)
kernel: Out of memory: Killed process 444878 (kiod5)
kernel: oom_reaper: reaped process 444878 (kiod5)
kernel: Out of memory: Killed process 4405 (pipewire)
kernel: oom_reaper: reaped process 4405 (pipewire)
kernel: Out of memory: Killed process 505026 (gvfs-udisks2-vo)
(it goes on...)
Out of memory: Killed process 4414 (dbus-daemon)
(...)
Out of memory: Killed process 4390 (systemd)

And behold, at the end it kills Wine process:
Out of memory: Killed process 731550 (metatester64.ex)
total-vm:4891544kB, anon-rss:3483116kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:7704kB oom_score_adj:0

It even says total-vm:4891544kB, but just before that it killed systemd
with total-vm:18760kB.

At this stage, system is completely crashed and I have to hard reset.

I'd appreciate any explanation to this situation and how to prevent it
in the future.
Please find journalctl result as compressed attachment (16 KB).

I didn't modified Debian in any way which can affect RAM and out of
memory situations, apart from increasing I/O buffers for better
performance (comments to changes are my own):

$ cat /etc/sysctl.conf
(...)
vm.dirty_background_ratio=20
# Writing starts after 20% of RAM is filled with data to write.

vm.dirty_ratio=40
# up to 40% of memory can be used as write buffers (more write requests
will cause I/O lock until enough data is flushed).

vm.dirty_expire_centisecs=30000
# data is allowed to sit in the buffers for max 5 minutes (max lost work
time)

vm.dirty_writeback_centisecs=6000
# how often to check for write data in buffers: 1 minute

Not sure if that causes OOM to kill entire system instead of one
offending process, I doubt it.

Thanks in advance for your comments friends!

--
With kindest regards, Piotr.

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
⠈⠳⣄⠀⠀⠀⠀

[toc] | [next] | [standalone]


#246708

FromSven Hoexter <sven@stormbind.net>
Date2022-03-29 12:10 +0200
Message-ID<E6gXv-4YQ5-13@gated-at.bofh.it>
In reply to#246707
On Tue, Mar 29, 2022 at 10:34:19AM +0100, piorunz wrote:
> Hello,
> 
> I use Debian Testing on AMD64, on a workstation with Ryzen 5800X - 16
> CPU cores and 64GB of ECC DDR4 RAM.
> 
> Today, Windows application I run on Wine for work has decided to eat all
> available memory, CPU and HDD I/O. I don't have swapfile, so Linux
> kernel must kill something to remain online when all RAM is taken by
> rogue application.
> That's where problem I noticed comes in - Debian oom-kill has killed
> EVERYTHING and actual offending memory hungry application at the end.
> Why?! It destroyed working KDE session and I had to hard reset the PC.

The in kernel oom killing is a constant issue. If you look through the
lwn.net articles of the past years there is work done to improve the
situation, but I believe that's not in a default setup yet.

E.g. we now have PSI as an information source
https://lwn.net/Articles/759781/
which can be used with the Facebook oomd or systemd-oomd to
have userland control over which process to kill.

If you really want to fine tune your system this should give
you a lead what to look for.

Sven

[toc] | [prev] | [next] | [standalone]


#246721

Frompiorunz <piorunz@gmx.com>
Date2022-03-29 20:40 +0200
Message-ID<E6oV3-53yz-13@gated-at.bofh.it>
In reply to#246708
On 29/03/2022 10:56, Sven Hoexter wrote:

> The in kernel oom killing is a constant issue. If you look through the
> lwn.net articles of the past years there is work done to improve the
> situation, but I believe that's not in a default setup yet.

Yes it's terrible, how this can be broken so badly? Logic is very simple
here, even back in Windows XP times it was already solved by Microsoft?
Just one simple thing, one line of logic:

1. If RAM memory usage > 95% AND no swap available AND I/O cache already
dropped THEN kill most memory hungry process
2. Repeat

Job done!

But instead I ended up with 400 KB of logs showing how kernel oomd was
busy destroying my working KDE session.

>
> E.g. we now have PSI as an information source
> https://lwn.net/Articles/759781/
> which can be used with the Facebook oomd or systemd-oomd to
> have userland control over which process to kill.
>
> If you really want to fine tune your system this should give
> you a lead what to look for.
>
> Sven
>
Thanks Sven, I will have a look.
Facebook-oomd, systemd-oomd? And built-in Linux solution? How many are
out there? :O

Do you reckon I can report this as a bug against Linux kernel in Debian
bug track?

--
With kindest regards, Piotr.

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
⠈⠳⣄⠀⠀⠀⠀

[toc] | [prev] | [next] | [standalone]


#246735

FromTixy <tixy@yxit.co.uk>
Date2022-03-30 10:20 +0200
Message-ID<E6BIB-5bBb-7@gated-at.bofh.it>
In reply to#246721
On Tue, 2022-03-29 at 19:39 +0100, piorunz wrote:
> On 29/03/2022 10:56, Sven Hoexter wrote:
> 
> > The in kernel oom killing is a constant issue. If you look through the
> > lwn.net articles of the past years there is work done to improve the
> > situation, but I believe that's not in a default setup yet.
> 
> Yes it's terrible, how this can be broken so badly? Logic is very simple
> here, even back in Windows XP times it was already solved by Microsoft?
> Just one simple thing, one line of logic:
> 
> 1. If RAM memory usage > 95% AND no swap available AND I/O cache already
> dropped THEN kill most memory hungry process
> 2. Repeat
> 
> Job done!

I may be wrong here, but I seem to remember that something like that
used to happen a long time ago, and it had a habit of picking the X
server as the first thing to kill, not very friendly for GUI users.

-- 
Tixy

[toc] | [prev] | [next] | [standalone]


#246797

Frompiorunz <piorunz@gmx.com>
Date2022-03-31 19:00 +0200
Message-ID<E76jn-5u50-9@gated-at.bofh.it>
In reply to#246735
On 30/03/2022 09:18, Tixy wrote:

>
> I may be wrong here, but I seem to remember that something like that
> used to happen a long time ago, and it had a habit of picking the X
> server as the first thing to kill, not very friendly for GUI users.

Well, nowadays oom killer is not so picky. It just kills (almost)
EVERYTHING and then offending memory hungry process as a last,
destroying entire work session.
And actually I found out X session still running with everything else
killed, even more laughs.

--
With kindest regards, Piotr.

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
⠈⠳⣄⠀⠀⠀⠀

[toc] | [prev] | [next] | [standalone]


#247375

FromJonathan Dowland <jon+debian-user@dow.land>
Date2022-04-20 12:30 +0200
Message-ID<EefKW-9YMZ-9@gated-at.bofh.it>
In reply to#246797
On Thu, Mar 31, 2022 at 05:53:20PM +0100, piorunz wrote:
>Well, nowadays oom killer is not so picky. It just kills (almost)
>EVERYTHING and then offending memory hungry process as a last,
>destroying entire work session.

You are extrapolating from your single experience to make a
generalisation that might not be true.

-- 
Please do not CC me for listmail.

👱🏻	Jonathan Dowland
✎	 jmtd@debian.org
🔗	https://jmtd.net

[toc] | [prev] | [next] | [standalone]


#247383

Frompiorunz <piorunz@gmx.com>
Date2022-04-20 17:30 +0200
Message-ID<Eekrf-a1GD-5@gated-at.bofh.it>
In reply to#247375
On 20/04/2022 11:28, Jonathan Dowland wrote:
> On Thu, Mar 31, 2022 at 05:53:20PM +0100, piorunz wrote:
>> Well, nowadays oom killer is not so picky. It just kills (almost)
>> EVERYTHING and then offending memory hungry process as a last,
>> destroying entire work session.
>
> You are extrapolating from your single experience to make a
> generalisation that might not be true.

Sorry but this happened to me a few times, each tome with the same Wine
program. it just loves to eat all available memory when its doing heavy
computations. When I am not careful and I start too many threads, is
eats all memory. My system gets killed every single time. I am not
extrapolating anything. Misbehaving app should get killed, but instead,
Linux kills itself.


--
With kindest regards, Piotr.

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
⠈⠳⣄⠀⠀⠀⠀

[toc] | [prev] | [next] | [standalone]


#247384

FromJonathan Dowland <jon+debian-user@dow.land>
Date2022-04-20 18:20 +0200
Message-ID<EeldD-a2da-3@gated-at.bofh.it>
In reply to#247383
On Wed, Apr 20, 2022 at 04:23:38PM +0100, piorunz wrote:
>Sorry but this happened to me a few times, each tome with the same Wine
>program. it just loves to eat all available memory when its doing heavy
>computations. When I am not careful and I start too many threads, is
>eats all memory. My system gets killed every single time. I am not
>extrapolating anything. Misbehaving app should get killed, but instead,
>Linux kills itself.

Ok, not one experience, but one scenario. And there are plausible
reasons why this might be happening as outlined in my other mail to the
thread (OOM adjustments to make the killer skip over the process).
Occam's razor suggests something about your particular setup versus
the OOM killer simply being as bad as you think it is. FWIW, I invoke
the OOM killer a lot recently (due to some scientific experiments) and
the sensible processes were killed in my case, every time, leaving my
desktop functional.

More data about your setup is needed to fathom out what's going on.

-- 
Please do not CC me for listmail.

👱🏻	Jonathan Dowland
✎	 jmtd@debian.org
🔗	https://jmtd.net

[toc] | [prev] | [next] | [standalone]


#247388

Frompiorunz <piorunz@gmx.com>
Date2022-04-20 19:50 +0200
Message-ID<EemCJ-a2Vs-11@gated-at.bofh.it>
In reply to#247384
On 20/04/2022 17:11, Jonathan Dowland wrote:
> On Wed, Apr 20, 2022 at 04:23:38PM +0100, piorunz wrote:
>> Sorry but this happened to me a few times, each tome with the same Wine
>> program. it just loves to eat all available memory when its doing heavy
>> computations. When I am not careful and I start too many threads, is
>> eats all memory. My system gets killed every single time. I am not
>> extrapolating anything. Misbehaving app should get killed, but instead,
>> Linux kills itself.
>
> Ok, not one experience, but one scenario. And there are plausible
> reasons why this might be happening as outlined in my other mail to the
> thread (OOM adjustments to make the killer skip over the process).
> Occam's razor suggests something about your particular setup versus
> the OOM killer simply being as bad as you think it is. FWIW, I invoke
> the OOM killer a lot recently (due to some scientific experiments) and
> the sensible processes were killed in my case, every time, leaving my
> desktop functional.
>
> More data about your setup is needed to fathom out what's going on.
>
I didn't configured OOM. I use default Debian Testing with KDE. I
changed nothing apart from user desktop things. MY situation is not "a
scenario". It's what everyone can reproduce, just install Wine and
MetaTester5 program, I can guide you. I cannot guarantee this will
happen with non-Wine programs.

So I just reproduced this again, system crashed totally because I forgot
to apply choom --adjust 1000 to all metatester processes. Reboot. After
reboot, I started MT again and applied the following:

renice 19 `pidof metatester64.ex`
choom -p `pidof metatester64.ex | awk {'print $1'}` --adjust 1000
choom -p `pidof metatester64.ex | awk {'print $2'}` --adjust 1000
choom -p `pidof metatester64.ex | awk {'print $3'}` --adjust 1000
choom -p `pidof metatester64.ex | awk {'print $4'}` --adjust 1000
choom -p `pidof metatester64.ex | awk {'print $5'}` --adjust 1000
choom -p `pidof metatester64.ex | awk {'print $6'}` --adjust 1000
choom -p `pidof metatester64.ex | awk {'print $7'}` --adjust 1000

I spawned more than 8 processes to saturate RAM quickly. Meaning, some
processes were not adjusted to 1000 score. After 20 seconds, OOM event
happened.
And what Linux does? It has at least 8 processes with 1000 score, but
let's see what is does:

$ sudo dmesg | grep "Out of memory"
[  704.018238] Out of memory: Killed process 39674 (metatester64.ex)
total-vm:2939256kB, anon-rss:250524kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:1344kB oom_score_adj:1000
[  708.528663] Out of memory: Killed process 25882 (QtWebEngineProc)
total-vm:5451920kB, anon-rss:20588kB, file-rss:0kB, shmem-rss:116kB,
UID:1000 pgtables:772kB oom_score_adj:300
[  709.811945] Out of memory: Killed process 23411 (krunner-keepass)
total-vm:52680kB, anon-rss:12860kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:140kB oom_score_adj:200
[  710.788999] Out of memory: Killed process 23413 (pipewire-media-)
total-vm:170052kB, anon-rss:9492kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:248kB oom_score_adj:200
[  712.748691] Out of memory: Killed process 23618 (kactivitymanage)
total-vm:550696kB, anon-rss:8848kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:244kB oom_score_adj:200
[  714.856773] Out of memory: Killed process 23414 (pulseaudio)
total-vm:1384772kB, anon-rss:8508kB, file-rss:0kB, shmem-rss:260kB,
UID:1000 pgtables:320kB oom_score_adj:200
[  716.051274] Out of memory: Killed process 23560 (kglobalaccel5)
total-vm:282368kB, anon-rss:6728kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:256kB oom_score_adj:200
[  716.082495] Out of memory: Killed process 23412 (pipewire)
total-vm:53124kB, anon-rss:4484kB, file-rss:0kB, shmem-rss:52kB,
UID:1000 pgtables:104kB oom_score_adj:200
[  717.965930] Out of memory: Killed process 23867 (ksystemstats)
total-vm:166540kB, anon-rss:3844kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:176kB oom_score_adj:200
[  718.775772] Out of memory: Killed process 23703 (kscreen_backend)
total-vm:227996kB, anon-rss:3060kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:164kB oom_score_adj:200
[  718.792339] Out of memory: Killed process 23417 (dbus-daemon)
total-vm:10552kB, anon-rss:1332kB, file-rss:0kB, shmem-rss:0kB, UID:1000
pgtables:60kB oom_score_adj:200
[  719.422552] Out of memory: Killed process 23714 (obexd)
total-vm:45836kB, anon-rss:732kB, file-rss:0kB, shmem-rss:0kB, UID:1000
pgtables:76kB oom_score_adj:200
[  719.423354] Out of memory: Killed process 23555 (dconf-service)
total-vm:157600kB, anon-rss:544kB, file-rss:0kB, shmem-rss:0kB, UID:1000
pgtables:64kB oom_score_adj:200
[  719.424525] Out of memory: Killed process 23397 ((sd-pam))
total-vm:169688kB, anon-rss:3724kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:92kB oom_score_adj:100
[  719.425356] Out of memory: Killed process 23396 (systemd)
total-vm:18600kB, anon-rss:1912kB, file-rss:0kB, shmem-rss:0kB, UID:1000
pgtables:76kB oom_score_adj:100
[  719.426070] Out of memory: Killed process 55808 (metatester64.ex)
total-vm:5510844kB, anon-rss:3909364kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:8516kB oom_score_adj:0

I lost Metatester (as expected), but I also lost sound and shortcuts and
KDE widgets. I need to reboot.

Look at memory sizes for last process and other processes killed. Last one:
Out of memory: Killed process 55808 (metatester64.ex)
total-vm:5510844kB, anon-rss:3909364kB, file-rss:0kB, shmem-rss:0kB,
UID:1000 pgtables:8516kB oom_score_adj:0

Why this is not killed right after 39674 (metatester64.ex)?
Why other 7 metatester64.ex processes for which I adjusted priority to
1000 were not killed before KDE processes were killed?

--
With kindest regards, Piotr.

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
⠈⠳⣄⠀⠀⠀⠀

[toc] | [prev] | [next] | [standalone]


#247391

FromDavid Wright <deblis@lionunicorn.co.uk>
Date2022-04-20 20:40 +0200
Message-ID<Eenp7-a3rp-11@gated-at.bofh.it>
In reply to#247388
On Wed 20 Apr 2022 at 18:45:47 (+0100), piorunz wrote:
> On 20/04/2022 17:11, Jonathan Dowland wrote:
> > On Wed, Apr 20, 2022 at 04:23:38PM +0100, piorunz wrote:
> > > Sorry but this happened to me a few times, each tome with the same Wine
> > > program. it just loves to eat all available memory when its doing heavy
> > > computations. When I am not careful and I start too many threads, is
> > > eats all memory. My system gets killed every single time. I am not
> > > extrapolating anything. Misbehaving app should get killed, but instead,
> > > Linux kills itself.
> > 
> > Ok, not one experience, but one scenario. And there are plausible
> > reasons why this might be happening as outlined in my other mail to the
> > thread (OOM adjustments to make the killer skip over the process).
> > Occam's razor suggests something about your particular setup versus
> > the OOM killer simply being as bad as you think it is. FWIW, I invoke
> > the OOM killer a lot recently (due to some scientific experiments) and
> > the sensible processes were killed in my case, every time, leaving my
> > desktop functional.
> > 
> > More data about your setup is needed to fathom out what's going on.
> > 
> I didn't configured OOM. I use default Debian Testing with KDE. I
> changed nothing apart from user desktop things. MY situation is not "a
> scenario". It's what everyone can reproduce, just install Wine and
> MetaTester5 program, I can guide you. I cannot guarantee this will
> happen with non-Wine programs.

With respect, the combination of testing + KDE + Wine + MetaTester5
is difficult to categorise as more than one scenario.

> So I just reproduced this again, system crashed totally because I forgot
> to apply choom --adjust 1000 to all metatester processes. Reboot. After
> reboot, I started MT again and applied the following:
> 
> renice 19 `pidof metatester64.ex`
> choom -p `pidof metatester64.ex | awk {'print $1'}` --adjust 1000
> choom -p `pidof metatester64.ex | awk {'print $2'}` --adjust 1000
> choom -p `pidof metatester64.ex | awk {'print $3'}` --adjust 1000
> choom -p `pidof metatester64.ex | awk {'print $4'}` --adjust 1000
> choom -p `pidof metatester64.ex | awk {'print $5'}` --adjust 1000
> choom -p `pidof metatester64.ex | awk {'print $6'}` --adjust 1000
> choom -p `pidof metatester64.ex | awk {'print $7'}` --adjust 1000
> 
> I spawned more than 8 processes to saturate RAM quickly. Meaning, some
> processes were not adjusted to 1000 score. After 20 seconds, OOM event
> happened.

When I start my window manager, I use a construction like this:

  if [ -x /usr/bin/fvwm ]; then
      mv -f "$HOME/.xsession-fvwm-$Displaynumber-log" "$HOME/.xsession-fvwm-$Displaynumber-log~"
      exec /usr/bin/fvwm > "$HOME/.xsession-fvwm-$Displaynumber-log" 2>&1 & Wmpid=$!
  elif [ -x …
      …
      …
  else
      printf '%s\n' "Error - no window manager found"
  fi

so that .xsession can avoid terminating with:

  wait $Wmpid

Would it be sensible for you to use a similar construction so that you
systematically catch /all/ of these processes with their choom commands.

> And what Linux does? It has at least 8 processes with 1000 score, but
> let's see what is does:
> 
> $ sudo dmesg | grep "Out of memory"
> [  704.018238] Out of memory: Killed process 39674 (metatester64.ex)
> total-vm:2939256kB, anon-rss:250524kB, file-rss:0kB, shmem-rss:0kB,
> UID:1000 pgtables:1344kB oom_score_adj:1000
> [  708.528663] Out of memory: Killed process 25882 (QtWebEngineProc)
> total-vm:5451920kB, anon-rss:20588kB, file-rss:0kB, shmem-rss:116kB,
> UID:1000 pgtables:772kB oom_score_adj:300
> [  709.811945] Out of memory: Killed process 23411 (krunner-keepass)
> total-vm:52680kB, anon-rss:12860kB, file-rss:0kB, shmem-rss:0kB,
> UID:1000 pgtables:140kB oom_score_adj:200
> [  710.788999] Out of memory: Killed process 23413 (pipewire-media-)
> total-vm:170052kB, anon-rss:9492kB, file-rss:0kB, shmem-rss:0kB,
> UID:1000 pgtables:248kB oom_score_adj:200
> [  712.748691] Out of memory: Killed process 23618 (kactivitymanage)
> total-vm:550696kB, anon-rss:8848kB, file-rss:0kB, shmem-rss:0kB,
> UID:1000 pgtables:244kB oom_score_adj:200
> [  714.856773] Out of memory: Killed process 23414 (pulseaudio)
> total-vm:1384772kB, anon-rss:8508kB, file-rss:0kB, shmem-rss:260kB,
> UID:1000 pgtables:320kB oom_score_adj:200
> [  716.051274] Out of memory: Killed process 23560 (kglobalaccel5)
> total-vm:282368kB, anon-rss:6728kB, file-rss:0kB, shmem-rss:0kB,
> UID:1000 pgtables:256kB oom_score_adj:200
> [  716.082495] Out of memory: Killed process 23412 (pipewire)
> total-vm:53124kB, anon-rss:4484kB, file-rss:0kB, shmem-rss:52kB,
> UID:1000 pgtables:104kB oom_score_adj:200
> [  717.965930] Out of memory: Killed process 23867 (ksystemstats)
> total-vm:166540kB, anon-rss:3844kB, file-rss:0kB, shmem-rss:0kB,
> UID:1000 pgtables:176kB oom_score_adj:200
> [  718.775772] Out of memory: Killed process 23703 (kscreen_backend)
> total-vm:227996kB, anon-rss:3060kB, file-rss:0kB, shmem-rss:0kB,
> UID:1000 pgtables:164kB oom_score_adj:200
> [  718.792339] Out of memory: Killed process 23417 (dbus-daemon)
> total-vm:10552kB, anon-rss:1332kB, file-rss:0kB, shmem-rss:0kB, UID:1000
> pgtables:60kB oom_score_adj:200
> [  719.422552] Out of memory: Killed process 23714 (obexd)
> total-vm:45836kB, anon-rss:732kB, file-rss:0kB, shmem-rss:0kB, UID:1000
> pgtables:76kB oom_score_adj:200
> [  719.423354] Out of memory: Killed process 23555 (dconf-service)
> total-vm:157600kB, anon-rss:544kB, file-rss:0kB, shmem-rss:0kB, UID:1000
> pgtables:64kB oom_score_adj:200
> [  719.424525] Out of memory: Killed process 23397 ((sd-pam))
> total-vm:169688kB, anon-rss:3724kB, file-rss:0kB, shmem-rss:0kB,
> UID:1000 pgtables:92kB oom_score_adj:100
> [  719.425356] Out of memory: Killed process 23396 (systemd)
> total-vm:18600kB, anon-rss:1912kB, file-rss:0kB, shmem-rss:0kB, UID:1000
> pgtables:76kB oom_score_adj:100
> [  719.426070] Out of memory: Killed process 55808 (metatester64.ex)
> total-vm:5510844kB, anon-rss:3909364kB, file-rss:0kB, shmem-rss:0kB,
> UID:1000 pgtables:8516kB oom_score_adj:0
> 
> I lost Metatester (as expected), but I also lost sound and shortcuts and
> KDE widgets. I need to reboot.
> 
> Look at memory sizes for last process and other processes killed. Last one:
> Out of memory: Killed process 55808 (metatester64.ex)
> total-vm:5510844kB, anon-rss:3909364kB, file-rss:0kB, shmem-rss:0kB,
> UID:1000 pgtables:8516kB oom_score_adj:0
> 
> Why this is not killed right after 39674 (metatester64.ex)?

We don't see all the process numbers listed for these tester programs,
but the 55808 could indicate that it missed being choomed. It might
help to be more systematic with starting the processes and with the
logged information.

> Why other 7 metatester64.ex processes for which I adjusted priority to
> 1000 were not killed before KDE processes were killed?

I've no idea. I don't run a DE, for starters.

Cheers,
David.

[toc] | [prev] | [next] | [standalone]


#246723

Frompiorunz <piorunz@gmx.com>
Date2022-03-29 20:50 +0200
Message-ID<E6p4K-53BU-11@gated-at.bofh.it>
In reply to#246708
On 29/03/2022 10:56, Sven Hoexter wrote:

And I wanted to highlight that this is not some trivial problem, at the
moment any rogue non-privileged process running on user account, or
simple coding error can destroy entire Linux session, be is a important
server with many services running or just a normal PC, computer is
totally unusable until manual intervention (hard reset by hand)!

--
With kindest regards, Piotr.

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
⠈⠳⣄⠀⠀⠀⠀

[toc] | [prev] | [next] | [standalone]


#246728

FromGreg Wooledge <greg@wooledge.org>
Date2022-03-29 21:20 +0200
Message-ID<E6pxL-541D-23@gated-at.bofh.it>
In reply to#246723
On Tue, Mar 29, 2022 at 07:41:53PM +0100, piorunz wrote:
> And I wanted to highlight that this is not some trivial problem, at the
> moment any rogue non-privileged process running on user account, or
> simple coding error can destroy entire Linux session, be is a important
> server with many services running or just a normal PC, computer is
> totally unusable until manual intervention (hard reset by hand)!

Resource limits are a thing.  man setrlimit (you may have to install
manpages-dev first).

Are they crude as hell?  Absolutely.  But they allow you to set up a
barrier that says "if you use more than x MB of memory, you die".  They
work well enough for some cases.

[toc] | [prev] | [next] | [standalone]


#246730

Frompiorunz <piorunz@gmx.com>
Date2022-03-29 22:10 +0200
Message-ID<E6qk9-54y8-11@gated-at.bofh.it>
In reply to#246728
On 29/03/2022 20:16, Greg Wooledge wrote:
> Resource limits are a thing.  man setrlimit (you may have to install
> manpages-dev first).
>
> Are they crude as hell?  Absolutely.  But they allow you to set up a
> barrier that says "if you use more than x MB of memory, you die".  They
> work well enough for some cases.

That's exactly what I need, nothing else! I don't understand why memory
limit like that is not built in to Debian, opening possibility to DoS
every system by allocating too much memory by unprivileged userland
process. Unless this is due to my configuration and everyone else is not
experiencing that?

How do I use it? I've read manual but that would need to written down as
a C++ program or something? ;(

--
With kindest regards, Piotr.

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
⠈⠳⣄⠀⠀⠀⠀

[toc] | [prev] | [next] | [standalone]


#246732

FromGreg Wooledge <greg@wooledge.org>
Date2022-03-29 22:20 +0200
Message-ID<E6qtP-54Bg-1@gated-at.bofh.it>
In reply to#246730
On Tue, Mar 29, 2022 at 09:03:14PM +0100, piorunz wrote:
> On 29/03/2022 20:16, Greg Wooledge wrote:
> > Resource limits are a thing.  man setrlimit (you may have to install
> > manpages-dev first).
> 
> How do I use it? I've read manual but that would need to written down as
> a C++ program or something? ;(

Usually things are launched from either a shell, or from systemd.

If the thing is launched from bash, then you can use bash's "ulimit"
command to set the resource limits before launching the thing.

If it's launched by systemd, look at systemd.exec(5) and search for
"Limit" to see how to specify resource limits in a unit file.

If it's launched from some shell that isn't bash, then you might be
able to find a "ulimit" or "limit" command that works in the other
shell.  Otherwise, you can use wrapper tools like DJB's "softlimit"
(from daemontools) to set resource limits and chain-load the desired thing.

> I need, nothing else! I don't understand why memory
> limit like that is not built in to Debian, opening possibility to DoS
> every system by allocating too much memory by unprivileged userland
> process.

Resource limits are not set by default, because they will cause processes
to die if they're set too low.  Only the person who needs them will be
able to determine which processes need a limit placed on them, and how
low the limit should be set.

[toc] | [prev] | [next] | [standalone]


#246725

Frompiorunz <piorunz@gmx.com>
Date2022-03-29 21:00 +0200
Message-ID<E6pep-53F8-13@gated-at.bofh.it>
In reply to#246708
On 29/03/2022 10:56, Sven Hoexter wrote:
> E.g. we now have PSI as an information source
> https://lwn.net/Articles/759781/
> which can be used with the Facebook oomd or systemd-oomd to
> have userland control over which process to kill.

Thanks, I've read this article. Unfortunately, this is just information
tool which can be used by engineers and developers so design their own
oomd. I am not a developer.

Is there any config file I can edit so just simply ask oomd to kill most
memory hugging process instead of entire system?

--
With kindest regards, Piotr.

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
⠈⠳⣄⠀⠀⠀⠀

[toc] | [prev] | [next] | [standalone]


#246727

FromNicholas Geovanis <nickgeovanis@gmail.com>
Date2022-03-29 21:20 +0200
Message-ID<E6pxL-541D-17@gated-at.bofh.it>
In reply to#246725

[Multipart message — attachments visible in raw view] — view raw

On Tue, Mar 29, 2022 at 1:59 PM piorunz <piorunz@gmx.com> wrote:

> On 29/03/2022 10:56, Sven Hoexter wrote:
> > E.g. we now have PSI as an information source
> > https://lwn.net/Articles/759781/
> > which can be used with the Facebook oomd or systemd-oomd to
> > have userland control over which process to kill.
>
> Thanks, I've read this article. Unfortunately, this is just information
> tool which can be used by engineers and developers so design their own
> oomd. I am not a developer.
>
> Is there any config file I can edit so just simply ask oomd to kill most
> memory hugging process instead of entire system?
>

Yes. Looks like oomd came from facebook :-)
https://github.com/facebookincubator/oomd/blob/main/docs/configuration.md

That's the kind of tool that gets turned-off in backend servers.

--
> With kindest regards, Piotr.
>
> ⢀⣴⠾⠻⢶⣦⠀
> ⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
> ⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
> ⠈⠳⣄⠀⠀⠀⠀
>
>

[toc] | [prev] | [next] | [standalone]


#246807

From<tomas@tuxteam.de>
Date2022-04-01 08:10 +0200
Message-ID<E7iDU-5CdC-1@gated-at.bofh.it>
In reply to#246725

[Multipart message — attachments visible in raw view] — view raw

On Tue, Mar 29, 2022 at 07:58:38PM +0100, piorunz wrote:
> On 29/03/2022 10:56, Sven Hoexter wrote:
> > E.g. we now have PSI as an information source
> > https://lwn.net/Articles/759781/
> > which can be used with the Facebook oomd or systemd-oomd to
> > have userland control over which process to kill.
> 
> Thanks, I've read this article. Unfortunately, this is just information
> tool which can be used by engineers and developers so design their own
> oomd. I am not a developer.
> 
> Is there any config file I can edit so just simply ask oomd to kill most
> memory hugging process instead of entire system?

See man (1) choom and search for oom in man (5) proc.

Cheers
-- 
t

[toc] | [prev] | [next] | [standalone]


#247227

Frompiorunz <piorunz@gmx.com>
Date2022-04-15 03:10 +0200
Message-ID<EciDf-8NOj-1@gated-at.bofh.it>
In reply to#246807
On 01/04/2022 07:08, tomas@tuxteam.de wrote:
> On Tue, Mar 29, 2022 at 07:58:38PM +0100, piorunz wrote:
>> On 29/03/2022 10:56, Sven Hoexter wrote:
>>> E.g. we now have PSI as an information source
>>> https://lwn.net/Articles/759781/
>>> which can be used with the Facebook oomd or systemd-oomd to
>>> have userland control over which process to kill.
>>
>> Thanks, I've read this article. Unfortunately, this is just information
>> tool which can be used by engineers and developers so design their own
>> oomd. I am not a developer.
>>
>> Is there any config file I can edit so just simply ask oomd to kill most
>> memory hugging process instead of entire system?
>
> See man (1) choom and search for oom in man (5) proc.
>
> Cheers

Thanks! That's exactly what I needed. Amazing.

Now, after I start the process, I run:

Adjust to maximum setting (kills first):
choom -p `pidof terminal64.exe` --adjust 1000

Lower priority of a processes:
renice 19 `pidof terminal64.exe`

That way offending process will never hang or destabilize my system.

I just wish that Linux kernel would give maximum oom score to process
with most memory, that's so obvious! But instead, defaults are so bad,
it kills everything but one offending process.
Can this be reported as a bug against linux kernel in Debian bug track?

--
With kindest regards, Piotr.

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
⠈⠳⣄⠀⠀⠀⠀

[toc] | [prev] | [next] | [standalone]


#247229

From<tomas@tuxteam.de>
Date2022-04-15 08:00 +0200
Message-ID<Ecn9T-8Qsj-1@gated-at.bofh.it>
In reply to#247227

[Multipart message — attachments visible in raw view] — view raw

On Fri, Apr 15, 2022 at 02:09:04AM +0100, piorunz wrote:
> On 01/04/2022 07:08, tomas@tuxteam.de wrote:

[...]

> > See man (1) choom and search for oom in man (5) proc.

> Thanks! That's exactly what I needed. Amazing.

You're welcome :)

> Now, after I start the process, I run:
> 
> Adjust to maximum setting (kills first):
> choom -p `pidof terminal64.exe` --adjust 1000
> 
> Lower priority of a processes:
> renice 19 `pidof terminal64.exe`
> 
> That way offending process will never hang or destabilize my system.
> 
> I just wish that Linux kernel would give maximum oom score to process
> with most memory, that's so obvious!

To a certain extent, it does (see below).

>                                       But instead, defaults are so bad,
> it kills everything but one offending process.
> Can this be reported as a bug against linux kernel in Debian bug track?

You can, of course, try. Your chances of success would be better
if you tried to understand what kind thoughts have gone into it,
and why it's not working for you. The way you are putting it, it
comes across (to me, at least!) as "what kind of stupid design
is this" (cf. your words: "obvious", "defaults so bad"). I know
you don't intend that, but, as you can imagine, it won't fly if
that's how others perceive it :-)

If you want to learn more about that, the Linux MM ("memory
management") people have set up a wiki for that:

  https://linux-mm.org/OOM_Killer

Enjoy :)

(and yes, on my box, more memory-hungry processes have a higher
/proc/<pid>/oom_score, so they seem to incur a higher risk of
being killed. The browser lies somewhat because it spawns quite
a few processes, but as a product of the propaganda industry,
lying is its second nature ;-)

Cheers
-- 
t


[toc] | [prev] | [next] | [standalone]


#247231

Frompiorunz <piorunz@gmx.com>
Date2022-04-15 12:10 +0200
Message-ID<Ecr3P-8T0A-9@gated-at.bofh.it>
In reply to#247229
On 15/04/2022 06:53, tomas@tuxteam.de wrote:
> If you want to learn more about that, the Linux MM ("memory
> management") people have set up a wiki for that:
>
>    https://linux-mm.org/OOM_Killer
>
> Enjoy:)
>
> (and yes, on my box, more memory-hungry processes have a higher
> /proc/<pid>/oom_score, so they seem to incur a higher risk of
> being killed. The browser lies somewhat because it spawns quite
> a few processes, but as a product of the propaganda industry,
> lying is its second nature;-)

Yes I understand that from my point of view everything is bad and I am
not objective.
However, I see how things are broken and that is indeed obvious.
I seen before my own eyes how Linux killed KDE session including Xorg
but left 10x wine's exe processes with 8GB RAM each as a last. Entire
tree with parent process was using all of the memory that would be 50+
GB. How this process tree is not having HIGHEST POSSIBLE OOM score and
is not being killed in first microsecond of OOM situation is beyond my
understanding.

--
With kindest regards, Piotr.

⢀⣴⠾⠻⢶⣦⠀
⣾⠁⢠⠒⠀⣿⡁ Debian - The universal operating system
⢿⡄⠘⠷⠚⠋⠀ https://www.debian.org/
⠈⠳⣄⠀⠀⠀⠀

[toc] | [prev] | [next] | [standalone]


Page 1 of 2  [1] 2  Next page →

Back to top | Article view | linux.debian.user


csiph-web