Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.project > #13910 > unrolled thread

Salsa CI overload

Started byIan Jackson <ijackson@chiark.greenend.org.uk>
First post2025-08-25 22:20 +0200
Last post2025-08-27 11:40 +0200
Articles 17 — 12 participants

Back to article view | Back to linux.debian.project


Contents

  Salsa CI overload Ian Jackson <ijackson@chiark.greenend.org.uk> - 2025-08-25 22:20 +0200
    Re: Salsa CI overload Louis-Philippe Véronneau <pollo@debian.org> - 2025-08-25 23:50 +0200
      Re: Salsa CI overload Simon Josefsson <simon@josefsson.org> - 2025-08-26 09:50 +0200
        Re: Salsa CI overload Louis-Philippe Véronneau <pollo@debian.org> - 2025-08-26 17:10 +0200
        Re: Salsa CI overload Otto Kekäläinen <otto@debian.org> - 2025-09-05 03:20 +0200
          Re: Salsa CI overload Lucas Nussbaum <lucas@debian.org> - 2025-09-05 15:40 +0200
            Re: Salsa CI overload Otto Kekäläinen <otto@debian.org> - 2025-09-05 19:40 +0200
              Re: Salsa CI overload Lucas Nussbaum <lucas@debian.org> - 2025-09-05 21:40 +0200
    Re: Salsa CI overload Alexander Wirt <formorer@debian.org> - 2025-08-27 00:30 +0200
      Re: Salsa CI overload Antoine Le Gonidec <vv221@debian.org> - 2025-08-27 01:30 +0200
        Re: Salsa CI overload Soren Stoutner <soren@debian.org> - 2025-08-27 23:50 +0200
        Re: Salsa CI overload Didier 'OdyX' Raboud <odyx@debian.org> - 2025-08-28 09:30 +0200
          salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload") Antonio Terceiro <terceiro@debian.org> - 2025-08-28 15:00 +0200
            Re: salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload") Jérémy Lal <kapouer@melix.org> - 2025-08-28 15:30 +0200
            Re: salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload") Marco d'Itri <md@Linux.IT> - 2025-08-28 15:30 +0200
              Re: salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload") Antonio Terceiro <terceiro@debian.org> - 2025-08-28 17:00 +0200
      Re: Salsa CI overload Ian Jackson <ijackson@chiark.greenend.org.uk> - 2025-08-27 11:40 +0200

#13910 — Salsa CI overload

FromIan Jackson <ijackson@chiark.greenend.org.uk>
Date2025-08-25 22:20 +0200
SubjectSalsa CI overload
Message-ID<LnLPz-ajYx-1@gated-at.bofh.it>
Hi.

(I wasn't able to find another mailing list for this.  I found a bit
of discussion on debian-devel about one specific overload incident.)

Since sid opened, people have naturally been hard at work.  Great!

But I notice that Salsa CI is now frequently very overloaded.  Even
when it's not half a day behind (!), I often see that my package is
able to run only about one job at a time.  I think this is slowing a
lot of us down quite considerably.

It seems to me that we (Debian) could probably acquire more computing
resources if we wanted.  Perhaps some of our jobs are wasteful, but I
think that overall, adding capacity would be well justified.

I haven't seen any discussion of this anywhere.  Is work ongoing to
try to add more capacity to Salsa CI?  What are the key difficulties?

If we knew what kind of expertise/assistance/input/sponsorship was
needed, I feel we could probably find it within our project.

Thanks,
Ian.

-- 
Ian Jackson <ijackson@chiark.greenend.org.uk>   These opinions are my own.  

Pronouns: they/he.  If I emailed you from @fyvzl.net or @evade.org.uk,
that is a private address which bypasses my fierce spamfilter.

[toc] | [next] | [standalone]


#13911

FromLouis-Philippe Véronneau <pollo@debian.org>
Date2025-08-25 23:50 +0200
Message-ID<LnNeF-akMb-7@gated-at.bofh.it>
In reply to#13910

[Multipart message — attachments visible in raw view] — view raw

On 2025-08-25 4 h 05 p.m., Ian Jackson wrote:
> Hi.
> 
> (I wasn't able to find another mailing list for this.  I found a bit
> of discussion on debian-devel about one specific overload incident.)
> 
> Since sid opened, people have naturally been hard at work.  Great!
> 
> But I notice that Salsa CI is now frequently very overloaded.  Even
> when it's not half a day behind (!), I often see that my package is
> able to run only about one job at a time.  I think this is slowing a
> lot of us down quite considerably.
> 
> It seems to me that we (Debian) could probably acquire more computing
> resources if we wanted.  Perhaps some of our jobs are wasteful, but I
> think that overall, adding capacity would be well justified.
> 
> I haven't seen any discussion of this anywhere.  Is work ongoing to
> try to add more capacity to Salsa CI?  What are the key difficulties?
> 
> If we knew what kind of expertise/assistance/input/sponsorship was
> needed, I feel we could probably find it within our project.
> 
> Thanks,
> Ian.
> 

I raised a similar issue during DebConf25 and it looks like some changes 
were made back then [1], as it seems to me Salsa CI is indeed faster 
than it used to be.

Generally though, Salsa CI is pretty underpowered, which definitely 
slows down my Debian work. Running the Lintian testsuite (which is a 
requirement to merge contributions) frequently takes more than 2 hours 
to run. On my local machine at home, it takes around 4 minutes :(

I feel like making people wait around for CI to run isn't a great use of 
our collective time.

IIUC, runners are currently sponsored machines running in the Google 
Cloud [2]. Maybe we could prod them again and ask for some more 
resources, pretty please?

[1]: many thanks to the folks who worked on that issue

[2]: 
https://salsa.debian.org/salsa/salsa-ansible/-/blob/master/inventories/prod/host_vars/salsa-runner.salsa-runner.debian.net.yml?ref_type=heads#L8-9

-- 
   ⢀⣴⠾⠻⢶⣦⠀
   ⣾⠁⢠⠒⠀⣿⡁  Louis-Philippe Véronneau
   ⢿⡄⠘⠷⠚⠋   pollo@debian.org / veronneau.org
   ⠈⠳⣄

[toc] | [prev] | [next] | [standalone]


#13912

FromSimon Josefsson <simon@josefsson.org>
Date2025-08-26 09:50 +0200
Message-ID<LnWBj-arAd-19@gated-at.bofh.it>
In reply to#13911

[Multipart message — attachments visible in raw view] — view raw

Louis-Philippe Véronneau <pollo@debian.org> writes:

> I feel like making people wait around for CI to run isn't a great use
> of our collective time.

+1 to more Salsa CI runner resources.

> Generally though, Salsa CI is pretty underpowered, which definitely
> slows down my Debian work. Running the Lintian testsuite (which is a
> requirement to merge contributions) frequently takes more than 2 hours
> to run. On my local machine at home, it takes around 4 minutes :(

What?  Do you have a pointer to a job where this happen?  Are you sure
you don't confuse queueing time with running time?  Salsa CI queueing
time can be hours/days for the last few days, but my experience is that
runtime is fairly okay once the job starts.

/Simon

[toc] | [prev] | [next] | [standalone]


#13913

FromLouis-Philippe Véronneau <pollo@debian.org>
Date2025-08-26 17:10 +0200
Message-ID<Lo3t7-awry-9@gated-at.bofh.it>
In reply to#13912
On 2025-08-26 3 h 29 a.m., Simon Josefsson wrote:
> Louis-Philippe Véronneau <pollo@debian.org> writes:
> 
>> I feel like making people wait around for CI to run isn't a great use
>> of our collective time.
> 
> +1 to more Salsa CI runner resources.
> 
>> Generally though, Salsa CI is pretty underpowered, which definitely
>> slows down my Debian work. Running the Lintian testsuite (which is a
>> requirement to merge contributions) frequently takes more than 2 hours
>> to run. On my local machine at home, it takes around 4 minutes :(
> 
> What?  Do you have a pointer to a job where this happen?  Are you sure
> you don't confuse queueing time with running time?  Salsa CI queueing
> time can be hours/days for the last few days, but my experience is that
> runtime is fairly okay once the job starts.
> 
> /Simon

The Lintian testsuite is pretty big (it builds ~1500 Debian packages and 
then runs Lintian on each of them) and the more CPU cores available, the 
fastest it is...

You'll find many examples of actual pipeline runtimes over 2h here: 
https://salsa.debian.org/lintian/lintian/-/pipelines

-- 
   ⢀⣴⠾⠻⢶⣦⠀
   ⣾⠁⢠⠒⠀⣿⡁  Louis-Philippe Véronneau
   ⢿⡄⠘⠷⠚⠋   pollo@debian.org / veronneau.org
   ⠈⠳⣄

[toc] | [prev] | [next] | [standalone]


#13933

FromOtto Kekäläinen <otto@debian.org>
Date2025-09-05 03:20 +0200
Message-ID<Lrthn-cSVL-1@gated-at.bofh.it>
In reply to#13912
Hi Denis,

> >> I feel like making people wait around for CI to run isn't a great use
> >> of our collective time.
> >
> > +1 to more Salsa CI runner resources.
>
> Is there a way for companies or institutes to contribute by supplying
> compute cycles from their clusters or data centres? I'm working for a
> research lab that has ample compute power and we would like to help out,
> but there may be some conditions that make this more complicated.

Do you represent and organization willing to donate capacity?

Some potential donors have expressed their willingness to contribute
resources tickets at https://salsa.debian.org/salsa/support, see e.g.
https://salsa.debian.org/salsa/support/-/issues/301, but based on what
I have understood from Salsa Admins is that it is easier for them to
simply expand the existing fleet at Google Cloud than to support hosts
across multiple cloud providers / datacenters. They have also been
increasing the runner fleet size in past years, and based on latest
email seems the salsa-wide CI system has up to 64 parallel runners
available at the moment.

Looking at the recently published https://salsa-status.debian.net/ on
the 30-day view, the spike we had 10 days ago due to Ruby team mass
changes is over and the base load seems to hover around 300 pipelines
per day.

We could further get this to drop to around 200 pipelines per day if
we promoted a culture that projects where Salsa CI is constantly
failing would turn it off, as CI as a regression testing system is
kind of moot if it constantly failing in a project and none of the
maintainers is fixing it, so it would be better to turn off to avoid
wasting resources. We have actually been doing a small email campaign
notifying Salsa projects that we have in Salsa CI stats seen been
running either scheduled pipelines that always fail or just very large
pipelines (build 1h+) with debian/latest failing asking them to turn
off Salsa CI until they can later re-enable it being in a passing
(green) state.


In the Salsa CI team we are also aware of some teams running excessive
amount of reverse builds on every git commit. We are soon announcing
https://salsa.debian.org/salsa-ci-team/pipeline/-/merge_requests/613
run reverse builds in a more controlled way hopefully helping to avoid
excessive number of builds.

- Otto

/Member of Salsa CI team, maintaining the pipeline code (but not in
Salsa Admin team and not maintaining the Salsa CI runners)

[toc] | [prev] | [next] | [standalone]


#13934

FromLucas Nussbaum <lucas@debian.org>
Date2025-09-05 15:40 +0200
Message-ID<LrEPv-d0xr-5@gated-at.bofh.it>
In reply to#13933
On 04/09/25 at 17:37 -0700, Otto Kekäläinen wrote:
> Looking at the recently published https://salsa-status.debian.net/ on
> the 30-day view, the spike we had 10 days ago due to Ruby team mass
> changes is over and the base load seems to hover around 300 pipelines
> per day.

I think that it would be better if we had infrastructure that encourages
teams and maintainers to perform work that improves standardization
across packages maintained by a team, which ultimately improves quality,
rather than send the message that such work is abnormal, and should be
spread over time sufficiently to limit impact.

Lucas

[toc] | [prev] | [next] | [standalone]


#13935

FromOtto Kekäläinen <otto@debian.org>
Date2025-09-05 19:40 +0200
Message-ID<LrIzL-d36u-7@gated-at.bofh.it>
In reply to#13934
> > Looking at the recently published https://salsa-status.debian.net/ on
> > the 30-day view, the spike we had 10 days ago due to Ruby team mass
> > changes is over and the base load seems to hover around 300 pipelines
> > per day.
>
> I think that it would be better if we had infrastructure that encourages
> teams and maintainers to perform work that improves standardization
> across packages maintained by a team, which ultimately improves quality,
> rather than send the message that such work is abnormal, and should be
> spread over time sufficiently to limit impact.

I didn't write in my message that it "should be spread over time", I
simply made a factual statement about how to read the statistics and
what conclusion to draw about the stats to estimate the base load in
the context of discussing Salsa CI runner load and potential need for
more hardware donations. I hope you do realize your statement "better
if .. rather than send the message" contains both your interpretation
and your reaction to your own interpretation. From my other messages
you surely know I am all in favor of simplifying and unifying
packaging practices across Debian, and having tooling to do it is
great.

[toc] | [prev] | [next] | [standalone]


#13936

FromLucas Nussbaum <lucas@debian.org>
Date2025-09-05 21:40 +0200
Message-ID<LrKrT-d4N9-1@gated-at.bofh.it>
In reply to#13935
On 05/09/25 at 10:14 -0700, Otto Kekäläinen wrote:
> > > Looking at the recently published https://salsa-status.debian.net/ on
> > > the 30-day view, the spike we had 10 days ago due to Ruby team mass
> > > changes is over and the base load seems to hover around 300 pipelines
> > > per day.
> >
> > I think that it would be better if we had infrastructure that encourages
> > teams and maintainers to perform work that improves standardization
> > across packages maintained by a team, which ultimately improves quality,
> > rather than send the message that such work is abnormal, and should be
> > spread over time sufficiently to limit impact.
> 
> I didn't write in my message that it "should be spread over time", I
> simply made a factual statement about how to read the statistics and
> what conclusion to draw about the stats to estimate the base load in
> the context of discussing Salsa CI runner load and potential need for
> more hardware donations. I hope you do realize your statement "better
> if .. rather than send the message" contains both your interpretation
> and your reaction to your own interpretation. From my other messages
> you surely know I am all in favor of simplifying and unifying
> packaging practices across Debian, and having tooling to do it is
> great.

I was referring to
https://lists.debian.org/debian-devel/2025/08/msg00547.html where you
asked:
> Could you perhaps limit your updates to maybe max 100 commits per day?

Note that the Ruby team is a nice case, with only 1249 packages: the
perl team (4089), python team (2858), go team (2441) and js team (1704)
have more.

Lucas

[toc] | [prev] | [next] | [standalone]


#13914

FromAlexander Wirt <formorer@debian.org>
Date2025-08-27 00:30 +0200
Message-ID<LoakW-aBBX-13@gated-at.bofh.it>
In reply to#13910

[Multipart message — attachments visible in raw view] — view raw

On Mon, Aug 25, 2025 at 09:05:14PM +0100, Ian Jackson wrote:
> Hi.
> 
> (I wasn't able to find another mailing list for this.  I found a bit
> of discussion on debian-devel about one specific overload incident.)
> 
> Since sid opened, people have naturally been hard at work.  Great!
> 
> But I notice that Salsa CI is now frequently very overloaded.  Even
> when it's not half a day behind (!), I often see that my package is
> able to run only about one job at a time.  I think this is slowing a
> lot of us down quite considerably.

I took some statistics. Those are only for our instance wide runner. 
Since beginning of the month we had roughly three times the amount of jobs. Today one project
(node-glob?!?) alone took 25% of all jobs (2700?!? I really have to check those numbers afer some sleep, 
but all others numbers look sane). salsa wasn't built for that. That is not only about more runners, the results and 
the artifacts have to be processed by godard too. So just waiving with _some_ money doesn't help. Sure, a complete overhaul would make sense
(2 recent servers, one just for postgresql, recent storage and so on). But that is not just some money, that would mean 
a lot money.



Today (UTC): 2025-08-26T00:00:00+00:00 .. 2025-08-27T00:00:00+00:00

Jobs processed today (instance-wide): 11142
Queue time today (s): avg=2008.9, median=185.1, p95=8466.9

Status distribution today:
  success: 9248
  failed: 1809
  canceled: 55
  running: 30

Top 5 projects today:
  js-team/node-glob: 2751
  aquilamacedo/pipeline: 305
  salsa-ci-team/pipeline: 293
  freexian-team/debusine: 288
  pochu/lts-pipeline: 237

Jobs by day (all scanned, not just today):
  2025-07-28: 3970
  2025-07-29: 3264
  2025-07-30: 3066
  2025-07-31: 5084
  2025-08-01: 4058
  2025-08-02: 5619
  2025-08-03: 3950
  2025-08-04: 3217
  2025-08-05: 3251
  2025-08-06: 3157
  2025-08-07: 4091
  2025-08-08: 3210
  2025-08-09: 4478
  2025-08-10: 6444
  2025-08-11: 7653
  2025-08-12: 8361
  2025-08-13: 5407
  2025-08-14: 4956
  2025-08-15: 5128
  2025-08-16: 5763
  2025-08-17: 5372
  2025-08-18: 5398
  2025-08-19: 5076
  2025-08-20: 5317
  2025-08-21: 5453
  2025-08-22: 5219
  2025-08-23: 5702
  2025-08-24: 6613
  2025-08-25: 14768
  2025-08-26: 11142

Alex

[toc] | [prev] | [next] | [standalone]


#13915

FromAntoine Le Gonidec <vv221@debian.org>
Date2025-08-27 01:30 +0200
Message-ID<LobgZ-aCe6-1@gated-at.bofh.it>
In reply to#13914

[Multipart message — attachments visible in raw view] — view raw

Le Wed, Aug 27, 2025 at 12:24:18AM +0200, Alexander Wirt a écrit :
> Today one project
> (node-glob?!?) alone took 25% of all jobs (2700?!? I really have to check those numbers afer some sleep, 
> but all others numbers look sane).

Probably not a mistake on your part. That package triggers jobs for each
of its reverse build-deps, and it has 1360 of these in unstable. Two
jobs for each of these and we get almost exactly to the 2700 number
you’re reporting.

[toc] | [prev] | [next] | [standalone]


#13917

FromSoren Stoutner <soren@debian.org>
Date2025-08-27 23:50 +0200
Message-ID<LowbL-aQSC-1@gated-at.bofh.it>
In reply to#13915

[Multipart message — attachments visible in raw view] — view raw

On Tuesday, August 26, 2025 6:43:30 PM Mountain Standard Time David Mulford 
wrote:
> All I know.. is ive been trying to unsubscribe from this group email for a
> while and its starting to drive me nuts. All the other debian group emails
> have auto unsubscribe working so maybe someone can fix that and take me off
> the mailing list. Greatly Appreciated.

Try using the web interface to unsubscribe:

https://lists.debian.org/debian-project/

If you are having issues unsubscribing, the most common problem is that you 
have subscribed with a different email address that is being forwarded to the 
email address you are reading.  If that is the case you can see the details by 
reviewing the email headers.  You need to unsubscribe with the same email 
address that is currently subscribed.

-- 
Soren Stoutner
soren@debian.org

[toc] | [prev] | [next] | [standalone]


#13918

FromDidier 'OdyX' Raboud <odyx@debian.org>
Date2025-08-28 09:30 +0200
Message-ID<LoFf3-aXGv-1@gated-at.bofh.it>
In reply to#13915

[Multipart message — attachments visible in raw view] — view raw

Le mercredi, 27 août 2025, 01.16:20 h heure d’été d’Europe centrale Antoine Le 
Gonidec a écrit :
> Le Wed, Aug 27, 2025 at 12:24:18AM +0200, Alexander Wirt a écrit :
> > Today one project
> > (node-glob?!?) alone took 25% of all jobs (2700?!? I really have to check
> > those numbers afer some sleep, but all others numbers look sane).
> 
> Probably not a mistake on your part. That package triggers jobs for each
> of its reverse build-deps, and it has 1360 of these in unstable. Two
> jobs for each of these and we get almost exactly to the 2700 number
> you’re reporting.

An important question to ask is: does it really need to proceed to trigger 
2000+ jobs _at every commit_ (= technically, at every branch ref update pushed 
to Salsa)?

Gitlab CI yaml syntax allows a lot of flexibility to determine when the CI 
gets ran automatically, manually, or not at all (see rules: in
https://salsa.debian.org/help/ci/yaml/_index.md#rules , it can also filter on 
commit message regexps, branch names, merge request conditions, etc), and I'd 
argue that the salsa-ci-team/pipeline should perhaps grow the capability (I 
have not checked, perhaps it exists already) to only launch on `debian/
latest`, or manually, or allow customizations.

Of course, having a full CI at every ref update can be useful (provided that 
someone actively cares when this fails), but there are many cases in which the 
pusher knows that their packaging change will not have an impact, and perhaps 
we should default to not having CI always be ran.

-- 
    OdyX

[toc] | [prev] | [next] | [standalone]


#13919 — salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload")

FromAntonio Terceiro <terceiro@debian.org>
Date2025-08-28 15:00 +0200
Subjectsalsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload")
Message-ID<LoKop-b17x-7@gated-at.bofh.it>
In reply to#13918

[Multipart message — attachments visible in raw view] — view raw

On Thu, Aug 28, 2025 at 09:00:41AM +0200, Didier 'OdyX' Raboud wrote:
> Le mercredi, 27 août 2025, 01.16:20 h heure d’été d’Europe centrale Antoine Le 
> Gonidec a écrit :
> > Le Wed, Aug 27, 2025 at 12:24:18AM +0200, Alexander Wirt a écrit :
> > > Today one project
> > > (node-glob?!?) alone took 25% of all jobs (2700?!? I really have to check
> > > those numbers afer some sleep, but all others numbers look sane).
> > 
> > Probably not a mistake on your part. That package triggers jobs for each
> > of its reverse build-deps, and it has 1360 of these in unstable. Two
> > jobs for each of these and we get almost exactly to the 2700 number
> > you’re reporting.
> 
> An important question to ask is: does it really need to proceed to trigger 
> 2000+ jobs _at every commit_ (= technically, at every branch ref update pushed 
> to Salsa)?

Not only it's not needed, it's also considered abuse of the salsa
infrastructure. I'm copying this message to all packages that seem to be
doing this (based on a search for debian/rdeps-ci.yml on codesearch).

Please note: if your package CI is triggering the rebuild of all reverse
dependencies on every single push, you are abusing the salsa
infrastructure. I do not imply malice, it's most probably just an
oversight.

I ask, however, that the maintainers of each package in Cc: to look into
it and take action to mitigate the abuse of shared resources (otherwise
I will have to). Maybe your packages does not have as many reverse
dependencies as node-glob, and it's fine, but please think about the
collective.

> Gitlab CI yaml syntax allows a lot of flexibility to determine when the CI 
> gets ran automatically, manually, or not at all (see rules: in
> https://salsa.debian.org/help/ci/yaml/_index.md#rules , it can also filter on 
> commit message regexps, branch names, merge request conditions, etc), and I'd 
> argue that the salsa-ci-team/pipeline should perhaps grow the capability (I 
> have not checked, perhaps it exists already) to only launch on `debian/
> latest`, or manually, or allow customizations.

There is a similar feature being developed for Salsa CI proper, as part
of a GSoC project. But in there, it needs to be explicitly triggered by
a person. I also suggested to the developer working on this to provide
an option of triggering the rebuild of a sample of the of reverse
dependencies, instead of all, and he said he'll work on it.

[toc] | [prev] | [next] | [standalone]


#13920 — Re: salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload")

FromJérémy Lal <kapouer@melix.org>
Date2025-08-28 15:30 +0200
SubjectRe: salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload")
Message-ID<LoKRr-b1yV-3@gated-at.bofh.it>
In reply to#13919

[Multipart message — attachments visible in raw view] — view raw

Le jeu. 28 août 2025 à 14:54, Antonio Terceiro <terceiro@debian.org> a
écrit :

> On Thu, Aug 28, 2025 at 09:00:41AM +0200, Didier 'OdyX' Raboud wrote:
> > Le mercredi, 27 août 2025, 01.16:20 h heure d’été d’Europe centrale
> Antoine Le
> > Gonidec a écrit :
> > > Le Wed, Aug 27, 2025 at 12:24:18AM +0200, Alexander Wirt a écrit :
> > > > Today one project
> > > > (node-glob?!?) alone took 25% of all jobs (2700?!? I really have to
> check
> > > > those numbers afer some sleep, but all others numbers look sane).
> > >
> > > Probably not a mistake on your part. That package triggers jobs for
> each
> > > of its reverse build-deps, and it has 1360 of these in unstable. Two
> > > jobs for each of these and we get almost exactly to the 2700 number
> > > you’re reporting.
> >
> > An important question to ask is: does it really need to proceed to
> trigger
> > 2000+ jobs _at every commit_ (= technically, at every branch ref update
> pushed
> > to Salsa)?
>
> Not only it's not needed, it's also considered abuse of the salsa
> infrastructure. I'm copying this message to all packages that seem to be
> doing this (based on a search for debian/rdeps-ci.yml on codesearch).
>
> Please note: if your package CI is triggering the rebuild of all reverse
> dependencies on every single push, you are abusing the salsa
> infrastructure. I do not imply malice, it's most probably just an
> oversight.
>
> I ask, however, that the maintainers of each package in Cc: to look into
> it and take action to mitigate the abuse of shared resources (otherwise
> I will have to). Maybe your packages does not have as many reverse
> dependencies as node-glob, and it's fine, but please think about the
> collective.
>


About node-glob:
another maintainer did setup the salsa-ci with all
reverse-build-dependencies tooling
installed, and we discussed about not allowing it by default... but I just
forgot to actually
change the default value.
I am very sorry about that, and it is done.

Jérémy

[toc] | [prev] | [next] | [standalone]


#13921 — Re: salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload")

FromMarco d'Itri <md@Linux.IT>
Date2025-08-28 15:30 +0200
SubjectRe: salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload")
Message-ID<LoKRr-b1yV-7@gated-at.bofh.it>
In reply to#13919

[Multipart message — attachments visible in raw view] — view raw

Can you tell me more about what is wrong with varnish and how to fix it?

I do not understand how it could stand out since it only has an handful 
of reverse dependencies and I rarely push to the repository before I am 
ready to upload.
Varnish is exactly the kind of package that can break its reverse 
dependencies because they are modules which often access its internals.
If using SALSA_CI_ENABLE_REVERSE_DEPENDENCY_BUILD is not acceptable then 
why is it available?

BTW, please note that after I took over the maintenance of Varnish 
I already cut a lot the CI time by disabling the very heavy internal 
test suite for the additional rebuilds:

https://salsa.debian.org/varnish-team/varnish/-/commit/f34d2997e77afc729b162041c0c511e5c5b41594

(And I still believe that nocheck should be set by default for these 
tests!)

-- 
ciao,
Marco

[toc] | [prev] | [next] | [standalone]


#13922 — Re: salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload")

FromAntonio Terceiro <terceiro@debian.org>
Date2025-08-28 17:00 +0200
SubjectRe: salsa: unconditional rebuild of all reverse dependencies on every push considered abuse (was "Salsa CI overload")
Message-ID<LoMgx-b2lP-1@gated-at.bofh.it>
In reply to#13921

[Multipart message — attachments visible in raw view] — view raw

Hi,

On Thu, Aug 28, 2025 at 03:15:46PM +0200, Marco d'Itri wrote:
> Can you tell me more about what is wrong with varnish and how to fix it?
>
> I do not understand how it could stand out since it only has an handful of
> reverse dependencies and I rarely push to the repository before I am ready
> to upload.
> Varnish is exactly the kind of package that can break its reverse
> dependencies because they are modules which often access its internals.
> If using SALSA_CI_ENABLE_REVERSE_DEPENDENCY_BUILD is not acceptable then why
> is it available?

In the varnish case, there are only 5 reverse build dependencies, so
it's mostly fine. The point is not having those rebuilds run on user
request, and not for every single push made to the git repository. 

As I mentioned, a better version os this job will be available from
salsa-ci proper at some point, and that will allow just this: reverse
dependency rebuilds are a manually triggered job that requires human
intervention to run. The problem with the version that is currently in
use by varnish and others, which trigger the rebuilds on every single
push.

For now, I *think* this should do it:

----------------8<----------------8<----------------8<----------------- 
diff --git i/debian/rdeps-ci.yml w/debian/rdeps-ci.yml
index 57b5670fa..ab2751669 100644
--- i/debian/rdeps-ci.yml
+++ w/debian/rdeps-ci.yml
@@ -1,5 +1,7 @@
 generate-config:
   image: $SALSA_CI_IMAGES_BASE
+  rules:
+    - when: manual
   needs:
     - job: aptly
       artifacts: true
----------------8<----------------8<----------------8<-----------------

Later when this feature is available in salsa-ci proper, then it should
be a matter of dropping that file entirely (and it's inclusion from
debian/salsa-ci.yml, and just relying on the better version from
salsa-ci.

> BTW, please note that after I took over the maintenance of Varnish I already
> cut a lot the CI time by disabling the very heavy internal test suite for
> the additional rebuilds:
> 
> https://salsa.debian.org/varnish-team/varnish/-/commit/f34d2997e77afc729b162041c0c511e5c5b41594
> 
> (And I still believe that nocheck should be set by default for these tests!)

While optimizations are always welcome, having a few jobs that can take
some time is fine from the infrastructure point of view, as several
others can still run at the time time. The salsa shared CI runner can
currently run up to 64 jobs concurrently.

The problem for everyone happens when a package triggers 2k jobs, as
node-glob did, because when that batch gets to the front of the queue,
it delays everyone else's jobs for quite some time (what prompted this
thread in the first place). Having this run manually on maintainer
request, once in a while, is also OK, but not unconditionally on every
push.

[toc] | [prev] | [next] | [standalone]


#13916

FromIan Jackson <ijackson@chiark.greenend.org.uk>
Date2025-08-27 11:40 +0200
Message-ID<LokNk-aIHe-17@gated-at.bofh.it>
In reply to#13914
Alexander Wirt writes ("Re: Salsa CI overload"):
> I took some statistics.

Thanks.  This is extremely illuminating.  Earlier I wrote this:

> > Perhaps some of our jobs are wasteful, but I think that overall,
> > adding capacity would be well justified.>

but these startling numbers suggest that I was wrong.

Is there any mechanism we could use for allocating capacity more
fairly?


If not, I think we must we rely on ad-hoc response to overload events,
and social pressure.  That is much less comfortable than an automatic
resource allocation system, as the latter (however imperfect) is
objective, and the message is delivered by a computer rather than an
exasperated human.

Anyway, in that case we don't want that to be a burden just on the
salsa admins, so it would be good to be able to see the numbers
publicly (at least, most of the numbers).

It should be noted of course that number of jobs is perhaps a poor
proxy for resource use, given that jobs can be of widely diffeerent
sizes.  But many jobs are so small that the overhead dominates.


> Top 5 projects today:
>   js-team/node-glob: 2751

I went and looked at this package.  I found that:

The pipeline has been failing on the main branch for some days, so it
doesn't seem like CI is being used to prevent landing broken changes
(or, it isn't effective in doing so).

I looked at one of the pipelines via the web UI and I didn't see the
thousands of jobs I was expecting based on the number above; I don't
know why.  Maybe the web UI just can't show so many.  I looked at two
arbitrary successful jobs.  They each took about 3 minutes.  Our
"docker-machine" executor container takes 1.5 minutes to set up per
job.  With small jobs like this, that time must be considered wasteful
overhead.  That's tolerable if there aren't very many.

But in this case *this project alone* has *wasted* 60 thread-hours'
worth of some kind of capacity, in a 24 hour period.  This is quite
extraordinary.

The remaining ones in the top 5, using 10x fewer jobs each, mostly
seem like repos we might expect to be under intense development.
Probably the computer resource usage there is more proportionate to
the humanm effort input.

Ian.

-- 
Ian Jackson <ijackson@chiark.greenend.org.uk>   These opinions are my own.  

Pronouns: they/he.  If I emailed you from @fyvzl.net or @evade.org.uk,
that is a private address which bypasses my fierce spamfilter.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.project


csiph-web