Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.project > #13699 > unrolled thread
| Started by | "M. Zhou" <lumin@debian.org> |
|---|---|
| First post | 2025-02-02 14:10 +0100 |
| Last post | 2025-04-21 02:30 +0200 |
| Articles | 9 — 6 participants |
Back to article view | Back to linux.debian.project
[draft] need your help on the AI-DFSG general resolution prepration "M. Zhou" <lumin@debian.org> - 2025-02-02 14:10 +0100
Re: [draft] need your help on the AI-DFSG general resolution prepration Holger Levsen <holger@layer-acht.org> - 2025-02-03 16:50 +0100
Re: [draft] need your help on the AI-DFSG general resolution prepration Stefano Zacchiroli <zack@debian.org> - 2025-02-04 15:40 +0100
Re: [draft] need your help on the AI-DFSG general resolution prepration Jacinto Dávila <jacinto.davila@gmail.com> - 2025-02-04 21:20 +0100
Re: [draft] need your help on the AI-DFSG general resolution prepration Ross Vandegrift <rvandegrift@debian.org> - 2025-02-07 22:30 +0100
when will we rebuild AI-based software from sources/datasets? Stefano Zacchiroli <zack@debian.org> - 2025-02-08 15:10 +0100
Re: when will we rebuild AI-based software from sources/datasets? Jonas Smedegaard <dr@jones.dk> - 2025-02-08 16:40 +0100
Re: [draft] need your help on the AI-DFSG general resolution prepration "M. Zhou" <lumin@debian.org> - 2025-04-13 09:10 +0200
Re: [draft] need your help on the AI-DFSG general resolution prepration "M. Zhou" <lumin@debian.org> - 2025-04-21 02:30 +0200
| From | "M. Zhou" <lumin@debian.org> |
|---|---|
| Date | 2025-02-02 14:10 +0100 |
| Subject | [draft] need your help on the AI-DFSG general resolution prepration |
| Message-ID | <KbBbr-dJNf-1@gated-at.bofh.it> |
Hi all, I heard that people were looking for me during FOSDEM. I spent a couple of hours and finally get something draft-ish for the previously mentioned general resolution on the software freedom interpolation with respect to AI software. https://salsa.debian.org/lumin/gr-ai-dfsg (I turned the issues on. Feel free to open issues there) This is an early draft. Before really posting to -vote, I need your help on the following aspect: (1) do you know any important but missing reference materials? (2) are the options clear enough for vote? Considering lots of the readers may not be faimiliar with how AI is created. I tried to explain it, as well as the implication if some components are missing. (3) is there anything unclear or ambiguous in the text for backgrounds and options? (4) is there anything else that should be added to the text? (5) what is the actionable outcome of this generaal resolution? (6) is a neutral tone necessary for a proposal? I have a clear tendency throughout the texts. (7) I have not yet asked ftp-master on their opinion. According to https://www.debian.org/vote/howto_proposal , there is a template https://www.debian.org/vote/sample_vote.template but I don't understand this XML dialect. How to use this XML file?
[toc] | [next] | [standalone]
| From | Holger Levsen <holger@layer-acht.org> |
|---|---|
| Date | 2025-02-03 16:50 +0100 |
| Subject | Re: [draft] need your help on the AI-DFSG general resolution prepration |
| Message-ID | <Kc6RY-e6PG-5@gated-at.bofh.it> |
| In reply to | #13699 |
[Multipart message — attachments visible in raw view] — view raw
hi, about https://salsa.debian.org/lumin/gr-ai-dfsg/-/blob/main/README.txt On Sun, Feb 02, 2025 at 12:56:59AM -0500, M. Zhou wrote: > (2) are the options clear enough for vote? Considering lots of the readers may > not be faimiliar with how AI is created. I tried to explain it, as well as > the implication if some components are missing. this explaination text is surely useful/well ment but also makes it somewhat hard to see where the GR text starts. > (3) is there anything unclear or ambiguous in the text for backgrounds and options? > (4) is there anything else that should be added to the text? > (5) what is the actionable outcome of this generaal resolution? > (6) is a neutral tone necessary for a proposal? I have a clear > tendency throughout the texts. as I understand your text and the GR process, you presented two (or three) GR options in one, which is a bit unusual, so maybe it will help to clearly to split this in (3 files) a.) GR background, explainations about AI b.) GR option 1 (called "Proposal A" in your text currently. c.) GR option 2 (called "Proposal B" in your text currently. Your Proposal C is not needed because "further discussion" is always an option in our GRs. Also you don't need to be neutral, though of course you can and maybe you should try. But my point is: anyone else can also propose GR options with a more or less neutral tone, so you can definitly be enthusiastic about your proposal! -- cheers, Holger ⢀⣴⠾⠻⢶⣦⠀ ⣾⠁⢠⠒⠀⣿⡁ holger@(debian|reproducible-builds|layer-acht).org ⢿⡄⠘⠷⠚⠋⠀ OpenPGP: B8BF54137B09D35CF026FE9D 091AB856069AAA1C ⠈⠳⣄ I have a joke about trickle down economics. 99% of you won’t ever get it.
[toc] | [prev] | [next] | [standalone]
| From | Stefano Zacchiroli <zack@debian.org> |
|---|---|
| Date | 2025-02-04 15:40 +0100 |
| Message-ID | <KcsfL-elQY-1@gated-at.bofh.it> |
| In reply to | #13699 |
[Multipart message — attachments visible in raw view] — view raw
Hello lumin, and thanks a lot for this work. On Sun, Feb 02, 2025 at 12:56:59AM -0500, M. Zhou wrote: > https://salsa.debian.org/lumin/gr-ai-dfsg > (I turned the issues on. Feel free to open issues there) before discussing specific details, my first reaction (and associated question) is that the text is *very* long for a vote. Do you plan to propose that full text as what will be voted on? In general, I think it's better to keep a GR text short and to the point, and post contextual material somewhere else. That somewhere else can be the ensuing discussions on -project/-vote, possibly summarized in periodic messages, but can also reside in dedicated web pages like your repo above. I'm not saying this not only for "voting overhead", in particular for voters who will discover this only at the time of call for votes, but also because the longer the text is, higher are the chances that people will disagree with at least some of the additional information material, and will hence decide to vote against the GR for side reasons. FWIW, I also agree with Holger that you should focus on the GR option that you *want* to see adopted. Leave to others the burden of proposing alternative options to be added to the ballot. Cheers -- Stefano Zacchiroli . zack@upsilon.cc . https://upsilon.cc/zack _. ^ ._ Full professor of Computer Science o o o \/|V|\/ Télécom Paris, Polytechnic Institute of Paris o o o </> <\> Co-founder & CSO Software Heritage o o o o /\|^|/\ Mastodon: https://mastodon.xyz/@zacchiro '" V "'
[toc] | [prev] | [next] | [standalone]
| From | Jacinto Dávila <jacinto.davila@gmail.com> |
|---|---|
| Date | 2025-02-04 21:20 +0100 |
| Message-ID | <Kcx5L-ep3I-1@gated-at.bofh.it> |
| In reply to | #13699 |
[Multipart message — attachments visible in raw view] — view raw
(1) do you know any important but missing reference materials? You may want to include references to currents cases in court, like: https://www.npr.org/2025/01/14/nx-s1-5258952/new-york-times-openai-microsoft Maybe not that particular one, but something to the effect. By supporting proposal B: "Toxic Candy" is free software, I believe one would be taking side on those disputes, against creators that believe that their work is being used as training data and has not been dutifully honored. Otherwise, the proposals look impeccable. Thank you On Sun, 2 Feb 2025 at 01:57, M. Zhou <lumin@debian.org> wrote: > Hi all, > > I heard that people were looking for me during FOSDEM. > > I spent a couple of hours and finally get something draft-ish > for the previously mentioned general resolution on the software > freedom interpolation with respect to AI software. > > https://salsa.debian.org/lumin/gr-ai-dfsg > (I turned the issues on. Feel free to open issues there) > > This is an early draft. Before really posting to -vote, I need your > help on the following aspect: > > (1) do you know any important but missing reference materials? > > (2) are the options clear enough for vote? Considering lots of the readers > may > not be faimiliar with how AI is created. I tried to explain it, as well as > the implication if some components are missing. > > (3) is there anything unclear or ambiguous in the text for backgrounds and > options? > > (4) is there anything else that should be added to the text? > > (5) what is the actionable outcome of this generaal resolution? > > (6) is a neutral tone necessary for a proposal? I have a clear > tendency throughout the texts. > > (7) I have not yet asked ftp-master on their opinion. > > > According to https://www.debian.org/vote/howto_proposal , > there is a template https://www.debian.org/vote/sample_vote.template > but I don't understand this XML dialect. How to use this XML file? > > -- Jacinto A. Dávila Quintero http://webdelprofesor.ula.ve/ingenieria/jacinto
[toc] | [prev] | [next] | [standalone]
| From | Ross Vandegrift <rvandegrift@debian.org> |
|---|---|
| Date | 2025-02-07 22:30 +0100 |
| Message-ID | <KdE5c-fa4i-27@gated-at.bofh.it> |
| In reply to | #13699 |
On Sun, Feb 02, 2025 at 12:56:59AM -0500, M. Zhou wrote: > (5) what is the actionable outcome of this generaal resolution? Is this actually unclear? The linked text claims this issue is urgent - usually, something is urgent because some issue demands action. > (6) is a neutral tone necessary for a proposal? I have a clear tendency > throughout the texts. I think it's okay to reveal a preference, but the "toxic candy" term is rather unfortunate. You've defined the debate with terms that dismiss the opposing viewpoint. Could you come up with a more neutral, descriptive name? Ross
[toc] | [prev] | [next] | [standalone]
| From | Stefano Zacchiroli <zack@debian.org> |
|---|---|
| Date | 2025-02-08 15:10 +0100 |
| Subject | when will we rebuild AI-based software from sources/datasets? |
| Message-ID | <KdTGV-fk2a-5@gated-at.bofh.it> |
| In reply to | #13699 |
Hello Mo, all, I've now read through the full GR text and commentary.
I've a bunch of comments, but I'll post them separately (and/or in MR).
In this mail I'd like to focus on one important aspect related to
implications:
On Sun, Feb 02, 2025 at 12:56:59AM -0500, M. Zhou wrote:
> (2) are the options clear enough for vote? Considering lots of the readers may
> not be faimiliar with how AI is created. I tried to explain it, as well as
> the implication if some components are missing.
I'd like to understand, in case option A passes, when and how Debian
will rebuild AI-models that are included in some free software that is
in the archive from its "source" (which will include the full training
dataset, as per option A indeed).
Concrete examples
-----------------
Let's take two simple examples that, size-wise, could fit in the Debian
archive together with training pipelines and datasets.
First, let's take a "small" image classification model that one day
might be included in Digikam or similar free software. Let's say the
trained model is ~1 GiB (like Moondream [1] today) and that the training
dataset is ~10 GiB (I've no idea if the Moondrean training dataset is
open data and I'm probably being very conservative with its size here;
just assume it is correct for now).
[1]: https://ollama.com/library/moondream:v2/blobs/e554c6b9de01
For a second, even smaller example, let's consider gnubg (GNU
backgammon) that contains today, in the Debian archive, a trained neural
network [2] of less than 1 MiB. Its training data is *not* in the
archive, but is available online without a license (AFAICT) [3] and
weights about ~80 MiB. The training code is available as well [4], even
though still in Python 2.
[2]: https://git.savannah.gnu.org/cgit/gnubg.git/log/gnubg.weights
[3]: https://alpha.gnu.org/gnu/gnubg/nn-training
[4]: https://git.savannah.gnu.org/cgit/gnubg/gnubg-nn.git
What do we put where?
---------------------
Regarding source packages, I suspect that most or our upstream authors
that will end up using free AI will *not* include training datasets in
distribution tarballs or Git repositories of the main software. So what
will we do downstream? Do we repack source packages to include the
training datasets? Do we create *separate* source packages for the
training datasets? Do we create a separate (ftp? git-annex? git-lfs?)
hosting place where to host large training datasets to avoid exploding
mirror sizes? Do we simply refer to external hosting places that are not
under Debian control?
Of course this would be irrelevant problem in the gnubg case (we can
just store everything in the source package), but it will become more
significant in the Digikam case, and the amount of those cases will
probably increase over time..
(Note that I do not have definitive answers to any of the questions in
this email. And also that I'm *not* raising them as counter arguments to
option A, which is my favorite one at the moment. I just want to make
sure that we have a rough idea of how we will in practice *implement*
option A in a way that fits Debian processes, rather than thinking about
them only after the vote. We have been there for a number of GRs in the
past, and it has not been fun.)
Regarding binary packages, the question applies for large trained AI
models too. On this front, we can rely entirely on upstream software to
"unbundle" trained datasets from their software, so that it is
downloaded on the fly on user machines and never enters the Debian
archive. Based on previous answers, I suspect this might be what Mo has
in mind. But I don't find it very satisfactory for a number of reasons:
it will not be universal, we might end up having to host some of the
large models ourselves at some point, and even when it is handled
upstream we will leave users on their own in terms of installation
risks, etc. (Yes, this is not a new problem, and applies to other
software that automatically download plugins and whatnot, but I still
don't like it.)
When do we retrain?
-------------------
The most difficult question for me is: when do we retrain AI models
(shipped in Debian binary packages) from their training datasets
(shipped in source packages)?
In some cases, as pointed out in the GR commentary, it will be
computationally impossible for Debian to do so. We can mostly ignore
these cases, but it hence begs the question: do we want to ship in
Debian trained AI models that *allegedly* have all their training
dataset and pipeline available under free licenses, if we cannot
verify/rebuild them ourselves? (If this smells like the XZ utils hack to
you, you're not alone!)
Let's focus now on the cases that *could* be retrained by Debian,
possibly after buying a dozen on GPUs to put on dedicated buildds.
Potential answers on when to retrain are:
- We never retrain.
We are now back to the already discussed XZ utils smell. I don't think
this would be wise/acceptable in case where it is feasible for us to
retrain.
- We retrain at every package build.
Why not, but it will be quite expensive. (Add here your favorite
environmental concerns.) It will also require some dedicated
scheduling to separate packages that need GPUs to build from others.
This will probably result in a natural separation between source
packages containing training datasets, that can be tagged as needing
GPUs to build, and result in binary packages that are dependencies for
the final software installed by users. Seems appealing to me.
- We retrain every now and then (e.g., once per release).
Compromise situation, between the previous two. Can be analogous to
bootstraping compilers, which we don't do systematically, but can be
done by motivated developers and porters. If we go down this path, we
probably want to standardize some debian/rules target ("bootstrap"?)
that recreate trained models from sources and can then be used to
update source packages that ship them.
Reproducible builds
-------------------
Side, but important consideration: retraining will in most cases not be
bitwise reproducible, as pointed out already in the GR commentary. The
practical consequence for Debian is that, all packages that will end up
containing the logic for retraining AI models will remain non-bitwise
reproducible for the foreseeable future. (Which is an additional good
argument for clearly separating those packages from others.)
Have I missed any other specific Debian process that will be impacted?
Cheers
--
Stefano Zacchiroli . zack@upsilon.cc . https://upsilon.cc/zack _. ^ ._
Full professor of Computer Science o o o \/|V|\/
Télécom Paris, Polytechnic Institute of Paris o o o </> <\>
Co-founder & CSO Software Heritage o o o o /\|^|/\
Mastodon: https://mastodon.xyz/@zacchiro '" V "'
[toc] | [prev] | [next] | [standalone]
| From | Jonas Smedegaard <dr@jones.dk> |
|---|---|
| Date | 2025-02-08 16:40 +0100 |
| Subject | Re: when will we rebuild AI-based software from sources/datasets? |
| Message-ID | <KdV61-fkMp-11@gated-at.bofh.it> |
| In reply to | #13730 |
[Multipart message — attachments visible in raw view] — view raw
Quoting Stefano Zacchiroli (2025-02-08 14:57:18) > Concrete examples > ----------------- > > Let's take two simple examples that, size-wise, could fit in the Debian > archive together with training pipelines and datasets. > > First, let's take a "small" image classification model that one day > might be included in Digikam or similar free software. Let's say the > trained model is ~1 GiB (like Moondream [1] today) and that the training > dataset is ~10 GiB (I've no idea if the Moondrean training dataset is > open data and I'm probably being very conservative with its size here; > just assume it is correct for now). > > [1]: https://ollama.com/library/moondream:v2/blobs/e554c6b9de01 > > For a second, even smaller example, let's consider gnubg (GNU > backgammon) that contains today, in the Debian archive, a trained neural > network [2] of less than 1 MiB. Its training data is *not* in the > archive, but is available online without a license (AFAICT) [3] and > weights about ~80 MiB. The training code is available as well [4], even > though still in Python 2. > > [2]: https://git.savannah.gnu.org/cgit/gnubg.git/log/gnubg.weights > [3]: https://alpha.gnu.org/gnu/gnubg/nn-training > [4]: https://git.savannah.gnu.org/cgit/gnubg/gnubg-nn.git Another example of seemingly "small" training data is Tesseract. DFSG of training data is tracked at https://bugs.debian.org/699609 with an optimistic view. A more pessimistic view seems indicated by an upstream mention of the training data being "all the WWW" and other comments mention the involvement of non-free fonts: https://github.com/tesseract-ocr/tesseract/issues/654#issuecomment-274574951 - Jonas -- * Jonas Smedegaard - idealist & Internet-arkitekt * Tlf.: +45 40843136 Website: http://dr.jones.dk/ * Sponsorship: https://ko-fi.com/drjones [x] quote me freely [ ] ask before reusing [ ] keep private
[toc] | [prev] | [next] | [standalone]
| From | "M. Zhou" <lumin@debian.org> |
|---|---|
| Date | 2025-04-13 09:10 +0200 |
| Message-ID | <KAZDz-dazz-5@gated-at.bofh.it> |
| In reply to | #13699 |
On Sun, 2025-02-02 at 00:56 -0500, M. Zhou wrote: > > https://salsa.debian.org/lumin/gr-ai-dfsg I overhauled the draft, and have split lots of rationales and background information to separate files as appendix. The current status looks like a release candidate to me. I'll further simplify the text, deleting more sentences as appropriate, and rephrasing redundant language, in order to further shrink its length while not losing key information. That means the remaining work is something like language polishing. When I'm feeling comfortable with the text length, the language, etc., I'll sign and post to -vote. That moment should be very soon.
[toc] | [prev] | [next] | [standalone]
| From | "M. Zhou" <lumin@debian.org> |
|---|---|
| Date | 2025-04-21 02:30 +0200 |
| Message-ID | <KDNcR-f4gx-3@gated-at.bofh.it> |
| In reply to | #13796 |
On Sun, 2025-04-13 at 00:57 -0400, M. Zhou wrote: > On Sun, 2025-02-02 at 00:56 -0500, M. Zhou wrote: > > > > https://salsa.debian.org/lumin/gr-ai-dfsg I signed and posted the proposal there: https://lists.debian.org/debian-vote/2025/04/msg00101.html This is my first time to create a GR. And I just learned from https://www.debian.org/vote/howto_follow that we need 5 sponsors before it is open for discussion. Much appreciated if folks are willing to second/sponsor or propose something different to push this forward!
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.project
csiph-web