Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]
Groups > linux.debian.project > #13644 > unrolled thread
| Started by | Gerardo Ballabio <gerardo.ballabio@gmail.com> |
|---|---|
| First post | 2024-11-04 12:00 +0100 |
| Last post | 2025-01-29 16:50 +0100 |
| Articles | 12 — 8 participants |
Back to article view | Back to linux.debian.project
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Gerardo Ballabio <gerardo.ballabio@gmail.com> - 2024-11-04 12:00 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Mo Zhou <lumin@debian.org> - 2024-11-05 03:20 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Sean Whitton <spwhitton@spwhitton.name> - 2025-01-25 13:20 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 "M. Zhou" <lumin@debian.org> - 2025-01-25 16:30 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Sean Whitton <spwhitton@spwhitton.name> - 2025-01-25 16:40 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 "M. Zhou" <lumin@debian.org> - 2025-01-25 17:20 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Sam Johnston <samj@samj.net> - 2025-01-25 17:20 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 "M. Zhou" <lumin@debian.org> - 2025-01-25 18:00 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Stefano Zacchiroli <zack@debian.org> - 2025-01-25 18:40 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Ilu <ilulu@gmx.net> - 2025-01-25 18:50 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Sam Johnston <samj@samj.net> - 2025-01-25 19:00 +0100
Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 nick black <dankamongmen@gmail.com> - 2025-01-29 16:50 +0100
| From | Gerardo Ballabio <gerardo.ballabio@gmail.com> |
|---|---|
| Date | 2024-11-04 12:00 +0100 |
| Subject | Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 |
| Message-ID | <JF2Yp-6y9e-11@gated-at.bofh.it> |
The OSAID 1.0 has now been released (with no modifications from the RC2). Are we still going to take any actions or will we let this go? Gerardo
[toc] | [next] | [standalone]
| From | Mo Zhou <lumin@debian.org> |
|---|---|
| Date | 2024-11-05 03:20 +0100 |
| Message-ID | <JFhkJ-6Hae-3@gated-at.bofh.it> |
| In reply to | #13644 |
I'm planning to draft a GR for this, but that is only going to happen after I get through some busy weeks. On 11/4/24 02:22, Gerardo Ballabio wrote: > The OSAID 1.0 has now been released (with no modifications from the RC2). > Are we still going to take any actions or will we let this go? > > Gerardo >
[toc] | [prev] | [next] | [standalone]
| From | Sean Whitton <spwhitton@spwhitton.name> |
|---|---|
| Date | 2025-01-25 13:20 +0100 |
| Message-ID | <K8NiN-bJXM-5@gated-at.bofh.it> |
| In reply to | #13645 |
[Multipart message — attachments visible in raw view] — view raw
Hello Lumin, On Mon 04 Nov 2024 at 05:52pm -08, Mo Zhou wrote: > I'm planning to draft a GR for this, but that is only going to happen after I > get through some busy weeks. Wondered if you'd had another chance to look at this. -- Sean Whitton
[toc] | [prev] | [next] | [standalone]
| From | "M. Zhou" <lumin@debian.org> |
|---|---|
| Date | 2025-01-25 16:30 +0100 |
| Message-ID | <K8QgG-bLNm-19@gated-at.bofh.it> |
| In reply to | #13671 |
On Sat, 2025-01-25 at 12:09 +0000, Sean Whitton wrote: > > Wondered if you'd had another chance to look at this. > Ummm... You know what may happen when there is no deadline. Did you ping this because there are some thoughts from the policy side? Or just curious?
[toc] | [prev] | [next] | [standalone]
| From | Sean Whitton <spwhitton@spwhitton.name> |
|---|---|
| Date | 2025-01-25 16:40 +0100 |
| Message-ID | <K8Qql-bLQG-29@gated-at.bofh.it> |
| In reply to | #13672 |
[Multipart message — attachments visible in raw view] — view raw
Hello, On Sat 25 Jan 2025 at 10:21am -05, M. Zhou wrote: > On Sat, 2025-01-25 at 12:09 +0000, Sean Whitton wrote: >> >> Wondered if you'd had another chance to look at this. >> > > Ummm... You know what may happen when there is no deadline. > > Did you ping this because there are some thoughts from the policy side? > Or just curious? Nothing to do with Debian Policy, no. I'm just interested in your thoughts on the matter. -- Sean Whitton
[toc] | [prev] | [next] | [standalone]
| From | "M. Zhou" <lumin@debian.org> |
|---|---|
| Date | 2025-01-25 17:20 +0100 |
| Message-ID | <K8R33-bMjQ-9@gated-at.bofh.it> |
| In reply to | #13673 |
On Sat, 2025-01-25 at 15:29 +0000, Sean Whitton wrote: > > Nothing to do with Debian Policy, no. > > I'm just interested in your thoughts on the matter. If we look at the history. Richard Stallman started the GNU project at a time point where individual developers and free software communities can really create software, and define what "free software" is. If "free software" was defined at a time point where individuals and free software communities could not write software on their own, that definition may turn into a Utopia declaration. Similarly, we are not yet in an era where individuals and free software communities can easily create original, useful, and competitive AI from scratch (to ensure, e.g., reproducibility and trustworthiness). Particularly for language models. Currently, the model creators can do any decision regarding their decisions, regardless of how we define terminologies. That said, with the advancements of research, and the whole communities appreciation and recognition on the value of free and open source, that day will come sooner or later for language models. Creating some simpler vision or language AIs by individual is already possible. So the hard deadline for this matter is the day when people can freely create and publish AIs. We will be too late at that time point. While I disagree on OSAID's requirement on training data, it is directing at the right direction. I can see OSI is making compromises due to the dilemma that they need something to apply in practice and take action, while free/open source communities can not yet create and own language models. What we have seen from OSI is possibly the best they can do at the current time point. >From the Debian side, my concerns are unchanged. The OSAID does not guarantee freedom to our user. I think FSF is keeping an eye on this. Let me try to make a draft within Feburary.
[toc] | [prev] | [next] | [standalone]
| From | Sam Johnston <samj@samj.net> |
|---|---|
| Date | 2025-01-25 17:20 +0100 |
| Message-ID | <K8R33-bMjQ-3@gated-at.bofh.it> |
| In reply to | #13672 |
On Sat, 25 Jan 2025 at 16:24, M. Zhou <lumin@debian.org> wrote: > > On Sat, 2025-01-25 at 12:09 +0000, Sean Whitton wrote: > > > > Wondered if you'd had another chance to look at this. > > Ummm... You know what may happen when there is no deadline. The best time to do this was last year around the OSAID 1.0 release. The next best time is now. Do you need our help? I'm working on an article about how the chickens have come home to roost with the VLC demo at CES 2025. With VLC advertising and users now expecting real-time AI subtitling that "appears to be built directly into the VLC app"[1], we have a situation where VLC is considered Open Source by the OSD, but NOT according to the OSAID and OSI leadership[2] because of Whisper being embedded. With more and more software being written by and incorporating AI, this situation is untenable. Distros like Debian would have to lobotomise popular apps like VLC, or accept more binary blobs. The OSI also just released a whitepaper[3] that further deliberately obfuscates the issue, prompting me to post this: The Open Source Initiative (OSI) goes to the effort of defining four classes of data *source* (hence the term!) in their Open Source AI Definition (OSAID) FAQ and again in the Open Future Foundation’s name in this new paper, only to then accept ANY of them… or NONE at all: - OPEN data under open licenses, which is the ONLY class that has any role in Open Source AI - PUBLIC data like Common Crawl Foundation dumps of the Internet, which are routinely ab/used without creators’ consent - OBTAINABLE data “including for a fee” like The New York Times articles and Adobe/Getty Images stock photos, which are guaranteed to get end users (but not necessarily vendors given limited liability clauses) sued - UNSHAREABLE NONPUBLIC data that obviously has no place in Open Source, like Facebook & Instagram feeds With the meaning of Open Source AI being defined solely by the LOWEST bar — no data delivered at all (which is allowed under the OSAID) — why bother with the smokescreen if not to deliberately deceive us users? An honest FAQ entry would have read like this: What kind of data should be required in the Open Source AI Definition? None. 1. https://hackaday.com/2025/01/15/floss-weekly-episode-816-open-source-ai/ 2. https://www.theverge.com/2025/1/9/24339817/vlc-player-automatic-ai-subtitling-translation 3. https://openfuture.eu/publication/data-governance-in-open-source-ai/
[toc] | [prev] | [next] | [standalone]
| From | "M. Zhou" <lumin@debian.org> |
|---|---|
| Date | 2025-01-25 18:00 +0100 |
| Message-ID | <K8RFM-bMxL-19@gated-at.bofh.it> |
| In reply to | #13674 |
On Sat, 2025-01-25 at 17:08 +0100, Sam Johnston wrote: > > The best time to do this was last year around the OSAID 1.0 release. > The next best time is now. Do you need our help? > I lean towards making things simpler. Yes I disagree with OSI's decision on OSAID, and the definition does not guarantee freedom at all. But a bold move towards picking a fight against OSI on this matter through Debian General Resolution sounds terrible and reckless to me. I'll focus on a simpler topic for the GR: "how does Debian community interpret DFSG and software freedom against the AI model and software?" I'll draft from a pure technical point of view. Neutral to individuals and organizations, without commenting on how others think and do. In that case it is as simply as elaborating the "toxic candy" case, and analyzing the OSAID's implication from a technical point of view. In that sense, things will be more constructive and doable. FSF will also able to learn from Debian's GR discussion. I'll put my limited energy on this matter towards such direction.
[toc] | [prev] | [next] | [standalone]
| From | Stefano Zacchiroli <zack@debian.org> |
|---|---|
| Date | 2025-01-25 18:40 +0100 |
| Message-ID | <K8Sit-bN0E-3@gated-at.bofh.it> |
| In reply to | #13676 |
[Multipart message — attachments visible in raw view] — view raw
On Sat, Jan 25, 2025 at 11:37:48AM -0500, M. Zhou wrote: > I'll focus on a simpler topic for the GR: > > "how does Debian community interpret DFSG and software freedom > against the AI model and software?" Clarifying this officially seem indeed very useful, both for Debian and for the free software world at large, due to how relevant Debian is as one of the important gatekeepers of what is free-software-ok and what is not. Many people out there, and possibly even within the project, are assuming that the current Debian AI Policy is an official project position, whereas it is not. It's important to ratify it somehow. Since the last time this was discussed, did anyone reach out to FTP master to understand *if* they currently have already an official answer to this question or not? I'm Cc:-ing them explicitly with this message, just in case. Aside from the potential constitutional implications of this (last time it felt like I was the only one worrying about this aspect, so I'm dropping it), it would still be interesting to know their take. And it's quite possible they would actively welcome/encourage an official project-wide vote on this matter. Cheers -- Stefano Zacchiroli . zack@upsilon.cc . https://upsilon.cc/zack _. ^ ._ Full professor of Computer Science o o o \/|V|\/ Télécom Paris, Polytechnic Institute of Paris o o o </> <\> Co-founder & CSO Software Heritage o o o o /\|^|/\ Mastodon: https://mastodon.xyz/@zacchiro '" V "'
[toc] | [prev] | [next] | [standalone]
| From | Ilu <ilulu@gmx.net> |
|---|---|
| Date | 2025-01-25 18:50 +0100 |
| Message-ID | <K8Ss9-bN3W-1@gated-at.bofh.it> |
| In reply to | #13676 |
When this discussion came up I immeditely thought of text-to-speech projects like piper, using AI generated voices derived from real persons voice data. The voice is a very specific attribute of every human. Its part of attributes that define that persons humanity. Same goes for the face, the fingerprint, the eyes, the movements, the genome. All these data are very specific and in summa defines a person (although there are probably more factors, just these came to mind). It also is already used to identify a person. If training data involves the core of a humans personhood - as mentioned above - it cannot be open-sourced or otherwise free. Consent does not matter. It belongs to that person and nobody else, period. I deeply believe that there are borders we are not allowed to cross, no matter how noble the cause. Do we still want to positively distinguish AI projects whose code, parameters, weights and adjustments are free? Yes I do and IMHO OSI has the same goal. That's why I think that OSAID is on the right track. They could have worded things better and they could have more precisely distinguished different types of training data but in light of time contraints (EU AI Act) I think they did the best they could. Will Debian accept a GR that requires all training data to be free, including training data that belongs to the core of human dignity? That would be disturbing. And in fact practically lobotomize good projects. Am 25.01.25 um 17:37 schrieb M. Zhou: > On Sat, 2025-01-25 at 17:08 +0100, Sam Johnston wrote: >> >> The best time to do this was last year around the OSAID 1.0 release. >> The next best time is now. Do you need our help? >> > > I lean towards making things simpler. > > Yes I disagree with OSI's decision on OSAID, and the definition > does not guarantee freedom at all. But a bold move towards picking > a fight against OSI on this matter through Debian General Resolution > sounds terrible and reckless to me. > > I'll focus on a simpler topic for the GR: > > "how does Debian community interpret DFSG and software freedom > against the AI model and software?" > > I'll draft from a pure technical point of view. Neutral to > individuals and organizations, without commenting on how others > think and do. In that case it is as simply as elaborating the > "toxic candy" case, and analyzing the OSAID's implication from > a technical point of view. > > In that sense, things will be more constructive and doable. > FSF will also able to learn from Debian's GR discussion. > I'll put my limited energy on this matter towards such direction. >
[toc] | [prev] | [next] | [standalone]
| From | Sam Johnston <samj@samj.net> |
|---|---|
| Date | 2025-01-25 19:00 +0100 |
| Message-ID | <K8SBP-bN6Z-1@gated-at.bofh.it> |
| In reply to | #13678 |
On Sat, 25 Jan 2025 at 18:42, Ilu <ilulu@gmx.net> wrote: > Will Debian accept a GR that requires all training data to be free, > including training data that belongs to the core of human dignity? That > would be disturbing. And in fact practically lobotomize good projects. Mozilla would like a word: https://commonvoice.mozilla.org/en/datasets "Each entry in the dataset consists of a unique MP3 and corresponding text file. Many of the 33,151 recorded hours in the dataset also include demographic metadata like age, sex, and accent that can help train the accuracy of speech recognition engines. The dataset currently consists of 22,109 validated hours in 133 languages, but we’re always adding more voices and languages. Take a look at our Languages page to request a language or start contributing." - samj
[toc] | [prev] | [next] | [standalone]
| From | nick black <dankamongmen@gmail.com> |
|---|---|
| Date | 2025-01-29 16:50 +0100 |
| Message-ID | <Kaiud-cS3J-1@gated-at.bofh.it> |
| In reply to | #13678 |
[Multipart message — attachments visible in raw view] — view raw
Ilu left as an exercise for the reader: > If training data involves the core of a humans personhood - as mentioned > above - it cannot be open-sourced or otherwise free. Consent does not > matter. It belongs to that person and nobody else, period. I deeply > believe that there are borders we are not allowed to cross, no matter > how noble the cause. i'm not well versed in these debates, but why am i prohibited from opening my personhood? are you claiming i can't, for instance, distribute a generative approximation to my voice as an open system? that seems such a strange position to me that i suspect i'm fundamentally misunderstanding you. -- nick black -=- https://nick-black.com to make an apple pie from scratch, you need first invent a universe.
[toc] | [prev] | [standalone]
Back to top | Article view | linux.debian.project
csiph-web