Groups | Search | Server Info | Keyboard shortcuts | Login | Register [http] [https] [nntp] [nntps]


Groups > linux.debian.project > #13644 > unrolled thread

Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2

Started byGerardo Ballabio <gerardo.ballabio@gmail.com>
First post2024-11-04 12:00 +0100
Last post2025-01-29 16:50 +0100
Articles 12 — 8 participants

Back to article view | Back to linux.debian.project


Contents

  Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Gerardo Ballabio <gerardo.ballabio@gmail.com> - 2024-11-04 12:00 +0100
    Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Mo Zhou <lumin@debian.org> - 2024-11-05 03:20 +0100
      Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Sean Whitton <spwhitton@spwhitton.name> - 2025-01-25 13:20 +0100
        Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 "M. Zhou" <lumin@debian.org> - 2025-01-25 16:30 +0100
          Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Sean Whitton <spwhitton@spwhitton.name> - 2025-01-25 16:40 +0100
            Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 "M. Zhou" <lumin@debian.org> - 2025-01-25 17:20 +0100
          Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Sam Johnston <samj@samj.net> - 2025-01-25 17:20 +0100
            Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 "M. Zhou" <lumin@debian.org> - 2025-01-25 18:00 +0100
              Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Stefano Zacchiroli <zack@debian.org> - 2025-01-25 18:40 +0100
              Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Ilu <ilulu@gmx.net> - 2025-01-25 18:50 +0100
                Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 Sam Johnston <samj@samj.net> - 2025-01-25 19:00 +0100
                Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2 nick black <dankamongmen@gmail.com> - 2025-01-29 16:50 +0100

#13644 — Re: Concerns regarding the "Open Source AI Definition" 1.0-RC2

FromGerardo Ballabio <gerardo.ballabio@gmail.com>
Date2024-11-04 12:00 +0100
SubjectRe: Concerns regarding the "Open Source AI Definition" 1.0-RC2
Message-ID<JF2Yp-6y9e-11@gated-at.bofh.it>
The OSAID 1.0 has now been released (with no modifications from the RC2).
Are we still going to take any actions or will we let this go?

Gerardo

[toc] | [next] | [standalone]


#13645

FromMo Zhou <lumin@debian.org>
Date2024-11-05 03:20 +0100
Message-ID<JFhkJ-6Hae-3@gated-at.bofh.it>
In reply to#13644
I'm planning to draft a GR for this, but that is only going to happen 
after I get through some busy weeks.

On 11/4/24 02:22, Gerardo Ballabio wrote:
> The OSAID 1.0 has now been released (with no modifications from the RC2).
> Are we still going to take any actions or will we let this go?
>
> Gerardo
>

[toc] | [prev] | [next] | [standalone]


#13671

FromSean Whitton <spwhitton@spwhitton.name>
Date2025-01-25 13:20 +0100
Message-ID<K8NiN-bJXM-5@gated-at.bofh.it>
In reply to#13645

[Multipart message — attachments visible in raw view] — view raw

Hello Lumin,

On Mon 04 Nov 2024 at 05:52pm -08, Mo Zhou wrote:

> I'm planning to draft a GR for this, but that is only going to happen after I
> get through some busy weeks.

Wondered if you'd had another chance to look at this.

-- 
Sean Whitton

[toc] | [prev] | [next] | [standalone]


#13672

From"M. Zhou" <lumin@debian.org>
Date2025-01-25 16:30 +0100
Message-ID<K8QgG-bLNm-19@gated-at.bofh.it>
In reply to#13671
On Sat, 2025-01-25 at 12:09 +0000, Sean Whitton wrote:
> 
> Wondered if you'd had another chance to look at this.
> 

Ummm... You know what may happen when there is no deadline.

Did you ping this because there are some thoughts from the policy side?
Or just curious?

[toc] | [prev] | [next] | [standalone]


#13673

FromSean Whitton <spwhitton@spwhitton.name>
Date2025-01-25 16:40 +0100
Message-ID<K8Qql-bLQG-29@gated-at.bofh.it>
In reply to#13672

[Multipart message — attachments visible in raw view] — view raw

Hello,

On Sat 25 Jan 2025 at 10:21am -05, M. Zhou wrote:

> On Sat, 2025-01-25 at 12:09 +0000, Sean Whitton wrote:
>>
>> Wondered if you'd had another chance to look at this.
>>
>
> Ummm... You know what may happen when there is no deadline.
>
> Did you ping this because there are some thoughts from the policy side?
> Or just curious?

Nothing to do with Debian Policy, no.

I'm just interested in your thoughts on the matter.

-- 
Sean Whitton

[toc] | [prev] | [next] | [standalone]


#13675

From"M. Zhou" <lumin@debian.org>
Date2025-01-25 17:20 +0100
Message-ID<K8R33-bMjQ-9@gated-at.bofh.it>
In reply to#13673
On Sat, 2025-01-25 at 15:29 +0000, Sean Whitton wrote:
> 
> Nothing to do with Debian Policy, no.
> 
> I'm just interested in your thoughts on the matter.

If we look at the history. Richard Stallman started the GNU
project at a time point where individual developers and free
software communities can really create software, and define
what "free software" is.

If "free software" was defined at a time point where individuals
and free software communities could not write software on their
own, that definition may turn into a Utopia declaration.

Similarly, we are not yet in an era where individuals and free
software communities can easily create original, useful, and
competitive AI from scratch (to ensure, e.g., reproducibility
and trustworthiness). Particularly for language models.
Currently, the model creators can do any decision regarding
their decisions, regardless of how we define terminologies.

That said, with the advancements of research, and the whole
communities appreciation and recognition on the value of free
and open source, that day will come sooner or later for language
models. Creating some simpler vision or language AIs by individual
is already possible.

So the hard deadline for this matter is the day when people
can freely create and publish AIs. We will be too late at that
time point.

While I disagree on OSAID's requirement on training data, it
is directing at the right direction. I can see OSI is making
compromises due to the dilemma that they need something to apply
in practice and take action, while free/open source communities
can not yet create and own language models. What we have seen
from OSI is possibly the best they can do at the current time
point.

>From the Debian side, my concerns are unchanged. The OSAID does
not guarantee freedom to our user.

I think FSF is keeping an eye on this. Let me try to make a draft
within Feburary.

[toc] | [prev] | [next] | [standalone]


#13674

FromSam Johnston <samj@samj.net>
Date2025-01-25 17:20 +0100
Message-ID<K8R33-bMjQ-3@gated-at.bofh.it>
In reply to#13672
On Sat, 25 Jan 2025 at 16:24, M. Zhou <lumin@debian.org> wrote:
>
> On Sat, 2025-01-25 at 12:09 +0000, Sean Whitton wrote:
> >
> > Wondered if you'd had another chance to look at this.
>
> Ummm... You know what may happen when there is no deadline.

The best time to do this was last year around the OSAID 1.0 release.
The next best time is now. Do you need our help?

I'm working on an article about how the chickens have come home to
roost with the VLC demo at CES 2025. With VLC advertising and users
now expecting real-time AI subtitling that "appears to be built
directly into the VLC app"[1], we have a situation where VLC is
considered Open Source by the OSD, but NOT according to the OSAID and
OSI leadership[2] because of Whisper being embedded. With more and
more software being written by and incorporating AI, this situation is
untenable. Distros like Debian would have to lobotomise popular apps
like VLC, or accept more binary blobs.

The OSI also just released a whitepaper[3] that further deliberately
obfuscates the issue, prompting me to post this:

The Open Source Initiative (OSI) goes to the effort of defining four
classes of data *source* (hence the term!) in their Open Source AI
Definition (OSAID) FAQ and again in the Open Future Foundation’s name
in this new paper, only to then accept ANY of them… or NONE at all:

- OPEN data under open licenses, which is the ONLY class that has any
role in Open Source AI
- PUBLIC data like Common Crawl Foundation dumps of the Internet,
which are routinely ab/used without creators’ consent
- OBTAINABLE data “including for a fee” like The New York Times
articles and Adobe/Getty Images stock photos, which are guaranteed to
get end users (but not necessarily vendors given limited liability
clauses) sued
- UNSHAREABLE NONPUBLIC data that obviously has no place in Open
Source, like Facebook & Instagram feeds

With the meaning of Open Source AI being defined solely by the LOWEST
bar — no data delivered at all (which is allowed under the OSAID) —
why bother with the smokescreen if not to deliberately deceive us
users? An honest FAQ entry would have read like this:

What kind of data should be required in the Open Source AI Definition?
None.

1. https://hackaday.com/2025/01/15/floss-weekly-episode-816-open-source-ai/
2. https://www.theverge.com/2025/1/9/24339817/vlc-player-automatic-ai-subtitling-translation
3. https://openfuture.eu/publication/data-governance-in-open-source-ai/

[toc] | [prev] | [next] | [standalone]


#13676

From"M. Zhou" <lumin@debian.org>
Date2025-01-25 18:00 +0100
Message-ID<K8RFM-bMxL-19@gated-at.bofh.it>
In reply to#13674
On Sat, 2025-01-25 at 17:08 +0100, Sam Johnston wrote:
> 
> The best time to do this was last year around the OSAID 1.0 release.
> The next best time is now. Do you need our help?
> 

I lean towards making things simpler.

Yes I disagree with OSI's decision on OSAID, and the definition
does not guarantee freedom at all. But a bold move towards picking
a fight against OSI on this matter through Debian General Resolution
sounds terrible and reckless to me.

I'll focus on a simpler topic for the GR:

"how does Debian community interpret DFSG and software freedom
against the AI model and software?"

I'll draft from a pure technical point of view. Neutral to
individuals and organizations, without commenting on how others
think and do. In that case it is as simply as elaborating the
"toxic candy" case, and analyzing the OSAID's implication from
a technical point of view.

In that sense, things will be more constructive and doable.
FSF will also able to learn from Debian's GR discussion.
I'll put my limited energy on this matter towards such direction.

[toc] | [prev] | [next] | [standalone]


#13677

FromStefano Zacchiroli <zack@debian.org>
Date2025-01-25 18:40 +0100
Message-ID<K8Sit-bN0E-3@gated-at.bofh.it>
In reply to#13676

[Multipart message — attachments visible in raw view] — view raw

On Sat, Jan 25, 2025 at 11:37:48AM -0500, M. Zhou wrote:
> I'll focus on a simpler topic for the GR:
> 
> "how does Debian community interpret DFSG and software freedom
> against the AI model and software?"

Clarifying this officially seem indeed very useful, both for Debian and
for the free software world at large, due to how relevant Debian is as
one of the important gatekeepers of what is free-software-ok and what is
not. Many people out there, and possibly even within the project, are
assuming that the current Debian AI Policy is an official project
position, whereas it is not. It's important to ratify it somehow.

Since the last time this was discussed, did anyone reach out to FTP
master to understand *if* they currently have already an official answer
to this question or not? I'm Cc:-ing them explicitly with this message,
just in case.

Aside from the potential constitutional implications of this (last time
it felt like I was the only one worrying about this aspect, so I'm
dropping it), it would still be interesting to know their take. And it's
quite possible they would actively welcome/encourage an official
project-wide vote on this matter.

Cheers
-- 
Stefano Zacchiroli . zack@upsilon.cc . https://upsilon.cc/zack  _. ^ ._
Full professor of Computer Science              o     o   o     \/|V|\/
Télécom Paris, Polytechnic Institute of Paris     o     o o    </>   <\>
Co-founder & CSO Software Heritage            o o o     o       /\|^|/\
Mastodon: https://mastodon.xyz/@zacchiro                        '" V "'

[toc] | [prev] | [next] | [standalone]


#13678

FromIlu <ilulu@gmx.net>
Date2025-01-25 18:50 +0100
Message-ID<K8Ss9-bN3W-1@gated-at.bofh.it>
In reply to#13676
When this discussion came up I immeditely thought of text-to-speech
projects like piper, using AI generated voices derived from real persons
voice data. The voice is a very specific attribute of every human. Its
part of attributes that define that persons humanity. Same goes for the
face, the fingerprint, the eyes, the movements, the genome. All these
data are very specific and in summa defines a person (although there are
probably more factors, just these came to mind). It also is already used
to identify a person.

If training data involves the core of a humans personhood - as mentioned
above - it cannot be open-sourced or otherwise free. Consent does not
matter. It belongs to that person and nobody else, period. I deeply
believe that there are borders we are not allowed to cross, no matter
how noble the cause.

Do we still want to positively distinguish AI projects whose code,
parameters, weights and adjustments are free? Yes I do and IMHO OSI has
the same goal. That's why I think that OSAID is on the right track. They
could have worded things better and they could have more precisely
distinguished different types of training data but in light of time
contraints (EU AI Act) I think they did the best they could.

Will Debian accept a GR that requires all training data to be free,
including training data that belongs to the core of human dignity? That
would be disturbing. And in fact practically lobotomize good projects.

Am 25.01.25 um 17:37 schrieb M. Zhou:
> On Sat, 2025-01-25 at 17:08 +0100, Sam Johnston wrote:
>>
>> The best time to do this was last year around the OSAID 1.0 release.
>> The next best time is now. Do you need our help?
>>
>
> I lean towards making things simpler.
>
> Yes I disagree with OSI's decision on OSAID, and the definition
> does not guarantee freedom at all. But a bold move towards picking
> a fight against OSI on this matter through Debian General Resolution
> sounds terrible and reckless to me.
>
> I'll focus on a simpler topic for the GR:
>
> "how does Debian community interpret DFSG and software freedom
> against the AI model and software?"
>
> I'll draft from a pure technical point of view. Neutral to
> individuals and organizations, without commenting on how others
> think and do. In that case it is as simply as elaborating the
> "toxic candy" case, and analyzing the OSAID's implication from
> a technical point of view.
>
> In that sense, things will be more constructive and doable.
> FSF will also able to learn from Debian's GR discussion.
> I'll put my limited energy on this matter towards such direction.
>

[toc] | [prev] | [next] | [standalone]


#13679

FromSam Johnston <samj@samj.net>
Date2025-01-25 19:00 +0100
Message-ID<K8SBP-bN6Z-1@gated-at.bofh.it>
In reply to#13678
On Sat, 25 Jan 2025 at 18:42, Ilu <ilulu@gmx.net> wrote:
> Will Debian accept a GR that requires all training data to be free,
> including training data that belongs to the core of human dignity? That
> would be disturbing. And in fact practically lobotomize good projects.

Mozilla would like a word: https://commonvoice.mozilla.org/en/datasets

"Each entry in the dataset consists of a unique MP3 and corresponding
text file. Many of the 33,151 recorded hours in the dataset also
include demographic metadata like age, sex, and accent that can help
train the accuracy of speech recognition engines. The dataset
currently consists of 22,109 validated hours in 133 languages, but
we’re always adding more voices and languages. Take a look at our
Languages page to request a language or start contributing."

 - samj

[toc] | [prev] | [next] | [standalone]


#13682

Fromnick black <dankamongmen@gmail.com>
Date2025-01-29 16:50 +0100
Message-ID<Kaiud-cS3J-1@gated-at.bofh.it>
In reply to#13678

[Multipart message — attachments visible in raw view] — view raw

Ilu left as an exercise for the reader:
> If training data involves the core of a humans personhood - as mentioned
> above - it cannot be open-sourced or otherwise free. Consent does not
> matter. It belongs to that person and nobody else, period. I deeply
> believe that there are borders we are not allowed to cross, no matter
> how noble the cause.

i'm not well versed in these debates, but why am i prohibited
from opening my personhood? are you claiming i can't, for instance,
distribute a generative approximation to my voice as an open
system? that seems such a strange position to me that i suspect
i'm fundamentally misunderstanding you.

-- 
nick black -=- https://nick-black.com
to make an apple pie from scratch,
you need first invent a universe.

[toc] | [prev] | [standalone]


Back to top | Article view | linux.debian.project


csiph-web