In theory, it’s all supposed to be permissively licenced code and the opt out is more than other models give. I saw Starcoder as one of the more ethical models. I thought the underlying principals to be fair at least.
I’m interested in this gut hostility to it regardless. Kind of shows how you can’t present LLMs in a positive angle no matter what. They shouldn’t be using anything GPL or similar.
Kind of shows how you can’t present LLMsharvesting peoples data without consent or even warning and making it difficult to impossible for people to avoid it in a positive angle no matter what.
Fixed it for you.
An LLM built from only consensually provided data would be perfectly fine, as long as it works without environmentally ruinous data centers.
In fact, that was how EVERY LLM was to begin with, until regulatory capture and corporate impunity reached the current crescendo.
But if it’s permissively licenced, couldn’t I just copy bits and pieces for my own project without asking?
Like I understand asking is always better and an opt in process for “the stack” or wtv would have been better received. Nevertheless, did they really have a legal obligation, rather than moral obligation, to ask given how this is licensed?
Taking GPL code would be a different situation right? They would have to also include it in their model creation… which theoretically they just could.
Most of those “permissive” licenses require redistributors to redistribute copies of the license texts in derivative works.
But I bet these AI models aren’t doing that. And it’s a damn neat certainty that the vibe coders who use the AI model are not attaching a license disclosure containing every permissive licenses in GitHub. Even if their vibe coded app is arguably a derivative work.
I do not believe a permissive license has any notion of consent. You can’t stop someone from forking your code as long as they follow the rules of the license, like crediting you.
Even in GPL, you can’t stop someone from using your code, see Gnome’s recent arguements with Mint over their usage of an old version of their Calender app.
Others in the thread have mentioned we’ll see permissive licenses with exceptions for Ai in the near future. It’s a solution because that door is currently open.
I do not believe a permissive license has any notion of consent. You can’t stop someone from forking your code as long as they follow the rules of the license, like crediting you.
Does the LLM ever credit the original author when it spits out code?
You can’t stop someone from forking your code as long as they follow the rules of the license, like crediting you
In that case,though, you’re still respecting the wishes of the person applying that permissive license consensually, though.
That’s VERY different from mass harvesting all data without permission (or credit) for profit, which would probably be against the terms of even the most permissive licenses.
Their claim was they followed the licenses and only used code they were implicitly allowed to use. I’d like to see if that was just astroturf and bullshit.
Honestly, though this is one of the reasons why people prefer copyleft and avoid permissive licenses, because yeah anyone can profit monetarily off your work otherwise.
Sure but you moved the goalpost from ‘‘they weren’t allowed’’ to “they aren’t being honest”. There’s people in this thread who are in the repos, so we can find that out without assuming.
Also, at least their model is opensource and can easily be selfhosted on consumer hardware. Helps with your datacenter point from earlier. Like if this project does what it says its supposed to do, it’s genuinely interesting. If they are pushing the boundaries of permissive licensing then yeah that is really concerning.
Well, they included some MPL 2.0 repos of mine at least, but skipped others that use GPL3. Yet another that is a mirror of some otherwise lost firmware files for early 2000s wifi cards (and definitely isn’t free software) is also included.
So possibly they filter out GPL2/3 specifically, rather than only include known permissive licenses. Which is a pretty bad way of doing it.
Okay, I see you, I think I may have misunderstood how you were phrasing it. Thinking you were saying they were scanning all of GitHub because it’s all permissive.
We’ll they said they only scanned stuff that would have allowed them implicitly, now did they really respect that? Their ‘‘stack’’ is public, so people can review it. It’s opensource, people can audit it at least.
The big giants can lift whatever they want from Github and we wouldn’t have the means to prove it. I’m sure Microsoft is using private repos as they like. It’s not a coincidence that Copilot was one of the earlier coding LLMs.
Machine learning is more than just “transformative use” and is not copying. Currently that is only like 98% true, because memorization does occur in a few cases. Like 0.8-2% and that number is probably less now 2 years later than that study. Ultimately once they fix the memorization issue and “purge” these memories and can no longer reproduce licensed code (which is mostly textbook examples and boilerplate code or very popular code) this will be true transformative learning.
Then they do not require any more permission to read and learn from a book or from code than a human would. As long as you own a book or have the right to read something, you’re allowed to do whatever you want with the knowledge you gained.
In theory, it’s all supposed to be permissively licenced code and the opt out is more than other models give. I saw Starcoder as one of the more ethical models. I thought the underlying principals to be fair at least.
I’m interested in this gut hostility to it regardless. Kind of shows how you can’t present LLMs in a positive angle no matter what. They shouldn’t be using anything GPL or similar.
Fixed it for you.
An LLM built from only consensually provided data would be perfectly fine, as long as it works without environmentally ruinous data centers.
In fact, that was how EVERY LLM was to begin with, until regulatory capture and corporate impunity reached the current crescendo.
But if it’s permissively licenced, couldn’t I just copy bits and pieces for my own project without asking?
Like I understand asking is always better and an opt in process for “the stack” or wtv would have been better received. Nevertheless, did they really have a legal obligation, rather than moral obligation, to ask given how this is licensed?
Taking GPL code would be a different situation right? They would have to also include it in their model creation… which theoretically they just could.
Most of those “permissive” licenses require redistributors to redistribute copies of the license texts in derivative works.
But I bet these AI models aren’t doing that. And it’s a damn neat certainty that the vibe coders who use the AI model are not attaching a license disclosure containing every permissive licenses in GitHub. Even if their vibe coded app is arguably a derivative work.
…do you know what the word “consensually” means? 🤦🏻
I do not believe a permissive license has any notion of consent. You can’t stop someone from forking your code as long as they follow the rules of the license, like crediting you.
Even in GPL, you can’t stop someone from using your code, see Gnome’s recent arguements with Mint over their usage of an old version of their Calender app.
Others in the thread have mentioned we’ll see permissive licenses with exceptions for Ai in the near future. It’s a solution because that door is currently open.
Does the LLM ever credit the original author when it spits out code?
In that case,though, you’re still respecting the wishes of the person applying that permissive license consensually, though.
That’s VERY different from mass harvesting all data without permission (or credit) for profit, which would probably be against the terms of even the most permissive licenses.
Their claim was they followed the licenses and only used code they were implicitly allowed to use. I’d like to see if that was just astroturf and bullshit.
Honestly, though this is one of the reasons why people prefer copyleft and avoid permissive licenses, because yeah anyone can profit monetarily off your work otherwise.
Was likely bullshit. Just like just about everything else people behind for profit LLMs say about their business practices.
Almost certainly
Sure but you moved the goalpost from ‘‘they weren’t allowed’’ to “they aren’t being honest”. There’s people in this thread who are in the repos, so we can find that out without assuming.
Also, at least their model is opensource and can easily be selfhosted on consumer hardware. Helps with your datacenter point from earlier. Like if this project does what it says its supposed to do, it’s genuinely interesting. If they are pushing the boundaries of permissive licensing then yeah that is really concerning.
You can host copyleft as well as all rights reserved code on GitHub. It’s not like Codeberg.
Yeah, they aren’t supposed to scrape that stuff. That’s kind of where I’m wondering if they limited their collections.
Well, they included some MPL 2.0 repos of mine at least, but skipped others that use GPL3. Yet another that is a mirror of some otherwise lost firmware files for early 2000s wifi cards (and definitely isn’t free software) is also included.
So possibly they filter out GPL2/3 specifically, rather than only include known permissive licenses. Which is a pretty bad way of doing it.
Okay, I see you, I think I may have misunderstood how you were phrasing it. Thinking you were saying they were scanning all of GitHub because it’s all permissive.
We’ll they said they only scanned stuff that would have allowed them implicitly, now did they really respect that? Their ‘‘stack’’ is public, so people can review it. It’s opensource, people can audit it at least.
The big giants can lift whatever they want from Github and we wouldn’t have the means to prove it. I’m sure Microsoft is using private repos as they like. It’s not a coincidence that Copilot was one of the earlier coding LLMs.
Machine learning is more than just “transformative use” and is not copying. Currently that is only like 98% true, because memorization does occur in a few cases. Like 0.8-2% and that number is probably less now 2 years later than that study. Ultimately once they fix the memorization issue and “purge” these memories and can no longer reproduce licensed code (which is mostly textbook examples and boilerplate code or very popular code) this will be true transformative learning.
Then they do not require any more permission to read and learn from a book or from code than a human would. As long as you own a book or have the right to read something, you’re allowed to do whatever you want with the knowledge you gained.