Would be a terrible shame if lots of people opted out.
If you are a EU citizen you might also want to write a complaint to privacy@huggingface.co because they are collecting your personally identifiable information in machine-readable form which they are distributing to third parties.
lol my first ever JavaScript project is in the stack. No wonder their models suck so much ass
my commits are the reason we haven’t achieved AGI
I never really needed any justification, but ever since AI companies just take stuff illegally, and that openly, it has become my justification to just pirate the shit out of everything.
(Excluding indie games of developers I like).
This specifically isn’t illegal though. They’ve only scraped public repos.
It should have been anyways. Like what the fuck were you all doing on the internet. Trying to sell your stupid art. Helping corporations steal data and build pay walls so you can earn $500/month selling your kitchy Garfield key chains. The internet should have always been about data replication and sharing. It’s too late to put that cat back in the bag. It was suppose to be stopped 20 years ago. Now we have all this stuff ruining the world and bringing back nazis. But at least the lefty capitalist artist sold some key chains
I don’t see how people selling keychains (physical items) helped corporations steal data.
Keychain just a stand in for anything. More like how someone making videos of making key chains used patent trolls to build tools to scan all the videos they can in order to find users they can take down. Just one example. The little lefty entrepreneur was used to beat the old school leftists into the dirty on behalf of the corporate interests. Now everything is owned by corporate and we lost the internet to racist nazis. There’s a moment 20 years ago things could have been different that we’ll never get back. Frontier was new and we needed to rally around the wagon to defend the frontier. Instead we sold out.
And now all the lefties are in the smallest corner of the internet and the rest of the internet and social media is owned by the capitalist and nazis. It all started by forgetting the roots and in favor of selling our key chains and trying to build a market rather than a collaborative space.
What’s with the obsession with “lefties” vs nazis?? Sounds like somebody stayed on the propaganda carousel a bit too long.
You choose what parts of the internet you want to use, so who cares if 99% is profit oriented garbage, honestly.
People choose to publicly share their stuff for free, so of course it’s gonna be taken advantage of. That’s literally capitalism.
A decade ago, if someone asked someone working on an Open Source project if they’d be ok with an AI reading their code and learning from it, they’d most likely say “yeah that sounds really cool!”
Somehow the tech-bros have fucked up AI so much that something that should be really cool seems creepy, lame, and nefarious all at once.
I mean they’d have been ok with it because tech bros were the ones automating other people out of jobs and never thought it would come for theirs.
The level of AI we have now was impossible science fiction a decade ago.
Doesn’t hugging face do open source AI?
Pretty sure all my repos are MIT licensed for the betterment of everyone, but I’m not on the list! So I guess I’m not good enough, or they are failing to follow the attribution clause of it.
My dotfile repo is there and it doesn’t have a licence. Meaning it’s technically not open source. Didn’t stop them
Strange because all of my MIT repos are listed - all the GPL ones are excluded.
Mine are all GPLv3 or forks of other repos and they are listed. I’m sanguine because the license allows for study and if it’s good enough for humans I don’t see why it’s not for clankers.
For me, and a few other projects I checked, it only has non-GPL repos. But it also does not have everything that isn’t GPL, despite the repos being much older than the cut-off date. But it does have repos without a license, which they are simply not allowed to copy.
I wonder if those repos have been deduplicated, and one of the forks (on some other person’s account) is included instead. Unfortunately you can only search the first 5M records via the website, and I don’t have time to play around with the API at the moment, so I could neither confirm nor deny that possibility
We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.
fuck them seriously; if you want to do that then don’t steal the repositories in the first place.
don’t steal the repositories
How is this remotely stealing?
It’s not, but it may be violation of licenses. And also, if it has personal information on it, that’s probably illegal under the GDPR.
Modern tech companies love using the rapsist’s model of consent.
Meanwhile at work we just had a training course that specifically said doing “opt out” instead of “opt in” violates the principle of informed consent.
Notice how a shit load of these people keep turning up to be rapists and pedophiles… They have no shits to give about informed consent. In their mind you don’t even have the right to the same agency they do.
Kinda hope it uses my code. It’s so terrible there’s no doubt it will make the resulting code from the model worse even if the impact is miniscule
I just opted out of all my repos except the God awful ones from middle and highschool. Those are basically prompt poison so fuckem
Unless I’m mistaken, this wasn’t written by the folks that scaped GitHub in the first place, someone just wrote a small tool to semi-automate the process of searching the scraped dats, and submitting a GitHub issue to have it removed.
The best way to opt out is by not using GitHub. Also opts you out of Copilot and a bunch of other stuff.
Other git hosts are also getting scraped, and have had to implement counters because of it. For example, this is the kind of thing Codeberg shows crawlers. I’ve even seen people who self-host complaining about getting overloaded because of bots scraping their forge
I’ve put Anubis before most of my website, including my forgejo instance. For the projects hosted there, which is not all, I can only hope that that’s enough.
I like to have the visibility and CI of GitHub. But this sucks ass.
Already did that some time ago but I am still in the dataset.
When I back something up, I save a copy and add “backup” to the name… Because I’m advanced.
If I’m feeling really good and healthy, I’ll even put it on a usb stick.
note that former users would have needed to remove their GitHub data before August 2025 to not be in this dataset
A little bit infuriating since huggingface itself requires login to access a large portion of the content on their site
Since they stole my paper on ethics in computer science, maybe the model will learn to act better than its owners
Wonder how many of those repos contain AI-generated code?
if I opt out, will the Roko’s basilisk come after me?
I think its hilarious that some people seem to actually take this science fiction variation on Pascal’s wager seriously.
Reported for cognitohazard. Delete this immediately.
(I kid, but someone really did report your comment.)
lmao
As long as you haven’t been told what you’re opting out of, you’re good.
Good thing I never finished a project!
tbh i’m thinking this alone isn’t that bad from an archiver/datahoarder perspective
As long as the datasets are open, it is our best hope. I know it doesn’t compensate the people whose work’s copyright and licenses have been violated, but I think it’s the only realistic hope we’ve got of getting out of this informational dystopia with a reasonably intact library of humanity’s knowledge that hasn’t been locked down and/or monetized. The AI scrapers and generators are in the process of burning down the great library of Alexandria that the Internet had become, and we are already starting to feel its loss. We cannot stop the wave of toxic pollution that is spreading through all our digital content now, but the archives from before this apocalypse started will become the most valuable thing humanity has ever produced. This is information war, and we are losing.
This is what confuses me. The internet archive has been archive the entire Internet for years. Yet AI does the same to make a way for people to code easier and it is a problem all of the sudden?
IA is a nonprofit and archives to preserve human history. shitty AI startups do this to monetize the data, and their end goal is to “replace” the people who made that data in the first place.
Thanks to these AI mfs the IA now prevents access to many items because they can be used as training material which fucking sucks
As a developer, you hold the copyright to your code. When you make it open-source, you grant a license to use the code and the resulting program under certain terms.
This is a contract. If you copy my code without following these terms, then that’s theft.
The Internet Archive’s use complies with these terms for all open-source licenses. These AI companies do not. In particular, here’s a quote from the MIT license, which you will find in a similar wording in all open-source licenses:
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
In effect, what this means, is that when you copy my code, I demand that you also copy the license text along with it, so that anyone else looking at this code knows the permissions I grant and the terms I require.
And now guess what these AI companies are doing. They copy my code and reproduce substantial portions upon a user asking, yet they do not include my license terms. They violate the contract under which they obtained my source code.
I suspect you don’t realize how shit that is, because source code is so abstract.
It’s like spending hundreds of hours painting a great artwork and then deciding that everyone should be able to give a copy to everyone they know, under the simple condition that they inform those people that they have this right as well.
And then comes along a company and sells my artwork for money, without informing their customers that they can pass it on for free. That’s, plain and simple, a criminal operation.See, the part where you lose me is where you start to talk about copyright. If code is open, then who gives a shit who “owns” it?
If nobody owns the code, then then nobody can enforce the terms of the license it was released under, and free software under the FSF definition becomes impossible. All you have is public domain.
For example, a company could take the Linux kernel, modify it and distribute it with their gadgets. And they could simply not release the modifications they’ve made, as is required by the GNU Public License. But nobody would be able to do anything about it. Currently, copyright laws allow the people who wrote the Linux kernel to sue the company for breaking the license and violating the authors’ copyrights
No one. What people give a shit about is the license that is supposed to keep the code open, which is being removed for profit without consequence.
In our current legal system, copyright is the basis for me to be able to set requirements on how my code can be shared. I do not care that I own it, I just care that it is shared under the conditions I set.
Without being able set these conditions, I would not open up my code.
Internet Archive exists as a reference for your edification on a donation basis; AI companies intend to initiate a top down societal restructuring of jobs and thus access to benefits, paywall access to your own collective information, fund themselves through ouroboros leveraged deals and VC money (value that’s been stolen from the general populace over the years)
I agree that for profit closed source AI companies are bad, but open weights models are a different thing, are they not?
Is there a form to ask to be included in the next stack? they seem to have missed me this time
“Oh no, people are using information I put publicly available on the internet for everyone to see!”
Morons. The lot of you.
I did it to invite collaboration and connect with other developers with similar interests. FOSS is more about building communities than building software, after all.
I did not anticipate that it could be (legally) used to dismantle the kinds of communities I wanted to build. (I did anticipate that it could be illegally used to that end, but historically that has tended to cause a Streisand Effect, so that risk seemed worth it.)
Sure. Let’s see if they use it for endeavours in the same spirit.
Or are you fine with they using this data for-profit without benefitting the public by also making it open?
(not talking about hf, I know starcoder. just in general)
Love you too, kitten. ♥️















