Would be a terrible shame if lots of people opted out.

If you are a EU citizen you might also want to write a complaint to privacy@huggingface.co because they are collecting your personally identifiable information in machine-readable form which they are distributing to third parties.

  • gedfromgont@piefed.ca
    link
    fedilink
    English
    arrow-up
    32
    ·
    5 hours ago

    I never really needed any justification, but ever since AI companies just take stuff illegally, and that openly, it has become my justification to just pirate the shit out of everything.

    (Excluding indie games of developers I like).

    • Melvin_Ferd@lemmy.world
      link
      fedilink
      arrow-up
      7
      arrow-down
      10
      ·
      edit-2
      5 hours ago

      It should have been anyways. Like what the fuck were you all doing on the internet. Trying to sell your stupid art. Helping corporations steal data and build pay walls so you can earn $500/month selling your kitchy Garfield key chains. The internet should have always been about data replication and sharing. It’s too late to put that cat back in the bag. It was suppose to be stopped 20 years ago. Now we have all this stuff ruining the world and bringing back nazis. But at least the lefty capitalist artist sold some key chains

        • Melvin_Ferd@lemmy.world
          link
          fedilink
          arrow-up
          6
          arrow-down
          6
          ·
          edit-2
          4 hours ago

          Keychain just a stand in for anything. More like how someone making videos of making key chains used patent trolls to build tools to scan all the videos they can in order to find users they can take down. Just one example. The little lefty entrepreneur was used to beat the old school leftists into the dirty on behalf of the corporate interests. Now everything is owned by corporate and we lost the internet to racist nazis. There’s a moment 20 years ago things could have been different that we’ll never get back. Frontier was new and we needed to rally around the wagon to defend the frontier. Instead we sold out.

          And now all the lefties are in the smallest corner of the internet and the rest of the internet and social media is owned by the capitalist and nazis. It all started by forgetting the roots and in favor of selling our key chains and trying to build a market rather than a collaborative space.

          • Ricaz@lemmy.dbzer0.com
            link
            fedilink
            English
            arrow-up
            1
            ·
            53 minutes ago

            What’s with the obsession with “lefties” vs nazis?? Sounds like somebody stayed on the propaganda carousel a bit too long.

            You choose what parts of the internet you want to use, so who cares if 99% is profit oriented garbage, honestly.

            People choose to publicly share their stuff for free, so of course it’s gonna be taken advantage of. That’s literally capitalism.

  • SpaceCowboy@lemmy.ca
    link
    fedilink
    arrow-up
    44
    arrow-down
    2
    ·
    9 hours ago

    A decade ago, if someone asked someone working on an Open Source project if they’d be ok with an AI reading their code and learning from it, they’d most likely say “yeah that sounds really cool!”

    Somehow the tech-bros have fucked up AI so much that something that should be really cool seems creepy, lame, and nefarious all at once.

    • FizzyOrange@programming.dev
      link
      fedilink
      arrow-up
      3
      ·
      2 hours ago

      I mean they’d have been ok with it because tech bros were the ones automating other people out of jobs and never thought it would come for theirs.

      The level of AI we have now was impossible science fiction a decade ago.

  • G_M0N3Y_2503@lemmy.zip
    link
    fedilink
    arrow-up
    12
    arrow-down
    1
    ·
    8 hours ago

    Pretty sure all my repos are MIT licensed for the betterment of everyone, but I’m not on the list! So I guess I’m not good enough, or they are failing to follow the attribution clause of it.

    • Miaou@jlai.lu
      link
      fedilink
      arrow-up
      1
      ·
      7 minutes ago

      My dotfile repo is there and it doesn’t have a licence. Meaning it’s technically not open source. Didn’t stop them

      • Alex@lemmy.ml
        link
        fedilink
        arrow-up
        4
        ·
        4 hours ago

        Mine are all GPLv3 or forks of other repos and they are listed. I’m sanguine because the license allows for study and if it’s good enough for humans I don’t see why it’s not for clankers.

        • fruitcantfly@programming.dev
          link
          fedilink
          arrow-up
          2
          ·
          2 hours ago

          For me, and a few other projects I checked, it only has non-GPL repos. But it also does not have everything that isn’t GPL, despite the repos being much older than the cut-off date. But it does have repos without a license, which they are simply not allowed to copy.

          I wonder if those repos have been deduplicated, and one of the forks (on some other person’s account) is included instead. Unfortunately you can only search the first 5M records via the website, and I don’t have time to play around with the API at the moment, so I could neither confirm nor deny that possibility

  • katy ✨@piefed.blahaj.zone
    link
    fedilink
    English
    arrow-up
    107
    arrow-down
    2
    ·
    12 hours ago

    We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.

    fuck them seriously; if you want to do that then don’t steal the repositories in the first place.

      • PlexSheep@infosec.pub
        link
        fedilink
        arrow-up
        4
        ·
        1 hour ago

        It’s not, but it may be violation of licenses. And also, if it has personal information on it, that’s probably illegal under the GDPR.

    • kibiz0r@midwest.social
      link
      fedilink
      English
      arrow-up
      70
      ·
      11 hours ago

      Meanwhile at work we just had a training course that specifically said doing “opt out” instead of “opt in” violates the principle of informed consent.

      • atomicbocks@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        14
        ·
        7 hours ago

        Notice how a shit load of these people keep turning up to be rapists and pedophiles… They have no shits to give about informed consent. In their mind you don’t even have the right to the same agency they do.

    • lps2@lemmy.ml
      link
      fedilink
      arrow-up
      32
      ·
      10 hours ago

      Kinda hope it uses my code. It’s so terrible there’s no doubt it will make the resulting code from the model worse even if the impact is miniscule

      • kboy101222@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        7
        arrow-down
        1
        ·
        7 hours ago

        I just opted out of all my repos except the God awful ones from middle and highschool. Those are basically prompt poison so fuckem

    • JakenVeina@midwest.social
      link
      fedilink
      arrow-up
      5
      ·
      10 hours ago

      Unless I’m mistaken, this wasn’t written by the folks that scaped GitHub in the first place, someone just wrote a small tool to semi-automate the process of searching the scraped dats, and submitting a GitHub issue to have it removed.

  • abacabadabacaba@infosec.pub
    link
    fedilink
    English
    arrow-up
    63
    arrow-down
    2
    ·
    12 hours ago

    The best way to opt out is by not using GitHub. Also opts you out of Copilot and a bunch of other stuff.

    • fruitcantfly@programming.dev
      link
      fedilink
      arrow-up
      5
      ·
      edit-2
      2 hours ago

      Other git hosts are also getting scraped, and have had to implement counters because of it. For example, this is the kind of thing Codeberg shows crawlers. I’ve even seen people who self-host complaining about getting overloaded because of bots scraping their forge

      • PlexSheep@infosec.pub
        link
        fedilink
        arrow-up
        2
        ·
        1 hour ago

        I’ve put Anubis before most of my website, including my forgejo instance. For the projects hosted there, which is not all, I can only hope that that’s enough.

        I like to have the visibility and CI of GitHub. But this sucks ass.

    • DarkCloud@lemmy.world
      link
      fedilink
      arrow-up
      4
      arrow-down
      1
      ·
      edit-2
      11 hours ago

      When I back something up, I save a copy and add “backup” to the name… Because I’m advanced.

      If I’m feeling really good and healthy, I’ll even put it on a usb stick.

    • Eager Eagle@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      edit-2
      11 hours ago

      note that former users would have needed to remove their GitHub data before August 2025 to not be in this dataset

  • chicken@lemmy.dbzer0.com
    link
    fedilink
    arrow-up
    29
    ·
    11 hours ago

    A little bit infuriating since huggingface itself requires login to access a large portion of the content on their site

  • goatbeard@beehaw.org
    link
    fedilink
    arrow-up
    9
    ·
    edit-2
    9 hours ago

    Since they stole my paper on ethics in computer science, maybe the model will learn to act better than its owners

    • cecilkorik@lemmy.ca
      link
      fedilink
      English
      arrow-up
      3
      ·
      edit-2
      9 hours ago

      As long as the datasets are open, it is our best hope. I know it doesn’t compensate the people whose work’s copyright and licenses have been violated, but I think it’s the only realistic hope we’ve got of getting out of this informational dystopia with a reasonably intact library of humanity’s knowledge that hasn’t been locked down and/or monetized. The AI scrapers and generators are in the process of burning down the great library of Alexandria that the Internet had become, and we are already starting to feel its loss. We cannot stop the wave of toxic pollution that is spreading through all our digital content now, but the archives from before this apocalypse started will become the most valuable thing humanity has ever produced. This is information war, and we are losing.

    • korendian@piefed.social
      link
      fedilink
      English
      arrow-up
      6
      arrow-down
      6
      ·
      8 hours ago

      This is what confuses me. The internet archive has been archive the entire Internet for years. Yet AI does the same to make a way for people to code easier and it is a problem all of the sudden?

      • hexagonwin@lemmy.today
        link
        fedilink
        arrow-up
        10
        arrow-down
        1
        ·
        5 hours ago

        IA is a nonprofit and archives to preserve human history. shitty AI startups do this to monetize the data, and their end goal is to “replace” the people who made that data in the first place.

        Thanks to these AI mfs the IA now prevents access to many items because they can be used as training material which fucking sucks

      • trem@lemmy.blahaj.zone
        link
        fedilink
        arrow-up
        9
        ·
        6 hours ago

        As a developer, you hold the copyright to your code. When you make it open-source, you grant a license to use the code and the resulting program under certain terms.

        This is a contract. If you copy my code without following these terms, then that’s theft.

        The Internet Archive’s use complies with these terms for all open-source licenses. These AI companies do not. In particular, here’s a quote from the MIT license, which you will find in a similar wording in all open-source licenses:

        The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

        https://mit-license.org/

        In effect, what this means, is that when you copy my code, I demand that you also copy the license text along with it, so that anyone else looking at this code knows the permissions I grant and the terms I require.

        And now guess what these AI companies are doing. They copy my code and reproduce substantial portions upon a user asking, yet they do not include my license terms. They violate the contract under which they obtained my source code.

        I suspect you don’t realize how shit that is, because source code is so abstract.
        It’s like spending hundreds of hours painting a great artwork and then deciding that everyone should be able to give a copy to everyone they know, under the simple condition that they inform those people that they have this right as well.
        And then comes along a company and sells my artwork for money, without informing their customers that they can pass it on for free. That’s, plain and simple, a criminal operation.

        • korendian@piefed.social
          link
          fedilink
          English
          arrow-up
          2
          arrow-down
          5
          ·
          4 hours ago

          See, the part where you lose me is where you start to talk about copyright. If code is open, then who gives a shit who “owns” it?

          • fruitcantfly@programming.dev
            link
            fedilink
            arrow-up
            2
            ·
            edit-2
            2 hours ago

            If nobody owns the code, then then nobody can enforce the terms of the license it was released under, and free software under the FSF definition becomes impossible. All you have is public domain.

            For example, a company could take the Linux kernel, modify it and distribute it with their gadgets. And they could simply not release the modifications they’ve made, as is required by the GNU Public License. But nobody would be able to do anything about it. Currently, copyright laws allow the people who wrote the Linux kernel to sue the company for breaking the license and violating the authors’ copyrights

          • Squirrelanna@lemmy.blahaj.zone
            link
            fedilink
            arrow-up
            5
            ·
            4 hours ago

            No one. What people give a shit about is the license that is supposed to keep the code open, which is being removed for profit without consequence.

          • trem@lemmy.blahaj.zone
            link
            fedilink
            arrow-up
            3
            ·
            3 hours ago

            In our current legal system, copyright is the basis for me to be able to set requirements on how my code can be shared. I do not care that I own it, I just care that it is shared under the conditions I set.

            Without being able set these conditions, I would not open up my code.

      • Jtotheb@lemmy.world
        link
        fedilink
        arrow-up
        3
        arrow-down
        1
        ·
        6 hours ago

        Internet Archive exists as a reference for your edification on a donation basis; AI companies intend to initiate a top down societal restructuring of jobs and thus access to benefits, paywall access to your own collective information, fund themselves through ouroboros leveraged deals and VC money (value that’s been stolen from the general populace over the years)

        • korendian@piefed.social
          link
          fedilink
          English
          arrow-up
          1
          arrow-down
          1
          ·
          4 hours ago

          I agree that for profit closed source AI companies are bad, but open weights models are a different thing, are they not?

  • SnailMagnitude@mander.xyz
    link
    fedilink
    English
    arrow-up
    3
    arrow-down
    9
    ·
    12 hours ago

    Is there a form to ask to be included in the next stack? they seem to have missed me this time

  • professor_prime@lemmychan.org
    link
    fedilink
    arrow-up
    8
    arrow-down
    25
    ·
    12 hours ago

    “Oh no, people are using information I put publicly available on the internet for everyone to see!”

    Morons. The lot of you.

    • kibiz0r@midwest.social
      link
      fedilink
      English
      arrow-up
      14
      ·
      11 hours ago

      I did it to invite collaboration and connect with other developers with similar interests. FOSS is more about building communities than building software, after all.

      I did not anticipate that it could be (legally) used to dismantle the kinds of communities I wanted to build. (I did anticipate that it could be illegally used to that end, but historically that has tended to cause a Streisand Effect, so that risk seemed worth it.)

    • lavember@programming.dev
      link
      fedilink
      arrow-up
      4
      ·
      10 hours ago

      Sure. Let’s see if they use it for endeavours in the same spirit.

      Or are you fine with they using this data for-profit without benefitting the public by also making it open?

      (not talking about hf, I know starcoder. just in general)