Skip to content

Why is this opt-out and not opt-in?? - #3724

Open
xx-kotas-xx wants to merge 1 commit into
bigcode-project:mainfrom
xx-kotas-xx:main
Open

Why is this opt-out and not opt-in??#3724
xx-kotas-xx wants to merge 1 commit into
bigcode-project:mainfrom
xx-kotas-xx:main

Conversation

@xx-kotas-xx

Copy link
Copy Markdown

Could you PLEASE make this opt-in instead? This is very scummy and I don't think as many people will be willing to support you if this is the case.

Why is this opt-out and not opt-in??
@CDSkyward

Copy link
Copy Markdown

100% should be opt-in instead of opt-out, makes you seem like a shady company to me.

@yxzzy-wtf

Copy link
Copy Markdown

Because it's a project run by greedy losers who don't understand consent.

@CDSkyward

Copy link
Copy Markdown

100%

@xx-kotas-xx

Copy link
Copy Markdown
Author

Because it's a project run by greedy losers who don't understand consent.

very true

@axialeaa

Copy link
Copy Markdown

also note that they have not been responding to (nor acting on) the past like 2k opt-out requests. i guess they weren't expecting this many people to hate generative AI and hate them by extension. what happens in the coming weeks will be even more of a litmus test.

@axialeaa

Copy link
Copy Markdown

100% should be opt-in instead of opt-out, makes you seem like a shady company to me.

they probably knew not enough people would like the idea enough to want to apply

@xx-kotas-xx

Copy link
Copy Markdown
Author

100% should be opt-in instead of opt-out, makes you seem like a shady company to me.

they probably knew not enough people would like the idea enough to want to apply

that's probably the reality of it

@yxzzy-wtf

Copy link
Copy Markdown

100% should be opt-in instead of opt-out, makes you seem like a shady company to me.

they probably knew not enough people would like the idea enough to want to apply

that's probably the reality of it

Laziness is almost certainly also a factor. Ethically reaching out to people to request opt-in takes time and is tough. This route lets them hoover up everything and then just ignore the requests to opt out, easy peasy.

@axialeaa

Copy link
Copy Markdown

it gets worse. it's effectively opt-in to look at the dataset, but opt-out for being included in it. the absolute gall of this fuckass company.

image

@xx-kotas-xx

xx-kotas-xx commented Jul 25, 2026

Copy link
Copy Markdown
Author

this cant be legal... at least in some regions, right?

@axialeaa

Copy link
Copy Markdown

it does seem like they've made an attempt to only scrape projects with permissive licenses, but that is demonstrably not a hard rule. some of my friends' ARR projects were included in the stack

@xx-kotas-xx

xx-kotas-xx commented Jul 25, 2026

Copy link
Copy Markdown
Author

hopefully someone is able to take legal action against this

in the meantime i guess i could create a license for whatever i release next
dunno if you can edit pre-existing licenses, but i hope i can

@axialeaa

Copy link
Copy Markdown

#3356 has a threat of litigation :3

@xx-kotas-xx

Copy link
Copy Markdown
Author

that's amazing!
hope something gets done about it

@CDSkyward

Copy link
Copy Markdown

I hope they get in legal trouble for this, because they 100% deserve it.

@loucyx

loucyx commented Jul 25, 2026

Copy link
Copy Markdown

Why is this opt-out and not opt-in??

Because they know the dataset would be tiny if they made it opt-in, because no one who cares (about their own cognitive capacity, or the access to information on the web, or the environment, or making good investments, or so many other things), would be part of this slop. They just applied the same principle others applied before when stealing art for their datasets: Not giving a **** about what anyone else thinks and just doing whatever they want.

Hopefully, the clock is ticking.

@sophuric

sophuric commented Jul 25, 2026

Copy link
Copy Markdown

I hope they get in legal trouble for this, because they 100% deserve it.

class action when™?

@sophuric

Copy link
Copy Markdown

If it was opt-in, the only people that would opt-in would most likely be vibecoders who use generative AI, causing their model to be trained on other AI generated code, i.e garbage in garbage out, but yeah I fucking loathe how they decided to make it opt-out.

@axialeaa

Copy link
Copy Markdown

If it was opt-in, the only people that would opt-in would most likely be vibecoders who use generative AI, causing their model to be trained on other AI generated code, i.e garbage in garbage out, but yeah I fucking loathe how they decided to make it opt-out.

garbage in, garbage out is better than honest stuff in, garbage out. at least it's self-contained garbage, then. the project would eat itself to death

@fishywitch

Copy link
Copy Markdown

im afraid class action in america against rich greedy assholes wont come far given how money runs everything in this world :D :D, worth trying though just for the possibilty alone

@soupermkc

soupermkc commented Jul 26, 2026

Copy link
Copy Markdown

it gets worse. it's effectively opt-in to look at the dataset, but opt-out for being included in it. the absolute gall of this fuckass company.
image

That's v1, the current iteration should be here (v3), which at least lets you view the thing.
Not that it's any better, it's still just redistributing code regardless of the licensing, violating copyright in the process.

Side note: They mention that they try to abide by the licenses included. The Stack includes code without a license, and did not try to contact the accounts hosting the code, therefore willingly including All Rights Reserved code.
They even admit that they have code without licensing as the type license_type has a state of no_license, and just from searching a couple pages a good chunk of them have no licensing. That is blatantly violating ARR.
It doesn't matter if it's not in the resulting datasets, it's still being redistributed through the publicly accessible training data repo.

@axialeaa

Copy link
Copy Markdown

oh i see, okay. thanks for following up on that

@axialeaa

Copy link
Copy Markdown

Why is this opt-out and not opt-in??

Because they know the dataset would be tiny if they made it opt-in, because no one who cares (about their own cognitive capacity, or the access to information on the web, or the environment, or making good investments, or so many other things), would be part of this slop. They just applied the same principle others applied before when stealing art for their datasets: Not giving a **** about what anyone else thinks and just doing whatever they want.

Hopefully, the clock is ticking.

also thank you very much for that link lol, i had not seen that before. maybe i will make it my screensaver. or make a short video with the william tell overture playing over the top. it's blissfully cathartic 😆

@matthewbaggett

Copy link
Copy Markdown

The AI sector doesn't understand consent or property rights, so we should just take their property in kind.

Clément Delangue, Thomas Wolf and Julien Chaumond are all people with property and exist in the real world, and that property can be seized.

@axialeaa

axialeaa commented Jul 26, 2026

Copy link
Copy Markdown

The AI sector doesn't understand consent or property rights, so we should just take their property in kind.

Clément Delangue, Thomas Wolf and Julien Chaumond are all people with property and exist in the real world, and that property can be seized.

much though i want to say "based" please be careful saying things like that. if this goes to court, it's not going to look good to have physical, in-world property theft threats from our end too. this is not me taking a "two wrongs don't make a right" high ground, i genuinely just think it serves to make our case slightly less legitimate in the eyes of the law.

i fully share the cynicism, let's just try to be as diplomatic as we can through this very reasonable anger :3

@capozi-devv

Copy link
Copy Markdown

For anyone who reads this thread, i found a repo with some anti-ai licenses with versions for most of the well known licenses
https://github.com/non-ai-licenses/non-ai-licenses/tree/main

@xx-kotas-xx

Copy link
Copy Markdown
Author

For anyone who reads this thread, i found a repo with some anti-ai licenses with versions for most of the well known licenses https://github.com/non-ai-licenses/non-ai-licenses/tree/main

nice! ty

@fishywitch

Copy link
Copy Markdown

do they even care abt these optout issues looks like the only time they gave a fuck was like 2yrs ago grrrrrr i hate ai chuds

@axialeaa

Copy link
Copy Markdown

looks like they don't. they're taking the "it'll blow over at some point" approach

@SomeAspy

SomeAspy commented Aug 1, 2026

Copy link
Copy Markdown

Github issues aren't exactly legally binding
I advise you waste their time, money, and man-power by emailing their legal teams:

legal@huggingface.co
dmca@huggingface.co

If you feel so inclined, you can even send them mail:
Hugging Face, Attn: DMCA Agent, 594 Broadway, Suite 1101, New York, NY 10012.

@loucyx

loucyx commented Aug 1, 2026

Copy link
Copy Markdown

Hi everyone, Opening issues and complaining doesn't work Github issues aren't exactly legally binding I advise you waste their time, money, and man-power by emailing their legal teams:

legal@huggingface.co dmca@huggingface.co

If you feel so inclined, you can even send them mail: Hugging Face, Attn: DMCA Agent, 594 Broadway, Suite 1101, New York, NY 10012.

You do understand that reads like: "Please don't complain publicly where everyone can see that no one is happy with this, and send us an email to hugging_face_lawyers@lol.com and hugging_face_turf@lol.com if you want us to do something".

Hugging Face folks keep offering paths with friction, which is such a cliche dark pattern is not even funny. If HF hasn't done anything ethical so far, why would we trust their emails to do something? I said it before, the best option is to take this down and make it opt-in. If HF truly believe in AI, then they should get thousands of people to volunteer their code, right? The second option is to inform everyone affected by this, and offer a one click "remove from everywhere" form. Any other path is scammy.

@SomeAspy

SomeAspy commented Aug 1, 2026

Copy link
Copy Markdown

You do understand that reads like: "Please don't complain publicly where everyone can see that no one is happy with this, and send us an email to our_lawyers@lol.com and our_turf@lol.com if you want us to do something".

You keep offering paths with friction, which is such a cliche dark pattern is not even funny. If you haven't done anything ethical so far, why would we trust your emails to do something? I said it before, the best option is to take this down and make it opt-in. If you truly believe in AI, then you should get thousands of people to volunteer their code, right? The second option is to inform everyone affected by this, and offer a one click "remove from everywhere" form. Any other path is scammy.

I agree. Do continue to complain publicly, and make a big deal out of it. but you should also send these emails since they are legally required to look over them and spend man hours processing them instead of some half-assed github action
I've updated my comment to make it seem more in line with what I was trying to say

@VerbenaIDK

Copy link
Copy Markdown

I have been on both the art side and development side of this whole AI mess, and have been against it since the beginning:
Fuck LLMs, the so called "AI", this project, this company and all other companies involved in this bubble.

I would love to pretend that the people involved care, but I genuinely do not believe they will fix the mess they've created and right their wrongs, until it starts to backfire, that is
So I cheer on people who are willing to litigate over this, in my opinion, legal threats and action on top of the backlash are the only real way forward to stop this kind of thing from happening, the people behind this can clearly see all the backlash, legal action is when they'll start to care and listen to the backlash.

So, here's to all of this backfiring, and for all the cases that may arise to go well!

@gudenau

gudenau commented Aug 3, 2026

Copy link
Copy Markdown

Leandro here, I led the BigCode project a few years ago and worked on Stack v3.

First of all I want to give a bit of context why we built The Stack (incl. v3): We built this dataset to foster an ecosystem of open models as an alternative to the closed proprietary models currently dominating the AI and coding agent worlds. I believe that the future is healthier, if there are open models that everyone can use for free and independently of any specific company. Unfortunately, to build practically useful and competitive open models there is no currently known way around gathering large and high quality training datasets. Providing such an artifact is the objective of The Stack, in order to enable an ecosystem of open models and alternatives to the proprietary models. Previous versions have successfully led to the creation of models such as:

* [StarCoder2](https://huggingface.co/bigcode/starcoder2-15b): a 15B code model, trained directly on [The Stack v2](https://huggingface.co/datasets/bigcode/the-stack-v2-train)

* [OLMo 2](https://huggingface.co/allenai/OLMo-2-1124-7B): AllenAI's fully open LLM, trained on [Dolma](https://huggingface.co/datasets/allenai/dolma), whose entire code split is The Stack (~411B tokens)

* [Apertus](https://huggingface.co/swiss-ai/Apertus-8B-2509): the open Swiss LLM from ETH Zurich and EPFL, trained on [StarCoderData](https://huggingface.co/datasets/bigcode/starcoderdata) and Stack-v2-Edu, both derived from The Stack

* [SmolLM2](https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B): an open model trained on FineWeb-Edu, DCLM and The Stack

While we believe coding models are here to stay and stand from the point of view that there is a need for non-closed models and open research in the field in addition to the dominant closed models, we understand that not everyone feels that way about the trade-offs in the field, or that more generally not everyone should agree to the use of their data. This is why we built an opt-out mechanism, in addition to filtering based on all licences we were able to find in each repository.

* we are currently running the opt-out pipeline and will output an updated version of the dataset within the coming 24-48h

* we will also keep running the opt-out pipeline again every couple of days depending on the number of requests

We would like to generally apologize for those would feel strongly about this decision and do hope that the general intent of our enterprise may come through even in this case. We'll keep you all posted very soon.

If you really want my code I can charge you for it. Considering how much money the LLM industry is burning and how unethical it all is, I'll quote you $1,000,000 US per-repo per-training run.

@modsbydreamCritting

modsbydreamCritting commented Aug 3, 2026

Copy link
Copy Markdown

Leandro here, I led the BigCode project a few years ago and worked on Stack v3.
First of all I want to give a bit of context why we built The Stack (incl. v3): We built this dataset to foster an ecosystem of open models as an alternative to the closed proprietary models currently dominating the AI and coding agent worlds. I believe that the future is healthier, if there are open models that everyone can use for free and independently of any specific company. Unfortunately, to build practically useful and competitive open models there is no currently known way around gathering large and high quality training datasets. Providing such an artifact is the objective of The Stack, in order to enable an ecosystem of open models and alternatives to the proprietary models. Previous versions have successfully led to the creation of models such as:

* [StarCoder2](https://huggingface.co/bigcode/starcoder2-15b): a 15B code model, trained directly on [The Stack v2](https://huggingface.co/datasets/bigcode/the-stack-v2-train)

* [OLMo 2](https://huggingface.co/allenai/OLMo-2-1124-7B): AllenAI's fully open LLM, trained on [Dolma](https://huggingface.co/datasets/allenai/dolma), whose entire code split is The Stack (~411B tokens)

* [Apertus](https://huggingface.co/swiss-ai/Apertus-8B-2509): the open Swiss LLM from ETH Zurich and EPFL, trained on [StarCoderData](https://huggingface.co/datasets/bigcode/starcoderdata) and Stack-v2-Edu, both derived from The Stack

* [SmolLM2](https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B): an open model trained on FineWeb-Edu, DCLM and The Stack

While we believe coding models are here to stay and stand from the point of view that there is a need for non-closed models and open research in the field in addition to the dominant closed models, we understand that not everyone feels that way about the trade-offs in the field, or that more generally not everyone should agree to the use of their data. This is why we built an opt-out mechanism, in addition to filtering based on all licences we were able to find in each repository.

* we are currently running the opt-out pipeline and will output an updated version of the dataset within the coming 24-48h

* we will also keep running the opt-out pipeline again every couple of days depending on the number of requests

We would like to generally apologize for those would feel strongly about this decision and do hope that the general intent of our enterprise may come through even in this case. We'll keep you all posted very soon.

If you really want my code I can charge you for it. Considering how much money the LLM industry is burning and how unethical it all is, I'll quote you $1,000,000 US per-repo per-training run.

"filtering based on all licences we were able to find in each repository." Yeah that doesn't work, numerous people have complained that their ARR licenced code is being used, myself included, even with code which has other licences most of the terms (such as attribution) are ignored.

"we understand that not everyone feels that way about the trade-offs in the field, or that more generally not everyone should agree to the use of their data. This is why we built an opt-out mechanism" Yeah this is better than the other models which simply steal code without even offering an opt out, but it still contributes to the problem of AI companies using other people's code without permission. Not being quite as unethical about it as everyone else yet still doing it isn't something to celebrate.

Numerous opt out requests have been sitting ignored for a week as well.

@IsaacMarovitz

Copy link
Copy Markdown

Numerous opt out requests have been sitting ignored for a week as well.

Agreed the time it takes for you opt-out request to get acknowledge let alone acted upon is unacceptable. Typical DMCA turn around time is 24-48 hours. That includes both acknowledging the request and acting upon it. My request has been open for a week now, and even then I will have to wait for the next release to have my repositories removed. AFAIK there is no policy of back-porting removals to old releases, so my code will forever be apart of them.

@Toby222

Toby222 commented Aug 4, 2026

Copy link
Copy Markdown

Github issues aren't exactly legally binding I advise you waste their time, money, and man-power by emailing their legal teams:

legal@huggingface.co dmca@huggingface.co

If you feel so inclined, you can even send them mail: Hugging Face, Attn: DMCA Agent, 594 Broadway, Suite 1101, New York, NY 10012.

I did (see one of my previous comments for the full email I sent to dmca@)
I still have not heard back from them over a week later

@Toby222

Toby222 commented Aug 4, 2026

Copy link
Copy Markdown
Hi Tobias, 

Thank you for your email and apologies for the delay here.

To opt out of The Stack v3, please use the official request form:

➜  Opt-out Request Form

You'll need to provide:

  • Your GitHub username
  • Whether you want all repos opted out or specific ones listed

The maintainers run the opt-out pipeline every few days, and updated datasets are typically published within 24–48 hours of processing.

Kind regards, 

The HuggingFace team

(... previous emails omitted for brevity ...)

Upon being asked again, I got sent this email
So any DMCA takedown request will just be met with a redirect to this dumpster fire here

@loucyx

loucyx commented Aug 4, 2026

Copy link
Copy Markdown

If not scam, why scam shaped?

@Tahirc1

Tahirc1 commented Aug 5, 2026

Copy link
Copy Markdown

i am thinking of sending a DMCA to legal@huggingface.co dmca@huggingface.co do they work ?

@lele394

lele394 commented Aug 5, 2026

Copy link
Copy Markdown

Yeah well, now I bulk pushed a license file to all my repos, did the opt-out thing, and am considering mailing their legal dpt.

Here's the (ironically) vibe coded script I used to push my license in bulk if anyone wants to do the same.

https://github.com/lele394/github-license-bulk-push

@axialeaa

axialeaa commented Aug 5, 2026

Copy link
Copy Markdown
image ehehehe i love that this is just at the bottom of the thread, almost taunting HF

@SomeAspy

SomeAspy commented Aug 5, 2026

Copy link
Copy Markdown

@anton-l @lvwerra
Given this is the v3 repo, how would I go about getting my licensed code out of v2 and v1? I'm assuming you just stole my code there too, no?

@Toby222 ask the legal team about their other ai models perhaps?

@Toby222

Toby222 commented Aug 5, 2026

Copy link
Copy Markdown

i am thinking of sending a DMCA to legal@huggingface.co dmca@huggingface.co do they work ?

See two comments above yours, @Tahirc1

@xx-kotas-xx

Copy link
Copy Markdown
Author

image ehehehe i love that this is just at the bottom of the thread, almost taunting HF

gonna keep that here until this conflict is resolved lol

@sylv256

sylv256 commented Aug 5, 2026

Copy link
Copy Markdown

Cool, so they're just blatantly infringing on copyright, and they will not respond to DMCA requests. Good to know!

@Toby222

Toby222 commented Aug 6, 2026

Copy link
Copy Markdown

I just bothered to look at the "opt-out pipeline"
And it's just ... completely stupid?
They seem to manually commit the list of "opted-out" things from the issues, and the "opt-out" workflow literally just closes the issues in bulk, doesn't even read the contents??

@Toby222

Toby222 commented Aug 6, 2026

Copy link
Copy Markdown

so what the f- even is the "opt-out pipeline"?

@lele394

lele394 commented Aug 6, 2026

Copy link
Copy Markdown

They seem to manually commit the list of "opted-out" things from the issues, and the "opt-out" workflow literally just closes the issues in bulk, doesn't even read the contents??

Wait, is our data actually being removed?

I'll wait until my issue is closed and I'll stream-sweep this dataset to look for my stuff.

@axialeaa

axialeaa commented Aug 6, 2026

Copy link
Copy Markdown

I just bothered to look at the "opt-out pipeline"
And it's just ... completely stupid?
They seem to manually commit the list of "opted-out" things from the issues, and the "opt-out" workflow literally just closes the issues in bulk, doesn't even read the contents??

ohhhh nohohohoho- this is a smoking gun

@Toby222

Toby222 commented Aug 6, 2026

Copy link
Copy Markdown

They seem to manually commit the list of "opted-out" things from the issues, and the "opt-out" workflow literally just closes the issues in bulk, doesn't even read the contents??

Wait, is our data actually being removed?

I'll wait until my issue is closed and I'll stream-sweep this dataset to look for my stuff.

Questionably, and extremely hard to verify (at least for me because I have absolutely no fucking clue how to even interact with that much data)
The base commit of the repo is 3 days old, and then there is one to "Apply opt-out removals", which is a squash commit, but the commit message links 14 separate commits which are all still valid and contain various levels of "opted-out" data (all of which will obviously also be in the base commit anyways)
You can also see at least two separate "opt-out runs" here and here, which both reference the unmodified dataset without the data removed, still containing it.

ohhhh nohohohoho- this is a smoking gun

? No idea what you mean, they've been pretty blatant from the get-go about rehosting unlicensed code without permission, them being incompetent when they allow you to request not to do it doesn't really change much.

@lele394

lele394 commented Aug 6, 2026

Copy link
Copy Markdown

I mean they have that /data/outputs.jsonl file tat I suppose they use internally to sweep and delete data?

Questionably, and extremely hard to verify (at least for me because I have absolutely no fucking clue how to even interact with that much data)

I'll actually give it a go. You apparently have the ability to stream the dataset from hugging face. I'll probably do that and sweep the metadata AND file content to see if any of my stuff is still in there. They could just be altering metadata, I don't trust them.

@axialeaa

axialeaa commented Aug 6, 2026

Copy link
Copy Markdown

? No idea what you mean, they've been pretty blatant from the get-go about rehosting unlicensed code without permission, them being incompetent when they allow you to request not to do it doesn't really change much.

oh i know but if it turns out they haven't been opting people out from the beginning and have instead just been closing reports to make it look like they're doing something, that's another level of fecklessness that wouldn't be surprising, but it would be damning

@Toby222

Toby222 commented Aug 6, 2026

Copy link
Copy Markdown

? No idea what you mean, they've been pretty blatant from the get-go about rehosting unlicensed code without permission, them being incompetent when they allow you to request not to do it doesn't really change much.

oh i know but if it turns out they haven't been opting people out from the beginning and have instead just been closing reports to make it look like they're doing something, that's another level of fecklessness that wouldn't be surprising, but it would be damning

No they have removed things from the repo, just with like, varying levels of competence.

@addisoncrump

addisoncrump commented Aug 6, 2026

Copy link
Copy Markdown

I mean they have that /data/outputs.jsonl file tat I suppose they use internally to sweep and delete data?

Yeah my issue requesting all my commits be removed (to catch e.g. forks) only seems to have duplicated my org request, according to that.

@Toby222

Toby222 commented Aug 6, 2026

Copy link
Copy Markdown

I mean they have that /data/outputs.jsonl file tat I suppose they use internally to sweep and delete data?

Yeah my issue requesting all my commits be removed (to catch e.g. forks) only seems to have duplicated my org request, according to that.

That's how I noticed, actually lel
I did the same thing and noticed it just removed my user again, so I checked how it was handled

@addisoncrump

Copy link
Copy Markdown

I have also sent a GDPR erasure notice, will let y'all know how that goes.

@Toby222

Toby222 commented Aug 11, 2026

Copy link
Copy Markdown

I have also sent a GDPR erasure notice, will let y'all know how that goes.

Any update? They seem to have given up doing anything

@lele394

lele394 commented Aug 11, 2026

Copy link
Copy Markdown

I have also sent a GDPR erasure notice, will let y'all know how that goes.

Any update? They seem to have given up doing anything

Sent one on the 05/08 too, no answers yet.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.