You have to be rather naive if you don't think these companies don't simply train on your prompts with or without your consent. They literally scrape everything - legal or not - and claim its fair use to train on, including straight piracy
The idea that they'll steal from everyone except you is just wishful thinking
I pay for a subscription to Duck.ai mainly because I don't want to be constantly fighting my vendor to protect my privacy.
Microsoft already did a rug pull on me and opted me in to training months after I signed up with Github Copilot. It exhausting and ultimately futile to monitor these companies.
It's not guaranteed that Duck.ai will continue to uphold its promise of not training on your sessions — if the company gets bought by Microsoft, it's only a matter of time before the switch to "you can opt out at any time". But since privacy is Duck.ai's brand, it will be somewhat harder for them to hide what they're doing should they betray their customers.
I also don't actually trust that Duck.ai sub-vendors OpenAI and Anthropic will uphold whatever contract they have with Duck.ai — the whole AI business model is built on lawless consumption of others work.
We'll ultimately have to run our own models locally, because it's impractical to defend against untrustworthy AI vendors.
This is a hugely misleading editorialised title.
The page title is "Can I opt out of my input or output data being used for training".
Right at the top of the page it says "In certain cases, your input and output data (such as conversations, documents, and other user-provided content) may be included in Mistral’s model training programs. You retain full control over this processing and have the right to opt out of these programs at any time."
For all the people here alleging that AI companies will train on your data even when you as a user explicitly opt out of training and their terms say they will respect that, etc - do you also believe that within a few years of them having harvested your data, you could perform "knowledge probing" on their models by prompting various questions that determine if they can near-verbatim reproduce your unique data, and then have enough other people do the same that you can then just launch a class-action lawsuit? Because if not... I've got a startup idea for you.
I could be mistaken, but wasn't Mistral openly championing themselves as a company and EU option that wouldn't do this/didn't do this?
I'd like to provide extra context to help with training. Like, here's my codebase and the logs and the docs, feel free to ask me questions about it... Whatever makes the next version more applicable to the problems I'm trying to solve would be excellent.
I just wish I could force them to share it with their competitors also.
There are so models that beats all of Mistral models, plus you can run many of them locally. Why would anyone run Mistral?
Submitted title was "Mistral now trains on user input by default, except on enterprise tier". THat's good information (if true) but best suited to a comment in the thread rather than the title (https://hn.algolia.com/?dateRange=all&page=0&prefix=false&so...).
We've changed it to the article title now per https://news.ycombinator.com/newsguidelines.html.
Summary of the different scenarios
Service: Vibe Plan: Non-Enterprise Default: Opted in Opt-out possible? Yes
Service: Vibe Plan: Enterprise Default: Opted out Opt-out possible? Yes (admin-managed)
Service: Mistral Studio/API Plan: Not specified Default: Not stated, I assume opted in Opt-out possible? Yes
> In the Admin panel[links to the admin panel], open the Privacy menu in the left-hand navigation bar.
Why not link to the https://admin.mistral.ai/plateforme/privacy panel directly
I just checked my admin dashboard and training on my data is (still?) off. Maybe I already set it to off when I signed up, I don't really remember.
I remember this being the default for a while already. When I signed up and got an API key, I remember I did turn off the training option specifically, so it must have already been a default to have training on input on.
Question: when AI companies say they don’t train on your input or output, do they mean that they don’t train on an extremely simple AI rework of your input or output as well?
> Vibe: users are not opted out by default
> Vibe (Enterprise): customers are opted out of training by default
So they made it opt in for enterprise (how it should be), but intentionally made opt out for regular user.Basically saying "screw you: to regular users.
Any self respecting user should stop using them.
This gets posted literally moments before I was about to pay them after ditching Claude. Thank you! I hate data collection in paid products.
I think I will be using Kagi Ultimate for the inference UI, so the data is somewhat anonymized before being collected.
Great news if you want your LLM to be horny, I guess.
Realistically, what are the risks? I get not wanting info you provide to AI being used against you in the future, like if you are gay and move to a country where that’s illegal (or live in one where it becomes illegal), but is “training” an actual risk here? I would assume that PII is redacted, and so beyond redaction failure, the real risk here seems to be to intellectual property and not to personal privacy.
Edit: Does “use for training” include “we store all your chat logs forever tied to your identity”?
I'd be interested in some legal/GDPR takes on PII handling in prompts. If someone enters PII into a prompt, and Mistral retains it for training, is it sufficient for them to say "don't enter PII into prompts?"
Of course with Claude and so on this bothers me too, but it doesn't seem like there's any real recourse under US law. But I would hope that "oh you shouldn't enter PII" isn't going to cut it under European law, that if I say "don't store my prompts, they include PII I don't want you storing" should be sufficient here under the GDPR and Mistral shouldn't be able to just store it anyway.
LeChat when it was launched as a consumer product was always going to be a data acquisition play.
This is them just making it very clear and disclosing as per European rules.
My bet is that mistral models will suddenly start to shine.
I assume they do this no before releasing new models. It feels like they expect higher user influx from their new models.
Just to achieve a great ideal: MEGA Make Europe Great Again.
> users are not opted out by default
Of course because you can’t be opt out by default that would be called opt in.
Come on, think how fast Grok evolves after SoaceXAI acquired Cursor. I just believe if you can't get enough real data from practice then you can't train a good model. And for Mistral collecting data is a must step no matter what approach it takes.
Most commenters here clutching their pearls as if Claude and Gemini Pro didn't do it already. In the latter you (as a paying customer) can't even store the chat history unless you agree to their 'improvement of services'. Do you have all your accounts paid for by the enterprise or you never check the settings?
Very disappointing. I really get the sense that all AI companies and anti-privacy parasites on society.
Very disappointing, but who doesn't train on user data? It's the only moat they have
i don't get the legal aspect of this
if you sign the contract (click I agree on Terms of Service), and it says that they do not train on the data, by what right can they backtrack on that?
Maybe it is assumed that they notified you and you have the right to terminate the contract? Relying on some clause where the contract can be updated at any time. Or terminated at any time (with the presumption that a clause change is a termination and an automatic signing of the new contract, but that's weak)
I wonder how much the progress is slowed down just because most companies don't train on consumer convos. Assuming they really don't.
I'm sure that the 3 users they have are very pissed about this news.
I'd send them a GDPR request, if I had an reason to use their fifth rate cloud models.
the rapid enshittification in LLM hype cycle is something else.
Got to love the private equity parasites ruining everything for the sake of profit.
I mean they all do it right ? Mistral is just the first one to publicly say it.
[flagged]
They wouldn’t dare train on language inputs from the wild, but they will happily strip out any code snippets you’re providing and sell/train on that à la GitHub/Copilot. File uploads too
[dead]
[dead]
[dead]
[dead]
Oh come on. So you should be fine with this because "hey we're the friendly europeans..."?
European innovation right here
you can't make this shit up
Context: After careful research our organization preferred a European partner with good central privacy controls. We landed on Mistral, after being disappointed that the Pro tier was opt-in to training on prompts by default we switched up to the Team tier which provides an organization dashboard with some relevant settings. As we did that Mistral changed these options and the Team tier was now also opt-in by default and at the same time seemed to have lost the ability to centrally disable training on prompts for your entire organization. This even caused some of our (testing) prompts to be used for training (which Mistral removed after we expressed our disappointment).
For some time these pages conflicted with what our users reported (they said that in contrast to what I stated to our management they found they were opted into training on prompts by default as per their own privacy page). Mistral just now corrected their docs. I'm not sure how long the conflicting situation has lasted, but at least for several days.
For contrast: Claude disables training on prompts for organizations starting from the 18 euro tier [0]. As a European I'm disappointed.
[0] https://claude.com/pricing#team-&-enterprise