At the end of the blog post we get this nugget.
> The pace of progress from here will be fast. Stay tuned.
Is the reason for the massive gains in certain benchmarks due to distillation from the other lead models hence the slightly "under" pattern seen in the comparison charts?
What experience have people had with Mistral models for cyber research? are they as constrainted as Anthropic models? I use claude day2day but have the need for a model with fewer guardrails to pentest our own APIs.
Excited to see this! Nice that they are saying this is just a first step.
Give them more compute!
I tried testing it, but reasoning effort parameter doesn't seem to be there and output is sort of broken because it reasons directly in the output tokens...
Awesome!
All things considered I'm more inclained to pay for EU-based AI in the end (supporting local company and most likely being more aligned with EU regulations…)
le chaton fat is real, my life is complete. Benches look crazy good for 1T.
no hugging face link :( ... but hey its on openrouter yay
Here are my results
https://dach.peerbench.ai/compare?models=mistralai%2Fmistral...
Looks like a bit better than the recent Kolibri-1 but still below Qwen3.8 27B
Seems like they've finally made a genuinely competitive model since the original LLM craze, congrats to them! Glad to see some diversification in open weights providers.
This sure is nice, but I've had less than satisfactory results with GLM5.3. I'd like Mistral to compete with Qwen3.8-Flash-Next a 120B class model that IMO is the first model that I can use for serious coding while running it locally.
I estimate it's coding ability on par with opus 4.6 (but opus definitely beats it on factual knowledge) Still it's a genuinely useful model, when everything else except Anthropic's models (and for only 3 weeks after it came out Google's Gemini 3 pro, before it got merged) are not to me.
I'd live to have one like that but EU made.
What's up with the name? It reminds me of my teenage self trying to speak in funny memes.
Without exaggeration, given a choice between models, I would pay for Mistral's model over Anthropic's based on the name alone, completely ignoring features or other technical considerations. The name is playful and is such a refreshing contrast to Anthropic's (and OpenAI's) doomsaying, scaremongering, and god-posturing.
Does it have the ability to capture the market like OpenAi or Anthropic ? My point is, Regular/Average users of AI do not really care about benchmarking. Marketing decides which company makes it to the phones or PCs.
if it's not available yet why have a 'try it today' header at all?
> "Try it today" > > There is still more to come. As we work toward releasing the weights, we will share further details on the model architecture, additional benchmarks, and our post-training methodology.
Since the Chinese companies publish their research it would have been odd if Mistral didn't start catching up.
I love the name! Teasing the ones making fun of them.
> Trained from scratch
How are they training without pirating the Z library corpus and all that?
I really hate this open weight but closed science approach. These companies just take from academics and all the Chinese companies that are doing good science, but without understanding the recipe it makes it hard to know where the failure points will be until your agent accidentally commits a crime.
Mistral doesn't publish the science.
Glad to see progress, despite the ever-increasing sabotage by the EU bureaucrats
So about 2 or 3 generations behind, just like they were a year ago?
Number one in Sovereign AI. Join our Discord.
Half price on open router right now
Is there consensus on if this was https://openrouter.ai/stealth/space-bunny-alpha ?
at 200M tokens for the full AI suite run its not token efficient at all
We actually got Le Chaton Fat before GTA 6
dont take it personally, i just dont understand why to release a model that is not showing new strong capabilities, why would anybody use this model and not Claude Opus.
I'm glad they're keeping at it!
I thought lechonk motto was just a meme!
Previous one is barely in top 50 on arena.ai
A bit disappointing to see it still lagging behind Chinese open models. Those Chinese models are pushing proprietary models to raise the bar, but we need equally strong non-Chinese open models to challenge the Chinese ones in turn.
Why is distillation weird?
I mean no disrespect but these are terrible numbers or am I missing something? It seems like Mistral continues to only be relevant for people that want a model trained in Europe. Too bad.
better than K3 and DS4, cool
Not bad only two major releases behind top tier. Edit : checked its rather 3 generations behind . Oh well
Looks like it's about a year behind still. i.e. its intelligence is behind models from roughly a year ago.
where does sit on the pareto distribution compered to Le Chaton Fat?
Amazing! just tested
impressive release this time by Mistral. bullish.
>> Unofficially ML4, very officially: le Chonk
Honestly just nice to see a leader in this space not take themselves so seriously.
1T parameters -- ugh, open models keep getting bigger and bigger! Running them at home is getting ever more unattainable, especially for those of us with bandwidth-poor hardware like Apple silicon -- please continue releasing smaller models, too!
europe finally getting into the race here.
Where can it be tested?
Massive fumble not to call it “le chaton fat”.
Too bad this got marked as a dupe, as it actually has benchmark info unlike the other page which is just docs.
The weird thing is how worse they are at things like coding than other open weights. You'd expect them to at least distill coding from other open weights to match them.
Pretty impressive. I genuinely wonder how Mistral hires talent when their salaries are so terrible. Guess there aren't many better places to work in Europe.
I tried a quick abbreviated kimbench on their playground before bothering to do the whole thing.
Maybe I didn't really select mistral 4? Either way, failed completely on question 1 and the next 2 questions were completely off base too. I didn't bother to finish.
Not suitable for my purposes I don't think.