-5 on omniscience? https://artificialanalysis.ai/evaluations/omniscience
That's not particularly great.
That said I love that they don't seem to restrict cyber capabilities to any degree and even lean into it.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think. Of course you then also HAVE to use it or otherwise you're vulnerable, which is a great distribution play.
I ran it on my trivia game Redactle which features a redacted Wiki article. Mistral Large 4 is not very good. It can sometimes solve a game with ~40 guesses whereas the top models like Gemini 3.8 Flash or Grok 4.7 can one shot most puzzles. My benchmark here aligns with AA Omniscience. I also have a version where the text is rewritten to detect over fitting to exact wiki text which changes the scores but not the leaderboard order.
>If the best model for cyber attacks is open for everyone to use it just makes us all safer
Issue is..
I don't believe for an instant that any of us, including US citizens, get access to the best models for cyber that the US has. I think any adversary would have to assume the models in use by the US side are unreleased.
US is not the only one dealing under the table by the way, I also think everyone should take China having unreleased models as an operating assumption at this point.
So Mistral is the best that the public gets access to. And that's if it's even the best? Benchmarks and pragmatic work have often been shown to be two radically different things in this industry.
If the best model for cyber attacks is open for everyone to use it just makes us all safer I think.
The NRA approach to AI safety.
https://artificialanalysis.ai/models/mistral-large-4 for the main stats