Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Also strong on cyber benchmarks (better than all chinese models), so this is a good defender model.
Lots of people shitting of Mistral for no reason imo. These are pretty good numbers across the board. Definitely good enough to use as a daily driver over other llms, if you have moral qualms with the others. For certain use cases, like cyber security, this may be the go to model.
I like to make fun of europe, but there's lots for mistral to be proud about in this release imo.
> Impressive vision benchmarking. If the vision model is truly as good as astra, that would make it best in the world.
Sure, it leads in one unpopular benchmark with internal numbers.
But look at the third benchmark. It only has Sol 5.6 from OpenAI and it shows a 12.8. I don't know this benchmark but AA's GDP.pdf listing has Sol 5.6 Max at a 27, even non-reasoning beats the 12.8. It's extremely weird to cherry pick Sol 5.6 and then also lie about the published third party benchmark score. I'd love AI companies to stop lying about this stuff.
AA - https://artificialanalysis.ai/evaluations/gdp-pdf?models=cla... Mistral announcement - https://mistral.ai/_astro/multimodal-benchmarks---gdp.pdf---...
Curious, what do you like to make fun of Europe about?
I'm still embedded in the OpenAI ecosystem, but man, when I try out Mistral, it is snappy! Like basically instant responses that seem faster and better in quality than ChatGPT on instant mode. I'm impressed.