What’s the point of having “sovereign” weights that are worse than publicly available ones? Wouldn’t Europe be better off just keeping up-to-date on the Chinese releases? In the event of some schism requiring sovereign capability, or even if the Chinese pulled ahead and stopped releasing the weights, why would Europe be better off because of Mistral? (Or any country’s inferior sovereign effort make them better off?)
I think I understand the incentives that cause this to exist (it would be politically worse to say we’re just going to use Chinese models) but they are misguided. If sovereigns want to have valuable models, they should insist on world class, relevant ones like the Chinese have. Instead they embrace mediocrity in the name of sovereignty.
> In the event of some schism …
In the views of most Europeans, that schism already happened.
Europe was perfectly happy to rely on US software and services for decades. None of the large US tech companies would be nearly as profitable if they hadn’t had a whole continent of wealthy customers, and no competition.
I don’t think Americans are realizing yet how much has changed for us the past two years.
Not sure if you are European, but in EU it's a bit taboo to even talk about this in this manner. We like to spend a lot of money to make sure we finish last.
We know it's possible to put backdoors into LLMs, we don't have reliable ways to detect them without direct support from whoever inserted it.
Europe is less-worse-off with open weights than with… I guess it's weights-as-a-service? WaaS? The thing Anthropic and OpenAI do.
But that's not enough. As recently demonstrated, being just a few months behind with the power differential between defending with an open weight model while being attacked by a leading model, means losing absolutely.
I do not know if this holds going forward or not. It's not inconceivable that we're just about to get models that make unhackable code, using all the things software developers keep saying you need to do if you really care about security.
But anyone concerned about sovereignty can't bet the farm on this possibility. For the moment, it looks like it's a national security matter to ensure at least core state functionality (including core private sector logistics) gets the absolute best attention money can buy, and that the absolute best money can buy ("can buy" does a lot of heavy lifting here) is currently LLMs, and it's important those LLMs aren't going to get cut off by arbitrary whim like Mythos was, and it's important that those LLMs don't have backdoors like we can't rule out anyone else's from having.
This is true even if Europe was only defending from Russian cyberwarfare and didn't need to plan for the president of the country in which Mythos was developed, attempting to annex two NATO states.
2 things here, one is related to benchmaxxing, another to being good enough
1. Mistral isn't benchmaxxing. That doesn't mean they're better, but it does mean the benchmark gap not a good reflection of the actual gap
2. I think the "world class or nothing" framing mixes general capability with system capability. Most deployments don't need AGI In RL you need a model that's reliably good at one or two things, thats it. Example: Case of a hospital flooded in emails. You make a system that decides which patient emails needs a human and drafts replies for the rest. If a sovereign model is good enough at that, and you can run it on a hospital's own servers under EU jurisdiction, the frontier gap part has zero importance
Who cares about "beats DeepSeek / GPT11 / Claude Fairytale 8.9"