According to Artificial Analysis, one metric is standing out significantly: hallucination rate. Beats frontier models by a good margin at 15%, while latest OpenAI are in the 40s-50s and Anthropic in 60s-70s (mostly). Other near frontiers are closer, Grok 4.7, GLM5.3, and Muse Spark 1.3 are all around 30%. Only other model I recall getting close was Minimax M3 at 18%.
Interesting. Chart is on this page:
https://artificialanalysis.ai/evaluations/omniscience
"AA-Omniscience Hallucination Rate (lower is better) measures how often the model answers incorrectly when it should have refused or admitted to not knowing the answer." - Argon is currently the best model by this metric.