I need someone to run actual benchmarks between the two.
Only relevant benchmarks are those you make yourself, targeted specifically for your workflows. Anything else is just number go up on a pretty graph, and every model out there is probably benchmaxxed to hell on the public ones anyway. Keep yours private.
These models have gotten a fair amount of attention -- we're hoping it's enough to get them added to some reliable inference providers and OpenRouter, at which point we'll run them on our full benchmark suite.
Benchmarks are the BMI of model evaluation.
They may have utility in trying to look at the whole landscape of models, but are very misleading when it comes to making 1:1 comparisons or in developing confidence at to how a given model will deliver on your workflow.