"Not really news" that a model from a smaller lab can beat weeks-old Fable 5.1 at 70% lower cost? What a time to be living in.
no, I mean not really news specifically on the number for terminal bench 2.
They did reinforcement learning on Kimi K3, probably specifically targeting the benchmarks.
Eh, it isn't news. I mean, it's just an echo bit of news to the great K3 release.
It’s very far from Fable on benchmarks that matter, like TB4.