After my and many others' experience with Claude Opus 5 being hot garbage for normal agentic programming use, I'm not sure benchmarks mean much anymore.
Much less Grok's, since they have a reputation for unethical benchmaxxing, among other things.