There are some benchmarks I personally trust. But the most reliable indicator other than public benchmarks is the quality of products people are working on, and the model they use for the work. Not a benchmark scoreboard or a one~few shot demo, but the product they have been building for weeks. The reality is, even the diehard advocates who build and sell tooling for open weights, still use proprietary frontier models for their jobs.
I've used Kimi K3 for my (hobby) Linux kernel work recently, because unlike the offerings from OpenAI and Anthropic it didn't flatly declare everything to be a cybersecurity issue.
(And, yes, I know we should blame C. C turns every bug into a cybersecurity nightmare.)