IMO it's not. It's benchmarking GPT 5.4 and Opus 4.6. It's also missing Claude Code... one of the most popular harnesses (the most?)
No Pi, no Aider either.
No Pi, no Aider either.