In my view, the useful like for like comparison is to the entire system that people actually use. Nobody uses a "naked model", so what is the point of these benchmarks that use them in that way?
I think the benchmarks should be trying to use realistic harnesses for both proprietary and open weight models.