I guess a weakness of the common benchmarks (artificial intelligence, gertlabs, and others) is that, from my understanding, they don't really cater to each language's specificity. If a language such as CL or Clojure boasts a superior REPL experience, that's entirely ignored by benchmarks, because they just look at the result of generic prompting. So maybe instead of writing countless blog posts on the superiority of such and such language, we should work on improving the benchmarking methodology. Then we would get more useful results.