Hopefully this wakes people up from this addiction to benchmarks when discussing various AI models. Different model families have strengths and weaknesses in various domains, but those are never discussed.
> but those are never discussed.
They are literally frequently discussed and it's why there are different benchmarks for different domains.
> but those are never discussed.
They are literally frequently discussed and it's why there are different benchmarks for different domains.