Would be cool if there was a benchmark to evaluate the “tool-like” quality of a model - its capability to quickly, cheaply, accurately, do exactly as it is asked.