"More tool work per model turn" could reduce the number of cache reads (or even cache misses) and associated cost?