Max effort is the way to give the highest perf, but not highest perf/$. Having claude (or other models) use a lower effort can often be 80% as smart but get to the results 10x faster for the problems where it works.
We've seen models perform worse at higher efforts in our vuln detection evals. For example IIRC gpt 5.5 and 5.6 both scored better or high as compared to xhigh.
We've seen models perform worse at higher efforts in our vuln detection evals. For example IIRC gpt 5.5 and 5.6 both scored better or high as compared to xhigh.