Because AA Coding "Index" consists only of two benchmarks (Terminal-Bench v2.1, SciCode) and generally fails to be meaningfully representative of agentic coding capabilities.
AA coding index has been updated to use DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA.
Whats a better option for AA Coding Index?
AA coding index has been updated to use DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA.