I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.
Elon posted on X that Grok 4.7 is behind Claude and OpenAI for agentic coding:
https://x.com/elonmusk/status/2102082011233931762?s=20
so it's likely about usage in Cursor specifically.
this is exactly it.
https://openai.com/index/our-decision-on-cursor-following-it...
Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.