The apparent advantage is exaggerated by them running Astra at six different effort levels, and almost everything else at just the maximum available effort.
I don't really understand why they keep doing this. Either run and report everyone at multiple effort levels, or run everyone at only one.
But alsi, token efficiency seems pretty artificial? For example tokenizers are different from model to model. The cost/perf Pareto frontier seems a lot more meaningful (and Astra does very well at that too, just to be clear. It seems to be a great model.)