I don't know why, but the benchmarks still fails to cover the difference between large models and small ones. The small ones are great for many things, including general coding, but the larger ones, like fable and astra, have some kind of intelligence that is not present in the small ones.
More parameters = more facts stored. Knowledges are almost incompressible, where strong reasoning only requires a 3B core or so.