Yes but that's because of the scaling laws for transistors. ML models seem to get better the bigger they are. If you want to compress the world's information, you need to look at all the information in the world.
Qwen 3.6 27b is way ahead of Llama 70b.
Qwen 3.6 27b is way ahead of Llama 70b.