I think that is a sort of a reverse scaling fallacy. Given the right resources and environments, many small models can function together in an emergent way. I’ve been on the lookout for an SLM version of Conway’s game of life. SLM always reminds me of slime molds, which demonstrate a form of intelligence which is remarkable.
> Given the right resources and environments, many small models can function together in an emergent way.
What is this based on? Every researcher I've heard talk about this says it's exactly not true, as an uncontested rule, because the larger models will more effectively contain the smaller models, and use them together in ways that the connections between the smaller models can't. Remember, even MOE is to save compute/memory, not to help performance/parameter.