>This one is 552B so almost twice that, so the huge gains in benchmark scores make sense - it's not really flash anymore, imo.
It uses fewer active parameters, though. (8B or 14B instead of always 13B)
So ... flash indeed.
200B of those 552B is PLE, which works more like a database that is read for each token, thus can be offloaded to a fast SSD.
200B of those 552B is PLE, which works more like a database that is read for each token, thus can be offloaded to a fast SSD.