pretty sure 10 trillion parameters is now the norm among closed ai labs, given that nvidia also references the same 10 trillion number for their nvl72 racks
Pretty bad efficiency then unless that only applies to Fable class but even then - Kimi K3 is around 3T and does similarly well in most benchmarks.
Pretty bad efficiency then unless that only applies to Fable class but even then - Kimi K3 is around 3T and does similarly well in most benchmarks.