I’m actually working on a small project that’s exactly this! Less quant so it’s only 150M parameters but this is amazing.
Soon ai in every lightbulb running Kubernetes
It is a bit of a bummer to see that the degree of 'compression' makes it a fancy llm noise-maker. It is still charming.
Gemma 4 when?
I'm curious how this would handle grammar checking on a basic word processor. Or maybe generate worlds for small text based games. I have no idea what the capabilities are of a cluster like this.