https://taalas.com/ has done it already for a wildly obsolete model. 14000 tokens per second.
https://chatjimmy.ai/ is their interactive. Tiny context, very dumb, but absurdly fast. Imagine this as a tool call for claude code for trivial changes - the tool call from the harness takes longer than the execution.
Wow. This is absolutely wild. I didn't expect that.
If we get to anywhere near this speed for the equivalent of the current models... I don't even know what to think about that future.
Holy crap, I was not prepared for how fast it responded. I just wrote "Just wanted to see how fast you are! Can you write me a quick story about a tiger who lives inside a block of cheese the size of a house?"
I pressed Enter, and the response was instant.
> Generated in 0.037s • 14,205 tok/s
This is unbelievable.
See Cerebras and Groq as well.
Wow! You weren't kidding,
I just tried it too and 14,098 tokens in .05 seconds, I barely blinked and it was done. There was no typing at all appearing on the screen. It just showed up.
https://chatjimmy.ai/chats/01dc66a4-4b1b-4dea-bb5f-926855e37...