Fast models is why I was hoping Taalas would get their butts into gear and eventually release a consumer priced card. I'd love to have a pcie card that screams along at 15k t/s even if on a heavily quantized 2026 level model forever.
Faster & cheaper tokens = more reasoning capability and more reasoning = better problem solving as far as I have seen.