I would assume asic based llm would work really well. Why did it not go well?
https://chatjimmy.ai/ runs Llama 3.1-8B on an ASIC as a demo by https://taalas.com/ I believe.
That's quite a few parameters shy of today's trillion-weight behemoths, but it is fast.
https://chatjimmy.ai/ runs Llama 3.1-8B on an ASIC as a demo by https://taalas.com/ I believe.
That's quite a few parameters shy of today's trillion-weight behemoths, but it is fast.