https://chatjimmy.ai/ runs Llama 3.1-8B on an ASIC as a demo by https://taalas.com/ I believe.
That's quite a few parameters shy of today's trillion-weight behemoths, but it is fast.
You are correct. I think this is the bull case. It seems like this would be useful right now for some things (eg moderation).
You are correct. I think this is the bull case. It seems like this would be useful right now for some things (eg moderation).