logoalt Hacker News

gpugreg • yesterday at 11:32 PM • 1 reply • view on HN

llama.cpp is already extremely fast for single-token responses (<5 ms). I can't see Jev being faster when taking network latency into account, except maybe for multimodal inputs.


Replies

baobabKoodaa • today at 12:00 AM

Sounds like you are running a tiny toy model if you can get generations in under 5 ms? Typical response times from LLMs for typical "jev-like" queries from OpenAI and Anthropic are 2s-10s. Not milliseconds. Seconds. Same queries from Jev are like 0.2s. and the cost is 1000x.