logoalt Hacker News

boredatoms • yesterday at 9:11 PM • 0 replies • view on HN

It also depends on the runtime, vllm is unbelievably slow at model loading compared to llama.cpp