logoalt Hacker News

ryan_glasstoday at 6:39 AM0 repliesview on HN

Ollama uses the llama.cpp backend for inference. I find Ollama noticably slower. Llama.cpp has had a built-in webui (used as llama-server) for a long time now so have owned the user experience too.