I use llama.cpp w/ llama-swap
https://github.com/ggml-org/llama.cpp https://github.com/mostlygeek/llama-swap
FYI, llama-server can now be run in router mode so llama-swap is probably only needed for more exotic scenarios.
Thanks for the link to llama-swap. Didn’t know about it and will definitely install it.
FYI, llama-server can now be run in router mode so llama-swap is probably only needed for more exotic scenarios.