I'm running it in router mode, but people on Reddit were recommending people use llama-swap instead. Am I missing something by using router mode?
I believe it's useful for running multiple llama.cpp forks at the same time (e.g. a model you want requires special patching)
I believe it's useful for running multiple llama.cpp forks at the same time (e.g. a model you want requires special patching)