What's everyones recommendation for one to run a good local LLM model on?
1-2 RTX5090 will be better value than Macs because they have the memory bandwidth for somewhat fast local inference.
If you don't want to wait for prefill, you're going to want a CUDA dGPU system.
1-2 RTX5090 will be better value than Macs because they have the memory bandwidth for somewhat fast local inference.