logoalt Hacker News

siavoshyesterday at 5:06 PM2 repliesview on HN

What's everyones recommendation for one to run a good local LLM model on?


Replies

manmalyesterday at 5:14 PM

1-2 RTX5090 will be better value than Macs because they have the memory bandwidth for somewhat fast local inference.

bigyabaiyesterday at 5:11 PM

If you don't want to wait for prefill, you're going to want a CUDA dGPU system.