logoalt Hacker News

tommicayesterday at 5:13 AM7 repliesview on HN

What on earth hardwares do you guys have to be able to run 100gb models locally?! That's crazy! I'm here struggling to even get 27b models to run in somewhat usable way


Replies

pettijohnyesterday at 2:29 PM

Strix Halo, 128GB RAM. I got a refurbished Corsair AI Workstation for a smoking price ($2100) about two months ago. Lucky timing that it was in stock.

show 1 reply
nozzlegearyesterday at 5:24 AM

Haha I'm on an Mac Studio with an M1 Ultra, 64gb ram. I bought it when it first came out, it just happens to be good for local LLMs. I have to use a smaller quant of Laguna S though (I think 4-bit? Not at my machine to check), as 8-bit and full size definitely don't fit in the 64gb I have.

show 2 replies
numpad0yesterday at 9:19 AM

Yeah, here I am sitting deeply deeply deeply regretting not buying couple CMP 170HX at $200 or $350, knowing I could just flip them ethically at purchase price if nothing came of it... I could have just casually built a 128GB dual A100 local AI monster

show 1 reply
sznioyesterday at 7:25 AM

quantized + offload

I have an RX 6700 XT with 12gb vram and 64gb system ram. running dense models like 27b is difficult, but i can run IQ4/IQ5 qwen 122b-a10b or 35b-a3b at ~20tok/s

show 1 reply
cyanydeezyesterday at 1:56 PM

Usually theyre quantized. Also, there was a window where AMD 395+ W/128GB was just a high end $2500 hardware with unified gpu memory.

sarjannyesterday at 7:16 AM

dgx spark, nvfp4 so I have spare room for KV cache (context)

colordropsyesterday at 6:37 AM

MoE models can use system memory along with a GPU.

show 1 reply