logoalt Hacker News

gesshatoday at 6:17 PM0 repliesview on HN

5090 is plenty for the Q4_K_M quantized version of 3.6 27B with reduced context size.

I run it on a 3090(24GB) and 64k context using GGUF format and llama-cpp. Double 3090 gives you 128k, quad 3090 gets you to full context - 256k.