logoalt Hacker News

Luker88 • today at 2:31 PM • 2 replies • view on HN

Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.

Surprisingly useful as long as you can leave it running a couple of hours at the very least.

While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.


Replies

londons_explore • today at 2:43 PM

Remember that a hosted AI company only needs ~1000 bytes per context token per user (ie. 100mb per user for typical coding - VRAM during inference, and moved to regular RAM or SSD whilst running a tool call)

Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.

➕ show 1 reply
bitexploder • today at 4:25 PM

Highly recommend the RCO-GSQ quant by ITSA btw. At IQ3_XSS it is within one point of the fully unquantized model.

➕ show 1 reply