logoalt Hacker News

xlaynyesterday at 7:06 PM1 replyview on HN

For the impatient, I merged llama.cpp tentative branches to get it running here https://github.com/alainnothere/llama.cpp/tree/disk-cache-ev..., thing runs at 23.54 token/sec and my setup runs at high 30 the 3.8 dense 27B.

and this is the pelican from the iq4_xs model https://github.com/alainnothere/llama.cpp/blob/disk-cache-ev...


Replies

cmrdporcupineyesterday at 7:45 PM

If you've got a DGX Spark try my little engine: https://github.com/rdaum/eider/

NVFP4 quant