logoalt Hacker News

kamranjonyesterday at 7:57 PM1 replyview on HN

What size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases overall quality.


Replies

beacon294today at 7:06 AM

Try the llama.cpp fork by thetom. It's called turboquant after the technique