alt
Hacker News
beacon294
•
today at 7:06 AM
•
0 replies
•
view on HN
Try the llama.cpp fork by thetom. It's called turboquant after the technique