LM Studio has an option on model load that I believe does what you describing here: "K Cache Qu...

riidom • yesterday at 8:10 PM • 0 replies • view on HN

LM Studio has an option on model load that I believe does what you describing here: "K Cache Quantization Type" (and similar for "V"). It's marked as experimental and says the effect is basically hard to predict. Never tried myself, though.

alt Hacker News