logoalt Hacker News

hypfertoday at 5:06 PM1 replyview on HN

Yes to both.

The thing is that I can either use the q8 context, or have not enough context window, so I just live with whatever degradation there is. The same can be said about the IQ4_NL. I would not go any lower though.

As for the draft count, indeed that depends on what you do with it, but for coding, reverse engineering and that kind of stuff it does pay off in my testing, though 5 is really pushing it, but the 4090 has so much compute.

Last logline I saw scroll by right now had 47% acceptance rate for 4th and 28% for 5th, but not sure how representative that is. I think when tuning 3.6, I saw more like 33%? But not 100% sure.


Replies

hedgehogtoday at 6:22 PM

Same, I have one workload where on 3.6 drafting 6 tokens is the fastest setting.

show 1 reply