logoalt Hacker News

mirekrusintoday at 10:57 AM1 replyview on HN

Just run /goal to optimise it and you should be good in less than an hour. Also best to use models that support speculative decoding.


Replies

LoganDarktoday at 11:51 AM

Optimize llama.cpp? Hmm.

WRT speculative decode, basically zero finetunes keep it. I'm testing with some ridiculous abliterated amalgamation so spec decode has been gone for most of its ancestry.

Fable recommended n-gram speculation so I'm working on that now.