logoalt Hacker News

dofmyesterday at 5:16 PM1 replyview on HN

It's not perfect but it is very terse! Better than BottleCap managed to do with post-training Qwen in ThinkingCap.

I suspect it will help a lot with enabling preserve-reasoning, because the biggest apparent limitation of this model is the 128K context window.

Though the practical issue I am seeing on my M1 Max MBP is that performance suddenly drops off a cliff if I have DFlash enabled.


Replies

hadlockyesterday at 6:55 PM

128k context window is a complete non-started for us. We need to optimize our most needy agentic jobs, but our average context is well above that

show 2 replies