I’ve been using v4 Flash 0731 a lot lately and you can’t beat the price performance. That said, it sometimes takes my prompts as more of a suggestion than a directive. I’ve found that introducing a reviewer subagent (even with the same model) helps push it back to what I’ve asked for. But makes every coding session a back and forth: “do X” -> “use a reviewer subagent to analyze whether you really did X as I asked”.
What model do you normally run the subagent on? You mentioned flash as well for that, but I wonder if a more 'strict' model would do a better job at pushing the main back on track.
deepseek-v4-flash-0731 has been awfully prone to infinite looping output in reasoning for me, and once it hallucinated in the middle of going in circles that I had instructed it to start using Yoda-speak, which… I don’t have any idea where that came from.