Just putting more details in, and giving it ways to check itself brings massive improvements. I often ask model to set up test for the problem before actually trying to solve it and it improves it a lot, both in how much babysitting is required (if it can test it itself quickly it goes faster), and the fact the context now contains more detailed description of the problem that came up when making tests.
Turns out TDD is far better for robots than humans, who knew