Yeah we heavily leverage coding agents for optimizing our kernels. Since it's highly verifiable and takes time to measure we often leave multiple running and improving performance on different model architectures.
Definitely still helps to reference relevant academic work as well, or even just encouraging the agent to make bigger structural leaps, otherwise it will often get stuck working on low impact micro-optimizations.
What kinds of prompting do you use to encourage structural leaps in your agents?
I don't let it commit the optimization unless it is larger than 5% and no regressions. If it can combine two optimizations and get above 10%, it is allowed.