Okay, so we’re using the same process — but your original message seemed to imply a sort of one shot no refinement iterations — which is what I was responding to as unrealistic (e.g. make a spec let goal run artifact is perfect)
Of course, all I’m saying is that you need to refine your sample! For instance: the allocation architecture is not correct, and one has to run a bunch of performance investigations and resolve it.
My responses are intending to convey that I don’t believe this is possible, no matter how good LMs get — and it seems like we are in agreement.
We're mixing two different things here though. You're saying that you would need to correct the compiler because the agent did wrong, I'm saying that you'll need to correct the compiler because your specification will be wrong. The agent does the correct thing, but the correct thing was wrong in some way, if that makes sense?
If your goal with building this compiler was performance, and this wasn't part of the initial specification, and the agent didn't assume it had to, is this what you're saying is a failure on the agents side?
There is no distribution to fight, is my hypothesis at least, if you're just a lot more clear exactly what you expect up front. Hence the whole "To correct those behaviors, you're going to write tools and skills" thing isn't even needed in the first place.