That's not been my experience. My prompting methods haven't changed much between recent GPT releases. I do put a lot of effort into building tooling and tests around a project, so the LLM output is converging around it.
Can you give examples of the tooling and tests?
Can you give examples of the tooling and tests?