Was the last time you've used LLM two years ago? Current SOTA makes good judgement calls and works well even with shitty prompts. And current SOTA is the worst these models will be.