The only thing faster moving that AI these days are the goalposts. Three years ago we would have been amazed if models were able to produce anything, now we have the luxury of nitpicking. Even the worst entries in the benchmark are quite impressive.
I remember getting wound up about latency and server issues playing counter-strike in the early '00s. At the same time though, it was hard to justify being angry because playing a multiplayer game with friends who were scattered all over town was something that had to be real magic.
I guess the wow!->adjust->complain->wow!->... cycle is endless as a human
No one asked for faster horses, they still became obsolete when cars came. Nothing new
Things mature, and expectations grow appropriately. That is true of more than just LLM performance.
Welcome to human nature.
Using reference images is a huge step for this sort of thing. The text-only approaches I've seen before were never going to be that good even with "perfect" AI, simply because describing 3D objects in text is not something that anyone is really any good at.