Not sure why you are being downvoted, but this is definitely a limitation of current models and difficult to optimize for. Anything you are not specifically optimizing for is essentially unconstrained: the model could learn it by pure chance, but the likelihood is exceedingly low.
Techniques, like the newly announced RL-XAR from Meta [1] are being developed that will likely improve reward models and guide RL training to optimize for metrics like software architecture that are hard to verify otherwise.