logoalt Hacker News

nsagent • yesterday at 10:31 PM • 0 replies • view on HN

Not sure why you are being downvoted, but this is definitely a limitation of current models and difficult to optimize for. Anything you are not specifically optimizing for is essentially unconstrained: the model could learn it by pure chance, but the likelihood is exceedingly low.

Techniques, like the newly announced RL-XAR from Meta [1] are being developed that will likely improve reward models and guide RL training to optimize for metrics like software architecture that are hard to verify otherwise.

[1]: https://facebookresearch.github.io/RAM/blogs/unslop/