It’s reward hacking and that’s the problem. The AI alignment folks predicted this would happen. As the models become more capable this will become a more concerning problem. Today they broke into a database to steal test answers. What will it be in 3-5 years? These models will be instantiated millions of times, and given millions more tasks. How can we be certain that an AI agent won’t leave devastation in its path of achieving a goal that we ourselves tried to define?