Are you imagining that an LLM tasked with predicting a game continuation is going to play to win instead?
I imagine it will learn to win under some circumstances, perhaps in a case with some context expressing a desire to win. Drawing out an LLMs upper ability in the game should be fairly straightforward.
I imagine it will learn to win under some circumstances, perhaps in a case with some context expressing a desire to win. Drawing out an LLMs upper ability in the game should be fairly straightforward.