logoalt Hacker News

ijustlovemath • yesterday at 2:07 PM • 2 replies • view on HN

I just think that with the vast amounts of compute involved and the tendency to reward hack, we can't assume the steps towards that formalization are without error until full human understanding of the formalization.


Replies

mkarrmann • yesterday at 2:25 PM

Repeating the above comment: most of the statements were already formalized prior OpenAI's work. So no, the statements were not "reward hacked".

➕ show 1 reply
engineeringwoke • yesterday at 7:54 PM

Humans have near-identical incentives to cheat compared to LLMs