Better prediction being better compression is Shannon 1948, and the link to machine learning is MacKay 2003 at Cambridge.