No mention of training/test/holdout dataset splits anywhere in this article - so how am I to know whether or not this is some overfit result?