GPT-5.6 Sol cheated so much on the METR benchmark that they couldn't assign an accurate time horizon.