logoalt Hacker News

kennywinkeryesterday at 5:02 PM0 repliesview on HN

Without access to reasoning traces, we can't know that - someone inside openai/anthropic would have to run the test - and we'd have to trust their results.

I would be curious to see how the open weight models do on a test like this - and then we'd be able to see the reasoning.