I guess the Riemann Hypothesis is an easy task then.
I still think that Anthropic went the wrong way. It would have been much more entertaining to ask the model to find a non trivial zero not on the line and give it encouragement. To see what exactly it will come up with.
Is it? How would you test an answer?
It is probably no coincidence that AI is exceedingly good at finding small counter examples. But for the Riemann hypothesis no such counter examples exist. And likely none exist.