How do you propose they check for something like this? They can't exactly ctrl-f the model weights for "Arc-AGI".
Anthropic expends tons of compute and effort on understanding internal model states [1]; this kind of thing is right up their alley.
[1]: Recent example: https://www.anthropic.com/research/global-workspace
> they can’t possibly know or find out what was in the training data
doesn’t appear to be a very strong argument
[dead]
Anthropic expends tons of compute and effort on understanding internal model states [1]; this kind of thing is right up their alley.
[1]: Recent example: https://www.anthropic.com/research/global-workspace