logoalt Hacker News

johnfnyesterday at 4:18 PM3 repliesview on HN

How do you propose they check for something like this? They can't exactly ctrl-f the model weights for "Arc-AGI".


Replies

vickychijwaniyesterday at 6:53 PM

Anthropic expends tons of compute and effort on understanding internal model states [1]; this kind of thing is right up their alley.

[1]: Recent example: https://www.anthropic.com/research/global-workspace

phoghedyesterday at 4:56 PM

> they can’t possibly know or find out what was in the training data

doesn’t appear to be a very strong argument

turing_curiousyesterday at 6:01 PM

[dead]