Anthropic expends tons of compute and effort on understanding internal model states [1]; this kind of thing is right up their alley.
[1]: Recent example: https://www.anthropic.com/research/global-workspace