Have you seen the news about decrypting the hidden COT in U.S. models? [0] The decoded logs revealed instances where Claude memorized answers to test questions beforehand while making its final output look like it had derived the answer step-by-step—hiding the memorization from the user.
pdf https://arxiv.org/pdf/2608.09867
for some reason I couldn't find any way to download it from that website.