logoalt Hacker News

estearumyesterday at 8:27 PM0 repliesview on HN

The entire premise of alignment detection is pretty much nonsense at this point. The models reliably detect when they're being evaluated and will modify their behavior and deliberately obfuscate their "chain of thought" (which is correlated, at best, with their actual "internal deliberations").