logoalt Hacker News

thereitgoes456yesterday at 8:11 PM0 repliesview on HN

I see, it’s a great point. I know some evals actually do use LLMs as a judge (e.g. those that try to measure debate skill), though the ways AI can try to cheat its way through every benchmark now are astoundingly varied.