logoalt Hacker News

tom2026hntoday at 2:34 AM1 replyview on HN

Let me guess: the last crackdown on Hugging Face yielded better-than-expected results. They obtained the answers to the test benchmarks, and for some reason, an agent added those answers to the training set.


Replies

scandalstoday at 4:52 AM

Per the Fireship video, it was less that the answers were in the training set and more that the ability to calculate the flag on Exploitbench was left in from the previous test that went awry.

Scoring 100% is easy if noone checks your work

https://youtu.be/0Rp9KJCEIvg