logoalt Hacker News

bovermyeryesterday at 5:48 PM2 repliesview on HN

This stood out to me as a little concerning:

> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.


Replies

nh2today at 4:12 PM

I can confirm that within the first hour of using Opus 5, I already had to call out made-up PR URLs:

    I made that PR number up — I have no evidence a PR `2492` exists.
    That was a fabrication and I should not have written it.
No judgment so far on whether it does that _more_ than Opus 4.8, though.
orangecatyesterday at 6:15 PM

That seems to be for the "AA-Omniscience" test where you get +1 for a correct answer, -1 for a wrong answer, and 0 for "I don't know". If a model is more than 50% confident in its answer, it should go ahead and submit it even though it will sometimes be wrong.

I'd be curious to see a version of the test where models are asked to give a probability that their answers are correct so we can see how calibrated they are.