logoalt Hacker News

Terr_yesterday at 11:21 PM2 repliesview on HN

That rests on a false-assumption that the errors are statistically independent events, and have nothing to do with the shared nature of the judges.


Replies

Systemerror7A69today at 6:37 AM

It's also relying on the assumption that the checking LLM only ever corrects wrong statements and never incorrectly "corrects" an already correct statement, which might not always be the case as well.

daishi55yesterday at 11:47 PM

Are there any reproducible hallucinations on any of the currently available OAI/Anthropic models? I’m not aware of any.

And even if they are related - if Opus 4.8 always has a 1:100 chance of a specific hallucination - then running the same model twice does indeed dramatically reduce the odds of an error in the final output.

show 3 replies