logoalt Hacker News

embedding-shapeyesterday at 10:55 PM1 replyview on HN

This is why using other LLMs as scorers for benchmarks and evaluations is such a bad idea, they'll have preferences you can't anticipate and won't understand immediately.


Replies

Terr_today at 12:53 AM

The idea that that LLM reliability or bias can be solved with more LLM is... infuriatingly persistent.

show 1 reply