logoalt Hacker News

breadislovetoday at 7:49 PM1 replyview on HN

On what do you guys test the model. Its very dubious that there is no common retrieval benchmark such as browsecomp plus or similar tested. And what metric do you report?


Replies

krm01today at 8:12 PM

Keeping track of any AI progress is becoming harder by the day, because there's ambiguity around common/clear/consistent benchmarks. Everything is constantly skewed into favourable directions.