logoalt Hacker News

breckenedgetoday at 5:33 PM2 repliesview on HN

Do these evaluations get re run a few weeks after launch? I started doing that yesterday for our internal dataset and found Sol’s performance had regressed to be equal to Luna’s. Granted this was one run, but something I’m becoming more concerned about, the model providers want to quickly prove they’re the best, people switch to them, then they pull the rug.


Replies

mnickytoday at 7:40 PM

Well there is at least the degradation tracker from Margin labs for Sol and Opus: https://marginlab.ai/trackers/codex/

echelontoday at 7:37 PM

These tests need to be sampled continuously.

Moreover, the tests should be randomized somehow to ensure the models don't memorize the answer.