logoalt Hacker News

plumb_samjitoday at 1:33 PM0 repliesview on HN

Interesting exploratory comparison, but I be cautious about treating it as a model benchmark With only three runs per model, the results are highly sensitive to randomness