logoalt Hacker News

Topfitoday at 4:41 PM1 replyview on HN

Honest question, how do you assess models this quickly? What metrics are you using? Would love to get my suite from multiple days and hundreds of prompts down to minutes. Got a few first pass tasks I run upon release for an initial experience, but those only work because even Fable and Sol fail despite objectively correct solutions existing, so it works because most models fail, but then, those are consciously not enough for coding, tool use, adherence or task specific inference and assessment…


Replies

satvikpendemtoday at 5:11 PM

What are you working on? That can dictate which models are best.

show 1 reply