logoalt Hacker News

objektiftoday at 2:45 AM1 replyview on HN

What are you basing how good they are on? Personal experience or some benchmarks?


Replies

a-t-c-gtoday at 3:55 AM

Benchmarks, we have internal ones testing reasoning fine-tuned v/s frontier + prompts

For some use cases it can be parity performance at 1/20th the cost up to exceeds at 1/10th the cost. Trade-off is ofc narrow applicability