logoalt Hacker News

fathermarz • today at 5:14 AM • 1 reply • view on HN

I’m surprised at the reasoning behind using the most expensive model which likely led to the “doing too much”. I think it would have been nice to have A/B tested different hospital setups in order to maximize not only efficiency but to see where the strength of each model started to emerge in a given role.

Fable feels like overkill for this also.


Replies

rafiss • today at 6:00 PM

We don’t run the most expensive model everywhere anymore. We started on Opus for every stage, then reshuffled several times. We landed on Fable doing planning and code review, because code review is the last gate and its bug findings were worth the most. Opus does coding and plan review, and Sonnet runs the nurse roles. The planner can also route mechanical changes to Sonnet. The plan reviewer re-checks that choice, and any rework goes back to Opus.

That’s tuning from the cost data, so it's not just totally based on feeling, but it's not a controlled A/B test. I agree that would be worth doing, and I think our experiment here is helping us make the case that it's worth spending more time and effort on running more controlled evaluations.

➕ show 1 reply