If you are training on data labeled by frontier models, how do you expect to exceed the performance of frontier models, other than in the cost dimension by recognizing simpler problems and routing to cheaper models?
Different frontier models are good at different things! We'll be the ones combining them optimally.
Different frontier models are good at different things! We'll be the ones combining them optimally.