The trick is to rarely switch, or switch at task boundaries. Often the conclusion of routing is actually "this one model is actually at the pareto front for this task, just use it always".
But then it's better to just not have a gateway switch models at all.
Just have the harness able to choose which model its sub-agents use, then tell it how to split up tasks and which models to use when doing so.
But then it's better to just not have a gateway switch models at all.
Just have the harness able to choose which model its sub-agents use, then tell it how to split up tasks and which models to use when doing so.