I assume I'm missing some context because you sound like you have a lot more experience here, but what kinds of capabilities are there that you can't do something like `interface LlmProvider { }` with a few basic methods for attachments and chat and then just implement from there? I'd even say you could not keep the ones you're not using (say, GPT is your go-to for a year) and update the other ones when you want to use them and swap out the `LlmProvider` implementation with a DI'd different implementation at your application's entrypoint.
Assume the APIs are equivalent across all providers and have exactly the same schemas. You still have this massive problem of alignment between the various reasoning models and the tools that are using those LlmProvider generic interfaces.
The various vendors behave differently enough to make a simple code contract swap largely infeasible for many kinds of domain-specific agents. I agree there are lots cases where this does work (probably most of them), but you sacrifice performance in the targeted scenarios where you need much more specific alignment and control.
The biggest example where this breaks down is with computer use and vision. I cannot take a targeted solution that works on OAI's stack and forklift it over to anthropic (or google) and expect anything to perform correctly. Some things will work some of the time, but it's basically starting all over again with alignment each time you move something this complex.
Having a generic interface for executing shell and writing code in an iterative ecosystem is not a hard problem to solve. Having a generic interface for reliably inspecting specific PDF documents in a very particular way in a single shot is a much more difficult problem to solve.
If you follow this argument to its logical conclusion, you will wind up implementing a provider-specific version of your tools. And the ones that demand this treatment will be the most complicated ones.