I'd start with something research focused like Undermind or Elicit. Although I don't think that the author is comfortable with using a tool that isn't produced by the model lab.
The planning model for tool use sounds something like CaMeL, which someone should really try implementing in a product.
CaMeL, from "Defeating Prompt Injections by Design", https://arxiv.org/abs/2503.18813
See also "CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents", https://arxiv.org/abs/2601.09923v1