logoalt Hacker News

paidxtoday at 1:07 AM1 replyview on HN

The interesting tradeoff here is not just better retrieval, but whether the pre-compiled wiki makes behavior more predictable over long-running tasks. It would be useful to see comparisons against a plain markdown repo plus Claude/Codex: task completion, citation accuracy, stale-context failures, and the cost of keeping the wiki updated. Multi-model support and bring-your-own-key options also seem important for enterprises that don't want their company's context tied to a single provider.


Replies

kushagrchitkartoday at 2:40 AM

We have been debating what the best benchmark to show would be. I'd love to test it out on long horizon tasks, but then we're thinking of best to make the base wiki to test it out