Rafi and I, who authored this post, will be hanging out here for any questions people may have.
Could you fix a bug in this codebase yourself, without the help of agents?
Maybe trying this would be a good experiment? Say route 10 or 20% of any future issues to the “Chief” to troubleshoot and fix directly, and record those for further work?
This, with an extremist take on code quality via linting in ci, is no doubt the future, at least for maintenance and extending the interface kind of work.
1. How does this work with greenfield lifts where the scope and final vision are not yet figured out?
2. Sorry if I missed this in the post, but will you open source this system?
How much of this is a snake eating its own tail against vs the control group?
are any of these artifacts public and available for inspection/use?
The blog post links to an issue in the Sinai repo, but it’s private or just doesn’t exist?
Which inherent limitations did you recognize in the metaphor before commencing this research?
[flagged]
Did you do any ablation studies on what is actually useful vs what happens to just work because these systems can work around whatever people do?
I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.