> it’s consumed just over $135,000 in Claude tokens
Ironically also just like a US hospital. Go in and come out with a six-figure bill.
I can think of nothing I'd like to use less than a database or filesystem vibecoded by Gas Town-flavored psychosis. Roleplaying with LLMs is not the secret to producing amazing code.
A hint at why GitHub actions have been unreliable. How many teams are running factories like this on GH infra now?
I really like this writeup. I find the metaphor satisfying, though there's a reason metaphor receded into the bushes as a practice in XP. A great metaphor is great to have, but they don't always exist. Fortunately with organizational systems, things will tend to rhyme. For example, a role that "Verifies that all processes have been followed before anything merges" seems like it could fit many different systems. We're at this fun stage where people are just trying stuff and I love it. I hope we get more info from this experiment.
The next role: Insurance Rep.
Ensures all other agents are operating efficiently and within reasonable levels of token usage given the expected level of effort to 'resolve' the patient. In moderate to severe cases, may lead to patient defenestration or agent revolt.
Or perhaps, the author considered this role but found it typically costs more than it saves.
I am curious why so many of these systems are based on GitHub issues. Why not use a proper ticket system?
I use Redmine (lightly customized) and have found that to work very well. Especially for performing review / UAT of the work or providing response to the agents questions when it wants me to pick an option.
The underlying data model of a ticket / project management system is so rich and well suited to bodies of work.
Rafi and I, who authored this post, will be hanging out here for any questions people may have.
I set different kinds of structures for different projects, Complete with setting and ambiance. It’s like agents work better if they are role-playing. It’s extremely disorienting and people with marginal mental stability are going to really have a bad time. What have we wrought?
Since we are all building our version of this, what mistakes did you made until you arrived to a well-balanced solution? I really like that you know how much an issue is, because then you can start optimizing.
Great post. Thank you for sharing it on HN.
Modern best practices for a medical team (standard roles, responsibilities, consequences, processes, etc.) have evolved through trial and error, since the advent of civilization, to prevent costly human error.
In hindsight, it's not too surprising that the same best practices can be applied by teams of AI agents to minimize costly AI error.
It's still incredible. We sure live in interesting times!
If I wanted to build something like this, at least conceptually with the roles, where’d I start? My first guess would be to give Claude your blog post; but any other pointers to make it work reliably? Do you happen to have the system open source?
Maybe I missed it, but I didn't really see anything about the long term quality or maintainability of the code. All I see is agent agent agent.
You give a very precise measure of redundancy in the skills. Can you give a bit more detail in how you decided you needed to audit them and how you went about it? Was it fully agent driven? Mostly human?
One role I didn't see was patient advocate/representative? That might be another approach to non-convergence - "how is this going?" and escalation.
Great post!
One thing I've been experimenting with in my own agent orchestration system is an agent that hangs around and does post-merge acceptance testing after the work ships. Any plans to add a follow-up phase? Travel nurse?
I’m surprised at the reasoning behind using the most expensive model which likely led to the “doing too much”. I think it would have been nice to have A/B tested different hospital setups in order to maximize not only efficiency but to see where the strength of each model started to emerge in a given role.
Fable feels like overkill for this also.
If the correct analogy for a team of SW agents isn’t a team of SW engineers why is that?
Humans: $160k / 9 months = $600/day
AI Software Factory: $4172 / 2 days = $2086/day
This seems unsustainable, unless you're also generating 3x the revenue.
Reminds me of the “surgical team” development model in the Mythical Man Month.
Nice writeup. Structured handoffs and external plan reviews are good takeaways.
I absolutely love how you are able to pull so much latent behavior from the underlying LLM. I wonder what other analogies can be pulled into agentic coding that come baked into the existing weights.
This is fantastic work. I've been exploring my own work system here: https://github.com/mas-bandwidth/nova-sprint and I'm adopting your ideas. Thanks!
Jesus Christ, we are just clowns aren’t we.
"Just lost another bug."
"It never gets any easier, huh?"
[flagged]
[flagged]
[flagged]
[dead]
[dead]
[flagged]
[flagged]
[flagged]
This feels to me like a lovely example of both high-quality systems design and information theory. From my perspective, one of the more abstractly interesting thoughts surfaced by this write-up is the degree to which you've "encoded" a highly complex set of relationships through use of metaphor.
That type of encoding (starting with the statement: you're at a teaching hospital) is profoundly more efficient than having to deeply and reliably articulate what each of the various roles your agents embody are, let alone their interactions.
Perhaps there is more room for any of us to consider what existing systems or identities are described abundantly in training sets and can be leveraged to ritualistically encode these social/civic/cultural dynamics saying: "act as though you're ___." Not exactly a new point, but one this write-up certainly is pulling me towards.
Thank you for sharing and the care you put into writing this!