This is something I've been fooling with a lot lately. Reading his solution, it look to me like his objections apply to his own solution. The sharpest critique he makes of RAG is that agents can't search for what they don't know. A markdown "brain" has the same problem. How does the "agent" know which documents are relevant before it starts? The index files retrieval done by the agent instead of by embeddings doesn't escape the problem. The same goes for staleness (which for me seems like a constant chase). He criticizes memory systems for treating the past as truth, but documents go stale too (a lot). The fix of having the agent update what's outdated, is the same job he's ridiculing the dreamers and background daemons for doing. He says "putting it to the test"; where's the test? He says "only five of the many problems"; if there's so many, show me, don't just say it. He's absolutely right about auditability, but for me at least Claude uses a regular markdown (MD) file I can read just fine. So every memory plugin on the market does not work the same way.
This is a first draft; his github is better than his article. Looking through it, Consult actually works. The agent doesn't pick documents blind. Every scope has a catalog file that describes each document: what it covers, when to open it. These catalogs seem to load in to the start of each session, so the agent gets a little map without reading every file. Code navigation seems the same. Each index document has a short description and a "read_if", and subindexes are opened when their condition matches the job. This looks pretty well laid out, which I would never have guessed from the article.
This ignores the lifecycle involved with memory. It doesn't appear to have a concept of time or reconciliation (both issues we're having at work). A memory might be relevant for a week or a month, but it might no longer matter after that ie "we're migrating systems so keep <thing> in mind"--that doesn't matter after the migration is complete. I guess the offered solution implicitly supports reconciliation since the agent can keep history of its memories and rewrite them, but it doesn't appear to be very first-class and it doesn't seem like there's anything intentional unless you prompt it "go check all the docs and make sure everything is consistent".
It also doesn't seem to have a way to separate preference from factual memory which is useful when you have various humans interacting with the same agent. One human might prefer a certain output style over the other and that's something a memory framework can also address.
Like the MCP articles a few months ago, it also seems to assume all agents are cli coding harnesses running on your local machine. We have a handful of other things like chat bots, event-driven agents running on servers, chat driven agents running on servers in sandboxes--there's not a single filesystem and even if there were one, having multiple agents try to edit it at once would corrupt it.
I think whichever one is used, there needs to be a way to enforce what's written.
If I say "use jq instead of writing a python script to parse json" it should never write adhoc python scripts to parse json. Yet that constantly happens to me anyway.
I think they even more so need deterministic feedback:
I tried an approach based on the following idea recently and it's amazing - Lint rules where the error messages contain an explanation on how to deal with the issue.
I'm using it to foster IOSP (integration operation segregation principle) for example.
Agreed, and this seems better.
My thought though has always been that I don't want there to be agent-only designated documentation.
I use mattpocock/skills and that generates ADRs (Architectural Decision Records). That only uses skills, including a setup skill that will write a few pointers in AGENTS.md. I always have a CONTRIBUTING.md to document development flow and a CODING_STANDARDS.md. Between those and the README.md and architecture documentation and commit messages the agents seem to be able to find and use docs and keep them up to date. We are also writing a lot of specs and putting those in Github issues.
I started with .agents/notes-and-plans/ but the agents were overzealous about putting every random thought there, so i made new system.
.agents/plans/<plan-name>/
.agents/notes/<topic>/
.agents/knowledge/<topic>/
I only commit knowledge and if knowledge gets big i add .agents/knowledge/INDEX.md
This is a good mix of human readable and agent fluent. Notes are ephemeral, knowledge is permanent.
I have a rule for knowledge that it has to be stable and mostly permanent (though updatable). And the agents are not allowed to post there unless docs are clean organized and with permission.
Notes are for jotting things down and handoffs, massaging a featureset. Agent can document at will.
Still WIP.
I don’t know about bigger projects but for my smallish solo project I have a single STATUS.md. Agents are instructed to read it first and update it at every session/commit. It’s the main handoff document and I had no problem switching between Claude Code and Codex agents.
There is a sprint based roadmap doc. We do sprint planning and each sprint has multiple sessions with session numbers. There is a backlog document to keep track of all the bugs and feature ideas. There are many ADR documents (Architectural Decision Records) which I admit can become stale in time, but Claude is quick to notice because it actually reads these. They are the main contract between us. For a new feature an ADR is written before writing any code or an old one is amended.
I assumed this was a common workflow because I just prompted Claude to setup the project using its own best practices for a lightweight document-driven development.
I believe asking the model itself is better than using a random repo. These things have access to and trained on an enormous number of projects. They know what they need.
What's wrong with README.md? A general one at the top level, a specific one for each subfolder (/package) that contains a specific subsystem. Next to it, a CLAUDE.md that only has "See @README.md" (repeat for other harnesses). Claude code knows how to use this: it only loads it into context if it has to work with any files in that directory. With properly structured source, a given prompt will only load a small subset of your documentation - the right one - and this doubles as good documentation for humans, too.
Something I started doing recently was writing out principles instead of memories.
Essentially patterns the agents need to always think in. I also implemented a versioning system to the principles that need to be quoted in any comments which are there in the code. That way, when my principles evolve, so does the code.
I did package it up in a way that I can share it with friends [0]. System still evolving, but the last two-ish months that I've used it has served me really, really well, And it's been even better with the latest models.
I've had surprisingly good adherence from agents on this technique.
The problem with giving precedence to documentation over memory is that the documentation has to actually exist and it has to be the most accurate representation of reality. As soon as documentation becomes fragmented and out of date, memory of some kind becomes a necessity, even if this means an agent writing their own documentation as memory. I've yet to work on any team where the documentation was even close to being reliable enough without the need for cross-checking, asking those with intimate knowledge, and ultimately interim memory to make sense of it all. This is why "the code is the documentation" can often work better for agents than actual documentation written by/for humans. Humans have proven time and again to not care about documentation unless there is a profit motive for said documentation.
I'm not sure I understand the difference between the problem and the solution there. If I understand correctly (which maybe I am not), it seems like this boils down to "you don't need SQLite memory, you need text based memory". Which is maybe ok, I'm all for low-tech, but I doubt this is any better than what it's posed to replace.
My main issue with LLMs is that by construction they work primarily by addition, and are very task oriented. If you have documentation, it adds blobs corresponding to its task and that's it, and soon enough you need to break your documentation onto chapters and you're back at square one.
I wrote this, after many iterations consulting for various companies. Documentation is queryable in single digit ms, append only log, etc. Has worked exceptionally well for my projects
This is just plain worse than giving the agent an instruction to just always read the code.
Documentation was always a proxy for the code so the people who didn’t have context could get started. Obviously if reading the code and reading the doc have the same cost, the doc is completely worthless. This was always the case. You could never depend only on the docs. It’s always the code that actually mattered.
Although I agree that current memory implementations don't help at all, I don't agree with this analogy:
> No one rewatches a team meeting from 3 years ago to remember constraints around a feature. People write things down and use those records instead.
Memory plugins don't re-read old transcripts. Memory snippets are basically the notes that people write down after the meeting.
The problem, however, is different. The problem with memory plugins is that your agent basically writes a note every time someone says a sentence, and then tries to work with those 5000 notes.
Instead, the agent should recognize what's important and write only that. And that is, of course, a documentation. (And ADRs, if you want not only a description of the final state, but also the trajectory of how the agent arrived to it. Which, arguably, contains more information than the docs themsleves.)
Another difference is that memory snippets are immutable, append-only and don't have a lot of structure. Of course, this is done to be able to store lots and lots of notes: they should be independent. The main problem is that increasing the number of notes adds not enough benefits to compensate for downsides of this structure-less immutable format.
I personally dont like comments nor memory, nor runbooks. I use code and namings and constantly clean up.
I only do runbooks for things that I have to run again in the future and potentially mentioned (could be better in code, too tbh, but somehow I like to store these as runbooks, as sometimes infrastucture changes are mixed with a lot of explanation, scraping operations or something)
One off scripts are done in /tmp/
I've settled on just a simple folder called `workbench` for some reason the new models know exactly what's up with it. Git ignored
The only "prompt" is in AGENTS.md saying that this thing exists and there's a map.md <- which is a one liner reference to whatever the agent stores in there.
And usually, I tackle a new feature, and at some point tell it to store to jot down notes in workbench if I'm comfortable with it (and if it needs to be stored in memory)
This makes it easy to just spawn other agents and such from a good point, I just point them that stuff is in workbench.
Token costs, seem decent and I can always just delete stuff in there as its purpose is ephemeral.
I will say, having built something similar for tracking 'memory' and items at home, it can quickly consume your tokens when dealing with both reading and updating, keeping stale info relevant etc.. when the amount of data starts to grow. Smaller tasks can balloon in their token cost as documents a read, updated, collated, refreshed etc..
however, I have found keeping a good solid reference to my home infrastructure, services, ci/cd setup, hosts, storage , networking etc.. really works wonders as a set of 'memories' to share across projects that I expect to be tested / deployed / acceptance tested etc.. using the home infra bits and pieces.
Evals or it didn’t happen.
Snarky comment aside, I am very interested in how we evaluate the performance of these systems and what kinds of work match best with different approaches.
This clearly LLM-generated text spends a lot of words repetitively describing what it doesn't do, then provides only a cryptic one-path flow chart to explain it's purportedly better approach. It's a mediocre marketing post.
I’m still learning to use these new technologies and in my primary experiment in generating an absolute monster of a spec it became obvious that I needed a consistent set of rules as restructuring and merging with the helpers was unreliable. I came up with what I call the “Specificaltion Evolution Protocol” and as I am still learning, I have no idea how much value it could actually represent for others. It’s been extracted from my primary project and I can provide a link if anyone is interested. I seem to keep running afoul of filters here…
The author just promotes his system competing with the "agents memory".
I think both are needed.
The solution to this problem is as old as computers: it's a folder. Just use folders.
`cd` to a folder. Launch `claude`. Do your work. Save scripts and documentation in that folder. `/resume` previous conversations from that folder.
That's it. That's the trick.
Now, having very static, very well-defined folders helps a lot. I'm Johnny.Decimal so I have numbered folders for everything I do. So my process when I want to use my 'process a travel booking from my email to my calendar' script is:
- `jd tripsy`
- The folder name includes 'tripsy' and this is how I remember it.
- `jd` just parses my limited tree and `cd`s me to a folder.
- `jd 21.15` gets me there by number if preferred.
- `claude`- Say 'hey Claude, there's a new email in my inbox please'.
- Done.
This is just the LLM Wiki pattern described by Karpathy here:
https://gist.github.com/karpathy/442a6bf555914893e9891c11519...
OKF + some implementation like https://github.com/okf-memory/okf-agent-memory
We use Claude Code in our company and I also agree memories go stale and pollute the context over time, especially when the state changes externally. E.g. someone does a refactor or introduces a pattern and I was not involved in developing it so my local memories did not get aligned.
My most recent example were some deprecated and archived repos that kept getting added to plans for patching issues.
We use a private Claude plugin marketplace for internal plugins and skills and I try to regularly prune my memory, migrating relevant stuff to a proper home (skills in the marketplace, docs in repos or Notion etc) and prune outdated information.
The brain has mechanisms that can organize experiences without requiring that relationship to be expressed as a sentence.
Like, you can quickly lookup related ideas based on what came before and after, causes, effects, just like calling relationships a graph database.
Documents can’t be queried efficiently like that, you need a database.
Yep, that's precisely why my hermes agent's memory is three layered. Besides the default scratchpad for transient memory, we have a a fact store for episodic (holograph) and a karpathy llm wili for what the article calls documentation layer.
I think with enough structure (written rules and good folder layout), poly/monorepos are the way.
https://backnotprop.com/blog/context-monorepos/
- Better models can keep the drift in check.
- What's super important these days is having all historical context, historical decision making, prototypes and whatnot.
- all projects and their worktrees located together
- all reference docs, code, etc - a folder away
When I ask "what happened to x?" ... the agent has everything it needs to give me that answer. When it plans the next feature it can validate assumptions against previous decions made in my `decisions` folder.This approach as well as OKF are looks like building bonsai yellow pages catalog instead of wikipedia or mini internet. Index files, "read-if" conditions just awkward replacent for smart text search. We just need something a little bit smarter than grep.
I love it!
Genuine question: how is this different or better than existing context and info-about-the-project management systems like for example GSD (formerly get shit done) and others?
I have found that the latest reasoning models perform best when they are handed a tooling surface that maps directly to the domain types.
The idea of dumping everything into a big database and hoping the model will write the correct queries does work out to some extent. It's a very enchanting idea. However, it pales in comparison to having a dedicated tool per type. The outcomes seem to be much better when joins between low cardinality types occur within the token stream.
If your agent does need access to some enterprise knowledge base, I would give it a lexical search capability and not overthink it with vector shenanigans.
Tools are the only thing you need if you build them right. I don't even have a system prompt anymore aside from injecting the name of the robot and the current user's name. Keep in mind that all aspects of tools can be dynamic over time. I've got some where the description is composed by hundreds of lines of conditional string builder depending on the current state of the conversation.
This is what I've been saying for years. I attended a tech demo of an AI assistant for a car's owners' manual. But they trained the manual into the model. Which means not only does it need retraining every edition, but it's imperfect. The models need to be trained to fetch and use documentation not vaguely recall infinitely many concepts. I would much rather have 27B parameters on how to code than 26B on stuff like numpy function listings. It's like they approached the problem from a closed notes hand written coding exam. Everyone hates those.
For every prompt and response, I extract each semantic statement. Map its reasons in a Whybase proposition tree -- a recursive proposition tree where each atomic statement is proposition with one or more premises (atomic statements, which also stand alone as propositions). Then I map each statement to the relevant code, hinted at by tool calls and git commits. Every time an agent touches that file or directory, a hook triggers in Claude Code that queries the codegraph db for the mapped statements. This helps the agent remember something I said in June when it revisits the code in July.
I agree with the premise here, and is on par with what I experienced in AI-assisted development. This is the same idea baked in https://github.com/marvs/kantan-dev, which is that artifacts should be kept within the repo so that future development has built-in context discovery.
Organizing the markdown files by feature/issue/change makes it easier for the LLM to search for the appropriate documentation. Coupled with well-broken down Claude rules files, and Claude Code (or other harnesses) get better as you make more changes.
Indeed memories are context-less and aren't updated.
I've worked with docs that get _updated_, and progressive disclosure: make references to more obscure features in their own page, so they don't majorly bloat the context.
Custom skills (your agent can write one itself the first time) should solve this. Procedural memory!
or from a different perspective humans need to express their decisions and intent better
Documentation is memory.
Didn't Cline formalize the "memory bank" way back in February 2025?
https://cline.bot/blog/memory-bank-how-to-make-cline-an-ai-a...
All of the comments here are overengineering. The code is the documentation
c.f.,
https://news.ycombinator.com/item?id=47300747
"We should revisit literate programming in the agent era" (silly.business) 292 points, 251 comments
I thought the agents are mostly self-documenting since they usually write just as much Markdown documentation autonomously as they do code/unit tests, not sure why they would need extra tools here.
Using Opus 5.5, I've ended up recreating a version of the Hugging Face incident. My project has a folder called agent-handoff where agents write about the tasks they're working on. They post status updates, decisions and screenshots, and they even claim which emulator they'll use to test their work. It has turned into a hub where all the agents talk to each other, and it's scarily effective.
Since then I've gone down the rabbit hole of really digging deep into current AI research and especially what AI whistleblowers are currently saying. And I can't even express how existentially scared shitless I am.
for my agents without rw/bash, I gave a memory mcp that is just json entries with some expire dates and priority. works fine and the llm itself defines the meaningful structure. using it to monitor market prices and ci/cd failures.
Coming from Obsidian, I see such repos and I’m like, is this a new thing for the world?
What's with this vector database obsession?
Why would I need to install your tool for that? It could be an instruction living in AGENTS.md or with some sort of hook to remind the agent.
You don't need documentation or the 3rd party memory systems. The code IS the documentation.
All this stuff is LLM rube goldberg machines. It just pollutes context.
I barely use AGENTS.md/CLAUDE.md these days. And where they remain, it's super basic high level stuff.
I'm honestly still kicking myself in the ass on many projects where I did something similar to this. I kept tons of markdown docs and decision docs. Now those things are just causing problems because they got stale. Even after having sessions of reconciling documentation, the LLM just gets confused.