logoalt Hacker News

Why Are Coding Agents So Dumb?

53 points • by mtlynch • today at 2:21 PM • 25 comments • view on HN

Comments

atleastoptimal • today at 10:11 PM

agents have progressed a LOT in the last year. Claiming they're still bad in the same way they were bad in 2025 is inaccurate.

athrowaway3z • today at 10:05 PM

> Is there a better coding agent for me? I’ve only tried Claude, Codex, OpenCode, Cline, and Pi.

I started to gloss over hard after a few paragraphs because I don't really have any the problems you describe anymor; at least not to the point i'd be blogging about them. Instead, I've just been iterating with an agent on various pi extensions that solve the issues.

As with operating systems - you can hold out in the hope somebody solves the right mix of issues in general, and they match your situation well-enough.

That's your choice.

I would note though that being the passive consumer gets you either Mac or Windows UX & prices.

➕ show 1 reply
giancarlostoro • today at 10:06 PM

The reason Claude looks up the docs of itself is because the model doesnt know about harness features that havent landed yet. The harness itself can be ahead of the model. Happens all the time.

polyterative • today at 9:12 PM

I see your point, but even the fact that I can speak to my computer and anything useful happens is still a miracle to me.I don't think I will ever get accustomed to how good the new models are.I just can't keep up.And I do this for a living.Ten hours a day.

A lot can be improved, but this is already so much speed.

mstank • today at 8:39 PM

I used to relate to this article quite a bit. In the last 3-4 months, not so much. I've found that the latest models -- Opus 5.5, Astra, etc. juggle multiple tasks, delegate exceptionally well and are very good at working independently.

I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.

ilamont • today at 8:46 PM

the agent never stops and says, “Wait, this is something another model could do cheaper and faster.” It just plows on with the slow, expensive model. Conversely, the agent never says, “This model is too dumb for this task. Let me tag in a smarter one.”

This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.

Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?

➕ show 2 replies
Rapzid • today at 9:35 PM

I have some significant experience in context engineering, but I'm most familiar with Codex as a coding harness right now. Sam Altman said "you don't need to write prompts anymore". This has widely been panned as something someone selling AI would say. If you care about the output and how much time/money it costs to produce it.. Just giving Codex an abstract tasks with zero extra guidance isn't going to produce the best results..

Codex Astra can do a great-(ish) job as a project coordinator dispatching tasks to a pool of 6.1 Sol sub agents. You can even give it an explicit goal and ownership over ensuring the work is carried out efficiently.

However the OOTB harness(and prompt) configuration may not do this for you. You'll have to provide guidance over how you want it to operate through your prompt, a skill, or etc.

And I'll say even though it's really good at this.. Having even more layers than 2 can help; a single agent given too many responsibilities will start to become fixated on a number of them while neglected others. You can check in occasionally to "nudge" it or you might need to split out responsibilities more..

I will say it's crazy Codex doesn't have more built-in task and sub agent management features. I almost wish that it had some stock orchestration patterns that worked OOTB, and then you could opt-in to a leaner setup where you provide more of the instruction.

Neywiny • today at 9:38 PM

I keep running into this. It's nice seeing others here struggle. I guess when all your doing is one-shot simple trivial tasks who cares. I've found they're great for that. But once I need to do real work, everybody makes their own esoteric abandonware that kinda works but kinda doesn't. Stars aren't a perfect indicator, but I haven't seen anything over a few hundred for these things I'm finding on GitHub. Same with downloads of plugins. It was very isolating feeling like I'm the only one not enamored by the state of this.

arjie • today at 9:28 PM

It always seems bizarre to me that people complain about software now. You can just write the thing you want. If you want to do things in parallel do them in parallel. Claude Code will allow you to run multiple instances in the same folder and let them communicate. Previously I used to let them intermediate through a communication bus but now they seem to be able to talk to each other.

I let most agents work asynchronously and don't pay attention so I don't care that much about the sequential nature. But if it's a problem for you then fix your harness. This is a bit like saying "Why are shoes so shit? There's a stone in one and it just gets stuck there and your foot steps on it and it hurts". Take off the shoe, and shake out the rock. Put the shoe back on. You have the power.

➕ show 3 replies
361994752 • today at 9:37 PM

I double checked the publish date before writing down this comment. Because I use the same harness (Opencode, to be specific) as the author, and some of the features are just right there. Like opencode can start multiple subagents for different tasks in parallel. Also I usually ask the main model (e.g. opus 5.5) to pick subagent models for me, and it has no problem identify the difficulty of work and delegate large portion of them to gpt-luna.

kgeist • today at 9:17 PM

AI models can multitask/use parallel subagents just fine; the issue is with harnesses that don't make it a priority via the default system prompt.

I run an LLM server with Qwen 3.6 in the office, and OpenCode, which the OP mentioned, usually defaults to sequential TODO lists, and it works fine with our little LLM server with 3-4 parallel users. But I noticed that once in a while the LLM got overloaded with requests in the queue, and you couldn't do anything for 20-30 minutes. My investigation led me to an employee who used QwenCode. I tried it myself then, and indeed, it immediately launched something like 6 parallel subagents, where OpenCode would have sequential TODOs with the same model by default.

So in the end, I had to detect QwenCode on the server side and serialize all its parallel requests into a single request queue, because it made life miserable for other OpenCode users :)

brandtcormorant • today at 9:08 PM

They are as dumb as their instructions.

Have you tried telling models about your dream agent environment?

They can build it.

cbrake • today at 5:37 PM

Enjoyed this article, lots I can relate to.

One thing that seems to help for me is to do the docs before plans (collaboratively edit with agent). Then I understand what this change is going to look like from the user's perspective before we start implementation. This seems to help keep things on track.

While I don't use this plugin a lot anymore, I think doc-driven development is one of the most effective ways to do development in any paradigm, I should probably refresh this plugin and use it more:

https://github.com/tmpdir-org/tmpdir-claude-code-marketplace...

peter_d_sherman • today at 10:02 PM

>https://mtlynch.io/why-are-coding-agents-so-dumb/#what-i-wis...

This is a great list for future Agent / Harness software engineers to read!

Oh sure, there may be some Agents/Harnesses that already accomplish some of these things -- but there doesn't seem to be one (as of the present day that I write this) that accomplish all of them...

As someone that watches the Agent/AI Harness (and related software) space, I will definitely be referring back to, and re-reading this list in the future!

An excellent post!

wrs • today at 9:06 PM

Actually, the last time I asked Claude Code about itself, it located and read its own minified source and told me something that wasn’t even in the docs.

➕ show 1 reply
scotty79 • today at 9:48 PM

> My dream agent

I have no idea what stops that person from just making it, with an agent of course.

imimayj1337 • today at 7:16 PM

It's true! It feels like we've been talking about harness optimisation and 'cool features' available in the cli tools for months at this point, but ostensibly there has not really been any significant upgrades to these harnesses since at least Claude Code imo. It does feel like a contrived way to harvest more and more information and test each conversation/action tool, to the detriment of those of us actually using them!

Kuyawa • today at 9:22 PM

Perhaps is not the agent that is dumb?

I asked DeepSeek to translate a page to five languages and it opened five subagents each one working independently on the translation, once they all finished the main agent informed me of the job completion with a bell. Fantastic!

Sooo, which agent?

chrisjj • today at 9:21 PM

> If I ask Claude how to use the features of Claude, it has to search online to figure out what this “Claude” thing is.

As expected. A model's knowledge is what was it ingested a creation.t

Unfortunately what we get is worse - for the same reason. Model version thinks it is its previous version.

decodingsi • today at 8:51 PM

[flagged]