It's worth noting that the context here is a system that was translated by an LLM which is the ideal use case for it. 90% of the thinking had already been done by developers of the existing system.
Most of this article went over my head, but I think this is key point:
if your prompt is not specific enough, many decisions are a coin flip.
If you are using AI, you still have to know what you want it to do. If you are a project manager and give your team incomplete requirements, the result may not be unlike this.
You need to understand the system you're working on enough to make well-informed decisions about its current and future states.
In some cases you can get away with not reading the code. And maybe in the future that will be more common. But for now, I personally prefer to read (or skim) the code, and not abdicate to AI.
Yep AI will just invent the missing stuff; that's not purely an AI thing, we do the same as well until we become experienced enough to know what questions to ask in order to narrow down vague requirements.
If I'm tired, yep I'm going to write more sloppy code, I'll put less thought into it and be more tempted to tell the AI 'just do it'. That hasn't changed. I'm not sure where this new breed of "don't read the code" is coming from when it's actual software engineers making the statement; maybe it's just over-excitement and it'll die down (hopefully).
Or it's just over-simplification -- I don't read ALL the code either anymore, but at the same time I'm reading SOME of the code. Tell your AI agent you're doing a code review workshop, ask it to select ten files at random and then go through them with you one at a time. Ask it to review common anti-patterns and put them in your coding standard. Then you can start asking it to review and fix code against your standard as a first pass. You'll still need to manually review the code to keep it in check, but at least you can make it a bit less daunting and a little more fun by actively collaborating on the AI with the review.
I'm learning it's just a 'different kind of tired'. AI has made some stuff much more efficient, including the choice between efficiency and carefulness. You can make that choice at a micro level and very quickly now if you want. There's still fatigue though, it's just a different kind of fatigue. You still get tired from all the thinking. You can still choose to put the work in or not and it still makes a difference. It's just a different kind of difficult. Doesn't mean that the thing to do is to throw your hands in the air and declare that your job shouldn't be difficult or include hard work anymore.
Yes I agree thinking is still needed off course, but it has never been easier to ship 'half baked' things that don't actually matter and just works. What I mean is, if something has little proven business value, why waste time making sure it is correctly implemented? I'd rather spend time thinking about things that has actual proven business value, otherwise I'm not opposed to shipping something that I understand only 60 or 70% and save energy. ie: how it works at a high level.
If we are going to the moon, then by all means I think we should probably scrutinize every single line of code, but most business aren't going to the moon.
All of this "you need to understand the code" discussion reminds me of when I was an undergrad. I worked in a computer with someone a fair bit older -- they had even worked with punchcard systems. He was always happy to oblige us with war stories from the olden times & we would lap them up.
He favorite axe to grind was how my generation put too much trust into the compiler. "Just because it compiles doesn't mean you're done. You need to look at the actual assembly. Compilers can be VERY inefficient."
This was in the 90s. There was some merit to his claim.
But today? How many people who write Javascript look at the actual assembly code that is run? Not many.
I can't help but wonder if this current "you need to understand the code" is the same thing all over again. And if in a few years, almost nobody will look at the generated Python, Ruby, etc.
> If you want a reliable system you should know when to fail
I think the main problem is that most people pushing for heavy AI delegation either
1. Don’t really care about system reliability or 2. Don’t understand the difference between code and system, and assume that a system must be reliable if some automated checks pass
I dread for the future where "you have to think" is considered a controversial statement.
It's not clear how the re-writes were done.
If it were me, I would start by getting a summary of what problems the existing solution solves. I would also make sure that summary had the non-functional requirement. Maybe write a test suite that opens the page and does all the stuff.
Then ask the LLM to implement in another language, with those solutions in mind, but first coming up with reasons why the target language might not be as good, and where it might have useful features that the original language didn't. Use the test suite to check if things are working.
I honestly don't think you would land far away. You would have to answer a few questions along the way, but it would mostly be plain sailing.
I would expect to be able to do this with very little human attention, whereas a language rewrite two years ago might take a whole month, not including learning the new language.
LLMs produce code faster than humans are able to check it. Today's systems are able to generate thousands of lines of code per day, no human team would be able to keep up with them. The problem, in my opinion, is the way LLMs are used. An LLM reasons and makes mistakes according to the logic it was created with, but different LLMs reason and make mistakes in different ways. I'll give an example: take 2 black boxes, each one with 2 inputs and 1 output. Inside, the boxes work in a different way: with the same inputs the outputs are different. The real key is to use the method of adversarial development: this reduces to a minimum (even if it doesn't eliminate completely) the possibility of errors. So to humans the task of checking the output, and to have quality output you need to spend, and a lot, at least 2 frontier LLMs. But the question is: is quality as a goal an expense or an investment?
With all these ports, I always wonder: Who implements new features, and in which code base?
It's easy to port a codebase (at least, if you have automated tests), but once you've done that, do you
- Implement features in the old code and port again?
- Implement them in the new, unfamiliar codebase?
- Just write a Jira ticket and let the agent YOLO it?
Informative thought process here. Pre-committing benchmarks to measure real efficiency gains might have helped, but hands-off complex rewrites are still not shovel ready.
These articles IMO suffer from good data. It seems we could probably at this point study open source projects longitudinally and see what we learn depending on their AI policies. Are there patterns in stability? Bugginess? Performance? Readability? Etc.
Or as someone in 2025 said..."it's better to be competent".
My preference is to not think, then act and then perpetually think about what I should have done instead
I agree what you have to think. But example is kind bad.
We got strictly better version of application for cheap. What is there even to complain about?
Really appreciate the time taken to dig into this with solid improvements to boot, great work
"If you looked at the code, and yeah, I know, we should not be reading code anymore"
I'm pretty sure this is sarcasm, but it's also a major factor in what differentiated slop from non-slop AI work: did a developer actually take the time to review and clean up generated code. Which is, itself, an extremely time consuming process
My suspicion is DHH is intentionally creating controversy to drive interest to Omarchy
If he truly believed in the idea of never reading the code, he'd fire all his developers and have the designers do it all - create very high level feature descriptions and iterate on them. Put your money where your mouth is
The quality of anything is proportional to the time people spend paying attention to quality measures they care about. If you care about the quality of the code (and I am extremely bimodal about when I do and do not), then spending time looking at the code, whether said code is written by a clanker or a coworker, will lead to tinkering with and improving it. If you care about UI functionality, and you spend time tweaking the UI functionality, then it will get better. If you care about performance, and you spend time measuring, displaying, improving, it will get better.
Clanker code is no exception. Slop is slop because it's low effort.
There are projects I've put most of my career into (necessarily pre-AI), and those I definitely care about the code quality. I look at every line and I don't take clanker code without going around the mill a few times with it. There are lots of things in the code I like and things I don't like and want to refactor. I care about the code quality. But there are projects I've put only weeks or month into, and some only minutes. For most of my vibe-coded projects, I've only cursorily glanced at the code. For these projects I don't care about code quality as much as the architecture and interaction. The ones where I consistently tweak it to work how I want are...just better.
This post resonates a lot with me and my approach to counter the idea that an LLM (as amazing as a tool it is) gives programmers permission to stop engineering. To me, executing a decision in the face of trade-offs is what the job is all about.
So, I want to keep pulling the thread: is it worth reading—meaning, understanding—the code in order to prevent p95 latency from skyrocketing, avoid OOM death, and keep our apps reliable for customers?
The cost of understanding the code is ostensibly very high relative to just using a clanker (citation needed), so does paying that cost translate to value—what is it worth? Concretely, if vibe coding increases bugs and enshittification, but decreases spending without decreasing revenue (read: customers suffer, but they don't leave), do we still need to pay the cost of reading code? Is it better to invest in stickiness, lobbying, and market capture?
This comment shouldn't be read as advocating for this; it's a thought experiment about pragmatism and trade-offs, and reflects what I'm witnessing companies (and "programmers") asking themselves. I personally care a lot about understanding systems and code, and I believe I'm paid to do exactly that (I very much enjoy understanding code and nobody is paying me to run a business, so I may be biased in this). In everyday SaaS-land—not talking about critical safety systems—we've seen production databases destroyed, personal data leaked, platforms unwittingly exploited, and UX bugs creep into our operating systems, and yet the companies involved keep on keeping on.
Is thinking, meaning taking the time to actually understand what our systems are doing, going to increase their shareholder value?
I dunno. You still have to test, but in my experience if I noticed unreliability during high load I could just tell Claude "I noticed unreliability during high load" and it would easily have set up a high load benchmark, discovered this broadcast channel issue and fixed it.
It definitely improves the results if you do think but I don't think "you still have to think" is a safe space that's going to save us from unemployment.
I think attempting few-shot agentic ports of an app as an exercise and having the results not be perfect is really a referendum on thinking or using AI in coding in any particularly meaningful way.
smokes cigarette
Life, eh?
I feel a bit sad DHH did the elixir one so dirty. It's the only language with the rule to follow rails to the T and of course Elixir can fly and be incredibly powerful without redis and other bullshit. It feels deliberately hamstrung. Almost maliciously because heavy hitters like Jose Valim have told him about this and he just ignores it.
Not a good look tbh
Oh wow, he's actually delusional:
"I didn't do anything. I asked two frontier agents to make the most of Elixir and this is what they came up with. Please do send a PR to speed things up! But also realize that it's not exactly extra points for Elixir if frontier agents can't find the magic "go fast" switches."
[flagged]
[flagged]
[flagged]
[dead]
This article just reminds me that workslop is so hard to counter because it takes no effort to produce and immense effort to debunk