“But that would be like telling a student in the 70s to pretend that calculators or computers don’t exist.”
If students of the 70s or today pretended they didn’t exist up to a certain point when they needed them to move forward, like bioinformatics or something, they 100% would be better off. There is plenty of research on off loading thinning providing a worse understanding of the material - eg side rules proving a better understanding than calculators.
I had the same thought and came to the exact same name: vibe crafting. Published a skill in the process, steps: https://github.com/scosman/vibe-crafting
In general, feeding the AI what you write for verification sounds like a good idea. But I think even this should only be done by someone who understands exactly what they're doing, how, and why, because otherwise they won't be able to analyze the accuracy of the AI's responses...
Insightful:
> the term “vibe-coding” suggests a kind of laissez-faire attitude where you don’t really care about the outcome and you’re just having fun. That’s what the phrase meant when it was coined, but the world has moved on. In many companies professional programmers are using AI in such a way that it’s impossible to imagine that they are also reading the resulting code in detail. This is what modern vibe-coding is. Deferring to the AI, not worrying about the individual lines of code, and keeping an eye on whether the code passes its tests and throws up any problems in production.
This way of working is the only one that justifies the trillion-dollar bet on the AI industry. I agree, this method should be called vibe-coding.
This article hits very close to what I have come up with myself and its good to see others thinking the same lines.
Core modules: Coded by myself, AI reviews and AI to discover/learn.
Stuff I don't care about Craft: API layer, CLI layer, Smoke tests, Integ tests - Dial AI heavy, and lighter human reviews accordingly
Obviously takes a lot of patience and very easy to sin, but on good days, its doable.I went from just coding manually, to using LLMs for "surgical edits", to "woah, AGI!", to "haha whoops, not even close", to just coding manually (my brain still works!), to "surgical edits" again.
My current approach is "ask for very small diffs" + "review them very carefully".
I'm not working a job though, I'm working on a multiplayer game.
Main findings: The frontier models can't reliably modify Pong without breaking it, so their skill appears to be quite domain-specific. (OK, to be fair, neither can I half the time!) This is probably because they are "time blind". I had one model try to test a game by running it at 0.1 frames per second and shoving each frame in the vision API...
If you leave any room for a misunderstanding, they will laser in on do it and do the stupidest thing possible. If you're not checking everything carefully, you will discover this later, and you will cry.
Formal proofs, oddly enough, do not improve the situation: they will simply prove mathematically that the absurd and pointless and backwards implementation is completely without defects. (It obviously does help within an implementation, though.)
They can't formally prove what the hell you meant when you told them to build something. That job remains frustratingly human!
Current dissatisfaction: (1) Harnesses are designed for super bloated codebases (i.e. designed to load as little context as possible) which make them pretty clunky for small repos and small edits. (I had a Surgical Edit Tool I need to bring back...), (2) Current LLMs are anal about verifying the most trivial change, even without prompting, even if it's impossible for them to verify it because they're blind so they start measuring pixel data in Python... Both of which eat up Speed and Cost, taking the work even further from Realtime/Interactive to Tedious/Sad.
I think one of the best use cases of AI is as a natural language interface to programming. The syntax of a programming language is an opinionated part of its design that often adds to the complexity of learning the language.
>It's a verbose coder, it overcomplicates things, and it works fast. Even if you had an inhuman level of attention, you couldn't keep up.
I wonder if that trick about prompting it to write like a 5th grader* would help here. Keep it simple!
*A trick which allows it to pass for human 70% of the time...
Following a few basic principles - I don't use paid-AI - only free services - so I regularly use ChatGPT to check the code generated by Claude and any others I might have limited access to.
I've recently settled on Deepseek4 for one of my projects, and had it review other code generated by other models, and .. yeah, that was quite eye-opening. Someone in the frontier-models part of the world is definitely paying attention to the AI slop generated by the other models, because having one AI checking the results of another AI has been quite fruitful, lately.
This approach feels just "add friction to your AI usage". It seems the worst of both worlds, both hand-coded and vibe-coded. You paste your code into a chat so it can tell you what to type yourself (the codebase access rule looks optional, but the copy-paste is one way by design). Replace "chat" with "Stack Overflow" and it'll sound familiar. You don't need to paste code into a chatbox if you're going to do an AI review later anyway.
The security argument is the strongest part of the post, and I don't disagree with it, but what it buys you is the review, and a review catches what you missed, whoever typed the characters. None of the ten dogmas follow from that.
My general approach is to design beforehand, do an adversarial review with AI, socialize it with humans (if needed), generate a plan, and start working item by item. Always keeping me, the human, in the loop (not that `/loop`), going through the steps generating code. Finally, a manual review and one AI adversarial review of the feature branch in a clean context, going section by section manually and discussing anything relevant, and off you go.
Writing code by hand feels great, but even local models can generate fine code. Deterministic linters and quality checks are what keep the quality in line. You can always modify things as long as you're in the process, but at the end of the day, you'll review more than you write. We're closer to being the assembly line's inspector than the crafters we once thought we were.
Very interesting. I find this is specially applicable to people learning or perfecting coding skills, not so much for developers already proficient in coding. What do you think?
[flagged]
[dead]
This misses the best use of AI in my opinion, which is to gain understanding. Whst is this bit of code doing? Is there a risk of data leaking here? Are permissions enforced downstream of this function?
Code review can go much deeper now if you use AI to aggressively attack a PR combined with your human insight. Same for planning a feature:
Can I consolidate this logic to a shared function? Does the error surface to the user and are there any gaps? What preexisting functionality is affected by this PR? Can this query be made more efficient?
LLMs are great with focused questions, up and down abstraction layers and across all kinds of concerns. Stack up these focused concerns into a rich understanding of what you are doing or writing.
Understanding is the real output, code is the byproduct.