logoalt Hacker News

pu_pe • today at 3:01 PM • 6 replies • view on HN

I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one.

Interpretability is the same, our abilities to do that have increased rather than decreased. I think a codebase generated by AI is actually more understandable than one generated by humans at this point, and you can ask clarifying questions whenever you get stuck.

TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.


Replies

eithed • today at 3:34 PM

If you can quantify quality/reliability/understandability, you can tell LLM what kind of code do you expect. If not, you get whatever.

At my current place we not only have automated tests, static analysis and static rector (linting, but also automatic pattern matcher for problematic code) but also: - architecture tests that define relationships between application layers - ADRs that guide developers (and agents as well) that communicate how new code should be written and how existing code should be treated

I find that "how code should look like"/"what code should do" is an ambiguous idea that always is preached, but never defined = everyone's idea of quality is slightly different and only looking at existing code you tend to align. Everyone's idea of what the product does/should do is kept within their heads. If we define this knowledge in writing LLMs can not only write code according to the patterns that are thus defined, review existing code based on these documents, but also actually read acceptance criteria documents to check if the code does what it's intended to do (gherkin)

Same goes for understandability - if LLM applies one pattern this time, another pattern another time, if you have multiple coding patterns then that hurts clarity. Sometimes LLMs work as common denominator thus achieving clarity, but I find that actually giving LLMs reference works.

munksbeer • today at 4:32 PM

> I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop

You're not going to get people to stop doing that by arguing on the internet, but in the end it won't matter, because it will stop, naturally.

In the future, you'll just get left behind and not hired if you're building code by hand, it's that simple. Even traditional code reviews are going to go away. It'll be more about the scope and then verifying correctness.

➕ show 1 reply
palmotea • today at 3:07 PM

> I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop. All these mission critical industries listed in the article rely on extensive testing for quality assurance, with human code review being a layer on top of all that, but far from the most critical one.

At least some places are abolishing formal QA because LLMs. There's a cult of speed uber alles that has a big intersection with LLM enthusiasm.

➕ show 2 replies
suddenlybananas • today at 3:06 PM

>TFA's points only make sense if the mental model the author has in mind is someone who writes a prompt then immediately puts an app into production without any thought behind it.

If coding were solved, then this would be true no?

➕ show 1 reply
westurner • today at 3:35 PM

> I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop

I agree. What does coverage-guided fuzzing fuzz if there is 100% test coverage?

So, then, 100% branch test coverage is not a sufficient metric (because it doesn't indicate whether the code is fuzzed or formally verified for example).

Would Branch coverage even be a sufficient software quality metric if we were to instead measure how many times each branch of code is covered by tests? How to verify that one test which executes 100% of the code and runs only one assertion on, say, a CLI utility exit code integer is actually sufficiently covering?

> I think a codebase generated by AI is actually more understandable than one generated by humans at this point,

From doing a larger port (of sphinx, docutils, myst-md-parser, pygments, to rust in westurner/dsport) with a lot of human in the loop and currently ~80% branch coverage, this seems to be at least initially true but just like real life there's drift from even a good plan that you pay a more expensive model to prepare.

I suppose it's the same challenge as architectural drift in open source non-LLM-assisted products and the solutions are pretty much the same: give better instructions (AGENTS.md,) and use better sufficiency criteria as an engineering manager (branch test coverage, fuzzing, formal methods, TLA+), and train and pay humans to do secure code review.

Sometimes the agent doesn't notice that the code already solves for that and implements its own implementation with tests and it's wastefully redundant when the code should be refactored and the tests should be refactored so that we can delete code in order to minimize bloat.

Unfortunately often, just like IRL software development, the response from the agent is not sufficient to close the issue.

One proposed solution for this that is in retrospect obvious and also essential to success in "normal"/"traditional"/"legacy" (non-AI) engineering projects, is to always verify whether the candidate solution satisfies the criteria;

From "Groundtruth – checks your AI coding agent's claims against the Git diff" https://news.ycombinator.com/item?id=48838209 :

> "Follow up to verify that the work was actually satisfactorily completed"

> Are there other sound management practices that aren't yet effectively implemented in current gen agents?

Oh, and always write tests, docs, commit messages, and changelog entries; but don't waste tokens on documenting something that doesn't verifiably pass sufficient tests.

bdcravens • today at 4:05 PM

> I agree, I think this false dichotomy between using LLMs and caring about quality/reliability needs to stop.

It's not just a false dichotomy, it's intellectual dishonesty. It wasn't that long that conversations about code quality, technical debt, etc were on the front page of HN on the regular. Whether it was coding bootcamp grads who had just enough confidence to be dangerous, "just ship it!" cargo culters, or the product of management breathing down the necks of otherwise good developers, there's plenty of "human slop" running in production across servers worldwide.

➕ show 1 reply