logoalt Hacker News

VCFundedGenYer • last Sunday at 10:37 PM • 7 replies • view on HN

LLMs still can't do math nor count letters in words. Nothing has changed there.


Replies

walrus01 • last Sunday at 11:20 PM

This is true but a sufficiently smart LLM (run in a harness like opencode, no special MCP, no customization done whatsoever) will quickly turn out a basic 1 to 2 page sized python script to do the math. They can't do the math with any guarantee of accuracy with their own internal reasoning since it's a language model.

But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.

Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: https://www.google.com/search?&q=karney+formula+geodetic+

reference: https://github.com/pbrod/karney

You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.

Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.

➕ show 1 reply
Kim_Bruning • yesterday at 6:35 AM

I could have sworn this had been fixed a while ago..

Just to check if I was actually crazy, I actually went and put a simple addition (7 digits + 7 digits) , and a simple letter counting question to Claude haiku(4.5) , sonnet(5), opus(5.5) and fable(5.1) . They all did just fine straight up.

If you don't mind spending the tokens, some older/other models can also arrive at the correct answer if you ask them to do the math in long form, since that fits nicely inside autoregression.

Not sure since when exactly, but letter-counting hasn't been a problem for a while now either. This used to be a problem due to the tokenizers used. Slightly older models can be asked to split the word out into letters, and then they can use autoregression to solve.

Edit: IMO google search uses a really dumb version of gemini, so I didn't expect it to straight up solve the problem; but it did it just as easily as the claude models. (tested 2026-09-28/eu)

➕ show 1 reply
jasonfarnon • yesterday at 12:20 AM

I wish I could not do math like LLMs

TeMPOraL • yesterday at 11:33 AM

> LLMs still can't do math nor count letters in words.

Humans still can't flap their hands and swim or fly.

Veedrac • yesterday at 5:11 AM

It's pretty wild that AI has solved a Millennium Prize problem and can accurately multiply two 40 digit numbers without tools and we still get stochastic parroting of claims like this.

➕ show 3 replies
fasterik • last Sunday at 10:45 PM

"LLMs can't do math" is a pretty hot take in September 2026.

➕ show 4 replies