logoalt Hacker News

Kim_Bruning • yesterday at 6:35 AM • 1 reply • view on HN

I could have sworn this had been fixed a while ago..

Just to check if I was actually crazy, I actually went and put a simple addition (7 digits + 7 digits) , and a simple letter counting question to Claude haiku(4.5) , sonnet(5), opus(5.5) and fable(5.1) . They all did just fine straight up.

If you don't mind spending the tokens, some older/other models can also arrive at the correct answer if you ask them to do the math in long form, since that fits nicely inside autoregression.

Not sure since when exactly, but letter-counting hasn't been a problem for a while now either. This used to be a problem due to the tokenizers used. Slightly older models can be asked to split the word out into letters, and then they can use autoregression to solve.

Edit: IMO google search uses a really dumb version of gemini, so I didn't expect it to straight up solve the problem; but it did it just as easily as the claude models. (tested 2026-09-28/eu)


Replies

leoedin • yesterday at 8:12 AM

Are the models "doing the calculation" or are they calling a calculator tool? There's a lot of talk about how models can do maths now, but I'm struggling to understand if that just means they just need to recognise that it's a maths problem and pass it to a tool, or if they're truly doing the numerical manipulation themselves.

➕ show 2 replies