logoalt Hacker News

danlitt • yesterday at 11:33 AM • 2 replies • view on HN

Not sure about multiplication, but letter counting is still constantly wrong if you can trick the LLM into not calling other tools. Obviously there are lots of mitigations on the server side to try to avoid this happening, but the underlying LLM has not got any better at this type of problem (and probably can't).


Replies

Veedrac • yesterday at 9:42 PM

I thirty (random long word, letter pairs) on a free model and it failed. I tested a SOTA model and it passed flawlessly. In both cases I denied tool use and spelling words out. So it seems obviously false that AI can't get better.