logoalt Hacker News

Veedrac • yesterday at 9:42 PM • 0 replies • view on HN

I thirty (random long word, letter pairs) on a free model and it failed. I tested a SOTA model and it passed flawlessly. In both cases I denied tool use and spelling words out. So it seems obviously false that AI can't get better.