logoalt Hacker News

beloch • yesterday at 9:42 PM • 14 replies • view on HN

This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words. LLM's don't "think" or "reason" in the normal definition of those terms. They can do some pretty amazing things, but still screw up basic things like telling you something that is obviously wrong and contradicts the top search results.

LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly. This may be why they are so difficult to constrain. You could give them something equivalent to the laws of robotics, but following laws requires thought processes they simply don't have.

I'm actually sort of amazed Google doesn't make people accept some kind of butt-covering EULA and post disclaimers about the inaccuracy of results before even showing you their AI's output. Are they not being sued over this kind of thing?


Replies

baubino • today at 1:39 AM

> LLM's, in their present stage of development, are sort of like a crack-addled idiot savant. Sometimes they are obviously insane, and sometimes they seem quite cogent, but you must never trust them implicitly.

I have nothing to add. Just wanted to save this quote for posterity. Thank you.

➕ show 1 reply
VCFundedGenYer • yesterday at 10:37 PM

LLMs still can't do math nor count letters in words. Nothing has changed there.

➕ show 7 replies
foobarbecue • yesterday at 10:31 PM

ChatGPT live mode still hallucinates letters in words like this. HuskIRL and FatherPhi on youtube have done some hilarious videos with it in the last couple of weeks. Beyond miscounting the Rs in strawberry, ChatGPT will say there are two Ds in "your mom" and one D in "uranus" . I tried it myself to check that the videos weren't fake and sure enough it still has this failure mode.

➕ show 2 replies
jonas21 • yesterday at 10:39 PM

LLMs might not reason exactly like humans, but they do produce much better results if you turn reasoning on.

The "crack-addled idiot savant" phase was really circa 2024, before the big labs figured this out.

I think the issue here is that Google decided that doing reasoning in the AI overviews in Google search would be too slow (and probably also too expensive), so it's still stuck making 2024-era mistakes.

GolfPopper • today at 6:52 AM

At this moment, ChatGPT still tells me raspberry has two p's in it.

➕ show 2 replies
Isamu • yesterday at 10:55 PM

>Are they not being sued over this kind of thing?

Maybe but you have to have deep pockets just to get to the starting line. And then you need standing, and some injury to argue.

Corporations have been remarkably successful at arguing they are operating within the bounds of free speech, whether or not what is said is factual, and whether or not any fact checking has been done.

nextaccountic • yesterday at 10:58 PM

> This is similar to how, not too long ago, LLM's had extreme difficulty counting the number of letters in some words.

The specific issue of Google is that they are using an underpowered model, not fit to task, and much prone to hallucination than either OpenAI or Anthropic free tier offerings.

Google should at least match the frontier labs at the free tier (with some limit; after that, degrade quality), ffs

➕ show 2 replies
somenameforme • today at 2:00 AM

Yip, I've gradually become quite optimistic about the future of LLMs, but the fact that they remain [very] glorified token probability prediction algorithms means that they will probably never be able to achieve meaningful 'intelligence.'

But that doesn't mean they won't be able to do a vast number of extremely intelligent seeming things. There's just so much information out there and any given human can never hold more than the most minuscule chunk of all of it in his mind, so they'll be able to connect lots of dots that we're missing simply because of our limited carrying capacity, but I still don't think they'll ever be able to create fundamentally new dots.

In other words:

- Solving extremely complex mathematical problems requiring extensive knowledge across multiple esoteric and complex domains? Yip.

- Creating math starting from a framework where math doesn't exist in any way, shape, or fashion? Nope.

Ironically, the more complex the cross-domain problems are, the more effective LLMs will seem to be, because you limit the number of humans who have any chance of internalizing everything across both domains, whereas for an LLM there's no such issue. This will create a perception of super intelligence, which will probably where any danger from LLMs would emerge. Doing things like using a token prediction algorithm to make war or other such strategic decisions, because of the misguided belief that it's not only intelligent but super intelligent. It's basically cargo cult logic.

jibal • today at 5:46 AM

> not too long ago

Like today. I asked Gemini for the longest state names with an even number of letters and it gave me North Carolina and South Carolina. When I complained that they are odd, it gave me North Dakota and South Dakota, which are both odd and not the longest. When I noted that, it went back to the Carolinas. Finally it appeared to switch to a different model that actually did counting and found Pennsylvania and West Virginia.

➕ show 1 reply
queenkjuul • today at 4:30 AM

They were sued over it in Germany and lost, which makes it all the more surprising they keep it up everywhere else tbh

➕ show 1 reply
mitxela • yesterday at 11:56 PM

LLMs are fundamentally predicting the next word to make coherent text. If you've ever played with a Markov chain text generator you've done this with a fairly dumb predictor that maintains coherence over a very short distance. Deep transformer neutral networks can do it with a much longer coherence distance but they are fundamentally performing the same operation. After "Question: Did the team make the playoffs? Answer:" a reasonable completion is "yes, the team made the playoffs". An early demonstration of GPT-2 was a fake news article about scientists discovering unicorns in Antarctica - the model doesn't "know" whether or not unicorns exist in Antarctica, but it's able to complete "Breaking news! Scientists have discovered a colony of English-speaking unicorns in Antarctica." by adding "The unicorns have a developed society with running water and electricity." because that's a sensible next sentence. (I didn't look up the actual text it wrote)

➕ show 3 replies
ssl-3 • yesterday at 11:13 PM

It says this this at the bottom of every one of the dumb responses that Google's trash-tier bot puts above the (deliberately awful, these days) search results:

AI can make mistakes, so double-check responses

➕ show 1 reply
fangspire • today at 4:48 AM

[dead]