logoalt Hacker News

gottorf • last Wednesday at 8:44 PM • 6 replies • view on HN

My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.


Replies

WarmWash • yesterday at 12:12 AM

The achilles heel of 3.8 flash is it's january 2025 knowledge cutoff date. Yes, almost 2 years ago.

I'm assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.

➕ show 2 replies
Gareth321 • yesterday at 8:29 AM

I agree. It's much worse than the cheap Chinese models. They appear to heavily bias parametric knowledge and discourage tool use. That's fine for things like "how do I perform CPR?" but worse than useless for any kind of research. [There is one benchmark showing far lower rates of hallucination, so let's see how accurate this is.](https://www.reddit.com/r/singularity/comments/1wuj72j/gemini...)

MILP • last Wednesday at 9:17 PM

I'm also not using it for coding but I've found Flash 3.8 to generate much better HTML output than Sonnet or Opus.

➕ show 1 reply
staticman2 • last Wednesday at 9:04 PM

The web version of Gemini is awful at search but I don't think that's the models fault.

esafak • yesterday at 5:32 AM

I would expect a flash model, with its reduced size, to suffer on tail tasks. That is the trade-off you make.

mattjoyce • last Wednesday at 9:22 PM

Hallucination seems a very dated term.

➕ show 3 replies