My experience with Gemini 3.8 Flash has been awful; it gives me the most hallucinations out of the major models. I'm not using it for coding, but general research on different topics.
I agree. It's much worse than the cheap Chinese models. They appear to heavily bias parametric knowledge and discourage tool use. That's fine for things like "how do I perform CPR?" but worse than useless for any kind of research. [There is one benchmark showing far lower rates of hallucination, so let's see how accurate this is.](https://www.reddit.com/r/singularity/comments/1wuj72j/gemini...)
I'm also not using it for coding but I've found Flash 3.8 to generate much better HTML output than Sonnet or Opus.
The web version of Gemini is awful at search but I don't think that's the models fault.
I would expect a flash model, with its reduced size, to suffer on tail tasks. That is the trade-off you make.
The achilles heel of 3.8 flash is it's january 2025 knowledge cutoff date. Yes, almost 2 years ago.
I'm assuming that Argon has at least a June 2026 date, but man, the 3 series models were a mess with newer information.