There's a benchmark for this and a lot of models get negative scores because they're so unreliable: https://artificialanalysis.ai/evaluations/omniscience
I wanted some examples they actually experienced. Because I use these things daily and haven’t seen a hallucination in a long long time.
I wanted some examples they actually experienced. Because I use these things daily and haven’t seen a hallucination in a long long time.