> out of 100 random questions I might think to ask, it’s likely to say something wrong or stupid a handful of times at least.
What are some examples?
There's a benchmark for this and a lot of models get negative scores because they're so unreliable: https://artificialanalysis.ai/evaluations/omniscience
There's a benchmark for this and a lot of models get negative scores because they're so unreliable: https://artificialanalysis.ai/evaluations/omniscience