I'll give you a recent example from my usage. Pi harness with extension for learning Chinese. When using it to feed drill questions to me and rate answers it would sometimes get lost in the sauce and start generating user aka me answer and then rate it and comment it. It's trivially wrong to the point that if a person would do that, they would be considered for some serious psych issues.
And it gets even better since when called out it wouldn't just take my word for it but only acknowledged the issue after parsing the log with clearly delineated user and model output.
So yeah while impressive things are able to be done, the current models are also dumb AF and an idiot savant is a pretty good label for them.
This is often (but not always, as it is often run with positive temperature which deliberately shifts the generation path) because of disconnected context and the issues around context compaction. Long context has always been critical to important work, but it remains a substantial challenge due to computational bottlenecks. It isn't a fundamental issue with the architecture, moreso the tricks to make them cheaper to use. In the long run, taking various concessions to make them cheaper to run is likely to be the limiting factor that prevents this "runaway explosion" that many people are worried about. That doesn't mean that the model design is inherently limited, like a lot of people seem to assume based on current failings.
Besides, have you spoken to someone lately that everyone would calls a genius? They can say the most braindead stuff sometimes. I wouldn't use worst-case performance as an indication of general capacity.