Much of what OAI and Anthropic are doing with LLMs has obvious failure modes.
The most obvious failure mode for their hacking evals was an improperly configured, tested and monitored sandbox.
Similarly, the very first question after an impressively correct result from any ML tool, LLM or not, is to see if the answer was already in the training data.
These companies don't even handle the blatantly obvious failure modes that do not kill people.
[flagged]